| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-03-31 | |||
| 19:52:34 | zigo | It's not as if it was so verbose ... :/ | |
| 19:57:19 | zigo | Oh ... | |
| 19:57:19 | zigo | <class 'nova.exception.GlanceConnectionFailed'> (HTTP 500) (Request-ID: req-df3f3de6-e2ba-41a7-b717-67cee688d832) | |
| 19:57:54 | zigo | I know what's going on then. :) | |
| 20:31:52 | sean-k-mooney | :) | |
| 20:32:11 | sean-k-mooney | ya if glance is unhappy then nova will be too | |
| 21:04:23 | openstackgerrit | Rodrigo Barbieri proposed openstack/nova master: [WIP] Error anti-affinity violation on live migrations https://review.opendev.org/c/openstack/nova/+/784166 | |
| 21:30:37 | openstackgerrit | Artom Lifshitz proposed openstack/nova master: WIP: Update SRIOV port pci_slot when unshelving https://review.opendev.org/c/openstack/nova/+/784168 | |
| 21:31:04 | artom | sean-k-mooney, ^^ \o/ | |
| 22:09:29 | zigo | sean-k-mooney: I belive I'm *very* close from booting my first VM with Wallaby. :) | |
| 23:10:47 | sean-k-mooney | artom: i need to digest this and think about it but i _think_ that works. i wpoudl prefer if you were not doing [0] but i get it | |
| 23:11:07 | artom | sean-k-mooney, yeah, there's some questionable code in there, I agree | |
| 23:11:19 | artom | I did mark it as WIP this time ;) | |
| 23:11:47 | artom | I want CI to do a full run to prove out the idea, then we can refine | |
| 23:11:49 | sean-k-mooney | you figured out how to get the requester_id | |
| 23:12:04 | sean-k-mooney | i looked at this a few times and missed that | |
| 23:12:10 | artom | Yeah, by setting it first :P | |
| 23:12:32 | artom | Which makes me think - VMs that got booted before that patch is applied will still have the original bug | |
| 23:12:46 | artom | Because they won't have requester_id on their PCIRequests | |
| 23:13:11 | sean-k-mooney | ture but you coudl do an online migratieon with a nova manage commnd | |
| 23:13:26 | sean-k-mooney | to correalte the exsitng allocation using the pci_slot | |
| 23:13:46 | sean-k-mooney | for currently shelved vms it wont help | |
| 23:13:55 | sean-k-mooney | but for any exisitng vm it would | |
| 23:13:56 | artom | It won't indeed | |
| 23:15:02 | sean-k-mooney | i wonder if gibi had any other logic reated to it being set | |
| 23:15:26 | sean-k-mooney | it was only set for port with resouce request before | |
| 23:15:40 | sean-k-mooney | i dont think it shoudl break anything but its one of teh things we shoudl check | |
| 23:16:56 | sean-k-mooney | anyway ya ill take a look at that again tomrrow | |
| 23:17:13 | artom | Yeah, no rush | |
| 23:17:22 | artom | I was just happy I managed to hax something together | |
| 23:18:00 | sean-k-mooney | yes its similar to what i tred to do before but i was trying to do it without moving settign the requester id | |
| 23:18:22 | sean-k-mooney | but i mean that works | |
| 23:19:05 | sean-k-mooney | i was basically workign the otherway and trying to put the request_id in the vif object in the network infor cache | |
| 23:19:10 | sean-k-mooney | but this is better | |
| 23:19:35 | sean-k-mooney | since it avoid the object change | |
| 23:21:43 | artom | Yeah, this should be backportable | |
| 23:26:54 | artom | And thanks for the requester_id pointer, btw :) | |
| 23:27:04 | artom | If it wasn't for that, I'd still be digging | |
| #openstack-nova - 2021-04-01 | |||
| 00:08:57 | ikla | can I use ceph for my instance and local disk for ephemeral storage? | |
| 00:09:10 | ikla | appears to be either or | |
| 00:09:24 | ikla | rbd or local | |
| 00:10:30 | ikla | or can I pass a local partition in? | |
| 06:58:07 | gibi_away | sean-k-mooney, artom: about requester_id in InstancePciRequest, this is the past info I have about deciding not to put the flavorid into the requester_id https://review.opendev.org/c/openstack/nova/+/625310/25/nova/pci/request.py#220 | |
| 07:27:24 | gibi | bauzas: hi! I've propose and RC2 patch and but a -W on it until next week https://review.opendev.org/c/openstack/releases/+/784201 | |
| 08:28:24 | gibi | lyarwood: I put a heavy -1 to https://review.opendev.org/c/openstack/nova/+/784130 | |
| 08:28:45 | sean-k-mooney | gibi: we could encode if its the flavorid/port id in the string | |
| 08:28:57 | sean-k-mooney | i do not belive its defined as a uuid | |
| 08:29:18 | sean-k-mooney | so we could do flavor/<uuid> and port/<uuid> if we needed too | |
| 08:29:45 | sean-k-mooney | normally i would say just add a requester_type | |
| 08:29:53 | sean-k-mooney | but that would be a db change | |
| 08:29:54 | gibi | sean-k-mooney: it wasn't filled not becasue we don't know what to put there, we did not put flavor there as nobody uses that information | |
| 08:30:19 | sean-k-mooney | ah yes | |
| 08:30:27 | sean-k-mooney | that makes sense i guess | |
| 08:30:32 | gibi | sean-k-mooney: and there later some source information is also added to the ovo in I292a0e2d840bbf657ba6d0932f9a3decbcb2778f | |
| 08:30:37 | gibi | https://review.opendev.org/c/openstack/nova/+/659101/4/nova/objects/instance_pci_requests.py | |
| 08:31:04 | sean-k-mooney | yes i remember that | |
| 08:31:19 | gibi | I have no issue if somebody starts adding flavorid tehre | |
| 08:31:20 | gibi | there | |
| 08:31:41 | sean-k-mooney | ya i think we can use store the uuid directly | |
| 08:31:49 | gibi | yepp | |
| 08:31:54 | sean-k-mooney | and use that source property to figure it out if we need too | |
| 08:32:09 | gibi | and we can add a fallback to the ovo that if the requester_id is None then it looks up the flavor stored in the instance | |
| 08:33:28 | sean-k-mooney | have you seen artoms change by the way https://review.opendev.org/c/openstack/nova/+/784168 | |
| 08:33:54 | sean-k-mooney | i still need to properly review it but at first glance this seams like a workable solution going forward | |
| 08:34:20 | sean-k-mooney | im still debating if we want to add somethign to populate the requeter id for exsiting recoreds | |
| 08:34:57 | sean-k-mooney | possibly as a nova manage command or something on init_host as a seperate patch | |
| 08:36:35 | sean-k-mooney | what i was thinkng was loop over all the claimed devices without requester_id set, then lookup the port info form the network info cache and try and use the pci slot to corralte the pci devices to the ports | |
| 08:39:16 | gibi | sean-k-mooney: haven't looked that review yet | |
| 08:42:35 | sean-k-mooney | its still WIP so not quite ready for review yet but the core chagne is there and the previous patch add a functional repoducer which passes after the change | |
| 08:45:20 | lyarwood | gibi: replied, not sure I agree tbh, seems like a trivial requirement to codify to avoid operators borking things | |
| 08:45:43 | lyarwood | gibi: we had an example downstream of evacuations from 2019 being cleaned up recently in 2021 that spawned this btw | |
| 08:46:01 | lyarwood | gibi: somehow in the iterim the instance(s) had been migrated back to the original host | |
| 08:46:19 | lyarwood | gibi: I assume with manual operator intervention to cleanup instance dirs etc | |
| 08:46:36 | lyarwood | gibi: thankfully no data was lost but I'm just trying to avoid a future situation where we don't get as lucky | |
| 08:47:02 | lyarwood | gibi: and again it's specifically for the force down, evacuate and force up usecase. Outside of that I agree that a periodic or something would be useful | |
| 08:47:11 | gibi | lyarwood: this is a slippery slope as then when the next bug cames that after force_down + evac the VM volume became corrupted due to two VMs writing it then somebody can say that we should not evacuate a VM if we cannot pervent VM duplication | |
| 08:47:58 | gibi | as soon as the admin used force_down we should stop caring | |
| 08:48:04 | lyarwood | gibi: this isn't trying to resolve any issues during the initial force down | |
| 08:48:13 | lyarwood | gibi: *only* when the operator forces the service back up | |
| 08:48:28 | lyarwood | gibi: and we are given the chance to check the state of the migrations | |
| 08:48:44 | sean-k-mooney | so the alternitive here would be to make the action for what happens when we discover vms in this state configureable | |
| 08:48:48 | lyarwood | gibi: anything that happens before the force up is out of scope for this | |
| 08:48:55 | sean-k-mooney | like the running_delete_instance_action | |
| 08:49:20 | sean-k-mooney | our current behavior is hardcoded reap the instance basically | |
| 08:49:22 | lyarwood | yeah as I said yesterday sean-k-mooney you would force up the compute service first before moving them | |
| 08:49:33 | lyarwood | moving them back* | |
| 08:49:45 | lyarwood | so checking during force up makes more sense tbh | |
| 08:50:19 | gibi | lyarwood: so the goal is to check if the system is in consistent state before we allow forecing the compute up | |
| 08:50:39 | sean-k-mooney | ya im not really against that check but i do kind fo agree with gibi that once they do force down then they are assuming resposiblity to some degree | |
| 08:50:47 | sean-k-mooney | gibi: yes | |
| 08:51:00 | lyarwood | gibi: correct, evacuation migration records in done with a source host pointing at the localhost meaning a service restart is required | |
| 08:51:24 | lyarwood | sean-k-mooney: yup agree but docs have been proven time and time again not to be enough to avoid PEBKAC | |
| 08:51:56 | sean-k-mooney | lyarwood: would an alternitive be to do an rpc on force up? | |
| 08:52:12 | sean-k-mooney | basically to trigger what woudl be done on service restart | |
| 08:52:32 | lyarwood | sean-k-mooney: yeah we could do that moving forward and backport this check? | |
| 08:52:47 | sean-k-mooney | well i dont know if that was a good idea or not | |
| 08:52:48 | lyarwood | sean-k-mooney: and/or have a periodic that triggers this | |
| 08:53:04 | sean-k-mooney | just tinking things through | |
| 08:53:06 | gibi | sean-k-mooney: while you could trigger the evac cleanup during force up we could think through what else needs to be done. Like clean up local deleted instances. | |
| 08:53:25 | lyarwood | we have a periodic for that right | |