| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-12-19 | |||
| 13:47:57 | mriedem | given the "inventory" in nova is the intersection of what's in the config and what's on the host | |
| 13:50:13 | mgariepy | my use case is more like: 1 vm / host with all the ressources, (pci passthrough ram and cpu), resize doesn't really work in that case. | |
| 13:51:32 | mgariepy | i noticed that the pci devices are re-created in the db, how is the link made to the computes ? | |
| 13:51:55 | mgariepy | pci passthrough is fun. it gave me quite a few issues lately. | |
| 13:52:22 | mriedem | so you can't cold migrate the vm because ther are no other available hosts with the same pci device? | |
| 13:52:26 | mriedem | *there | |
| 13:52:35 | mgariepy | no there isn't | |
| 13:54:16 | mriedem | hmm, i wonder if you could trick the resize to same host though by creating a private duplicate flavor with some bogus extra spec like foo=bar | |
| 13:54:34 | mgariepy | in nova.pci_devices why isn't it link to the compute node id ? | |
| 13:54:40 | mriedem | you'd have to of course enable this option https://docs.openstack.org/nova/latest/configuration/config.html#DEFAULT.allow_resize_to_same_host | |
| 13:54:52 | mgariepy | when recreating the device | |
| 13:55:26 | mriedem | i'm really not the person that can answer low level questions about how the pci device code works within nova, | |
| 13:55:36 | mriedem | hopefully sean-k-mooney and/or stephenfin could help there | |
| 13:56:01 | mriedem | iow i have to look all of the code up every time i need to investigate it | |
| 13:56:28 | mriedem | the PciDevice object does have a compute_node_id field | |
| 13:56:30 | sean-k-mooney | mriedem: sorry was reading bug. what is the context | |
| 13:56:53 | mgariepy | as far as i'm concern, if i have a compute node, and the devices is the same in the same pci address, it should be ""undeleted"" instead of created again. | |
| 13:57:31 | mgariepy | sean-k-mooney, https://bugs.launchpad.net/nova/+bug/1809040 | |
| 13:57:31 | openstack | Launchpad bug 1809040 in OpenStack Compute (nova) "pci device lost when error in the configuration file " [Undecided,New] | |
| 13:58:14 | mgariepy | sean-k-mooney, in nova.pci_devices why isn't it link to the compute node id ? | |
| 13:58:45 | mgariepy | https://paste.ubuntu.com/p/Pn76QVmwqr/ | |
| 13:59:19 | mriedem | they are linked to compute node 23 there | |
| 13:59:34 | sean-k-mooney | mgariepy: they are | |
| 14:00:09 | mgariepy | yep, but the issue is that if i remove the passthrough config from nova.conf, it get the deleted_at but if I re-add it it create new one. | |
| 14:00:17 | mgariepy | in the same host, same address. etc.. | |
| 14:00:36 | sean-k-mooney | yes | |
| 14:00:41 | sean-k-mooney | that is expected | |
| 14:01:29 | sean-k-mooney | the id filed is an auto incremting filed and the uuid is randomly generated if the device id not found in the database | |
| 14:01:34 | mgariepy | wouldn't be better to re-use the old entry if all the info matches? | |
| 14:02:06 | sean-k-mooney | mgariepy: if we have already deleted it no | |
| 14:02:06 | stephenfin | mgariepy: For what it's worth, that also confused me but it is expected | |
| 14:02:23 | mgariepy | the 2 first are the ""original"" one, then the 2 other are the new one created. | |
| 14:02:31 | mriedem | it seems that things break down on reboot here https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L5350 | |
| 14:02:42 | mriedem | pci_manager.get_instance_pci_devs(instance) must be returning [] | |
| 14:02:45 | sean-k-mooney | mgariepy: we do it this way because you could have pull the card and install a different one in the same slot | |
| 14:02:59 | mgariepy | yes. | |
| 14:03:24 | mgariepy | then the product id would have changed | |
| 14:03:26 | sean-k-mooney | that could break the guest if we jsut blindly reused it | |
| 14:03:29 | mriedem | and i think that's probably [] because i think instance.pci_devices is set during a resource claim, which doesn't happen on reboot | |
| 14:03:42 | sean-k-mooney | no the product ids could be the same but if they were nics | |
| 14:03:52 | sean-k-mooney | that were usidn for pf passthouhg the mac woudl have changed | |
| 14:03:58 | openstackgerrit | Merged openstack/nova master: Fix a broken-link in nova doc https://review.openstack.org/626110 | |
| 14:04:10 | mgariepy | yes for a nic, that's true. | |
| 14:04:25 | mgariepy | that's my testbed , the other system i have uses graphic card :D | |
| 14:04:26 | mgariepy | haha | |
| 14:04:41 | sean-k-mooney | event for gpus we only the vendor id and produc id are recored | |
| 14:04:47 | sean-k-mooney | not the subvend id | |
| 14:05:32 | sean-k-mooney | so all GTX 1080s have the same vendor id and product id but an EVGA or asus one has different subvendor ids | |
| 14:05:48 | sean-k-mooney | not they should be identical but no all such product are. | |
| 14:06:09 | sean-k-mooney | well actully the clock speed/ram could chagne | |
| 14:06:28 | mgariepy | yeah but the drive would manage that for you. | |
| 14:07:32 | sean-k-mooney | it may but the point is if we deteact the device was removed we cannot trust that it si the same and cant reuse the entry | |
| 14:07:53 | sean-k-mooney | we do not recerate tehm on every reboot | |
| 14:08:18 | sean-k-mooney | we only do it if the agent does a pci scan and did not find them | |
| 14:08:49 | mgariepy | yeah, i never had issue with reboot before :D haha | |
| 14:09:22 | mgariepy | anyway at least now i know, and i'll be more careful next time. | |
| 14:10:30 | sean-k-mooney | so out of interst the device was allocated when it was removed | |
| 14:10:40 | sean-k-mooney | what was teh state of the vm | |
| 14:10:44 | stephenfin | mriedem: Yeah, I'm not actually sure how else we could resolve that besides a cold migrate/rebuild. It's a mismatch between two sources of truth: what the libvirt driver is finding on the host (based on the whitelist) and what the instance is saying it's using | |
| 14:11:25 | mgariepy | as long as you don't restart it's ok | |
| 14:11:29 | sean-k-mooney | mgariepy: i would have assumed you would have migrated the vm off the host before upgrade but if you didnt have you tried doing a openstack server reboot --hard to try and fix the issue | |
| 14:11:59 | mgariepy | when you restart, the libvirt config generated doesn't have the pci devices but boots anyway | |
| 14:12:27 | sean-k-mooney | right ok then the only way to fix that is likely to shelve and unshelve | |
| 14:13:00 | sean-k-mooney | e.g. to free up the node and its pci device and then recreate the vm with its data on the same node | |
| 14:13:05 | mgariepy | the reboot --hard doesn't work, since i guess the data is pulled from the DB and my ""new"" device is not allocated | |
| 14:13:09 | mgariepy | yeah. | |
| 14:13:35 | mgariepy | i'll have my client to rebuild his cluster. | |
| 14:14:13 | mgariepy | sean-k-mooney, i do inplace upgrades, it's not a big deal ;P | |
| 14:14:27 | sean-k-mooney | mgariepy: well you could also manually fix this in the db without too much hasel but i guess | |
| 14:14:53 | sean-k-mooney | on a larger cluster it may be more complicated however | |
| 14:15:10 | sean-k-mooney | mgariepy: i have done that but usally i create a test env first. | |
| 14:15:11 | mgariepy | i don't like updating the db. sometimes it comes back to bite me .. | |
| 14:15:58 | sean-k-mooney | mgariepy: yes it can it can be less painful then redeploying the cluster however | |
| 14:16:12 | sean-k-mooney | unless you jsut ment the applicateion in the cluster. | |
| 14:16:13 | mgariepy | it's a "special" contributed cluster part of another one | |
| 14:16:45 | sean-k-mooney | deleting the vm and recreating it would have the same effect but i assume this happend on all nodes so you would have to delete all the vms with passthough and recreate them | |
| 14:16:46 | mgariepy | he runs some kind of hpc cluster on kubernetes on the vms | |
| 14:17:56 | sean-k-mooney | so to be clear. you use openstack to spwan 1 vm per host that uses all the host resouces then they use kubernetes to run a distibuted hpc application on the vms | |
| 14:18:10 | mgariepy | anyway, not a big deal, i'll be more careful next time. i just messed up the nova.conf pci config on the upgrade | |
| 14:18:12 | mgariepy | yep. | |
| 14:18:15 | mgariepy | haha :D | |
| 14:18:32 | sean-k-mooney | that seam overly complicated but if ti work it works i guess :) | |
| 14:18:59 | mgariepy | it's shared, and this way the client doesn't really have to deal with the hardware... | |
| 14:19:17 | mgariepy | and have some other benefit like access to some storage and so on. | |
| 14:19:57 | sean-k-mooney | yep i totally get why you would do it its just you have at least 3 layers fo orchestration there | |
| 14:20:18 | mgariepy | yep, not all the same person do all 3. | |
| 14:20:33 | sean-k-mooney | openstack orchestrting the vms, kubernets orchestatin the hpc cluster and spark or whatever orchestatign the hpc jobs on the cluster | |
| 14:21:06 | mgariepy | probably slurm, but i'm not 100% sure. | |
| 14:21:33 | mgariepy | anyway, thanks a lot for your time and help. | |
| 14:21:42 | sean-k-mooney | if you deploy your openstack with kubernets you can make it nice an inceptione | |
| 14:21:55 | sean-k-mooney | no worres are you ok with me closing https://bugs.launchpad.net/nova/+bug/1809040 | |
| 14:21:56 | openstack | Launchpad bug 1809040 in OpenStack Compute (nova) "pci device lost when error in the configuration file " [Undecided,New] | |
| 14:22:18 | mgariepy | the question is ,will I be able to remove the physical server at some point. | |
| 14:22:20 | mgariepy | yep | |
| 14:25:07 | stephenfin | sean-k-mooney: Perhaps we could add a nova-compute start up check to see if there are unrecognized PCI devices attached to running instances and fail to start if so? | |
| 14:25:24 | stephenfin | sean-k-mooney: Thought I guess by then the old PCI devs in the manager would have been marked deleted and new ones created | |
| 14:25:48 | stephenfin | Unless we did it realllly early, but that would involve duplicating a lot of logic | |
| 14:26:01 | stephenfin | really early = before the PCI manager stuff kicks off | |
| 14:26:17 | sean-k-mooney | stephenfin: well i was going to suggest in the but if some wanted to retarget the bug to allow a hard reboot to fix it then it would be fine | |
| 14:26:26 | jaypipes | jackding: will try my best | |