| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-06-05 | |||
| 11:15:53 | pvc | hi anyone | |
| 11:15:54 | pvc | can help me | |
| 11:17:44 | gibi | pvc: "but there is a time that my hypervisor is going down" this is not specific enough for me to be able to help troubleshooting it | |
| 11:18:24 | pvc | gibi i have an instance residing on that hypervisor | |
| 11:18:33 | pvc | with gpu on it | |
| 11:18:49 | pvc | then my hypervisor always restarting i dont know why | |
| 11:19:46 | pvc | compute host that have a PCI passthrough | |
| 11:26:42 | openstackgerrit | Chen Hanxiao proposed openstack/nova master: sync_guest_time: use the proper errno https://review.openstack.org/572346 | |
| 11:41:15 | gibi | pvc: have you tried to dig into your hypervisor logs to see why it restarted? | |
| 12:50:21 | ttsiouts | gibi: Thanks for the comments on the PENDING state spec! | |
| 12:50:53 | ttsiouts | gibi: Whenever you have some time I wanted to ask you something before I update it.. | |
| 12:51:36 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Transform instance.exists notification https://review.openstack.org/403660 | |
| 12:51:54 | gibi | ttsiouts: I have now 10 minutes before I have to join a call | |
| 12:52:30 | ttsiouts | I'll try to make it quick | |
| 12:52:35 | ttsiouts | :) | |
| 12:52:49 | ttsiouts | I wanted to ask about the note on scheduler hints.. | |
| 12:53:14 | ttsiouts | The intention was to not include them in the payload | |
| 12:53:35 | ttsiouts | but it would be really helpful to have the forced_nodes/hosts | |
| 12:54:10 | gibi | ttsiouts: I think it is a reasonable need for your Reaper to know those | |
| 12:54:34 | ttsiouts | gibi: awesome! | |
| 12:54:43 | gibi | ttsiouts: if other cores are OK to include the hints I will not block this on the hints | |
| 12:55:08 | ttsiouts | gibi: thank you very much!! I will update the spec | |
| 12:58:01 | gibi | ttsiouts: cool. thanks for pushing that spec forward | |
| 13:00:39 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081 | |
| 13:00:42 | bhagyashri_s | efried: Hi, | |
| 13:01:12 | gibi | ttsiouts: you can ignore my rebuild in cell0 question in the PENDING spec as I now found your other spec about rebuild in cell0 :) | |
| 13:01:38 | jaypipes | gibi: crap... just saw your comments on the cpu-resources spec. :) | |
| 13:01:48 | jaypipes | gibi: will answer on the last revision. | |
| 13:05:38 | gibi | jaypipes: no problem | |
| 13:07:19 | efried | bhagyashri_s: Hello | |
| 13:07:25 | bhagyashri_s | efried, jaypipes, bauzas: Addressed review comments on https://review.openstack.org/#/c/560459/ could you please help to review Thank you :) | |
| 13:07:33 | efried | ack | |
| 13:07:45 | bauzas | ack | |
| 13:09:18 | bhagyashri_s | efried, jaypipes, bauzas: thank you :) | |
| 13:17:11 | jangutter | Good $timezone! In the spirit of spec review day, is anyone interested in reviewing a spec for a $network $thingy? As a pleasant side effect, hopefully one more legacy VIF can be moved to os-vif. https://review.openstack.org/#/c/567148/ | |
| 13:43:08 | mriedem | johnthetubaguy: tssurya: dansmith: we might need to setup a call at some point for the 'handling a down' cell spec https://review.openstack.org/#/c/557369/ | |
| 13:43:14 | mriedem | lots of questions for me on that right now | |
| 13:43:31 | dansmith | okay | |
| 13:45:03 | mriedem | dansmith: you should run through the comments and discussion in there and then see what you think about a call | |
| 13:46:28 | dansmith | mriedem: okay, before I do, do you know anything about the functional.libvirt.test_pci_sriov_servers tests? | |
| 13:46:46 | dansmith | I seem to have destabilized them as they pass in isolation but not on top of my patch when run with everything else | |
| 13:47:00 | dansmith | looks like vladik wrote them initially | |
| 13:47:06 | dansmith | maybe stephenfin knows about them? | |
| 13:47:41 | stephenfin | dansmith: It's been a while since I touched them | |
| 13:47:50 | dansmith | stephenfin: http://logs.openstack.org/95/572195/1/check/nova-tox-functional/cbf4f55/testr_results.html.gz | |
| 13:48:08 | dansmith | get that most of the time when running in parallel, but they pass when I run them by themselves | |
| 13:48:11 | dansmith | which smells like a bad mock | |
| 13:48:37 | stephenfin | Yeah, looks like there's some global state getting trashed | |
| 13:49:55 | naichuans | efried: Hi, Eric, do you have any suggestion about vgpu rp delete? I posted a new patch | |
| 13:53:15 | stephenfin | dansmith: Nothing jumps out at me :/ | |
| 13:53:52 | dansmith | nothing about that patch should be affecting global state to make that a legit failure, so I dunno | |
| 13:54:05 | dansmith | if I run just that one module in parallel it fails, | |
| 13:54:10 | dansmith | so it must be something internal to that | |
| 13:54:14 | dansmith | and not from another test | |
| 13:54:18 | stephenfin | It's always _that_ test that fails, right? | |
| 13:54:34 | stephenfin | Or a variety of the tests from that module? | |
| 13:54:34 | dansmith | yes, but only in parallel with the other tests in that file | |
| 13:54:38 | efried | naichuans: not having looked at your update yet, the way I figured you would do it is: | |
| 13:54:38 | efried | - Construct a set of all providers in the tree whose names start with your prefix (GPUG_ or VGPU_ or whatever it was). | |
| 13:54:38 | efried | - Construct a set of all providers you *expect* there to be - i.e. the ones you've discovered from the host_data. | |
| 13:54:38 | efried | - Subtract the second set from the first. This is the set of providers you need to delete. | |
| 13:54:38 | efried | - For each provider that needs to be deleted, check for allocations. If you don't find any, you can delete the provider. If you do find allocations... I think you just have to set reserved=total and leave the provider there. | |
| 13:54:38 | dansmith | if I run just that test, it passes | |
| 13:54:45 | stephenfin | gotcha | |
| 13:55:08 | dansmith | stephenfin: there's another one in another file under libvirt/ there that also fails occasionally, but they seem to be unrelated | |
| 13:56:50 | dansmith | stephenfin: is that message indicating that some mock for the system's pinnable CPUs is not set and so it's empty? or something? | |
| 13:57:20 | stephenfin | That message indicates you're trying to pin an instance but the CPUs aren't available | |
| 13:57:29 | stephenfin | Generally because they're pinned to something else | |
| 13:57:31 | naichuans | efried: OK, looks a acceptable choice, thanks. | |
| 13:57:48 | dansmith | stephenfin: that list will be empty like that if the cpus are pinned elsewhere? | |
| 13:57:57 | stephenfin | Yup | |
| 13:58:40 | dansmith | so like another test not having deleted a server or something? | |
| 13:58:43 | mriedem | stephenfin: if you want, i can address my nits in https://review.openstack.org/#/c/541290/ | |
| 13:58:45 | dansmith | although they should all be getting an empty db | |
| 13:58:50 | stephenfin | You're saying I want to take CPUs Y from the total host set of X | |
| 13:59:00 | stephenfin | mriedem: Go for it | |
| 13:59:05 | stephenfin | mriedem: and thanks | |
| 14:00:05 | stephenfin | dansmith: It could also be that we're booting two instances consecutively from different tests | |
| 14:00:17 | dansmith | but they should have different databases and not see each other | |
| 14:00:19 | stephenfin | except, yeah, each test gets its own sqlite DB | |
| 14:00:20 | stephenfin | Hmm | |
| 14:00:52 | mriedem | do the fake hosts/nodes in the tests have the same name? | |
| 14:01:02 | mriedem | because the nova.tests.unit.virt.fake set_nodes or whatever is global | |
| 14:01:15 | mriedem | one of the tests might not be resetting the fake node on cleanup? | |
| 14:01:19 | dansmith | yeah | |
| 14:01:49 | dansmith | I don't see them using fake set_nodes | |
| 14:01:54 | dansmith | unless it's somewhere else | |
| 14:02:00 | mriedem | parent class? | |
| 14:02:42 | dansmith | I don't think so.. they | |
| 14:02:49 | dansmith | are starting their own compute service for some reason | |
| 14:02:58 | dansmith | maybe so they can control the topology | |
| 14:03:07 | stephenfin | dansmith: Yeah, that | |
| 14:04:12 | stephenfin | That's stored in nova.objects.compute_node.ComputeNode.numa_topology, which I think is created on startup for nova-compute | |
| 14:04:51 | dansmith | and they don't do fake node set in the parent class(es) | |
| 14:06:33 | tssurya | mriedem: yea that spec is going crazy | |
| 14:06:49 | dansmith | stephenfin: hmm, just got two failures from that module in one run, so something is definitely wonky there | |
| 14:07:03 | stephenfin | dansmith: Yeah, I'm taking a look at it now | |
| 14:07:15 | dansmith | I really can't think of why my patch would be affecting this kind of thing | |
| 14:09:41 | openstackgerrit | Matt Riedemann proposed openstack/nova-specs master: Add 'numa-aware-vswitches' spec https://review.openstack.org/541290 | |
| 14:09:57 | mriedem | stephenfin: fyi https://review.openstack.org/#/c/541290/17..18/specs/rocky/approved/numa-aware-vswitches.rst | |