| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-08-21 | |||
| 14:47:56 | mriedem | maybe? | |
| 14:48:09 | mriedem | maybe neutron f's up and screws up the active source host port binding? | |
| 14:48:11 | sean-k-mooney | ill apply the patch in anycase and see what happens | |
| 14:50:20 | openstack | Launchpad bug 1788149 in OpenStack Compute (nova) "nova list too slow" [Undecided,Incomplete] | |
| 14:50:20 | s10 | mriedem: bug https://bugs.launchpad.net/nova/+bug/1788149 could be eventlet related and maybe caused by nova-neutron connection on every nova show/nova list (https://bugs.launchpad.net/nova/+bug/1567655) | |
| 14:50:22 | openstack | Launchpad bug 1567655 in OpenStack Compute (nova) "500 error when trying to list instances and neutron-server is down" [Medium,Confirmed] | |
| 14:51:00 | mriedem | s10: hmm yeah good point re eventlet | |
| 14:51:13 | mriedem | also if they are running nova-api via wsgi in pike we aren't monkey patching eventlet | |
| 14:52:09 | openstack | Launchpad bug 1567655 in OpenStack Compute (nova) "500 error when trying to list instances and neutron-server is down" [Medium,Confirmed] | |
| 14:52:09 | mriedem | and yeah https://bugs.launchpad.net/nova/+bug/1567655 came up again with my product team overlords last week | |
| 14:52:15 | mriedem | regarding perf/scaling issues with nova api | |
| 14:52:51 | mriedem | tl;dr we should cache the port security group information in instance.info_cache like we do for other port information | |
| 14:53:41 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge config drive extension response into server controller https://review.openstack.org/584223 | |
| 14:53:42 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge extended server attributes extension response https://review.openstack.org/584590 | |
| 14:53:43 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge server usage extension response into server view builder https://review.openstack.org/585262 | |
| 14:53:43 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge keypair extension response into server view builder https://review.openstack.org/584748 | |
| 14:53:44 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge security groups extension response into server view builder https://review.openstack.org/585475 | |
| 14:53:45 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge extended_status extension response into server view builder https://review.openstack.org/592092 | |
| 14:56:29 | openstackgerrit | Dan Smith proposed openstack/nova master: WIP: Record cell success/failure/timeout in CrossCellLister https://review.openstack.org/594265 | |
| 14:56:33 | dansmith | tssurya: ^ | |
| 14:56:51 | dansmith | tssurya: what if we do that, and then in get_instance_objects_sorted() (or above) we just get a handle on InstanceLister itself | |
| 14:57:14 | dansmith | tssurya: then we can construct the missing instances from the failed cells separately from the intricate multi-cell-listing logic? | |
| 14:57:48 | tssurya | dansmith: sounds good, guess its better to move it out of the generator | |
| 14:58:19 | sean-k-mooney | mriedem: this is really annoying. im not seeing the issue going from ovs-dpdk to kernel ovs. the migration failes but the binding are fine. i was hitting this going form ovs to ovs-dpdk when i reported the bug so ill test that next. | |
| 14:58:23 | dansmith | tssurya: yeah I actually would like it higher than get_instance_objects_sorted(), but definitely want it out of the low-level logic if possible | |
| 15:01:13 | tssurya | dansmith: when you say higher than get_instance_objects_sorted(), then you mean we return the list of non-responsive cells instead from ^^ patch and constrcut it in compute/api ? | |
| 15:01:33 | tssurya | construct* | |
| 15:01:43 | dansmith | tssurya: yes, I'd like that better personally, just to keep "api logic" closer to the api | |
| 15:02:10 | dansmith | tssurya: maybe just return a tuple from get_instance_objects_sorted() indicating (failed_cell_uuids, instances_i_actually_got) | |
| 15:03:38 | tssurya | dansmith: okay looks doable, I will try it out | |
| 15:04:00 | dansmith | cool | |
| 15:04:03 | tssurya | thanks for the review btw, was getting kind of lost in the details | |
| 15:04:11 | tssurya | :) | |
| 15:04:20 | dansmith | tssurya: no problem, this is complicated and I've been neglecting it too long | |
| 15:04:33 | tssurya | hehe :D | |
| 15:04:49 | openstackgerrit | Jose Castro Leon proposed openstack/nova master: Add extend in-use volumes support for RBD https://review.openstack.org/594273 | |
| 15:05:18 | jroll | mriedem: yeah, death to the baremetal filters | |
| 15:05:40 | mriedem | i left comments on the bug, | |
| 15:05:56 | mriedem | was mostly looking for input on how people have done vm/bm in a single compute endpoint, i assume host aggregates | |
| 15:06:07 | mriedem | but i've heard there are also quota issues when doing it that way | |
| 15:15:01 | sean-k-mooney | mriedem: you can use capablityes in the flavor to avoid the need for host aggregates | |
| 15:16:05 | sean-k-mooney | not sure how many people go that route vs AZs or host aggregates | |
| 15:17:21 | sean-k-mooney | atully with more recent releases you can just use resouce classes + dedicated baremetal or vm flavor and let placement handel it | |
| 15:18:01 | mriedem | that's why the baremetal filter options were deprecated in pike and removed in rocky | |
| 15:18:06 | mriedem | which is why i marked the bug as won't fix | |
| 15:20:25 | dansmith | mriedem: do you know if Kevin_Zheng and yikun_ are still working on tests? because I think if they have no ERROR instances, I could hack up a generator they could test to compare apples to apples on whether or not that object list loop could go faster | |
| 15:20:53 | jrock_cfdg | hello - i'm trying to add a serial device with specific paramaters to an instance at creation time (source mode=connect host=0.0.0.0 port=4555) ; I think i've narrowed it down to these 3 scripts (/usr/lib/python-2.7/site-packages/nova/virt/libvirt/{config,driver,guest}.py - which is the correct place to make this change? And has anyone here done anything like this and maybe have some examples? | |
| 15:20:54 | mriedem | idk | |
| 15:20:56 | dansmith | meaning still have their profiling setup accessible or whatevef | |
| 15:20:59 | dansmith | okay | |
| 15:21:02 | mriedem | i'm sure it's still setup | |
| 15:21:17 | Kevin_Zheng | we can still test | |
| 15:21:19 | mriedem | it's just a bash script on a devstack deploy on a baremetal host | |
| 15:21:25 | dansmith | Kevin_Zheng: ohai | |
| 15:21:30 | mriedem | the lurker | |
| 15:21:54 | mriedem | oh right, monday, tuesday thursday are work late days for kevin and yikun | |
| 15:22:00 | dansmith | ah | |
| 15:22:11 | dansmith | Kevin_Zheng: I assume all your test instances are ACTIVE or something right? | |
| 15:22:22 | Kevin_Zheng | Yes | |
| 15:22:26 | Kevin_Zheng | All active | |
| 15:22:51 | dansmith | Kevin_Zheng: so right here, we iterate all the instances: https://github.com/openstack/nova/blob/master/nova/compute/instance_list.py#L124-L126 | |
| 15:23:27 | dansmith | Kevin_Zheng: and so I'm wondering if removing that would also help your perf a bit.. the problem is we have to handle faults, which are handled in that list method right now | |
| 15:23:56 | mriedem | maciejjozefczyk: if you're around https://review.openstack.org/#/c/591607/ | |
| 15:23:57 | dansmith | Kevin_Zheng: but with the batching, we *might* be better off doing that in the batches instead of at the top to reduce latency | |
| 15:24:08 | mriedem | maciejjozefczyk: our public cloud ops team reported the same issue | |
| 15:24:12 | Kevin_Zheng | So instead all instances, we do what? | |
| 15:24:55 | dansmith | Kevin_Zheng: well, we'd do it in the batch handler, so we fill faults on ~100 instances at a time in "parallel" instead of on 1000 instances serially | |
| 15:25:12 | mriedem | efried: i guess we can land this now huh https://review.openstack.org/#/c/520024/ | |
| 15:25:16 | Kevin_Zheng | Oh OK | |
| 15:25:47 | dansmith | Kevin_Zheng: sounds like if I come up with a test patch you could run it again and compare to without the patch just to see if it helps or hurts? | |
| 15:25:59 | Kevin_Zheng | Guess I have to generate some error instance then | |
| 15:26:14 | cdent | yay! on 520024 | |
| 15:26:15 | Kevin_Zheng | Yeah we can do it | |
| 15:26:28 | mriedem | you insert them right into the cell db right? | |
| 15:26:35 | Kevin_Zheng | Yes | |
| 15:26:41 | dansmith | Kevin_Zheng: well, the first test would still be all active, just to measure what the perf impact of unrolling that loop is | |
| 15:27:05 | dansmith | Kevin_Zheng: then we'd test a patch with some error instances to see if we lose all of that with the fault handling, or only a fraction of the gain we made | |
| 15:27:50 | Kevin_Zheng | Ok | |
| 15:30:20 | dansmith | Kevin_Zheng: I'll try cooking something up and will add you to the review | |
| 15:30:57 | Kevin_Zheng | Cool, I will go to bed and check in the morning | |
| 15:31:01 | dansmith | thanks | |
| 15:33:01 | sean-k-mooney | mriedem: so regarding the live migration bug. the source node is activating the binding on the dest host binding after the migration aborts and this is also racing with the deltion of the binding on the dest host ... | |
| 15:33:12 | efried | mriedem: Yes, on 024, thanks. | |
| 15:37:58 | mriedem | sean-k-mooney: so we're hitting post-copy and then aborting? | |
| 15:38:19 | mriedem | there are only 2 places that live migration activates the dest host port binding: | |
| 15:38:27 | mriedem | 1. post-copy event callback from libvirt | |
| 15:38:38 | mriedem | 2. _post_live_migration after the hypervisor said the live migration was successful | |
| 15:42:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Explicitly fail if trying to attach SR-IOV port https://review.openstack.org/591898 | |
| 15:42:26 | sean-k-mooney | mriedem: http://paste.openstack.org/show/728534/ | |
| 15:42:27 | dansmith | ugh, the expectation that we return an instancelist from get_all makes this harder than I thought | |
| 15:42:53 | sean-k-mooney | i think we are geting an updat form neutron and that is trigering the activate. perhaps hitting the _pos_live_migration code | |
| 15:45:05 | sean-k-mooney | mriedem: lines 43-50 are teh ones im suspicous of | |
| 15:45:43 | mriedem | a neutron event wouldn't make us activate a port | |
| 15:45:47 | mriedem | just refresh the info cache | |
| 15:46:17 | mriedem | Aug 21 16:19:17 devstack2 nova-compute[25894]: WARNING nova.compute.manager [None req-594840ec-7af2-47d2-929b-cef9dda07bb8 service nova] [instance: fead1ca6-beab-4c47-a73e-a3ab7f7c4de2] Received unexpected event network-vif-unplugged-ef02ea3f-9a11-4519-bcd3-2bfca97edf26 for instance with vm_state active and task_state migrating. | |
| 15:46:20 | mriedem | means we ignore it | |
| 15:47:23 | sean-k-mooney | hum ok well on line 50 we activate the port binding for devstack5 which was the destination node. but the migration has already aborted | |
| 15:47:58 | mriedem | Aug 21 16:19:17 devstack2 nova-compute[25894]: DEBUG nova.network.neutronv2.api [None req-c8b07cbc-52f7-4d20-aacc-f3036ad90c8d None None] Activated binding for port ef02ea3f-9a11-4519-bcd3-2bfca97edf26 and host devstack5. {{(pid=25894) activate_port_binding /opt/stack/nova/nova/network/neutronv2/api.py:1352}} | |
| 15:47:58 | mriedem | hmm | |
| 15:48:00 | mriedem | indeed | |