| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-08-21 | |||
| 14:35:07 | dansmith | we're still iterating all of the instances from a cell before returning them in order to do the fault stuff | |
| 14:35:29 | mriedem | dansmith: i think when you were adding instance lister, | |
| 14:35:40 | mriedem | you pre-joined faults in the db api and it didn't seem to make a difference in perf | |
| 14:35:48 | mriedem | and it might have caused some other issue | |
| 14:36:05 | dansmith | so I'm surprised we gained as much as we did by my batching, and so I wonder if we push the faults into the batches if it would help | |
| 14:36:10 | dansmith | mriedem: yeah, I remember that now | |
| 14:36:13 | mriedem | we also only show fault if the vm state is ERROR or DELETED | |
| 14:36:16 | dansmith | right | |
| 14:37:12 | mriedem | well this bug says nova list is too slow https://bugs.launchpad.net/nova/+bug/1788149 | |
| 14:37:12 | openstack | Launchpad bug 1788149 in OpenStack Compute (nova) "nova list too slow" [Undecided,New] | |
| 14:37:15 | mriedem | so there is that | |
| 14:37:57 | dansmith | heh | |
| 14:38:41 | tssurya | nice | |
| 14:41:28 | mriedem | alex_xu: did you say gmann was on vacation? https://review.openstack.org/#/c/584223/ | |
| 14:41:45 | mriedem | (8:09:57 AM) alex_xu: gmann: enjoy your vacation! | |
| 14:41:50 | tssurya | mriedem: yes untill 31st | |
| 14:42:17 | tssurya | until* | |
| 14:45:21 | sean-k-mooney | mriedem: regarding the migration issue i dont see "Binding ports to destination host" in either the source or dest compute logs | |
| 14:46:08 | sean-k-mooney | the dest does have "Plugging VIFs using destination host port bindings before live migration." and "Deleted binding for port 3218fd70-ea82-4ee1-9a5b-2d3c9d8b9fa0 and host devstack2." | |
| 14:46:47 | mriedem | the former is when we do pre_live_migration on the dest host, | |
| 14:46:54 | mriedem | at that point port bindings are still active for the source host | |
| 14:47:03 | mriedem | the latter is when we're rolling back after the failed migration | |
| 14:47:21 | mriedem | so i'm not sure that my patch would fix your issue if we never deactivated the source host bindings | |
| 14:47:42 | mriedem | that's why i was saying i'd be surprised if it fixed it b/c it would mean the failure happened after post-copy | |
| 14:47:47 | sean-k-mooney | ya. i expected to see boot. could deleteing the port bindings be the root of the issue | |
| 14:47:56 | mriedem | maybe? | |
| 14:48:09 | mriedem | maybe neutron f's up and screws up the active source host port binding? | |
| 14:48:11 | sean-k-mooney | ill apply the patch in anycase and see what happens | |
| 14:50:20 | s10 | mriedem: bug https://bugs.launchpad.net/nova/+bug/1788149 could be eventlet related and maybe caused by nova-neutron connection on every nova show/nova list (https://bugs.launchpad.net/nova/+bug/1567655) | |
| 14:50:20 | openstack | Launchpad bug 1788149 in OpenStack Compute (nova) "nova list too slow" [Undecided,Incomplete] | |
| 14:50:22 | openstack | Launchpad bug 1567655 in OpenStack Compute (nova) "500 error when trying to list instances and neutron-server is down" [Medium,Confirmed] | |
| 14:51:00 | mriedem | s10: hmm yeah good point re eventlet | |
| 14:51:13 | mriedem | also if they are running nova-api via wsgi in pike we aren't monkey patching eventlet | |
| 14:52:09 | mriedem | and yeah https://bugs.launchpad.net/nova/+bug/1567655 came up again with my product team overlords last week | |
| 14:52:09 | openstack | Launchpad bug 1567655 in OpenStack Compute (nova) "500 error when trying to list instances and neutron-server is down" [Medium,Confirmed] | |
| 14:52:15 | mriedem | regarding perf/scaling issues with nova api | |
| 14:52:51 | mriedem | tl;dr we should cache the port security group information in instance.info_cache like we do for other port information | |
| 14:53:41 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge config drive extension response into server controller https://review.openstack.org/584223 | |
| 14:53:42 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge extended server attributes extension response https://review.openstack.org/584590 | |
| 14:53:43 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge keypair extension response into server view builder https://review.openstack.org/584748 | |
| 14:53:43 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge server usage extension response into server view builder https://review.openstack.org/585262 | |
| 14:53:44 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge security groups extension response into server view builder https://review.openstack.org/585475 | |
| 14:53:45 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Merge extended_status extension response into server view builder https://review.openstack.org/592092 | |
| 14:56:29 | openstackgerrit | Dan Smith proposed openstack/nova master: WIP: Record cell success/failure/timeout in CrossCellLister https://review.openstack.org/594265 | |
| 14:56:33 | dansmith | tssurya: ^ | |
| 14:56:51 | dansmith | tssurya: what if we do that, and then in get_instance_objects_sorted() (or above) we just get a handle on InstanceLister itself | |
| 14:57:14 | dansmith | tssurya: then we can construct the missing instances from the failed cells separately from the intricate multi-cell-listing logic? | |
| 14:57:48 | tssurya | dansmith: sounds good, guess its better to move it out of the generator | |
| 14:58:19 | sean-k-mooney | mriedem: this is really annoying. im not seeing the issue going from ovs-dpdk to kernel ovs. the migration failes but the binding are fine. i was hitting this going form ovs to ovs-dpdk when i reported the bug so ill test that next. | |
| 14:58:23 | dansmith | tssurya: yeah I actually would like it higher than get_instance_objects_sorted(), but definitely want it out of the low-level logic if possible | |
| 15:01:13 | tssurya | dansmith: when you say higher than get_instance_objects_sorted(), then you mean we return the list of non-responsive cells instead from ^^ patch and constrcut it in compute/api ? | |
| 15:01:33 | tssurya | construct* | |
| 15:01:43 | dansmith | tssurya: yes, I'd like that better personally, just to keep "api logic" closer to the api | |
| 15:02:10 | dansmith | tssurya: maybe just return a tuple from get_instance_objects_sorted() indicating (failed_cell_uuids, instances_i_actually_got) | |
| 15:03:38 | tssurya | dansmith: okay looks doable, I will try it out | |
| 15:04:00 | dansmith | cool | |
| 15:04:03 | tssurya | thanks for the review btw, was getting kind of lost in the details | |
| 15:04:11 | tssurya | :) | |
| 15:04:20 | dansmith | tssurya: no problem, this is complicated and I've been neglecting it too long | |
| 15:04:33 | tssurya | hehe :D | |
| 15:04:49 | openstackgerrit | Jose Castro Leon proposed openstack/nova master: Add extend in-use volumes support for RBD https://review.openstack.org/594273 | |
| 15:05:18 | jroll | mriedem: yeah, death to the baremetal filters | |
| 15:05:40 | mriedem | i left comments on the bug, | |
| 15:05:56 | mriedem | was mostly looking for input on how people have done vm/bm in a single compute endpoint, i assume host aggregates | |
| 15:06:07 | mriedem | but i've heard there are also quota issues when doing it that way | |
| 15:15:01 | sean-k-mooney | mriedem: you can use capablityes in the flavor to avoid the need for host aggregates | |
| 15:16:05 | sean-k-mooney | not sure how many people go that route vs AZs or host aggregates | |
| 15:17:21 | sean-k-mooney | atully with more recent releases you can just use resouce classes + dedicated baremetal or vm flavor and let placement handel it | |
| 15:18:01 | mriedem | that's why the baremetal filter options were deprecated in pike and removed in rocky | |
| 15:18:06 | mriedem | which is why i marked the bug as won't fix | |
| 15:20:25 | dansmith | mriedem: do you know if Kevin_Zheng and yikun_ are still working on tests? because I think if they have no ERROR instances, I could hack up a generator they could test to compare apples to apples on whether or not that object list loop could go faster | |
| 15:20:53 | jrock_cfdg | hello - i'm trying to add a serial device with specific paramaters to an instance at creation time (source mode=connect host=0.0.0.0 port=4555) ; I think i've narrowed it down to these 3 scripts (/usr/lib/python-2.7/site-packages/nova/virt/libvirt/{config,driver,guest}.py - which is the correct place to make this change? And has anyone here done anything like this and maybe have some examples? | |
| 15:20:54 | mriedem | idk | |
| 15:20:56 | dansmith | meaning still have their profiling setup accessible or whatevef | |
| 15:20:59 | dansmith | okay | |
| 15:21:02 | mriedem | i'm sure it's still setup | |
| 15:21:17 | Kevin_Zheng | we can still test | |
| 15:21:19 | mriedem | it's just a bash script on a devstack deploy on a baremetal host | |
| 15:21:25 | dansmith | Kevin_Zheng: ohai | |
| 15:21:30 | mriedem | the lurker | |
| 15:21:54 | mriedem | oh right, monday, tuesday thursday are work late days for kevin and yikun | |
| 15:22:00 | dansmith | ah | |
| 15:22:11 | dansmith | Kevin_Zheng: I assume all your test instances are ACTIVE or something right? | |
| 15:22:22 | Kevin_Zheng | Yes | |
| 15:22:26 | Kevin_Zheng | All active | |
| 15:22:51 | dansmith | Kevin_Zheng: so right here, we iterate all the instances: https://github.com/openstack/nova/blob/master/nova/compute/instance_list.py#L124-L126 | |
| 15:23:27 | dansmith | Kevin_Zheng: and so I'm wondering if removing that would also help your perf a bit.. the problem is we have to handle faults, which are handled in that list method right now | |
| 15:23:56 | mriedem | maciejjozefczyk: if you're around https://review.openstack.org/#/c/591607/ | |
| 15:23:57 | dansmith | Kevin_Zheng: but with the batching, we *might* be better off doing that in the batches instead of at the top to reduce latency | |
| 15:24:08 | mriedem | maciejjozefczyk: our public cloud ops team reported the same issue | |
| 15:24:12 | Kevin_Zheng | So instead all instances, we do what? | |
| 15:24:55 | dansmith | Kevin_Zheng: well, we'd do it in the batch handler, so we fill faults on ~100 instances at a time in "parallel" instead of on 1000 instances serially | |
| 15:25:12 | mriedem | efried: i guess we can land this now huh https://review.openstack.org/#/c/520024/ | |
| 15:25:16 | Kevin_Zheng | Oh OK | |
| 15:25:47 | dansmith | Kevin_Zheng: sounds like if I come up with a test patch you could run it again and compare to without the patch just to see if it helps or hurts? | |
| 15:25:59 | Kevin_Zheng | Guess I have to generate some error instance then | |
| 15:26:14 | cdent | yay! on 520024 | |
| 15:26:15 | Kevin_Zheng | Yeah we can do it | |
| 15:26:28 | mriedem | you insert them right into the cell db right? | |
| 15:26:35 | Kevin_Zheng | Yes | |