| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-05-17 | |||
| 16:42:51 | edleafe | dansmith: ok | |
| 16:45:57 | belmorei_ | dansmith: me and tssurya were debugging the scheduler time. It's related with https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L772 | |
| 16:46:20 | belmorei_ | creating a dict with all instances info per host | |
| 16:48:01 | dansmith | per host we're considering each time we run through yeah? | |
| 16:48:14 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: Localdisk https://review.openstack.org/549300 | |
| 16:48:20 | belmorei_ | yeah | |
| 16:49:21 | dansmith | belmorei_: what is the slow part exactly? it looks like it's hitting a cell db once per host instead of grabbing a bunch at a time maybe? | |
| 16:49:46 | melwitt | I've actually seen that complaint before and I don't understand how compute isn't sending updates to the scheduler. or does that not happen anymore because of placement? | |
| 16:49:56 | mriedem | it does | |
| 16:49:58 | dansmith | melwitt: that would be an upcall | |
| 16:49:59 | mriedem | if configured to do so | |
| 16:50:04 | dansmith | right | |
| 16:50:05 | melwitt | I see | |
| 16:50:09 | mriedem | it's an rpc cast from all computes to the scheduler topic | |
| 16:50:14 | mriedem | it's only used for the affinity filters | |
| 16:50:29 | mriedem | track_instance_changes or something like that | |
| 16:50:44 | dansmith | yeah, and what I'm not getting is why this would be different for them in the new arrangement,m | |
| 16:50:47 | mriedem | https://docs.openstack.org/nova/latest/configuration/config.html#filter_scheduler.track_instance_changes | |
| 16:50:55 | dansmith | because they're only considering hosts from one cell at this point anyway | |
| 16:50:56 | mriedem | dansmith: i'm wondering the same | |
| 16:51:01 | melwitt | yeah. lemme see if I can find the bug I'm thinking of | |
| 16:51:51 | bauzas | is someone able to tell me how to simply just associate an instance and a floating IP, without going thru the list of ports ? | |
| 16:52:07 | bauzas | we had that in the past, but we removed that here | |
| 16:52:12 | melwitt | this one https://bugs.launchpad.net/nova/+bug/1737465 | |
| 16:52:13 | openstack | Launchpad bug 1737465 in OpenStack Compute (nova) "[cellv2] the performance issue of cellv2 when creating 500 instances concurrently" [Undecided,Confirmed] - Assigned to Jiang (jiangpf) | |
| 16:52:20 | belmorei_ | dansmith not following... hosts from one cell? | |
| 16:52:28 | mriedem | bauzas: I CAN | |
| 16:52:39 | mriedem | bauzas: https://docs.openstack.org/python-openstackclient/pike/cli/command-objects/floating-ip.html#floating-ip-set | |
| 16:52:45 | mriedem | get the port id from the instance | |
| 16:52:51 | mriedem | then associate the port to hte floating ip | |
| 16:52:55 | mriedem | or create the floating ip with the port | |
| 16:52:55 | bauzas | mriedem: port is required | |
| 16:52:57 | dansmith | belmorei_: that code should be building a dict of host info, but only from the list of uuids that we got back from placement, which should be only one cell, and only the subset of those that are candidates | |
| 16:53:05 | mriedem | bauzas: yes the port is what links the instance to the floating ip | |
| 16:53:07 | bauzas | mriedem: so I need to openstack port list before | |
| 16:53:23 | bauzas | I was thinking we could somehow magically avoid one roundtrip | |
| 16:53:33 | bauzas | with openstack CLI doing that in secret for me | |
| 16:53:38 | mriedem | you can do it on a single command line if you want... | |
| 16:53:46 | bauzas | awk my love | |
| 16:53:50 | mriedem | no | |
| 16:53:51 | mriedem | NO | |
| 16:54:01 | mriedem | PORT1=$(openstack port list --device-id $SERVER1 -c ID -f value) | |
| 16:54:07 | mriedem | FLOATINGIP1=$(openstack floating ip create --port $PORT1 public -c floating_ip_address -f value) | |
| 16:54:13 | mriedem | do that in a single line | |
| 16:54:22 | bauzas | meh | |
| 16:54:23 | belmorei_ | dansmith: yes, but takes a lot of time even for 200 nodes | |
| 16:54:24 | mriedem | then $$$ me | |
| 16:54:34 | mriedem | bauzas: meh? | |
| 16:54:40 | dansmith | belmorei_: but it should have been doing that before in the same way in the cell scheduler I mean | |
| 16:54:41 | bauzas | mriedem: yeah, meh. | |
| 16:54:46 | mriedem | bauzas: see channel topic | |
| 16:54:50 | bauzas | mriedem: hahah | |
| 16:54:53 | mriedem | love my idea or get out | |
| 16:54:54 | dansmith | mriedem: nice | |
| 16:54:56 | bauzas | anyway, you're right | |
| 16:55:11 | bauzas | that's not a dev call :p | |
| 16:55:31 | bauzas | mriedem: FWIW, you were confused with the status of libvirt vGPUs | |
| 16:55:45 | bauzas | mriedem: the branch is up for reviews, and I just updated the spec | |
| 16:55:54 | dansmith | belmorei_: maybe it was taking just as long before in the cell, but it wasn't as obvious because it was spread out? | |
| 16:55:59 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM snapshot cleanup https://review.openstack.org/568988 | |
| 16:56:01 | melwitt | dansmith: what do you mean by "cell scheduler"? | |
| 16:56:05 | dansmith | belmorei_: or did you have track_instance_changes enabled? I wouldn't think affinity would work with cellsv1 so I would have assumed no | |
| 16:56:11 | dansmith | melwitt: in cellsv1 there is a scheduler in the cell | |
| 16:56:11 | bauzas | mriedem: jay had some concerns with traits, but I think we should do that in a separate spec | |
| 16:56:20 | belmorei_ | dansmith: maybe | |
| 16:56:26 | belmorei_ | dansmith: track_instance_changes is enabled | |
| 16:56:34 | melwitt | okay. well, with cells v1, wouldn't compute be reporting updates to the scheduler and then it wouldn't build that dict? | |
| 16:56:42 | mriedem | belmorei_: was it disabled when you were doing cells v1? | |
| 16:56:50 | dansmith | melwitt: that's what I just said above | |
| 16:56:59 | bauzas | melwitt: dansmith: catching up the convo, what's the problem with computes updating the scheduler ? | |
| 16:57:05 | dansmith | but I thought we still had to hit the db once per host even with that enbabled | |
| 16:57:09 | bauzas | there is a fanout upcall | |
| 16:57:10 | melwitt | ? so doesn't that make it obvious why it was fast with cells v1 and slow with cells v2? | |
| 16:57:31 | melwitt | because with cells v2 it's building the dict every time in the absence of updates? I'm confused | |
| 16:57:32 | dansmith | melwitt: I had asked him if they had it enabled | |
| 16:57:46 | dansmith | melwitt: affinity doesn't work with cellsv1 so I would have expected they didn't have that enabled | |
| 16:57:50 | dansmith | (afaik) | |
| 16:58:03 | melwitt | okay, I see | |
| 16:58:09 | dansmith | and, I thought we still hit the db once per host even with it enabled for the compute info, but not for the instance info part | |
| 16:58:15 | dansmith | so I wouldn't expect a lage difference | |
| 16:58:16 | melwitt | it probably works for them by way of patches, maybe | |
| 16:58:17 | dansmith | *large | |
| 16:58:23 | dansmith | well, then they're cheating :) | |
| 16:59:05 | melwitt | what's the compute info part? | |
| 17:00:01 | melwitt | the normal host states stuff? | |
| 17:00:13 | dansmith | yeah | |
| 17:01:07 | belmorei_ | ok, we will continue to dig on this | |
| 17:01:13 | belmorei_ | thanks | |
| 17:01:29 | belmorei_ | I'm leaving now | |
| 17:01:59 | mriedem | random thinking out loud, i think the instance info list only uses the uuid from the instances for the affinity filters, | |
| 17:02:07 | melwitt | belmorei_: seeya next week | |
| 17:02:15 | mriedem | we could probably use placement to just pull uuids of consumers that have allocations on a given node and replace that | |
| 17:02:29 | mriedem | would be a lot fast | |
| 17:02:31 | mriedem | *faster | |
| 17:02:39 | dansmith | well, it depends on which piece of his linked function is taking a while | |
| 17:02:46 | dansmith | we do hit the db for the hosts every time, | |
| 17:02:55 | dansmith | but the instance list is the thing we would skip with track_changes | |
| 17:02:59 | dansmith | so yeah, maybe it's just that | |
| 17:03:09 | mriedem | https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L753 could also avoid some default joins... | |
| 17:03:16 | dansmith | still once call to placement per instance would be not great | |