| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-11-26 | |||
| 14:29:34 | alex_xu | mriedem: do you know resize works with numa topology change in the new flavor? I can't found we parse the numa topo from new flavor. Try to figure out it maybe a bug or something we don't support well yet. | |
| 14:31:31 | mriedem | alex_xu: i think it's supported...we'll do the move_claim during resize which gets the new_flavor off the migration record in the resource tracker | |
| 14:31:53 | mriedem | and the request spec uses the new flavor (i think) when it runs through the scheduler (numa topo filter) during scheduling for the resize | |
| 14:32:08 | alex_xu | mriedem: yea...I saw the move claim code also. | |
| 14:32:40 | alex_xu | mriedem: but the request spec wont extract new numa topo from the new flavor. I guess we missed something | |
| 14:34:28 | mriedem | cfriesen might know off the top of his head faster than me | |
| 14:34:38 | mriedem | but i thought cold migration worked | |
| 14:35:40 | mriedem | alex_xu: you're saying we don't update https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L55 for the new flavor right? | |
| 14:36:24 | alex_xu | mriedem: yes | |
| 14:36:53 | alex_xu | mriedem: we generate new numa_topology in move claim. but we still use old numa_topology in the scheduler | |
| 14:38:40 | mriedem | now i'm having a hard time finding where we set the new_flavor on the request spec before calling the scheduler | |
| 14:39:11 | alex_xu | hah, probably because we don't have that :) | |
| 14:40:25 | mriedem | heh, no i think that happens b/c i just wrote a functional regression test that relies on it | |
| 14:40:37 | mriedem | https://review.openstack.org/#/c/619123/ | |
| 14:41:58 | alex_xu | oh, yea, I miss read that. but still not parse the numa stuff from flavor | |
| 14:43:47 | alex_xu | mriedem: it's late for me, I will dig into more tomorrow. thanks for the info, good to know that isn't something we don't support, then it is probably a bug... | |
| 14:43:50 | mriedem | alex_xu: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L316 is where we set the new flavor on reqspec | |
| 14:44:01 | mriedem | alex_xu: sure, i'll ping you if i figure something out :) | |
| 14:44:07 | mriedem | good night | |
| 14:48:21 | mriedem | stephenfin: are you aware of cold migration/resize support for numa topology changes ^ ? | |
| 14:48:27 | mriedem | i thought that was all baked in long ago | |
| 14:49:28 | kashyap | mriedem: IIRC, he's out for a few more days. | |
| 14:49:46 | mriedem | ok i'll wait for cfriesen then | |
| 14:54:53 | mriedem | dansmith: do you think https://review.openstack.org/#/c/607735/ is worth sending to rocky? | |
| 14:55:02 | mriedem | it's a pretty latent issue, so not really sure it's worth it | |
| 14:55:41 | dansmith | mriedem: it is, but it's a pretty trivial thing to push back and we know it's a problem for people | |
| 14:55:48 | dansmith | we also know we can't push it back any farther, | |
| 14:55:58 | dansmith | so it's not like it can go back to kilo or anything | |
| 14:56:02 | mriedem | sure | |
| 14:56:03 | mriedem | ok | |
| 15:04:41 | cfriesen | mriedem: alex_xu: pretty sure cold migration *was* working before all the placement stuff, but I seem to remember seeing at least one bug saying it's currently broken. | |
| 15:05:55 | cfriesen | we're on pike at the moment and it seems to be working, but we do have a few patches in that area for other features. | |
| 15:06:07 | sean-k-mooney | cold migration in what context? | |
| 15:06:41 | sean-k-mooney | mriedem: stephenfin is on PTO until next week | |
| 15:06:54 | sean-k-mooney | mriedem: he left his znc bouncer running | |
| 15:07:34 | mriedem | cfriesen: as alex pointed out, | |
| 15:07:45 | mriedem | i don't see the request spec numa topologies field updated before it goes through the scheduler | |
| 15:08:01 | sean-k-mooney | mriedem: oh alex_xu was asking about numa toplogy changes for resize/cold migration it "should" work upstream | |
| 15:08:29 | mriedem | https://github.com/openstack/nova/blob/master/nova/scheduler/filters/numa_topology_filter.py#L74 | |
| 15:08:45 | mriedem | i don't see RequestSpec.numa_topology updated from the new flavor *before* we hit the numa filter | |
| 15:09:04 | mriedem | so i don't see how the scheduling is working | |
| 15:09:31 | sean-k-mooney | hum so you think its using the old flavor perhaps | |
| 15:09:47 | sean-k-mooney | i can try testing this in a hour or so | |
| 15:10:03 | jmlowe | Has anybody had trouble with the placement client in nova compute hanging? | |
| 15:10:04 | sean-k-mooney | i jsut need to get my dev enviornemt running after the weekend | |
| 15:10:25 | mriedem | jmlowe: never heard of that | |
| 15:10:45 | jmlowe | Specifically nova.compute.resource_tracker._update_available_resource was holding a lock for about 2 min | |
| 15:11:05 | jmlowe | Sometimes it would do it and sometimes not | |
| 15:11:33 | jmlowe | no errors anywhere | |
| 15:12:21 | cdent | jmlowe: how many instances on that compute node? | |
| 15:12:42 | jmlowe | It did wreak havoc with instance launches | |
| 15:12:43 | cdent | are you sure it was talking to placement where things were stuck, or just somewhere in _udpate*? | |
| 15:12:58 | cdent | alos | |
| 15:13:05 | cdent | ^walso | |
| 15:13:07 | cfriesen | mriedem: sorry, had to answer a call from my boss. let me take a quick look | |
| 15:14:22 | jmlowe | I got it down to 2 min to run _update_inventory in nova/scheduler/client/report.py | |
| 15:15:07 | mriedem | jmlowe: how many instances on that compute? | |
| 15:15:10 | jmlowe | it seems to have made the http put call correctly, but then just waits for tcp timeout | |
| 15:15:13 | mriedem | libvirt or vcenter or ironic? | |
| 15:15:36 | jmlowe | it's all of my 280 computes, so 0 - 24 | |
| 15:15:42 | jmlowe | libvirt | |
| 15:16:04 | mriedem | hmm, well there is a lock in the resource tracker when that is called, which will make things in nova-compute slow to a crawl, | |
| 15:16:11 | mriedem | but why the placement response would be so slow idk, | |
| 15:16:21 | mriedem | have you traced the request via request ID in the placement api logs? | |
| 15:16:27 | mriedem | also, which release? | |
| 15:16:44 | jmlowe | yes, nothing in placement takes more than a second or so | |
| 15:16:57 | jmlowe | queens | |
| 15:18:20 | jmlowe | seems like the there's a bug in some underlying client library that is occasionally not returning from a http call without a timeout | |
| 15:18:38 | mriedem | hmm, nova is using keystoneauth1 to send the requests to placement | |
| 15:18:57 | mriedem | you might be able to enable some debug logging there | |
| 15:19:45 | cdent | jmlowe: I think you should make jeremy find and fix this | |
| 15:20:28 | cdent | not you fungi | |
| 15:20:48 | jmlowe | I did, then went on a 6 week road trip | |
| 15:20:55 | fungi | yay! for once something's not my fault | |
| 15:20:56 | jmlowe | then he went | |
| 15:21:08 | cdent | The jeremy of which I speak is an old friend (on the order of 30 years) | |
| 15:21:13 | cdent | Typical of him. | |
| 15:21:17 | cfriesen | mriedem: alex_xu: I think it's this code that updates the flavor on a resize: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L304-L316 | |
| 15:22:27 | mriedem | cfriesen: but that doesn't update the RequestSpec.numa_topology field | |
| 15:22:44 | mriedem | the else block | |
| 15:22:57 | mriedem | or RequestSpec.pci_requests for that matter | |
| 15:23:44 | mriedem | so if resize with a new numa topology works today, it's getting lucky b/c the scheduler picks a host that fits the original topology, and then the RT.move_claim on the compute fits the new requested numa topology | |
| 15:23:53 | mriedem | as far as i can tell anyway | |
| 15:24:42 | sean-k-mooney | mriedem: well we trow away all the host pinning info and recalulate it on the compute node anyway so it is likely working becaue it gets lucky or we retry | |
| 15:24:54 | cdent | jmlowe: do you have load balancers or proxies between n-cpu placement? | |
| 15:25:03 | jmlowe | I do | |
| 15:25:10 | cdent | connections perhaps not closing? | |
| 15:25:13 | mriedem | cdent: fwiw, that ENABLED_PYTHON3_PACKAGES variable in devstack looks like it should be killed now | |
| 15:25:26 | mriedem | "# Special case some services that have experimental | |
| 15:25:27 | mriedem | # support for python3 in progress, but don't claim support | |
| 15:25:27 | mriedem | # in their classifier" | |
| 15:25:28 | cdent | mriedem: it does come into play | |
| 15:25:33 | mriedem | yeah i see where it's used | |
| 15:25:39 | mriedem | but i'm just thinking it shouldn't exist anymore | |
| 15:25:44 | jmlowe | probably, I'd expect the http client to close the connection once it received a 200 responxe | |
| 15:25:49 | mriedem | given the py3 first goal | |
| 15:26:02 | mriedem | e.g. neutron isn't in that list | |
| 15:26:04 | mriedem | nor keystone | |
| 15:26:17 | sean-k-mooney | mriedem: well it either shouldn not exists or default to ture right | |
| 15:26:21 | cdent | mriedem: yes, I agree that it appears to be an anachronism that shouldn't be used at all, but that it sometimes does get used | |