Earlier  
Posted Nick Remark
#openstack-nova - 2018-11-26
14:31:53 mriedem and the request spec uses the new flavor (i think) when it runs through the scheduler (numa topo filter) during scheduling for the resize
14:32:08 alex_xu mriedem: yea...I saw the move claim code also.
14:32:40 alex_xu mriedem: but the request spec wont extract new numa topo from the new flavor. I guess we missed something
14:34:28 mriedem cfriesen might know off the top of his head faster than me
14:34:38 mriedem but i thought cold migration worked
14:35:40 mriedem alex_xu: you're saying we don't update https://github.com/openstack/nova/blob/master/nova/objects/request_spec.py#L55 for the new flavor right?
14:36:24 alex_xu mriedem: yes
14:36:53 alex_xu mriedem: we generate new numa_topology in move claim. but we still use old numa_topology in the scheduler
14:38:40 mriedem now i'm having a hard time finding where we set the new_flavor on the request spec before calling the scheduler
14:39:11 alex_xu hah, probably because we don't have that :)
14:40:25 mriedem heh, no i think that happens b/c i just wrote a functional regression test that relies on it
14:40:37 mriedem https://review.openstack.org/#/c/619123/
14:41:58 alex_xu oh, yea, I miss read that. but still not parse the numa stuff from flavor
14:43:47 alex_xu mriedem: it's late for me, I will dig into more tomorrow. thanks for the info, good to know that isn't something we don't support, then it is probably a bug...
14:43:50 mriedem alex_xu: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L316 is where we set the new flavor on reqspec
14:44:01 mriedem alex_xu: sure, i'll ping you if i figure something out :)
14:44:07 mriedem good night
14:48:21 mriedem stephenfin: are you aware of cold migration/resize support for numa topology changes ^ ?
14:48:27 mriedem i thought that was all baked in long ago
14:49:28 kashyap mriedem: IIRC, he's out for a few more days.
14:49:46 mriedem ok i'll wait for cfriesen then
14:54:53 mriedem dansmith: do you think https://review.openstack.org/#/c/607735/ is worth sending to rocky?
14:55:02 mriedem it's a pretty latent issue, so not really sure it's worth it
14:55:41 dansmith mriedem: it is, but it's a pretty trivial thing to push back and we know it's a problem for people
14:55:48 dansmith we also know we can't push it back any farther,
14:55:58 dansmith so it's not like it can go back to kilo or anything
14:56:02 mriedem sure
14:56:03 mriedem ok
15:04:41 cfriesen mriedem: alex_xu: pretty sure cold migration *was* working before all the placement stuff, but I seem to remember seeing at least one bug saying it's currently broken.
15:05:55 cfriesen we're on pike at the moment and it seems to be working, but we do have a few patches in that area for other features.
15:06:07 sean-k-mooney cold migration in what context?
15:06:41 sean-k-mooney mriedem: stephenfin is on PTO until next week
15:06:54 sean-k-mooney mriedem: he left his znc bouncer running
15:07:34 mriedem cfriesen: as alex pointed out,
15:07:45 mriedem i don't see the request spec numa topologies field updated before it goes through the scheduler
15:08:01 sean-k-mooney mriedem: oh alex_xu was asking about numa toplogy changes for resize/cold migration it "should" work upstream
15:08:29 mriedem https://github.com/openstack/nova/blob/master/nova/scheduler/filters/numa_topology_filter.py#L74
15:08:45 mriedem i don't see RequestSpec.numa_topology updated from the new flavor *before* we hit the numa filter
15:09:04 mriedem so i don't see how the scheduling is working
15:09:31 sean-k-mooney hum so you think its using the old flavor perhaps
15:09:47 sean-k-mooney i can try testing this in a hour or so
15:10:03 jmlowe Has anybody had trouble with the placement client in nova compute hanging?
15:10:04 sean-k-mooney i jsut need to get my dev enviornemt running after the weekend
15:10:25 mriedem jmlowe: never heard of that
15:10:45 jmlowe Specifically nova.compute.resource_tracker._update_available_resource was holding a lock for about 2 min
15:11:05 jmlowe Sometimes it would do it and sometimes not
15:11:33 jmlowe no errors anywhere
15:12:21 cdent jmlowe: how many instances on that compute node?
15:12:42 jmlowe It did wreak havoc with instance launches
15:12:43 cdent are you sure it was talking to placement where things were stuck, or just somewhere in _udpate*?
15:12:58 cdent alos
15:13:05 cdent ^walso
15:13:07 cfriesen mriedem: sorry, had to answer a call from my boss. let me take a quick look
15:14:22 jmlowe I got it down to 2 min to run _update_inventory in nova/scheduler/client/report.py
15:15:07 mriedem jmlowe: how many instances on that compute?
15:15:10 jmlowe it seems to have made the http put call correctly, but then just waits for tcp timeout
15:15:13 mriedem libvirt or vcenter or ironic?
15:15:36 jmlowe it's all of my 280 computes, so 0 - 24
15:15:42 jmlowe libvirt
15:16:04 mriedem hmm, well there is a lock in the resource tracker when that is called, which will make things in nova-compute slow to a crawl,
15:16:11 mriedem but why the placement response would be so slow idk,
15:16:21 mriedem have you traced the request via request ID in the placement api logs?
15:16:27 mriedem also, which release?
15:16:44 jmlowe yes, nothing in placement takes more than a second or so
15:16:57 jmlowe queens
15:18:20 jmlowe seems like the there's a bug in some underlying client library that is occasionally not returning from a http call without a timeout
15:18:38 mriedem hmm, nova is using keystoneauth1 to send the requests to placement
15:18:57 mriedem you might be able to enable some debug logging there
15:19:45 cdent jmlowe: I think you should make jeremy find and fix this
15:20:28 cdent not you fungi
15:20:48 jmlowe I did, then went on a 6 week road trip
15:20:55 fungi yay! for once something's not my fault
15:20:56 jmlowe then he went
15:21:08 cdent The jeremy of which I speak is an old friend (on the order of 30 years)
15:21:13 cdent Typical of him.
15:21:17 cfriesen mriedem: alex_xu: I think it's this code that updates the flavor on a resize: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L304-L316
15:22:27 mriedem cfriesen: but that doesn't update the RequestSpec.numa_topology field
15:22:44 mriedem the else block
15:22:57 mriedem or RequestSpec.pci_requests for that matter
15:23:44 mriedem so if resize with a new numa topology works today, it's getting lucky b/c the scheduler picks a host that fits the original topology, and then the RT.move_claim on the compute fits the new requested numa topology
15:23:53 mriedem as far as i can tell anyway
15:24:42 sean-k-mooney mriedem: well we trow away all the host pinning info and recalulate it on the compute node anyway so it is likely working becaue it gets lucky or we retry
15:24:54 cdent jmlowe: do you have load balancers or proxies between n-cpu placement?
15:25:03 jmlowe I do
15:25:10 cdent connections perhaps not closing?
15:25:13 mriedem cdent: fwiw, that ENABLED_PYTHON3_PACKAGES variable in devstack looks like it should be killed now
15:25:26 mriedem "# Special case some services that have experimental
15:25:27 mriedem # in their classifier"
15:25:27 mriedem # support for python3 in progress, but don't claim support
15:25:28 cdent mriedem: it does come into play
15:25:33 mriedem yeah i see where it's used
15:25:39 mriedem but i'm just thinking it shouldn't exist anymore
15:25:44 jmlowe probably, I'd expect the http client to close the connection once it received a 200 responxe
15:25:49 mriedem given the py3 first goal
15:26:02 mriedem e.g. neutron isn't in that list
15:26:04 mriedem nor keystone
15:26:17 sean-k-mooney mriedem: well it either shouldn not exists or default to ture right
15:26:21 cdent mriedem: yes, I agree that it appears to be an anachronism that shouldn't be used at all, but that it sometimes does get used
15:26:33 mriedem i shall put out the dhellmann signal
15:26:33 cdent it is currently what makes my fix work

Earlier   Later