Earlier  
Posted Nick Remark
#openstack-nova - 2018-11-26
15:12:21 cdent jmlowe: how many instances on that compute node?
15:12:42 jmlowe It did wreak havoc with instance launches
15:12:43 cdent are you sure it was talking to placement where things were stuck, or just somewhere in _udpate*?
15:12:58 cdent alos
15:13:05 cdent ^walso
15:13:07 cfriesen mriedem: sorry, had to answer a call from my boss. let me take a quick look
15:14:22 jmlowe I got it down to 2 min to run _update_inventory in nova/scheduler/client/report.py
15:15:07 mriedem jmlowe: how many instances on that compute?
15:15:10 jmlowe it seems to have made the http put call correctly, but then just waits for tcp timeout
15:15:13 mriedem libvirt or vcenter or ironic?
15:15:36 jmlowe it's all of my 280 computes, so 0 - 24
15:15:42 jmlowe libvirt
15:16:04 mriedem hmm, well there is a lock in the resource tracker when that is called, which will make things in nova-compute slow to a crawl,
15:16:11 mriedem but why the placement response would be so slow idk,
15:16:21 mriedem have you traced the request via request ID in the placement api logs?
15:16:27 mriedem also, which release?
15:16:44 jmlowe yes, nothing in placement takes more than a second or so
15:16:57 jmlowe queens
15:18:20 jmlowe seems like the there's a bug in some underlying client library that is occasionally not returning from a http call without a timeout
15:18:38 mriedem hmm, nova is using keystoneauth1 to send the requests to placement
15:18:57 mriedem you might be able to enable some debug logging there
15:19:45 cdent jmlowe: I think you should make jeremy find and fix this
15:20:28 cdent not you fungi
15:20:48 jmlowe I did, then went on a 6 week road trip
15:20:55 fungi yay! for once something's not my fault
15:20:56 jmlowe then he went
15:21:08 cdent The jeremy of which I speak is an old friend (on the order of 30 years)
15:21:13 cdent Typical of him.
15:21:17 cfriesen mriedem: alex_xu: I think it's this code that updates the flavor on a resize: https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L304-L316
15:22:27 mriedem cfriesen: but that doesn't update the RequestSpec.numa_topology field
15:22:44 mriedem the else block
15:22:57 mriedem or RequestSpec.pci_requests for that matter
15:23:44 mriedem so if resize with a new numa topology works today, it's getting lucky b/c the scheduler picks a host that fits the original topology, and then the RT.move_claim on the compute fits the new requested numa topology
15:23:53 mriedem as far as i can tell anyway
15:24:42 sean-k-mooney mriedem: well we trow away all the host pinning info and recalulate it on the compute node anyway so it is likely working becaue it gets lucky or we retry
15:24:54 cdent jmlowe: do you have load balancers or proxies between n-cpu placement?
15:25:03 jmlowe I do
15:25:10 cdent connections perhaps not closing?
15:25:13 mriedem cdent: fwiw, that ENABLED_PYTHON3_PACKAGES variable in devstack looks like it should be killed now
15:25:26 mriedem "# Special case some services that have experimental
15:25:27 mriedem # support for python3 in progress, but don't claim support
15:25:27 mriedem # in their classifier"
15:25:28 cdent mriedem: it does come into play
15:25:33 mriedem yeah i see where it's used
15:25:39 mriedem but i'm just thinking it shouldn't exist anymore
15:25:44 jmlowe probably, I'd expect the http client to close the connection once it received a 200 responxe
15:25:49 mriedem given the py3 first goal
15:26:02 mriedem e.g. neutron isn't in that list
15:26:04 mriedem nor keystone
15:26:17 sean-k-mooney mriedem: well it either shouldn not exists or default to ture right
15:26:21 cdent mriedem: yes, I agree that it appears to be an anachronism that shouldn't be used at all, but that it sometimes does get used
15:26:33 cdent it is currently what makes my fix work
15:26:33 mriedem i shall put out the dhellmann signal
15:26:48 sean-k-mooney mriedem: there are some project that dont rung on python3 i belive
15:26:57 jmlowe Testing now, but it seems I can mitigate all of this with the timeout setting in the placement section of nova.conf
15:27:46 mriedem cdent: a peach?
15:28:27 mriedem jmlowe: do you have to set connection timeouts for other nova interactions, like with glance?
15:28:42 jmlowe I've never needed them before
15:29:39 mriedem weird. i think in queens we were going through ksa for glance interactions as well
15:29:45 mriedem so it should all be the same
15:29:54 mriedem from a client perspective i mean,
15:29:57 mriedem just different api servers
15:31:31 jmlowe I have been seeing the occasional ssl error in neutron clients, so could even be as simple as some sort of bug in haproxy
15:31:39 cfriesen mriedem: sean-k-mooney: yeah, I think you might be right, and most of the time it gets fixed up by the claim. ick.
15:33:23 sean-k-mooney cfriesen: mriedem we need to do a cleanup/audit of the migraion codepaths in general. artoms live migration spec will help but we likely have some cleanup to do for resize/cold migrate too.
15:34:07 cdent jmlowe: what's hosting your placement service? mod_wsgi, uwsgi, something else?
15:35:05 sean-k-mooney cfriesen: mriedem i know in pratcice cold migrate and resize "usually" work and results in the numa toplogy being recaluated but it may be just getting lucky or relying on retries. when we model numa in placemnt this will cange and we will have to correct the behavior for that
15:36:00 sean-k-mooney cfriesen: by the way i dont know if you saw my ML post regarding the pci numa affitiy policies
15:36:40 cdent efried: thanks for the +w on the external placement in nova change. I assume we want to let its child ( https://review.openstack.org/#/c/618215/ ) sit until more things have settled
15:36:43 cfriesen sean-k-mooney: glanced at it. haven't had a chance to dig in to it, they've got my focused on other stuff at the moment.
15:37:04 sean-k-mooney cfriesen: i did not test all the ploices but at least the prefer policy does not work as intended so we will need to fix that
15:37:14 efried cdent: Been avoiding that one until "WIP" goes away.
15:37:26 sean-k-mooney + add the ablity for it to work with neutron sriov ports
15:38:01 efried cdent: Basically taking any excuse to defer reviews while I try to get my feet back under me.
15:38:24 cdent efried: i've left it wip in expectation of it needing to hold. How do I mark something "i'd like reviews but this isn't done"? But yeah, understand the need to defer.
15:38:45 cdent (sometimes its the concept not the implementation that needs the review)
15:39:20 efried cdent: If you'd like me to -2 it, I could do that. Or you could -W it. Not sure either would help get it more attention, though.
15:39:28 mriedem cdent: should probably send a separate email to the ML about https://review.openstack.org/#/c/617941/ so nova people not paying much attention to placement know what's going on with tests now
15:39:30 efried what are we actually waiting on?
15:39:44 cdent mriedem: aye aye
15:40:02 mriedem b/c when i rebase and my tests start failing i'm going to wonder wtf
15:41:22 sean-k-mooney cdent: mriedem efried is https://review.openstack.org/#/c/599208/ still considerd a requiremetn before the placement extraction is complete
15:41:32 cdent yes
15:42:04 sean-k-mooney is there a list of what is out standing beyond that or is that the final feature
15:42:14 mriedem https://etherpad.openstack.org/p/BER-placement-extract
15:42:22 sean-k-mooney mriedem: cool thanks
15:43:06 mriedem bauzas: did you get a functional test written that does a reshape and then schedules to the child provider inventory resource?
15:44:32 mriedem cdent: why were these tests removed? https://review.openstack.org/#/c/617941/21/nova/tests/unit/cmd/test_status.py
15:44:35 mnaser https://review.openstack.org/#/c/619349/ simple backport if anyone has a second (i'll babysit the rest)
15:46:12 mriedem lyarwood: ^ just consider my backport a proxy +2 there
15:46:31 cdent mriedem: because the tests uses the rp_objects directly
15:46:48 cdent the follow on patch may even remove the status command since it can no longer work
15:47:08 cdent there's a fixme added in cmd/status.py
15:47:09 openstackgerrit Balazs Gibizer proposed openstack/nova master: Send RP uuid in the port binding https://review.openstack.org/569459
15:47:10 openstackgerrit Balazs Gibizer proposed openstack/nova master: Test boot with more ports with bandwidth request https://review.openstack.org/573317
15:47:23 mriedem cdent: ok but the rp objects aren't removed in that patch, so it seemed out of place
15:47:30 mriedem we could count rps using the placement api right?
15:47:33 mriedem or using the fixture
15:48:37 cdent I think that one can still work, but not the inventory-related one
15:48:54 cdent I went through so many iterations on that series of changes, I may have lost my place

Earlier   Later