Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-18
16:26:26 dansmith melwitt: yep
16:26:32 melwitt mriedem: ^
16:26:56 melwitt whew, good.
16:27:33 openstackgerrit Merged openstack/nova master: [placement] fix 500 error when allocating to bad class https://review.openstack.org/484162
16:27:40 mriedem melwitt: you mean the recheck performed in conductor?
16:27:48 melwitt mriedem: yeah
16:28:05 gibi jaypipes: http://paste.openstack.org/show/615755/
16:28:33 mriedem yeah, my point earlier today was the difference in ux - before counting qoutas, you'd go overquota and fail the api with a 403 and no instances are created - after we if we catch it in conductor, the instances all get put into ERROR state and you have to clean them up
16:28:38 mriedem melwitt: not the end of the world, but it's different
16:29:32 dansmith now I'm forgetting,
16:29:45 dansmith but the first check is done in api, and the extra if configured check is in conductor, right?
16:29:57 melwitt dansmith: yes
16:29:58 mriedem yeah
16:30:03 mriedem we recheck for cells v1 in the api
16:30:06 dansmith right, so the ux change is only across the boundary
16:30:07 mriedem b/c that's where the instances are created
16:30:10 mriedem but...cellsv1
16:30:11 dansmith which isn't that big of a deal, IMHO
16:30:41 mriedem it's not the end of the world,
16:30:41 dansmith quota is so leaky right now,
16:30:57 mriedem i only ever worry about the NFV robots hammering things in unexpected ways
16:30:58 jaypipes gibi: are you calling GET /allocation_candidates?resources=CUSTOM_MAGIC:512 or are you including VCPU and MEMORY_MB in the request as well?
16:31:08 dansmith we could have added another check to the current stuff where we decided to fail your instance because we healed the quota and the numbers no longer work out
16:31:32 openstackgerrit Merged openstack/nova master: Support tag instances when boot(4/4) https://review.openstack.org/469800
16:31:38 melwitt every quota.reserve() call heals before checking
16:31:54 mriedem heal == refresh()?
16:31:55 melwitt er, it heals YOUR quota. but not others in your project. so you could still be screwed
16:32:26 dansmith melwitt: it doesn't heal negative things, right?
16:32:42 dansmith like when things get all out of whack and require admin and nova-manage intervention
16:32:45 dansmith that's what I meant
16:32:49 gibi jaypipes: I'm calling just MAGIC:512 first, that results in HTTP 500 then I call with both MAGIC and VCPU + MEMORY that goes through but results in an empty response
16:32:51 melwitt I don't think it stores negative usage
16:33:23 jaypipes gibi: can you paste the logs from the second request for VCPU, MAGIC and MEMORY_MB?
16:33:24 melwitt yeah, you need nova-manage when not all users in the cloud are active in booting instances and their quota is out of whack and affecting those who are active
16:33:41 dansmith melwitt: that's all I mean.. when the current broke-ass shit gets broke-ass
16:33:47 melwitt yeah
16:34:11 mriedem can i borrow $5 for some quota?
16:34:39 dansmith mriedem: you're just going to spend it on booze, so no
16:35:14 mriedem s/itches/scratches/
16:35:16 mriedem sorry
16:35:51 melwitt heh. I was thinking about that. lots of people where I grew up did say you itch an itch. which never made sense to me
16:36:38 gibi jaypipes: that generate a lot less log: http://paste.openstack.org/show/615756/
16:37:31 jaypipes dansmith, mriedem: so, here's an interesting little wrinkle... if I request *only* resources that are provided by a sharing provider, and don't request any resources that say, a compute node, would have, there's no way to return an AllocationRequest that would include the compute node resource provider UUID.
16:38:00 dansmith why would you
16:38:01 dansmith ?
16:38:14 dansmith if you don't request any resources from the compute node... you don't get a compute node
16:39:12 jaypipes dansmith: but the compute node is a target because the shared resource is shared with it.
16:39:40 dansmith can you give a concrete example?
16:39:55 dansmith because shared disk should be allocate-able without a compute node if you don't care about CPU/RAM right?
16:40:51 jaypipes dansmith: I actually can't think of a concrete use-case here.
16:40:52 mriedem if you request things that don't result in compute nodes, you just get NoValidHost
16:40:59 mriedem which is fine for nova scheduler right?
16:41:04 dansmith mriedem: from scheduler, but not from placement
16:41:06 mriedem we can't build you a server on a ceph pool alone
16:41:09 mriedem dansmith: sure
16:41:19 jaypipes mriedem: well, technically you get a KeyError right now :) but, even then, I'm not sure.
16:41:27 dansmith right, but it should be fine for cinder to ask for an allocation against a storage provider with nothing else
16:41:34 mriedem yeah so let's fix the KeyError in placement
16:41:39 mriedem right
16:42:23 jaypipes mriedem: yeah, I'm working on fixing that :)
16:49:57 melwitt mriedem: something I was wondering, I noticed the func tests you wrote are in novaclient, so we wouldn't catch a problem unless we look at novaclient jobs
16:50:04 jaypipes https://bugs.launchpad.net/nova/+bug/1705071
16:50:05 openstack Launchpad bug 1705071 in OpenStack Compute (nova) "[placement] Attempting to find allocation candidates for shared-only resources results in KeyError" [High,Triaged]
16:50:11 mriedem notifications meeting in #openstack-meeting-4 in 10 minutes
16:50:36 mriedem melwitt: yup, i did that because they run single tenant in serial
16:50:46 mriedem melwitt: unlike the tempest dsvm jobs
16:51:16 mriedem but idk maybe that doesn't make sense,
16:51:22 mriedem we could also do a functional test in tree
16:51:22 melwitt mriedem: okay, just wanted to make sure that was the intention. I was thinking of the overhead vs the fact that a failure there might go unnoticed for some time
16:51:44 mriedem my goal was the change that runs each test 15 times
16:51:48 mriedem to check for leaks
16:52:07 mriedem but we could also do that in tree tests i reckon
16:52:52 melwitt right. I was looking at them yesterday and wasn't sure whether to +W them once I noticed they're in novaclient. so I wanted to ask you first
16:53:41 mriedem ask and ye shall receive...a meh of an answer
16:53:48 melwitt :)
17:11:14 moshele jaypipes: hi
17:11:34 jaypipes moshele: well hello there :)
17:12:34 moshele jaypipes: can we do goolge hangout ? I have some question regarding the designer/vnic_type/vif/plugins
17:13:18 jaypipes moshele: yes, OK with me. I had asked you, jangutter and sean-k-mooney earlier today to do one.
17:13:20 moshele jaypipes: around https://review.openstack.org/#/c/484197/ and also https://review.openstack.org/#/c/398265/
17:13:57 moshele jaypipes: ha ok I guess I missed that
17:15:09 moshele sean-k-mooney, jangutter : can we do a call tomorrow?
17:15:40 sean-k-mooney moshele: sure but we are cutting it rather close if we want to have support in pike for ovs with hardware offload
17:16:33 moshele sean-k-mooney, jaypipes: what is blocking this than https://review.openstack.org/#/c/398265/
17:17:19 moshele sean-k-mooney: we can do the hw_veb refactor (https://review.openstack.org/#/c/484197/) in queen
17:17:21 sean-k-mooney moshele: it need rodolfos feature based scheduling
17:19:16 moshele sean-k-mooney: I thought we agreed that it will be just limitation with working with SR-IOV. and we will fix it in queens
17:19:32 sean-k-mooney moshele: jaypipes to supprot https://review.openstack.org/#/c/398265 we basically need https://review.openstack.org/#/c/449257/ and https://review.openstack.org/#/c/451777/
17:20:53 sean-k-mooney moshele: i dont think we should merge https://review.openstack.org/#/c/398265 without adressing the sriov work as the port creation will change. e.g. in queens you would need to set the feature request in teh port binings so it existing vms would be broken on upgrade
17:21:02 jaypipes moshele: I'm more concerned about os-vif patches that are dependencies.
17:21:14 jaypipes moshele: since those have a hard freeze date of this Thursday
17:21:21 jaypipes moshele: do we have a list of those patches?
17:21:56 sean-k-mooney jaypipes: you should have an email titled [openstack-dev][os-vif] 1.6.1 release for pike. in your inbox
17:22:22 sean-k-mooney jaypipes: i belive the should have section are the required os-vif patches
17:22:41 jaypipes sean-k-mooney: yes, been going through that
17:23:10 jaypipes sean-k-mooney: the representor ones, yeah?
17:23:21 sean-k-mooney yep these Improve OVS Representor Lookup https://review.openstack.org/#/c/484051/
17:23:22 sean-k-mooney Add support for VIFPortProfileOVSRepresentor https://review.openstack.org/#/c/483921/
17:23:24 sean-k-mooney unplug_vf_passthrough: don't try to delete representor netdev https://review.openstack.org/#/c/478820/
17:23:25 jaypipes gotcha

Earlier   Later