| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-18 | |||
| 16:26:08 | melwitt | dansmith: I guess the rechecking quota stuff takes care of the gap between build requests and instances ... in that after creating the instance objects we check quota again. so maybe that's how not counting build requests can be fine (if someone is configured to be strict about quotas) | |
| 16:26:26 | dansmith | melwitt: yep | |
| 16:26:32 | melwitt | mriedem: ^ | |
| 16:26:56 | melwitt | whew, good. | |
| 16:27:33 | openstackgerrit | Merged openstack/nova master: [placement] fix 500 error when allocating to bad class https://review.openstack.org/484162 | |
| 16:27:40 | mriedem | melwitt: you mean the recheck performed in conductor? | |
| 16:27:48 | melwitt | mriedem: yeah | |
| 16:28:05 | gibi | jaypipes: http://paste.openstack.org/show/615755/ | |
| 16:28:33 | mriedem | yeah, my point earlier today was the difference in ux - before counting qoutas, you'd go overquota and fail the api with a 403 and no instances are created - after we if we catch it in conductor, the instances all get put into ERROR state and you have to clean them up | |
| 16:28:38 | mriedem | melwitt: not the end of the world, but it's different | |
| 16:29:32 | dansmith | now I'm forgetting, | |
| 16:29:45 | dansmith | but the first check is done in api, and the extra if configured check is in conductor, right? | |
| 16:29:57 | melwitt | dansmith: yes | |
| 16:29:58 | mriedem | yeah | |
| 16:30:03 | mriedem | we recheck for cells v1 in the api | |
| 16:30:06 | dansmith | right, so the ux change is only across the boundary | |
| 16:30:07 | mriedem | b/c that's where the instances are created | |
| 16:30:10 | mriedem | but...cellsv1 | |
| 16:30:11 | dansmith | which isn't that big of a deal, IMHO | |
| 16:30:41 | dansmith | quota is so leaky right now, | |
| 16:30:41 | mriedem | it's not the end of the world, | |
| 16:30:57 | mriedem | i only ever worry about the NFV robots hammering things in unexpected ways | |
| 16:30:58 | jaypipes | gibi: are you calling GET /allocation_candidates?resources=CUSTOM_MAGIC:512 or are you including VCPU and MEMORY_MB in the request as well? | |
| 16:31:08 | dansmith | we could have added another check to the current stuff where we decided to fail your instance because we healed the quota and the numbers no longer work out | |
| 16:31:32 | openstackgerrit | Merged openstack/nova master: Support tag instances when boot(4/4) https://review.openstack.org/469800 | |
| 16:31:38 | melwitt | every quota.reserve() call heals before checking | |
| 16:31:54 | mriedem | heal == refresh()? | |
| 16:31:55 | melwitt | er, it heals YOUR quota. but not others in your project. so you could still be screwed | |
| 16:32:26 | dansmith | melwitt: it doesn't heal negative things, right? | |
| 16:32:42 | dansmith | like when things get all out of whack and require admin and nova-manage intervention | |
| 16:32:45 | dansmith | that's what I meant | |
| 16:32:49 | gibi | jaypipes: I'm calling just MAGIC:512 first, that results in HTTP 500 then I call with both MAGIC and VCPU + MEMORY that goes through but results in an empty response | |
| 16:32:51 | melwitt | I don't think it stores negative usage | |
| 16:33:23 | jaypipes | gibi: can you paste the logs from the second request for VCPU, MAGIC and MEMORY_MB? | |
| 16:33:24 | melwitt | yeah, you need nova-manage when not all users in the cloud are active in booting instances and their quota is out of whack and affecting those who are active | |
| 16:33:41 | dansmith | melwitt: that's all I mean.. when the current broke-ass shit gets broke-ass | |
| 16:33:47 | melwitt | yeah | |
| 16:34:11 | mriedem | can i borrow $5 for some quota? | |
| 16:34:39 | dansmith | mriedem: you're just going to spend it on booze, so no | |
| 16:35:14 | mriedem | s/itches/scratches/ | |
| 16:35:16 | mriedem | sorry | |
| 16:35:51 | melwitt | heh. I was thinking about that. lots of people where I grew up did say you itch an itch. which never made sense to me | |
| 16:36:38 | gibi | jaypipes: that generate a lot less log: http://paste.openstack.org/show/615756/ | |
| 16:37:31 | jaypipes | dansmith, mriedem: so, here's an interesting little wrinkle... if I request *only* resources that are provided by a sharing provider, and don't request any resources that say, a compute node, would have, there's no way to return an AllocationRequest that would include the compute node resource provider UUID. | |
| 16:38:00 | dansmith | why would you | |
| 16:38:01 | dansmith | ? | |
| 16:38:14 | dansmith | if you don't request any resources from the compute node... you don't get a compute node | |
| 16:39:12 | jaypipes | dansmith: but the compute node is a target because the shared resource is shared with it. | |
| 16:39:40 | dansmith | can you give a concrete example? | |
| 16:39:55 | dansmith | because shared disk should be allocate-able without a compute node if you don't care about CPU/RAM right? | |
| 16:40:51 | jaypipes | dansmith: I actually can't think of a concrete use-case here. | |
| 16:40:52 | mriedem | if you request things that don't result in compute nodes, you just get NoValidHost | |
| 16:40:59 | mriedem | which is fine for nova scheduler right? | |
| 16:41:04 | dansmith | mriedem: from scheduler, but not from placement | |
| 16:41:06 | mriedem | we can't build you a server on a ceph pool alone | |
| 16:41:09 | mriedem | dansmith: sure | |
| 16:41:19 | jaypipes | mriedem: well, technically you get a KeyError right now :) but, even then, I'm not sure. | |
| 16:41:27 | dansmith | right, but it should be fine for cinder to ask for an allocation against a storage provider with nothing else | |
| 16:41:34 | mriedem | yeah so let's fix the KeyError in placement | |
| 16:41:39 | mriedem | right | |
| 16:42:23 | jaypipes | mriedem: yeah, I'm working on fixing that :) | |
| 16:49:57 | melwitt | mriedem: something I was wondering, I noticed the func tests you wrote are in novaclient, so we wouldn't catch a problem unless we look at novaclient jobs | |
| 16:50:04 | jaypipes | https://bugs.launchpad.net/nova/+bug/1705071 | |
| 16:50:05 | openstack | Launchpad bug 1705071 in OpenStack Compute (nova) "[placement] Attempting to find allocation candidates for shared-only resources results in KeyError" [High,Triaged] | |
| 16:50:11 | mriedem | notifications meeting in #openstack-meeting-4 in 10 minutes | |
| 16:50:36 | mriedem | melwitt: yup, i did that because they run single tenant in serial | |
| 16:50:46 | mriedem | melwitt: unlike the tempest dsvm jobs | |
| 16:51:16 | mriedem | but idk maybe that doesn't make sense, | |
| 16:51:22 | melwitt | mriedem: okay, just wanted to make sure that was the intention. I was thinking of the overhead vs the fact that a failure there might go unnoticed for some time | |
| 16:51:22 | mriedem | we could also do a functional test in tree | |
| 16:51:44 | mriedem | my goal was the change that runs each test 15 times | |
| 16:51:48 | mriedem | to check for leaks | |
| 16:52:07 | mriedem | but we could also do that in tree tests i reckon | |
| 16:52:52 | melwitt | right. I was looking at them yesterday and wasn't sure whether to +W them once I noticed they're in novaclient. so I wanted to ask you first | |
| 16:53:41 | mriedem | ask and ye shall receive...a meh of an answer | |
| 16:53:48 | melwitt | :) | |
| 17:11:14 | moshele | jaypipes: hi | |
| 17:11:34 | jaypipes | moshele: well hello there :) | |
| 17:12:34 | moshele | jaypipes: can we do goolge hangout ? I have some question regarding the designer/vnic_type/vif/plugins | |
| 17:13:18 | jaypipes | moshele: yes, OK with me. I had asked you, jangutter and sean-k-mooney earlier today to do one. | |
| 17:13:20 | moshele | jaypipes: around https://review.openstack.org/#/c/484197/ and also https://review.openstack.org/#/c/398265/ | |
| 17:13:57 | moshele | jaypipes: ha ok I guess I missed that | |
| 17:15:09 | moshele | sean-k-mooney, jangutter : can we do a call tomorrow? | |
| 17:15:40 | sean-k-mooney | moshele: sure but we are cutting it rather close if we want to have support in pike for ovs with hardware offload | |
| 17:16:33 | moshele | sean-k-mooney, jaypipes: what is blocking this than https://review.openstack.org/#/c/398265/ | |
| 17:17:19 | moshele | sean-k-mooney: we can do the hw_veb refactor (https://review.openstack.org/#/c/484197/) in queen | |
| 17:17:21 | sean-k-mooney | moshele: it need rodolfos feature based scheduling | |
| 17:19:16 | moshele | sean-k-mooney: I thought we agreed that it will be just limitation with working with SR-IOV. and we will fix it in queens | |
| 17:19:32 | sean-k-mooney | moshele: jaypipes to supprot https://review.openstack.org/#/c/398265 we basically need https://review.openstack.org/#/c/449257/ and https://review.openstack.org/#/c/451777/ | |
| 17:20:53 | sean-k-mooney | moshele: i dont think we should merge https://review.openstack.org/#/c/398265 without adressing the sriov work as the port creation will change. e.g. in queens you would need to set the feature request in teh port binings so it existing vms would be broken on upgrade | |
| 17:21:02 | jaypipes | moshele: I'm more concerned about os-vif patches that are dependencies. | |
| 17:21:14 | jaypipes | moshele: since those have a hard freeze date of this Thursday | |
| 17:21:21 | jaypipes | moshele: do we have a list of those patches? | |
| 17:21:56 | sean-k-mooney | jaypipes: you should have an email titled [openstack-dev][os-vif] 1.6.1 release for pike. in your inbox | |
| 17:22:22 | sean-k-mooney | jaypipes: i belive the should have section are the required os-vif patches | |
| 17:22:41 | jaypipes | sean-k-mooney: yes, been going through that | |
| 17:23:10 | jaypipes | sean-k-mooney: the representor ones, yeah? | |
| 17:23:21 | sean-k-mooney | yep these Improve OVS Representor Lookup https://review.openstack.org/#/c/484051/ | |
| 17:23:22 | sean-k-mooney | Add support for VIFPortProfileOVSRepresentor https://review.openstack.org/#/c/483921/ | |
| 17:23:24 | sean-k-mooney | unplug_vf_passthrough: don't try to delete representor netdev https://review.openstack.org/#/c/478820/ | |