| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-08-09 | |||
| 18:47:19 | mriedem | wrong one | |
| 18:47:19 | mriedem | https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes | |
| 18:47:25 | mriedem | The scheduler now requests allocation candidates from the Placement service during scheduling. The allocation candidates information was introduced in the Placement API 1.10 microversion, so you should upgrade the placement service before the Nova scheduler service so that the scheduler can take advantage of the allocation candidate information. | |
| 18:47:42 | mriedem | the code totally falls back though if 1.10 isn't available | |
| 18:48:08 | mriedem | if we want to make 1.10 required, that's fine, but we need to also update nova-status, which i could do lickety split | |
| 18:48:10 | dansmith | yeah I thought it was in the prelude that I read this morning | |
| 18:48:28 | cdent | I think we should require 1.10 | |
| 18:48:34 | dansmith | there's something that specifically says "be sure to upgrade placement first" | |
| 18:48:41 | dansmith | I really thought that was the prelude | |
| 18:48:48 | mriedem | that's https://docs.openstack.org/releasenotes/nova/unreleased.html#upgrade-notes | |
| 18:49:04 | mriedem | the language there is more clear, the behavior in the scheduler code is not | |
| 18:49:09 | dansmith | ah it's thjs: https://review.openstack.org/#/c/491900/2/doc/source/user/placement.rst | |
| 18:49:17 | dansmith | pike placement upgrade notes | |
| 18:49:32 | cdent | if we haven’t got 1.10, we haven’t got alloc_candidates and we’ve not forced things forward, let’s do that | |
| 18:49:43 | mriedem | i'll update nova-status quick | |
| 18:50:23 | cdent | If nobody gets to it tonight, I can and will look into the stuff where tempest is spewing ERROR tomorrow morning | |
| 18:50:38 | mriedem | it's because we have a small flavor, 1 VCPU | |
| 18:50:39 | dansmith | mriedem: did you say you think the errors during single-host migration come from the bottom patch? | |
| 18:50:53 | mriedem | when we resize on the same host, that subtracts the flavor from itself, resulting in 0 VCPU | |
| 18:50:55 | mriedem | which is an invalid allocation | |
| 18:51:03 | mriedem | dansmith: from the middle patch | |
| 18:51:07 | mriedem | the pep8 is on the bottom patch | |
| 18:51:08 | dansmith | okay I was going to say.. | |
| 18:51:21 | dansmith | ah | |
| 18:51:44 | mriedem | so we'd need some additional floor type logic | |
| 18:51:46 | cdent | isn’t the allocation supposed to be doubling to 2 already? where’s that getting lost? | |
| 18:51:48 | mriedem | such that an allocation can't dip below 0 | |
| 18:52:08 | mriedem | cdent: good question | |
| 18:52:24 | mriedem | Instance 6fa7b953-c1fe-4520-8564-aeba8e90aece has resources on 1 compute nodes | |
| 18:52:28 | mriedem | Original resources from same-host allocation: {u'VCPU': 1, u'MEMORY_MB': 64} | |
| 18:52:33 | mriedem | Subtracting old resources from same-host allocation: {u'VCPU': 0, u'MEMORY_MB': 0, 'DISK_GB': 0} | |
| 18:52:36 | mriedem | then kablammo | |
| 18:52:39 | dansmith | mriedem: we're doing the -1 merge which should be subtracting us from the doubled allocation | |
| 18:53:20 | mriedem | yeah the allocation doesn't appear to be doubled | |
| 18:53:21 | mriedem | is the problem | |
| 18:53:44 | dansmith | so | |
| 18:53:45 | dansmith | in that patch, | |
| 18:53:52 | dansmith | we're still healing on both sides with reckless abandon | |
| 18:53:56 | dansmith | does it show up in thelast patch? | |
| 18:53:59 | dansmith | that's where we stop doing that | |
| 18:54:09 | dansmith | although that patch has a typo in it, so probably isn't running properly anyway | |
| 18:54:23 | dansmith | er I mean, had until I updated it a few minutes ago | |
| 18:56:50 | mriedem | umm | |
| 18:56:51 | mriedem | http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-cpu.txt.gz#_Aug_09_16_27_56_298572 | |
| 18:56:58 | mriedem | Sending updated allocation [{'resource_provider': {'uuid': '97ce076a-2644-465b-95dc-fc0674152976'}, 'resources': {u'VCPU': 0, u'MEMORY_MB': 0, 'DISK_GB': 0}}] for instance 6fa7b953-c1fe-4520-8564-aeba8e90aece after removing resources for 97ce076a-2644-465b-95dc-fc0674152976. {{(pid=19567) remove_provider_from_instance_allocation /opt/stack/new/nova/nova/scheduler/client/report.py:1091}} | |
| 18:56:58 | cdent | ah, so doubling gets cleaned up too soon (until later in the series)? | |
| 18:57:12 | mriedem | oh nvm | |
| 18:57:17 | mriedem | that's known, i was thinking that was the bottom change | |
| 18:59:17 | mriedem | the current_allocs get logged but they don't have 2 VCPU | |
| 18:59:20 | jaypipes | dansmith: are you actively making changes to these patches? if not, I can address mriedem's comments. | |
| 18:59:26 | mriedem | so the doubling isn't happening it appears | |
| 19:00:43 | mriedem | this is the resize in the scheduler http://logs.openstack.org/10/488510/30/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/5a36c66/logs/screen-n-sch.txt.gz#_Aug_09_16_27_37_363471 | |
| 19:00:53 | mriedem | req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 | |
| 19:01:12 | openstackgerrit | Merged openstack/nova master: doc: provide more details on scheduling with placement https://review.openstack.org/491900 | |
| 19:01:49 | mriedem | req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 | |
| 19:01:51 | mriedem | oops | |
| 19:01:56 | mriedem | Aug 09 16:27:37.633395 ubuntu-xenial-internap-mtl01-10347573 nova-scheduler[17280]: DEBUG nova.scheduler.client.report [None req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 tempest-MigrationsAdminTest-402050297 tempest-MigrationsAdminTest-402050297] Doubling-up allocation request for move operation. {{(pid=17280) _move_operation_alloc_request /opt/stack/new/nova/nova/scheduler/client/report.py:162}} | |
| 19:02:01 | mriedem | it says it's double stuffing it | |
| 19:02:05 | openstackgerrit | Merged openstack/nova master: Add a prelude section for Pike https://review.openstack.org/491424 | |
| 19:02:19 | mriedem | and it looks like it does | |
| 19:02:20 | mriedem | Aug 09 16:27:37.633725 ubuntu-xenial-internap-mtl01-10347573 nova-scheduler[17280]: DEBUG nova.scheduler.client.report [None req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 tempest-MigrationsAdminTest-402050297 tempest-MigrationsAdminTest-402050297] New allocation request containing both source and destination hosts in move operation: {'allocations': [{'resource_provider': {'uuid': u'97ce076a-2644-465b-95dc-fc0674152976'}, 'resource | |
| 19:02:20 | mriedem | {u'VCPU': 2, u'MEMORY_MB': 192}}]} {{(pid=17280) _move_operation_alloc_request /opt/stack/new/nova/nova/scheduler/client/report.py:202}} | |
| 19:02:58 | openstackgerrit | Merged openstack/nova master: Mark max microversion for Pike in history doc https://review.openstack.org/491581 | |
| 19:03:10 | mriedem | but then i see this | |
| 19:03:10 | mriedem | Aug 09 16:27:37.713717 ubuntu-xenial-internap-mtl01-10347573 nova-scheduler[17280]: DEBUG nova.scheduler.filter_scheduler [None req-7cbe9eeb-dc55-4986-be30-99bd7a03d2c4 tempest-MigrationsAdminTest-402050297 tempest-MigrationsAdminTest-402050297] Successfully claimed resources for instance 6fa7b953-c1fe-4520-8564-aeba8e90aece using allocation request {u'allocations': [{u'resource_provider': {u'uuid': u'97ce076a-2644-465b-95dc- | |
| 19:03:10 | mriedem | 74152976'}, u'resources': {u'VCPU': 1, u'MEMORY_MB': 128}}]} {{(pid=17280) _claim_resources /opt/stack/new/nova/nova/scheduler/filter_scheduler.py:289}} | |
| 19:03:22 | mriedem | which is the single allocation again | |
| 19:03:22 | mriedem | wtf | |
| 19:03:42 | openstackgerrit | Merged openstack/nova master: Document service layout for consoles with cells https://review.openstack.org/491914 | |
| 19:03:45 | mriedem | OHHHHHHHHH | |
| 19:03:47 | mriedem | i know why | |
| 19:04:02 | mriedem | mfing pass by reference | |
| 19:04:14 | openstackgerrit | Merged openstack/nova master: Cleanup release note about ignoring allow_same_net_traffic https://review.openstack.org/491855 | |
| 19:04:22 | jaypipes | mriedem: that's passing by reference. | |
| 19:04:25 | mriedem | dansmith: jaypipes: cdent: problem is right here https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L198 | |
| 19:04:31 | mriedem | return new_alloc_req | |
| 19:04:36 | mriedem | but we're not touching new_alloc_req | |
| 19:04:52 | mriedem | so we double stuff but don't persist | |
| 19:05:07 | jaypipes | yep | |
| 19:05:54 | cdent | oops | |
| 19:06:25 | cdent | does any of gibi’s pending functional stuff cover that? | |
| 19:06:35 | mriedem | not sure | |
| 19:06:45 | jaypipes | func or unit should have covered this... | |
| 19:06:50 | mriedem | anyway, i'll get this nova-status thing done | |
| 19:06:58 | mriedem | and then could push a fix for that on top | |
| 19:07:01 | jaypipes | mriedem: k, I'll fix the above. | |
| 19:07:02 | mriedem | and then we rebase the series on top of those | |
| 19:07:04 | mriedem | or that | |
| 19:07:13 | jaypipes | mriedem: you tell me. what do you want? | |
| 19:07:38 | cdent | I really need to go, but please, if stuff’s not stable by the time you guys expire tonight, let me know so I can continue it in the morning | |
| 19:07:45 | jaypipes | mriedem: I can tack on a fix on the bottom of this series? | |
| 19:07:54 | mriedem | jaypipes: yes it would have to be | |
| 19:08:08 | mriedem | note i've got other issues in that 2nd change though, | |
| 19:08:12 | mriedem | which i guess could be follow ups | |
| 19:08:22 | mriedem | i'm just kind of getting tired of making myself TODOs about chasing follow ups | |
| 19:08:25 | mriedem | but realize we're short on time | |
| 19:08:58 | jaypipes | mriedem: I'm happy to address your comments on the second patch. just need to know if dansmith is actively working on these. | |
| 19:09:08 | dansmith | nope | |
| 19:09:44 | jaypipes | dansmith: gotcha. I'll get on this then. | |
| 19:10:37 | cdent | k, in that case I’ll look at the existing main stack in the morning and see where we’re at | |