| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-28 | |||
| 16:30:59 | dansmith | cfriesen: is that in the resource tracker? | |
| 16:31:44 | cfriesen | nova/compute/manager.py | |
| 16:31:56 | dansmith | cfriesen: oh, look at the function | |
| 16:31:58 | dansmith | cfriesen: it returns the first | |
| 16:32:06 | dansmith | "for old compat" | |
| 16:32:35 | dansmith | and since ironic doesn't support live migration...YET | |
| 16:38:14 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: set/check if inventory change in tree https://review.openstack.org/470575 | |
| 16:38:15 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: add nested resource providers https://review.openstack.org/377138 | |
| 16:38:15 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: integrate ProviderTree to report client https://review.openstack.org/415921 | |
| 16:38:16 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: adds REST API for nested providers https://review.openstack.org/384807 | |
| 16:38:16 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: allow filter providers in tree https://review.openstack.org/377215 | |
| 16:38:17 | openstackgerrit | Jay Pipes proposed openstack/nova master: placement: update client to set parent provider https://review.openstack.org/385693 | |
| 16:43:04 | mriedem | stvnoyes: the multinode live migration grenade job failed on your live migration new style attach patch, but it failed here http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/subnode-2/screen-n-cpu.txt.gz?level=TRACE#_Sep_26_14_28_11_637958 | |
| 16:43:25 | mriedem | and looking at logstash, that happens on a ton of patches, so it's probably just a 100% failure in that job right now | |
| 16:43:42 | dansmith | jaypipes: I'm missing the reasoning for the complicated "compare two dicts" method in that first patch | |
| 16:44:38 | jaypipes | dansmith: yeah, it's a long conversation on previous revisions... I'm going to change it to only compare the total field value and ignore the other things like reserved/min_unit/allocation_ratio, etc | |
| 16:45:14 | dansmith | jaypipes: oh, sorry I can re-read.. but, is this not to be used for "know when I need to report new inventory" sort of thing? | |
| 16:45:18 | jaypipes | dansmith: the reason it's like that is because we still don't have any consistency and agreement on who (virt driver or resource tracker) owns various fields about the inventory like alloc ratio etc | |
| 16:45:28 | efried | dansmith The history is here: https://review.openstack.org/#/c/470575/2/nova/compute/provider_tree.py@98 | |
| 16:46:08 | dansmith | jaypipes: ah, to avoid us reporting new just because the virt driver thinks the allocation_ratio should be different? | |
| 16:46:21 | jaypipes | yup | |
| 16:47:05 | dansmith | hmm | |
| 16:47:16 | dansmith | the virt driver needs to report total and reserved | |
| 16:47:42 | dansmith | but allocation_ratio is clearly the purview of the compute manager | |
| 16:47:56 | dansmith | min_unit is probably virt driver I guess | |
| 16:48:15 | dansmith | taking the scaleio example of 8GB minimum slice or whatever | |
| 16:51:06 | mriedem | stvnoyes: yeah the grenade live migration job is just totally busted http://tinyurl.com/y9ca6zox | |
| 16:51:53 | mriedem | something changed in august http://tinyurl.com/yby7m5p8 | |
| 16:52:27 | stvnoyes | that's pretty bad | |
| 16:53:45 | efried | johnthetubaguy jaypipes Where does ironic (or ftm any virt driver) set up its resource providers today? | |
| 16:53:59 | jaypipes | efried: it doesn't. the RT does. | |
| 16:54:06 | jaypipes | efried: because there's only one RP. | |
| 16:54:32 | mriedem | stvnoyes: looking at http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/console.html#_2017-09-26_14_28_03_894125 | |
| 16:54:37 | efried | jaypipes In ironic, only one RP per compute host? Not one RP per ironic node? | |
| 16:54:41 | mriedem | tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok | |
| 16:54:50 | mriedem | tempest.api.compute.admin.test_live_migration.LiveAutoBlockMigrationV225Test.test_live_block_migration [22.911966s] ... ok | |
| 16:54:55 | mriedem | oops | |
| 16:55:02 | jaypipes | efried: we want to get to where there is a sort of mediation/negotiation between the RT and the virt driver and the generic device manager as each of them "processes" the ProviderTree and sets inventory records and traits based on their own information. | |
| 16:55:03 | mriedem | tempest.api.compute.admin.test_live_migration.LiveMigrationTest.test_live_block_migration [10.052345s] ... FAILED | |
| 16:55:17 | mriedem | so it works with the 2.25 microversion, where we pass block_migration=auto | |
| 16:55:21 | mriedem | but fails before that | |
| 16:55:22 | jaypipes | efried: in ironic, each baremetal Ironic node is its own RP in placementy. | |
| 16:55:48 | mriedem | so i bet something with the auto changes broke how we do live block migration between n-1 hosts | |
| 16:56:25 | stvnoyes | what were the auto changes? | |
| 16:56:26 | efried | jaypipes Aha. Then bug 1714248 exists because... the compute host is creating its own separate RP for some reason? | |
| 16:56:28 | openstack | bug 1714248 in OpenStack Compute (nova) "Compute node HA for ironic doesn't work due to the name duplication of Resource Provider " [High,Confirmed] https://launchpad.net/bugs/1714248 | |
| 16:56:53 | jaypipes | efried: not sure, lemme read | |
| 16:57:18 | efried | oh, or because the RT is trying to create a new RP for the node it's taking over... | |
| 16:57:29 | mriedem | stvnoyes: https://docs.openstack.org/nova/pike/reference/api-microversion-history.html#maximum-in-mitaka | |
| 16:57:29 | jaypipes | efried: note that nova-compute was *never* intended for "HA setups"... | |
| 16:57:39 | mriedem | ^ predates the spike in failures though, which started around 8/18 | |
| 16:57:48 | efried | ...And that's out of the control of ironic-specific code... | |
| 16:58:28 | efried | which is why my comment #5 is a non-starter | |
| 16:58:42 | efried | until, as you say, there's some coordination between RT and virt to manage the RPs. | |
| 16:59:00 | stvnoyes | mriedem: brb- grabbing lunch | |
| 17:00:53 | mriedem | stvnoyes: ok. it also fails on the pike compute... | |
| 17:04:22 | johnthetubaguy | efried: I have to head out now and cook my dinner, but I think this maps nicely to how we setup traits, I am leaning towards ironic doing things in placement | |
| 17:05:02 | efried | johnthetubaguy Okay. I've got a passel of draft comments on your spec, which should be ready for your perusal next time you're on. | |
| 17:05:13 | johnthetubaguy | cool, thanks | |
| 17:06:12 | dansmith | jaypipes: in the db migration, why aren't we referencing the parent provider by id instead of uuid? | |
| 17:06:36 | dansmith | jaypipes: and, isn't recording the root just going to limit us later when we need to restructure a tree? | |
| 17:06:39 | efried | johnthetubaguy (Looking back over 'em, removing the ones we've talked about in here, the only things left are typos ): | |
| 17:07:39 | jaypipes | dansmith: a good question on the parent provider UUID thing. not entirely sure why I did that. | |
| 17:09:22 | dansmith | jaypipes: I can't tell you how elated I feel -1ing a db schema change of yours for performance reasons | |
| 17:09:34 | jaypipes | dansmith: recording the root is an optimization to avoid needing to do hierarchical queries. and we don't restructure the root, only potentially children within the tree (in other words, root_provider_id won't change. | |
| 17:09:41 | jaypipes | dansmith: :) | |
| 17:10:05 | dansmith | jaypipes: right, it won't for compute nodes, but it could for other types of resources | |
| 17:10:22 | dansmith | jaypipes: like you move a disk shelf from one NAS device to another | |
| 17:11:08 | jaypipes | dansmith: possibly, sure, but those kinds of moves are few and far between in comparison to the 10 or 100 times as many read requests for tree data | |
| 17:11:44 | dansmith | hmm | |
| 17:13:33 | dansmith | jaypipes: well, I commented for later | |
| 17:13:45 | jaypipes | okey dokey | |
| 17:14:24 | dansmith | jaypipes: I haven't gotten to how we tell placement about our parent, but presumably we're forbidden from trying to reparent a first-level provider? | |
| 17:15:05 | jaypipes | dansmith: haven't gotten to that either. | |
| 17:15:54 | dansmith | jaypipes: also, in case you're wondering | |
| 17:16:11 | dansmith | jaypipes: yes, it's amazingly beautiful out here on the deck... 71F and clear skies | |
| 17:16:20 | jaypipes | dansmith: lol | |
| 17:26:57 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Alternate Hosts https://review.openstack.org/504275 | |
| 17:32:52 | mriedem | stvnoyes: aha | |
| 17:32:58 | mriedem | it's a ci job configuration issue | |
| 17:33:06 | mriedem | the nodes are configured differently for block migration | |
| 17:36:05 | stvnoyes | mriedem: that's good to hear. much better than a code/upgrade issue. | |
| 17:36:46 | mriedem | i'm pretty sure i've had to fix this before... | |
| 17:44:37 | eandersson | Is versioned notifications properly implemented in Mitaka? We don't see any versioned notifications being sent. | |
| 17:44:40 | openstackgerrit | Dan Smith proposed openstack/nova master: Use improved instance_list module in compute API https://review.openstack.org/505418 | |
| 17:44:41 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix minor input items from previous patches https://review.openstack.org/506416 | |
| 17:44:41 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix CellDatabases fixture swallowing exceptions https://review.openstack.org/506312 | |
| 17:56:40 | mriedem | eandersson: probably not at that point | |
| 17:57:11 | mriedem | introduced in newton https://specs.openstack.org/openstack/nova-specs/specs/newton/implemented/versioned-notification-transformation-newton.html | |
| 17:57:37 | mriedem | the framework code was in mitaka | |
| 17:58:03 | eandersson | I see - thanks mriedem | |
| 18:11:13 | mriedem | mtreinish: do you know anything about how tempest is upgraded in grenade? | |
| 18:12:13 | mtreinish | mriedem: it's not, the code and config should be the same between versions | |
| 18:12:20 | mtreinish | everything should be master for tempest | |
| 18:12:20 | sean-k-mooney | mriedem: i taught tempest was not upgreaed in grenade because it was not ment to be version specific | |
| 18:12:33 | mriedem | mtreinish: the config is different between pike and queens | |
| 18:12:43 | mtreinish | mriedem: links? | |
| 18:12:47 | mriedem | at least in this multinode grenade job | |
| 18:12:48 | mriedem | http://logs.openstack.org/87/463987/20/check/gate-grenade-dsvm-neutron-multinode-live-migration-nv/ae8875f/logs/old/tempest_conf.txt.gz | |
| 18:12:51 | mriedem | ^ pike | |