| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-26 | |||
| 21:10:10 | mriedem | *needed | |
| 21:10:21 | mriedem | melwitt: plus i think we did one or more of these same granularity policy things in ocata, | |
| 21:10:32 | melwitt | mriedem: okay, cool. thanks, I learned a thing | |
| 21:10:34 | mriedem | the key is it must be backward compatible with an existing policy json i think | |
| 21:10:54 | mriedem | so an operator would need to opt into the more granular rules | |
| 21:11:19 | melwitt | right. I think they are taking care of that in the patch | |
| 21:11:23 | melwitt | cool | |
| 21:15:06 | openstackgerrit | melanie witt proposed openstack/nova master: deprecate ``wsgi_log_format`` config variable https://review.openstack.org/486623 | |
| 21:24:33 | bauzas | jaypipes: just a question about https://review.openstack.org/#/c/487589/1/nova/scheduler/client/report.py | |
| 21:25:15 | bauzas | jaypipes: when we self-heal by the RT, we remove allocations that are not related to the existing instances, right? | |
| 21:25:34 | mriedem | jaypipes: issues in https://review.openstack.org/#/c/487589 | |
| 21:27:39 | dansmith | bauzas: that's what I said in my comment | |
| 21:27:48 | jaypipes | bauzas: when the move_claim() completes on the destination host, it will overwrite the allocations to only be the ones on the destination host, yes. I think that's what you're asking? | |
| 21:27:58 | jaypipes | mriedem: blasted? | |
| 21:28:10 | dansmith | jaypipes: no | |
| 21:28:26 | dansmith | jaypipes: he's asking about regular RT healing on the source node while the migration is going on, erasing the double claim | |
| 21:28:57 | bauzas | dansmith: my question is about if either the source or the target RT removes the allocations for the moving instance given the instance.host is not related to it | |
| 21:29:00 | dansmith | jaypipes: which was in my comment about the plan.. we need to make sure we don't heal over that | |
| 21:29:10 | openstackgerrit | Eric Fried proposed openstack/nova master: WIP: Set Adapter interface defaults in conf https://review.openstack.org/487581 | |
| 21:29:14 | dansmith | bauzas: instance.host will be related to the source until it completes | |
| 21:29:21 | dansmith | bauzas: and it would remove the double claim once it runs | |
| 21:29:28 | bauzas | dansmith: that's my question | |
| 21:29:33 | dansmith | right | |
| 21:29:56 | mriedem | jaypipes: with questions | |
| 21:30:05 | bauzas | dansmith: because if instance.host is set to the source host, wouldn't the target RT removing the allocations for the target host ? | |
| 21:30:18 | dansmith | again, yes | |
| 21:31:16 | edleafe | mriedem: regarding https://bugs.launchpad.net/nova/+bug/1706772 - if we catch that and move on, the flavors will never migrate. What's the alternative? | |
| 21:31:17 | openstack | Launchpad bug 1706772 in OpenStack Compute (nova) "InternalServerError: Internal Server Error (HTTP 500) in n-cpu logs on startup with Ironic driver" [High,Confirmed] | |
| 21:31:21 | jaypipes | dansmith: do you want me to put something in the RT's update_available_resource() method that basically says "oh, this is migrating? fuck it, don't touch placement"? | |
| 21:31:32 | dansmith | jaypipes: we have to do something yeah | |
| 21:31:45 | mriedem | i had a comment in the patch related to this: "when the move_claim() completes on the destination host, it will overwrite the allocations to only be the ones on the destination host, yes. " | |
| 21:31:49 | bauzas | dansmith: sorry, I misunderstood your comment, I thought you were saying it wasn't a problem | |
| 21:31:51 | jaypipes | dansmith: we already jump through a shit-ton of hoops for move operations in the RT... | |
| 21:31:54 | jaypipes | what's one more... | |
| 21:32:04 | dansmith | jaypipes: well, it's either correct or it's not... | |
| 21:32:13 | mriedem | is PUT /allocations/<intsance uuid> completely overwriting the allocations for that instance? | |
| 21:32:21 | dansmith | mriedem: yes | |
| 21:32:23 | bauzas | mriedem: that is correct AFAIK | |
| 21:32:24 | jaypipes | dansmith: no, for move operations the definition of "correct" is fuzzy. | |
| 21:32:34 | dansmith | jaypipes: I disagree :) | |
| 21:33:18 | mriedem | ah i see | |
| 21:35:44 | mriedem | so this works the same if you're doing a resize and revert back to the source host i think | |
| 21:36:17 | jaypipes | I'm wondering if instances that are migrating are in RT.tracked_instances... | |
| 21:36:22 | jaypipes | if they aren't, we're good. | |
| 21:36:42 | dansmith | jaypipes: well, the existing RT stuff is not even correct, as you know | |
| 21:36:49 | bauzas | jaypipes: well, I don't think so | |
| 21:37:07 | bauzas | jaypipes: IIRC, tracked_instances is for existing instances | |
| 21:37:13 | dansmith | jaypipes: if you have the instance, you should be able to check instance.migration_context to know if it's moving | |
| 21:37:15 | bauzas | not for migrating ones | |
| 21:37:23 | jaypipes | ok, guys, I think we're good... | |
| 21:37:25 | dansmith | bauzas: for migrating ones on the source they should be there right? | |
| 21:37:26 | jaypipes | lemme explain. | |
| 21:37:27 | jaypipes | pls. | |
| 21:37:51 | bauzas | dansmith: for the source RT, yeah they should be there AFAIK | |
| 21:38:01 | dansmith | bauzas: right, destination node does not matter | |
| 21:38:05 | jaypipes | so, in update_available_resource(), we call _update_usage_from_instance(). this is the "auto-heal" thing. | |
| 21:38:13 | bauzas | correct | |
| 21:38:35 | jaypipes | within that method, we only delete the allocation if the instance is in DELETED or SHELVE_OFFLOADED state | |
| 21:38:40 | bauzas | the problem is how we could possibly have duplicate allocations for both target and source if source just removes the target allocations ? | |
| 21:38:47 | jaypipes | otherwise we don't touch the allocations. | |
| 21:38:53 | dansmith | jaypipes: eh? | |
| 21:39:07 | dansmith | jaypipes: we compare the generated ones to the ones from placement and re-put them if they differ | |
| 21:39:17 | dansmith | that's how we get allocations now | |
| 21:39:19 | bauzas | jaypipes: PUT /allocations/<instance_uuid> is cleaning up existing allocs, nope ? | |
| 21:39:49 | bauzas | FWIW, it's becoming late and I could be wrong | |
| 21:40:00 | dansmith | jaypipes: this: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L890-L904 | |
| 21:40:16 | mriedem | edleafe: need to find out when the nova-compute service with the ironic driver is able to talk to ironic | |
| 21:40:57 | jaypipes | dansmith: when does update_instance_allocation() get called though? | |
| 21:41:03 | dansmith | jaypipes: oh, are you looking at the "is new or is old" bit of RT? | |
| 21:41:10 | jaypipes | dansmith: CORRECT! | |
| 21:41:33 | dansmith | jaypipes: okay, so that's a problem for the failure case then, | |
| 21:41:43 | jaypipes | dansmith: so we don't actually call that update_instance_allocation() unless it's either a brand new instance or it's DELETED/SHEVE_OFFLOADED | |
| 21:41:52 | dansmith | jaypipes: because that means we'll never heal over the double allocation on the source when the migration is canceled | |
| 21:42:05 | edleafe | mriedem: sure, but my point was that since the code runs in init_host, it won't get a second chance to run *after* the ironic service starts up | |
| 21:42:58 | mriedem | edleafe: at some point on startup the compute manager is calling get_inventory which has to refresh the node list | |
| 21:43:05 | mriedem | b/c of the pre_start_hook in the compute manager | |
| 21:43:15 | mriedem | so does that just not work today? or are we just not waiting long enough? | |
| 21:43:25 | dansmith | mriedem: edleafe: riht, that probably should run when we get a new node | |
| 21:43:33 | dansmith | because you can get/lose nodes on ironic at runtime | |
| 21:44:14 | dansmith | because we could have started with one node, migrated those, and then gained a couple more nodes later when someone shuts down another ironic compute in the hash ring | |
| 21:44:20 | edleafe | dansmith: I thougth a new node would be empty to start | |
| 21:44:25 | edleafe | no instance | |
| 21:44:26 | dansmith | and if that is done during upgrade, which it is, then you need to migrate it | |
| 21:44:29 | dansmith | edleafe: no, because ^ | |
| 21:45:04 | dansmith | edleafe: nova-computes shard off the full set of ironic nodes, | |
| 21:45:15 | dansmith | and if you were to upgrade one ironic compute, then shut down your old one, | |
| 21:45:31 | dansmith | you'd start with some nodes, and then later get a bunch more when the ring rebalances | |
| 21:45:44 | dansmith | after init_host(), at runtime, old nodes with instances that you now own and need to migrate | |
| 21:46:18 | edleafe | dansmith: so then what IYO would be a better place for this? | |
| 21:46:50 | dansmith | edleafe: I'd have to go dig just like you, but there's a place in there where we rebalance the ring (or balance it for the first time) and get a list of the nodes we own | |
| 21:46:51 | cfriesen_ | jaypipes: back at the last PTG did you arrive at any conclusions on how to handle the Intel CAT stuff? | |
| 21:46:54 | dansmith | edleafe: so .. there. :) | |
| 21:47:06 | jaypipes | cfriesen_: bad time to bring that question :) | |
| 21:47:14 | jaypipes | cfriesen_: how about discuss tomorrow? | |
| 21:48:36 | dansmith | edleafe: _refresh_hash_ring() is a good place to start | |
| 21:48:43 | cfriesen_ | jaypipes: sure | |
| 21:49:00 | dansmith | after you get the hash ring you could probably spawn your thread to go examine the instances on what the hash ring says are your nodes | |
| 21:49:48 | edleafe | dansmith: looking at that now... | |
| 21:50:05 | dansmith | edleafe: obviously I wasn't thinking about this possibility either, as I'm used to the world before this was here and there was pretty much only one ironic compute ever | |
| 21:51:23 | mriedem | it's curious that this blows up during init_host in the driver, but not when driver.get_available_nodes is called | |