| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-02-15 | |||
| 19:06:02 | melwitt | mriedem: FYI, digging into the difference between the cleanup volumes vs ports bugs today | |
| 19:06:18 | melwitt | based on your comments in the patch | |
| 19:08:50 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/queens: DNM: Test LM with encrypted volumes https://review.openstack.org/545093 | |
| 19:09:38 | jroll | dansmith: ah yep, the subnode n-cpu considers itself down at this point, I believe http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/subnode-2/screen-n-cpu.txt.gz#_Feb_15_17_59_05_738069 | |
| 19:10:06 | jroll | ironic is unreachable for like 5 minutes | |
| 19:10:22 | jroll | or rather a full resource tracker run and then some | |
| 19:11:19 | openstackgerrit | Jackie Truong proposed openstack/python-novaclient master: Microversion 2.61 - Add trusted_image_certificates https://review.openstack.org/500396 | |
| 19:12:05 | mriedem | cfriesen: then what you have here https://review.openstack.org/#/c/525253/1/nova/conductor/manager.py doesn't help you | |
| 19:12:10 | mriedem | cfriesen: so i'm confused | |
| 19:12:42 | mriedem | cfriesen: is https://review.openstack.org/#/c/528385/ what you are looking for? | |
| 19:12:56 | dansmith | jroll: yeah, but with johnthetubaguy's reuse-compute-node patch I would think this wouldn't be a problem right? | |
| 19:12:58 | jroll | aaaand we have some problems deleting RPs: http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-n-cpu.txt.gz#_Feb_15_17_58_22_418875 | |
| 19:13:00 | jroll | (wtf) | |
| 19:13:22 | jroll | dansmith: I would think so too, just confirming there is likely a rebalance, so they could be disagreeing | |
| 19:13:32 | dansmith | yeah | |
| 19:13:50 | dansmith | jroll: ah, that 503 during deleting is weird | |
| 19:14:11 | jroll | dansmith: indeed | |
| 19:14:20 | dansmith | jroll: do you guys have to restart apache? | |
| 19:14:39 | TheJulia | dansmith: we do | |
| 19:14:44 | dansmith | okay | |
| 19:14:45 | dansmith | also | |
| 19:14:47 | TheJulia | we update the configuration to load a vhost | |
| 19:15:02 | TheJulia | we also shutdown services at 17:55 for the upgrade, nova would have remained running | |
| 19:15:09 | dansmith | if placement is crashing the same way as conductor, maybe placement is dead under apache, hence the 503? | |
| 19:15:12 | jroll | good lord keystone db migrations are spammy | |
| 19:15:13 | dansmith | until you restart? | |
| 19:15:20 | TheJulia | so 17:58 is when everything is down | |
| 19:15:31 | dansmith | thus nova never got to delete the RPs | |
| 19:15:55 | TheJulia | I think the restart is during the ironic upgrade, checking to see when it actually occured | |
| 19:16:09 | dansmith | right, but if placement started crashing during the upgrade of packages, | |
| 19:16:21 | dansmith | which is when nova would have noticed ironic went away and tried to delete RPs or something, | |
| 19:16:24 | jroll | OH | |
| 19:16:28 | jroll | keystone is upgrading there | |
| 19:16:35 | jroll | and so placement can't validate the token | |
| 19:16:36 | dansmith | and then you restart apache.. | |
| 19:16:38 | dansmith | ohh | |
| 19:16:43 | jroll | http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-placement-api.txt.gz#_Feb_15_17_58_22_463228 | |
| 19:17:02 | mriedem | melwitt: i'm likely also going to start writing a functional test for the case that we delete a build request for a bfv instance during local delete, because we aren't cleaning up volumes there either | |
| 19:17:08 | dansmith | and you get a bug, and you get a bug, and you get a bug... | |
| 19:17:21 | dansmith | jroll: good catch | |
| 19:17:28 | melwitt | mriedem: ack | |
| 19:17:50 | TheJulia | http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/grenade.sh.txt.gz#_2018-02-15_18_00_28_322 is when we restart apache | |
| 19:17:53 | dansmith | jroll: surely you're not intending to have keystone be upgradingin this task right? | |
| 19:17:59 | dansmith | s/task/job/ | |
| 19:18:30 | jroll | dansmith: I think we just haven't cared either way in the past | |
| 19:18:44 | jroll | keystone's one of those hand-wavy things to me, it's always there, don't care what version | |
| 19:18:46 | dansmith | jroll: do you care now? :P | |
| 19:18:49 | jroll | heh | |
| 19:21:37 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add admin guide doc on volume multiattach support https://review.openstack.org/544090 | |
| 19:21:48 | TheJulia | jroll: we could just skip restarting apache and fire up the local ironic-api, however then we're not forcing traffic through an older API endpoint which changes the scenario, although we would still have the rpc version pin | |
| 19:21:56 | mriedem | bauzas: since you helped review the multiattach series, can you check out ^ so we can backport that for queens RC2? | |
| 19:22:23 | jroll | TheJulia: eh, I'd rather not | |
| 19:22:28 | jroll | also these times aren't lining up :/ | |
| 19:22:32 | TheJulia | actually, other services will still likely need to restart things | |
| 19:22:44 | TheJulia | so we shouldn't try to avoid restarting apache | |
| 19:22:59 | jroll | oh it does line up, okay | |
| 19:23:16 | jroll | TheJulia: maybe we um, try to make apache not need 66 seconds to restart | |
| 19:23:18 | TheJulia | heh, 28 apache restarts in the grenade log | |
| 19:23:27 | jroll | jesus | |
| 19:23:31 | dansmith | wow | |
| 19:23:33 | dansmith | that's impressive | |
| 19:23:57 | mriedem | the edge people said openstack needed to be slimmed down | |
| 19:24:13 | TheJulia | yeah.... | |
| 19:24:15 | cfriesen | mriedem: I think you're right, I got messed up with which patch was fixing what. :) the fixes at https://review.openstack.org/#/c/528385 and https://review.openstack.org/#/c/340614/ look like they might do the trick. | |
| 19:24:31 | mriedem | cfriesen: cool | |
| 19:24:36 | mriedem | it is confusing | |
| 19:24:55 | mriedem | cfriesen: the point of creating the bdms in cell0 also was so that the local delete in the api can remove them | |
| 19:25:06 | mriedem | and thus avoid us having nova-compute, nova-api AND nova-conductor doing volume cleanup | |
| 19:25:49 | TheJulia | jroll: what makes you think 66 seconds? | |
| 19:26:27 | jroll | TheJulia: this line to the next one: http://logs.openstack.org/50/544750/8/check/ironic-grenade-dsvm-multinode-multitenant/5713fb8/logs/screen-keystone.txt.gz#_Feb_15_17_57_44_538878 | |
| 19:26:31 | jroll | it isn't 66 seconds to restart | |
| 19:26:48 | jroll | nor is it an apache thing | |
| 19:26:50 | jroll | but we shut down keystone for the entire keystone upgrade | |
| 19:26:58 | TheJulia | ahh, yeah | |
| 19:27:02 | jroll | which is silly | |
| 19:27:22 | TheJulia | Well.... are we sure it is not silly? | |
| 19:28:39 | jroll | I would think keystone should be able to do rolling upgrades very easily | |
| 19:28:47 | jroll | but they don't have the magic tag, so who knows | |
| 19:30:24 | dansmith | well, it's not really placement, it's nova-compute right? | |
| 19:30:39 | dansmith | we're probably not handling the 503 very gracefully in compute whilst trying to do things | |
| 19:33:30 | mriedem | is the 503 this? https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L93 | |
| 19:34:13 | dansmith | I woudn't think so, we're getting a response | |
| 19:34:46 | oomichi | hi, can someone take a look at https://review.openstack.org/#/c/532822 ? that is easy and completely the same patch is posted. so want to avoid the same thing | |
| 19:34:48 | dansmith | we just raise and kill the RT | |
| 19:37:45 | oomichi | thanks, mriedem | |
| 19:40:41 | mriedem | melwitt: i went ahead and created this for tracking the community wide goal and linked it to the other 4 bp's we've had in the past for the same thing https://blueprints.launchpad.net/nova/+spec/mox-removal | |
| 19:40:57 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: Account for host agg allocation ratio in placement https://review.openstack.org/544683 | |
| 19:41:00 | mriedem | and marked the other as done https://review.openstack.org/545100 | |
| 19:41:21 | melwitt | mriedem: great, thank you | |
| 19:41:23 | jaypipes | bauzas, cdent, edleafe, dansmith, mriedem, efried: ^^ round two on the aggregate allocation ratio one. | |
| 19:43:31 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: Account for host agg allocation ratio in placement https://review.openstack.org/544683 | |
| 19:44:09 | mriedem | melwitt: i also assume we'll want 2 blueprints for removing nova-network and cellsv1 | |
| 19:44:17 | mriedem | as those likely won't be a single change | |
| 19:44:27 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: mirror nova host aggregates to placement API https://review.openstack.org/545057 | |
| 19:44:34 | melwitt | mriedem: yeah, I don't think they would be a single change | |
| 19:44:44 | mriedem | i'm not sure on the order there though, probably nova-network first since you can't start nova-network w/o cells v1 enabled | |
| 19:45:01 | mriedem | and some people run cellsv1 with neutron | |
| 19:46:32 | openstackgerrit | Jay Pipes proposed openstack/nova-specs master: Account for host agg allocation ratio in placement https://review.openstack.org/544683 | |
| 19:50:00 | efried | jaypipes: Ya done yet? | |
| 19:58:07 | jaypipes | efried: yeah, sorry :) | |