| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-10-07 | |||
| 16:36:41 | sean-k-mooney | cool | |
| 16:37:12 | efried | stephenfin, mriedem_afk: Can/should we touch the admin guide as well? (SEV caveat) | |
| 16:37:40 | sean-k-mooney | im wondering if we should rename reserved_huge_pages to reserved_mem_pages give it can be used to reserve mempages of any size not just hugepages | |
| 16:37:54 | stephenfin | fine by me | |
| 16:38:18 | efried | stephenfin: yeah, there's already a section on "Impermanent Limitations" | |
| 16:38:46 | efried | https://docs.openstack.org/nova/latest/admin/configuration/hypervisor-kvm.html#impermanent-limitations | |
| 16:38:52 | sean-k-mooney | stephenfin: efried did you see my comment on what actully would be broken in the reno patch | |
| 16:38:59 | efried | sean-k-mooney: yeah, he fixed that part | |
| 16:39:05 | efried | would like you to reviow | |
| 16:39:07 | efried | review | |
| 16:39:15 | sean-k-mooney | cool im finding it now | |
| 16:39:16 | efried | but also want to get the admin doc updated in the same patch. | |
| 16:39:24 | efried | sean-k-mooney: https://review.opendev.org/686414 | |
| 16:40:08 | sean-k-mooney | thanks i just found it | |
| 16:40:49 | eandersson | dking_desktop, I'll let you know if I find anything | |
| 16:41:02 | sean-k-mooney | am the only comment i would make is ``hw_scsi_model=virtio-scsi`` and ``hw_disk_bus=scsi`` or | |
| 16:41:04 | sean-k-mooney | ``hw_cdrom_bus=scsi`` | |
| 16:41:15 | eandersson | efried, you wouldn't happen to know if there has been any recent fixes to instances getting stuck in building / scheduling? | |
| 16:41:27 | sean-k-mooney | should be ``hw_scsi_model=virtio-scsi`` and (``hw_disk_bus=scsi`` or ``hw_cdrom_bus=scsi``) | |
| 16:41:32 | efried | eandersson: depends what you mean by recent, but there have been, yes, stand by... | |
| 16:41:38 | eandersson | :D | |
| 16:41:42 | stephenfin | sean-k-mooney: how about 'and either...'? | |
| 16:41:49 | stephenfin | X and either Y or Z | |
| 16:41:56 | sean-k-mooney | sure i just want to make that clear | |
| 16:41:56 | stephenfin | I want to use English :D | |
| 16:41:59 | stephenfin | yup, agreed | |
| 16:42:00 | eandersson | We are seeing frequent issues with instances stuck in building... with Rocky | |
| 16:42:13 | sean-k-mooney | e.g. hw_cdrom_bugs=scsi wont break it by itself | |
| 16:42:17 | efried | eandersson: oh, I thought it was merged: https://review.opendev.org/#/c/666857/ | |
| 16:42:28 | efried | eandersson: does that ^ match your bug/symptom? | |
| 16:42:37 | openstack | Launchpad bug 1833581 in OpenStack Compute (nova) "instance stuck in BUILD state if nova-compute is restarted" [Low,In progress] - Assigned to Balazs Gibizer (balazs-gibizer) | |
| 16:42:37 | efried | eandersson: https://bugs.launchpad.net/nova/+bug/1833581 | |
| 16:42:42 | eandersson | Yea that sounds like it | |
| 16:43:06 | eandersson | Took you 30s :D | |
| 16:44:07 | efried | eandersson: you could hit the "affects me" button. Not sure if that sends gibi an email or not... | |
| 16:44:48 | efried | cdent: skimming the comments, it looks like maybe we were waiting for you to take a look at that fix --^ | |
| 16:45:14 | openstackgerrit | Stephen Finucane proposed openstack/nova stable/train: docs: Highlight the current broken state of SEV https://review.opendev.org/686414 | |
| 16:45:19 | cdent | hello, what? which? | |
| 16:45:55 | efried | cdent: https://review.opendev.org/#/c/666857/ | |
| 16:46:13 | efried | I didn't look at why, just saw a comment from mriedem_afk to that effect | |
| 16:46:42 | cdent | k, thanks, reading | |
| 16:46:58 | efried | "Let's wait to get some more feedback from cdent on this for the vcenter case" | |
| 16:47:10 | efried | PS11 9/18 | |
| 16:49:56 | efried | 0_0 | |
| 16:50:35 | mriedem | https://data.whicdn.com/images/98587406/superthumb.jpg?t=1390970016 | |
| 16:51:19 | mriedem | gonna get lunch and check back in a bit | |
| 16:51:51 | cdent | efried: left a +1 | |
| 16:51:58 | efried | thanks cdent | |
| 16:55:29 | eandersson | That bug sure looks like the issue. The only thing that concerns me is that we saw this hours after compute restarts, but might just have been delayed reports. | |
| 17:06:24 | efried | eandersson: I was +2 at some point in the past on that fix, and mriedem was just waiting for cdent to ack (I think), so we should be able to merge it fairly soon I would think. Will have to wait until next week to start merging backports, though. | |
| 17:07:02 | eandersson | Sounds good - we do large rolling deployments of computes | |
| 17:07:38 | eandersson | takes forever | |
| 17:08:15 | eandersson | So knowing that it's safe to do so helps a lot. Otherwise we need to disable the apis | |
| 17:25:34 | donnyd | sean-k-mooney: Yea that makes sense to me | |
| 17:42:00 | efried | stephenfin: I really appreciate the cleanup of oldy moldy change sets, but would you mind leaving a one-line reason when you abandon please? | |
| 17:47:10 | mriedem | efried: i'm +2 on gibi's interrupted build fix now https://review.opendev.org/#/c/666857/ | |
| 17:47:28 | efried | mriedem: cool, thanks, I'll get back to that in the next few minutes. | |
| 17:57:51 | efried | eandersson: How many instances are you seeing stuck, max? | |
| 17:58:07 | efried | just curious, because it seems like it would be pretty hard to have more than a small number. | |
| 17:58:50 | eandersson | 10+ | |
| 17:59:29 | mriedem | ricolin: your fix looks fine but i want to check what tests could be updated to make sure that's not regressed again | |
| 17:59:40 | eandersson | We had a RabbitMQ outage that could have caused some of them as well. | |
| 17:59:58 | mriedem | eandersson: i know you also reported issues with rescheduling failures in conductor recently | |
| 18:00:00 | efried | eandersson: all from the same compute host? | |
| 18:00:07 | mriedem | and we don't put the instances into ERROR state in all cases there | |
| 18:00:09 | eandersson | different hosts | |
| 18:00:42 | efried | okay cool. Trying to assess the probability of the "slow startup" thing noted by cdent and gibi. Seems like it would be pretty low. | |
| 18:00:48 | ricolin | mriedem, thanks | |
| 18:01:34 | eandersson | I was trying to figure out what happened, but couldn't really find anything. | |
| 18:01:59 | eandersson | The VM was in placement, but instance info was not yet updated. | |
| 18:02:23 | eandersson | btw also a lot of VMs stuck in deleting | |
| 18:02:26 | mriedem | ricolin: tbc those jobs were going to fail anyway because of http://status.openstack.org/elastic-recheck/#1844929 but definitely need to get that unboundlocalerror fixed | |
| 18:02:31 | eandersson | but looks like a different issue | |
| 18:02:44 | eandersson | because when VMs are stuck deleting, restarting nova-compute works everytime. | |
| 18:03:47 | mriedem | because the restart code checks for those | |
| 18:04:10 | eandersson | Been unable to figure out why they got stuck in the first place. | |
| 18:04:36 | eandersson | They spiked the hours after a RabbitMQ cluster failed, but seen it before that as well. | |
| 18:04:39 | mriedem | something blow up between the time the api marks the instance as deleting and then rpc casts to the compute? | |
| 18:04:58 | eandersson | Not that I could find. | |
| 18:06:10 | eandersson | I see RabbitMQ connection errors hours, days before that | |
| 18:06:32 | eandersson | but I feel like it shouldn't be the cause, unless RabbitMQ connections just hang dead | |
| 18:06:38 | eandersson | Maybe due to something like https://github.com/openstack/oslo.messaging/commit/26fccea84371a3690a4d7fe970f8d6c5ba9c154b | |
| 18:07:50 | eandersson | I'll try to dig a bit deeper next time it happens. | |
| 18:09:03 | efried | eandersson, mriedem, gibi: +A that fix, but I think we're going to miss the boat on VirtDriverNotReady more frequently than anticipated in ironic cases. | |
| 18:12:41 | openstackgerrit | Rico Lin proposed openstack/nova master: [Gate fix] Avoid use cell_uuid before assignment https://review.opendev.org/686996 | |
| 18:13:14 | ricolin | mriedem, just add bug link to https://review.opendev.org/#/c/686996/ . | |
| 18:14:26 | mriedem | ricolin: yeah i'm working on fixing the test that was masking the regression though | |
| 18:14:35 | ricolin | mriedem, feel free to update that patch to update the test/ whatever needed | |
| 18:15:04 | mriedem | efried: just depends on order of operations - devstack ironic jobs start ironic-api after nova-compute | |
| 18:15:08 | mriedem | b/c of the devstack plugin structure | |
| 18:15:15 | ricolin | it's pretty late for me, so will check on this when I awake tomorrow | |
| 18:15:15 | efried | yes | |
| 18:15:28 | mriedem | but in a real world deployment doohicky ironic might be up before nova | |
| 18:15:34 | mriedem | ricolin: sure np | |
| 18:15:36 | efried | *might* be | |
| 18:16:05 | efried | I guess we'll see. | |
| 18:27:32 | mriedem | melwitt: dansmith: am i overcomplicating this, or should the try/except here https://github.com/openstack/nova/blob/2c6542948f1f57c1a8a8faec1ba61800eaddff2a/nova/context.py#L442 be within the while loop? | |
| 18:27:48 | mriedem | e.g. the first result is a celltimeout but the 2nd yields results | |
| 18:28:41 | dansmith | mriedem: it won't be in the queue until it has timed out so I don't think it matters | |
| 18:29:01 | mriedem | ok so maybe just something i'm seeing in a unit test that is a problem | |