| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-11-04 | |||
| 14:51:12 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add upgrade check about old computes https://review.opendev.org/760520 | |
| 15:02:45 | openstackgerrit | Takashi Natsume proposed openstack/nova stable/victoria: Fix a hacking test https://review.opendev.org/758112 | |
| 15:05:44 | openstackgerrit | Balazs Gibizer proposed openstack/nova stable/victoria: [doc]: Fix glance image_metadata link https://review.opendev.org/761423 | |
| 15:07:10 | openstackgerrit | Balazs Gibizer proposed openstack/nova stable/victoria: Use cell targeted context to query BDMs for metadata https://review.opendev.org/761424 | |
| 15:29:19 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Bump the lowest eventlet version to 0.26.1 https://review.opendev.org/761427 | |
| 15:55:10 | openstackgerrit | Balazs Gibizer proposed openstack/nova-specs master: [trivial]: replace NUMNA with NUMA https://review.opendev.org/761436 | |
| 16:19:43 | openstackgerrit | Merged openstack/nova-specs master: [trivial]: replace NUMNA with NUMA https://review.opendev.org/761436 | |
| 16:26:56 | openstack | Launchpad bug 1902276 in OpenStack Compute (nova) "libvirtd going into a tight loop causing instances to not transition to ACTIVE" [Undecided,New] | |
| 16:26:56 | gibi | somebody with connection to the libvirt maintainer should look at this nova bug https://bugs.launchpad.net/nova/+bug/1902276 | |
| 16:28:04 | melwitt | kashyap: ^ | |
| 16:28:26 | kashyap | melwitt: Yeah, familar with it, as I worked with the reporter here the other day | |
| 16:28:39 | melwitt | ah ok, cool | |
| 16:28:43 | kashyap | melwitt: I asked one of the libvirt devs on Friday, but my timing wasn't right | |
| 16:28:52 | gibi | kashyap: thanks! | |
| 16:28:57 | kashyap | I'll check again | |
| 16:29:18 | kashyap | gibi: It looks fishy, as it's not 100% reproducible ... as the reporter says "a few minutes later things go back to normal" | |
| 16:29:26 | kashyap | But we don't know what changed :-( | |
| 16:31:47 | gibi | kashyap: the nova image download take ~ 300 seconds for the VM that then triggers the loop in libvirtd so it might be that the hypervisor host has high load | |
| 16:31:57 | kashyap | gibi: Yeah, just reading your report :) | |
| 16:31:58 | gibi | but I was not able to confirm it from the logs | |
| 16:33:13 | kashyap | I see. Hypervisor load sounds plausible - as we've hit load-related (CI) issues libvirt driver before. But still let me check w/ Dan or someone from upstream libvirt | |
| 16:34:25 | gibi | kashyap: thanks for taking this up with the libvirt maintainers | |
| 16:35:34 | kashyap | gibi: Just posted on #virt, OFTC network. | |
| 16:36:18 | kashyap | gibi: Is this blocking patch merges? | |
| 16:39:05 | melwitt | I've been struggling for a couple of days trying to get an approved patch through the gate, but I'm not sure whether that particular bug is involved. I would need to re-look at the logs to verify | |
| 16:39:25 | kashyap | (I've reposted the looping libvirtd log bits as a plain text, as the pastebins expire) | |
| 16:40:58 | kashyap | melwitt: Noted; Michael, the reporter, was saying on last Friday that it's "intermittent", which makes it a bit more difficult to debug | |
| 16:41:56 | melwitt | yeah, that's been the theme of all of the gate bugs I'm aware of. intermittent and thus hard to troubleshoot :( | |
| 16:42:06 | melwitt | *current gate bugs I'm aware of | |
| 16:43:01 | kashyap | Yeah, matches my past experience | |
| 16:43:53 | kashyap | melwitt: In the same vein as how Twitter (I'm not on it) seems to label Trump's tweets as misleading, wonder we should adapt that text for these intermittent bugs :D | |
| 16:44:26 | kashyap | - Some or all of the content shared in this Tweet is disputed and might be misleading about an election or other civic process. | |
| 16:44:29 | kashyap | + Some or all of the content shared in this bug is disputed and might be misleading due to intermittent failures. | |
| 16:46:02 | kashyap | gibi: melwitt: More seriously, can I "subscribe" (Cc) someone else to a LaunchPad, right? | |
| 16:46:21 | kashyap | IIRC, yes. /me tries | |
| 16:46:26 | melwitt | I think you can | |
| 16:48:11 | kashyap | melwitt: I can't :-( I wanted to Cc Michal from libvirt but it says "No items matched <email ID>" | |
| 16:48:44 | melwitt | do you know his launchpad id? | |
| 16:48:55 | kashyap | melwitt: Oh, having a Launchpad ID is mandatory? | |
| 16:49:19 | melwitt | it might be, that's the only way I've seen subscribing | |
| 16:50:44 | kashyap | Ah, noted. I don't think he has one - searching doesn't show up anything. | |
| 16:51:41 | kashyap | melwitt: I pointed to him on IRC; he's taking a look | |
| 16:53:00 | melwitt | thanks! | |
| 16:55:23 | kashyap | melwitt: gibi: That's quick -- Michal (Privoznik) says it looks like a genuine bug. I'll update the bug once we get more details | |
| 16:56:35 | melwitt | sounds great, thank you kashyap | |
| 17:06:24 | kashyap | melwitt: So, the libvirt version in the logs above is 5.4.0; but havne't we switche dalready to libvirt-6.0.0? | |
| 17:06:39 | kashyap | lyarwood: --^ (By "we", I mean upstream CI) | |
| 17:08:11 | kashyap | So Michal says, there were improvements in libvirt-6.1.0 release on this area of event loops. | |
| 17:13:52 | stephenfin | lyarwood: Could you cast an eye over https://review.opendev.org/#/c/631053/ this evening, please? | |
| 17:14:04 | stephenfin | It's been around for quite a while :-D | |
| 17:18:01 | mloza | hello, is it possible to update the video model to vmga of an existing instance? | |
| 17:19:11 | mloza | if i edit /etc/libvirt/qemu/instance-, it reverts to default when the instance is hard rebooted | |
| 17:20:04 | stephenfin | mloza: Outside of rebuilding to a new image, no. We don't support setting it via the flavor so resize isn't an option | |
| 17:21:47 | stephenfin | mloza: You'll have to modify the DB manually if you want to avoid the rebuild | |
| 17:22:18 | mloza | can you tell me which table do I need modify | |
| 17:22:23 | mloza | to modif* | |
| 17:24:03 | stephenfin | iirc, we persist image metadata properties for an image in the instance_system_metadata table | |
| 17:25:16 | stephenfin | in case it wasn't obvious, back up the DB first and note that any support guarantees are gone out the window if you modify the DB manually | |
| 17:37:27 | sean-k-mooney | i think the table name is system_metadata not instance_system metadata but yes we do | |
| 17:37:34 | sean-k-mooney | with an img_ prefix | |
| 17:38:04 | sean-k-mooney | os if it was hw_video_model it woudl be img_hw_video_model in the db | |
| 17:38:48 | sean-k-mooney | mloza: were you asking about this on the mailing list too? we basically said the same in our replies | |
| 17:49:45 | bauzas | gibi: stephenfin: fwiw, you accepted a breaking RPC change with https://review.opendev.org/#/c/715326/29/nova/compute/manager.py@3327 by not accepting a nullable accels argument | |
| 17:50:06 | bauzas | sean-k-mooney: ^ | |
| 17:52:41 | bauzas | if a compute client is sending a 5.0 cast to a compute service, then there won't have a accels argument, so the manager will return an exception | |
| 17:56:26 | dansmith | bauzas: good catch | |
| 17:56:39 | dansmith | bauzas: not too late to fix that | |
| 17:57:44 | bauzas | dansmith: I wonder whether we should fix it by the compute v5 proxy I write or having another change we could backport to victoria ? | |
| 17:58:00 | dansmith | bauzas: another change that we backport | |
| 17:58:09 | bauzas | ack, doing it then | |
| 17:58:19 | dansmith | before people try to upgrade to victoria | |
| 17:58:52 | dansmith | technically, this shouldn't be a problem if people upgrade their controllers first, but if they don't, they'll get an explosion that won't be easy to decipher | |
| 17:58:59 | bauzas | yeah | |
| 17:59:01 | dansmith | well, no, actuall,y | |
| 17:59:16 | dansmith | it would blow up for anyone with an old compute if the version is pinned | |
| 17:59:21 | bauzas | given most of the operators upgrade first their conductors, it shouldn't be a problem | |
| 17:59:33 | bauzas | but in case they pin it, yes | |
| 17:59:44 | dansmith | no, because the client was done properly, it will break | |
| 17:59:52 | dansmith | even if they don't pin, assuming they use =auto | |
| 17:59:56 | bauzas | if you pin the API, right? | |
| 18:00:07 | bauzas | why then for auto ? | |
| 18:00:18 | dansmith | the default is =auto, which will select the lowest version supported by all computes, | |
| 18:00:24 | bauzas | ahah | |
| 18:00:32 | dansmith | so if you have one old compute, api, conductor, etc will all choose the older and will not send that argument, | |
| 18:00:33 | bauzas | I see | |
| 18:00:35 | dansmith | and thus it'll explode | |
| 18:00:49 | bauzas | TIL | |
| 18:00:54 | bauzas | about how auto works | |
| 18:00:59 | dansmith | so s/good catch/great catch/ :) | |
| 18:01:21 | dansmith | easy backport to fix it though, luckily, before people start rolling to V | |
| 18:01:28 | bauzas | yup, writing it now | |
| 18:01:36 | bauzas | git stash first tho :) | |
| 18:01:46 | dansmith | heh yeah I bet :) | |
| 18:04:47 | bauzas | dansmith: do we have some documentation about pin=auto ? maybe on your blog ? | |
| 18:04:53 | dansmith | lol | |
| 18:05:07 | dansmith | I guess I didn't think it was really something people often confused | |
| 18:05:25 | dansmith | I'd expect the config doc to be accurate, but let me look | |
| 18:05:53 | bauzas | nevermind, found it ;) https://docs.openstack.org/nova/latest/user/upgrade.html | |
| 18:06:51 | dansmith | the config doc says "don't worry, we'll handle it" without any detail | |
| 18:07:07 | bauzas | yeah, looking at the config option text | |