Earlier  
Posted Nick Remark
#openstack-nova - 2020-11-04
14:03:43 openstackgerrit Balazs Gibizer proposed openstack/nova master: Reproduce bug 1896463 in func env https://review.opendev.org/754100
14:05:10 openstackgerrit Balazs Gibizer proposed openstack/nova master: Set instance host and drop migration under lock https://review.opendev.org/754815
14:07:17 openstack bug 1896463 in OpenStack Compute (nova) rocky "evacuation failed: Port update failed : Unable to correlate PCI slot " [Low,In progress] https://launchpad.net/bugs/1896463
14:07:17 openstackgerrit Balazs Gibizer proposed openstack/nova master: Reproduce bug 1896463 in func env https://review.opendev.org/754100
14:08:38 openstackgerrit Balazs Gibizer proposed openstack/nova master: Set instance host and drop migration under lock https://review.opendev.org/754815
14:51:12 openstackgerrit Balazs Gibizer proposed openstack/nova master: Add upgrade check about old computes https://review.opendev.org/760520
15:02:45 openstackgerrit Takashi Natsume proposed openstack/nova stable/victoria: Fix a hacking test https://review.opendev.org/758112
15:05:44 openstackgerrit Balazs Gibizer proposed openstack/nova stable/victoria: [doc]: Fix glance image_metadata link https://review.opendev.org/761423
15:07:10 openstackgerrit Balazs Gibizer proposed openstack/nova stable/victoria: Use cell targeted context to query BDMs for metadata https://review.opendev.org/761424
15:29:19 openstackgerrit Balazs Gibizer proposed openstack/nova master: Bump the lowest eventlet version to 0.26.1 https://review.opendev.org/761427
15:55:10 openstackgerrit Balazs Gibizer proposed openstack/nova-specs master: [trivial]: replace NUMNA with NUMA https://review.opendev.org/761436
16:19:43 openstackgerrit Merged openstack/nova-specs master: [trivial]: replace NUMNA with NUMA https://review.opendev.org/761436
16:26:56 openstack Launchpad bug 1902276 in OpenStack Compute (nova) "libvirtd going into a tight loop causing instances to not transition to ACTIVE" [Undecided,New]
16:26:56 gibi somebody with connection to the libvirt maintainer should look at this nova bug https://bugs.launchpad.net/nova/+bug/1902276
16:28:04 melwitt kashyap: ^
16:28:26 kashyap melwitt: Yeah, familar with it, as I worked with the reporter here the other day
16:28:39 melwitt ah ok, cool
16:28:43 kashyap melwitt: I asked one of the libvirt devs on Friday, but my timing wasn't right
16:28:52 gibi kashyap: thanks!
16:28:57 kashyap I'll check again
16:29:18 kashyap gibi: It looks fishy, as it's not 100% reproducible ... as the reporter says "a few minutes later things go back to normal"
16:29:26 kashyap But we don't know what changed :-(
16:31:47 gibi kashyap: the nova image download take ~ 300 seconds for the VM that then triggers the loop in libvirtd so it might be that the hypervisor host has high load
16:31:57 kashyap gibi: Yeah, just reading your report :)
16:31:58 gibi but I was not able to confirm it from the logs
16:33:13 kashyap I see. Hypervisor load sounds plausible - as we've hit load-related (CI) issues libvirt driver before. But still let me check w/ Dan or someone from upstream libvirt
16:34:25 gibi kashyap: thanks for taking this up with the libvirt maintainers
16:35:34 kashyap gibi: Just posted on #virt, OFTC network.
16:36:18 kashyap gibi: Is this blocking patch merges?
16:39:05 melwitt I've been struggling for a couple of days trying to get an approved patch through the gate, but I'm not sure whether that particular bug is involved. I would need to re-look at the logs to verify
16:39:25 kashyap (I've reposted the looping libvirtd log bits as a plain text, as the pastebins expire)
16:40:58 kashyap melwitt: Noted; Michael, the reporter, was saying on last Friday that it's "intermittent", which makes it a bit more difficult to debug
16:41:56 melwitt yeah, that's been the theme of all of the gate bugs I'm aware of. intermittent and thus hard to troubleshoot :(
16:42:06 melwitt *current gate bugs I'm aware of
16:43:01 kashyap Yeah, matches my past experience
16:43:53 kashyap melwitt: In the same vein as how Twitter (I'm not on it) seems to label Trump's tweets as misleading, wonder we should adapt that text for these intermittent bugs :D
16:44:26 kashyap - Some or all of the content shared in this Tweet is disputed and might be misleading about an election or other civic process.
16:44:29 kashyap + Some or all of the content shared in this bug is disputed and might be misleading due to intermittent failures.
16:46:02 kashyap gibi: melwitt: More seriously, can I "subscribe" (Cc) someone else to a LaunchPad, right?
16:46:21 kashyap IIRC, yes. /me tries
16:46:26 melwitt I think you can
16:48:11 kashyap melwitt: I can't :-( I wanted to Cc Michal from libvirt but it says "No items matched <email ID>"
16:48:44 melwitt do you know his launchpad id?
16:48:55 kashyap melwitt: Oh, having a Launchpad ID is mandatory?
16:49:19 melwitt it might be, that's the only way I've seen subscribing
16:50:44 kashyap Ah, noted. I don't think he has one - searching doesn't show up anything.
16:51:41 kashyap melwitt: I pointed to him on IRC; he's taking a look
16:53:00 melwitt thanks!
16:55:23 kashyap melwitt: gibi: That's quick -- Michal (Privoznik) says it looks like a genuine bug. I'll update the bug once we get more details
16:56:35 melwitt sounds great, thank you kashyap
17:06:24 kashyap melwitt: So, the libvirt version in the logs above is 5.4.0; but havne't we switche dalready to libvirt-6.0.0?
17:06:39 kashyap lyarwood: --^ (By "we", I mean upstream CI)
17:08:11 kashyap So Michal says, there were improvements in libvirt-6.1.0 release on this area of event loops.
17:13:52 stephenfin lyarwood: Could you cast an eye over https://review.opendev.org/#/c/631053/ this evening, please?
17:14:04 stephenfin It's been around for quite a while :-D
17:18:01 mloza hello, is it possible to update the video model to vmga of an existing instance?
17:19:11 mloza if i edit /etc/libvirt/qemu/instance-, it reverts to default when the instance is hard rebooted
17:20:04 stephenfin mloza: Outside of rebuilding to a new image, no. We don't support setting it via the flavor so resize isn't an option
17:21:47 stephenfin mloza: You'll have to modify the DB manually if you want to avoid the rebuild
17:22:18 mloza can you tell me which table do I need modify
17:22:23 mloza to modif*
17:24:03 stephenfin iirc, we persist image metadata properties for an image in the instance_system_metadata table
17:25:16 stephenfin in case it wasn't obvious, back up the DB first and note that any support guarantees are gone out the window if you modify the DB manually
17:37:27 sean-k-mooney i think the table name is system_metadata not instance_system metadata but yes we do
17:37:34 sean-k-mooney with an img_ prefix
17:38:04 sean-k-mooney os if it was hw_video_model it woudl be img_hw_video_model in the db
17:38:48 sean-k-mooney mloza: were you asking about this on the mailing list too? we basically said the same in our replies
17:49:45 bauzas gibi: stephenfin: fwiw, you accepted a breaking RPC change with https://review.opendev.org/#/c/715326/29/nova/compute/manager.py@3327 by not accepting a nullable accels argument
17:50:06 bauzas sean-k-mooney: ^
17:52:41 bauzas if a compute client is sending a 5.0 cast to a compute service, then there won't have a accels argument, so the manager will return an exception
17:56:26 dansmith bauzas: good catch
17:56:39 dansmith bauzas: not too late to fix that
17:57:44 bauzas dansmith: I wonder whether we should fix it by the compute v5 proxy I write or having another change we could backport to victoria ?
17:58:00 dansmith bauzas: another change that we backport
17:58:09 bauzas ack, doing it then
17:58:19 dansmith before people try to upgrade to victoria
17:58:52 dansmith technically, this shouldn't be a problem if people upgrade their controllers first, but if they don't, they'll get an explosion that won't be easy to decipher
17:58:59 bauzas yeah
17:59:01 dansmith well, no, actuall,y
17:59:16 dansmith it would blow up for anyone with an old compute if the version is pinned
17:59:21 bauzas given most of the operators upgrade first their conductors, it shouldn't be a problem
17:59:33 bauzas but in case they pin it, yes
17:59:44 dansmith no, because the client was done properly, it will break
17:59:52 dansmith even if they don't pin, assuming they use =auto
17:59:56 bauzas if you pin the API, right?
18:00:07 bauzas why then for auto ?
18:00:18 dansmith the default is =auto, which will select the lowest version supported by all computes,
18:00:24 bauzas ahah
18:00:32 dansmith so if you have one old compute, api, conductor, etc will all choose the older and will not send that argument,
18:00:33 bauzas I see
18:00:35 dansmith and thus it'll explode
18:00:49 bauzas TIL
18:00:54 bauzas about how auto works
18:00:59 dansmith so s/good catch/great catch/ :)
18:01:21 dansmith easy backport to fix it though, luckily, before people start rolling to V
18:01:28 bauzas yup, writing it now
18:01:36 bauzas git stash first tho :)
18:01:46 dansmith heh yeah I bet :)
18:04:47 bauzas dansmith: do we have some documentation about pin=auto ? maybe on your blog ?
18:04:53 dansmith lol

Earlier   Later