Earlier  
Posted Nick Remark
#openstack-nova - 2019-11-01
15:03:19 mriedem ok ok 5 days
15:03:20 dansmith my notifications thing has been going all week.. maybe since last week
15:03:49 mriedem for big series like this it's always best to work from the bottom up where everything is noop until "turned on" in the API at the end if possible
15:03:52 dansmith yeah, since monday
15:04:01 mriedem so there is less risk in merging the bottom noop stuff
15:07:45 Sundar mriedem: The code flow starts from the API, such as getting device profile info based on the extra specs. If the API stuff came later, wouldn't the patch sequence be out of order? We can only do UT for the earlier patches then.
15:08:02 cdent mriedem: I assume you saw my response on the gate+placement email? Sorry I've not been able to dig more deeply. I've been engaged elsewhere...To such an extent that I may come off the mailing list pretty soon to maintain sanity.
15:09:01 efried Sundar: in this case the "master switch" could conceivably be the part that processes the device profile out of the flavor.
15:09:28 mriedem cdent: yup, and replied, and thanks
15:09:45 mriedem cdent: i think the likeliest culprit is noise neighbors on oversubscribed nodes as you said
15:17:28 cdent mriedem: the gist of what I can find on that innodb error (now that I've managed to actually read your responses) is "you disk i/o is way constrained"
15:19:32 mriedem cdent: heh, yeah looking at this dstat graph, i/o is just 0 from 16:52 to 16:55
15:19:42 cdent oh dear
15:20:32 cdent there should be some kind of "how am I supposed to work in these conditions" meme we can interject here
15:21:08 efried https://makeameme.org/meme/i-cant-work-2sl5sn
15:21:11 mriedem https://toddroblog.files.wordpress.com/2017/02/drunk-hobo-flips-double-bird-1.jpg?w=809 ?
15:21:22 mriedem oh blurred out?!
15:23:09 mriedem i wish i knew about https://lamada.eu/dstat-graph/# years ago
15:26:44 cdent yeah, why don't things like that get passed around more easily?
15:27:16 mriedem probably because i was too shy to say "i'm dumb, how can i see this dstat output in a graph?"
15:27:44 mriedem i assume all of my co-workers to write tools to generate graphs on the fly for everything they do
15:29:50 cdent https://www.youtube.com/watch?v=MKIaS0lh-uo
15:30:48 mriedem heh, it's rare to get a repo man reference in here
15:33:25 cdent what provider is that 0 I/O on?
15:33:40 mriedem inap
15:35:29 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: link to nova code review guide from dev policies https://review.opendev.org/692569
15:54:17 efried I already rechecked a few dansmith
15:54:27 openstackgerrit Matt Riedemann proposed openstack/nova master: Document CD mentality policy for nova contributors https://review.opendev.org/692572
15:54:34 mriedem dansmith: i took a crack at documenting that ^
15:54:37 dansmith efried: in that case, https://www.youtube.com/watch?v=0u8teXR8VE4
15:56:23 mriedem dansmith: thoughts on the no graceful shutdown thing i mentioned here? http://lists.openstack.org/pipermail/openstack-discuss/2019-October/010495.html - wondering if we should go back to like a 10 second shutdown timeout in case that is screwing up the guest during things that transfer the root disk or snapshot it (shelve/unshelve) leading to ssh failures later
15:57:09 mriedem i had made that change in devstack b/c of a real bug where we see the 60 second shutdown timeout in nova not doing anything, potentially leading to overall job timeouts if we're waiting a full minute to power off guests during a tempest run
15:57:52 dansmith mriedem: I doubt that cirros is responding to the graceful shutdown request at all
15:58:01 dansmith mriedem: so I'm not sure it matters
15:58:18 dansmith if you have a devstack up you should be able to test
15:58:46 dansmith even still, unless it's writing to the disk when the destroy happens, it wouldn't likely do any damage (and shouldn't really anyway)
15:59:12 dansmith don't we get the console of the guest we fail to ssh into if we do? unless it's sitting at a panic or some unhappy state, I'd not be concerned
15:59:28 mriedem yes we do
16:00:16 sean-k-mooney mriedem: are you suspecting disk curruption?
16:00:19 mriedem there are several guest ssh failure bugs in the gate right now, many different reasons (dhcp lease fails, out of space on the guest disk [growroot fails], kernel panic)
16:00:28 sean-k-mooney due to an agressive shutdown?
16:00:29 mriedem i'm grasping at straws
16:01:08 mriedem in one case that i caught in the gate where the 60 second shutdown timed out, i got the guest console log http://paste.openstack.org/show/752002/
16:01:25 mriedem cdent wondered (in the bug) if the metadata api retry loop in the guest was making it ignore the shutdown request
16:01:32 openstack Launchpad bug 1829896 in OpenStack Compute (nova) "libvirt: "Instance failed to shutdown in 60 seconds." in the gate" [Undecided,New]
16:01:32 mriedem https://bugs.launchpad.net/nova/+bug/1829896
16:02:05 sean-k-mooney well in that case it got a dhcp lease but then could not hit the metadata service
16:02:10 mriedem i'm also wondering if there are ideas on figuring out what is eating up all of the disk in these guests when growroot fails
16:02:17 sean-k-mooney so that looks like a neutron issue with the metadata proxy
16:03:05 sean-k-mooney if glean/cloud init is waithing for the metadata then
16:03:11 sean-k-mooney the ssh service may not have started
16:03:45 openstackgerrit Lee Yarwood proposed openstack/nova stable/queens: Add 'path' query parameter to console access url https://review.opendev.org/692573
16:06:30 lyarwood mriedem: yup, working on the other change now thanks
16:07:42 sean-k-mooney mriedem: i do know we have had issue with hot plug if the early init in the guest takes too long
16:08:05 mriedem lyarwood: i'm sort of inclined to say you guys (RH) should keep that one downstream
16:08:06 sean-k-mooney but we also see Starting acpid: OK
16:08:17 mriedem since i don't see people clamoring for using the bleeding edge novnc stuff back on queens
16:09:06 mriedem i'm not like -2, but those are my "feelings"
16:09:15 mriedem what you humans would call feelings anyway
16:09:53 lyarwood haha yeah no issues, thought I'd at least post it upstream first to see if anyone wanted it
16:10:02 sean-k-mooney ah ha proof mriedem is actully an irc bot !!!
16:10:54 sean-k-mooney mriedem: on other random straw is there seams to be very littly entropy in the guest
16:10:59 sean-k-mooney random: dd urandom read with 31 bits of entropy available
16:12:14 sean-k-mooney the topic of enabling the virtio random number generator by default has come up before. i wonder if it could be related to that.
16:14:09 mriedem i seem to remember a recentish patch of kashyap's related to this
16:15:21 mriedem i'm just thinking of this https://review.opendev.org/#/c/577385/
16:16:17 sean-k-mooney ya there is that one but we also talked a little about adding the random number gereate to the libvirt xml by defualt rather then needing to set the image property
16:16:18 mriedem we don't use that in our devstack guests though since we don't have hw_rng_model=virtio in the image
16:16:24 mriedem oh
16:16:37 mriedem well, i have to go to what you humans would call lunch now
16:28:35 openstackgerrit Lee Yarwood proposed openstack/nova stable/queens: Reduce scope of 'path' query parameter to noVNC consoles https://review.opendev.org/692581
17:19:32 dansmith cripes
17:19:42 dansmith cache notifications patch is gonna fail again
17:21:15 openstackgerrit John Garbutt proposed openstack/nova-specs master: Add Unified Limits Spec https://review.opendev.org/602201
17:34:41 mriedem dansmith: it's because of your sin
17:35:59 efried dansmith: I'm +A on mriedem's CD doc patch, but the predecessor (link shuffle) needs to be sent as well
17:36:07 efried https://review.opendev.org/#/c/692569/1
17:36:34 dansmith efried: oh sorry I missed that earlier I guess
17:36:46 dansmith I dun getified it
17:40:13 mriedem dansmith: i looked at a dstat from one of the 'timed out waiting for cell' failed jobs and the only thing around that time is a spike in paging and networking which doesn't tell me much
17:40:32 dansmith mriedem: yeah I looked pretty in depth at one dstat from a cell fail job
17:40:36 dansmith and it looked overly normal to me
17:40:53 dansmith mysql wasn't spiking that I could tell, it was the big memory user, but only like 400MB at the time
17:41:02 dansmith IO wasn't crazy, etc
17:41:30 mriedem yeah nothing obvious like the 0 i/ops in the other thing earlier today
17:42:06 dansmith I didn't catch the zero iops thing
17:42:08 mriedem and there is no rabbit involved b/c we're going straight to the db yeah?
17:42:14 dansmith you were assuming zero iops meant bad?
17:42:30 mriedem the 0 iops thing was related to another bug in the ML where we get long timeouts
17:42:42 mriedem which triggered the resize long_rpc_timeout discussion
17:43:08 mriedem and a very obvious message in the mysql error log
17:43:29 dansmith you mean zero iops for a long period of time or something I assume/
17:43:35 mriedem 3 minutes yeah
17:43:54 mriedem while placement was processing a POST /allocations request
17:43:57 dansmith okay, yeah, one or two samples with zero iops wouldn't be alarming to me, but minutes of it, for sure
17:46:49 openstack Launchpad bug 1825584 in OpenStack Compute (nova) "eventlet monkey-patching breaks AMQP heartbeat on uWSGI" [Undecided,Confirmed]
17:46:49 mriedem and https://bugs.launchpad.net/nova/+bug/1825584 shouldn't be related to the scheduler timing out waiting for a response from cell1 b/c there is no uwsgi involved there
17:46:53 mriedem just scheduler running with eventlet
18:30:42 openstackgerrit Merged openstack/nova stable/rocky: libvirt: Ignore volume exceptions during post_live_migration https://review.opendev.org/691283
18:36:07 openstackgerrit Matt Riedemann proposed openstack/nova stable/queens: libvirt: Ignore volume exceptions during post_live_migration https://review.opendev.org/691284

Earlier   Later