Earlier  
Posted Nick Remark
#openstack-nova - 2019-11-01
15:08:02 cdent mriedem: I assume you saw my response on the gate+placement email? Sorry I've not been able to dig more deeply. I've been engaged elsewhere...To such an extent that I may come off the mailing list pretty soon to maintain sanity.
15:09:01 efried Sundar: in this case the "master switch" could conceivably be the part that processes the device profile out of the flavor.
15:09:28 mriedem cdent: yup, and replied, and thanks
15:09:45 mriedem cdent: i think the likeliest culprit is noise neighbors on oversubscribed nodes as you said
15:17:28 cdent mriedem: the gist of what I can find on that innodb error (now that I've managed to actually read your responses) is "you disk i/o is way constrained"
15:19:32 mriedem cdent: heh, yeah looking at this dstat graph, i/o is just 0 from 16:52 to 16:55
15:19:42 cdent oh dear
15:20:32 cdent there should be some kind of "how am I supposed to work in these conditions" meme we can interject here
15:21:08 efried https://makeameme.org/meme/i-cant-work-2sl5sn
15:21:11 mriedem https://toddroblog.files.wordpress.com/2017/02/drunk-hobo-flips-double-bird-1.jpg?w=809 ?
15:21:22 mriedem oh blurred out?!
15:23:09 mriedem i wish i knew about https://lamada.eu/dstat-graph/# years ago
15:26:44 cdent yeah, why don't things like that get passed around more easily?
15:27:16 mriedem probably because i was too shy to say "i'm dumb, how can i see this dstat output in a graph?"
15:27:44 mriedem i assume all of my co-workers to write tools to generate graphs on the fly for everything they do
15:29:50 cdent https://www.youtube.com/watch?v=MKIaS0lh-uo
15:30:48 mriedem heh, it's rare to get a repo man reference in here
15:33:25 cdent what provider is that 0 I/O on?
15:33:40 mriedem inap
15:35:29 openstackgerrit Matt Riedemann proposed openstack/nova master: doc: link to nova code review guide from dev policies https://review.opendev.org/692569
15:54:17 efried I already rechecked a few dansmith
15:54:27 openstackgerrit Matt Riedemann proposed openstack/nova master: Document CD mentality policy for nova contributors https://review.opendev.org/692572
15:54:34 mriedem dansmith: i took a crack at documenting that ^
15:54:37 dansmith efried: in that case, https://www.youtube.com/watch?v=0u8teXR8VE4
15:56:23 mriedem dansmith: thoughts on the no graceful shutdown thing i mentioned here? http://lists.openstack.org/pipermail/openstack-discuss/2019-October/010495.html - wondering if we should go back to like a 10 second shutdown timeout in case that is screwing up the guest during things that transfer the root disk or snapshot it (shelve/unshelve) leading to ssh failures later
15:57:09 mriedem i had made that change in devstack b/c of a real bug where we see the 60 second shutdown timeout in nova not doing anything, potentially leading to overall job timeouts if we're waiting a full minute to power off guests during a tempest run
15:57:52 dansmith mriedem: I doubt that cirros is responding to the graceful shutdown request at all
15:58:01 dansmith mriedem: so I'm not sure it matters
15:58:18 dansmith if you have a devstack up you should be able to test
15:58:46 dansmith even still, unless it's writing to the disk when the destroy happens, it wouldn't likely do any damage (and shouldn't really anyway)
15:59:12 dansmith don't we get the console of the guest we fail to ssh into if we do? unless it's sitting at a panic or some unhappy state, I'd not be concerned
15:59:28 mriedem yes we do
16:00:16 sean-k-mooney mriedem: are you suspecting disk curruption?
16:00:19 mriedem there are several guest ssh failure bugs in the gate right now, many different reasons (dhcp lease fails, out of space on the guest disk [growroot fails], kernel panic)
16:00:28 sean-k-mooney due to an agressive shutdown?
16:00:29 mriedem i'm grasping at straws
16:01:08 mriedem in one case that i caught in the gate where the 60 second shutdown timed out, i got the guest console log http://paste.openstack.org/show/752002/
16:01:25 mriedem cdent wondered (in the bug) if the metadata api retry loop in the guest was making it ignore the shutdown request
16:01:32 mriedem https://bugs.launchpad.net/nova/+bug/1829896
16:01:32 openstack Launchpad bug 1829896 in OpenStack Compute (nova) "libvirt: "Instance failed to shutdown in 60 seconds." in the gate" [Undecided,New]
16:02:05 sean-k-mooney well in that case it got a dhcp lease but then could not hit the metadata service
16:02:10 mriedem i'm also wondering if there are ideas on figuring out what is eating up all of the disk in these guests when growroot fails
16:02:17 sean-k-mooney so that looks like a neutron issue with the metadata proxy
16:03:05 sean-k-mooney if glean/cloud init is waithing for the metadata then
16:03:11 sean-k-mooney the ssh service may not have started
16:03:45 openstackgerrit Lee Yarwood proposed openstack/nova stable/queens: Add 'path' query parameter to console access url https://review.opendev.org/692573
16:06:30 lyarwood mriedem: yup, working on the other change now thanks
16:07:42 sean-k-mooney mriedem: i do know we have had issue with hot plug if the early init in the guest takes too long
16:08:05 mriedem lyarwood: i'm sort of inclined to say you guys (RH) should keep that one downstream
16:08:06 sean-k-mooney but we also see Starting acpid: OK
16:08:17 mriedem since i don't see people clamoring for using the bleeding edge novnc stuff back on queens
16:09:06 mriedem i'm not like -2, but those are my "feelings"
16:09:15 mriedem what you humans would call feelings anyway
16:09:53 lyarwood haha yeah no issues, thought I'd at least post it upstream first to see if anyone wanted it
16:10:02 sean-k-mooney ah ha proof mriedem is actully an irc bot !!!
16:10:54 sean-k-mooney mriedem: on other random straw is there seams to be very littly entropy in the guest
16:10:59 sean-k-mooney random: dd urandom read with 31 bits of entropy available
16:12:14 sean-k-mooney the topic of enabling the virtio random number generator by default has come up before. i wonder if it could be related to that.
16:14:09 mriedem i seem to remember a recentish patch of kashyap's related to this
16:15:21 mriedem i'm just thinking of this https://review.opendev.org/#/c/577385/
16:16:17 sean-k-mooney ya there is that one but we also talked a little about adding the random number gereate to the libvirt xml by defualt rather then needing to set the image property
16:16:18 mriedem we don't use that in our devstack guests though since we don't have hw_rng_model=virtio in the image
16:16:24 mriedem oh
16:16:37 mriedem well, i have to go to what you humans would call lunch now
16:28:35 openstackgerrit Lee Yarwood proposed openstack/nova stable/queens: Reduce scope of 'path' query parameter to noVNC consoles https://review.opendev.org/692581
17:19:32 dansmith cripes
17:19:42 dansmith cache notifications patch is gonna fail again
17:21:15 openstackgerrit John Garbutt proposed openstack/nova-specs master: Add Unified Limits Spec https://review.opendev.org/602201
17:34:41 mriedem dansmith: it's because of your sin
17:35:59 efried dansmith: I'm +A on mriedem's CD doc patch, but the predecessor (link shuffle) needs to be sent as well
17:36:07 efried https://review.opendev.org/#/c/692569/1
17:36:34 dansmith efried: oh sorry I missed that earlier I guess
17:36:46 dansmith I dun getified it
17:40:13 mriedem dansmith: i looked at a dstat from one of the 'timed out waiting for cell' failed jobs and the only thing around that time is a spike in paging and networking which doesn't tell me much
17:40:32 dansmith mriedem: yeah I looked pretty in depth at one dstat from a cell fail job
17:40:36 dansmith and it looked overly normal to me
17:40:53 dansmith mysql wasn't spiking that I could tell, it was the big memory user, but only like 400MB at the time
17:41:02 dansmith IO wasn't crazy, etc
17:41:30 mriedem yeah nothing obvious like the 0 i/ops in the other thing earlier today
17:42:06 dansmith I didn't catch the zero iops thing
17:42:08 mriedem and there is no rabbit involved b/c we're going straight to the db yeah?
17:42:14 dansmith you were assuming zero iops meant bad?
17:42:30 mriedem the 0 iops thing was related to another bug in the ML where we get long timeouts
17:42:42 mriedem which triggered the resize long_rpc_timeout discussion
17:43:08 mriedem and a very obvious message in the mysql error log
17:43:29 dansmith you mean zero iops for a long period of time or something I assume/
17:43:35 mriedem 3 minutes yeah
17:43:54 mriedem while placement was processing a POST /allocations request
17:43:57 dansmith okay, yeah, one or two samples with zero iops wouldn't be alarming to me, but minutes of it, for sure
17:46:49 mriedem and https://bugs.launchpad.net/nova/+bug/1825584 shouldn't be related to the scheduler timing out waiting for a response from cell1 b/c there is no uwsgi involved there
17:46:49 openstack Launchpad bug 1825584 in OpenStack Compute (nova) "eventlet monkey-patching breaks AMQP heartbeat on uWSGI" [Undecided,Confirmed]
17:46:53 mriedem just scheduler running with eventlet
18:30:42 openstackgerrit Merged openstack/nova stable/rocky: libvirt: Ignore volume exceptions during post_live_migration https://review.opendev.org/691283
18:36:07 openstackgerrit Matt Riedemann proposed openstack/nova stable/queens: libvirt: Ignore volume exceptions during post_live_migration https://review.opendev.org/691284
18:48:25 mriedem interesting that we don't have a task_state change while confirming a resized server
18:48:35 mriedem to prevent other actions from happening while still confirming and cleaning up the source host
18:57:59 mriedem but we do have task_states.RESIZE_REVERTING for reverting a resize
19:36:52 openstackgerrit Merged openstack/nova master: doc: link to nova code review guide from dev policies https://review.opendev.org/692569
19:36:57 openstackgerrit Merged openstack/nova master: Document CD mentality policy for nova contributors https://review.opendev.org/692572
19:46:58 mriedem dansmith: MessagingTimeouts in nova-next in the latest failure in https://review.opendev.org/#/c/691390/

Earlier   Later