| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-11-01 | |||
| 15:01:00 | dansmith | sounds like a good thing to do tome | |
| 15:01:18 | Sundar | I have generally thought of the patch series as a unified whole. Merging individual parts will not give you any useful functionality. The base patch has a procedural -2, so that the whole series will merge together. | |
| 15:01:40 | dansmith | it's fine to not have incremental functionality, | |
| 15:01:46 | Sundar | That said, I don't believe in merging 'broken patches' either | |
| 15:01:52 | dansmith | it's not fine to merge things that use something before the thing is available | |
| 15:02:12 | dansmith | and we can't arrange for the whole series to merge as a whole, | |
| 15:02:34 | dansmith | which means the -2 holds things until all the pieces are ready, but they still have to be mergeable in isolation, because they will | |
| 15:02:54 | mriedem | especially since it takes about 3 days of rechecks to merge anything in nova right now | |
| 15:03:08 | dansmith | three days? please | |
| 15:03:19 | mriedem | ok ok 5 days | |
| 15:03:20 | dansmith | my notifications thing has been going all week.. maybe since last week | |
| 15:03:49 | mriedem | for big series like this it's always best to work from the bottom up where everything is noop until "turned on" in the API at the end if possible | |
| 15:03:52 | dansmith | yeah, since monday | |
| 15:04:01 | mriedem | so there is less risk in merging the bottom noop stuff | |
| 15:07:45 | Sundar | mriedem: The code flow starts from the API, such as getting device profile info based on the extra specs. If the API stuff came later, wouldn't the patch sequence be out of order? We can only do UT for the earlier patches then. | |
| 15:08:02 | cdent | mriedem: I assume you saw my response on the gate+placement email? Sorry I've not been able to dig more deeply. I've been engaged elsewhere...To such an extent that I may come off the mailing list pretty soon to maintain sanity. | |
| 15:09:01 | efried | Sundar: in this case the "master switch" could conceivably be the part that processes the device profile out of the flavor. | |
| 15:09:28 | mriedem | cdent: yup, and replied, and thanks | |
| 15:09:45 | mriedem | cdent: i think the likeliest culprit is noise neighbors on oversubscribed nodes as you said | |
| 15:17:28 | cdent | mriedem: the gist of what I can find on that innodb error (now that I've managed to actually read your responses) is "you disk i/o is way constrained" | |
| 15:19:32 | mriedem | cdent: heh, yeah looking at this dstat graph, i/o is just 0 from 16:52 to 16:55 | |
| 15:19:42 | cdent | oh dear | |
| 15:20:32 | cdent | there should be some kind of "how am I supposed to work in these conditions" meme we can interject here | |
| 15:21:08 | efried | https://makeameme.org/meme/i-cant-work-2sl5sn | |
| 15:21:11 | mriedem | https://toddroblog.files.wordpress.com/2017/02/drunk-hobo-flips-double-bird-1.jpg?w=809 ? | |
| 15:21:22 | mriedem | oh blurred out?! | |
| 15:23:09 | mriedem | i wish i knew about https://lamada.eu/dstat-graph/# years ago | |
| 15:26:44 | cdent | yeah, why don't things like that get passed around more easily? | |
| 15:27:16 | mriedem | probably because i was too shy to say "i'm dumb, how can i see this dstat output in a graph?" | |
| 15:27:44 | mriedem | i assume all of my co-workers to write tools to generate graphs on the fly for everything they do | |
| 15:29:50 | cdent | https://www.youtube.com/watch?v=MKIaS0lh-uo | |
| 15:30:48 | mriedem | heh, it's rare to get a repo man reference in here | |
| 15:33:25 | cdent | what provider is that 0 I/O on? | |
| 15:33:40 | mriedem | inap | |
| 15:35:29 | openstackgerrit | Matt Riedemann proposed openstack/nova master: doc: link to nova code review guide from dev policies https://review.opendev.org/692569 | |
| 15:54:17 | efried | I already rechecked a few dansmith | |
| 15:54:27 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Document CD mentality policy for nova contributors https://review.opendev.org/692572 | |
| 15:54:34 | mriedem | dansmith: i took a crack at documenting that ^ | |
| 15:54:37 | dansmith | efried: in that case, https://www.youtube.com/watch?v=0u8teXR8VE4 | |
| 15:56:23 | mriedem | dansmith: thoughts on the no graceful shutdown thing i mentioned here? http://lists.openstack.org/pipermail/openstack-discuss/2019-October/010495.html - wondering if we should go back to like a 10 second shutdown timeout in case that is screwing up the guest during things that transfer the root disk or snapshot it (shelve/unshelve) leading to ssh failures later | |
| 15:57:09 | mriedem | i had made that change in devstack b/c of a real bug where we see the 60 second shutdown timeout in nova not doing anything, potentially leading to overall job timeouts if we're waiting a full minute to power off guests during a tempest run | |
| 15:57:52 | dansmith | mriedem: I doubt that cirros is responding to the graceful shutdown request at all | |
| 15:58:01 | dansmith | mriedem: so I'm not sure it matters | |
| 15:58:18 | dansmith | if you have a devstack up you should be able to test | |
| 15:58:46 | dansmith | even still, unless it's writing to the disk when the destroy happens, it wouldn't likely do any damage (and shouldn't really anyway) | |
| 15:59:12 | dansmith | don't we get the console of the guest we fail to ssh into if we do? unless it's sitting at a panic or some unhappy state, I'd not be concerned | |
| 15:59:28 | mriedem | yes we do | |
| 16:00:16 | sean-k-mooney | mriedem: are you suspecting disk curruption? | |
| 16:00:19 | mriedem | there are several guest ssh failure bugs in the gate right now, many different reasons (dhcp lease fails, out of space on the guest disk [growroot fails], kernel panic) | |
| 16:00:28 | sean-k-mooney | due to an agressive shutdown? | |
| 16:00:29 | mriedem | i'm grasping at straws | |
| 16:01:08 | mriedem | in one case that i caught in the gate where the 60 second shutdown timed out, i got the guest console log http://paste.openstack.org/show/752002/ | |
| 16:01:25 | mriedem | cdent wondered (in the bug) if the metadata api retry loop in the guest was making it ignore the shutdown request | |
| 16:01:32 | openstack | Launchpad bug 1829896 in OpenStack Compute (nova) "libvirt: "Instance failed to shutdown in 60 seconds." in the gate" [Undecided,New] | |
| 16:01:32 | mriedem | https://bugs.launchpad.net/nova/+bug/1829896 | |
| 16:02:05 | sean-k-mooney | well in that case it got a dhcp lease but then could not hit the metadata service | |
| 16:02:10 | mriedem | i'm also wondering if there are ideas on figuring out what is eating up all of the disk in these guests when growroot fails | |
| 16:02:17 | sean-k-mooney | so that looks like a neutron issue with the metadata proxy | |
| 16:03:05 | sean-k-mooney | if glean/cloud init is waithing for the metadata then | |
| 16:03:11 | sean-k-mooney | the ssh service may not have started | |
| 16:03:45 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/queens: Add 'path' query parameter to console access url https://review.opendev.org/692573 | |
| 16:06:30 | lyarwood | mriedem: yup, working on the other change now thanks | |
| 16:07:42 | sean-k-mooney | mriedem: i do know we have had issue with hot plug if the early init in the guest takes too long | |
| 16:08:05 | mriedem | lyarwood: i'm sort of inclined to say you guys (RH) should keep that one downstream | |
| 16:08:06 | sean-k-mooney | but we also see Starting acpid: OK | |
| 16:08:17 | mriedem | since i don't see people clamoring for using the bleeding edge novnc stuff back on queens | |
| 16:09:06 | mriedem | i'm not like -2, but those are my "feelings" | |
| 16:09:15 | mriedem | what you humans would call feelings anyway | |
| 16:09:53 | lyarwood | haha yeah no issues, thought I'd at least post it upstream first to see if anyone wanted it | |
| 16:10:02 | sean-k-mooney | ah ha proof mriedem is actully an irc bot !!! | |
| 16:10:54 | sean-k-mooney | mriedem: on other random straw is there seams to be very littly entropy in the guest | |
| 16:10:59 | sean-k-mooney | random: dd urandom read with 31 bits of entropy available | |
| 16:12:14 | sean-k-mooney | the topic of enabling the virtio random number generator by default has come up before. i wonder if it could be related to that. | |
| 16:14:09 | mriedem | i seem to remember a recentish patch of kashyap's related to this | |
| 16:15:21 | mriedem | i'm just thinking of this https://review.opendev.org/#/c/577385/ | |
| 16:16:17 | sean-k-mooney | ya there is that one but we also talked a little about adding the random number gereate to the libvirt xml by defualt rather then needing to set the image property | |
| 16:16:18 | mriedem | we don't use that in our devstack guests though since we don't have hw_rng_model=virtio in the image | |
| 16:16:24 | mriedem | oh | |
| 16:16:37 | mriedem | well, i have to go to what you humans would call lunch now | |
| 16:28:35 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/queens: Reduce scope of 'path' query parameter to noVNC consoles https://review.opendev.org/692581 | |
| 17:19:32 | dansmith | cripes | |
| 17:19:42 | dansmith | cache notifications patch is gonna fail again | |
| 17:21:15 | openstackgerrit | John Garbutt proposed openstack/nova-specs master: Add Unified Limits Spec https://review.opendev.org/602201 | |
| 17:34:41 | mriedem | dansmith: it's because of your sin | |
| 17:35:59 | efried | dansmith: I'm +A on mriedem's CD doc patch, but the predecessor (link shuffle) needs to be sent as well | |
| 17:36:07 | efried | https://review.opendev.org/#/c/692569/1 | |
| 17:36:34 | dansmith | efried: oh sorry I missed that earlier I guess | |
| 17:36:46 | dansmith | I dun getified it | |
| 17:40:13 | mriedem | dansmith: i looked at a dstat from one of the 'timed out waiting for cell' failed jobs and the only thing around that time is a spike in paging and networking which doesn't tell me much | |
| 17:40:32 | dansmith | mriedem: yeah I looked pretty in depth at one dstat from a cell fail job | |
| 17:40:36 | dansmith | and it looked overly normal to me | |
| 17:40:53 | dansmith | mysql wasn't spiking that I could tell, it was the big memory user, but only like 400MB at the time | |
| 17:41:02 | dansmith | IO wasn't crazy, etc | |
| 17:41:30 | mriedem | yeah nothing obvious like the 0 i/ops in the other thing earlier today | |
| 17:42:06 | dansmith | I didn't catch the zero iops thing | |
| 17:42:08 | mriedem | and there is no rabbit involved b/c we're going straight to the db yeah? | |
| 17:42:14 | dansmith | you were assuming zero iops meant bad? | |
| 17:42:30 | mriedem | the 0 iops thing was related to another bug in the ML where we get long timeouts | |
| 17:42:42 | mriedem | which triggered the resize long_rpc_timeout discussion | |
| 17:43:08 | mriedem | and a very obvious message in the mysql error log | |