| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-11-01 | |||
| 15:54:08 | artom | Chicken, meet again. Cart, meet horse. | |
| 15:54:17 | artom | "again"? I mean egg | |
| 15:57:57 | mriedem | the only thing i could think there is we could try updating the instance within the build request, but that's pretty shitty | |
| 15:59:04 | sean-k-mooney | mriedem: in this case however it seams like they are booting with an invalid flavor id right? | |
| 15:59:14 | mriedem | no | |
| 15:59:22 | mriedem | he's injecting some kind of fault into the code | |
| 15:59:24 | mriedem | to trigger the db error | |
| 15:59:35 | mriedem | if you try to boot with an invalid flavor id, you'll get a 404 in the api looking up the flavor | |
| 16:00:16 | sean-k-mooney | oh ok i was trying to figure out how the create a flavor with id 1E+22 but then failed to boot with that flavor | |
| 16:00:18 | mriedem | so, i mean, your cell db could drop right when we're trying to create the server i guess, that would do it as well | |
| 16:00:24 | mriedem | but are we going to handle that scenario everywhere in nova? | |
| 16:01:11 | sean-k-mooney | ok right so in that case the insnace would be in the api db but fail to insert into the cell db | |
| 16:02:08 | sean-k-mooney | me moved the instance staus into the api db recnetly right so ya in that case we coudl set error on the api db i guess but there are a tone of other edgecase like that we dont handel | |
| 16:04:18 | sean-k-mooney | mriedem: the other thing we coudl do is have a periodic task that just updates the status of perptually building instance to error after some time e.g a day or rety limit*build timeout or something | |
| 16:05:14 | openstack | Launchpad bug 1800204 in OpenStack Compute (nova) "n-cpu.service consuming 100% of CPU indeterminately" [Undecided,New] | |
| 16:05:14 | mriedem | this guy has been busy https://bugs.launchpad.net/nova/+bug/1800204 https://bugs.launchpad.net/nova/+bug/1799949 | |
| 16:05:15 | openstack | Launchpad bug 1799949 in OpenStack Compute (nova) "VM instance building forever when an RPC error occurs" [Undecided,New] | |
| 16:06:20 | openstack | Launchpad bug 1800204 in OpenStack Compute (nova) "n-cpu.service consuming 100% of CPU indeterminately" [Undecided,New] | |
| 16:06:20 | sean-k-mooney | mriedem: actull https://bugs.launchpad.net/nova/+bug/1800204 seams familar there was a similar bug report a few monts back around the rocky release | |
| 16:13:15 | sean-k-mooney | oh wait that n-cpu not the conductor never mind | |
| 16:16:10 | sean-k-mooney | mriedem: do you think we will actully adress any of those bugs. | |
| 16:16:34 | stephenfin | artom: So, do I need to review https://code.engineering.redhat.com/gerrit/#/c/154627/2 yet? | |
| 16:16:45 | sean-k-mooney | stephenfin: wrong irc | |
| 16:16:51 | stephenfin | ta :) | |
| 16:19:14 | mriedem | sean-k-mooney: probably not | |
| 16:19:35 | mriedem | unless there is a more obvious way to create those faults with injecting code into the path and blow up the system | |
| 16:19:42 | mriedem | *without | |
| 16:22:10 | sean-k-mooney | mriedem: i was just debating if we shoudl triage them as incomplete or wontfix unless a different way to reporduce can be provided | |
| 16:23:48 | mriedem | i marked one of them as opinion | |
| 16:25:01 | johnthetubaguy | FWIW, I always wanted to be able to "timeout" tasks to try and catch that pending forever case. They caused me endless pain at Rackspace (I think mostly in the migrate/resize code path). The difference was they were more expected / user triggered errors. | |
| 16:25:06 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add --before to nova-manage db archive_deleted_rows https://review.openstack.org/556751 | |
| 16:26:07 | mriedem | johnthetubaguy: how much of that was resolved with service user tokens though? | |
| 16:26:14 | mriedem | or the long_rpc_timeout we have since rocky | |
| 16:26:23 | mriedem | which we're using now in the live migration flows that do rpc calls | |
| 16:27:17 | johnthetubaguy | mriedem: yeah, I saw that go in. Although most of those cases it went to Error (eventually) when it didn't have to. | |
| 16:27:34 | mriedem | that's a different bug then | |
| 16:29:05 | sean-k-mooney | johnthetubaguy: i have seen this happen with rabbitmq restarte in the past where when perstiency was disabled on instance build and a few other cases. | |
| 16:29:40 | sean-k-mooney | i never really considerd that a nova bug however because i cased the issue by restarting rabbit | |
| 16:31:04 | sean-k-mooney | johnthetubaguy: but ya its proably more complcated then jsut set to error after x time as some request could still be in flight | |
| 16:36:00 | melwitt | mriedem: yes, it completely slipped my mind :( and I'm not done going through the entire list of the schedule yet | |
| 16:40:12 | mriedem | melwitt: i added several sessions in there based on my schedule | |
| 16:41:09 | melwitt | ok, thank you. that's helpful | |
| 16:47:07 | johnthetubaguy | sean-k-mooney: yeah, its hard to get right | |
| 16:56:53 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: API microversion bump for handling-down-cell https://review.openstack.org/591657 | |
| 16:56:54 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add DownCellFixture https://review.openstack.org/614810 | |
| 16:56:57 | mriedem | tssurya: ^ | |
| 16:57:10 | tssurya | looking, thanks | |
| 17:04:43 | tssurya | mriedem: okay I am going to write the tests here https://review.openstack.org/#/c/591657/12/nova/tests/functional/api_sample_tests/test_servers.py based on your fixture | |
| 17:13:08 | openstackgerrit | Merged openstack/nova master: Make ResourceTracker.tracked_instances a set https://review.openstack.org/608781 | |
| 17:18:31 | dansmith | mriedem: tssurya I'm explaining the host-status concept to someone right now, and why an instance state doesn't go to STOPPED just because the compute node is down | |
| 17:18:47 | dansmith | mriedem: I wonder if it would make sense to integrate the use of the UNKNOWN state we're adding here with that feature, | |
| 17:18:59 | dansmith | so that in the same microversion, instances with a down host show up as UNKNOWN as well | |
| 17:25:59 | tssurya | dansmith: you mean you want to add a new "UNKNOWN" vm_state ? | |
| 17:26:20 | dansmith | tssurya: you're already doing that from the external view right now | |
| 17:26:25 | tssurya | yea | |
| 17:26:39 | dansmith | we would do a similar thing for real instances we can look up just fine, but which have down hosts | |
| 17:27:05 | dansmith | the only problem would be that right now UNKNOWN means "the rest of the instance details aren't there" which would be slightly more ambiguous in this case | |
| 17:28:27 | tssurya | hmm, makes sense to make the instance state UNKNOWN since we don't know the host state, I mean I guess "UNKNOWN" could mean unknown details/state right ? | |
| 17:28:54 | belmoreira | dansmith mriedem should placement/nova issues be discussed here or in placement channel | |
| 17:28:57 | cfriesen | dansmith: for what it's worth, in our environment if a compute node goes down an external entity sets all of the instances to the "error" state, until they were automatically recovered. | |
| 17:29:30 | dansmith | that seems like an improper use of the error state to me | |
| 17:29:48 | dansmith | not to mention that nova on its own won't know whether they're still up and fine or not | |
| 17:29:50 | cfriesen | if the host is "down", then we fence it off and force a reboot. those instances are guaranteed to be toast | |
| 17:29:53 | dansmith | which is why we don't call them "stopped" | |
| 17:30:15 | dansmith | cfriesen: okay, well, that's better in that case but vanilla nova can't do or know that | |
| 17:30:42 | cfriesen | agreed, nova itself can't know the bigger picture | |
| 17:30:47 | dansmith | belmoreira: depends on what it is.. if it's integration issues then probably here | |
| 17:32:03 | belmoreira | is the increase of the number of requests to placement | |
| 17:32:16 | cfriesen | dansmith: although, if an external entity uses the nova API to tell nova that the compute node is "down", it's supposed to have already fenced off the node to prevent instances from (eg) talking to volumes. | |
| 17:32:44 | belmoreira | have a look into: https://docs.google.com/document/d/1d5k1hA3DbGmMyJbXdVcekR12gyrFTaj_tJdFwdQy-8E/edit?usp=sharing | |
| 17:32:49 | cfriesen | otherwise you could evacuate and then have two copies of an instance trying to access the same cinder volume | |
| 17:33:17 | dansmith | cfriesen: yeah, true.. I guess I just prefer something less overloaded like UNKNOWN than saying it's stopped or error | |
| 17:33:18 | belmoreira | this is the number of requests to placement when compute-nodes get upgrade to Rocky | |
| 17:33:28 | dansmith | efried: ^ | |
| 17:33:39 | efried | how far back am I reading? | |
| 17:33:44 | dansmith | efried: one line | |
| 17:33:52 | dansmith | and the url he posted | |
| 17:35:08 | dansmith | belmoreira: is it really increasing, or is that you bringing nodes on over time? | |
| 17:36:00 | belmoreira | the increase of requests shows the compute nodes being upgraded over time (queens -> rocky) | |
| 17:36:10 | dansmith | okay | |
| 17:36:15 | dansmith | (ouch) | |
| 17:36:22 | efried | What's happening at that cat-head bump? | |
| 17:36:31 | efried | or possibly batman | |
| 17:36:34 | dansmith | online migrations? | |
| 17:37:22 | efried | those look like trait requests | |
| 17:37:34 | efried | if I'm reading this right. | |
| 17:37:57 | belmoreira | no, it must be a cell that upgraded and then stopped nova-compute | |
| 17:38:31 | belmoreira | so, in the second graph we can see all the new requests | |
| 17:39:25 | belmoreira | UUID/trais ; ?in_tree; UUID/aggregates; ... | |
| 17:40:54 | efried | right, so it looks to me like, before rocky, we weren't calling ?in_tree, UUID/aggregates, or ?member_of at all. Which makes sense. | |
| 17:41:27 | belmoreira | yes, and this seems to be the reason of the increase of requests | |
| 17:42:10 | efried | but also increased number of requests for inventories. | |
| 17:42:16 | belmoreira | but is a huge increase. Just added another graph with the response time of my placement infrastructure | |
| 17:42:56 | efried | I have to say, this isn't all that surprising. | |
| 17:44:39 | efried | although, hm, I would have expected this jump in queens | |
| 17:44:57 | efried | belmoreira: Was this an upgrade from queens, or from earlier? | |
| 17:45:33 | belmoreira | efried from queens | |
| 17:47:46 | belmoreira | I could handle it creating more placement nodes (x3). But looks too much... | |
| 17:48:15 | efried | belmoreira: Can you give me a sense of what this timeline represents? At what point are all the upgrades done and the cloud in stable state? | |
| 17:50:34 | belmoreira | efried the nova/placement control plane was upgraded between 8:00 and 9:00. ~12:00 the compute nodes started to upgrade (this takes 24h for all of them upgrade) | |