Earlier  
Posted Nick Remark
#openstack-nova - 2019-09-23
16:08:01 mriedem "oh hi your cloud provider is doing an upgrade and you shouldn't be trying to resize your instance right now bob!"
16:08:43 mriedem stephenfin: where have we dropped support for old microversions in novaclient?
16:08:52 gibi mriedem: the funny thing is that I think we cannot hit that error after the fix
16:09:04 gibi mriedem: so we could even remove it
16:09:21 stephenfin mriedem: e.g. 01fb16533bf562f39fe822bc12b9cc34b8580359
16:09:30 mriedem gibi: well, technically your fix is in conductor checking it's config but the computes involved in the resize/cold migrate could have different config with different pins
16:09:53 mriedem stephenfin: that one broke osc
16:09:55 gibi mriedem: can we pin our computes differently? how will they talk to each other?
16:10:08 mriedem and they had to fix to avoid using novaclient and hit the api directly
16:10:26 mriedem gibi: well i meant if your conductor was unpinned but the computes were, or not restarted yet or something
16:10:28 mriedem after unpinning them
16:10:32 mriedem i'm nto saying that's normal
16:10:47 mriedem but i'd probably leave the error checking in place until we have a compute rpc api major version bump to 6.0
16:10:53 gibi mriedem: OK. I so it is a sort transinet. I can accept that
16:10:57 stephenfin Yup, I recall that coming up now. Any reason we opted to do that instead of reverting the changes to 'novaclient/v2/client.py'?
16:11:04 stephenfin that = call the API directly
16:11:22 mriedem gibi: so if you wanted, you could write a separate patch which (1) logs the upgrade stuff but doesn't include it in the PortUpdateFailed message, (2) adds a TODO to that PortUpdateFailed code to say we can drop it after compute RPC API is bumped to 6.0
16:11:50 gibi mriedem: sure. I can do that
16:11:51 mriedem stephenfin: it was released late in novaclient, end of the milestone before FF and i think osc found the problem too late
16:12:18 mriedem stephenfin: so a mixture of releasing breaking / major version things in novaclient too late in the release to catch them and (2) osc historically not doing functional testing
16:12:30 mriedem well, not doing good enough coverage with functional testing
16:13:40 mriedem and (3) me not realizing osc was going to be broken
16:14:06 stephenfin that makes sense
16:14:10 mriedem that was probably before i cared more about osc and just figured, "this has been deprecated in novaclient since newton, we can surely drop it and no one will care"
16:14:35 mriedem so having been burned before, and burning myself and others, that's why i'm less cavalier about just removing old stuff
16:15:40 stephenfin Yeah, certainly true for something like this
16:15:50 mriedem speaking of an upgrade-related burning sensation https://review.opendev.org/#/q/topic:bug/1843090+(status:open+OR+status:merged)
16:16:36 stephenfin I was holding off on that til some people had looked at the oslo.messaging bug
16:16:44 stephenfin though maybe that doesn't make sense
16:17:28 mriedem you mean gibi's RequestSpecImageSerializationFixture in the functoinal test?
16:17:33 mriedem it's an existing pattern / known issue
16:18:28 mriedem could also summon dansmith to review those since it's rpc pin related
16:18:33 mriedem and that's dan's middle name
16:18:50 mriedem dan rpc-pin-and-sometimes-evacuate smith
16:19:15 stephenfin If dansmith doesn't get to it by tomorrow, I will
16:29:48 mriedem gibi: i think we could drop the new and redundant unit test in https://review.opendev.org/#/c/683948/
16:29:52 mriedem the functional test covers it
17:24:01 openstackgerrit Merged openstack/nova master: Add note about needing noVNC >= v1.1.0 with using ESX https://review.opendev.org/682946
17:30:21 bbobrov_ hi! In https://opendev.org/openstack/nova/src/branch/master/nova/servicegroup/drivers/db.py#L99 messaging.MessagingTimeout is getting caught. My understanding is that service_ref.save() only interacts with the database. How can MessagingTimeout happen?
17:38:48 sean-k-mooney bbobrov: if this is executing on the compute node then the save call would invoke the db update via rpc as the comptue services do not have database access
17:41:09 sean-k-mooney bbobrov: so on the concoctor,api and schduler this shoudl not raise a messaging timeout but it can in the compute agents
17:41:32 bbobrov sean-k-mooney: understood, thanks
17:45:55 openstackgerrit Matt Riedemann proposed openstack/nova master: Rename Claims resources to compute_node https://review.opendev.org/679470
17:50:20 openstackgerrit Matt Riedemann proposed openstack/nova master: Add a prelude for the Train release https://review.opendev.org/683327
17:53:47 bbobrov #2 then. there is a `periodic_enable` option, and in nova codebase it seems to affect only https://opendev.org/openstack/nova/src/branch/master/nova/service.py#L205 . if true, self.periodic_tasks is scheduled to be run by add_dynamic_timer
17:54:08 bbobrov But there is no implementation of periodic_tasks or run_periodic_tasks anywhere in Nova. The method in the base scheduler class does `pass`. Are there any implementation of run_periodic_tasks in Nova that i miss or they can come from somewhere else?
17:56:58 sean-k-mooney we use the one form oslo service https://github.com/openstack/oslo.service/blob/master/oslo_service/periodic_task.py
18:07:56 mriedem melwitt: do you know if tripleo has any kind of tooling that pokes a cell mq/db on upgrade to make sure it's ok before saying everything is good to go in the api?
18:22:38 bbobrov sean-k-mooney: thanks, it makes sense now
18:32:57 lyarwood mriedem: we are only just introducing full support for multi cell deployments now in 16 but I can ask around about that in the morning.
18:33:48 lyarwood owalsh: ^ unless you're around and know?
18:34:52 mriedem lyarwood: yeah in grenade upstream jobs we don't have multiple cells, just cell0 and cell1
18:34:57 mriedem trying to debug https://bugs.launchpad.net/nova/+bug/1844929
18:34:58 openstack Launchpad bug 1844929 in OpenStack Compute (nova) "grenade jobs failing due to "Timed out waiting for response from cell" in scheduler" [High,Confirmed]
18:35:35 mriedem it looks like we upgrade code to train, restart scheduler, it starts up, even hits the cell1 database to pull compute nodes and instances for it's in-memory cache, and then on the first scheduling attempt after that we hang
18:37:45 dansmith mriedem: scheduler doesn't need rabbit to do that, so the rabbit errors in your bug report would be unrelated I'd think
18:39:41 mriedem true, didn't think of that
18:39:42 mriedem obviously
18:39:49 mriedem i just left a comment with a bunch of log links,
18:40:17 mriedem but it looks like we start up ok, hit the cell1 db to pull compute nodes and instances, then 4 minutes later is the first scheduling attempt after the upgrade and at that point we're timing out
18:41:24 sean-k-mooney dansmith: we do eventurlly need to adress the rabbit mq heartbeat issue even if it is unrelated to this issue
18:41:32 mriedem it also looks like this goes back to 9/14 which is a few days before i originally thought
18:41:35 dansmith mriedem: that timeout is related to the scatter
18:42:04 dansmith mriedem: that's not new information, I'm just saying
18:42:29 dansmith mriedem: I was expecting to see a trace, but I think the scatter squashes that
18:43:15 sean-k-mooney i belive if any of the request raise an excpeiton we catch it and return the exception instead of raising it or something like that
18:43:20 mriedem it just logs a warning without a trace
18:43:31 dansmith mriedem: right
18:43:35 sean-k-mooney so i dont think we actuly log the traceback the way we normally would
18:43:52 dansmith sean-k-mooney: didn't I just say that?
18:44:34 sean-k-mooney yes i was halfway through typing it when you did so i finished it and hit enter
18:44:34 mriedem what i'm concerned about is if something in https://review.opendev.org/#/c/641907/ which merged and was released late is causing some side effect with the restart
18:44:48 dansmith mriedem: you said it was able to prime its cache on startup right? so this is it getting all nodes from all cells on the schedule run?
18:45:11 mriedem we're getting through this https://github.com/openstack/nova/blob/597b34cd87ac349c0f3702a872630f3c830b1483/nova/scheduler/host_manager.py#L413
18:45:18 mriedem and then 4 minutes later the first scheduling request comes,
18:45:25 mriedem and it's going back to the cell to pull compute nodes by uuid
18:45:27 mriedem and times out
18:45:31 dansmith mriedem: I guess I thought we were stopping with systemd in that case, so no sighup involed yeah?
18:47:40 sean-k-mooney https://github.com/openstack/grenade/blob/master/projects/60_nova/shutdown.sh#L21
18:47:57 sean-k-mooney i have not checked devstack but yes
18:48:16 sean-k-mooney i think tha tdoes systemctl stop devstack@n=* effectivly
18:48:45 dansmith right
18:49:03 sean-k-mooney ya it does https://github.com/openstack/devstack/blob/master/lib/nova#L1047-L1059
18:49:49 mriedem Sep 22 00:37:32.126839 ubuntu-bionic-ovh-gra1-0011664420 systemd[1]: Stopped Devstack devstack@n-sch.service.
18:49:53 mriedem Sep 22 00:45:55.359862 ubuntu-bionic-ovh-gra1-0011664420 systemd[1]: Started Devstack devstack@n-sch.service.
18:50:42 dansmith mriedem: do we even stop/start mysql in grenade during the upgrade?
18:50:44 mriedem Sep 22 00:37:27.786606 ubuntu-bionic-ovh-gra1-0011664420 nova-scheduler[25563]: INFO oslo_service.service [None req-91e88f0d-9b5c-4cb7-a5e9-e7309f922832 None None] Caught SIGTERM, stopping children
18:50:46 mriedem yeah not a HUP
18:50:54 mriedem dansmith: pretty sure we don't
18:51:21 dansmith mriedem: yeah, so, pretty weird that it connects and works once and then times out later, because it's not like we restart mysql after that point or something
18:51:35 dansmith mriedem: so I wonder if we're actually doing something other than what we think
18:51:48 dansmith and since we don't log a trace there we don't know where it's actually timing out
18:52:46 openstackgerrit Dustin Cowles proposed openstack/nova-specs master: Spec: Provider config YAML file https://review.opendev.org/680471
18:59:17 openstackgerrit Matt Riedemann proposed openstack/nova master: Log CellTimeout traceback in scatter_gather_cells https://review.opendev.org/684118
18:59:18 mriedem dansmith: thinking like this? ^
19:00:13 dansmith mriedem: yeah, I mean we kinda specifically decided not to explode there so we're tolerant of transient failures
19:00:26 dansmith I wonder if we should log debug with exc_info=True and warning without or something
19:00:31 dansmith but yeah, something
19:01:01 mriedem was thinking about passing down a kwarg to scatter_gather_cells to tell it what to do, but that's probably icky
19:01:12 mriedem trace_on_timeout = kwargs.pop('trace_on_timeout', False)

Earlier   Later