Earlier  
Posted Nick Remark
#openstack-nova - 2021-04-28
21:54:00 dansmith we want to not regress the minimum version by starting something old, nor do we want to migrate some data that an older service that might get turned on and be confused
21:54:00 andrewbogott they were hidden in the DB
21:54:11 dansmith hidden in the db?
21:54:11 dansmith hidden in the db?
21:54:26 andrewbogott just -- I found references to hosts in the db that didn't appear in 'openstack compute service list'
21:54:26 dansmith if it's not ignoring deleted records then that's definitely a bug
21:54:26 dansmith if it's not ignoring deleted records then that's definitely a bug
21:54:27 andrewbogott just -- I found references to hosts in the db that didn't appear in 'openstack compute service list'
21:54:57 andrewbogott but I guess since I didn't restart all services after every db change I don't know which bit was causing the problem :(
21:54:57 andrewbogott but I guess since I didn't restart all services after every db change I don't know which bit was causing the problem :(
21:55:44 dansmith well, there might still be an issue with considering deleted records it sounds like, which would be a bug if you want to file it
21:55:44 dansmith well, there might still be an issue with considering deleted records it sounds like, which would be a bug if you want to file it
21:56:13 andrewbogott at the very least there's an issue with <unknown> service still reporting the version of a deleted service, at least until it's restarted
21:56:13 andrewbogott at the very least there's an issue with <unknown> service still reporting the version of a deleted service, at least until it's restarted
21:56:35 dansmith well, do we know if it's deleted or just not being shown in the API?
21:56:35 dansmith well, do we know if it's deleted or just not being shown in the API?
21:56:45 andrewbogott good point, it could be either
21:56:46 andrewbogott good point, it could be either
21:56:47 dansmith I don't know how the api behaves
21:56:47 dansmith I don't know how the api behaves
21:56:53 andrewbogott I guess I need to figure out if I can reproduce it
21:56:53 andrewbogott I guess I need to figure out if I can reproduce it
21:56:58 dansmith but yes, if it's that, then definitely a bug
21:56:58 dansmith but yes, if it's that, then definitely a bug
21:57:16 andrewbogott dansmith: does that minimum_version_check happen via RPC? I stuck a bunch of debug lines in the version check on my compute node and they were never traversed.
21:57:17 andrewbogott dansmith: does that minimum_version_check happen via RPC? I stuck a bunch of debug lines in the version check on my compute node and they were never traversed.
21:57:44 dansmith computes can't talk to the database, so yes
21:57:44 dansmith computes can't talk to the database, so yes
21:57:45 dansmith but other services go straight to the db for it
21:57:45 dansmith but other services go straight to the db for it
21:58:45 andrewbogott I don't think I've ever seen @base.remotable_classmethod but I take that to mean 'this happens on a totally different system'
21:58:45 andrewbogott I don't think I've ever seen @base.remotable_classmethod but I take that to mean 'this happens on a totally different system'
21:58:52 andrewbogott or at least can
21:58:53 andrewbogott or at least can
21:59:04 dansmith it does, if coming from compute,
21:59:05 dansmith it does, if coming from compute,
21:59:11 andrewbogott ok
21:59:11 andrewbogott ok
21:59:18 andrewbogott that solves another mystery then
21:59:18 clarkb gmann: ya and it works like 90% of the time
21:59:18 andrewbogott that solves another mystery then
21:59:18 clarkb gmann: ya and it works like 90% of the time
21:59:19 dansmith and yes, you might be right.. you might need to restart conductor which does that on behalf of the compute
21:59:19 dansmith and yes, you might be right.. you might need to restart conductor which does that on behalf of the compute
21:59:38 clarkb gmann: that is why I suspect maybe it has to do with the cache refresh since in theory that only happens if you exceed timeouts or similar
21:59:38 clarkb gmann: that is why I suspect maybe it has to do with the cache refresh since in theory that only happens if you exceed timeouts or similar
21:59:45 dansmith yeah, I hadn't really thought of that, but I guess we'd get conductor's in-memory "cache" of that version
21:59:46 dansmith yeah, I hadn't really thought of that, but I guess we'd get conductor's in-memory "cache" of that version
22:00:12 andrewbogott That seems like the simplest explanation. Not easy to fix though
22:00:12 andrewbogott That seems like the simplest explanation. Not easy to fix though
22:00:20 andrewbogott other than having the 'delete' command print
22:00:20 andrewbogott other than having the 'delete' command print
22:00:23 dansmith well,
22:00:26 andrewbogott "ok now restart your conductor"
22:00:26 andrewbogott "ok now restart your conductor"
22:00:27 dansmith not sure that needs a fix
22:00:27 dansmith not sure that needs a fix
22:00:56 andrewbogott d'you think the conductor would've refreshed its state eventually?
22:00:56 andrewbogott d'you think the conductor would've refreshed its state eventually?
22:01:17 clarkb gmann: fungi mentioned it could possibly be an openstack(client|sdk) regression where it isn't passing that info properly
22:01:17 dansmith no, all the nodes will hold that value in memory until restart, that's kinda the point
22:01:17 dansmith no, all the nodes will hold that value in memory until restart, that's kinda the point
22:01:17 clarkb gmann: fungi mentioned it could possibly be an openstack(client|sdk) regression where it isn't passing that info properly
22:01:28 dansmith API and scheduler will do the same
22:01:28 dansmith API and scheduler will do the same
22:01:46 dansmith it's just that restarting an api worker will resolve it, whereas computes are dependent on the conductors for it
22:01:46 dansmith it's just that restarting an api worker will resolve it, whereas computes are dependent on the conductors for it
22:03:21 andrewbogott It's really the conductor's job to provide /current/ database state though isn't it?
22:03:21 andrewbogott It's really the conductor's job to provide /current/ database state though isn't it?
22:03:40 andrewbogott Are there reasons why we would want the conductor to not re-read for every query? (Other than performance)
22:03:40 andrewbogott Are there reasons why we would want the conductor to not re-read for every query? (Other than performance)
22:04:16 dansmith yeah, it's expensive, and it also controls conductor's RPC pin, which we don't want to shift at runtime unless you restart or HUP it
22:04:16 dansmith yeah, it's expensive, and it also controls conductor's RPC pin, which we don't want to shift at runtime unless you restart or HUP it
22:05:43 dansmith the conductor also translates things according to the version pin and the versions supported by the compute that is asking for a thing
22:05:43 dansmith the conductor also translates things according to the version pin and the versions supported by the compute that is asking for a thing
22:05:47 dansmith it's not just a transparent proxy
22:05:47 dansmith it's not just a transparent proxy
22:06:00 andrewbogott hmmmm
22:06:00 andrewbogott hmmmm
22:06:49 andrewbogott So really it should be pushed -- when a service is deleted it should force some kind of reload (rather than the conductor re-reading the db periodically)
22:06:49 andrewbogott So really it should be pushed -- when a service is deleted it should force some kind of reload (rather than the conductor re-reading the db periodically)
22:06:59 andrewbogott that sounds like a pain to implement for a tiny gain
22:06:59 andrewbogott that sounds like a pain to implement for a tiny gain
22:07:44 dansmith you want to fan-out to every service in the cluster every time a service is deleted just in case some of them need to re-calculate the min version?
22:07:45 dansmith you want to fan-out to every service in the cluster every time a service is deleted just in case some of them need to re-calculate the min version?
22:07:50 dansmith seems like a lot of pain for little gain
22:07:50 dansmith seems like a lot of pain for little gain
22:07:56 dansmith or maybe that's what you meant.
22:07:56 dansmith or maybe that's what you meant.
22:08:07 andrewbogott yeah, a lot of pain
22:08:07 andrewbogott yeah, a lot of pain
22:09:18 andrewbogott So a real bug but not one that's worth fixing
22:09:18 andrewbogott So a real bug but not one that's worth fixing
22:09:37 openstackgerrit Dmitrii Shcherbakov proposed openstack/nova-specs master: Introduce Transport Nodes https://review.opendev.org/c/openstack/nova-specs/+/787458
22:10:24 andrewbogott dansmith: I need to run -- thanks again for getting me unstuck!
22:10:24 andrewbogott dansmith: I need to run -- thanks again for getting me unstuck!
22:15:41 dansmith yup
22:15:42 dansmith yup
#openstack-nova - 2021-04-29
00:34:11 brinzhang gibi, gmann: the nova-grenade-multinode task always failed, is there a bug tracing?
00:34:11 brinzhang gibi, gmann: the nova-grenade-multinode task always failed, is there a bug tracing?

Earlier   Later