| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-02-04 | |||
| 18:51:19 | dansmith | I have stuff that has been in the check queue for three hours and it hasn't started to run a single thing | |
| 18:51:30 | sean-k-mooney | for prioritisation beteween pipliens | |
| 18:51:38 | dansmith | I've already talked to infra about this, | |
| 18:51:43 | dansmith | and the other precedences are used for things | |
| 18:51:49 | dansmith | check and experimental are the same even | |
| 18:51:50 | sean-k-mooney | yep they are | |
| 18:51:58 | sean-k-mooney | yep both low | |
| 18:52:07 | sean-k-mooney | althoguh experimtal will report back first | |
| 18:52:20 | dansmith | again, I think we can do a lot without making this an infra problem | |
| 18:52:25 | sean-k-mooney | they have teh same precidence but are in differnet queues | |
| 18:52:33 | sean-k-mooney | dansmith: oh ya i know | |
| 18:52:33 | dansmith | and just making it so we get fast check in two hours and slow check in 24 hours isn't really going to help | |
| 18:52:47 | sean-k-mooney | its more if we run out of room with your current effort | |
| 18:52:54 | sean-k-mooney | there are other things we can do with infra | |
| 18:52:57 | sean-k-mooney | but its more involed | |
| 18:53:46 | sean-k-mooney | im not suggesting we start with infra changes just pointing out we can do things via infra changes if its still a proablem | |
| 18:53:47 | dansmith | there's lots we could ask infra to do, but relative to the staffing of the top five projects, I mean.. :) | |
| 18:54:34 | sean-k-mooney | the other thing too is you were just looking at 1st party ci | |
| 18:54:36 | dansmith | I get the impression the things we "could do with infra" require a very wide-scope of potential considerations, more than we can just do in our job defs, and likely would need zuul changes | |
| 18:54:59 | sean-k-mooney | dansmith: yes the infra chagnes are openstack wide | |
| 18:55:08 | sean-k-mooney | requireign both gerrit and zull configuration chagnes | |
| 18:55:17 | sean-k-mooney | so very big/wide reaching hammer | |
| 18:55:48 | sean-k-mooney | not running 8 almost identicaly jobs is relitivly local in contrast | |
| 18:56:01 | dansmith | yeah | |
| 18:56:32 | dansmith | I wish I could help accelerate us not running the two grenades because those are fairly heavy and really duplicative | |
| 18:56:51 | sean-k-mooney | well we can stop that in nova | |
| 18:57:02 | sean-k-mooney | right now if we want too | |
| 18:57:16 | dansmith | I know, but I think we agreed to wait until the ceph and zuulv3 thing was resolved | |
| 18:57:25 | dansmith | I already proposed it with -W to wait on that | |
| 18:58:44 | sean-k-mooney | well mor i ment we can drop integrated-gate-compute template then contol which grenddade jobs run rom the check and gate pipelines | |
| 18:58:51 | sean-k-mooney | https://github.com/openstack/nova/blob/master/.zuul.yaml#L421 | |
| 18:59:04 | sean-k-mooney | which means we coudl jsut run - nova-grenade-multinode | |
| 18:59:21 | dansmith | https://review.opendev.org/c/openstack/tempest/+/771499 | |
| 18:59:43 | dansmith | we agreed we would wait to do that until the ceph multinode zuulv3 thing was resolved | |
| 19:00:15 | sean-k-mooney | oh i know i just was pointing out we could do it via nova if we had resovled it and not need a tempest patch | |
| 19:01:02 | dansmith | and when we discussed, gmann wanted it changed there ^, but yes, the mechanics aren't hard, it's the agreement required, and in this case, blocking on the zuulv3 conversion | |
| 19:01:26 | alexe9191 | good day everyone:) | |
| 19:01:34 | sean-k-mooney | alexe9191: o/ | |
| 19:01:35 | alexe9191 | I have a question about the retry filter in nova | |
| 19:01:50 | alexe9191 | def host_passes(self, host_state, spec_obj): | |
| 19:01:50 | alexe9191 | I am wondering where does it get it's spec_obj from ? specefically this piece of code here: | |
| 19:01:51 | alexe9191 | retry = spec_obj.retry | |
| 19:01:51 | alexe9191 | """Skip nodes that have already been attempted.""" | |
| 19:01:52 | sean-k-mooney | alexe9191: it has not been required for quite some time | |
| 19:02:24 | sean-k-mooney | alexe9191: its passed in by the filter schduler | |
| 19:02:33 | alexe9191 | but where is it stored? memory or db? | |
| 19:02:58 | sean-k-mooney | its the request spec | |
| 19:03:04 | sean-k-mooney | its builts in the api | |
| 19:03:12 | sean-k-mooney | then passed to the conductor and scudler | |
| 19:03:22 | sean-k-mooney | i belive we might have it in the api db | |
| 19:03:35 | alexe9191 | Interesting, let me check | |
| 19:03:43 | dansmith | yes, api_db | |
| 19:04:30 | alexe9191 | ok, so if a host fail, it will register it's state here in that table. | |
| 19:04:32 | alexe9191 | request_specs | |
| 19:04:45 | sean-k-mooney | no | |
| 19:04:57 | sean-k-mooney | we dont commit that to the db | |
| 19:05:21 | sean-k-mooney | we only track that in memroy i belive | |
| 19:05:31 | alexe9191 | that table is quite loaded though. | |
| 19:05:47 | sean-k-mooney | the request_sepc is used for other things | |
| 19:05:53 | sean-k-mooney | so it is saved in the db | |
| 19:06:11 | sean-k-mooney | but wee dont save the failed host in the request spec during schdule and commit that back | |
| 19:06:30 | alexe9191 | so that part right here: | |
| 19:06:31 | alexe9191 | hosts, spec_obj, index) | |
| 19:06:31 | alexe9191 | return self.filter_handler.get_filtered_objects(self.enabled_filters, | |
| 19:07:02 | alexe9191 | i see it's passing hosts but the retry filter is kicking out all of the hosts in the aggregate I am trying to schedule in. They are all healthy and they are hosting VMs, they probably had issues at some time. | |
| 19:07:21 | alexe9191 | The interesting thing, this happens only with one flavor, other flavors are returning a different results for the retryfilter. | |
| 19:08:23 | alexe9191 | I restarted nova-compute on all of those hosts but that did not change the result of the scheduling:) so I am wondering to be honest where does it saves the host state. | |
| 19:08:29 | sean-k-mooney | alexe9191: what releast of nova are you using by the way | |
| 19:08:32 | alexe9191 | rocky:) | |
| 19:08:39 | sean-k-mooney | so you have placment | |
| 19:08:59 | alexe9191 | yes | |
| 19:09:10 | sean-k-mooney | then you can disable the retry filter entirly | |
| 19:09:19 | sean-k-mooney | i belive rocky is the release we stopped using it | |
| 19:09:25 | sean-k-mooney | thats what im checking now | |
| 19:10:12 | sean-k-mooney | ah it was queens https://github.com/openstack/nova/blob/master/releasenotes/notes/deprecate-retry-filter-4d1dba39a2c21836.yaml | |
| 19:10:34 | sean-k-mooney | as part of https://specs.openstack.org/openstack/nova-specs/specs/queens/implemented/return-alternate-hosts.html | |
| 19:10:52 | sean-k-mooney | alexe9191: so on rocky you can and should disable the retry filter | |
| 19:11:22 | alexe9191 | Ok! that's good to know | |
| 19:12:03 | alexe9191 | Is there a way to mitigate the effect of the retry filter right now? restarting the nova scheduler right now is probably something that's gonna cause a lot of grief | |
| 19:12:12 | alexe9191 | We have about 800~ hosts | |
| 19:12:20 | alexe9191 | 9 schedulers | |
| 19:12:33 | sean-k-mooney | really why so manny? | |
| 19:12:49 | alexe9191 | we're thinking about cells but this is in the future plans | |
| 19:13:01 | alexe9191 | it's a big infrastructure | |
| 19:13:24 | alexe9191 | that's why I was wondering if I can empty that spec_obj from a cache/db table | |
| 19:13:34 | sean-k-mooney | still scduling is typeiclaly not the largets part of a but | |
| 19:13:43 | sean-k-mooney | infacti its typeiclly quite a small amount | |
| 19:13:53 | sean-k-mooney | 9 schdulers is quite a lot | |
| 19:14:27 | sean-k-mooney | alexe9191: no the previsouly tired host are only updated in memory | |
| 19:14:51 | alexe9191 | in the nova-scheduler I am guessing | |
| 19:15:02 | sean-k-mooney | yep | |
| 19:15:25 | openstackgerrit | Merged openstack/nova master: Fix invalid argument formatting in exception messages https://review.opendev.org/c/openstack/nova/+/763511 | |
| 19:15:27 | sean-k-mooney | they only are tracked during the singel scheduling request | |
| 19:15:29 | alexe9191 | one more question, is that spec_obj also tied to the flavor ? | |
| 19:15:53 | sean-k-mooney | kind of | |
| 19:16:00 | dansmith | sean-k-mooney: should the retry filter even be used anymore? don't we pass selected_hosts down and have the conductor just iterate over them and then declare it dead? | |
| 19:16:35 | sean-k-mooney | dansmith: we do from queens | |
| 19:16:42 | sean-k-mooney | dansmith: thats whyi said to remove it | |
| 19:16:48 | dansmith | oh I see you said that above | |
| 19:16:57 | sean-k-mooney | dansmith: also i delete it on master a few release ago | |