| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-02-08 | |||
| 14:54:36 | mriedem | i want to get this bug fixed and move on | |
| 14:54:38 | bauwser | hah | |
| 14:55:05 | bauwser | anyway, I need to go to a meeting | |
| 14:55:10 | bauwser | :/ | |
| 15:21:25 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Trim fake_deserialize_context in test_conductor https://review.openstack.org/635859 | |
| 15:21:27 | mriedem | another subunit parser thing ^ based on http://logs.openstack.org/27/619527/16/gate/openstack-tox-lower-constraints/cd2889d/job-output.txt.gz#_2019-02-08_12_32_26_104934 | |
| 15:26:34 | mriedem | dansmith: so wrt upgrades we just assert that conductor always goes before everything else yeah? and api is last. which is why if we're making changes to both api and conductor at the same time w/o changing rpc method interfaces, but like through fields on the RequestSpec (behavior changes), we aren't doing any strict version checking in the API to make sure conductor is new enough to handle that behavior change | |
| 15:26:41 | mriedem | we just assume that conductor will be upgraded before api | |
| 15:27:06 | dansmith | mriedem: api isn't last, compute is last | |
| 15:27:08 | mriedem | unlike how the api will check min versions on computes for behavior | |
| 15:27:19 | dansmith | mriedem: we pretty much have to assume that conductor, scheduler, and api all go at the same time | |
| 15:27:19 | mriedem | ok but conductor is first | |
| 15:27:27 | mriedem | yeah | |
| 15:27:53 | mriedem | was just thinking about https://review.openstack.org/#/c/635668/2/nova/conductor/tasks/migrate.py@200 | |
| 15:28:07 | mriedem | which is changes to conductor based on changes the api will make to the request spec, which is all assumed to be upgraded at the same time | |
| 15:28:30 | dansmith | yeah | |
| 15:28:48 | mriedem | i must be having some sort of existential crisis this morning | |
| 15:29:22 | dansmith | mriedem: go to your happy place | |
| 15:29:40 | spatel | sean-k-mooney: morning | |
| 15:30:33 | sean-k-mooney | spatel: o/ | |
| 15:30:43 | spatel | openstack hypervisor stats show command " count | 179" is this number of hypervisor right? | |
| 15:31:09 | melwitt | mriedem: looking at https://review.openstack.org/635859, is it that the assertions are constantly failing and being swallowed? is it bad that the assertions were failing? | |
| 15:32:29 | spatel | why hypervisor stats showing count 179 & "openstack hypervisor list" showing 193 | |
| 15:32:42 | spatel | which one i should consider right? | |
| 15:32:44 | spatel | sean-k-mooney: ^ | |
| 15:33:25 | spatel | http://paste.openstack.org/show/744746/ | |
| 15:36:49 | mriedem | melwitt: yes and idk | |
| 15:37:01 | mriedem | i do'nt know where the 'fake_user' is coming from b/c it's not coming from anything in test_conductor.py that i can tell | |
| 15:40:08 | mriedem | i assume that any test which really cares about the context being serialized/deserialized is checking that explicitly, otherwise we use fake user/project ids all the time | |
| 15:41:49 | sean-k-mooney | spatel: i know you had that issue in the past but i can rememebr what it was again | |
| 15:42:29 | melwitt | mriedem: yeah, fair enough | |
| 15:59:22 | spatel | i think stats output doing something which list not doing.. that is why its difference | |
| 16:08:01 | sean-k-mooney | spatel: the hypervior stats i think only show running vms | |
| 16:08:33 | sean-k-mooney | nova list will also show shelved instnaces and ones in error or power off state | |
| 16:08:47 | spatel | if that is true then it make sense | |
| 16:09:14 | spatel | let me drill and see | |
| 16:11:04 | sean-k-mooney | oh wait you are not asking about nova list | |
| 16:11:27 | sean-k-mooney | you are asing about count in openstack hypervisor stats show vs the outpu of openstack hypervior list | |
| 16:11:47 | spatel | yes count | |
| 16:11:57 | spatel | now you are on same page :) | |
| 16:12:13 | spatel | sean-k-mooney: ^^ | |
| 16:16:28 | sean-k-mooney | spatel: try openstack hypervisor list -c "Host IP" -c "Hypervisor Type" -f value | grep QEMU | sort | uniq | wc -l | |
| 16:17:00 | spatel | 193 | |
| 16:17:33 | sean-k-mooney | ok i was wondering if you had duplicate entries | |
| 16:17:57 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Set the conductor indirection API when running nova-metadata under uwsgi https://review.openstack.org/635577 | |
| 16:18:10 | spatel | i verified already basic stuff | |
| 16:18:41 | spatel | i think openstack hypervisor stats show ( only showing hypervisor in use ) | |
| 16:19:14 | spatel | let me do experiment and see... let me empty or spin up vm and see count change or not | |
| 16:22:27 | sean-k-mooney | spatel: well teh api doces say it should the number of hyperviors | |
| 16:22:29 | sean-k-mooney | https://developer.openstack.org/api-ref/compute/?expanded=show-hypervisor-statistics-detail#show-hypervisor-statistics | |
| 16:22:47 | sean-k-mooney | so if it the number of active or in use hyperviors there is at least a docs bug | |
| 16:23:39 | spatel | let me prove it and open bug | |
| 16:23:53 | spatel | did you try to test these command in your cloud? | |
| 16:24:08 | mriedem | dansmith: melwitt: i'm +2 on the bottom 3 down cell changes in that series now https://review.openstack.org/#/c/635120/ | |
| 16:24:10 | sean-k-mooney | yes and they are the same | |
| 16:24:23 | mriedem | splitting those out from the big patch was definitely the way to go | |
| 16:24:24 | sean-k-mooney | ill disabel a compute service and see if it changes | |
| 16:25:43 | sean-k-mooney | spatel: ok i can repoduce | |
| 16:25:59 | sean-k-mooney | if the comptue node is disabled hypervior stats does not report it | |
| 16:26:19 | sean-k-mooney | so that is a bug either in the code or in the docs | |
| 16:26:54 | spatel | oh!!! | |
| 16:27:32 | melwitt | mriedem: ack | |
| 16:28:20 | spatel | sean-k-mooney: you are right!! i have around 10 or 12 disabled compute node | |
| 16:28:44 | sean-k-mooney | spatel: what release are you running again? | |
| 16:28:50 | spatel | pike | |
| 16:28:54 | spatel | sorry queens | |
| 16:29:35 | sean-k-mooney | i am not sure if we changed it in rocky or queens but we used to auto disable compute nodes after 3-5 failed builts without a successful build | |
| 16:29:49 | sean-k-mooney | if you manually did not disable them that coudl be the cause | |
| 16:30:35 | spatel | i didn't remember when and who did disable it.. but at least clear that stats not showing disable | |
| 16:31:04 | spatel | Do you know how to list disable hypervisor from command line? | |
| 16:32:57 | dansmith | mriedem: okay cool, glad to see that, I'll look a bit later | |
| 16:47:25 | sean-k-mooney | spatel: openstack compute service list | |
| 16:48:13 | spatel | thanks | |
| 16:51:47 | spatel | sean-k-mooney: 179 + 14 (disable) = 193 | |
| 16:52:09 | spatel | I got my answer :) thanks | |
| 16:55:47 | sean-k-mooney | spatel: can you file a bug just so we dont forget | |
| 16:56:09 | sean-k-mooney | as i said we either should update the docs or the code | |
| 17:02:09 | spatel | i think we should update code and add one more row ( disabled compute node ) ;) | |
| 17:21:00 | mriedem | ew another subunit parser giant log capture http://logs.openstack.org/27/619527/16/check/openstack-tox-py35/ab8a233/job-output.txt.gz#_2019-02-08_14_57_03_415588 | |
| 17:21:41 | mriedem | gross that's from o.vo | |
| 17:22:06 | mriedem | so i suppose we need to set log levels to debug for oslo.versionedobjects and oslo.messaging in our test runs | |
| 17:22:16 | mriedem | s/debug/warning/ | |
| 17:22:35 | mriedem | although for oslo.messaging we're tracing exceptions in those conductor unit tests... | |
| 17:22:44 | sean-k-mooney | mriedem: one question do we know what the cause is? i was wondering if it could be a locale issue or something like that | |
| 17:24:12 | melwitt | sean-k-mooney: I don't think so. see the comments in this earlier bug for more info https://bugs.launchpad.net/cinder/+bug/1728640 | |
| 17:24:14 | openstack | Launchpad bug 1728640 in Cinder "py35 unit test subunit.parser failures" [Critical,Fix released] - Assigned to Sean McGinnis (sean-mcginnis) | |
| 17:25:08 | mriedem | i think the summary is we're sending a shit load of content to the subunit output stream capture which blows it up | |
| 17:25:52 | sean-k-mooney | ya that was the other ting i was wondering but i guess that means we are not mocking out the loggers enough in the unit tests | |
| 17:26:14 | melwitt | yeah. from mtreinish: "we may be exceeding the max attachment size in subunit" | |
| 17:26:35 | melwitt | (from a comment in that bug) | |
| 17:27:21 | sean-k-mooney | on the plus side if we do reduce the log output it will both help with space on logs.openstack.org and maybe speed up the tests | |
| 17:27:36 | sean-k-mooney | all that io is proably slowing them down | |
| 17:34:38 | mriedem | well normally i don't think you get a lot of this output unless a test fails | |
| 17:34:52 | mriedem | or the subunit parser blows up | |
| 17:35:38 | openstackgerrit | Elod Illes proposed openstack/nova stable/queens: Handle IndexError in _populate_neutron_binding_profile https://review.openstack.org/635897 | |
| 17:59:31 | mnaser | mriedem, cdent, dansmith: https://review.openstack.org/#/c/635852/ has an openstack ansible cross repo job (which you can see installs from the zuul cloned nova -- http://logs.openstack.org/52/635852/1/check/nova-openstack-ansible-cross-repo/9738be9/logs/ara-report/result/113b6602-4a28-4512-964b-4174593eb507/ "Processing /home/zuul/src/git.openstack.org/openstack/nova") | |
| 17:59:55 | mnaser | i will add another tox env which will run it making sure that we *skip* placement deploy to merge that eventually | |
| 18:00:15 | mnaser | and then the normal functional test will run *with* out-of-repo placement | |
| 18:00:26 | cdent | awesome, thanks | |
| 18:02:53 | cdent | I'll look more closely soon, probably monday | |