| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-11-21 | |||
| 13:26:27 | efried | okay. Pretty sure the docs were the first thing stephenfin hit -- but yeah, they wouldn't mention a microversion cutover if we didn't do that. | |
| 13:26:46 | johnthetubaguy | efried: we can go back an add that for sure | |
| 13:26:47 | efried | So you're looking for a patch that creates a new "signal microversion" and updates the docs accordingly. | |
| 13:26:56 | efried | that makes sense ++ | |
| 13:27:07 | johnthetubaguy | yeah, thanks, that clears it up in my head too | |
| 13:27:08 | johnthetubaguy | phew | |
| 13:27:37 | efried | you can still have the interop snafu I mentioned earlier. But at least in that scenario your "signal" is the 410. | |
| 13:28:32 | sean-k-mooney | zigo: the nviad-t4 apparently supprot sriov so you have to set teh device_type to type-PF | |
| 13:28:40 | johnthetubaguy | efried: yeah, its that you could have known better | |
| 13:28:53 | sean-k-mooney | stephenfin: ^ that is the thin you were writing the docs patch for right | |
| 13:29:07 | sean-k-mooney | https://review.opendev.org/#/c/694522/ | |
| 13:29:09 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Specify what RPs _ensure_resource_provider collects https://review.opendev.org/695429 | |
| 13:29:12 | zigo | sean-k-mooney: The issue is with my p1000, the t4 looks like working ... | |
| 13:29:18 | zigo | But thanks, I'll try. | |
| 13:29:32 | gibi | efried: fixed up the comment in _ensure_resource_provider_collects https://review.opendev.org/695429 | |
| 13:29:46 | sean-k-mooney | zigo: what is the issue you are having specificly | |
| 13:30:00 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Force config drive in nova-next multinode job https://review.opendev.org/695431 | |
| 13:30:00 | zigo | sean-k-mooney: PCI alias nvidia-p1000 is not defined (HTTP 400) | |
| 13:30:07 | zigo | (when trying to spawn my instance) | |
| 13:30:21 | gmann | efried: but from nova ussuri onwards, any microversion will be no-nova-net-api for all cloud so interop things needs to be taken care by discoverability of new microversion we will introduce instead of code handle that. | |
| 13:30:27 | sean-k-mooney | have you set the alias on both the compute nodes and the controler nodes | |
| 13:30:35 | zigo | sean-k-mooney: Yeah, I did that ... | |
| 13:30:47 | zigo | I did with puppet, so normally, it will have restart all nova services. | |
| 13:31:41 | sean-k-mooney | and you have it set in the [pci] section not default | |
| 13:32:06 | gmann | efried: RE: provider as 'openstack:' these are log it return from 60_nova/resources.sh:_get_inventory_value for VCPU etc: https://zuul.opendev.org/t/openstack/build/323afb9d5fd94f62b0bba4bac6004442/log/logs/grenade.sh.txt.gz#28851 | |
| 13:32:36 | gmann | and that is heppning when octavia grenade job is moving to py3. i ma not sure how it has to do with py3 htings. | |
| 13:32:38 | gmann | thins | |
| 13:32:54 | zigo | sean-k-mooney: That's all I did yes... | |
| 13:32:54 | gmann | https://review.opendev.org/#/c/693486/ | |
| 13:33:49 | gmann | mriedem: have you seen this before in any grenade job - https://zuul.opendev.org/t/openstack/build/323afb9d5fd94f62b0bba4bac6004442/log/logs/grenade.sh.txt.gz#28854 | |
| 13:33:59 | zigo | Oh, another thing which is very annoying, I keep having in my logs: | |
| 13:33:59 | zigo | Instance 2aab6469-4292-4d04-80de-2ae2a7174b3a has been moved to another host clint1-compute-2.infomaniak.ch(clint1-compute-2.infomaniak.ch). There are allocations remaining against the source host that might need to be removed: {'resources': {'DISK_GB': 80, 'MEMORY_MB': 24576, 'VCPU': 8}}. | |
| 13:34:05 | zigo | Many of this ... | |
| 13:34:29 | zigo | Is this a known issue with Rocky? | |
| 13:34:33 | sean-k-mooney | that usually means you have not compuleted migrations using resize-verify | |
| 13:34:45 | zigo | sean-k-mooney: It's mostly all live migrations. | |
| 13:34:58 | efried | gmann: those CI results are still not loading up for me :( | |
| 13:35:13 | gmann | oh | |
| 13:35:24 | zigo | sean-k-mooney: IMO, the only thing that remains is the placement record... | |
| 13:35:29 | efried | gmann: if you have them open, maybe you could pastebin the relevant chunk? | |
| 13:35:34 | zigo | I could easily write a clean-up script I suppose. | |
| 13:35:34 | sean-k-mooney | efried: https://zuul.opendev.org/t/openstack/build/323afb9d5fd94f62b0bba4bac6004442/log/logs/grenade.sh.txt.gz#28854 loaded for me fine | |
| 13:36:16 | sean-k-mooney | zigo: it should be updated automatically | |
| 13:36:37 | efried | gmann: ah, it's working now. | |
| 13:36:53 | mriedem | gmann: the broken pipe? no | |
| 13:38:34 | gmann | it return 'openstack:' as provider name from 60_nova/resources.sh:_get_inventory_value | |
| 13:39:18 | mriedem | there is a broken pipe right before that | |
| 13:39:34 | mriedem | 2019-11-20 17:52:23.572 | +++ /opt/stack/new/grenade/projects/60_nova/resources.sh:_get_inventory_value:57 : head -n1 2019-11-20 17:52:23.573 | +++ /opt/stack/new/grenade/projects/60_nova/resources.sh:_get_inventory_value:57 : openstack resource provider list -f value 2019-11-20 17:52:23.573 | +++ /opt/stack/new/grenade/projects/60_nova/resources.sh:_get_inventory_value:57 : cut -d ' ' -f 1 2019-11-20 17:52:24.613 | Exc | |
| 13:39:34 | mriedem | on raised: [Errno 32] Broken pipe | |
| 13:39:40 | mriedem | so parsing the output is failing | |
| 13:40:37 | mriedem | openstack resource provider list -f value | |
| 13:40:40 | mriedem | oops | |
| 13:40:42 | efried | gmann: yeah, this is gonna have nothing to do with a provider named 'openstack:'. It looks to me like we're parsing error output from the openstack command (which would start with 'openstack: $something_went_wrong') | |
| 13:40:45 | mriedem | provider=$(openstack resource provider list -f value | head -n1 | cut -d ' ' -f 1) | |
| 13:40:54 | mriedem | that's the command that's fialing | |
| 13:41:00 | mriedem | well, parsing that's failing | |
| 13:41:31 | mriedem | if there was just one provider we could do: | |
| 13:41:40 | mriedem | provider=$(openstack resource provider list -f value -c uuid) | |
| 13:41:56 | mriedem | but if it's a multinode grenade job then there will be more than one and that doesn't work | |
| 13:42:16 | efried | mriedem: well, we should do that anyway, and head -n1 it | |
| 13:42:20 | efried | i.e. don't do the cut | |
| 13:42:32 | efried | not that that would help here, because clearly the command is failing. | |
| 13:42:52 | efried | But why is the error output going to stdout rather than stderr? | |
| 13:43:01 | mriedem | of that provider list command allowed passing a --name for filtering, we could pass the local fqdn to get 1 result back... | |
| 13:43:03 | efried | Whole point of stderr is so exactly this doesn't happen, and you can see what actually went wrong. | |
| 13:43:28 | zigo | mriedem: A much nicer way using http://harelba.github.io/q/ : provider=$(openstack resource provider list --format csv | q -H -d, "SELECT uuid FROM - LIMIT 1") | |
| 13:43:30 | zigo | ;) | |
| 13:43:51 | efried | mriedem: to that point, it looks like we don't care *which* provider we're grabbing? That seems... weird. | |
| 13:43:56 | zigo | (q-text-as-data is such a nice tool...) | |
| 13:43:59 | mriedem | efried: we don't, | |
| 13:44:12 | mriedem | it's a smoke test to make sure that we can save off some inventory before upgrading and that after the upgrade it's still there | |
| 13:44:13 | gmann | pick up anything should be fine | |
| 13:44:25 | efried | okay. | |
| 13:45:31 | zigo | mriedem: Then you could do: provider=$(openstack resource provider list --format csv | q -H -d, "SELECT uuid FROM - WHERE name='something-you-want'") | |
| 13:45:35 | efried | anyway, the problem here seems to be that the openstack command -- the first thing in the pipe -- is failing, printing its error to stdout. | |
| 13:47:32 | mriedem | zigo: or i could just do provider=$(openstack resource provider list -f value -c uuid --name `hostname -f`) | |
| 13:47:44 | mriedem | and not rely on pipes and other tooling | |
| 13:47:58 | mriedem | but that provider list command doesn't support --name (yet - that's easy to add) | |
| 13:48:48 | gmann | API has the name filter ? | |
| 13:48:51 | mriedem | yes | |
| 13:49:06 | efried | one way to get output like 'openstack: $stuff' is if the openstack command doesn't exist. But that output goes to stderr like it should. | |
| 13:50:58 | sean-k-mooney | gmann: its how neutron identifies the compute node resouce provider without needing to known the compute node uuid | |
| 13:51:10 | efried | I can't find anything in the code that's joining stderr to stdout. Unless the job itself is doing that. | |
| 13:51:10 | sean-k-mooney | it looks up the RP by hostname | |
| 13:52:17 | efried | Okay, apparently I'm the only one concerned about the fact that `openstack resource provider list` is producing bogus output in gmann's case, so I must be misunderstanding what we're actually trying to solve here. /me stfu, call if you need me. | |
| 13:53:48 | gmann | sean-k-mooney: and there it is working fine? failure case of octavia grenade job on py3. | |
| 13:54:36 | sean-k-mooney | they do that in code not via osc | |
| 13:54:47 | gmann | ohk | |
| 13:59:09 | sean-k-mooney | is looking up the provider by hostname all that is breakign the job? | |
| 13:59:57 | sean-k-mooney | it would be quick to fix osc-plamcenet to support that but equally quick to just do it with curl | |
| 14:01:32 | gmann | sean-k-mooney: fixing osc might take time with release etc until octavia job can install it from source. | |
| 14:01:54 | sean-k-mooney | do we know what cause the broken pipes? | |
| 14:02:03 | gmann | no. | |
| 14:02:21 | haleyb | sean-k-mooney: would adding osc-placement to requirements in octavia fix it as well? | |
| 14:02:53 | sean-k-mooney | i think mriedem said list does not currently support --name | |
| 14:04:02 | sean-k-mooney | i was looking at the greand job logs by the way gmann haleyb do you have the link to the octavia job | |
| 14:04:33 | johnsom | haleyb Yes, that is the error output you are seeing. Installing osc-placement should fix it. | |
| 14:04:49 | gmann | sean-k-mooney: https://review.opendev.org/#/c/693486/ | |
| 14:04:50 | johnsom | https://www.irccloud.com/pastebin/PrDd7zG6/ | |
| 14:05:08 | openstackgerrit | Merged openstack/nova stable/pike: Delete instance_id_mappings record in instance_destroy https://review.opendev.org/684658 | |