| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-03-12 | |||
| 12:41:59 | openstackgerrit | sean mooney proposed openstack/nova master: Provider Config File: Function to further validate and retrieve configs https://review.opendev.org/676029 | |
| 13:12:33 | openstackgerrit | John Garbutt proposed openstack/nova master: Update quota_class APIs for db and api limits https://review.opendev.org/712143 | |
| 13:17:38 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Ensures that COMPUTE_RESOURCE_SEMAPHORE usage is fair https://review.opendev.org/712674 | |
| 13:17:45 | gibi | melwitt: ^^ | |
| 13:19:31 | gibi | melwitt: I created some test and I'm +2 on Jason's patch | |
| 13:22:25 | masterpe | gibi: I have just rebooted the compute node and search in the nova-compute.log, but I could not find any "Cleaning up allocations" message | |
| 13:22:39 | gibi | masterpe: interesting | |
| 13:25:36 | gibi | masterpe: did you managed to find the allocations in placement either via the CLI or in the db? | |
| 13:27:12 | gibi | masterpe: are the evacuated instances still exists (not deleted) on compute3 ? | |
| 13:27:17 | masterpe | cli does not work, I will search in the DB. | |
| 13:27:43 | masterpe | yes | |
| 13:31:51 | dansmith | gibi: I think it's safe to merge the vmware deprecation package now.. quite a bit of support on it | |
| 13:33:06 | openstackgerrit | Brin Zhang proposed openstack/python-novaclient master: Microversion 2.83: Add volume-patch CLI https://review.opendev.org/712651 | |
| 13:34:21 | brinzhang | dansmith: can you leave some comments in https://review.opendev.org/#/c/694428/3/nova/objects/instance_action.py@201 | |
| 13:34:48 | brinzhang | dansmith: I am confusing of mriedem said in comments ^^ | |
| 13:34:59 | gibi | dansmith: done | |
| 13:35:17 | dansmith | brinzhang: yeah I remember not fully grasping that as well, but I will read and try to grok | |
| 13:35:19 | dansmith | gibi: thanks | |
| 13:35:34 | gibi | dansmith: thanks for proposing the deprecation | |
| 13:35:41 | dansmith | np | |
| 13:36:44 | brinzhang | dansmith: thanks, I know little about serialize_args. | |
| 13:37:55 | masterpe | gibi: yes the instance are still on compute3 | |
| 13:38:59 | gibi | masterpe: was the evacuation successfull? what does nova instance-action-list <uuid> shows regarding the evacuation? | |
| 13:40:09 | lyarwood | sean-k-mooney / stephenfin / kashyap: https://review.opendev.org/#/c/696834/ - would you be able to hit this today if you have time? | |
| 13:41:00 | kashyap | lyarwood: I have already read bits of the patch yesterday, will respond | |
| 13:41:51 | masterpe | nova instance-action-list b4a3a998-e8b4-482c-ae4d-dddaee707ff1 ( the uuid is from a instance ) | |
| 13:41:57 | masterpe | evacuate | req-552e7f33-8f94-4bc3-a992-8f59b1f92f11 | - | 2020-03-06T17:26:14.000000 | |
| 13:42:10 | lyarwood | kashyap: thanks | |
| 13:43:11 | gibi | masterpe: yeah so the evacuation was successful. I have no further idea how https://github.com/openstack/nova/blob/stable/rocky/nova/compute/manager.py#L729 is not triggered. | |
| 13:44:20 | gibi | masterpe: if you find the offending allocations against compute2 in allocations table of the nova_api db then you can deleted them manually. But I have no idea how to prevent them to happen | |
| 13:44:22 | dansmith | brinzhang: so, the thing I don't get about what he said is the "pass a format_exc_val kwarg" which isn't really an option here | |
| 13:44:48 | masterpe | I can not find that code in the file | |
| 13:45:04 | dansmith | brinzhang: but we can make serialize_args pass the cleaned-up exception message-or-type thing for exc_val | |
| 13:48:28 | masterpe | gibi: In my manager.py the commit https://opendev.org/openstack/nova/commit/09de94e39bbcfc7f8130638e73a8248e49cb6ab7 is missing, so probebly the file is from before that. | |
| 13:50:05 | brinzhang | dansmith: what do you mean? Can you elaborate? I still don't quite understand. In other words, I don't know what I should do about @serialize_args | |
| 13:50:14 | gibi | masterpe: ahh. you have 18.2.3 and that patch is in 18.3 | |
| 13:50:17 | dansmith | brinzhang: I'm cooking up a patch for you | |
| 13:50:23 | gibi | masterpe: that explains what you see | |
| 13:50:44 | brinzhang | dansmith: thanks | |
| 13:51:28 | masterpe | when I do select * from allocations where consumer_id = 'b4a3a998-e8b4-482c-ae4d-dddaee707ff1'; | |
| 13:51:28 | masterpe | I see 4 enty's | |
| 13:51:45 | masterpe | the collum resource_provider_id is the compute node? | |
| 13:52:35 | gibi | masterpe: yes, the resource_provider_id is the compute node uuid | |
| 13:52:55 | gibi | masterpe: most probably you have allocation for that instance on two computes, compute2 and compute3 | |
| 13:53:13 | gibi | the bugfix that is missing from your rocky version is fixing that | |
| 13:53:37 | gibi | if you could upgrade to 18.3 and restart the compute that would also fix your deployment as well | |
| 14:03:22 | dansmith | gibi: fwiw, normally I can make the 1600 meeting but not today because I already have a thing I have to do | |
| 14:03:31 | dansmith | er, s/fwiw/fyi/ | |
| 14:03:43 | gibi | dansmith: ack, thanks for the heads up | |
| 14:04:23 | masterpe | The id in resource_providers is linked with resource_provider_id of allocations | |
| 14:05:15 | gibi | masterpe: yes sorry, the resource_provider_id is an internal id pointing to the resource_providers table and there is an uuid in that table that points to the compute node | |
| 14:05:49 | masterpe | resource_provider_id has id 9 | |
| 14:06:24 | masterpe | so I can delete all records in allocations with resource_provider_id = 9? | |
| 14:08:51 | gibi | masterpe: if you evacuated all the servers from compute2 then yes | |
| 14:09:37 | masterpe | yes compute2 is emty | |
| 14:23:05 | dSrinivas | Hi all, while attaching the volume to the Nova intance it is failing with this errorhttp://paste.openstack.org/show/790596/ can anyone please me | |
| 14:25:11 | dSrinivas | For FC not for iscsi | |
| 14:31:18 | sean-k-mooney | so that look like an os_brick issue but have you check if the FC deivce is mounted properly on the host at the expected path | |
| 14:31:54 | dSrinivas | FC devices are Online | |
| 14:32:23 | dSrinivas | It is attaching the volume and again it is deleting after 20seconds | |
| 14:32:45 | sean-k-mooney | and if you run the grep command yourself does it find the deivice | |
| 14:32:51 | dSrinivas | Cinder volume logs it is showing attachment updated successful | |
| 14:32:58 | sean-k-mooney | hum well the deleing it proably part of the clean up | |
| 14:33:47 | dSrinivas | Systool -c fc_host -v | grep Online it is showing the ports as online | |
| 14:34:04 | sean-k-mooney | on the storage backend | |
| 14:34:32 | sean-k-mooney | that does not mean that the compute node was able to connect to it and mount it right? | |
| 14:34:46 | sean-k-mooney | it just means the sotrage backend thing it available to be mounted? | |
| 14:34:51 | dSrinivas | At storage side it is showing attached again after 20se | |
| 14:34:52 | sean-k-mooney | FC is not my thing | |
| 14:35:00 | lyarwood | dSrinivas: which version is this? | |
| 14:35:01 | dSrinivas | Deleting | |
| 14:35:07 | dSrinivas | Master | |
| 14:35:49 | openstackgerrit | Dan Smith proposed openstack/nova master: Make serialize_args handle exception messages safely https://review.opendev.org/712697 | |
| 14:35:52 | lyarwood | yeah sorry no idea then, can you raise this in an os-brick bug and ask for help in #openstack-cinder ? | |
| 14:35:53 | sean-k-mooney | dSrinivas: the only thing that comes to mind is maybe it is timing out waithg for it to be avialble | |
| 14:35:56 | dansmith | brinzhang: mriedem ^ | |
| 14:36:27 | mriedem | ooo | |
| 14:36:31 | dSrinivas | How to increase the retries | |
| 14:38:04 | sean-k-mooney | dSrinivas: im not sure. | |
| 14:38:29 | dSrinivas | It is attaching and then deleting immediately | |
| 14:38:54 | sean-k-mooney | and you can see that on the compute node not the storage backend | |
| 14:39:54 | lyarwood | dSrinivas: ignore the volume state, the underlying block devices aren't present. | |
| 14:40:18 | lyarwood | dSrinivas: raise an os-brick bug for this and reach out to the folks in #openstack-cinder | |
| 14:40:31 | lyarwood | dSrinivas: we don't directly support the os-brick code in this channel | |
| 14:40:34 | openstackgerrit | sean mooney proposed openstack/nova master: Provider Config File: Functions to merge provider configs to provider tree https://review.opendev.org/676522 | |
| 14:41:05 | dSrinivas | I have check the n-cpu logs and pasted the link above | |
| 14:41:40 | lyarwood | dSrinivas: yes, n-cpu is calling os-brick and it is failing to connect the volume here | |
| 14:41:44 | sean-k-mooney | dSrinivas: which is showing an error form os-bricks which is a lib that is maintained by both cinder and nova | |
| 14:42:05 | lyarwood | sean-k-mooney: well we chip in with the odd bugfix ;) | |
| 14:42:16 | sean-k-mooney | yes | |
| 14:42:42 | sean-k-mooney | well you chip in with the odd bugfix i avoid it | |
| 14:44:17 | openstackgerrit | sean mooney proposed openstack/nova master: Provider Config File: Enable loading and merging of provider configs https://review.opendev.org/693460 | |
| 14:44:53 | kplant | anyone know what would cause "No valid host was found. No valid host found for resize " when trying to resize an instance? | |
| 14:45:05 | kplant | i have plenty of resources available and can deploy new instances with the same flavor i'm trying to resize to | |
| 14:45:36 | mriedem | dansmith: left some comments, thanks | |
| 14:47:02 | dansmith | mriedem: thanks | |
| 14:50:09 | openstackgerrit | Takashi NATSUME proposed openstack/python-novaclient master: Microversion 2.82 - nova cyborg interaction https://review.opendev.org/712700 | |
| 14:51:08 | jkulik | kplant, if you only have one host, you'd explicitly have to allow resize on the same host, iirc | |
| 14:51:22 | kplant | i did see that, i have more than one host though | |
| 14:51:44 | jkulik | then you'll have to look into the logs of nova-scheduler I guess | |