| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-11-22 | |||
| 17:08:33 | gregwork | i was at that summit, i wish i attended that session | |
| 17:08:39 | gregwork | ill check it out | |
| 17:09:08 | gregwork | wait cross_az_attach defaults to true | |
| 17:09:16 | gregwork | at least thats how it looks? | |
| 17:09:47 | gregwork | "By default there is no availability zone restruction on volume attach." | |
| 17:09:51 | mriedem | correct | |
| 17:10:33 | gregwork | i see in cinder that the volumes i wanted to attach were created, but they just didnt attach | |
| 17:10:37 | gregwork | also heat didnt throw an error | |
| 17:10:39 | mriedem | read https://docs.openstack.org/nova/latest/admin/availability-zones.html#resource-affinity | |
| 17:12:34 | gregwork | mriedem: i read that but since my cross_az_attach is unset (defaulted to true) that doesnt explain why the volume silently didnt attach. it seems to list the caveats with setting it to false | |
| 17:17:10 | mriedem | are you talking about attaching volumes to an existing server or boot from volume where nova creates the volumes and attaches them to the server being created? | |
| 17:17:34 | openstackgerrit | Stephen Finucane proposed openstack/nova master: trivial: Resolve (most) flake8 3.x issues https://review.opendev.org/695732 | |
| 17:17:34 | openstackgerrit | Stephen Finucane proposed openstack/nova master: WIP: Switch to flake8 3.x https://review.opendev.org/695733 | |
| 17:19:22 | mriedem | if cross_az_attach=True (default) then nova won't specify any AZ when creating volumes during boot from volume, | |
| 17:19:36 | mriedem | if the attach is failing, it's likely due to something with the setup/storage backend etc | |
| 17:19:42 | mriedem | so you'd have to dig into that | |
| 17:21:34 | stephenfin | have a good weekend | |
| 17:24:05 | openstackgerrit | Merged openstack/nova stable/train: Added openssh-client into bindep https://review.opendev.org/691808 | |
| 17:51:03 | openstackgerrit | Merged openstack/nova stable/stein: Use admin neutron client to query ports for binding https://review.opendev.org/694665 | |
| 17:51:10 | openstackgerrit | Merged openstack/nova master: Abort live-migration during instance_init https://review.opendev.org/678016 | |
| 17:59:07 | gregwork | mriedem: so you were right, it was something failing with cinder, however (and this is probably a bug) the volume attachment would fail because nova running "multipathd show status" failed (multipathd was not started on the compute node). the issue is that neither heat, nor nova, nor manually doing it in horizon caught the exception. | |
| 17:59:18 | gregwork | it was only when i looked at nova-compute.log on the compute node that i saw the traceback | |
| 17:59:31 | gregwork | im using queens so maybe this is fixed upstream | |
| 18:00:00 | gregwork | just took the request and happily said it was done | |
| 18:13:57 | sean-k-mooney | well on one hand managing the lifecycle of multipathd is out of scope of nova | |
| 18:15:07 | sean-k-mooney | im not sure if os-brick would have detechted the fact that multipathd was not running or if nova would to be able to raise it and im not sure if we would want to expose that to a non admin if it did | |
| 18:15:27 | sean-k-mooney | heat cant know this unless cinder/nova told it as heat has no acess to the compute node | |
| 18:15:59 | gregwork | sean-k-mooney: well from the end user experience take horizon for example.. doing volumes -> manage attachments and picking my instance .. that should throw an error | |
| 18:16:18 | gregwork | but it doesnt, it accepts the request and nothing happens | |
| 18:16:28 | gregwork | a user should be informed what they asked for failed at least | |
| 18:16:33 | gregwork | maybe not the precise details | |
| 18:16:49 | sean-k-mooney | well it shoudl go to an error state or not be marked as attached unless it completes | |
| 18:17:08 | gregwork | but it doesnt which is why im mentioning it here | |
| 18:17:37 | sean-k-mooney | so nova could do two things at that point | |
| 18:17:46 | sean-k-mooney | it could put the instance in an error state | |
| 18:17:53 | sean-k-mooney | or it could role back the attachment | |
| 18:17:55 | gregwork | (to be clear it doesnt mark it as attached, it just returns an unmodified horizon screen) | |
| 18:18:15 | sean-k-mooney | ok well that is different | |
| 18:18:23 | sean-k-mooney | what is the volume status | |
| 18:18:33 | gregwork | available | |
| 18:18:42 | gregwork | it doesnt go into an error state | |
| 18:18:44 | sean-k-mooney | then it sound like its doing the right thing | |
| 18:18:52 | gregwork | its like you didnt do anything at all | |
| 18:19:06 | sean-k-mooney | yes because it tried to attach and failed | |
| 18:19:07 | melwitt | did you check the instance actions api? :P | |
| 18:19:17 | sean-k-mooney | i was just going to mention that | |
| 18:19:28 | melwitt | I won! | |
| 18:19:33 | melwitt | jk | |
| 18:19:37 | sean-k-mooney | that is where the error if it was going to be reported would likely end up | |
| 18:19:47 | gregwork | in cinder or nova ? | |
| 18:19:51 | melwitt | nova | |
| 18:20:39 | melwitt | I don't know how it looks on horizon, but it's like 'openstack server event list' on the openstackclient | |
| 18:21:40 | gregwork | so i cant openstack server event list <that instance> because we fixed it just now (configuring multipathd) and redeployed the stack | |
| 18:21:40 | sean-k-mooney | in horizon there is the action log for the instace | |
| 18:21:47 | gregwork | where would it be in the infra | |
| 18:21:51 | sean-k-mooney | when you click into it | |
| 18:21:52 | gregwork | compute or control plane | |
| 18:22:04 | gregwork | the logs you are looking for | |
| 18:22:42 | sean-k-mooney | gregwork: if you have not deleted the instnace it would still be in the instance event log | |
| 18:22:54 | melwitt | by redeployed do you mean wiped the database? yeah, what sean-k-mooney said | |
| 18:22:57 | gregwork | it was deleted when the stack was cleaned up | |
| 18:23:07 | gregwork | so all i have are the overcloud logs on the various nodes | |
| 18:23:18 | melwitt | I thought you could still see for deleted no? because one of the actions will be 'deleted' | |
| 18:23:29 | melwitt | that's how you can check who deleted an instance and what time, for example | |
| 18:23:42 | sean-k-mooney | if you have not done a db archive/purge it should still be there | |
| 18:23:53 | gregwork | id need the uuid of the instance wouldnt i ? the instances were redployed from a templated name | |
| 18:23:59 | gregwork | or is there a list all events | |
| 18:24:19 | sean-k-mooney | yes you would | |
| 18:25:03 | melwitt | you need uuid | |
| 18:25:30 | sean-k-mooney | you would be hitting this api endpoint https://docs.openstack.org/api-ref/compute/?&expanded=list-actions-for-server-detail#list-actions-for-server | |
| 18:26:31 | sean-k-mooney | speaking of which is the volume attachmet api call blocking or async https://docs.openstack.org/api-ref/compute/?&expanded=attach-a-volume-to-an-instance-detail#attach-a-volume-to-an-instance | |
| 18:27:11 | sean-k-mooney | it returns a 200 so i assume its blocking but i expected it to be async | |
| 18:27:36 | melwitt | gregwork: ok, well, just fyi the "action log" 'openstack server event list' is a place to see failed actions that did not actually touch the instance. and there's a new change proposed to add more detail to errors that go in there, to aid in communication to the user https://review.opendev.org/694428 | |
| 18:28:14 | efried | melwitt or mriedem or dansmith: Could I please get +A on three simple test-only tech debt reductions (johnthetubaguy just +2ed 'em so they're freshhhh): https://review.opendev.org/#/q/topic:use-base-test-code+(status:open+OR+status:merged) | |
| 18:28:30 | efried | johnthetubaguy: thanks for that btw | |
| 18:28:30 | melwitt | gregwork: this came up recently with a change in behavior in the resize and migration APIs too http://lists.openstack.org/pipermail/openstack-discuss/2019-November/011040.html | |
| 18:29:57 | sean-k-mooney | melwitt: ya although in that context that is because we changed it to async so you have to check the action log going forward | |
| 18:30:22 | sean-k-mooney | melwitt: if attach is currently blocking the we woudl expect the error to be returned there | |
| 18:30:33 | melwitt | yeah. I'm just saying it brought the action log to my attention as a thing to use more regularly | |
| 18:30:33 | sean-k-mooney | if its async then it would only be in the action log | |
| 18:30:46 | sean-k-mooney | yep | |
| 18:30:58 | melwitt | because prior to that, the only time I've seen action log used is when you're like "who deleted this instance and on what date" | |
| 18:31:29 | sean-k-mooney | did that come up on the mailing list | |
| 18:31:45 | sean-k-mooney | the fact the userid was missing or somthing like that | |
| 18:31:46 | melwitt | and I know, some other clouds use it on a regular basis, just in yahoo when I was there, they didn't. so maybe other people don't | |
| 18:31:57 | gregwork | so there is an instance id in the python traceback of nova-compute.log | |
| 18:32:06 | gregwork | failed to attach xxx to for id | |
| 18:32:10 | gregwork | let me try that | |
| 18:34:40 | gregwork | im guessing the instance id in this line is not actually the instance uuid ? | |
| 18:34:44 | gregwork | 2019-11-22 12:26:56.077 8 ERROR nova.compute.manager [req-a25aba4e-952f-431e-bb85-1e2a45f62e3b d16f2bc26e91f8221042c91c61d0c4f7e7ae57a09507d54809f0279e3d6ee29e f4e6856a85b14ecc99b494159f28e425 - 424dfc884bcd46e7915e62495d6b569c 424dfc884bcd46e7915e62495d6b569c] [instance: 1e65006f-58ee-470d-9111-522d5c63094d] Failed to attach 176cca17-3f02-4fcd-884d-a4bd74ad5943 at /dev/vdb: ProcessExecutionError: Unexpected | |
| 18:34:45 | gregwork | error while running command. | |
| 18:34:56 | gregwork | doing an openstack server even list 1e65006f-58ee-470d-9111-522d5c63094d | |
| 18:34:59 | gregwork | says its unknown | |
| 18:35:15 | gregwork | *event | |
| 18:35:31 | sean-k-mooney | am you might need to pass a --deleted flag | |
| 18:35:35 | sean-k-mooney | if we support that | |
| 18:36:04 | gregwork | there is no --deleted flag | |
| 18:36:11 | gregwork | at least not in my osc | |
| 18:36:29 | sean-k-mooney | looking at the api i dont see one either | |
| 18:36:50 | gregwork | melwitt: seemed to think this was a thing, hence the "show when something was deleted" | |