| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2021-01-15 | |||
| 10:24:21 | brinzhang0 | gibi: thanks for check the cyborg shelve/unshelve patch ^^ | |
| 11:29:47 | openstackgerrit | Lee Yarwood proposed openstack/nova master: api: Log os-resetState as an instance action https://review.opendev.org/c/openstack/nova/+/770926 | |
| 11:43:24 | gibi | bauzas: are your concerns answered in https://review.opendev.org/c/openstack/nova/+/729563 ? I'm holding the +W until ack it | |
| 12:36:00 | sean-k-mooney | songwenping__: i was suggesting that gets shoudl not create teh mdev ever. the cyborg agent should create the mdevs for all bound arqs on start up. ideally it would set the binding state to "provisioning" on start up also and we would have anohter state "unknown" that the api would report if the cyborg agent misses its heart beat | |
| 12:37:12 | sean-k-mooney | songwenping__: so on start up nova would see 1 of 3 states, bound if the cyborg agent started first and hadn completed bininding, unknown if the cyborg agent has not heartbeat to the cyborg conductor yet | |
| 12:37:36 | sean-k-mooney | or provisioning if the agent is in the process of creating the mdev | |
| 12:38:32 | sean-k-mooney | if cyborg was to send the same bining complete event when it change teh status form provisioning to bound as it does during normal arq binding | |
| 12:38:45 | sean-k-mooney | the process in nova awould be the same | |
| 12:39:34 | sean-k-mooney | e.g. set up event handeler, check if its already bound if not wait for binding complete event, if it si bound cancel the even waiter and proceed with boot | |
| 12:40:55 | sean-k-mooney | gibi: bauzas does that sound like a resonable approch to ye instead of have get arq sometimes do an rpc call to the cybrog agent to create mdevs | |
| 12:41:37 | bauzas | gibi: will look at https://review.opendev.org/c/openstack/nova/+/729563 | |
| 12:42:28 | gibi | bauzas: thanks | |
| 12:42:57 | bauzas | and for the vGPU support in Cyborg, I'll also look at the new revision today | |
| 12:43:41 | gibi | sean-k-mooney: do we have this event handling at nova compute startup for neutron ports too? | |
| 12:44:48 | sean-k-mooney | gibi: we do not need to rebinding them in the neutron case just replug them. | |
| 12:45:18 | sean-k-mooney | in the neutorn case os-vif is the thing that attaches the ports to the network backend in general | |
| 12:45:37 | sean-k-mooney | which si the equivalent of creating the mdev | |
| 12:45:45 | gibi | sean-k-mooney: I see | |
| 12:47:47 | sean-k-mooney | i just find it unsetaling that we would considre change a get form a simple db lookup into an rpc call | |
| 12:48:25 | gibi | sean-k-mooney: so at compute startup nova gets the binding state from cyborg, if it is unknow or provisioning then keep the guest power state off but set up an event waiter. If the arq state is "OK" then nova would start the guest during compute startup. | |
| 12:48:45 | sean-k-mooney | gibi: not quite | |
| 12:48:50 | gibi | correct me please | |
| 12:49:30 | sean-k-mooney | i was thinking we would always set up the waiter and early out if it was bound like we do for normal spwan | |
| 12:50:02 | gibi | can nova we loose an event during the compute reboot? | |
| 12:50:29 | gibi | if yes then that guest would be stuck waiting of the event that was sent by cyborg but lost in the comptue restart | |
| 12:51:03 | gibi | if the event is never lost then I'm OK to wait for the event | |
| 12:51:05 | sean-k-mooney | the event would hit the api and then be enqued to the compute node topic queue | |
| 12:51:15 | sean-k-mooney | so i dont think it would be lost | |
| 12:51:15 | gibi | sean-k-mooney: cool | |
| 12:51:20 | gibi | that seem OK | |
| 12:51:34 | gibi | hm | |
| 12:52:12 | gibi | so in this case the sending the event is triggered by cyborg agent restart, in any other case sending the event is triggered by a nova bind request | |
| 12:52:39 | sean-k-mooney | yes | |
| 12:53:09 | gibi | so there are extra cases to handle. 1) a single cyborg agent restart will send events and if the compute service was not restarted then these events needs to be consumed but ignored | |
| 12:53:14 | sean-k-mooney | we could just call bind if we wanted too and not require teh cyborg agent to auto create them but i think the auto create would be more efficent | |
| 12:54:00 | sean-k-mooney | gibi: we have unexpeted event handeling in nova already | |
| 12:54:05 | gibi | cool | |
| 12:54:19 | sean-k-mooney | if we dont have a waiter when we deque it we just log it and discard | |
| 12:54:28 | gibi | that seems OK too then | |
| 12:54:34 | sean-k-mooney | which si ok because we will check the state when we get to that part of the code | |
| 12:55:44 | sean-k-mooney | basically im just suggesting using the exact saem event system we use of inital sapwn after where we start the binidng in the conductor then wait for it with an early out in the compute | |
| 12:56:03 | gibi | OK, I don't have a #2 actually :) | |
| 12:56:06 | sean-k-mooney | but in this case the cyborg agent would start the bind on start up | |
| 12:56:25 | gibi | sean-k-mooney: so far what you suggest feels OK to me | |
| 12:57:25 | sean-k-mooney | ill find the time stamp for this and add it to the reveiw | |
| 12:57:31 | gibi | cool | |
| 12:57:32 | gibi | thanks | |
| 12:57:55 | sean-k-mooney | bauzas: if you have time to read scool back and find anything concering with that please chime in | |
| 13:26:56 | bauzas | sean-k-mooney: looking | |
| 13:33:44 | bauzas | sean-k-mooney: are you talking about creating the mdevs in sysfs or binding them to the instance by modifying the XML ? | |
| 13:35:54 | sean-k-mooney | bauzas: sysfs | |
| 13:36:01 | bauzas | ack | |
| 13:36:16 | bauzas | if so, I agree, Cyborg should create them | |
| 13:36:22 | sean-k-mooney | specifically cyborg creating them | |
| 13:36:28 | bauzas | (the agent) | |
| 13:36:43 | sean-k-mooney | right but it shoudl do it automatically rather then as a result of a GET ot ARQ show | |
| 13:37:23 | sean-k-mooney | well GET /ARQ/<uuid> | |
| 13:55:34 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: WIP/DNM libvirt: Start emitting DeviceRemovedEvent and DeviceRemovalFailedEvent events https://review.opendev.org/c/openstack/nova/+/749929 | |
| 13:55:35 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: DNM try to replace retry with libvirt event in detach https://review.opendev.org/c/openstack/nova/+/770246 | |
| 13:59:34 | bauzas | sean-k-mooney: yeah, provisioning them directly | |
| 14:00:00 | bauzas | once the operator modifies the config | |
| 14:00:22 | sean-k-mooney | what config? | |
| 14:00:40 | bauzas | their own config for telling which vgpu type for each pGPU | |
| 14:00:45 | sean-k-mooney | we are talking about recreating them after a host reboot | |
| 14:01:14 | sean-k-mooney | the cyborg spec currently say after a host reboot when we reboot the instance that we will do a arq show | |
| 14:01:25 | bauzas | I haven't seen it | |
| 14:01:26 | sean-k-mooney | and that show will do an rpc to the agent to create the mdev | |
| 14:01:28 | bauzas | if so, -1 for me | |
| 14:01:59 | sean-k-mooney | im suggesting that instead on start up the agent shoucl check what instance are on the current host and ensure there mdevs exist | |
| 14:02:15 | sean-k-mooney | and that arqs should have 2 newe states | |
| 14:02:25 | sean-k-mooney | unknon meaning the agent missed it heart beat | |
| 14:02:35 | sean-k-mooney | and provisioning meaning its currenlty seting up the mdev | |
| 14:03:18 | sean-k-mooney | so when we do the show arq binding call if its in provisioning or unknow we wait for the async event | |
| 14:03:27 | sean-k-mooney | if its in bound we know cyborg is done and we proceed | |
| 14:10:53 | bauzas | sean-k-mooney: then I agree with you | |
| 14:11:21 | bauzas | have you provided those comments in the spec ? | |
| 14:11:43 | sean-k-mooney | yes although not that cohently initally so i reference the irc logs above | |
| 14:11:59 | sean-k-mooney | i had give that feedbac in patch set 7 or 9 too | |
| 14:22:30 | bauzas | ++ | |
| 14:22:41 | bauzas | I'll then add my comments then too | |
| 14:39:41 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/train: trivial: Resolve (most) flake8 3.x issues https://review.opendev.org/c/openstack/nova/+/770943 | |
| 14:39:42 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/train: WIP/DNM Switch to hacking 2.x https://review.opendev.org/c/openstack/nova/+/770944 | |
| 14:47:45 | stephenfin | lyarwood: Does this need a blueprint, I wonder? It is kind of feature'ish https://review.opendev.org/c/openstack/nova/+/770926 | |
| 15:03:28 | lyarwood | stephenfin: likely, I was procrastinating this morning and wrote that without thinking | |
| 15:07:14 | lyarwood | gibi: ^ re this, should I create a blueprint for this? | |
| 15:08:33 | gibi | lyarwood: yeah, a bp would be good for that it is adding a feature basically | |
| 15:09:20 | gibi | or we can phrase the whole thing as a bug | |
| 15:09:34 | gibi | we forget to log resetState into the instance actions | |
| 15:09:40 | gibi | so meh, either a bp or a bug would be good | |
| 15:10:05 | gibi | no structural API impact so definetly not a spec | |
| 15:19:55 | lyarwood | gibi: ack let me spin this into a bug | |
| 15:20:04 | gibi | works for me | |
| 15:21:25 | openstackgerrit | Stephen Finucane proposed openstack/nova master: api-ref: Clarify 'all_tenants' command https://review.opendev.org/c/openstack/nova/+/770947 | |
| 15:21:38 | stephenfin | That's the easiest "bugfix" anyone will see this week ^ | |
| 15:24:39 | gibi | stephenfin: +@ | |
| 15:24:41 | gibi | stephenfin: +2 | |
| 15:24:58 | stephenfin | thanks :) | |
| 15:30:29 | dansmith | ahh, the rare but coveted +@ vote | |
| 15:30:36 | openstackgerrit | Lee Yarwood proposed openstack/nova master: api: Log os-resetState as an instance action https://review.opendev.org/c/openstack/nova/+/770926 | |
| 15:39:05 | gibi | dansmith: :) | |