| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-12-02 | |||
| 16:32:16 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Implement reschedule logic for cross-cell resize/migrate https://review.opendev.org/696213 | |
| 16:32:16 | openstackgerrit | Matt Riedemann proposed openstack/nova master: WIP: Add negative test to delete server during cross-cell resize claim https://review.opendev.org/688832 | |
| 16:32:59 | openstackgerrit | Stephen Finucane proposed openstack/nova master: docs: Clarify configuration steps for PF devices https://review.opendev.org/694522 | |
| 16:33:03 | stephenfin | Sure, done | |
| 16:33:26 | stephenfin | efried: Some real nice docs there ^ :) | |
| 16:35:28 | efried | thanks, and noted. | |
| 16:36:06 | dansmith | efried: did you see my comment about the event getting 422? | |
| 16:36:14 | efried | not yet | |
| 16:36:42 | dansmith | okay I don't want sundar to be held up on a decision there or anything | |
| 16:37:04 | efried | reading now | |
| 16:38:41 | efried | dansmith: oh, that's a good point. If cyborg gets this particular 422 it should mean the bind finished before we even got to the host, so the poll inside wait_for_instance_event will trigger and prevent us from waiting (for an event that won't come) | |
| 16:39:16 | dansmith | efried: right, if it gets a 404 then it should fail the bind, but if it gets 422, it shouldn't freak out | |
| 16:39:25 | efried | ++ | |
| 16:39:37 | dansmith | in fact, 422 confirms, in a weird way, that the instance is still there, just not on a host yet | |
| 16:40:16 | dansmith | so if you could chuck that ++ into the review so sundar isn't waiting that would be good | |
| 16:43:30 | efried | dansmith: done | |
| 16:43:52 | dansmith | efried: spanks | |
| 16:44:05 | efried | comfortable and slimming | |
| 16:44:07 | dansmith | efried: have you chatted with him to know that he's responning? | |
| 16:44:11 | dansmith | *respinning | |
| 16:44:18 | efried | haltingly | |
| 16:44:30 | dansmith | haltingly? | |
| 16:44:41 | efried | think voicemail tag | |
| 16:44:45 | efried | but with slack | |
| 16:46:07 | dansmith | I, uh, will trust you | |
| 16:49:53 | stephenfin | dansmith: Can I remove remotable methods on objects if there are no callers left in-tree | |
| 16:50:11 | stephenfin | Or are they versioned since they're RPC'y | |
| 16:50:26 | dansmith | stephenfin: not according to the letter of the law | |
| 16:50:27 | dansmith | and yes, they're RPC by virtue of being remotable | |
| 16:50:43 | dansmith | stephenfin: I would just replace their bodies with raise NotImplementedError() | |
| 16:51:16 | stephenfin | Yeah, this is for Network.(create|destroy|associate|disassociate) | |
| 16:51:19 | stephenfin | I'll do just that | |
| 16:51:24 | dansmith | aye | |
| 16:59:27 | efried | stephenfin: +2 on that docs patch, but a pile of nits if you're so inclined. | |
| 17:00:24 | efried | gibi or stephenfin: are you on board with https://review.opendev.org/#/c/695985/ (ability to cancel instance events)? | |
| 17:01:21 | melwitt | stephenfin: did you see my note here that other quotas related to nova-network will be able to be removed? https://review.opendev.org/#/c/686812/7/nova/api/openstack/compute/quota_classes.py@38 | |
| 17:01:23 | stephenfin | That's way too much context to load up at 5pm /o\ I can hit it in the morning though | |
| 17:01:39 | efried | ack, thx | |
| 17:02:00 | stephenfin | melwitt: Sure did. It's on the list of stuff still to clean up | |
| 17:02:23 | melwitt | stephenfin: k, coolness | |
| 17:10:45 | Sundar | Responded to the discussion on exit_wait_early for instance events in https://review.opendev.org/#/c/631244 | |
| 17:12:33 | Sundar | efried: stephenfin: ^ relates to https://review.opendev.org/#/c/695985/ | |
| 17:12:34 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/stein: libvirt: Add a rbd_connect_timeout configurable https://review.opendev.org/669167 | |
| 17:12:36 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/rocky: libvirt: Add a rbd_connect_timeout configurable https://review.opendev.org/669168 | |
| 17:12:39 | openstackgerrit | Lee Yarwood proposed openstack/nova stable/queens: libvirt: Add a rbd_connect_timeout configurable https://review.opendev.org/669169 | |
| 17:12:46 | lyarwood | melwitt: ^ reopened and rebased | |
| 17:16:48 | melwitt | lyarwood: ack thanks | |
| 17:19:21 | mriedem | lyarwood: did anyone that reported that issue actually say the fix resolved it for them? | |
| 17:19:39 | lyarwood | mriedem: we've just sent a test build out to a stable/queens user | |
| 17:20:14 | lyarwood | mriedem: it's going to fix their issues but will at least confirm that their network is borked | |
| 17:31:08 | efried | Sundar: back atcha. To be clear, I'm not seeing a problem; are you? | |
| 17:33:36 | Sundar | efried: The notification from Cyborg will fail in many/most cases; ignoring that error could be considered a hack, rather than a solution. After all, folks looking at the n-api logs will see an error. Are we ok with that? | |
| 17:33:57 | sean-k-mooney | gibi: this is the bug for the issue with live migrating https://bugs.launchpad.net/nova/+bug/1854844. ill try to think about ways to adress this in a backportable way. i think i know of one way but this might just be something we have to fix on master only | |
| 17:33:58 | openstack | Launchpad bug 1854844 in OpenStack Compute (nova) "libvirt: tx/rx queue lenght and max queues are not updtated on live migration" [Low,Triaged] - Assigned to sean mooney (sean-k-mooney) | |
| 17:34:28 | efried | Sundar: Only if they're looking at debug logs :) | |
| 17:35:00 | efried | The cyborg logs should of course explain what's happening as well. | |
| 17:35:17 | efried | And yes, I'm fine with this. | |
| 17:37:49 | Sundar | efried: From Cyborg's POV, the notification after binding would be best-effort: ignore any failures and complete the binding. The Cyborg logs will indicate as much. | |
| 17:38:25 | sean-k-mooney | gibi: we might be able to store the queue paramater in the port binding profile which is an unversioned dict of strigns filed and was intend to store hypervior specific info for a port. | |
| 17:39:13 | efried | Sundar: Not *any* failures. | |
| 17:39:18 | Sundar | If we're all ok with that, I'm good. But there are a few nuances to consider. First, the function to get the resolved ARQs is now defined in https://review.opendev.org/631245, while it needs to be used in the previous patch https://review.opendev.org/631244. | |
| 17:39:24 | sean-k-mooney | gibi: the use of multiple port bindings for live migration and the rx/tx queue size options were both intoduced in rocky so it should be backportable to all affected version if i do it right | |
| 17:40:00 | efried | Sundar: As dansmith suggested (oh, maybe it was in IRC, not in the patch) you should still fail and clean up the binding if you get e.g. a 404. | |
| 17:40:10 | efried | meaning the instance was deleted by the time you went to emit the notification. | |
| 17:40:55 | sean-k-mooney | gibi: well assuming the call order works out ill follow up with this later | |
| 17:41:25 | efried | So, to be clear, the 422 is a special case, with a prominent NOTE and a nice debug log stating that the binding completed before the instance landed on the host, so it's acceptable that the notification is dropped. | |
| 17:42:06 | efried | Sundar: refactoring method X from patch Z into patch Y because that's where it's now going to be used, yeah, that's SOP. | |
| 17:42:19 | Sundar | efried: There could be a bunch of different failure modes: cannot locate n-api service, located but it timed out?, etc. We need to distinguish that category from 'called n-api and it returned an error'. In the latter case, is there granularity of error codes for why it failed? | |
| 17:42:50 | efried | Sundar: Yes, the 422 is a special case; any other error is an error. | |
| 17:43:41 | Sundar | Need to ceck if 422 is used for any other error case. | |
| 17:43:44 | Sundar | *check | |
| 17:47:10 | Sundar | Re. the patch series, I think we need to squash https://review.opendev.org/631245 into the previous patch https://review.opendev.org/631244. Because once the get_resolved_arqs() moves to previous patch, the only things left behind are the virt driver changes. It would arguably be easier to understand those changes together with the changes to get | |
| 17:47:11 | Sundar | the arqs in compute/manager.py. | |
| 17:50:49 | mriedem | nice https://docs.openstack.org/nova/latest/admin/#overview - "nova-api Receives XML requests" | |
| 17:56:17 | openstack | Launchpad bug 1289064 in OpenStack Compute (nova) "live migration of instance should claim resources on target compute node" [Medium,In progress] - Assigned to Artom Lifshitz (notartom) | |
| 17:56:17 | melwitt | stephenfin: I noticed a new recent comment on this bug https://bugs.launchpad.net/nova/+bug/1289064 saying, maybe it can be closed bc numa aware lm was merged? seems so to me but wanted to check with you | |
| 17:56:35 | Sundar | Only one use of 422 error code in create: https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/server_external_events.py . So, we are probably good. | |
| 17:56:58 | efried | yup (I checked that earlier) | |
| 17:57:06 | artom | melwitt, stephenfin, oh, well - not entirely - only NUMA LM does claims | |
| 17:57:19 | artom | Then again, only NUMA LM has non-placement resources, so maybe yes? | |
| 17:57:26 | artom | SRIOV is handled as well, albeit separately | |
| 17:58:26 | melwitt | er, or artom, sorry ^ (I pinged stephen bc his patches about rejecting lm requests that have numa topology were merged as Related-Bug) | |
| 17:59:28 | melwitt | hm, ok. I don't really understand what that means so I dunno either | |
| 17:59:29 | sean-k-mooney | melwitt: ya so stephens patch just reject live migration if the guest has a numa toplogy at the api | |
| 18:00:07 | sean-k-mooney | sriov live migration directly allocates and claims the devices during live migration | |
| 18:00:22 | sean-k-mooney | numa live migration uses a move claim to claim resouces | |
| 18:00:58 | sean-k-mooney | (cpus and mempages/hugepages) | |
| 18:02:03 | artom | sean-k-mooney, so, could we say that all non-placement resources are not accounted correctly during live migration? | |
| 18:02:15 | sean-k-mooney | no | |
| 18:02:33 | sean-k-mooney | i dont think this is an issue on master | |
| 18:02:47 | artom | Err, *are accounted | |
| 18:02:50 | eandersson | Looks good gibi | |
| 18:03:18 | eandersson | btw mriedem did you see the minor bug I mentioned over the weekend | |
| 18:03:47 | sean-k-mooney | artom: right non-placement rescoues (cpu,mempages,sriov devices) should be accounted for in train+ and we blocked numa/sriov migration before that retoactivly | |
| 18:04:06 | openstack | Launchpad bug 1289064 in OpenStack Compute (nova) "live migration of instance should claim resources on target compute node" [Medium,In progress] - Assigned to Artom Lifshitz (notartom) | |
| 18:04:06 | sean-k-mooney | artom: melwitt so we might just be able to close https://bugs.launchpad.net/nova/+bug/1289064 | |
| 18:04:28 | eandersson | https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L2075 <-- eventlet.Timeout (based on BaseException, not Exception) may be raised here causing an UnboundError. | |
| 18:05:32 | mriedem | eandersson: nope | |
| 18:06:03 | melwitt | sean-k-mooney: if there's not anything more to do there, I'd think it would be nice to close it | |
| 18:07:10 | sean-k-mooney | well there are still cases where live migration is not supported but on master i think we account for all resouces correctly in the cases where live migration is supported | |