Earlier  
Posted Nick Remark
#openstack-nova - 2018-06-04
15:25:33 efried bauzas: I'm saying with the naming convention being used, you can totally do that.
15:26:42 naichuans efried: currently yes, we detect the configure changes by checking inventory changes. Like bauzas said, we also can communicate with hypervisor to check previous allocated vgpu instances
15:27:26 bauzas naichuans: we could pass the allocations to the virt driver with init_host()
15:27:36 bauzas exactly like we do for spawn() or other virt methods
15:27:47 mnaser melwitt, dansmith: https://bugs.launchpad.net/nova/+bug/1581977
15:27:48 bauzas naichuans: but for the moment, I think we can just document this
15:27:48 openstack Launchpad bug 1581977 in OpenStack Compute (nova) "Invalid input for dns_name when spawning instance with .number at the end" [Undecided,Invalid]
15:29:42 naichuans bauzas: yes, we can do it. could we use it to determin vgpu type current used? Looks difficult
15:29:59 naichuans previously used
15:30:13 bauzas naichuans: I did that for reboot :)
15:30:34 bauzas I'm just looking up the existing instances and check whether the mdev is there or not
15:30:38 bauzas and just recreate it
15:32:07 naichuans bauzas: currently, we defined vgpu type in nova.conf. these types is same with the types in hpyervisor record, so it may easier to compare with hypervisor
15:35:18 naichuans and vgpu type is transparent for placement
15:40:19 openstackgerrit Matt Riedemann proposed openstack/nova-specs master: Spec for volume multiattach enhancements https://review.openstack.org/552078
15:43:32 naichuans bauzas: efried: It is late at night in China now, I'm going to sleep. I will keep close touch with you in the email.
15:44:00 openstackgerrit Dan Smith proposed openstack/nova master: Use oslo.messaging per-call monitoring https://review.openstack.org/566696
15:44:18 efried naichuans: Thanks for helping me understand this stuff. I think with some updates on the comments we might be there now.
15:44:47 naichuans efried: Np, thank you very much for the review :)
15:53:10 openstackgerrit MultipleCrashes proposed openstack/nova master: Retry decorator fix for autoscale delete https://review.openstack.org/570370
17:39:52 karimull Hi ...need your reviews on https://review.openstack.org/#/c/569498/
18:17:59 mriedem karimull: i suggest putting that into the runways review queue https://etherpad.openstack.org/p/nova-runways-rocky
18:43:12 efried karimull: Reviewed.
18:43:57 efried karimull: But I agree with mriedem - this definitely needs eyes from the likes of jaypipes, dansmith, mriedem.
18:51:42 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix typo in enable_certificate_validation config option help https://review.openstack.org/572185
18:55:17 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081
19:47:53 openstackgerrit Dan Smith proposed openstack/nova master: Change consecutive build failure limit to a weigher https://review.openstack.org/572195
20:06:21 openstackgerrit Matt Riedemann proposed openstack/nova-specs master: Remove device from volume attach requests (spec) https://review.openstack.org/452546
20:06:28 mriedem dansmith: finally got around to updating this old spec ^
20:18:40 openstackgerrit Jay Pipes proposed openstack/nova-specs master: Standardize CPU resource tracking https://review.openstack.org/555081
20:43:57 lyarwood efried: https://bugs.launchpad.net/nova/+bug/1775075 - afternoon, if you have a second could you take a look at this bug and the comment I left on the original change? https://review.openstack.org/#/c/510947/1/nova/context.py@121
20:43:58 openstack Launchpad bug 1775075 in OpenStack Compute (nova) "EndpointNotFound raised by Pike n-cpu when running alongside Queens n-api" [Undecided,New]
20:45:32 efried lyarwood: looking...
20:45:47 lyarwood efried: thanks :)
20:48:21 efried lyarwood: Okay, I think this may come down to what version of ksa exists on pike, and whether os-service-types is available.
20:50:02 lyarwood efried: 3.1.0 iirc
20:50:40 lyarwood efried: adding volumev2 back in nova/context.py is enough in my local env FWIW
20:50:57 lyarwood efried: on the n-api host that is
20:51:19 cdent is tomorrow spec sprint again?
20:51:34 efried lyarwood: I'm fine with that as a resolution; but we should probably bring mordred in on this issue, since he was basically leading me through a lot of the service catalog work.
20:51:43 melwitt cdent: yes
20:51:43 efried cdent: yes, afaik
20:51:49 cdent cool, thanks
20:52:08 melwitt I'll send another email to remind to the dev ML late today
20:52:30 lyarwood efried: ack, I'll push something to stable/queens now
20:52:40 lyarwood efried: thanks :)
20:53:02 efried lyarwood: os-service-types is not in requirements in pike, and I think that's probably the source of the problem.
20:54:26 lyarwood efried: hmm okay, AFAICT the Queens n-api was just stripping the volumev2 endpoints from the request context it was sending to the Pike compute, I'm not sure how os-service-types could help there but I've never really touched this area before.
20:55:37 efried lyarwood: The endpoints themselves should be getting picked up correctly on the queens side. If we're sending actual endpoints across the wire, that should work fine.
20:56:05 efried lyarwood: Because the queens side will have os-service-types, which means it'll find the correct endpoints (whether they're called volumev2 or whatever) for block-storage.
20:57:05 mordred efried: aroo?
20:59:46 efried mordred: See https://bugs.launchpad.net/nova/+bug/1775075 and https://review.openstack.org/#/c/510947/1/nova/context.py@121
20:59:47 openstack Launchpad bug 1775075 in OpenStack Compute (nova) "EndpointNotFound raised by Pike n-cpu when running alongside Queens n-api" [Undecided,New]
21:00:59 efried mordred: What lyarwood is finding is that when he adds 'volumev2' back in on the queens side, it fixes the problem. Which I don't quite understand why doing that on the *queens* side would fix. Because that guy has os-service-types and the right ksa... oh, unless this is one of those ksa bugs we've since fixed, where endpoint lookup isn't happening right.
21:02:10 lyarwood efried: so the Pike side is looking for the cinderv2 endpoint in the catalog stashed in the request context
21:02:28 lyarwood efried: the cinderv3 endpoint is there correctly, but without the cinderv2 endpoint it just fails
21:02:37 mordred efried: probably. when you say "queens side" and "pike side" - what does that mean here?
21:02:51 lyarwood mordred: queens n-api, pike n-cpu
21:02:56 mordred gotcha
21:03:49 mordred yah - I think you'll need the volumev2 in there until everything is on queens - I think efried is correct about there being missing fixes
21:05:11 lyarwood oh I forgot to say that our P n-cpu's have catalog_info=volumev2:cinderv2:internal in nova.conf forcing them to look for cinderv2
21:06:19 efried ah
21:06:31 efried well then yeah
21:06:38 lyarwood yeah my bad, the default is cinderv3 right?
21:06:48 efried so is this a bug in the code, or is this something that should be fixed via conf on site?
21:07:21 lyarwood I'd say code still if possible
21:08:10 efried lyarwood: Is there a similar argument for restoring 'volume' ?
21:08:41 lyarwood efried: yeah assuming you could set catalog_info to look for it
21:11:12 efried lyarwood: tbc, this patch doesn't exist on pike, right?
21:11:19 efried https://review.openstack.org/#/c/510947/ that is
21:11:28 lyarwood efried: correct
21:11:31 lyarwood efried: just Queens
21:12:13 efried lyarwood: Okay, well, restoring 'volume' is not necessary because https://review.openstack.org/#/c/409904/
21:12:45 efried lyarwood: But volumev2 was dropped in q (https://review.openstack.org/#/c/501874/)
21:13:09 efried lyarwood: So I guess unless queens n-api can talk to ocata n-cpu? Is that a thing?
21:13:40 lyarwood efried: nope, shouldn't be.
21:14:17 efried okay. Then I suppose restoring volumev2 on the premise that q can talk to p is a reasonable thing. Though I'm still confused as to why the problem would manifest if the endpoint lookup is being done on the q side.
21:14:52 efried lyarwood: Do you have the ability to check whether upgrading ksa (without reintroducing 'volumev2' in that RequestContext filter) also fixes the problem?
21:15:13 lyarwood efried: on the Q n-api host?
21:15:38 efried yeah - upgrade ksa and service-types-authoryt.
21:15:42 efried authority
21:15:58 efried lyarwood: That would give us the info as to whether the issue lies with missing ksa fixes, as mordred quasi-corroborated.
21:16:13 lyarwood efried: just to the latest stable versions I assume?
21:16:20 lyarwood efried: I can give that a try
21:16:26 efried lyarwood: yeah, that should be adequate.
21:20:51 mordred latest ksa SHOULD make this happier
21:27:02 lyarwood efried / mordred ; still fails with ksa 3.7.0
21:27:42 efried lyarwood: Okay, that's an interesting data point. Are you sure the RequestContext itself is really going across the wire, complete with its service catalog?
21:28:28 efried lyarwood: If the cached service catalog is being stripped, but the remainder of the context is being trasmitted, then that would certainly cause the problem. Because the pike side doesn't have os-service-types so it doesn't know what to do with 'block-storage'.
21:29:10 efried or maybe we haven't looked up (the cinder endpoint in) the service catalog yet by the time it's being sent over.
21:29:39 lyarwood efried: yeah I'm sure, cinderv3 is there when it's going over the wire from Q
21:30:06 efried lyarwood: Ohh, because the queens side would be using the cinderv3 endpoint.
21:30:12 efried lyarwood: But the pike side wants the v2 endpoint.
21:30:23 lyarwood efried: yeah indeed
21:30:51 efried So if the q request context only looked up v3, even if it sent the cached service catalog over the wire, the p side wouldn't see the v2 endpoint in it.
21:31:51 lyarwood efried: it's actually finding both before we hit nova/context.py iirc
21:32:06 efried lyarwood: But would have no reason to load up the discovery info for v2
21:32:08 lyarwood efried: on Q, but we then strip cinderv2 out causing the problem
21:32:20 efried oh, I see, yes.
21:32:42 efried lyarwood: IMO the bug sounds like we shouldn't be sending the RequestContext from Q to P. But we're not gonna fix that at this point. Adding volumev2 back in sounds like the less intrusive fix.

Earlier   Later