| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-04-23 | |||
| 20:10:00 | artom | ^^^ is about boot-time device tagging, but it explains the "why (is that useless to the guest OS)" question | |
| 20:10:05 | artom | jaypipes, it doens't | |
| 20:10:09 | dansmith | by not supporting more than one virt method? | |
| 20:10:31 | artom | jaypipes, wait, ec2... the metadata API? | |
| 20:10:48 | dansmith | he means EC2 the service I believe | |
| 20:10:52 | jaypipes | dansmith: no, how can an EC2 API user inform their guest that a particular NIC is "for this specific network"? | |
| 20:11:15 | dansmith | jaypipes: that's totally different than applying a tag to a thing | |
| 20:11:33 | artom | dansmith, to be fair, that's kinda what tags were made for | |
| 20:11:42 | dansmith | jaypipes: you may impute some meaning from the tag, but that's different | |
| 20:11:49 | jaypipes | exactly what they were made for, actually. | |
| 20:12:09 | dansmith | jaypipes: so lets use the volume case | |
| 20:12:23 | dansmith | jaypipes: doesn't ec2 let you say "this volume will be vdb"? | |
| 20:13:38 | jaypipes | dansmith: you can specify one, but it might rename it behind the scenes. | |
| 20:13:50 | artom | Also, and this may be stupid, but what does EC2 have to do with any of this? | |
| 20:14:06 | dansmith | artom: zero, because they don't do many hypervisors | |
| 20:14:29 | dansmith | jaypipes: can we jump on a hangout here to figure out what the real concern is? because I feel like something must be confused if you're really worried about this | |
| 20:15:33 | dansmith | jaypipes: https://hangouts.google.com/call/Jcmcsrr1qa7qmYfuZ_myAAEE ? | |
| 20:16:26 | dansmith | artom: of course, if jaypipes'll join | |
| 20:16:41 | artom | Of course what? | |
| 20:17:08 | dansmith | artom: join ^ | |
| 20:17:09 | jaypipes | I'm on. | |
| 20:17:46 | artom | dansmith, uh, sure, except it keeps telling me I'm not allowed | |
| 20:18:12 | dansmith | artom: we're on, it's just public | |
| 20:18:41 | artom | Oh hey, the gmail account works | |
| 20:23:33 | melwitt | mriedem: yeah, I guess so. I was thinking similar because it's not clear to me if this has utility beyond abort cold migrate. the alternatives section is empty https://specs.openstack.org/openstack/nova-specs/specs/rocky/approved/list-show-all-server-migration-types.html#alternatives | |
| 20:23:55 | mriedem | melwitt: exactly. ok i'll start a thread. | |
| 20:24:01 | melwitt | thanks | |
| 20:28:39 | melwitt | arvindn0_, mriedem: which thing are y'all talking about "moving to the end of the runway" earlier? | |
| 20:28:55 | mriedem | glance-image-traits | |
| 20:28:58 | mriedem | L62 | |
| 20:29:04 | mriedem | end-o-queue | |
| 20:29:08 | arvindn0_ | Support Traits in Glance: https://blueprints.launchpad.net/nova/+spec/glance-image-traits | |
| 20:29:22 | melwitt | oh, I see | |
| 20:29:36 | melwitt | I got confused with the rebuild talk right after it | |
| 20:29:45 | arvindn0_ | we are discussing an ammendement to the originally approved spec... | |
| 20:30:03 | melwitt | gotcha. I see the note in the runways etherpad | |
| 20:30:09 | arvindn0_ | rebuild is related to the amendment :) | |
| 20:30:19 | melwitt | ah, okay | |
| 20:38:20 | melwitt | mriedem: agreed that the host/hostId to instance action events API is the only thing that looks ready for a runway, so I'm gonna move it there | |
| 20:38:35 | mriedem | cool | |
| 20:38:36 | mriedem | http://lists.openstack.org/pipermail/openstack-dev/2018-April/129736.html | |
| 20:38:49 | melwitt | thanks | |
| 20:41:54 | arvindn0_ | mriedem: wanted to quick check something on the scenario 2 i detailed in the comment | |
| 20:43:08 | arvindn0_ | a host with two SRIOV nic. One is normal SRIOV nic, another one with some kind of offload feature. | |
| 20:43:46 | melwitt | yikun: hi, I know you're not around right now but FYI your blueprint "Add host/hostId to instance action events API" has been added to a review runway https://etherpad.openstack.org/p/nova-runways-rocky | |
| 20:44:37 | arvindn0_ | Initial instance launch happens with SRIOV_VF:1 allocated, rebuild lauches with modified request with traits=HW_NIC_OFFLOAD_X, so basically we want the instance to be allocated the second nic | |
| 20:45:30 | arvindn0_ | but the original allocation happens against nic1 and since in rebuild the original allocations are not changed, we have wrong allocations | |
| 20:46:09 | arvindn0_ | mriedem: is the above scenario an issue if we use GET /resource_providers/{rp_uuid}/traits approch? | |
| 20:49:28 | mriedem | arvindn0_: did you see efried's reply in the mailing list? | |
| 20:50:07 | mriedem | even with efried's suggestion, that scenario is likely a gap yes | |
| 20:50:40 | arvindn0_ | mriedem: trying to figure out the mailing list...i only get digests....do you have a link? | |
| 20:50:44 | mriedem | we need to have some way of knowing, will a rebuild with a new image result in new allocations and if so, fail the rebuild | |
| 20:51:13 | mriedem | you can change your subscription to not be digests :) | |
| 20:51:24 | mriedem | http://lists.openstack.org/pipermail/openstack-dev/2018-April/129734.html | |
| 20:51:57 | mriedem | arvindn0_: also, what you're describing above is this bug https://bugs.launchpad.net/nova/+bug/1763766 | |
| 20:51:57 | openstack | Launchpad bug 1763766 in OpenStack Compute (nova) "nova needs to disallow topology changes on image rebuild" [Medium,Triaged] | |
| 20:51:58 | efried | mriedem, arvindn0_: With what I suggested, you will still know that. Because you know the RPs you've already allocated from (and which you haven't). | |
| 20:52:23 | mriedem | efried: i'm not sure we do, we just have the root provider uuid in the scheduler filter | |
| 20:52:36 | mriedem | unless we are going to build a ProviderTree object or something | |
| 20:52:55 | efried | mriedem: Do you have access to a SchedulerReportClient from wherever you are? | |
| 20:53:01 | mriedem | sure | |
| 20:53:03 | efried | mriedem: If so, then yeah, you have that ability with a single call. | |
| 20:54:07 | efried | mriedem: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L986 would be that call. | |
| 20:54:36 | efried | The 'ensure_root' part is slightly uncomfortable, but you know the root exists when you call this. | |
| 20:54:42 | efried | And we can factor that out of there if it's an issue. | |
| 20:57:06 | munimehan | All, when nova generates VM definition it don't generate the PCIs for guest addresses, is there a way to manipulate those to make it in sync for the interface order | |
| 20:57:58 | arvindn0_ | efried: good suggestion...will look into that. | |
| 20:58:59 | arvindn0_ | btw currently in the image_props_filter.py we dont have SchedulerReportClient...we can create and get access...but now we have 2 API calls from the filter.... | |
| 21:00:15 | efried | arvindn0_: Do you have a resourcetracker? | |
| 21:00:23 | efried | arvindn0_: or a schedulerclient? | |
| 21:01:00 | arvindn0_ | nope...currently the filters make 0 api calls....so no clients of any sort | |
| 21:01:44 | efried | arvindn0_: But we were going to have to call into placement one way or another (unless we went with option 1 and just blocked the whole thing). | |
| 21:02:15 | efried | arvindn0_: So yeah, you're not getting around that. And those should be pretty quick calls, one would hope. | |
| 21:02:49 | efried | arvindn0_: The set logic around required vs forbidden will be a little tricky. | |
| 21:02:54 | arvindn0_ | yup...i was just pointing out that we might now have to make 2 calls, but like you said they should be quick since they are focussed on only the current compute node | |
| 21:03:26 | arvindn0_ | we dont support forbidden traits in images i think... | |
| 21:04:58 | efried | arvindn0_: If you did, I think you would probably have to do two separate GET /resource_providers calls (so three placement calls total) - one for required and one for forbidden. | |
| 21:05:24 | efried | arvindn0_: Sorry, I take that back - I would have to think that through when I'm not so distracted. | |
| 21:06:04 | arvindn0_ | efried: will save you time, we dont support forbidden traits in image...we only support required traits | |
| 21:06:14 | efried | k | |
| 21:07:05 | efried | arvindn0_: I'm thinking through the logic, though, trying to figure out how you will actually know you're good, even with required. | |
| 21:08:36 | efried | arvindn0_: Aha, actually, it's easier than I thought. | |
| 21:09:48 | efried | arvindn0_: You actually *only* need the call to get_provider_tree_and_ensure_root. You can walk that guy and collect the set of all traits from all the RPs you're already allocated from. Subtract that from the set of traits in your image. If there's anything left over, fail. Otherwise you're good. | |
| 21:10:42 | arvindn0_ | how do we know our allocated RP's? | |
| 21:10:48 | arvindn0_ | is that part of the provider tree? | |
| 21:11:15 | efried | Yes, you pass in your compute node UUID | |
| 21:11:21 | efried | You'll get back a ProviderTree. | |
| 21:11:33 | efried | wait, hold on. | |
| 21:12:07 | arvindn0_ | efried: not too familiar with the provider tree object...my understanding provider tree will return all traits etc | |
| 21:12:28 | arvindn0_ | but not which RP is allocated to us... | |
| 21:12:33 | efried | yes it will. For the whole tree, and associated sharing providers. And you need ... | |
| 21:12:34 | efried | exactly. | |
| 21:12:54 | efried | arvindn0_: Now, we're focused on an instance here, right? | |
| 21:12:59 | efried | arvindn0_: Not the whole host. | |
| 21:13:01 | arvindn0_ | yup | |
| 21:13:44 | arvindn0_ | this is rebuild of that instance...so we need to know what resource was allocated with what traits... | |
| 21:14:06 | efried | So what you'll *actually* want to get is the *allocations* for the instance. From that you can get the providers associated with your instance. Walk the ProviderTree on just those, peeling out the traits. Subtract that from the set of traits in your image. If there's anything left over, fail. Otherwise, you're good. | |
| 21:14:08 | arvindn0_ | we cant also do an aggregate of all the traits, because we want to ensure a 1:1 mapping | |
| 21:14:23 | efried | arvindn0_: That doesn't make sense. | |
| 21:14:39 | efried | arvindn0_: What do you mean by 1:1 mapping? | |