| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2019-11-11 | |||
| 15:47:14 | efried | sean-k-mooney: yeah, that's what I have in /usr/lib as well, point is that the libvirt-bin and n-cpu services were looking in the x86... dir. | |
| 15:47:26 | sean-k-mooney | oh ok | |
| 15:47:57 | efried | sean-k-mooney: iow I assume we should either be installing there, or be blowing those away with symlinks. | |
| 15:48:50 | sean-k-mooney | it might be best to just drop the --system | |
| 15:49:03 | sean-k-mooney | it clearly not doing it write | |
| 15:49:22 | sean-k-mooney | then it will install in /usr/local/lib which is in the lib path | |
| 15:49:43 | efried | I want to say other things blew up when I tried that | |
| 15:50:30 | efried | anyway, hopefully this is all moot irl. Unless I want to try to run a CI job around this thing. And I'm not sure if there's going to be value in that. | |
| 15:50:51 | sean-k-mooney | well ill try and spend some more time on this and see if i can make it a bit more robost. | |
| 15:52:21 | efried | sean-k-mooney: I appreciate that. Though it's probably not worth trying to make it perfect. I think the sweet spot will be having it work in devstack CI (again, if I end up wanting to do something there), which is of course a pretty narrow subset of usages. | |
| 15:53:13 | canori01 | efried: Is there a way, I coudl just force the instance to take a new allocation and spawn somewhere? I have plenty of spare capacity | |
| 15:53:13 | efried | ...excluding, among other things, "works in my dirty devstack, containing who-knows-what packages, that I've restacked a zillion times" | |
| 15:53:23 | sean-k-mooney | efried: ya i would like to get a ci job in that plugin repo at a minium but i have need to use this plugin alot in the past when enableing things for openstack so i want it to be atleast usable for that | |
| 15:53:36 | efried | canori01: At this point you should still be able to evacuate it, no? | |
| 15:54:06 | efried | oh, you mean YOU have use cases too?? Well, then, by all means :P | |
| 15:54:20 | sean-k-mooney | well not currently | |
| 15:54:53 | sean-k-mooney | i orginally wrote the first version of that plugin back in 2014 | |
| 15:55:06 | sean-k-mooney | i have used it for at least 4 feature at this point | |
| 15:55:07 | canori01 | efried: I've tried resetting the state of the instance and evacuating again, but it consistently gives me this http://paste.openstack.org/show/785945/ | |
| 15:55:21 | sean-k-mooney | so i like keping it working for when i need it | |
| 15:56:12 | efried | canori01: Oh, hm, that's a reproducible failure? That is not expected. I think we need a bug. | |
| 15:56:55 | canori01 | ok, thanks | |
| 15:57:24 | sean-k-mooney | canori01: you dont have two copies of the nova compute agent running on the compute node by any chacce? | |
| 15:58:08 | sean-k-mooney | oh never mind | |
| 15:58:13 | sean-k-mooney | this is in the scduler | |
| 16:07:47 | dansmith | cdent: don't we have some guideline around not making fundamental api behaviors change based on policy or config? (well, I know config) | |
| 16:08:40 | dansmith | cdent: meaning something like a user is granted some policy element, which enables a deep feature that they don't even know about nor can detect, which makes all calls to a specific API suddenly return 202 and be async, where normally they are 200 and sync | |
| 16:08:43 | cdent | dansmith: there's a "what should cause a microversion bump" guideline somewhere that I can dig up. Is that what you're thinkin gof? | |
| 16:09:01 | dansmith | cdent: and two users with differing policies for this element would see the different sync/async behavior | |
| 16:09:23 | cdent | I don't know of anythign with that level of detail, but that sounds like a definite no no | |
| 16:09:27 | dansmith | cdent: no because this would differ within one microversion for two users with different endowments | |
| 16:09:40 | dansmith | cdent: seems like it to me as well | |
| 16:10:09 | dansmith | cdent: might be more than you want to digest at the moment, but this is the reference: https://review.opendev.org/#/c/635684/53/nova/compute/api.py@3852 | |
| 16:10:09 | cdent | does this apply to something you're looking at (I just got back from being away, so not in context) | |
| 16:10:14 | cdent | jinx | |
| 16:10:44 | dansmith | cdent: the shortcut is: we do a cast if allow_cross_cell_resize is true, and (in a later patch) that will be true based on your policy | |
| 16:11:03 | dansmith | and that seems bad | |
| 16:11:59 | cdent | so cloud A wouild be cast and B would be call, dependent on policy settings? would it be discoverable in any way? | |
| 16:12:58 | cdent | hmmm. however, is a user ever supposed to that cells exist? | |
| 16:13:35 | cdent | because this isn't preventing or allowing resize, correct? just resize of a particular type? | |
| 16:14:15 | dansmith | cdent: no, it'd be *User* A and user B that see differences, in the same cloud | |
| 16:14:39 | dansmith | cdent: and no, a user would not know they have this ability, nor know why they see different behavior from their peer | |
| 16:15:06 | cdent | hmmm. is there a short explanation for the "why" on controlling by policy? | |
| 16:15:13 | dansmith | cdent: it's literally gating whether or not the scheduler will return all hosts, or only hosts within the current cell, not something the user knows about | |
| 16:15:30 | cdent | (I'm trying to read the code, but there's a lot here) | |
| 16:15:35 | dansmith | cdent: the why is so you can control it per-user, where config would be per-cloud | |
| 16:15:59 | cdent | yes, but _why_ would you control is per user? | |
| 16:16:00 | cdent | s/is/it/ | |
| 16:16:12 | dansmith | cdent: because it | |
| 16:16:24 | dansmith | is more expensive, and/or you may have other reasons, | |
| 16:17:10 | dansmith | it's a little weird that it's a policy knob, but it is that way just so that it can be controlled per user for whatever reason | |
| 16:17:34 | cdent | I can see that people would want that, even though they shouldn't ;) | |
| 16:17:45 | cdent | everybody wants the knobs | |
| 16:18:05 | cdent | okay, I think I have enough info to think about this with some level of legitimacy | |
| 16:18:09 | dansmith | well, that's a different argument I guess, but ... yeah. | |
| 16:18:20 | dansmith | my other argument would be that this is potentially more failure-prone, | |
| 16:18:37 | dansmith | so you may not want to give it to "simple people" or something | |
| 16:18:45 | dansmith | or just let admins do it so they can move things around the cloud | |
| 16:18:58 | dansmith | like the policy knob that lets you direct migrations to a specific host vs. just use the scheduler | |
| 16:19:16 | dansmith | but regardless, that discussion is somewhat separate from the api behavior thing | |
| 16:19:44 | cdent | yeah, just trying to add a bit of definition so I'm not filling in the blanks poorly | |
| 16:19:57 | dansmith | yup | |
| 16:20:23 | dansmith | explaining it to you here makes me feel more like my +0 should be a -1 | |
| 16:21:01 | cdent | what it is doing seems like the right thing, but I agree that it should have a microversion for the reason you state: change in behavior when same code has a different user using it | |
| 16:21:06 | cdent | i'll say as much on the review | |
| 16:21:09 | sean-k-mooney | dansmith: i dont know if you could take extention list approch like neutron | |
| 16:21:23 | sean-k-mooney | dansmith: and have that reterun different values based on policy | |
| 16:21:31 | cdent | what neutron does is terrifying from an api standpoint | |
| 16:21:44 | sean-k-mooney | well at least you can detect it | |
| 16:21:51 | dansmith | cdent: the microversion would have to say "after microversion 2.X, the method could be sync or async so watch your return code" | |
| 16:21:52 | sean-k-mooney | but yes | |
| 16:21:59 | sean-k-mooney | it still feels weird | |
| 16:22:00 | dansmith | cdent: which I thought is one of the reasons we *don't* add a microversion | |
| 16:22:22 | dansmith | sean-k-mooney: we specifically do not version our api via extensions anymore, so I'm not sure why you'd suggest that | |
| 16:22:27 | dansmith | that's the whole point of microversions :) | |
| 16:22:39 | sean-k-mooney | dansmith: neutron uses bot | |
| 16:22:42 | sean-k-mooney | *both | |
| 16:22:52 | sean-k-mooney | on your return code point | |
| 16:23:08 | sean-k-mooney | it woudl be 200 for syc and 202 or 203 for async? | |
| 16:23:30 | dansmith | sean-k-mooney: I said that above | |
| 16:23:58 | sean-k-mooney | would they send the same payload in both cases? | |
| 16:24:23 | dansmith | sean-k-mooney: you might want to read the backscroll and the code/comment I linked above | |
| 16:24:25 | sean-k-mooney | i kind of feel like haveing a flag in the post body would be cleaner | |
| 16:24:28 | dansmith | it'd be good context for the conversation | |
| 16:24:47 | dansmith | sean-k-mooney: because I've said several times that this is not something user knows about or can/would opt into | |
| 16:24:49 | sean-k-mooney | am i will after my next 1:1 in 5 mins | |
| 16:24:59 | sean-k-mooney | dansmith: ah ok | |
| 16:26:07 | sean-k-mooney | oh this is for cross cell resize | |
| 16:27:57 | sean-k-mooney | oh its related to the downcall to the compute. hum ya thats tricky | |
| 16:28:47 | sean-k-mooney | well maybe "down call" is not right but the thing we were talking to matt about last week | |
| 19:12:49 | efried | jroll: progress: using ephemeral="yes" nothing shows up on the disk. This means we'll have to define & set the secret on the fly once per guest per service-process-instance, which seems reasonable. | |
| 19:13:45 | efried | jroll: Next sticky design point: bootstrapping the secret UUID. | |
| 19:14:17 | efried | We don't want it to be in the flavor, because then you need a new flavor per secret. | |
| 19:15:52 | efried | I think we can do it this way: Create it with a random secret in the key manager. I guess neither the VM user nor the operator needs to know the secret value; they just have to be able to access it from the keymgr. | |
| 19:16:34 | efried | So we just have to figure out a way to associate the secret UUID with the instance. | |
| 19:16:45 | efried | We could use the instance UUID, but that seems fragile. | |
| 19:17:21 | efried | otherwise I suppose we need to use a new field in the instance. Which makes me a little bit sad, but it's not the end of the world. | |
| 19:17:35 | efried | dansmith: do you have any other suggestions? | |
| 19:18:10 | efried | TL;DR: I'm going to generate a UUID when I create the instance, and need to be able to associate that UUID with the instance for the life of the instance. | |
| 19:18:31 | dansmith | efried: system_metadata ? | |