| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-06 | |||
| 13:53:09 | bauzas | mriedem: yep, I'm wondering if we should reconsider the opportunity to have claims done by the conductor for the reasons you mentioned | |
| 13:53:26 | bauzas | #1 if you force, then you need to reconcile claims | |
| 13:53:39 | mriedem | i think we can still do force and have the scheduler handle the claim | |
| 13:53:45 | mriedem | i'm going to write a bp for that | |
| 13:53:54 | mriedem | since it's going to be an rpc change and require some consideration | |
| 13:53:57 | bauzas | #2 for move operations, there are a list of corner cases depending on the success of the move that require the conductor to reconcile allocations as well | |
| 13:54:16 | mriedem | sure, but for #2 we can also fail on the compute and need to cleanup allocations, | |
| 13:54:24 | bauzas | good point | |
| 13:54:30 | mriedem | which we're seeing with prep_resize failing, evacuate moveclaim failing, and unshelve claim failing | |
| 13:54:44 | mriedem | i do'nt really see this any different from cleaning up ports and volume attachments | |
| 13:54:52 | mriedem | which happen both in conductor sometimes and in the compute | |
| 13:54:59 | bauzas | about the force thingy, what if placement returns "Sorry"' to the scheduler if it claims ? | |
| 13:55:12 | mriedem | then we fail | |
| 13:55:21 | bauzas | I'm not sure operators would accept that | |
| 13:55:26 | mriedem | they are going to have to | |
| 13:55:30 | bauzas | LOL | |
| 13:55:50 | mriedem | if you force to a compute today, the claim on the compute could still fail | |
| 13:56:00 | mriedem | except live migration doesn't do a claim | |
| 13:56:07 | mriedem | but evacuate does | |
| 13:56:15 | bauzas | that's more complicated : | |
| 13:56:22 | mriedem | maybe those claims never fail because we pass an empty limits dict | |
| 13:56:33 | mriedem | so the claim just considers unlimited cpu/ram/disk | |
| 13:56:37 | bauzas | since we don't call the scheduler, the scheduler isn't passing limits to the conductor which eventually gives them to the compute | |
| 13:56:51 | bauzas | so we don't really verify the resource usage | |
| 13:57:00 | mriedem | yeah i just said the same thing | |
| 13:57:13 | mriedem | honestly i didn't realize that's how things worked until last week when writing some functional tests for this | |
| 13:57:15 | jaypipes | mriedem, bauzas, alex_xu: https://twitter.com/jaypipes/status/905429571012616192 | |
| 13:57:15 | bauzas | you typed too fast, damn you | |
| 13:57:44 | mriedem | jaypipes: i can't tell if they look freaked out like that because of the storm or if that's just normal | |
| 13:57:45 | bauzas | jaypipes: orly? :( you're going to be impacted ? :( | |
| 13:58:06 | jaypipes | bauzas: yeah. we're probably going to evacuate tonight or tomorrow morning. | |
| 13:58:11 | bauzas | jaypipes: le woof is hugging them | |
| 13:58:13 | jaypipes | mriedem: just normal. | |
| 13:58:29 | jaypipes | heh | |
| 13:58:32 | mriedem | bauzas: point is, we need the resource allocation representation in placement to be accurate, | |
| 13:58:40 | mriedem | and since the RT isn't adjusting claims anymore, | |
| 13:58:45 | mriedem | after the fact i mean, | |
| 13:58:53 | mriedem | we can't just not create the allocations, | |
| 13:59:35 | mriedem | with the force stuff before with claims, the RT would still report the usage on the compute node to the scheduler and if it was full then that compute node is out of the running for placement decisions | |
| 13:59:47 | mriedem | if we don't report the allocations to placement, then things are incorrect for the scheduler | |
| 13:59:53 | mriedem | regardless of force | |
| 14:00:04 | mriedem | plus, as i said in the ML, i think force is a bad idea anyway | |
| 14:00:27 | mriedem | i'm semi ok with continuing to ignore the filters (even though the live migration force scenario still checks ram and compute filters, but in conductor code) | |
| 14:00:34 | mriedem | but i'm not ok with bypassing placement | |
| 14:01:57 | mriedem | bauzas: so maybe what you want to see in https://review.openstack.org/#/c/499399/ is a release note that says, evacuate with force=True may fail due to placement | |
| 14:02:18 | mriedem | and at some point an update to the api-ref for the force flag on evacuate and live migrate | |
| 14:02:44 | mriedem | i.e. the force flag bypasses the scheduler filters but can still fail due to placement | |
| 14:03:57 | mriedem | we merged the same for live migrate for the pike GA but didn't have a release note for that https://review.openstack.org/#/c/496727/ | |
| 14:04:05 | mriedem | i could add a release note for the 16.0.1 release | |
| 14:04:42 | bauzas | mriedem: so, I agree with you about having "something" that would reconcile allocations if you force | |
| 14:05:06 | mriedem | since the RT isn't going to do that once all computes are upgraded to pike, | |
| 14:05:18 | bauzas | the fact is, some operators for example want to live-migrate a whole host to another one with the price of degradation | |
| 14:05:22 | dansmith | we can audit without heal | |
| 14:05:24 | mriedem | it's better to try and allocate in the controller and fail fast than do it in the compute after we've already live migrated or evacuated something | |
| 14:05:41 | bauzas | because that's an emergency situation | |
| 14:05:53 | bauzas | and they don't want their customers to know that the host is bad | |
| 14:06:07 | mriedem | if they are live migrating, the customer doesn't know | |
| 14:06:11 | mriedem | if it moved or not | |
| 14:06:38 | dansmith | well, the customer will know of course | |
| 14:06:58 | dansmith | minimum disruption, but that's all | |
| 14:06:58 | bauzas | sure, but they want to migrate sooner than later and somehow want to migrate the host to some other they know | |
| 14:07:24 | bauzas | agreed | |
| 14:07:37 | mriedem | btw, unrelated, but can we get this merged https://review.openstack.org/#/c/500968/ to unblock the requirements upper-constraints change from merging? | |
| 14:07:46 | dansmith | I'm not sure I understand what the concern is, I just started reading | |
| 14:07:46 | mriedem | ^ is needed to get nova unit tests to pass with the latest os-xenapi lib | |
| 14:07:48 | dansmith | all we have to do is make sure force does a caim first, and if that fails, fail, right? | |
| 14:08:21 | mriedem | dansmith: which is exactly what this does https://review.openstack.org/#/c/496727/ | |
| 14:08:34 | mriedem | and https://review.openstack.org/#/c/499399/ | |
| 14:08:38 | dansmith | right | |
| 14:09:10 | mriedem | we're really talking about two things and diverged from the first, which is bauzas wants to consider moving the claim code out of the scheduler and into conductor | |
| 14:09:16 | mriedem | and then we went down the force flag rathole | |
| 14:09:25 | alex_xu | jamespage: bauzas there also a proposal in the intel about integrate an agent which managing cpu l3 cache into the nova. I'm also thinking report the l3 cache inventory by the agent to the placement directly is option, but there missed a way for the agent to know new instance boot up in the node except polling the libvirt | |
| 14:09:43 | mriedem | alex_xu: listen for notifications | |
| 14:09:46 | alex_xu | jaypipes: ^ | |
| 14:09:55 | alex_xu | jamespage: sorry for pointed to wrong person | |
| 14:10:15 | alex_xu | mriedem: but the notification isn't for single host, we must filter the info | |
| 14:10:18 | jaypipes | alex_xu: you mean RDT/CAT, right? | |
| 14:10:18 | mriedem | alex_xu: if you want a hypervisor agnostic solution, listen for notifications rather than polling libvirt directly | |
| 14:10:25 | alex_xu | jaypipes: yea | |
| 14:11:24 | mriedem | alex_xu: the instance.create.end payload would have the host/node information in it so yeah you'd have to filter but that doesn't seem hard | |
| 14:11:31 | bauzas | I don't get why it requires libvirt to be polled | |
| 14:12:02 | bauzas | and I second mriedem about subscribing to the notifications bus | |
| 14:12:30 | kashyap | bauzas: What requires libvirt to be polled? /me read it out of context; will read the scrollback | |
| 14:12:37 | alex_xu | mriedem: yea, that won't be so hard, that probably each node have a agent to listen the notification | |
| 14:12:41 | mriedem | bauzas: xenserver ci isn't going to care about unit test changes https://review.openstack.org/#/c/500968/ | |
| 14:12:51 | bauzas | holy ship, you're right | |
| 14:13:05 | alex_xu | bauzas: kashyap an external agent try to know there is new instance boot up | |
| 14:13:16 | mriedem | alex_xu: plus then you can use the fancy new versioned notifications and provide feedback to gibi | |
| 14:13:39 | alex_xu | mriedem: yeah | |
| 14:14:02 | gibi | yeah :) | |
| 14:14:03 | kashyap | alex_xu: Ah, it's the L3 CAT interface. I see some in-progress design work of exposing it in libvirt upstream | |
| 14:14:23 | alex_xu | kashyap: yea, looks like libvirt interface is slow | |
| 14:14:49 | kashyap | alex_xu: Yeah, there's a new thread that started a day ago | |
| 14:14:56 | kashyap | Did you see that? About the new RFC / design | |
| 14:15:36 | alex_xu | kashyap: not yet, another team member follow that, probably I will tell him | |
| 14:16:35 | mdbooth | melwitt: You about? Spotted something potentially funky looking at your change https://review.openstack.org/#/c/498983/ | |
| 14:16:48 | mdbooth | Not directly related to your patch | |
| 14:17:13 | alex_xu | kashyap: thanks for the info | |
| 14:17:23 | mdbooth | I couldn't see where we wrap swap_volume in any kind of lock | |
| 14:17:41 | kashyap | alex_xu: Here it is: "Yet another RFC for [Intel] CAT" - https://www.redhat.com/archives/libvir-list/2017-September/msg00072.html | |