Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-06
13:51:39 bauzas honestly, dansmith proposed to have the scheduler code to be run by the conductor service, and I just think all of this tends to that
13:51:41 mriedem the issue with force is that we bypass the scheduler altogether
13:52:13 bauzas yup
13:52:22 mriedem so are you asking if we should stop doing the claim in the scheduler and move that to conductor now in queens?
13:53:09 bauzas mriedem: yep, I'm wondering if we should reconsider the opportunity to have claims done by the conductor for the reasons you mentioned
13:53:26 bauzas #1 if you force, then you need to reconcile claims
13:53:39 mriedem i think we can still do force and have the scheduler handle the claim
13:53:45 mriedem i'm going to write a bp for that
13:53:54 mriedem since it's going to be an rpc change and require some consideration
13:53:57 bauzas #2 for move operations, there are a list of corner cases depending on the success of the move that require the conductor to reconcile allocations as well
13:54:16 mriedem sure, but for #2 we can also fail on the compute and need to cleanup allocations,
13:54:24 bauzas good point
13:54:30 mriedem which we're seeing with prep_resize failing, evacuate moveclaim failing, and unshelve claim failing
13:54:44 mriedem i do'nt really see this any different from cleaning up ports and volume attachments
13:54:52 mriedem which happen both in conductor sometimes and in the compute
13:54:59 bauzas about the force thingy, what if placement returns "Sorry"' to the scheduler if it claims ?
13:55:12 mriedem then we fail
13:55:21 bauzas I'm not sure operators would accept that
13:55:26 mriedem they are going to have to
13:55:30 bauzas LOL
13:55:50 mriedem if you force to a compute today, the claim on the compute could still fail
13:56:00 mriedem except live migration doesn't do a claim
13:56:07 mriedem but evacuate does
13:56:15 bauzas that's more complicated :
13:56:22 mriedem maybe those claims never fail because we pass an empty limits dict
13:56:33 mriedem so the claim just considers unlimited cpu/ram/disk
13:56:37 bauzas since we don't call the scheduler, the scheduler isn't passing limits to the conductor which eventually gives them to the compute
13:56:51 bauzas so we don't really verify the resource usage
13:57:00 mriedem yeah i just said the same thing
13:57:13 mriedem honestly i didn't realize that's how things worked until last week when writing some functional tests for this
13:57:15 jaypipes mriedem, bauzas, alex_xu: https://twitter.com/jaypipes/status/905429571012616192
13:57:15 bauzas you typed too fast, damn you
13:57:44 mriedem jaypipes: i can't tell if they look freaked out like that because of the storm or if that's just normal
13:57:45 bauzas jaypipes: orly? :( you're going to be impacted ? :(
13:58:06 jaypipes bauzas: yeah. we're probably going to evacuate tonight or tomorrow morning.
13:58:11 bauzas jaypipes: le woof is hugging them
13:58:13 jaypipes mriedem: just normal.
13:58:29 jaypipes heh
13:58:32 mriedem bauzas: point is, we need the resource allocation representation in placement to be accurate,
13:58:40 mriedem and since the RT isn't adjusting claims anymore,
13:58:45 mriedem after the fact i mean,
13:58:53 mriedem we can't just not create the allocations,
13:59:35 mriedem with the force stuff before with claims, the RT would still report the usage on the compute node to the scheduler and if it was full then that compute node is out of the running for placement decisions
13:59:47 mriedem if we don't report the allocations to placement, then things are incorrect for the scheduler
13:59:53 mriedem regardless of force
14:00:04 mriedem plus, as i said in the ML, i think force is a bad idea anyway
14:00:27 mriedem i'm semi ok with continuing to ignore the filters (even though the live migration force scenario still checks ram and compute filters, but in conductor code)
14:00:34 mriedem but i'm not ok with bypassing placement
14:01:57 mriedem bauzas: so maybe what you want to see in https://review.openstack.org/#/c/499399/ is a release note that says, evacuate with force=True may fail due to placement
14:02:18 mriedem and at some point an update to the api-ref for the force flag on evacuate and live migrate
14:02:44 mriedem i.e. the force flag bypasses the scheduler filters but can still fail due to placement
14:03:57 mriedem we merged the same for live migrate for the pike GA but didn't have a release note for that https://review.openstack.org/#/c/496727/
14:04:05 mriedem i could add a release note for the 16.0.1 release
14:04:42 bauzas mriedem: so, I agree with you about having "something" that would reconcile allocations if you force
14:05:06 mriedem since the RT isn't going to do that once all computes are upgraded to pike,
14:05:18 bauzas the fact is, some operators for example want to live-migrate a whole host to another one with the price of degradation
14:05:22 dansmith we can audit without heal
14:05:24 mriedem it's better to try and allocate in the controller and fail fast than do it in the compute after we've already live migrated or evacuated something
14:05:41 bauzas because that's an emergency situation
14:05:53 bauzas and they don't want their customers to know that the host is bad
14:06:07 mriedem if they are live migrating, the customer doesn't know
14:06:11 mriedem if it moved or not
14:06:38 dansmith well, the customer will know of course
14:06:58 dansmith minimum disruption, but that's all
14:06:58 bauzas sure, but they want to migrate sooner than later and somehow want to migrate the host to some other they know
14:07:24 bauzas agreed
14:07:37 mriedem btw, unrelated, but can we get this merged https://review.openstack.org/#/c/500968/ to unblock the requirements upper-constraints change from merging?
14:07:46 dansmith I'm not sure I understand what the concern is, I just started reading
14:07:46 mriedem ^ is needed to get nova unit tests to pass with the latest os-xenapi lib
14:07:48 dansmith all we have to do is make sure force does a caim first, and if that fails, fail, right?
14:08:21 mriedem dansmith: which is exactly what this does https://review.openstack.org/#/c/496727/
14:08:34 mriedem and https://review.openstack.org/#/c/499399/
14:08:38 dansmith right
14:09:10 mriedem we're really talking about two things and diverged from the first, which is bauzas wants to consider moving the claim code out of the scheduler and into conductor
14:09:16 mriedem and then we went down the force flag rathole
14:09:25 alex_xu jamespage: bauzas there also a proposal in the intel about integrate an agent which managing cpu l3 cache into the nova. I'm also thinking report the l3 cache inventory by the agent to the placement directly is option, but there missed a way for the agent to know new instance boot up in the node except polling the libvirt
14:09:43 mriedem alex_xu: listen for notifications
14:09:46 alex_xu jaypipes: ^
14:09:55 alex_xu jamespage: sorry for pointed to wrong person
14:10:15 alex_xu mriedem: but the notification isn't for single host, we must filter the info
14:10:18 jaypipes alex_xu: you mean RDT/CAT, right?
14:10:18 mriedem alex_xu: if you want a hypervisor agnostic solution, listen for notifications rather than polling libvirt directly
14:10:25 alex_xu jaypipes: yea
14:11:24 mriedem alex_xu: the instance.create.end payload would have the host/node information in it so yeah you'd have to filter but that doesn't seem hard
14:11:31 bauzas I don't get why it requires libvirt to be polled
14:12:02 bauzas and I second mriedem about subscribing to the notifications bus
14:12:30 kashyap bauzas: What requires libvirt to be polled? /me read it out of context; will read the scrollback
14:12:37 alex_xu mriedem: yea, that won't be so hard, that probably each node have a agent to listen the notification
14:12:41 mriedem bauzas: xenserver ci isn't going to care about unit test changes https://review.openstack.org/#/c/500968/
14:12:51 bauzas holy ship, you're right
14:13:05 alex_xu bauzas: kashyap an external agent try to know there is new instance boot up
14:13:16 mriedem alex_xu: plus then you can use the fancy new versioned notifications and provide feedback to gibi
14:13:39 alex_xu mriedem: yeah
14:14:02 gibi yeah :)
14:14:03 kashyap alex_xu: Ah, it's the L3 CAT interface. I see some in-progress design work of exposing it in libvirt upstream
14:14:23 alex_xu kashyap: yea, looks like libvirt interface is slow
14:14:49 kashyap alex_xu: Yeah, there's a new thread that started a day ago
14:14:56 kashyap Did you see that? About the new RFC / design
14:15:36 alex_xu kashyap: not yet, another team member follow that, probably I will tell him
14:16:35 mdbooth melwitt: You about? Spotted something potentially funky looking at your change https://review.openstack.org/#/c/498983/

Earlier   Later