Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-06
14:01:57 mriedem bauzas: so maybe what you want to see in https://review.openstack.org/#/c/499399/ is a release note that says, evacuate with force=True may fail due to placement
14:02:18 mriedem and at some point an update to the api-ref for the force flag on evacuate and live migrate
14:02:44 mriedem i.e. the force flag bypasses the scheduler filters but can still fail due to placement
14:03:57 mriedem we merged the same for live migrate for the pike GA but didn't have a release note for that https://review.openstack.org/#/c/496727/
14:04:05 mriedem i could add a release note for the 16.0.1 release
14:04:42 bauzas mriedem: so, I agree with you about having "something" that would reconcile allocations if you force
14:05:06 mriedem since the RT isn't going to do that once all computes are upgraded to pike,
14:05:18 bauzas the fact is, some operators for example want to live-migrate a whole host to another one with the price of degradation
14:05:22 dansmith we can audit without heal
14:05:24 mriedem it's better to try and allocate in the controller and fail fast than do it in the compute after we've already live migrated or evacuated something
14:05:41 bauzas because that's an emergency situation
14:05:53 bauzas and they don't want their customers to know that the host is bad
14:06:07 mriedem if they are live migrating, the customer doesn't know
14:06:11 mriedem if it moved or not
14:06:38 dansmith well, the customer will know of course
14:06:58 dansmith minimum disruption, but that's all
14:06:58 bauzas sure, but they want to migrate sooner than later and somehow want to migrate the host to some other they know
14:07:24 bauzas agreed
14:07:37 mriedem btw, unrelated, but can we get this merged https://review.openstack.org/#/c/500968/ to unblock the requirements upper-constraints change from merging?
14:07:46 dansmith I'm not sure I understand what the concern is, I just started reading
14:07:46 mriedem ^ is needed to get nova unit tests to pass with the latest os-xenapi lib
14:07:48 dansmith all we have to do is make sure force does a caim first, and if that fails, fail, right?
14:08:21 mriedem dansmith: which is exactly what this does https://review.openstack.org/#/c/496727/
14:08:34 mriedem and https://review.openstack.org/#/c/499399/
14:08:38 dansmith right
14:09:10 mriedem we're really talking about two things and diverged from the first, which is bauzas wants to consider moving the claim code out of the scheduler and into conductor
14:09:16 mriedem and then we went down the force flag rathole
14:09:25 alex_xu jamespage: bauzas there also a proposal in the intel about integrate an agent which managing cpu l3 cache into the nova. I'm also thinking report the l3 cache inventory by the agent to the placement directly is option, but there missed a way for the agent to know new instance boot up in the node except polling the libvirt
14:09:43 mriedem alex_xu: listen for notifications
14:09:46 alex_xu jaypipes: ^
14:09:55 alex_xu jamespage: sorry for pointed to wrong person
14:10:15 alex_xu mriedem: but the notification isn't for single host, we must filter the info
14:10:18 jaypipes alex_xu: you mean RDT/CAT, right?
14:10:18 mriedem alex_xu: if you want a hypervisor agnostic solution, listen for notifications rather than polling libvirt directly
14:10:25 alex_xu jaypipes: yea
14:11:24 mriedem alex_xu: the instance.create.end payload would have the host/node information in it so yeah you'd have to filter but that doesn't seem hard
14:11:31 bauzas I don't get why it requires libvirt to be polled
14:12:02 bauzas and I second mriedem about subscribing to the notifications bus
14:12:30 kashyap bauzas: What requires libvirt to be polled? /me read it out of context; will read the scrollback
14:12:37 alex_xu mriedem: yea, that won't be so hard, that probably each node have a agent to listen the notification
14:12:41 mriedem bauzas: xenserver ci isn't going to care about unit test changes https://review.openstack.org/#/c/500968/
14:12:51 bauzas holy ship, you're right
14:13:05 alex_xu bauzas: kashyap an external agent try to know there is new instance boot up
14:13:16 mriedem alex_xu: plus then you can use the fancy new versioned notifications and provide feedback to gibi
14:13:39 alex_xu mriedem: yeah
14:14:02 gibi yeah :)
14:14:03 kashyap alex_xu: Ah, it's the L3 CAT interface. I see some in-progress design work of exposing it in libvirt upstream
14:14:23 alex_xu kashyap: yea, looks like libvirt interface is slow
14:14:49 kashyap alex_xu: Yeah, there's a new thread that started a day ago
14:14:56 kashyap Did you see that? About the new RFC / design
14:15:36 alex_xu kashyap: not yet, another team member follow that, probably I will tell him
14:16:35 mdbooth melwitt: You about? Spotted something potentially funky looking at your change https://review.openstack.org/#/c/498983/
14:16:48 mdbooth Not directly related to your patch
14:17:13 alex_xu kashyap: thanks for the info
14:17:23 mdbooth I couldn't see where we wrap swap_volume in any kind of lock
14:17:41 kashyap alex_xu: Here it is: "Yet another RFC for [Intel] CAT" - https://www.redhat.com/archives/libvir-list/2017-September/msg00072.html
14:17:43 mdbooth e.g. by setting task_state
14:17:53 mdbooth I expect I just missed it, though.
14:18:14 mdbooth Do you happen to know if there is any locking around that method?
14:19:00 mdbooth If there isn't, the save/restore xml after a long-running rebase mechanism would be potentially pretty bad.
14:19:14 kashyap mdbooth: I read your comment this morning, and did find your task exclusion observation interesting
14:19:25 gibi mriedem: could you confirm that we want to delete the allocation of the instance during _destroy_evacuated_instances if the instance failed an evacuation?
14:19:35 gibi mriedem: context is in https://review.openstack.org/#/c/499237/
14:19:36 kashyap Hmm, although the patch is a strict improvement as-is, now I wonder about your comment - where is the bug lurking...
14:19:50 dansmith mriedem: I'm going to be in an airport when the cells meeting is going on
14:19:57 dansmith mriedem: I can try to run it from there, but no promises
14:20:03 alex_xu kashyap: thanks!
14:20:25 mriedem dansmith: let's just skip
14:20:35 dansmith mriedem: cool
14:20:56 gibi mriedem: I'm asking because then a subsequent rebuild needs to allocate resources again
14:22:56 mriedem gibi: hmm, and a rebuild to the same host won't do a claim, and the RT won't auto-heal the allocations on the node
14:23:05 mriedem and rebuild (not evacuate) bypasses the scheduler
14:23:22 mriedem that's a good point
14:23:39 gibi mriedem: yes, that is my fear
14:24:03 gibi mriedem: in the other hand we could put the instance to error state after the failed evac and keep the allocation
14:24:17 gibi mriedem: this way the rebuild doesn't need to allocate
14:24:23 mriedem i'm in the middle of something and will have to digest the replies in that review later
14:24:30 mriedem gibi: we might also just have to leave this until the ptg
14:24:37 gibi mriedem: sure
14:24:59 mriedem could you add it to the ptg etherpad in case it has to wait until next week?
14:25:10 gibi mriedem: adding...
14:26:23 hogepodge mriedem: initial schedule is up, let me know if times work for you, leave notes if things need to be shuffled. It's all preliminary right now and subject to change. https://etherpad.openstack.org/p/InteropDenver2017PTG
14:27:52 mriedem hogepodge: ok thanks
14:32:31 gibi mriedem: I added it to the etherpad
14:36:40 mriedem gibi: thanks
14:40:14 hogepodge dtantsur|afk: ^^
14:41:46 openstackgerrit Hongbin Lu proposed openstack/nova master: Handle exception on adding secgroup https://review.openstack.org/465173
14:41:51 openstackgerrit Merged openstack/nova master: rbd: Remove unnecessary 'encode' calls https://review.openstack.org/412356
14:42:32 stephenfin Hurrah ^
14:44:13 openstackgerrit Merged openstack/nova master: Remove qpid description in doc https://review.openstack.org/499087
14:54:19 dansmith mriedem: dtantsur|afk: see this ironic migration thing? https://review.openstack.org/#/c/501025/2
14:56:57 mriedem now i do
14:58:31 esberglu I'm getting the following unauthorized command error from nova-rootwrap intermittently
14:58:32 esberglu http://paste.openstack.org/show/620547/
14:58:58 mriedem sdague: looks like https://review.openstack.org/#/c/457636/ is happy - the devstack patch to install the osc-placement plugin
14:59:01 mriedem http://logs.openstack.org/36/457636/9/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/ad2b234/logs/pip2-freeze.txt.gz
14:59:08 esberglu In /etc/nova/rootwrap.d/compute.filters I have the following line
14:59:09 esberglu tee: CommandFilter, tee, root
14:59:25 esberglu Anyone have any idea why sometimes that filter is failing to match?
15:03:25 sdague mriedem: yep

Earlier   Later