| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-06 | |||
| 14:01:57 | mriedem | bauzas: so maybe what you want to see in https://review.openstack.org/#/c/499399/ is a release note that says, evacuate with force=True may fail due to placement | |
| 14:02:18 | mriedem | and at some point an update to the api-ref for the force flag on evacuate and live migrate | |
| 14:02:44 | mriedem | i.e. the force flag bypasses the scheduler filters but can still fail due to placement | |
| 14:03:57 | mriedem | we merged the same for live migrate for the pike GA but didn't have a release note for that https://review.openstack.org/#/c/496727/ | |
| 14:04:05 | mriedem | i could add a release note for the 16.0.1 release | |
| 14:04:42 | bauzas | mriedem: so, I agree with you about having "something" that would reconcile allocations if you force | |
| 14:05:06 | mriedem | since the RT isn't going to do that once all computes are upgraded to pike, | |
| 14:05:18 | bauzas | the fact is, some operators for example want to live-migrate a whole host to another one with the price of degradation | |
| 14:05:22 | dansmith | we can audit without heal | |
| 14:05:24 | mriedem | it's better to try and allocate in the controller and fail fast than do it in the compute after we've already live migrated or evacuated something | |
| 14:05:41 | bauzas | because that's an emergency situation | |
| 14:05:53 | bauzas | and they don't want their customers to know that the host is bad | |
| 14:06:07 | mriedem | if they are live migrating, the customer doesn't know | |
| 14:06:11 | mriedem | if it moved or not | |
| 14:06:38 | dansmith | well, the customer will know of course | |
| 14:06:58 | bauzas | sure, but they want to migrate sooner than later and somehow want to migrate the host to some other they know | |
| 14:06:58 | dansmith | minimum disruption, but that's all | |
| 14:07:24 | bauzas | agreed | |
| 14:07:37 | mriedem | btw, unrelated, but can we get this merged https://review.openstack.org/#/c/500968/ to unblock the requirements upper-constraints change from merging? | |
| 14:07:46 | mriedem | ^ is needed to get nova unit tests to pass with the latest os-xenapi lib | |
| 14:07:46 | dansmith | I'm not sure I understand what the concern is, I just started reading | |
| 14:07:48 | dansmith | all we have to do is make sure force does a caim first, and if that fails, fail, right? | |
| 14:08:21 | mriedem | dansmith: which is exactly what this does https://review.openstack.org/#/c/496727/ | |
| 14:08:34 | mriedem | and https://review.openstack.org/#/c/499399/ | |
| 14:08:38 | dansmith | right | |
| 14:09:10 | mriedem | we're really talking about two things and diverged from the first, which is bauzas wants to consider moving the claim code out of the scheduler and into conductor | |
| 14:09:16 | mriedem | and then we went down the force flag rathole | |
| 14:09:25 | alex_xu | jamespage: bauzas there also a proposal in the intel about integrate an agent which managing cpu l3 cache into the nova. I'm also thinking report the l3 cache inventory by the agent to the placement directly is option, but there missed a way for the agent to know new instance boot up in the node except polling the libvirt | |
| 14:09:43 | mriedem | alex_xu: listen for notifications | |
| 14:09:46 | alex_xu | jaypipes: ^ | |
| 14:09:55 | alex_xu | jamespage: sorry for pointed to wrong person | |
| 14:10:15 | alex_xu | mriedem: but the notification isn't for single host, we must filter the info | |
| 14:10:18 | mriedem | alex_xu: if you want a hypervisor agnostic solution, listen for notifications rather than polling libvirt directly | |
| 14:10:18 | jaypipes | alex_xu: you mean RDT/CAT, right? | |
| 14:10:25 | alex_xu | jaypipes: yea | |
| 14:11:24 | mriedem | alex_xu: the instance.create.end payload would have the host/node information in it so yeah you'd have to filter but that doesn't seem hard | |
| 14:11:31 | bauzas | I don't get why it requires libvirt to be polled | |
| 14:12:02 | bauzas | and I second mriedem about subscribing to the notifications bus | |
| 14:12:30 | kashyap | bauzas: What requires libvirt to be polled? /me read it out of context; will read the scrollback | |
| 14:12:37 | alex_xu | mriedem: yea, that won't be so hard, that probably each node have a agent to listen the notification | |
| 14:12:41 | mriedem | bauzas: xenserver ci isn't going to care about unit test changes https://review.openstack.org/#/c/500968/ | |
| 14:12:51 | bauzas | holy ship, you're right | |
| 14:13:05 | alex_xu | bauzas: kashyap an external agent try to know there is new instance boot up | |
| 14:13:16 | mriedem | alex_xu: plus then you can use the fancy new versioned notifications and provide feedback to gibi | |
| 14:13:39 | alex_xu | mriedem: yeah | |
| 14:14:02 | gibi | yeah :) | |
| 14:14:03 | kashyap | alex_xu: Ah, it's the L3 CAT interface. I see some in-progress design work of exposing it in libvirt upstream | |
| 14:14:23 | alex_xu | kashyap: yea, looks like libvirt interface is slow | |
| 14:14:49 | kashyap | alex_xu: Yeah, there's a new thread that started a day ago | |
| 14:14:56 | kashyap | Did you see that? About the new RFC / design | |
| 14:15:36 | alex_xu | kashyap: not yet, another team member follow that, probably I will tell him | |
| 14:16:35 | mdbooth | melwitt: You about? Spotted something potentially funky looking at your change https://review.openstack.org/#/c/498983/ | |
| 14:16:48 | mdbooth | Not directly related to your patch | |
| 14:17:13 | alex_xu | kashyap: thanks for the info | |
| 14:17:23 | mdbooth | I couldn't see where we wrap swap_volume in any kind of lock | |
| 14:17:41 | kashyap | alex_xu: Here it is: "Yet another RFC for [Intel] CAT" - https://www.redhat.com/archives/libvir-list/2017-September/msg00072.html | |
| 14:17:43 | mdbooth | e.g. by setting task_state | |
| 14:17:53 | mdbooth | I expect I just missed it, though. | |
| 14:18:14 | mdbooth | Do you happen to know if there is any locking around that method? | |
| 14:19:00 | mdbooth | If there isn't, the save/restore xml after a long-running rebase mechanism would be potentially pretty bad. | |
| 14:19:14 | kashyap | mdbooth: I read your comment this morning, and did find your task exclusion observation interesting | |
| 14:19:25 | gibi | mriedem: could you confirm that we want to delete the allocation of the instance during _destroy_evacuated_instances if the instance failed an evacuation? | |
| 14:19:35 | gibi | mriedem: context is in https://review.openstack.org/#/c/499237/ | |
| 14:19:36 | kashyap | Hmm, although the patch is a strict improvement as-is, now I wonder about your comment - where is the bug lurking... | |
| 14:19:50 | dansmith | mriedem: I'm going to be in an airport when the cells meeting is going on | |
| 14:19:57 | dansmith | mriedem: I can try to run it from there, but no promises | |
| 14:20:03 | alex_xu | kashyap: thanks! | |
| 14:20:25 | mriedem | dansmith: let's just skip | |
| 14:20:35 | dansmith | mriedem: cool | |
| 14:20:56 | gibi | mriedem: I'm asking because then a subsequent rebuild needs to allocate resources again | |
| 14:22:56 | mriedem | gibi: hmm, and a rebuild to the same host won't do a claim, and the RT won't auto-heal the allocations on the node | |
| 14:23:05 | mriedem | and rebuild (not evacuate) bypasses the scheduler | |
| 14:23:22 | mriedem | that's a good point | |
| 14:23:39 | gibi | mriedem: yes, that is my fear | |
| 14:24:03 | gibi | mriedem: in the other hand we could put the instance to error state after the failed evac and keep the allocation | |
| 14:24:17 | gibi | mriedem: this way the rebuild doesn't need to allocate | |
| 14:24:23 | mriedem | i'm in the middle of something and will have to digest the replies in that review later | |
| 14:24:30 | mriedem | gibi: we might also just have to leave this until the ptg | |
| 14:24:37 | gibi | mriedem: sure | |
| 14:24:59 | mriedem | could you add it to the ptg etherpad in case it has to wait until next week? | |
| 14:25:10 | gibi | mriedem: adding... | |
| 14:26:23 | hogepodge | mriedem: initial schedule is up, let me know if times work for you, leave notes if things need to be shuffled. It's all preliminary right now and subject to change. https://etherpad.openstack.org/p/InteropDenver2017PTG | |
| 14:27:52 | mriedem | hogepodge: ok thanks | |
| 14:32:31 | gibi | mriedem: I added it to the etherpad | |
| 14:36:40 | mriedem | gibi: thanks | |
| 14:40:14 | hogepodge | dtantsur|afk: ^^ | |
| 14:41:46 | openstackgerrit | Hongbin Lu proposed openstack/nova master: Handle exception on adding secgroup https://review.openstack.org/465173 | |
| 14:41:51 | openstackgerrit | Merged openstack/nova master: rbd: Remove unnecessary 'encode' calls https://review.openstack.org/412356 | |
| 14:42:32 | stephenfin | Hurrah ^ | |
| 14:44:13 | openstackgerrit | Merged openstack/nova master: Remove qpid description in doc https://review.openstack.org/499087 | |
| 14:54:19 | dansmith | mriedem: dtantsur|afk: see this ironic migration thing? https://review.openstack.org/#/c/501025/2 | |
| 14:56:57 | mriedem | now i do | |
| 14:58:31 | esberglu | I'm getting the following unauthorized command error from nova-rootwrap intermittently | |
| 14:58:32 | esberglu | http://paste.openstack.org/show/620547/ | |
| 14:58:58 | mriedem | sdague: looks like https://review.openstack.org/#/c/457636/ is happy - the devstack patch to install the osc-placement plugin | |
| 14:59:01 | mriedem | http://logs.openstack.org/36/457636/9/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/ad2b234/logs/pip2-freeze.txt.gz | |
| 14:59:08 | esberglu | In /etc/nova/rootwrap.d/compute.filters I have the following line | |
| 14:59:09 | esberglu | tee: CommandFilter, tee, root | |
| 14:59:25 | esberglu | Anyone have any idea why sometimes that filter is failing to match? | |
| 15:03:25 | sdague | mriedem: yep | |