Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-28
12:52:15 s-dean ok thanks, im trying to track down why i can list my instances with openstack server list
12:52:17 sdague cdent: ok, so I'm going to close it as Opinion Wishlist then, as it's one of those notes to self that's getting tracked in other ways (like via the etherpad)
12:52:32 cdent sdague: seems fair
12:53:28 efried s-dean There's a whole thing I don't understand about deleted instances not really being deleted, some kind of "soft" delete.
12:54:51 s-dean No i haven't deleted the instances yet, the command is hanging and not returning anything however the API server is returning 200 OK
12:56:19 efried s-dean does `nova list` work?
12:57:08 s-dean also in the instances table(nova DB), the cell_name filed is NULL
12:57:21 s-dean nope
12:57:57 s-dean that is also hanging, its like it cant access the database;
13:00:51 s-dean also there is no response
13:07:49 bauzas sdague: looking up bugs ? I'm about to do that too
13:08:18 sdague bauzas: yeh, I've been trying to burn the New list down
13:08:21 sdague getting close
13:08:24 bauzas k
13:09:00 bauzas I'll also look at the in-progress bugs
13:09:10 bauzas to make sure we won't miss any important one
13:29:46 mriedem pike-3 tag is in https://review.openstack.org/#/c/488218/
13:35:59 openstackgerrit Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187
13:36:45 cdent mriedem: can you summarize what you leared while doing that ^?
13:37:30 mriedem cdent: well it does two things
13:38:12 mriedem 1. when the RT (source node during a migration) untracks it's instances and attempts to delete allocations, we get the current allocations for the instance for all providers and if the source node rp uuid isn't in the current list of allocations, it noops
13:38:37 mriedem 2. if the source node rp uuid is in the list of current allocations but there is also at least one other VCPU resource provider, it logs a message before deleting the allocations
13:38:48 mriedem i have seen #2 in the live migration CI runs,
13:38:51 mriedem i haven't seen #1
13:39:04 mriedem http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/107a810/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_27_22_21_14_766229
13:39:45 dansmith yeah so #2 will get reported as a bug and tell us that we need to skip instead of log and continue there
13:39:52 mriedem i really just pushed this to see if things were really bad, like if we hit #1 a lot
13:40:46 dansmith mriedem: you could try setting the heal interval really small so it runs more often during the few migrations we actually do
13:40:50 cdent do you feel better or worse?
13:41:03 mriedem i feel better
13:41:18 mriedem #1 would be a scarier race imo
13:41:37 mriedem and could still happen, but chances are slim - at least in our CI env which isn't that load intensive
13:41:54 mriedem #1 happens after the source node RT pulls the allocations for it's own uuid,
13:42:19 mriedem so to hit #1, that means the allocations are gone for that node between the time it pulls them and the time it processes it's list of untracked instances
13:42:45 mriedem i think #2 is fixed by making the RT check the vm_state of the instance before trying to delete it's allocations
13:42:49 mriedem which is something jay was working on yesterday
13:42:51 leakypipes for the record, guys, I should be done with my patch shortly.
13:43:00 bauzas mriedem: yup, saw the pike-3 change this morning
13:43:02 leakypipes just working on unit tests now
13:43:15 bauzas it was already merged
13:43:19 dansmith mriedem: insert some artificial rpc delay into the periodic and see if #1 happens
13:43:41 mriedem i can do that
13:44:30 mriedem cdent: bauzas: also https://bugs.launchpad.net/nova/+bug/1707071 if you haven't seen that yet
13:44:31 openstack Launchpad bug 1707071 in OpenStack Compute (nova) ocata "Compute nodes will fight over allocations during migration" [Medium,Confirmed]
13:44:40 bauzas not yet indeed
13:44:57 bauzas mriedem: dansmith: leakypipes: I thought about something this night
13:45:04 bauzas cdent: ^
13:45:13 leakypipes dansmith: you need to change your clothes.
13:45:31 bauzas what if we should just get the current allocations before deleting them by the scheduler, so in case we know we have a problem, we could put them again ?
13:45:32 dansmith heh
13:45:42 bauzas oh snap, FF
13:46:07 cdent thanks mriedem
13:46:48 superdan bauwser: get + check is just as safe as get + delete + re-put, except you never have to delete
13:46:49 superdan not sure why we would do the latter
13:47:06 bauwser mriedem: superdan: leakypipes: oh, and me and cdent just discussed about https://review.openstack.org/#/c/427200/ : it could be important for operators
13:47:50 bauwser superdan: sure, was just wondering if it was simplier than just trying to get two allocations for both source and target when moving
13:48:00 bauwser I know it's the consensus
13:48:05 superdan I don't see how it helps
13:48:06 bauwser and I totally agree with it
13:48:06 mriedem bauwser: yes there is an odd scenario in there which i commented on
13:48:14 superdan without two allocations we're not accounting for the resources used by a moving instance
13:48:19 bauwser but I do wonder if we could simplify the problem
13:48:33 bauwser at least for the races we know of
13:48:48 superdan well, the double allocation. there's still only one allocation for the instance
13:49:08 bauwser the problem here is that we have 3 different services looking at placement when moving : #1 scheduler, #2 source compute, #3 target compute
13:49:21 superdan that's kindof the point of placement right?
13:49:40 bauwser yeah, I know, but that means we have like shared information
13:49:56 bauwser but I think we discussed that yesterday
13:49:56 mriedem that's the point
13:49:57 mriedem the global view
13:50:03 superdan yeah, that's the whole thing
13:50:11 bauwser anyway, I don't want to nitpick
13:50:26 mriedem but yes that means old code that assumes it's all local and owns everything at any given time has to change
13:50:32 superdan it's not the shared state/view/responsibility that concerns me, it's that the current RT was designed for a different model
13:50:38 bauwser I'm just trying to see how to help with the problems we know about multiple services taking the same allocations
13:51:43 leakypipes hopefully my new patch's code comments explain the situation well enough.
13:51:47 bauwser ideally, eventually, I'm not sure we need resouretrackers for computes
13:52:01 superdan bauwser: we don't in the long run
13:52:01 bauwser if allocations would only be done by schedulers
13:52:15 superdan we still need some of that code on the computes,
13:52:19 superdan but not the full RT
13:52:23 bauwser I'd see RTs just *deleting* allocations if something goes mad
13:52:26 superdan that's what I was saying yesterday
13:52:28 bauwser superdan: yeah, I know
13:52:46 bauwser I'm just thinking out loud to try to identify how we could simplify
13:52:48 superdan so.... :)
13:52:57 superdan okay :)
13:53:01 bauwser but I agree with you, using the existing RT means tech deby
13:53:22 superdan Tech Deby.. that's like the host of a kids show about computers?
13:54:25 bauwser :)
13:54:43 bauwser my keyboard is AZERYT :p
13:55:37 superdan T and Y are together on both
13:55:49 superdan I thought it was AZERTY?
13:56:59 bauwser rather, yting
13:59:59 cdent can someone merge this https://review.openstack.org/#/c/488363/ looking at that warning is getting tiresome
14:05:28 kashyap superdan: Thanks for the review on this: https://review.openstack.org/#/c/485752/
14:05:38 kashyap superdan: And thanks for the little snark, too :P
14:06:10 superdan heh
14:06:56 kashyap Can anyone +W it? I can also backport it to the revlevant upstream branches
14:09:28 superdan kashyap: sdague loves +Wing patches like that

Earlier   Later