Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-28
11:49:01 openstackgerrit Chris Dent proposed openstack/nova master: [placement] gabbi tests for shared custom resource class https://review.openstack.org/485209
11:50:46 cdent bauzas: there’s quite a lot of small things that either fix minor bugs, cover areas not previously covered, or document things. all of which would be nice for pike. update 30 will be ready in a few minutes. If you and others can get to some of the things in the “other” section that would be very useful.
11:51:21 bauzas cdent: sorry, I don't understand you
11:51:30 bauzas cdent: you want me to do what ? review ?
11:52:20 cdent bauzas: a) I’m agree with hou on the import to merge 427200 and b) saying there’s lots of similar things, which will be listed in the # Other section rp update 30 coming out soon, which will need review and ought to be merged
11:52:40 bauzas ah, yes, agreed on that
11:53:05 bauzas we need to make sure we use those two weeks correctly, ie. not for features but rather bugs and/or docs
11:53:19 bauzas tbc, I'll mostly do bug triage next week
11:53:28 bauzas and the rest of my time will be for reviewing
12:03:47 sdague cdent / bauzas / jay - just an FYI of something I closed as won't fix - https://bugs.launchpad.net/nova/+bug/1707085 - in case you have other opinions
12:03:48 openstack Launchpad bug 1707085 in OpenStack Compute (nova) "Max_unit should account for allocation_ratio" [Undecided,Won't fix]
12:09:35 cdent sdague: your response makes sense to me, and max_unit is being set as it for the reasons you describe
12:23:26 cdent edleafe: left a comment on https://review.openstack.org/#/c/487954/ Just trying to make your life harder (as usual)
12:24:27 edleafe cdent: saw it. I'm not sure that any instance on an ironic node would ever have any other custom resource class. I mean, that's the point, right?
12:24:55 cdent edleafe: dunno
12:25:13 cdent it’s the internet. if it is possible to do something, people will
12:25:28 edleafe cdent: what I *do* know is that operators can change the CRC of a node
12:25:39 edleafe cdent: and like you said, if they can, they will :)
12:26:41 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/477091
12:27:35 openstackgerrit Doug Hellmann proposed openstack/nova master: add a redirect for the old cells landing page https://review.openstack.org/487932
12:48:11 sdague cdent: is this still a thing? https://bugs.launchpad.net/nova/+bug/1659647
12:48:12 openstack Launchpad bug 1659647 in OpenStack Compute (nova) "The resource tracker clears the tracked_instances dictionary on every periodic job" [Undecided,New]
12:49:13 cdent sdague: yeah, that came up in the hangout last night
12:49:47 cdent fixing it as described there would break “healing” but it’s now no longer clear that perfect healing is what we want
12:50:06 cdent it did get brought up in atlanta, but unresolved. will probably need to come up again in denver
12:50:46 s-dean Hi there, in the nova_cell0 database is the instanaces tables supposed to be populated?
12:51:18 efried s-dean As I understand it, that table only houses instances that failed to be scheduled.
12:51:43 efried s-dean I learned that from https://review.openstack.org/#/c/487183/
12:52:15 s-dean ok thanks, im trying to track down why i can list my instances with openstack server list
12:52:17 sdague cdent: ok, so I'm going to close it as Opinion Wishlist then, as it's one of those notes to self that's getting tracked in other ways (like via the etherpad)
12:52:32 cdent sdague: seems fair
12:53:28 efried s-dean There's a whole thing I don't understand about deleted instances not really being deleted, some kind of "soft" delete.
12:54:51 s-dean No i haven't deleted the instances yet, the command is hanging and not returning anything however the API server is returning 200 OK
12:56:19 efried s-dean does `nova list` work?
12:57:08 s-dean also in the instances table(nova DB), the cell_name filed is NULL
12:57:21 s-dean nope
12:57:57 s-dean that is also hanging, its like it cant access the database;
13:00:51 s-dean also there is no response
13:07:49 bauzas sdague: looking up bugs ? I'm about to do that too
13:08:18 sdague bauzas: yeh, I've been trying to burn the New list down
13:08:21 sdague getting close
13:08:24 bauzas k
13:09:00 bauzas I'll also look at the in-progress bugs
13:09:10 bauzas to make sure we won't miss any important one
13:29:46 mriedem pike-3 tag is in https://review.openstack.org/#/c/488218/
13:35:59 openstackgerrit Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187
13:36:45 cdent mriedem: can you summarize what you leared while doing that ^?
13:37:30 mriedem cdent: well it does two things
13:38:12 mriedem 1. when the RT (source node during a migration) untracks it's instances and attempts to delete allocations, we get the current allocations for the instance for all providers and if the source node rp uuid isn't in the current list of allocations, it noops
13:38:37 mriedem 2. if the source node rp uuid is in the list of current allocations but there is also at least one other VCPU resource provider, it logs a message before deleting the allocations
13:38:48 mriedem i have seen #2 in the live migration CI runs,
13:38:51 mriedem i haven't seen #1
13:39:04 mriedem http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/107a810/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_27_22_21_14_766229
13:39:45 dansmith yeah so #2 will get reported as a bug and tell us that we need to skip instead of log and continue there
13:39:52 mriedem i really just pushed this to see if things were really bad, like if we hit #1 a lot
13:40:46 dansmith mriedem: you could try setting the heal interval really small so it runs more often during the few migrations we actually do
13:40:50 cdent do you feel better or worse?
13:41:03 mriedem i feel better
13:41:18 mriedem #1 would be a scarier race imo
13:41:37 mriedem and could still happen, but chances are slim - at least in our CI env which isn't that load intensive
13:41:54 mriedem #1 happens after the source node RT pulls the allocations for it's own uuid,
13:42:19 mriedem so to hit #1, that means the allocations are gone for that node between the time it pulls them and the time it processes it's list of untracked instances
13:42:45 mriedem i think #2 is fixed by making the RT check the vm_state of the instance before trying to delete it's allocations
13:42:49 mriedem which is something jay was working on yesterday
13:42:51 leakypipes for the record, guys, I should be done with my patch shortly.
13:43:00 bauzas mriedem: yup, saw the pike-3 change this morning
13:43:02 leakypipes just working on unit tests now
13:43:15 bauzas it was already merged
13:43:19 dansmith mriedem: insert some artificial rpc delay into the periodic and see if #1 happens
13:43:41 mriedem i can do that
13:44:30 mriedem cdent: bauzas: also https://bugs.launchpad.net/nova/+bug/1707071 if you haven't seen that yet
13:44:31 openstack Launchpad bug 1707071 in OpenStack Compute (nova) ocata "Compute nodes will fight over allocations during migration" [Medium,Confirmed]
13:44:40 bauzas not yet indeed
13:44:57 bauzas mriedem: dansmith: leakypipes: I thought about something this night
13:45:04 bauzas cdent: ^
13:45:13 leakypipes dansmith: you need to change your clothes.
13:45:31 bauzas what if we should just get the current allocations before deleting them by the scheduler, so in case we know we have a problem, we could put them again ?
13:45:32 dansmith heh
13:45:42 bauzas oh snap, FF
13:46:07 cdent thanks mriedem
13:46:48 superdan bauwser: get + check is just as safe as get + delete + re-put, except you never have to delete
13:46:49 superdan not sure why we would do the latter
13:47:06 bauwser mriedem: superdan: leakypipes: oh, and me and cdent just discussed about https://review.openstack.org/#/c/427200/ : it could be important for operators
13:47:50 bauwser superdan: sure, was just wondering if it was simplier than just trying to get two allocations for both source and target when moving
13:48:00 bauwser I know it's the consensus
13:48:05 superdan I don't see how it helps
13:48:06 mriedem bauwser: yes there is an odd scenario in there which i commented on
13:48:06 bauwser and I totally agree with it
13:48:14 superdan without two allocations we're not accounting for the resources used by a moving instance
13:48:19 bauwser but I do wonder if we could simplify the problem
13:48:33 bauwser at least for the races we know of
13:48:48 superdan well, the double allocation. there's still only one allocation for the instance
13:49:08 bauwser the problem here is that we have 3 different services looking at placement when moving : #1 scheduler, #2 source compute, #3 target compute
13:49:21 superdan that's kindof the point of placement right?
13:49:40 bauwser yeah, I know, but that means we have like shared information
13:49:56 mriedem that's the point
13:49:56 bauwser but I think we discussed that yesterday
13:49:57 mriedem the global view
13:50:03 superdan yeah, that's the whole thing
13:50:11 bauwser anyway, I don't want to nitpick

Earlier   Later