Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-28
10:54:36 bhagyashris cdent: It address the disk used against compute node resource provider.
10:58:23 vdrok good morning everyone!
10:58:38 vdrok sdague: small request, if you have some time :) https://review.openstack.org/480624
11:00:47 vdrok thanks!
11:01:01 cdent bhagyashris: ah, right, because the compue is reporting it’s own disk too, so is considered a valid target for all the resources. /me thinks
11:01:30 cdent bhagyashris: are you working from today’s master or something earlier?
11:02:29 bhagyashris cdent: yeah. I am working from todays master commit-id: 23c4eb34380bdf3eece11abbe0f6ccb68c060f47 "claim resources in placement API during schedule()"
11:07:10 cdent bhagyashris: do you have the log of it making the PUT /allocations requests? It might be that the scheduler is claimng one set of allocations, but later (in the periodic update) the resource tracker is changing things. if you’re going to be using the CUSTOM_DISK resource class you need to make sure your flavor doesn’t indicate any requests for disk_gb
11:07:53 cdent I think over the next few days things are likely to be quite unstable with claims, and there will be a few more changes coming
11:09:48 bhagyashris cdent: ohk.
11:10:28 cdent bhagyashris: it’s great that you’re doing this investigating, but right now is going to be a chaotic time
11:13:15 cdent bauzas: if this still alive? https://review.openstack.org/#/c/427200/
11:13:24 bauzas cdent: looking
11:13:47 bauzas cdent: well, maybe we should discuss with matt
11:13:58 bauzas I can rebase it
11:14:05 bauzas meanwhile
11:14:26 cdent k, thanks
11:16:26 bhagyashris cdent: ohk. Just to inform you I have checked the allocations table entries and in that I saw the disk_gb reported against the compute node resource provider and not the shared resource provider.
11:16:59 bhagyashris cdent: Thank you for your time :)
11:17:19 cdent yeah, in the case where you have two different provider of DISK_GB it is likely that the non-shared provider will be picked when there are multiple options.
11:24:56 bhagyashris cdent: ohk.
11:25:58 openstackgerrit Sean Dague proposed openstack/nova master: always show urls in list_cells https://review.openstack.org/487860
11:36:10 gibi cdent: yes, I have the same feeling about the host name confusion
11:36:27 gibi cdent: but I cannot spend more time on it today. I will continue looking into it on Monday
11:36:42 cdent gibi: I also left a comment on your test about another thing worth looking into
11:36:56 cdent if I have time today I may experiment some more, but not sure I’ll have time
11:37:06 cdent thanks for messing with it
11:37:38 gibi thanks for thelping
11:38:21 gibi if you make some progress still today then do not hesistate to update any of my reviews
11:43:58 openstackgerrit Sylvain Bauza proposed openstack/nova master: Add a status check for enabled filters https://review.openstack.org/427200
11:44:24 bauzas cdent: ^
11:44:34 cdent ✔
11:44:40 bauzas after reviewing it, I think it's important to merge it
11:44:43 bauzas for Pike
11:46:28 openstackgerrit Chris Dent proposed openstack/nova master: Add functional test for two-cell scheduler behaviors https://review.openstack.org/452006
11:49:01 openstackgerrit Chris Dent proposed openstack/nova master: [placement] gabbi tests for shared custom resource class https://review.openstack.org/485209
11:50:46 cdent bauzas: there’s quite a lot of small things that either fix minor bugs, cover areas not previously covered, or document things. all of which would be nice for pike. update 30 will be ready in a few minutes. If you and others can get to some of the things in the “other” section that would be very useful.
11:51:21 bauzas cdent: sorry, I don't understand you
11:51:30 bauzas cdent: you want me to do what ? review ?
11:52:20 cdent bauzas: a) I’m agree with hou on the import to merge 427200 and b) saying there’s lots of similar things, which will be listed in the # Other section rp update 30 coming out soon, which will need review and ought to be merged
11:52:40 bauzas ah, yes, agreed on that
11:53:05 bauzas we need to make sure we use those two weeks correctly, ie. not for features but rather bugs and/or docs
11:53:19 bauzas tbc, I'll mostly do bug triage next week
11:53:28 bauzas and the rest of my time will be for reviewing
12:03:47 sdague cdent / bauzas / jay - just an FYI of something I closed as won't fix - https://bugs.launchpad.net/nova/+bug/1707085 - in case you have other opinions
12:03:48 openstack Launchpad bug 1707085 in OpenStack Compute (nova) "Max_unit should account for allocation_ratio" [Undecided,Won't fix]
12:09:35 cdent sdague: your response makes sense to me, and max_unit is being set as it for the reasons you describe
12:23:26 cdent edleafe: left a comment on https://review.openstack.org/#/c/487954/ Just trying to make your life harder (as usual)
12:24:27 edleafe cdent: saw it. I'm not sure that any instance on an ironic node would ever have any other custom resource class. I mean, that's the point, right?
12:24:55 cdent edleafe: dunno
12:25:13 cdent it’s the internet. if it is possible to do something, people will
12:25:28 edleafe cdent: what I *do* know is that operators can change the CRC of a node
12:25:39 edleafe cdent: and like you said, if they can, they will :)
12:26:41 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/477091
12:27:35 openstackgerrit Doug Hellmann proposed openstack/nova master: add a redirect for the old cells landing page https://review.openstack.org/487932
12:48:11 sdague cdent: is this still a thing? https://bugs.launchpad.net/nova/+bug/1659647
12:48:12 openstack Launchpad bug 1659647 in OpenStack Compute (nova) "The resource tracker clears the tracked_instances dictionary on every periodic job" [Undecided,New]
12:49:13 cdent sdague: yeah, that came up in the hangout last night
12:49:47 cdent fixing it as described there would break “healing” but it’s now no longer clear that perfect healing is what we want
12:50:06 cdent it did get brought up in atlanta, but unresolved. will probably need to come up again in denver
12:50:46 s-dean Hi there, in the nova_cell0 database is the instanaces tables supposed to be populated?
12:51:18 efried s-dean As I understand it, that table only houses instances that failed to be scheduled.
12:51:43 efried s-dean I learned that from https://review.openstack.org/#/c/487183/
12:52:15 s-dean ok thanks, im trying to track down why i can list my instances with openstack server list
12:52:17 sdague cdent: ok, so I'm going to close it as Opinion Wishlist then, as it's one of those notes to self that's getting tracked in other ways (like via the etherpad)
12:52:32 cdent sdague: seems fair
12:53:28 efried s-dean There's a whole thing I don't understand about deleted instances not really being deleted, some kind of "soft" delete.
12:54:51 s-dean No i haven't deleted the instances yet, the command is hanging and not returning anything however the API server is returning 200 OK
12:56:19 efried s-dean does `nova list` work?
12:57:08 s-dean also in the instances table(nova DB), the cell_name filed is NULL
12:57:21 s-dean nope
12:57:57 s-dean that is also hanging, its like it cant access the database;
13:00:51 s-dean also there is no response
13:07:49 bauzas sdague: looking up bugs ? I'm about to do that too
13:08:18 sdague bauzas: yeh, I've been trying to burn the New list down
13:08:21 sdague getting close
13:08:24 bauzas k
13:09:00 bauzas I'll also look at the in-progress bugs
13:09:10 bauzas to make sure we won't miss any important one
13:29:46 mriedem pike-3 tag is in https://review.openstack.org/#/c/488218/
13:35:59 openstackgerrit Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187
13:36:45 cdent mriedem: can you summarize what you leared while doing that ^?
13:37:30 mriedem cdent: well it does two things
13:38:12 mriedem 1. when the RT (source node during a migration) untracks it's instances and attempts to delete allocations, we get the current allocations for the instance for all providers and if the source node rp uuid isn't in the current list of allocations, it noops
13:38:37 mriedem 2. if the source node rp uuid is in the list of current allocations but there is also at least one other VCPU resource provider, it logs a message before deleting the allocations
13:38:48 mriedem i have seen #2 in the live migration CI runs,
13:38:51 mriedem i haven't seen #1
13:39:04 mriedem http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/107a810/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_27_22_21_14_766229
13:39:45 dansmith yeah so #2 will get reported as a bug and tell us that we need to skip instead of log and continue there
13:39:52 mriedem i really just pushed this to see if things were really bad, like if we hit #1 a lot
13:40:46 dansmith mriedem: you could try setting the heal interval really small so it runs more often during the few migrations we actually do
13:40:50 cdent do you feel better or worse?
13:41:03 mriedem i feel better
13:41:18 mriedem #1 would be a scarier race imo
13:41:37 mriedem and could still happen, but chances are slim - at least in our CI env which isn't that load intensive
13:41:54 mriedem #1 happens after the source node RT pulls the allocations for it's own uuid,
13:42:19 mriedem so to hit #1, that means the allocations are gone for that node between the time it pulls them and the time it processes it's list of untracked instances
13:42:45 mriedem i think #2 is fixed by making the RT check the vm_state of the instance before trying to delete it's allocations
13:42:49 mriedem which is something jay was working on yesterday
13:42:51 leakypipes for the record, guys, I should be done with my patch shortly.

Earlier   Later