| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-28 | |||
| 11:01:30 | cdent | bhagyashris: are you working from today’s master or something earlier? | |
| 11:02:29 | bhagyashris | cdent: yeah. I am working from todays master commit-id: 23c4eb34380bdf3eece11abbe0f6ccb68c060f47 "claim resources in placement API during schedule()" | |
| 11:07:10 | cdent | bhagyashris: do you have the log of it making the PUT /allocations requests? It might be that the scheduler is claimng one set of allocations, but later (in the periodic update) the resource tracker is changing things. if you’re going to be using the CUSTOM_DISK resource class you need to make sure your flavor doesn’t indicate any requests for disk_gb | |
| 11:07:53 | cdent | I think over the next few days things are likely to be quite unstable with claims, and there will be a few more changes coming | |
| 11:09:48 | bhagyashris | cdent: ohk. | |
| 11:10:28 | cdent | bhagyashris: it’s great that you’re doing this investigating, but right now is going to be a chaotic time | |
| 11:13:15 | cdent | bauzas: if this still alive? https://review.openstack.org/#/c/427200/ | |
| 11:13:24 | bauzas | cdent: looking | |
| 11:13:47 | bauzas | cdent: well, maybe we should discuss with matt | |
| 11:13:58 | bauzas | I can rebase it | |
| 11:14:05 | bauzas | meanwhile | |
| 11:14:26 | cdent | k, thanks | |
| 11:16:26 | bhagyashris | cdent: ohk. Just to inform you I have checked the allocations table entries and in that I saw the disk_gb reported against the compute node resource provider and not the shared resource provider. | |
| 11:16:59 | bhagyashris | cdent: Thank you for your time :) | |
| 11:17:19 | cdent | yeah, in the case where you have two different provider of DISK_GB it is likely that the non-shared provider will be picked when there are multiple options. | |
| 11:24:56 | bhagyashris | cdent: ohk. | |
| 11:25:58 | openstackgerrit | Sean Dague proposed openstack/nova master: always show urls in list_cells https://review.openstack.org/487860 | |
| 11:36:10 | gibi | cdent: yes, I have the same feeling about the host name confusion | |
| 11:36:27 | gibi | cdent: but I cannot spend more time on it today. I will continue looking into it on Monday | |
| 11:36:42 | cdent | gibi: I also left a comment on your test about another thing worth looking into | |
| 11:36:56 | cdent | if I have time today I may experiment some more, but not sure I’ll have time | |
| 11:37:06 | cdent | thanks for messing with it | |
| 11:37:38 | gibi | thanks for thelping | |
| 11:38:21 | gibi | if you make some progress still today then do not hesistate to update any of my reviews | |
| 11:43:58 | openstackgerrit | Sylvain Bauza proposed openstack/nova master: Add a status check for enabled filters https://review.openstack.org/427200 | |
| 11:44:24 | bauzas | cdent: ^ | |
| 11:44:34 | cdent | ✔ | |
| 11:44:40 | bauzas | after reviewing it, I think it's important to merge it | |
| 11:44:43 | bauzas | for Pike | |
| 11:46:28 | openstackgerrit | Chris Dent proposed openstack/nova master: Add functional test for two-cell scheduler behaviors https://review.openstack.org/452006 | |
| 11:49:01 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] gabbi tests for shared custom resource class https://review.openstack.org/485209 | |
| 11:50:46 | cdent | bauzas: there’s quite a lot of small things that either fix minor bugs, cover areas not previously covered, or document things. all of which would be nice for pike. update 30 will be ready in a few minutes. If you and others can get to some of the things in the “other” section that would be very useful. | |
| 11:51:21 | bauzas | cdent: sorry, I don't understand you | |
| 11:51:30 | bauzas | cdent: you want me to do what ? review ? | |
| 11:52:20 | cdent | bauzas: a) I’m agree with hou on the import to merge 427200 and b) saying there’s lots of similar things, which will be listed in the # Other section rp update 30 coming out soon, which will need review and ought to be merged | |
| 11:52:40 | bauzas | ah, yes, agreed on that | |
| 11:53:05 | bauzas | we need to make sure we use those two weeks correctly, ie. not for features but rather bugs and/or docs | |
| 11:53:19 | bauzas | tbc, I'll mostly do bug triage next week | |
| 11:53:28 | bauzas | and the rest of my time will be for reviewing | |
| 12:03:47 | sdague | cdent / bauzas / jay - just an FYI of something I closed as won't fix - https://bugs.launchpad.net/nova/+bug/1707085 - in case you have other opinions | |
| 12:03:48 | openstack | Launchpad bug 1707085 in OpenStack Compute (nova) "Max_unit should account for allocation_ratio" [Undecided,Won't fix] | |
| 12:09:35 | cdent | sdague: your response makes sense to me, and max_unit is being set as it for the reasons you describe | |
| 12:23:26 | cdent | edleafe: left a comment on https://review.openstack.org/#/c/487954/ Just trying to make your life harder (as usual) | |
| 12:24:27 | edleafe | cdent: saw it. I'm not sure that any instance on an ironic node would ever have any other custom resource class. I mean, that's the point, right? | |
| 12:24:55 | cdent | edleafe: dunno | |
| 12:25:13 | cdent | it’s the internet. if it is possible to do something, people will | |
| 12:25:28 | edleafe | cdent: what I *do* know is that operators can change the CRC of a node | |
| 12:25:39 | edleafe | cdent: and like you said, if they can, they will :) | |
| 12:26:41 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/477091 | |
| 12:27:35 | openstackgerrit | Doug Hellmann proposed openstack/nova master: add a redirect for the old cells landing page https://review.openstack.org/487932 | |
| 12:48:11 | sdague | cdent: is this still a thing? https://bugs.launchpad.net/nova/+bug/1659647 | |
| 12:48:12 | openstack | Launchpad bug 1659647 in OpenStack Compute (nova) "The resource tracker clears the tracked_instances dictionary on every periodic job" [Undecided,New] | |
| 12:49:13 | cdent | sdague: yeah, that came up in the hangout last night | |
| 12:49:47 | cdent | fixing it as described there would break “healing” but it’s now no longer clear that perfect healing is what we want | |
| 12:50:06 | cdent | it did get brought up in atlanta, but unresolved. will probably need to come up again in denver | |
| 12:50:46 | s-dean | Hi there, in the nova_cell0 database is the instanaces tables supposed to be populated? | |
| 12:51:18 | efried | s-dean As I understand it, that table only houses instances that failed to be scheduled. | |
| 12:51:43 | efried | s-dean I learned that from https://review.openstack.org/#/c/487183/ | |
| 12:52:15 | s-dean | ok thanks, im trying to track down why i can list my instances with openstack server list | |
| 12:52:17 | sdague | cdent: ok, so I'm going to close it as Opinion Wishlist then, as it's one of those notes to self that's getting tracked in other ways (like via the etherpad) | |
| 12:52:32 | cdent | sdague: seems fair | |
| 12:53:28 | efried | s-dean There's a whole thing I don't understand about deleted instances not really being deleted, some kind of "soft" delete. | |
| 12:54:51 | s-dean | No i haven't deleted the instances yet, the command is hanging and not returning anything however the API server is returning 200 OK | |
| 12:56:19 | efried | s-dean does `nova list` work? | |
| 12:57:08 | s-dean | also in the instances table(nova DB), the cell_name filed is NULL | |
| 12:57:21 | s-dean | nope | |
| 12:57:57 | s-dean | that is also hanging, its like it cant access the database; | |
| 13:00:51 | s-dean | also there is no response | |
| 13:07:49 | bauzas | sdague: looking up bugs ? I'm about to do that too | |
| 13:08:18 | sdague | bauzas: yeh, I've been trying to burn the New list down | |
| 13:08:21 | sdague | getting close | |
| 13:08:24 | bauzas | k | |
| 13:09:00 | bauzas | I'll also look at the in-progress bugs | |
| 13:09:10 | bauzas | to make sure we won't miss any important one | |
| 13:29:46 | mriedem | pike-3 tag is in https://review.openstack.org/#/c/488218/ | |
| 13:35:59 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187 | |
| 13:36:45 | cdent | mriedem: can you summarize what you leared while doing that ^? | |
| 13:37:30 | mriedem | cdent: well it does two things | |
| 13:38:12 | mriedem | 1. when the RT (source node during a migration) untracks it's instances and attempts to delete allocations, we get the current allocations for the instance for all providers and if the source node rp uuid isn't in the current list of allocations, it noops | |
| 13:38:37 | mriedem | 2. if the source node rp uuid is in the list of current allocations but there is also at least one other VCPU resource provider, it logs a message before deleting the allocations | |
| 13:38:48 | mriedem | i have seen #2 in the live migration CI runs, | |
| 13:38:51 | mriedem | i haven't seen #1 | |
| 13:39:04 | mriedem | http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/107a810/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_27_22_21_14_766229 | |
| 13:39:45 | dansmith | yeah so #2 will get reported as a bug and tell us that we need to skip instead of log and continue there | |
| 13:39:52 | mriedem | i really just pushed this to see if things were really bad, like if we hit #1 a lot | |
| 13:40:46 | dansmith | mriedem: you could try setting the heal interval really small so it runs more often during the few migrations we actually do | |
| 13:40:50 | cdent | do you feel better or worse? | |
| 13:41:03 | mriedem | i feel better | |
| 13:41:18 | mriedem | #1 would be a scarier race imo | |
| 13:41:37 | mriedem | and could still happen, but chances are slim - at least in our CI env which isn't that load intensive | |
| 13:41:54 | mriedem | #1 happens after the source node RT pulls the allocations for it's own uuid, | |
| 13:42:19 | mriedem | so to hit #1, that means the allocations are gone for that node between the time it pulls them and the time it processes it's list of untracked instances | |
| 13:42:45 | mriedem | i think #2 is fixed by making the RT check the vm_state of the instance before trying to delete it's allocations | |
| 13:42:49 | mriedem | which is something jay was working on yesterday | |
| 13:42:51 | leakypipes | for the record, guys, I should be done with my patch shortly. | |
| 13:43:00 | bauzas | mriedem: yup, saw the pike-3 change this morning | |
| 13:43:02 | leakypipes | just working on unit tests now | |
| 13:43:15 | bauzas | it was already merged | |
| 13:43:19 | dansmith | mriedem: insert some artificial rpc delay into the periodic and see if #1 happens | |
| 13:43:41 | mriedem | i can do that | |