| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-10-09 | |||
| 14:47:01 | mriedem | maybe | |
| 14:47:07 | belmorei_ | mriedem dansmith jroll thanks for the help | |
| 14:47:13 | mriedem | jroll: around the time we deprecated the core/ram/disk filters | |
| 14:47:51 | jroll | mriedem: yeah I don't remember what either, just going by irc logs | |
| 14:47:52 | openstack | Launchpad bug 1787910 in OpenStack Compute (nova) rocky "OVB overcloud deploy fails on nova placement errors" [High,Fix committed] - Assigned to Matt Riedemann (mriedem) | |
| 14:47:52 | mriedem | yeah https://bugs.launchpad.net/tripleo/+bug/1787910 | |
| 14:52:16 | dansmith | mriedem: so they still had ram required in those flavors and failed when ironic stopped reporting ram inventory | |
| 14:52:17 | dansmith | yeah? | |
| 14:52:44 | dansmith | we have to cut over at some point and I thought we already had.. workaround config flag to let them get over the hump seems like the best thing at this point | |
| 14:54:27 | mriedem | dansmith: they being tripleo in that bug? | |
| 14:54:36 | dansmith | well, ovb but yeah | |
| 14:55:05 | mriedem | yeah looks like it based on https://review.openstack.org/#/c/596093/ | |
| 15:05:57 | dtantsur | belmorei_: a workaround may be to remove memory_mb and vcpus from ironic nodes properties | |
| 15:06:27 | dtantsur | with something like $ openstack baremetal node unset <node> --property memory_mb | |
| 15:07:20 | dtantsur | mriedem: this may be a bit easier than hacking nova ^^ | |
| 15:08:02 | belmorei_ | dtantsur: thanks, but the problem is the number of baremetal nodes that we have. Also, we would need to change the commission procedure to include that. | |
| 15:08:38 | belmorei_ | for now I will just patch this in nova | |
| 15:08:56 | dtantsur | belmorei_: do you use something like inspection to populate these properties? | |
| 15:09:36 | dtantsur | also before resource classes I used to use host aggregates to more or less separate bm and vm nodes on the same nova | |
| 15:11:35 | belmorei_ | dtantsur: yes, inspection populate them | |
| 15:13:47 | mriedem | belmorei_: out of curiosity, before cells v2, could your vm flavors get scheduled to bm cells/ | |
| 15:13:48 | mriedem | ? | |
| 15:14:05 | mriedem | or were the flavors segregated at the top cell layer? | |
| 15:17:22 | mriedem | belmorei_: also fyi, you can't set total=0 for inventory on the resource class as i said above, placement api will reject that since total must be >=1 | |
| 15:17:36 | mriedem | so need to just omit posting those non-custom-resource-class inventories | |
| 15:20:42 | belmorei_ | mriedem: with cellsV1 we were using the baremetal filters, so they will not be schedule to an already deployed node. But yes, if a user used a vm flavor a think it would be the same (it would use the physical node to create the vm flavor instance) | |
| 15:21:34 | belmorei_ | mriedem: thanks for the heads up for the patch | |
| 15:32:40 | openstackgerrit | Martin Midolesov proposed openstack/nova master: vmware:PropertyCollector for caching instance properties https://review.openstack.org/608278 | |
| 15:33:35 | mriedem | dansmith: belmorei_: fyi i'm working on a rocky patch with the workaround option | |
| 15:37:34 | dansmith | mriedem: cool | |
| 15:39:22 | belmorei_ | mriedem: thanks | |
| 15:39:25 | belmorei_ | mriedem: https://bugs.launchpad.net/nova/+bug/1796920 | |
| 15:39:27 | openstack | Launchpad bug 1796920 in OpenStack Compute (nova) "Baremetal nodes should not be exposing non-custom-resource-class (vcpu, ram, disk)" [Undecided,New] | |
| 15:58:26 | mriedem | dansmith: looks like zuulv3 status something or other changed and now openstack-gerrit-dashboard is getting NoneType errors - you see the same? | |
| 15:59:08 | dansmith | mriedem: I noticed it was failing this morning but didn't go to look if zuul was down. usually that's the reason | |
| 15:59:32 | mriedem | i'm guessing API change http://zuul.openstack.org/status | |
| 15:59:43 | mriedem | not sure, but the dashboard is different | |
| 15:59:57 | dansmith | ah yeah | |
| 16:03:06 | imacdonn | dansmith: could you take a peek at this, please? https://review.openstack.org/608091 | |
| 16:08:43 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/rocky: [stable-only] Add report_ironic_standard_resource_class_inventory option https://review.openstack.org/609043 | |
| 16:08:50 | mriedem | dansmith: jroll: dtantsur: ^ belmiro took off....would be nice if he can confirm that fixes his problem | |
| 16:09:34 | jroll | thanks | |
| 16:10:44 | imacdonn | that's one long option name :) | |
| 16:10:56 | mriedem | suggestions welcome | |
| 16:11:16 | mriedem | i figured do_the_dew wouldn't be helpful | |
| 16:11:18 | dansmith | imacdonn: done | |
| 16:11:29 | dansmith | imacdonn: mriedem should look at that too | |
| 16:11:33 | dansmith | or rather | |
| 16:11:39 | mriedem | i did once.. | |
| 16:11:40 | dansmith | mriedem should look at and agree with me on that too | |
| 16:13:51 | imacdonn | I do see your point | |
| 16:14:26 | imacdonn | not sure if anyone is actually doing the "keep hammering on it until it concedes" approach, but yeah | |
| 16:14:37 | edmondsw | and that notification having the message is also important for PowerVC, since it has means to present errors from notifications in the PowerVC GUI | |
| 16:14:53 | dansmith | imacdonn: I expect everyone is | |
| 16:15:08 | edmondsw | oops, ignore ^, somehow jumped channels | |
| 16:15:39 | imacdonn | my suspicion is that some people are running it once, and missing the fact that there are failures, and maybe others are not running it at all | |
| 16:16:44 | dansmith | people have to run this at various times or things won't work | |
| 16:17:09 | imacdonn | that may not be immediately obvious | |
| 16:17:41 | imacdonn | I've upgraded at least pike -> queens -> rocky without doing any online migrations, and nothing obviously didn't work | |
| 16:18:09 | dansmith | we have some db migrations which have blocked if you haven't run these to completion | |
| 16:18:14 | dansmith | maybe none since pike, but.. | |
| 16:18:17 | mriedem | http://git.openstack.org/cgit/openstack/openstack-ansible-os_nova/tree/tasks/nova_db_setup.yml#n98 | |
| 16:18:25 | mriedem | osa is certainly using it | |
| 16:18:51 | dansmith | I guess the default now is to run until completion, which is probably what people are doing I guess | |
| 16:19:09 | dansmith | but I know the return value here was critical earlier when people were running batches themselves | |
| 16:19:24 | mriedem | imacdonn: we also migrate some stuff online outside of the command | |
| 16:19:29 | mriedem | like on read from the db | |
| 16:19:32 | mriedem | or new resource create | |
| 16:20:16 | imacdonn | yeah, I know ... my point is that it's possible to get away without running the command, at least in some circumstances | |
| 16:20:41 | dansmith | imacdonn: I'm not sure what that has to do with anything | |
| 16:20:53 | dansmith | OSA and tripleo, and I expect other systems run this explicitly | |
| 16:20:55 | mriedem | if it's possible it's by chance | |
| 16:21:09 | dansmith | if you don't and it doesn't break in the versions you use, then you got lucky, | |
| 16:21:13 | mriedem | like dan said, we probably just haven't had a blocker migration in awhile | |
| 16:21:17 | dansmith | but that doesn't really mean anything for how important this is to notbreak | |
| 16:21:30 | mriedem | also depends on how old your data is, | |
| 16:21:50 | mriedem | i plan on dropping our request spec compat from newton in stein and if you don't have that migration done you'll fail to do things like migrate instances | |
| 16:21:53 | imacdonn | OK, nevermind .. I wasn't disgreeing that it's important to solve ... | |
| 16:22:17 | mriedem | so just make this return 2, doc and reno it and we're happy right? | |
| 16:23:02 | imacdonn | yeah, that'd work for this particular problem ... although it's probably not backportable ? | |
| 16:23:20 | imacdonn | (since it'll break things that don't know to check for 2) | |
| 16:23:58 | mriedem | i'm not sure; if things are failing but tooling is not aware of it, i think it's probably better to opt to the side of putting an upgrade release note and saying this will fail now | |
| 16:24:08 | mriedem | but i'd rather know something isn't working explicitly | |
| 16:24:19 | mriedem | dansmith: agree? ^ | |
| 16:24:31 | imacdonn | but if the automation is just repeating infinitely until it gets a 0, it'll .... repeat infinitely | |
| 16:25:33 | dansmith | what if 2 means "I didn't do anything but there were exceptions", 1 means "I did things, maybe there were some exceptions too", 0 means "I didn't do anything, but no errors" | |
| 16:25:48 | dansmith | repeat on 1, done on 0, 2 means we hit terminal fail state | |
| 16:26:24 | dansmith | either way people that loop on nonzero will break with anything you're going to do, which is why reno and doc is super important | |
| 16:26:39 | imacdonn | right | |
| 16:27:37 | dansmith | not sure how I feel about changing behavior in a backport with a retroactive reno, but mriedem is the authority here, so I'd do whatever he says | |
| 16:28:21 | imacdonn | I'm thinking that most people only read release notes for a new release, not for errata updates | |
| 16:28:43 | dansmith | unfortunately I don't think they even read them for new releases, but.. yeah | |
| 16:28:49 | imacdonn | heh yeah, there is that | |
| 16:30:25 | mriedem | i'll defer to tonyb | |
| 16:31:21 | dansmith | the AUD stops with tonyb | |
| 16:36:50 | melwitt | . | |
| 16:43:21 | mriedem | i get it | |
| 16:43:24 | mriedem | took me awhile | |
| 16:48:42 | imacdonn | dansmith: I'm on the fence about your last suggestion .... I think I tend towards exceptions being bad, requiring some problem to be addressed immediately ... but then there was a suggestion that some migrations may raise exceptions "by design" if some other migration had not yet been completed | |
| 16:49:26 | dansmith | imacdonn: well, by design or not, we've had some that won't complete until others do | |