Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-29
18:20:42 mriedem rust belt baby
18:21:00 melwitt ok, hm
18:23:15 mriedem this is likely related since the report client still relies on a provider tree for caching
18:24:05 melwitt yeah, like you mentioned in the bug. but how could the cache think the allocation ratio is already set to 16.0, for example? and think nothing changed
18:24:41 melwitt if it's currently 0.0
18:26:22 sean-k-mooney mriedem: well we dont call prov_tree.update_inventory(nodename, inv_data) in the except block you mentioned that updated the internal view
18:26:23 mriedem so my logging patch is most likely not going to help the xen case here because it's not going down that alternative route, i'll update in a bit
18:26:37 mriedem sean-k-mooney: no but the report client does internally
18:26:38 cdent mriedem: what are the odds it was related to this change: https://review.openstack.org/#/c/520024/
18:26:43 mriedem on "it's" version
18:26:56 mriedem cdent: heh jaypipes already brought that up :)
18:27:02 mriedem and it was the first thing i thought of today
18:27:11 mriedem but that isn't on rocky and zigo said he hit this on rocky
18:27:14 cdent oh, sorry, I was out walking and it crossed my mind
18:27:32 mriedem unless of course zigo is actually hitting stein code somehow
18:27:40 cdent I'm not certain that what zigo is experiencing is the same as what the xen stuff is experiencing
18:27:54 melwitt do we see this log message, LOG.warning('Unable to refresh my resource provider record')? says "# NOTE(danms): Either we failed to fetch/create the RP on our first attempt, or a previous attempt had to invalidate the cache, and we were unable to refresh it. Bail and try again next time."
18:27:56 cdent because it's not clear what setup zigo has or hasn't done
18:28:17 jaypipes yeah, cdent, that was SOOOOO one hour ago, geez! :P
18:28:29 cdent heh, I was fetching elderberries
18:28:30 cdent which is like
18:28:32 cdent so british
18:28:34 jaypipes :)
18:29:13 cdent it was the first thing I came across when trying to figure out why the ComputeNode after get_inventory might be wrong
18:30:00 cdent because tracing the code the ComputeNode at that point has to have a cpu_allocation_ratio of 0 for us to see the results that are happening
18:30:29 cdent (was also prepared to fetch blackberries but they aren't ready)
18:31:54 mriedem are dingleberries in season yet?
18:32:02 jaypipes cdent: yeah, after finding out that zigo was on rocky, I went and reviewed the last 3 months of patches to rocky that could have anything at all do with allocation ratio or inventory setting and couldn't find anything at all. Also note that the patch you mentioned above isn't in Rocky
18:32:09 jaypipes mriedem: always.
18:32:13 mriedem ha
18:33:00 mriedem melwitt: http://logs.openstack.org/41/590041/17/check/tempest-full/b3f9ddd/controller/logs/screen-n-cpu.txt.gz#_Aug_27_14_18_24_078058 is a failed xen run if you want to dig for logs
18:33:11 melwitt thx
18:36:24 openstackgerrit Matt Riedemann proposed openstack/nova master: Add contributor guide for upgrade status checks https://review.openstack.org/596902
18:41:18 melwitt not seeing this message in the log LOG.debug('Updated inventory for %s at generation %i', which should be there if we've ever successfully updated inventory. which supports the theory that self._provider_tree.has_inventory_changed is returning False
18:41:46 melwitt I don't see any of the error log messages associated with a failure to update inventory
18:43:09 melwitt looks like it would be helpful to have a debug message "Inventory has not changed, skipping update" when it skips
18:43:10 sean-k-mooney its kind of dump but im going to hard code a not implmented excpetion here https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L895 and restack and see if i get the same behavior
18:44:46 mriedem melwitt: yeah that's what my debug patch is doing
18:44:51 sean-k-mooney melwitt: i think thi is where we check if things have changed https://github.com/openstack/nova/blob/722d5b477219f0a2435a9f4ad4d54c61b83219f1/nova/scheduler/client/report.py#L865
18:45:18 mriedem but now i'm distracted by fracas in the tc channel
18:45:24 melwitt mriedem: yeah, just saw that and was about to say, that's what you're already doing and are going to update it to also do it for the non-provider tree route
18:45:45 melwitt since that's where we're going for xen anyway
18:45:51 mriedem yes in progress
18:45:56 mriedem despite fracas
18:46:00 melwitt sean-k-mooney: yes, that's it
18:51:46 openstackgerrit Matt Riedemann proposed openstack/nova master: Add debug logs for when provider inventory changes https://review.openstack.org/597560
18:51:48 mriedem updated for alt path ^ note hyperv and vmware etc would all be failing from this as well if that alternate path is the issue
18:54:02 sean-k-mooney mriedem: i wonder if the hyperv ci also hardcordes the allocation ratios in the conf
18:56:31 mriedem we can find out
18:58:14 sean-k-mooney mriedem: the hyperv ones look ok
18:58:25 sean-k-mooney as in not set
18:58:29 melwitt oh, well, if xen is hard-coding values that match the 0.0 defaults, then placement would _not_ be updated right?
18:59:03 melwitt oh, nevermind. reportclient should be comparing with what's in placement
18:59:03 sean-k-mooney mlavalle: xen were hardcoding real ratios
18:59:40 sean-k-mooney * melwitt: ^
18:59:49 sean-k-mooney melwitt: http://dd6b71949550285df7dc-dda4e480e005aaa13ec303551d2d8155.r49.cf1.rackcdn.com/60/597560/1/check/dsvm-tempest-neutron-network/0b2f0a8/logs/local.conf.txt.gz
19:00:10 sean-k-mooney [DEFAULT]
19:00:11 sean-k-mooney disk_allocation_ratio = 2.0
19:00:13 sean-k-mooney ram_allocation_ratio = 1.5
19:00:15 sean-k-mooney cpu_allocation_ratio = 16.0
19:00:38 melwitt yeah. I was trying to think if that would appear to reportclient as "no change" and therefore not update placement
19:00:41 sean-k-mooney all vaild that said i would never advise setting the disk_allocation_ration over 1
19:01:20 melwitt but, reportclient should be comparing what placement has with those new values, so it should see a change. but from what we know so far, it looks like it isn't seeing a change. mriedem's debug logs will confirm
19:01:30 sean-k-mooney melwitt: if you could some how get the value to the report clinet as 0.0 then yes
19:02:34 melwitt right, yeah
19:03:31 mriedem sean-k-mooney: they are *now*
19:03:35 mriedem they weren't when they reported the issue
19:03:42 mriedem they are hard-coding them in config as a workaround for the CI failure
19:03:54 mriedem which is why i've reverted that change to try and actually get a recreate with logging
19:04:03 sean-k-mooney right so before they were not set
19:04:24 cdent mriedem: have you added logs which watch the value of the cn.*_allocation_ratio in some way?
19:05:02 cdent i'm looking at https://review.openstack.org/#/c/597560/2/nova/compute/resource_tracker.py,unified and wonder if we want more info about the state of the compute node along the way
19:06:04 cdent when cn.save() is called if those values are weird for some reason the ratio adjustment stuff in _from_db_object _might_ no be behaving as expected
19:06:15 mriedem i could add that
19:06:17 cdent (of course you have probably already analyzed this while I was getting elderberries)
19:06:25 melwitt sean-k-mooney: yeah, so focusing on the values somehow being 0.0 _after_ the normalize from compute node object, that's what got mentioned earlier, how that could possibly happen
19:07:11 melwitt the normalize function is supposed to be filling in with the defaults 16.0 etc
19:08:36 sean-k-mooney ... so my raise NotImplemented to force the alt path with the libvirt driver sill resulted in the correct vaules
19:09:17 melwitt ? so how is xen failing? I thought it was taking the alt path
19:09:47 sean-k-mooney melwitt: it is but taking the alt path is apparently not enough
19:09:57 melwitt oh
19:11:40 sean-k-mooney i have matt's debug patch applied also but im not sure that is going to show where the default values got applied
19:12:58 openstackgerrit Matt Riedemann proposed openstack/nova master: Add debug logs for when provider inventory changes https://review.openstack.org/597560
19:13:05 mriedem cdent: like this? ^
19:18:52 cdent mriedem: yeah, nice. that combined with the other stuff ought to help see the flow
19:19:18 cdent s/see/better see/
19:19:37 cdent The difficulty with creating an MTC for this makes me anxious
19:26:31 sean-k-mooney im restacking in offline mode (with libvirt) we are expecting to see the defaulting to ... message if the compute node object is setting the defaults right
19:28:34 sean-k-mooney i can deploy a xen node tommorow if needed to see if i can reporduce
19:41:30 openstackgerrit Matt Riedemann proposed openstack/nova master: Revert "libvirt: add method to configure migration speed" https://review.openstack.org/590814
19:43:53 cfriesen jaypipes: re the cold migration with PCI devices. were you talking about the difference between it being theoretically supported and actually doing it? StarlingX integration tests do cold migration with PCI/SRIOV regularly, but I realize that doesn't answer the question for upstream.
19:45:34 sean-k-mooney cfriesen: i think i have done it in the past also i had tought it was ment to be supported. that said not sure it updated teh resouce tracker correctly
19:45:49 jaypipes cfriesen: yes, I'm referring to real-world deployments who do migrations where the instances hold on to their IP addresses, GPUs, and everything else and are migrated to a different rack/region whatever
19:46:19 jaypipes cfriesen: but whatever, I'm running from that conversation screaming.
19:46:25 cfriesen jaypipes: :)
19:55:35 openstackgerrit Matt Riedemann proposed openstack/nova master: Add functional test for live migrate with anti-affinity group https://review.openstack.org/588935
19:56:22 mriedem cfriesen: upstream supports cold migration with pci devices
19:56:31 mriedem remember moshe and ludovic got that working
19:57:03 mriedem there was also 3rd party ci from mellanox for it at one time i think
19:57:06 mriedem but that might be dead now

Earlier   Later