| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-08-29 | |||
| 18:27:02 | mriedem | and it was the first thing i thought of today | |
| 18:27:11 | mriedem | but that isn't on rocky and zigo said he hit this on rocky | |
| 18:27:14 | cdent | oh, sorry, I was out walking and it crossed my mind | |
| 18:27:32 | mriedem | unless of course zigo is actually hitting stein code somehow | |
| 18:27:40 | cdent | I'm not certain that what zigo is experiencing is the same as what the xen stuff is experiencing | |
| 18:27:54 | melwitt | do we see this log message, LOG.warning('Unable to refresh my resource provider record')? says "# NOTE(danms): Either we failed to fetch/create the RP on our first attempt, or a previous attempt had to invalidate the cache, and we were unable to refresh it. Bail and try again next time." | |
| 18:27:56 | cdent | because it's not clear what setup zigo has or hasn't done | |
| 18:28:17 | jaypipes | yeah, cdent, that was SOOOOO one hour ago, geez! :P | |
| 18:28:29 | cdent | heh, I was fetching elderberries | |
| 18:28:30 | cdent | which is like | |
| 18:28:32 | cdent | so british | |
| 18:28:34 | jaypipes | :) | |
| 18:29:13 | cdent | it was the first thing I came across when trying to figure out why the ComputeNode after get_inventory might be wrong | |
| 18:30:00 | cdent | because tracing the code the ComputeNode at that point has to have a cpu_allocation_ratio of 0 for us to see the results that are happening | |
| 18:30:29 | cdent | (was also prepared to fetch blackberries but they aren't ready) | |
| 18:31:54 | mriedem | are dingleberries in season yet? | |
| 18:32:02 | jaypipes | cdent: yeah, after finding out that zigo was on rocky, I went and reviewed the last 3 months of patches to rocky that could have anything at all do with allocation ratio or inventory setting and couldn't find anything at all. Also note that the patch you mentioned above isn't in Rocky | |
| 18:32:09 | jaypipes | mriedem: always. | |
| 18:32:13 | mriedem | ha | |
| 18:33:00 | mriedem | melwitt: http://logs.openstack.org/41/590041/17/check/tempest-full/b3f9ddd/controller/logs/screen-n-cpu.txt.gz#_Aug_27_14_18_24_078058 is a failed xen run if you want to dig for logs | |
| 18:33:11 | melwitt | thx | |
| 18:36:24 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add contributor guide for upgrade status checks https://review.openstack.org/596902 | |
| 18:41:18 | melwitt | not seeing this message in the log LOG.debug('Updated inventory for %s at generation %i', which should be there if we've ever successfully updated inventory. which supports the theory that self._provider_tree.has_inventory_changed is returning False | |
| 18:41:46 | melwitt | I don't see any of the error log messages associated with a failure to update inventory | |
| 18:43:09 | melwitt | looks like it would be helpful to have a debug message "Inventory has not changed, skipping update" when it skips | |
| 18:43:10 | sean-k-mooney | its kind of dump but im going to hard code a not implmented excpetion here https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L895 and restack and see if i get the same behavior | |
| 18:44:46 | mriedem | melwitt: yeah that's what my debug patch is doing | |
| 18:44:51 | sean-k-mooney | melwitt: i think thi is where we check if things have changed https://github.com/openstack/nova/blob/722d5b477219f0a2435a9f4ad4d54c61b83219f1/nova/scheduler/client/report.py#L865 | |
| 18:45:18 | mriedem | but now i'm distracted by fracas in the tc channel | |
| 18:45:24 | melwitt | mriedem: yeah, just saw that and was about to say, that's what you're already doing and are going to update it to also do it for the non-provider tree route | |
| 18:45:45 | melwitt | since that's where we're going for xen anyway | |
| 18:45:51 | mriedem | yes in progress | |
| 18:45:56 | mriedem | despite fracas | |
| 18:46:00 | melwitt | sean-k-mooney: yes, that's it | |
| 18:51:46 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add debug logs for when provider inventory changes https://review.openstack.org/597560 | |
| 18:51:48 | mriedem | updated for alt path ^ note hyperv and vmware etc would all be failing from this as well if that alternate path is the issue | |
| 18:54:02 | sean-k-mooney | mriedem: i wonder if the hyperv ci also hardcordes the allocation ratios in the conf | |
| 18:56:31 | mriedem | we can find out | |
| 18:58:14 | sean-k-mooney | mriedem: the hyperv ones look ok | |
| 18:58:25 | sean-k-mooney | as in not set | |
| 18:58:29 | melwitt | oh, well, if xen is hard-coding values that match the 0.0 defaults, then placement would _not_ be updated right? | |
| 18:59:03 | melwitt | oh, nevermind. reportclient should be comparing with what's in placement | |
| 18:59:03 | sean-k-mooney | mlavalle: xen were hardcoding real ratios | |
| 18:59:40 | sean-k-mooney | * melwitt: ^ | |
| 18:59:49 | sean-k-mooney | melwitt: http://dd6b71949550285df7dc-dda4e480e005aaa13ec303551d2d8155.r49.cf1.rackcdn.com/60/597560/1/check/dsvm-tempest-neutron-network/0b2f0a8/logs/local.conf.txt.gz | |
| 19:00:10 | sean-k-mooney | [DEFAULT] | |
| 19:00:11 | sean-k-mooney | disk_allocation_ratio = 2.0 | |
| 19:00:13 | sean-k-mooney | ram_allocation_ratio = 1.5 | |
| 19:00:15 | sean-k-mooney | cpu_allocation_ratio = 16.0 | |
| 19:00:38 | melwitt | yeah. I was trying to think if that would appear to reportclient as "no change" and therefore not update placement | |
| 19:00:41 | sean-k-mooney | all vaild that said i would never advise setting the disk_allocation_ration over 1 | |
| 19:01:20 | melwitt | but, reportclient should be comparing what placement has with those new values, so it should see a change. but from what we know so far, it looks like it isn't seeing a change. mriedem's debug logs will confirm | |
| 19:01:30 | sean-k-mooney | melwitt: if you could some how get the value to the report clinet as 0.0 then yes | |
| 19:02:34 | melwitt | right, yeah | |
| 19:03:31 | mriedem | sean-k-mooney: they are *now* | |
| 19:03:35 | mriedem | they weren't when they reported the issue | |
| 19:03:42 | mriedem | they are hard-coding them in config as a workaround for the CI failure | |
| 19:03:54 | mriedem | which is why i've reverted that change to try and actually get a recreate with logging | |
| 19:04:03 | sean-k-mooney | right so before they were not set | |
| 19:04:24 | cdent | mriedem: have you added logs which watch the value of the cn.*_allocation_ratio in some way? | |
| 19:05:02 | cdent | i'm looking at https://review.openstack.org/#/c/597560/2/nova/compute/resource_tracker.py,unified and wonder if we want more info about the state of the compute node along the way | |
| 19:06:04 | cdent | when cn.save() is called if those values are weird for some reason the ratio adjustment stuff in _from_db_object _might_ no be behaving as expected | |
| 19:06:15 | mriedem | i could add that | |
| 19:06:17 | cdent | (of course you have probably already analyzed this while I was getting elderberries) | |
| 19:06:25 | melwitt | sean-k-mooney: yeah, so focusing on the values somehow being 0.0 _after_ the normalize from compute node object, that's what got mentioned earlier, how that could possibly happen | |
| 19:07:11 | melwitt | the normalize function is supposed to be filling in with the defaults 16.0 etc | |
| 19:08:36 | sean-k-mooney | ... so my raise NotImplemented to force the alt path with the libvirt driver sill resulted in the correct vaules | |
| 19:09:17 | melwitt | ? so how is xen failing? I thought it was taking the alt path | |
| 19:09:47 | sean-k-mooney | melwitt: it is but taking the alt path is apparently not enough | |
| 19:09:57 | melwitt | oh | |
| 19:11:40 | sean-k-mooney | i have matt's debug patch applied also but im not sure that is going to show where the default values got applied | |
| 19:12:58 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add debug logs for when provider inventory changes https://review.openstack.org/597560 | |
| 19:13:05 | mriedem | cdent: like this? ^ | |
| 19:18:52 | cdent | mriedem: yeah, nice. that combined with the other stuff ought to help see the flow | |
| 19:19:18 | cdent | s/see/better see/ | |
| 19:19:37 | cdent | The difficulty with creating an MTC for this makes me anxious | |
| 19:26:31 | sean-k-mooney | im restacking in offline mode (with libvirt) we are expecting to see the defaulting to ... message if the compute node object is setting the defaults right | |
| 19:28:34 | sean-k-mooney | i can deploy a xen node tommorow if needed to see if i can reporduce | |
| 19:41:30 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Revert "libvirt: add method to configure migration speed" https://review.openstack.org/590814 | |
| 19:43:53 | cfriesen | jaypipes: re the cold migration with PCI devices. were you talking about the difference between it being theoretically supported and actually doing it? StarlingX integration tests do cold migration with PCI/SRIOV regularly, but I realize that doesn't answer the question for upstream. | |
| 19:45:34 | sean-k-mooney | cfriesen: i think i have done it in the past also i had tought it was ment to be supported. that said not sure it updated teh resouce tracker correctly | |
| 19:45:49 | jaypipes | cfriesen: yes, I'm referring to real-world deployments who do migrations where the instances hold on to their IP addresses, GPUs, and everything else and are migrated to a different rack/region whatever | |
| 19:46:19 | jaypipes | cfriesen: but whatever, I'm running from that conversation screaming. | |
| 19:46:25 | cfriesen | jaypipes: :) | |
| 19:55:35 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add functional test for live migrate with anti-affinity group https://review.openstack.org/588935 | |
| 19:56:22 | mriedem | cfriesen: upstream supports cold migration with pci devices | |
| 19:56:31 | mriedem | remember moshe and ludovic got that working | |
| 19:57:03 | mriedem | there was also 3rd party ci from mellanox for it at one time i think | |
| 19:57:06 | mriedem | but that might be dead now | |
| 20:16:11 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Delete instance_group_member records from API DB during archive https://review.openstack.org/588943 | |
| 20:16:12 | mriedem | melwitt: just noticed you had commented on this ^ test should cover the case you noted now | |
| 20:17:08 | melwitt | ok, will look | |
| 20:20:18 | mriedem | crap forgot to update the bug reference in the commit message | |
| 20:21:05 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Delete instance_group_member records from API DB during archive https://review.openstack.org/588943 | |
| 20:28:26 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove old check_attach version check in API https://review.openstack.org/588348 | |
| 20:39:26 | dansmith | melwitt: mriedem: this is going to pass tests in a few minutes: https://review.openstack.org/#/c/597206 | |
| 20:39:49 | dansmith | if you could ack it with a +1 (or tell me what you want changed), I will go about trying to figure out how I'm going to get that merged :) | |
| 20:40:06 | melwitt | will do | |
| 20:43:05 | dansmith | also i was going to verify resource providers before/after and then realized we can't really do that since other projects might create providers, and we have no "service type" field | |
| 20:43:37 | mriedem | one hack way to determine a compute node provider is via the VCPU inventory | |