Earlier  
Posted Nick Remark
#openstack-nova - 2018-04-30
14:32:31 stephenfin kashyap: Can do
14:33:37 kashyap Gracias. I've got 3 more in that same vein, will post them soon.
14:33:59 owalsh jaypipes: not sure how that could be, there is no upgrade...
14:35:27 owalsh jaypipes: also wouldn't that result in a KeyError?
14:36:49 jaypipes owalsh: sec, on standup
14:38:58 jaypipes 2018-04-25 05:40:18.617 21167 WARNING nova.scheduler.client.report [req-6b36b888-fa33-4301-aedf-3389020fe8d8 - - - - -] Discovering suitable URL for placement API failed.: DiscoveryFailure: Could not determine a suitable URL for the plugin
14:38:58 jaypipes owalsh: this is occurring:
14:39:14 jaypipes owalsh: that is the root of the issue, I believe.
14:39:28 jaypipes owalsh: something up with the service catalog discovery of placement maybe?
14:39:34 owalsh jaypipes: not up yet...
14:39:46 owalsh jaypipes: returning 404s later https://logs.rdoproject.org/openstack-periodic-24hr/periodic-tripleo-ci-centos-7-ovb-1ctlr_1comp-featureset002-pike-upload/7dadafb/overcloud-controller-0/var/log/nova/nova-placement-api.log.txt.gz#_2018-04-25_05_48_32_956
14:43:27 jaypipes owalsh: just realized something...
14:43:48 jaypipes owalsh: there is no self._resource_providers object any more. That has been replaced by self._provider_tree.
14:44:12 jaypipes owalsh: lemme look further into this. This is Pike, yeah?
14:44:19 owalsh jaypipes: yea, pike
14:44:25 owalsh jaypipes: thanks!
14:44:29 jaypipes owalsh: k, thx. gimme a few to track down.
14:45:39 jaypipes owalsh: no, that's not it... we switched to provider_tree in Queens. must be something else. :(
14:50:08 bauzas jaypipes: owalsh: sorry, I'm back
14:53:37 bauzas owalsh: jaypipes: I wonder if the root cause is https://logs.rdoproject.org/openstack-periodic-24hr/periodic-tripleo-ci-centos-7-ovb-1ctlr_1comp-featureset002-pike-upload/7dadafb/overcloud-novacompute-0/var/log/nova/nova-compute.log.txt.gz#_2018-04-25_05_40_03_801
14:53:51 bauzas if so, we're not registering the root RP
14:54:52 jaypipes bauzas: https://github.com/openstack/nova/blob/stable/pike/nova/scheduler/client/report.py#L514-L523
14:55:12 jaypipes bauzas: for some reason, we're either not creating or not getting the rp record for the compute node.
14:55:36 jaypipes bauzas: maybe a check for whether rp is None is needed before line 518.
14:55:43 jaypipes bauzas: and raise some exception.
14:55:52 openstackgerrit Aditya Vaja proposed openstack/nova master: remove IVS plug/unplug as they're moved to separate plugin https://review.openstack.org/534371
14:57:22 openstackgerrit Matthew Booth proposed openstack/nova master: libvirt: Fix misleading debug msg "Instance is running" https://review.openstack.org/565234
14:58:27 bauzas jaypipes: tbh, I think that if https://github.com/openstack/nova/blob/stable/pike/nova/scheduler/client/report.py#L516 is not working, it's an operator issue
14:58:33 bauzas owalsh: ^
14:59:21 owalsh bauzas: ack, yea... don't see any POST requests getting through to placement
15:00:30 bauzas owalsh: see also https://logs.rdoproject.org/openstack-periodic-24hr/periodic-tripleo-ci-centos-7-ovb-1ctlr_1comp-featureset002-pike-upload/7dadafb/overcloud-novacompute-0/var/log/nova/nova-compute.log.txt.gz#_2018-04-25_05_40_03_854
15:00:38 bauzas owalsh: looks like it's a keystone issue
15:00:58 jaypipes bauzas: @safe_connect is hiding the connection issue.
15:01:16 bauzas owalsh: jaypipes: so, IMHO, when we try to register the root RP (by creating it using the Placement API), we have a keystone problem
15:01:27 bauzas at least a connection problem
15:01:41 bauzas that's why the root RP is None
15:02:13 jaypipes bauzas: right. and @safe_connect's "only do something if X number of warnings happened" is hiding issues with connectivity. that's what I think at l;east.
15:02:26 bauzas yup
15:02:36 bauzas jaypipes: anyway, thanks for the help
15:02:44 bauzas owalsh: like I said, I think it's not a nova problem
15:02:45 jaypipes not sure I helped much :)
15:03:03 bauzas owalsh: rather a configuration issue because of the HTTP503
15:03:23 bauzas jaypipes: you did, dude ,)
15:03:44 bauzas jaypipes: I'm a bit said to no longer be a Placement expert
15:04:01 bauzas jaypipes: so having you telling me if I'm right is definitely helping me :)
15:05:27 owalsh jaypipes, bauzas: thanks guys, helps me a lot if we can rule out placement as the root cause :-)
15:05:45 bauzas owalsh: I think Placement is a canary
15:06:02 bauzas like we had NoValidHost for something else
15:06:20 owalsh bauzas: yea, was just about to say... NoValidHost is the canary
15:06:48 bauzas fortunately, we now have a separate exception
15:07:59 bauzas efried: jaypipes: stephenfin: oh btw. thanks for having reviewed my vGPU series. FWIW, https://twitter.com/sylvainbauza/status/990884997010685953 :)
15:10:38 melwitt mriedem: I noticed the novaclient change on "Add host/hostId to instance action events API" https://review.openstack.org/#/c/564667 has merged. is everything done for that bp now and time to remove from runway?
15:11:08 mriedem melwitt: yeah, i marked the bp complete on friday i think, but forgot to remove it from runways
15:11:35 melwitt mriedem: k, cool
15:19:04 stephenfin bauzas: :)
15:23:52 dansmith jaypipes: why are we not migrating the records for allocations that were created before user/project were required?
15:24:27 melwitt dansmith: those should be auto-healed right, by compute updates
15:24:38 dansmith melwitt: I don't think so
15:25:00 melwitt dansmith: that was the thinking as to why no migration was added when user/project were added
15:25:02 dansmith making them auto-heal is what I mean by migrating
15:25:24 melwitt any update to allocations should add user/project if not already existing
15:25:48 dansmith right, but does the reportclient re-write the allocation if just project/user is missing/
15:25:53 dansmith I thought it just counted resources
15:26:42 dansmith reportclient has become too complicated for me to be able to reasonably look I think
15:26:44 jaypipes dansmith: cdent doesn't think user and project should be NOT NULL...
15:26:58 dansmith jaypipes: because why?
15:27:04 melwitt I'm gonna look again
15:27:27 dansmith jaypipes: I was asking because we're making things nullable to add the consumer generation,
15:27:32 jaypipes dansmith: I'm not sure. he thinks that consumers in placement shouldn't need a project or user. so that placement can be used "for more things than just nova" was his answer.
15:27:42 dansmith jaypipes: but I guess unless we bump the minimum microversion we have to support those continuing to be created
15:27:55 jaypipes dansmith: yes, I would like to see those be NOT NULL, but there was stiff resistance from both cdent and edleafe
15:28:00 dansmith jaypipes: user/project is required in later microversions right?
15:28:39 dansmith if clients use older microversions just to get themselves an allocation without user/project, things are going to fall apart pretty quick
15:29:05 jaypipes dansmith: 1.8 added them. 1.12 made them required for allocations.
15:29:21 jaypipes dansmith: yes, I've argued this with both cdent and edleafe.
15:30:38 dansmith jaypipes: yeah, so not enforcing them in schema because people could be using older microversions is valid, but expecting people to use 1.5 going forward just because the want to create non-multitenant allocations is crazypants
15:33:25 jaypipes dansmith: I agree with you.
15:34:44 openstackgerrit Vladyslav Drok proposed openstack/nova master: ironic: Report resources as reserved when needed https://review.openstack.org/517921
15:35:31 fishbone_ hello all, I receive an error in the instance log when loading windows instances: pywintypes.com_error: (-2147352567, 'Exception occurred.', (0, 'Session', 'Access is denied. ', None, 0, -2147024891), None)
15:36:15 fishbone_ My first assumption is updating the cloudbase-init package on the images but would anyone know another possible cause?
15:37:38 melwitt dansmith: looks like any update of allocations will delete the already existing allocations, so because user/project is required >= 1.8, compute updates should result in ensuring user/project exist for allocations https://github.com/openstack/nova/blob/master/nova/api/openstack/placement/objects/resource_provider.py#L2065
15:38:01 dansmith melwitt: right, but computes don't just update allocations all the time
15:38:36 openstackgerrit Kashyap Chamarthy proposed openstack/nova master: libvirt: Remove support for Intel CMT `perf` events https://review.openstack.org/565242
15:38:54 melwitt okay, I had thought there was a periodic update, but that was temporary right? I think that was happening back when microversion 1.8 was added
15:39:59 dansmith melwitt: we don't heal active instances since ocata
15:41:12 melwitt dansmith: okay, I think 1.8 was added in pike. so it sounds like we are missing a migration of already existing instances
15:41:39 dansmith melwitt: well, it doesn't matter if we're not going to do the needful on the placement side
15:41:45 dansmith I mean, doesn't matter for my question above
15:41:55 dansmith matters for us using that data for quotas later, but not what I was asking about
15:42:07 dansmith hmm, I have op for some reason.. do we need a topic update before I drop it?
15:42:24 melwitt yeah, need to swap a runway
15:42:41 melwitt one of them has merged as of friday
15:43:11 dansmith it would help if you could make sure the actual blueprint tag is in the runway description somewhere
15:43:16 dansmith so I can just copy that in and not have to look it up
15:43:25 melwitt okay, can do
15:43:30 dansmith s/you/whoever is doing the runway jostling/
15:46:57 melwitt thanks
15:47:04 dansmith aye
15:50:07 melwitt yikun: fyi, your blueprint add-host-to-instance-action-events has been removed from the review runway as all the related code has merged. please feel free to add feedback about your experience with review runways at L186 https://etherpad.openstack.org/p/nova-runways-rocky
16:38:37 openstackgerrit Aditya Vaja proposed openstack/nova master: remove IVS plug/unplug as they're moved to separate plugin https://review.openstack.org/534371

Earlier   Later