Earlier  
Posted Nick Remark
#openstack-nova - 2018-01-30
20:59:21 mriedem melwitt: yes on the last question, that needs to go on top of https://review.openstack.org/#/c/536936/
20:59:47 mriedem as for the former question, i tried that in https://review.openstack.org/#/c/536798/ but my patch to not run cellsv1 in stable landed first so the job didn't run on that patch
20:59:51 melwitt ah, okay. thanks
20:59:56 mriedem we'd have to do some depends-on chicanery there
21:00:03 melwitt gotcha
21:00:43 mriedem cfriesen: so on the perf thing, you just found out that compute was using more cpu than before?
21:01:02 mriedem that was likely because in newton the computes started posting inventory information to placement from every RT update
21:01:19 mriedem but you said server creates were taking a lot longer
21:03:37 openstackgerrit melanie witt proposed openstack/nova stable/ocata: Bumping functional test job timeouts https://review.openstack.org/539320
21:05:01 cfriesen mriedem: yeah, on same hardware instance creation on newton was 32 sec and pike was 230. but it's hard to break down exactly what's causing it since anything audit-driven will also show higher usage.
21:06:41 cfriesen mriedem: it should be noted this is on an all-in-one system, so the openstack services are constrained to only two CPUs, and those were pretty much pinned
21:06:50 cfriesen ie at 100% usage
21:09:14 mriedem cfriesen: can't run osprofiler + rally or something to see at least where the majority of the time is being spent?
21:12:28 cfriesen mriedem: we've got traces from intel's vtune analyzer showing how much time is spent where, but just about everything shows increases. there's no single smoking gun.
21:14:53 openstackgerrit Ed Leafe proposed openstack/nova master: Make the InstanceMapping marker UUID-like https://review.openstack.org/539323
21:16:04 mriedem cfriesen: huh
21:18:22 openstackgerrit Eric Fried proposed openstack/nova master: Make generation optional in ProviderTree https://review.openstack.org/539324
21:20:48 mriedem melwitt: looks like you have real test failures in https://review.openstack.org/#/c/539013/
21:21:42 melwitt mriedem: ah, thank you. I hadn't looked at the detail yet. I shall fix that up
21:21:47 sean-k-mooney cfriesen: dumb question but could it be related to meltdow/specter patches?
21:23:42 cfriesen sean-k-mooney: not so dumb. :) but no, this was a load from before the meltdown/spectre patches were applied.
21:25:40 mriedem cfriesen: and you've got the latest stable/pike release?
21:26:28 mriedem like i was wondering if any of the RequestContext changes might be related https://github.com/openstack/nova/commits/stable/pike/nova/context.py
21:26:34 mriedem since the context is used everywhere
21:27:51 cfriesen mriedem: not the latest, no. originally from 16.0.2, with some stable/pike stuff since then. not sure exactly what, I've been busy with other stuff.
21:28:46 mriedem cfriesen: ok would be good to know if it's still the same issue after you've got the latest pike fixes applied
21:28:57 mriedem https://docs.openstack.org/releasenotes/nova/pike.html
21:30:15 mriedem full change log http://paste.openstack.org/show/658066/
21:30:46 cfriesen don't see anything related to "context" in there, got any pointers?
21:31:52 cfriesen last stable/pike changes to nova/context.py were from October
21:32:27 mriedem nothing in that changelog jumps out at me as a perf related fix
21:34:35 cfriesen does the upstream CI environment do performance tests of common operations?
21:35:56 openstackgerrit Eric Fried proposed openstack/nova master: Remove compute nodes arg from ProviderTree init https://review.openstack.org/539330
21:38:31 mriedem cfriesen: no
21:38:44 mriedem too much variance from node to node
21:39:04 mriedem not sure what can be pulled perf-trend wise from openstack-health
21:39:32 mriedem http://status.openstack.org/openstack-health/#/
21:40:30 mriedem like, i don't know how to take that and see how long a simple tempest create server test has taken over the last 12 monhts
21:40:32 mriedem *months
21:40:36 mriedem mtreinish can maybe help
21:40:57 openstackgerrit melanie witt proposed openstack/nova stable/ocata: Stop globally caching host states in scheduler HostManager https://review.openstack.org/539013
21:49:05 mriedem mmedvede: powerkvm ci seems pretty unhappy
21:49:07 mriedem is that a known issue?
21:51:30 openstackgerrit Sylvain Bauza proposed openstack/nova master: Provide support matrix and doc for VGPU https://review.openstack.org/539266
21:56:07 mriedem dansmith: want to hit these backports? I didn't realize those weren't merged by now https://review.openstack.org/#/q/Ie70c77db753711e1449e99534d3b83669871943f+status:open
21:56:48 mmedvede mriedem: I do not see anything too far out of ordinary, which powerkvm ci unhappiness are you referring to?
21:57:00 mmedvede double checking now
21:57:12 mriedem mmedvede: https://review.openstack.org/#/c/538510/
21:57:18 mriedem https://dal05.objectstorage.softlayer.net/v1/AUTH_3d8e6ecb-f597-448c-8ec2-164e9f710dd6/pkvmci/nova/10/538510/3/check/tempest-dsvm-full-xenial/41f1c6d/
22:01:15 edleafe efried: PUT {} is not semantically the same as DELETE, even if in most cases the result is the same
22:04:15 mriedem cfriesen: speaking of perf, this is an easy fix for an RT perf issue if you're building several instances on the same compute host at once https://review.openstack.org/#/q/Ib588c31a4d2075f8730409d50c99dfb04180a9cd+status:open
22:04:32 mriedem our operations people were hitting perf issues with the big RT update lock
22:09:08 dansmith mriedem: got em
22:10:05 efried edleafe: Oh? How not?
22:10:26 efried edleafe: Oh, you mean in the general case, where None and {} aren't the same thing.
22:10:50 prometheanfire win 30
22:11:51 edleafe efried: No. Sometimes you need to indicate if anything was in fact deleted. PUT {} can't do that; DELETE can
22:12:12 efried edleafe: It can? How?
22:12:43 openstackgerrit Matt Riedemann proposed openstack/nova stable/ocata: Pass the correct image to build_request_spec in conductor.rebuild_instance https://review.openstack.org/516404
22:14:12 mmedvede mriedem: that pkvmci failure seems to have been a fluke on that VM, one of a kind. Sorry for the false negative.
22:14:56 edleafe efried: In cases where that distinction is important, you can return a 404 if the thing you're trying to delete is not there
22:15:17 sean-k-mooney the ibm powerkvm ci is still broken currently correct?
22:15:25 edleafe efried: Like I said, it isn't usually necessary. When you want to delete something, you usually just want it gone.
22:15:58 efried edleafe: Okay, then I'm specifically talking about resource provider inventories, traits, allocations, and aggregates.
22:16:42 efried edleafe: I can provide a good argument for why we should favor PUT <empty> over DELETE (at least DELETE as currently implemented). I'm trying to figure out if there's an argument for the other side.
22:16:49 edleafe efried: for those cases, I simply prefer the grammar of DELETE
22:17:07 mmedvede sean-k-mooney: I am confused as to why you think so. It did have a few failures, but failure rate is within normal
22:17:26 efried edleafe: The grammar of the request (as opposed to the (lack of) response), right?
22:17:28 edleafe efried: That's what DELETE is designed to do.
22:17:52 edleafe efried: No, PUT {} is an awkward way of saying DELETE
22:18:49 efried edleafe: Yeah, I get it. Is it "wrong" (in the annals of HTTP, or REST, or APIs, or whatever) for a DELETE API to return a payload?
22:20:30 sean-k-mooney efried: yes i belive it is not ment to have a payload generally
22:20:46 edleafe efried: Unless the response is 204.
22:21:01 mmedvede sean-k-mooney: am I missing something? Both http://ci-watch.tintri.com/project?project=nova and https://dal05.objectstorage.softlayer.net/v1/AUTH_3d8e6ecb-f597-448c-8ec2-164e9f710dd6/pkvmci/index.html do not indicate a systemic failure on nova patches
22:21:09 mmedvede for powerkvm ci
22:21:40 sean-k-mooney mmedvede: there was a message a week or two saying it was i was not sure if it was fixed or not
22:21:56 edleafe efried: Othewise, you can return either a description of the deleted resource (200) or a URL to check for success/failure on a 202.
22:22:28 efried edleafe: So it'd be acceptable for a DELETE to return 200 with a payload?
22:23:01 edleafe efried: it's required
22:23:07 mriedem sean-k-mooney: are you thinking of the zvm ci?
22:23:23 mriedem or zkvm i mean
22:23:58 efried edleafe: Sorry, I mean I get that it's cool for DELETE to respond 204 with no content; I'm asking whether there's any restriction - standard-wise or cultural - against a DELETE responding 200 with a payload.
22:24:07 mriedem mmedvede: i had seen some other pkvm ci failures in stable branches, but those might have been old/transient
22:24:11 efried and I think you've said that's acceptable.
22:24:46 edleafe efried: yes, it's acceptable, although it isn't very common
22:26:14 sean-k-mooney mriedem: yes i was https://www.mail-archive.com/openstack-dev@lists.openstack.org/msg115082.html
22:27:29 mriedem sean-k-mooney: mixing up one of the dozen ibm 3rd party CIs is grounds for pistols at dawn
22:28:02 mmedvede mriedem: yes, stable branches have high rate of failure unfortunately, I'll shift some time to look at those.
22:28:53 sean-k-mooney mriedem: haha well the grenade job is failing because of a ubuntu keyring missing on the powervm ci too but that could be intermitent
22:29:33 mmedvede sean-k-mooney: that is intermittent, there is a bug in ubuntu somewhere that we reported
22:29:59 mmedvede it happens only last 10 minutes of any hour
22:30:21 sean-k-mooney mmedvede: ya the patch i noticed it on is for rocky anyway so im not going to waste ci time rechecking
22:32:47 sean-k-mooney anyway i have fixed my unrelated ovs db socket somehow became a directory and broke everything issue with kolla so im going to head home for the evening
22:32:58 efried edleafe: btw, in case it wasn't obvious, this is pursuant to what we were discussing the other day. Without a response payload, we have to assume things about the effect of DELETE on provider generation. One possible solution is to use PUT <empty> where available, which it happens to be for all of these. Another is to implement DELETE with a response payload from which we can glean the new generation.
22:36:16 edleafe efried: Aren't we sending the generation along with the PUT/DELETE request?
22:37:00 efried edleafe: With PUT, yes. Not with DELETE, which doesn't accept a payload. The latter is a definite (but separate) problem.
22:37:37 efried edleafe: But even the former only guarantees that we're deleting what we thought we were deleting. The lack of generation in the return is a problem for *subsequent* updates.
22:38:11 efried ...unless we continue to make assumptions about how placement does generations. Which IMO is wrong.
22:38:40 edleafe efried: So say I get the generation back from the PUT/DELETE. Right after that, other requests modify the resource. What good does getting back gen+1 from the request do me then?
22:39:54 efried edleafe: In that scenario, it doesn't save you anything, because your next update will 409 and you have to re-GET the provider and its associated stuff before you redrive your update.

Earlier   Later