Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-26
23:31:14 dansmith an allocation will increment the generation of the provider on the placement side
23:31:23 dansmith and a transaction will abort if we try to do it at the same time as something else
23:31:26 melwitt ah, right. that's what I was thinking of
23:31:52 dansmith placement should really retry those for us server-side I would think, for an allocation type request
23:31:58 melwitt I couldn't remember where the RP generation was related there
23:32:22 dansmith doesn't matter, but I would think it would be better, especially given the advice for that error is "try exactly the same thing again"
23:32:24 melwitt yeah, I was thinking the same
23:32:29 mriedem this was the error from the server
23:32:30 mriedem There was a conflict when trying to complete your request.\n\n Inventory changed while attempting to allocate: Another thread concurrently updated the data. Please retry your update
23:32:38 dansmith right
23:32:41 mriedem i'm not sure why inventory would change
23:32:45 mriedem that's static in this case
23:33:01 dansmith any allocation will change the generation
23:33:04 mriedem the update_available_resource periodic will post inventory, but only if it changes
23:33:19 dansmith so any two allocations can conflict
23:33:21 melwitt I'm +1 on the idea of retrying server-side
23:34:16 melwitt I'm a little worried how many conflicts can we get in real life by ppl trying to create 500 servers and the scheduler is doing a "pack" pattern
23:35:10 melwitt as far as how to tune how many retries
23:35:23 melwitt to allow
23:35:33 dansmith this is also a highly synthetic scenario with a "virt driver" that doesn't get looked at much.. it could be doing something to re-stab inventory for no reason or something
23:36:00 mriedem for each instance, we go through the filters
23:36:17 mriedem and doesn't the scheduler have some kind of tracking on the HostState objects themselves for chosen hosts?
23:36:22 edleafe mriedem: that error message is poorly worded. It should be something like "available inventory has changed"
23:36:25 melwitt oh, the "inventory changed" yeah, I don't know anything about that. so that means it wasn't an allocation writing conflict?
23:36:41 dansmith melwitt: it's al related
23:36:45 dansmith *all
23:37:07 dansmith edleafe: that's not really accurate, AFACT, since allocations will cause the generation to increase
23:37:16 dansmith edleafe: and thus it could be nothing changed with inventory to cause that
23:37:34 edleafe dansmith: uh, that's why I added "available".
23:37:46 edleafe dansmith: IOW, some inventory has been allocated
23:37:47 mriedem this is basically the inventory reporting for the fake driver https://github.com/openstack/nova/blob/master/nova/virt/fake.py#L111
23:37:51 edleafe and changed the generation
23:38:01 dansmith edleafe: okay, I wouldn't word it that way for clarity, but okay :)
23:38:47 edleafe I wouldn't word it that way either, but I was guessing the author's intent
23:39:07 edleafe *cough* cdent *cough*
23:39:08 dansmith mriedem: that's the info that we use to generate it, yes
23:39:33 melwitt yeah, if it can happen without inventory (total possible capacity) changing, then that error message is confusing to me
23:40:05 edleafe melwitt: yeah, it's worded very poorly
23:40:10 dansmith mriedem: I'd look to see if the compute is hitting placement /inventory ever after the first go, and maybe check the nothing-changed short-circuit to make sure we're never going through it
23:40:32 mriedem the inventory nothing changed?
23:40:37 melwitt I dunno, I think I know just enough for it to be confusing. for an end user, it might not be confusing
23:40:47 dansmith mriedem: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L571-L572
23:40:56 mriedem melwitt: end user won't see it, they'll see NoValidHost on 500 instances
23:41:01 mriedem the operator will see it
23:41:04 melwitt good point
23:41:47 dansmith mriedem: or just look at placement logs to see if inventory is hit any time after compute startup
23:46:59 mriedem gdi, how do i regex search with grep
23:47:06 mriedem sudo journalctl -a -u devstack@placement-api.service | grep '.*PUT.*\/inventories.*'
23:48:12 dansmith PUT.*invent should be all you need
23:49:29 mriedem doesn't work
23:51:13 mriedem ah, well,
23:51:16 mriedem PUT.*alloc works
23:51:21 mriedem so it probably just wrapped
23:51:28 mriedem and it's not updating inventory, as it shouldn't
23:52:16 mriedem got my 1000 instances now, so will do the test stuff once i'm done with dinner
23:53:00 dansmith so I'd also check to make sure compute isn't doing the ocata healing during boot or something like that
23:57:18 takashin Spec cores, could you review https://review.openstack.org/#/c/489029/ ? It got one +2.
#openstack-nova - 2017-09-27
00:13:05 mriedem dansmith: interesting, listing without details, 500 error (cell0) and 500 active (cell1) is a lot faster than the 1000 active,
00:13:13 mriedem i suppose because we don't have as much to join
00:13:27 dansmith mriedem: with my patch or before?
00:13:52 mriedem oh shit, nvm - copy paste error
00:13:59 mriedem was using the compute endpoint url from my other devstack :)
00:14:05 mriedem "wow this is fast!"
00:15:09 mriedem hah, here we go, nice and slow
00:15:16 mriedem fault loading mofos
00:20:55 mriedem 4.495s with GET /servers, 1000 ACTIVE vms. 11.185s with 500 error, 500 active
00:30:39 mriedem 24.125s to list them with details
00:42:34 openstackgerrit Michael Still proposed openstack/nova master: Move ploop commands to privsep. https://review.openstack.org/492325
00:42:34 openstackgerrit Michael Still proposed openstack/nova master: Read from console ptys using privsep. https://review.openstack.org/489486
00:42:35 openstackgerrit Michael Still proposed openstack/nova master: Don't shell out to mkdir, use ensure_tree() https://review.openstack.org/492326
00:42:35 openstackgerrit Michael Still proposed openstack/nova master: Cleanup mount / umount and associated rmdir calls https://review.openstack.org/494423
00:42:36 openstackgerrit Michael Still proposed openstack/nova master: Move lvm handling to privsep. https://review.openstack.org/495516
00:42:36 openstackgerrit Michael Still proposed openstack/nova master: Move shred to privsep. https://review.openstack.org/495537
00:42:37 openstackgerrit Michael Still proposed openstack/nova master: Move xend existence probes to privsep. https://review.openstack.org/495538
00:42:37 openstackgerrit Michael Still proposed openstack/nova master: Move the idmapshift binary into privsep. https://review.openstack.org/495541
00:42:38 openstackgerrit Michael Still proposed openstack/nova master: Move loopback setup and removal to privsep. https://review.openstack.org/495664
00:42:38 openstackgerrit Michael Still proposed openstack/nova master: Move nbd commands to privsep. https://review.openstack.org/500351
00:42:39 openstackgerrit Michael Still proposed openstack/nova master: Move kpartx calls to privsep. https://review.openstack.org/500354
00:42:39 openstackgerrit Michael Still proposed openstack/nova master: Move blkid calls to privsep. https://review.openstack.org/500398
00:44:13 mriedem interesting, listing with details and microversion 2.53 is not much worse than with microversion 2.1 for the error/active mix case - it was nearly double between microversions when all were active
00:47:48 mriedem dansmith: time for your change, do i need https://review.openstack.org/#/c/505456/ or just the one below it?
00:48:17 dansmith mriedem: the one below it should orphan those so they're never called
00:48:25 dansmith so you shouldn't notice any difference afaik
00:48:44 mriedem ok
01:38:29 mriedem dansmith: ok i have results in https://etherpad.openstack.org/p/nova-instance-list
01:38:31 mriedem with your change
01:39:18 dansmith is that faster/same except for details?
01:39:25 mriedem compared to w/o your change, (1) GET /servers with microversion 2.1 is slightly faster
01:39:42 mriedem GET /servers/detail with microversion is about the same, a bit faster
01:39:45 mriedem 2.1
01:39:56 mriedem but, GET /server/details with microversion 2.53 is slower
01:40:01 mriedem not a ton, but it's slower
01:40:10 mriedem 25.78 compared to 30.10
01:40:28 mriedem but, it's not a huge different
01:40:30 dansmith oh only detail with the later microversion
01:40:33 mriedem right
01:40:43 mriedem something about >2.1 always makes listing with details slower
01:40:44 dansmith and there's some fault handling behavior difference?
01:40:53 mriedem at least because of the joins on the (1) services table and (2) tags table
01:41:17 mriedem i don't think there is any fault handling behavior differences with microversion >2.1

Earlier   Later