| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 23:12:21 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Log instance uuid when retrying claims in the scheduler https://review.openstack.org/507705 | |
| 23:16:59 | openstackgerrit | Moshe Levi proposed openstack/nova master: Don't overwrite binding-profile https://review.openstack.org/505613 | |
| 23:24:42 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Return Selection Objects https://review.openstack.org/498830 | |
| 23:25:48 | dansmith | I'm actually not sure why we'd be hitting concurrent updates during an allocation event, | |
| 23:25:54 | dansmith | given that we're not passing rp_generation | |
| 23:26:08 | dansmith | either the allocation fits or doesn't | |
| 23:27:39 | melwitt | hm, yeah | |
| 23:27:43 | dansmith | it must be that placement doesn't hide transaction commits from us | |
| 23:27:46 | dansmith | based on the commit that added it | |
| 23:28:12 | dansmith | which would just be "tons of churn on one provider" as the reason | |
| 23:29:57 | melwitt | what does that mean? placement not hiding transaction commits | |
| 23:31:14 | dansmith | an allocation will increment the generation of the provider on the placement side | |
| 23:31:23 | dansmith | and a transaction will abort if we try to do it at the same time as something else | |
| 23:31:26 | melwitt | ah, right. that's what I was thinking of | |
| 23:31:52 | dansmith | placement should really retry those for us server-side I would think, for an allocation type request | |
| 23:31:58 | melwitt | I couldn't remember where the RP generation was related there | |
| 23:32:22 | dansmith | doesn't matter, but I would think it would be better, especially given the advice for that error is "try exactly the same thing again" | |
| 23:32:24 | melwitt | yeah, I was thinking the same | |
| 23:32:29 | mriedem | this was the error from the server | |
| 23:32:30 | mriedem | There was a conflict when trying to complete your request.\n\n Inventory changed while attempting to allocate: Another thread concurrently updated the data. Please retry your update | |
| 23:32:38 | dansmith | right | |
| 23:32:41 | mriedem | i'm not sure why inventory would change | |
| 23:32:45 | mriedem | that's static in this case | |
| 23:33:01 | dansmith | any allocation will change the generation | |
| 23:33:04 | mriedem | the update_available_resource periodic will post inventory, but only if it changes | |
| 23:33:19 | dansmith | so any two allocations can conflict | |
| 23:33:21 | melwitt | I'm +1 on the idea of retrying server-side | |
| 23:34:16 | melwitt | I'm a little worried how many conflicts can we get in real life by ppl trying to create 500 servers and the scheduler is doing a "pack" pattern | |
| 23:35:10 | melwitt | as far as how to tune how many retries | |
| 23:35:23 | melwitt | to allow | |
| 23:35:33 | dansmith | this is also a highly synthetic scenario with a "virt driver" that doesn't get looked at much.. it could be doing something to re-stab inventory for no reason or something | |
| 23:36:00 | mriedem | for each instance, we go through the filters | |
| 23:36:17 | mriedem | and doesn't the scheduler have some kind of tracking on the HostState objects themselves for chosen hosts? | |
| 23:36:22 | edleafe | mriedem: that error message is poorly worded. It should be something like "available inventory has changed" | |
| 23:36:25 | melwitt | oh, the "inventory changed" yeah, I don't know anything about that. so that means it wasn't an allocation writing conflict? | |
| 23:36:41 | dansmith | melwitt: it's al related | |
| 23:36:45 | dansmith | *all | |
| 23:37:07 | dansmith | edleafe: that's not really accurate, AFACT, since allocations will cause the generation to increase | |
| 23:37:16 | dansmith | edleafe: and thus it could be nothing changed with inventory to cause that | |
| 23:37:34 | edleafe | dansmith: uh, that's why I added "available". | |
| 23:37:46 | edleafe | dansmith: IOW, some inventory has been allocated | |
| 23:37:47 | mriedem | this is basically the inventory reporting for the fake driver https://github.com/openstack/nova/blob/master/nova/virt/fake.py#L111 | |
| 23:37:51 | edleafe | and changed the generation | |
| 23:38:01 | dansmith | edleafe: okay, I wouldn't word it that way for clarity, but okay :) | |
| 23:38:47 | edleafe | I wouldn't word it that way either, but I was guessing the author's intent | |
| 23:39:07 | edleafe | *cough* cdent *cough* | |
| 23:39:08 | dansmith | mriedem: that's the info that we use to generate it, yes | |
| 23:39:33 | melwitt | yeah, if it can happen without inventory (total possible capacity) changing, then that error message is confusing to me | |
| 23:40:05 | edleafe | melwitt: yeah, it's worded very poorly | |
| 23:40:10 | dansmith | mriedem: I'd look to see if the compute is hitting placement /inventory ever after the first go, and maybe check the nothing-changed short-circuit to make sure we're never going through it | |
| 23:40:32 | mriedem | the inventory nothing changed? | |
| 23:40:37 | melwitt | I dunno, I think I know just enough for it to be confusing. for an end user, it might not be confusing | |
| 23:40:47 | dansmith | mriedem: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L571-L572 | |
| 23:40:56 | mriedem | melwitt: end user won't see it, they'll see NoValidHost on 500 instances | |
| 23:41:01 | mriedem | the operator will see it | |
| 23:41:04 | melwitt | good point | |
| 23:41:47 | dansmith | mriedem: or just look at placement logs to see if inventory is hit any time after compute startup | |
| 23:46:59 | mriedem | gdi, how do i regex search with grep | |
| 23:47:06 | mriedem | sudo journalctl -a -u devstack@placement-api.service | grep '.*PUT.*\/inventories.*' | |
| 23:48:12 | dansmith | PUT.*invent should be all you need | |
| 23:49:29 | mriedem | doesn't work | |
| 23:51:13 | mriedem | ah, well, | |
| 23:51:16 | mriedem | PUT.*alloc works | |
| 23:51:21 | mriedem | so it probably just wrapped | |
| 23:51:28 | mriedem | and it's not updating inventory, as it shouldn't | |
| 23:52:16 | mriedem | got my 1000 instances now, so will do the test stuff once i'm done with dinner | |
| 23:53:00 | dansmith | so I'd also check to make sure compute isn't doing the ocata healing during boot or something like that | |
| 23:57:18 | takashin | Spec cores, could you review https://review.openstack.org/#/c/489029/ ? It got one +2. | |
| #openstack-nova - 2017-09-27 | |||
| 00:13:05 | mriedem | dansmith: interesting, listing without details, 500 error (cell0) and 500 active (cell1) is a lot faster than the 1000 active, | |
| 00:13:13 | mriedem | i suppose because we don't have as much to join | |
| 00:13:27 | dansmith | mriedem: with my patch or before? | |
| 00:13:52 | mriedem | oh shit, nvm - copy paste error | |
| 00:13:59 | mriedem | was using the compute endpoint url from my other devstack :) | |
| 00:14:05 | mriedem | "wow this is fast!" | |
| 00:15:09 | mriedem | hah, here we go, nice and slow | |
| 00:15:16 | mriedem | fault loading mofos | |
| 00:20:55 | mriedem | 4.495s with GET /servers, 1000 ACTIVE vms. 11.185s with 500 error, 500 active | |
| 00:30:39 | mriedem | 24.125s to list them with details | |
| 00:42:34 | openstackgerrit | Michael Still proposed openstack/nova master: Read from console ptys using privsep. https://review.openstack.org/489486 | |
| 00:42:34 | openstackgerrit | Michael Still proposed openstack/nova master: Move ploop commands to privsep. https://review.openstack.org/492325 | |
| 00:42:35 | openstackgerrit | Michael Still proposed openstack/nova master: Cleanup mount / umount and associated rmdir calls https://review.openstack.org/494423 | |
| 00:42:35 | openstackgerrit | Michael Still proposed openstack/nova master: Don't shell out to mkdir, use ensure_tree() https://review.openstack.org/492326 | |
| 00:42:36 | openstackgerrit | Michael Still proposed openstack/nova master: Move shred to privsep. https://review.openstack.org/495537 | |
| 00:42:36 | openstackgerrit | Michael Still proposed openstack/nova master: Move lvm handling to privsep. https://review.openstack.org/495516 | |
| 00:42:37 | openstackgerrit | Michael Still proposed openstack/nova master: Move the idmapshift binary into privsep. https://review.openstack.org/495541 | |
| 00:42:37 | openstackgerrit | Michael Still proposed openstack/nova master: Move xend existence probes to privsep. https://review.openstack.org/495538 | |
| 00:42:38 | openstackgerrit | Michael Still proposed openstack/nova master: Move nbd commands to privsep. https://review.openstack.org/500351 | |
| 00:42:38 | openstackgerrit | Michael Still proposed openstack/nova master: Move loopback setup and removal to privsep. https://review.openstack.org/495664 | |
| 00:42:39 | openstackgerrit | Michael Still proposed openstack/nova master: Move blkid calls to privsep. https://review.openstack.org/500398 | |
| 00:42:39 | openstackgerrit | Michael Still proposed openstack/nova master: Move kpartx calls to privsep. https://review.openstack.org/500354 | |
| 00:44:13 | mriedem | interesting, listing with details and microversion 2.53 is not much worse than with microversion 2.1 for the error/active mix case - it was nearly double between microversions when all were active | |
| 00:47:48 | mriedem | dansmith: time for your change, do i need https://review.openstack.org/#/c/505456/ or just the one below it? | |
| 00:48:17 | dansmith | mriedem: the one below it should orphan those so they're never called | |
| 00:48:25 | dansmith | so you shouldn't notice any difference afaik | |
| 00:48:44 | mriedem | ok | |
| 01:38:29 | mriedem | dansmith: ok i have results in https://etherpad.openstack.org/p/nova-instance-list | |
| 01:38:31 | mriedem | with your change | |
| 01:39:18 | dansmith | is that faster/same except for details? | |
| 01:39:25 | mriedem | compared to w/o your change, (1) GET /servers with microversion 2.1 is slightly faster | |
| 01:39:42 | mriedem | GET /servers/detail with microversion is about the same, a bit faster | |