| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 23:40:32 | mriedem | the inventory nothing changed? | |
| 23:40:37 | melwitt | I dunno, I think I know just enough for it to be confusing. for an end user, it might not be confusing | |
| 23:40:47 | dansmith | mriedem: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L571-L572 | |
| 23:40:56 | mriedem | melwitt: end user won't see it, they'll see NoValidHost on 500 instances | |
| 23:41:01 | mriedem | the operator will see it | |
| 23:41:04 | melwitt | good point | |
| 23:41:47 | dansmith | mriedem: or just look at placement logs to see if inventory is hit any time after compute startup | |
| 23:46:59 | mriedem | gdi, how do i regex search with grep | |
| 23:47:06 | mriedem | sudo journalctl -a -u devstack@placement-api.service | grep '.*PUT.*\/inventories.*' | |
| 23:48:12 | dansmith | PUT.*invent should be all you need | |
| 23:49:29 | mriedem | doesn't work | |
| 23:51:13 | mriedem | ah, well, | |
| 23:51:16 | mriedem | PUT.*alloc works | |
| 23:51:21 | mriedem | so it probably just wrapped | |
| 23:51:28 | mriedem | and it's not updating inventory, as it shouldn't | |
| 23:52:16 | mriedem | got my 1000 instances now, so will do the test stuff once i'm done with dinner | |
| 23:53:00 | dansmith | so I'd also check to make sure compute isn't doing the ocata healing during boot or something like that | |
| 23:57:18 | takashin | Spec cores, could you review https://review.openstack.org/#/c/489029/ ? It got one +2. | |
| #openstack-nova - 2017-09-27 | |||
| 00:13:05 | mriedem | dansmith: interesting, listing without details, 500 error (cell0) and 500 active (cell1) is a lot faster than the 1000 active, | |
| 00:13:13 | mriedem | i suppose because we don't have as much to join | |
| 00:13:27 | dansmith | mriedem: with my patch or before? | |
| 00:13:52 | mriedem | oh shit, nvm - copy paste error | |
| 00:13:59 | mriedem | was using the compute endpoint url from my other devstack :) | |
| 00:14:05 | mriedem | "wow this is fast!" | |
| 00:15:09 | mriedem | hah, here we go, nice and slow | |
| 00:15:16 | mriedem | fault loading mofos | |
| 00:20:55 | mriedem | 4.495s with GET /servers, 1000 ACTIVE vms. 11.185s with 500 error, 500 active | |
| 00:30:39 | mriedem | 24.125s to list them with details | |
| 00:42:34 | openstackgerrit | Michael Still proposed openstack/nova master: Read from console ptys using privsep. https://review.openstack.org/489486 | |
| 00:42:34 | openstackgerrit | Michael Still proposed openstack/nova master: Move ploop commands to privsep. https://review.openstack.org/492325 | |
| 00:42:35 | openstackgerrit | Michael Still proposed openstack/nova master: Cleanup mount / umount and associated rmdir calls https://review.openstack.org/494423 | |
| 00:42:35 | openstackgerrit | Michael Still proposed openstack/nova master: Don't shell out to mkdir, use ensure_tree() https://review.openstack.org/492326 | |
| 00:42:36 | openstackgerrit | Michael Still proposed openstack/nova master: Move shred to privsep. https://review.openstack.org/495537 | |
| 00:42:36 | openstackgerrit | Michael Still proposed openstack/nova master: Move lvm handling to privsep. https://review.openstack.org/495516 | |
| 00:42:37 | openstackgerrit | Michael Still proposed openstack/nova master: Move the idmapshift binary into privsep. https://review.openstack.org/495541 | |
| 00:42:37 | openstackgerrit | Michael Still proposed openstack/nova master: Move xend existence probes to privsep. https://review.openstack.org/495538 | |
| 00:42:38 | openstackgerrit | Michael Still proposed openstack/nova master: Move nbd commands to privsep. https://review.openstack.org/500351 | |
| 00:42:38 | openstackgerrit | Michael Still proposed openstack/nova master: Move loopback setup and removal to privsep. https://review.openstack.org/495664 | |
| 00:42:39 | openstackgerrit | Michael Still proposed openstack/nova master: Move blkid calls to privsep. https://review.openstack.org/500398 | |
| 00:42:39 | openstackgerrit | Michael Still proposed openstack/nova master: Move kpartx calls to privsep. https://review.openstack.org/500354 | |
| 00:44:13 | mriedem | interesting, listing with details and microversion 2.53 is not much worse than with microversion 2.1 for the error/active mix case - it was nearly double between microversions when all were active | |
| 00:47:48 | mriedem | dansmith: time for your change, do i need https://review.openstack.org/#/c/505456/ or just the one below it? | |
| 00:48:17 | dansmith | mriedem: the one below it should orphan those so they're never called | |
| 00:48:25 | dansmith | so you shouldn't notice any difference afaik | |
| 00:48:44 | mriedem | ok | |
| 01:38:29 | mriedem | dansmith: ok i have results in https://etherpad.openstack.org/p/nova-instance-list | |
| 01:38:31 | mriedem | with your change | |
| 01:39:18 | dansmith | is that faster/same except for details? | |
| 01:39:25 | mriedem | compared to w/o your change, (1) GET /servers with microversion 2.1 is slightly faster | |
| 01:39:42 | mriedem | GET /servers/detail with microversion is about the same, a bit faster | |
| 01:39:45 | mriedem | 2.1 | |
| 01:39:56 | mriedem | but, GET /server/details with microversion 2.53 is slower | |
| 01:40:01 | mriedem | not a ton, but it's slower | |
| 01:40:10 | mriedem | 25.78 compared to 30.10 | |
| 01:40:28 | mriedem | but, it's not a huge different | |
| 01:40:30 | dansmith | oh only detail with the later microversion | |
| 01:40:33 | mriedem | right | |
| 01:40:43 | mriedem | something about >2.1 always makes listing with details slower | |
| 01:40:44 | dansmith | and there's some fault handling behavior difference? | |
| 01:40:53 | mriedem | at least because of the joins on the (1) services table and (2) tags table | |
| 01:41:17 | mriedem | i don't think there is any fault handling behavior differences with microversion >2.1 | |
| 01:41:22 | mriedem | if there was, that might explain it | |
| 01:41:27 | dansmith | okay I thought you were saying there was | |
| 01:41:41 | dansmith | I dunno why because I'm pre-joining it when we were loading them separate | |
| 01:41:48 | mriedem | the only other joins i can think of right now with microversion >2.1 is on the services table (2.16) and tags able (2.26) | |
| 01:42:26 | mriedem | still, it's a difference of about 4 seconds, which isn't huge here | |
| 01:42:35 | dansmith | so aside from fault, there's no difference in what I'm doing vs what we do currently, | |
| 01:42:42 | dansmith | other than we're not serializing the queries | |
| 01:43:18 | dansmith | without my change we issue the cell0 one and then the cell1 one, where now we're doing both at once | |
| 01:43:28 | dansmith | is this a devstack vm on your laptop or something better? | |
| 01:43:38 | mriedem | it's in a vexxhost vm | |
| 01:44:43 | mriedem | the fault stuff is the only major difference i can think of, since we'll be joining on fault all the time, rather than just for instances in ERROR state | |
| 01:45:09 | mriedem | maybe that is equaling things out somehow, idk, like if i had 1000 all in ERROR state before/after your change, that might be different in favor of yours | |
| 01:45:31 | dansmith | hmm, yeah, I guess maybe that might be it | |
| 01:46:08 | mriedem | i do have the numbers from yesterday before your change with 1000 ACTIVE, | |
| 01:46:20 | mriedem | so tomorrow i could run yours through with all active and see if there is a bigger difference because of the fault join | |
| 01:46:26 | dansmith | well, I guess we could go back to the not automatic loading of fault | |
| 01:46:48 | mriedem | i'll run that all active scenario tomorrow to see if it could be the fault stuff, | |
| 01:46:54 | mriedem | it's nearly 9pm so i'm not going to do it tonight | |
| 01:47:09 | dansmith | there was something the API was doing that made it seem way better to do this than what it was doing | |
| 01:47:19 | dansmith | but it's been a while now | |
| 01:48:34 | dansmith | we could also plumb the logic of when to load the fault into the lower layers | |
| 01:49:14 | mriedem | yup i was thinking that too | |
| 01:49:34 | mriedem | another thing that might be causing the microversion bloat, is maybe the microversion to pull the embedded flavor out of the instance | |
| 01:49:42 | mriedem | added in pike | |
| 01:50:01 | dansmith | you could run through each microversion and see where the spike is | |
| 01:50:33 | dansmith | the sorting layer on top of this really has nothing to do with what we're sorting though | |
| 01:50:42 | dansmith | it doesn't make any more copies of things, nor iterate the list more times | |
| 01:50:47 | mriedem | 2.47 | |
| 01:51:59 | dansmith | so, the change right before the switchover should do the fault loading but not the sorting, so you could run against that and see if it's more like the earlier or more like the later | |
| 01:53:49 | mriedem | https://review.openstack.org/#/c/506774/ ? | |
| 01:53:59 | mriedem | like, revert that on top of the change that uses the new code in the API | |
| 01:54:00 | mriedem | ? | |
| 01:54:14 | mriedem | oh, nvm, | |
| 01:54:15 | dansmith | oh, I guess you were running on master already? | |
| 01:54:23 | mriedem | yes, new devstack as of today | |
| 01:54:38 | dansmith | yeah, okay | |
| 01:55:19 | mriedem | so 2.16 makes us join on services, 2.26 makes us join on tags, 2.47 returns instance.flavor, and your change always joins on faults | |
| 01:55:25 | mriedem | 2.47 is suspicious | |
| 01:55:31 | mriedem | since that's from instance_extra | |