Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-26
23:40:37 melwitt I dunno, I think I know just enough for it to be confusing. for an end user, it might not be confusing
23:40:47 dansmith mriedem: https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L571-L572
23:40:56 mriedem melwitt: end user won't see it, they'll see NoValidHost on 500 instances
23:41:01 mriedem the operator will see it
23:41:04 melwitt good point
23:41:47 dansmith mriedem: or just look at placement logs to see if inventory is hit any time after compute startup
23:46:59 mriedem gdi, how do i regex search with grep
23:47:06 mriedem sudo journalctl -a -u devstack@placement-api.service | grep '.*PUT.*\/inventories.*'
23:48:12 dansmith PUT.*invent should be all you need
23:49:29 mriedem doesn't work
23:51:13 mriedem ah, well,
23:51:16 mriedem PUT.*alloc works
23:51:21 mriedem so it probably just wrapped
23:51:28 mriedem and it's not updating inventory, as it shouldn't
23:52:16 mriedem got my 1000 instances now, so will do the test stuff once i'm done with dinner
23:53:00 dansmith so I'd also check to make sure compute isn't doing the ocata healing during boot or something like that
23:57:18 takashin Spec cores, could you review https://review.openstack.org/#/c/489029/ ? It got one +2.
#openstack-nova - 2017-09-27
00:13:05 mriedem dansmith: interesting, listing without details, 500 error (cell0) and 500 active (cell1) is a lot faster than the 1000 active,
00:13:13 mriedem i suppose because we don't have as much to join
00:13:27 dansmith mriedem: with my patch or before?
00:13:52 mriedem oh shit, nvm - copy paste error
00:13:59 mriedem was using the compute endpoint url from my other devstack :)
00:14:05 mriedem "wow this is fast!"
00:15:09 mriedem hah, here we go, nice and slow
00:15:16 mriedem fault loading mofos
00:20:55 mriedem 4.495s with GET /servers, 1000 ACTIVE vms. 11.185s with 500 error, 500 active
00:30:39 mriedem 24.125s to list them with details
00:42:34 openstackgerrit Michael Still proposed openstack/nova master: Move ploop commands to privsep. https://review.openstack.org/492325
00:42:34 openstackgerrit Michael Still proposed openstack/nova master: Read from console ptys using privsep. https://review.openstack.org/489486
00:42:35 openstackgerrit Michael Still proposed openstack/nova master: Don't shell out to mkdir, use ensure_tree() https://review.openstack.org/492326
00:42:35 openstackgerrit Michael Still proposed openstack/nova master: Cleanup mount / umount and associated rmdir calls https://review.openstack.org/494423
00:42:36 openstackgerrit Michael Still proposed openstack/nova master: Move lvm handling to privsep. https://review.openstack.org/495516
00:42:36 openstackgerrit Michael Still proposed openstack/nova master: Move shred to privsep. https://review.openstack.org/495537
00:42:37 openstackgerrit Michael Still proposed openstack/nova master: Move xend existence probes to privsep. https://review.openstack.org/495538
00:42:37 openstackgerrit Michael Still proposed openstack/nova master: Move the idmapshift binary into privsep. https://review.openstack.org/495541
00:42:38 openstackgerrit Michael Still proposed openstack/nova master: Move loopback setup and removal to privsep. https://review.openstack.org/495664
00:42:38 openstackgerrit Michael Still proposed openstack/nova master: Move nbd commands to privsep. https://review.openstack.org/500351
00:42:39 openstackgerrit Michael Still proposed openstack/nova master: Move kpartx calls to privsep. https://review.openstack.org/500354
00:42:39 openstackgerrit Michael Still proposed openstack/nova master: Move blkid calls to privsep. https://review.openstack.org/500398
00:44:13 mriedem interesting, listing with details and microversion 2.53 is not much worse than with microversion 2.1 for the error/active mix case - it was nearly double between microversions when all were active
00:47:48 mriedem dansmith: time for your change, do i need https://review.openstack.org/#/c/505456/ or just the one below it?
00:48:17 dansmith mriedem: the one below it should orphan those so they're never called
00:48:25 dansmith so you shouldn't notice any difference afaik
00:48:44 mriedem ok
01:38:29 mriedem dansmith: ok i have results in https://etherpad.openstack.org/p/nova-instance-list
01:38:31 mriedem with your change
01:39:18 dansmith is that faster/same except for details?
01:39:25 mriedem compared to w/o your change, (1) GET /servers with microversion 2.1 is slightly faster
01:39:42 mriedem GET /servers/detail with microversion is about the same, a bit faster
01:39:45 mriedem 2.1
01:39:56 mriedem but, GET /server/details with microversion 2.53 is slower
01:40:01 mriedem not a ton, but it's slower
01:40:10 mriedem 25.78 compared to 30.10
01:40:28 mriedem but, it's not a huge different
01:40:30 dansmith oh only detail with the later microversion
01:40:33 mriedem right
01:40:43 mriedem something about >2.1 always makes listing with details slower
01:40:44 dansmith and there's some fault handling behavior difference?
01:40:53 mriedem at least because of the joins on the (1) services table and (2) tags table
01:41:17 mriedem i don't think there is any fault handling behavior differences with microversion >2.1
01:41:22 mriedem if there was, that might explain it
01:41:27 dansmith okay I thought you were saying there was
01:41:41 dansmith I dunno why because I'm pre-joining it when we were loading them separate
01:41:48 mriedem the only other joins i can think of right now with microversion >2.1 is on the services table (2.16) and tags able (2.26)
01:42:26 mriedem still, it's a difference of about 4 seconds, which isn't huge here
01:42:35 dansmith so aside from fault, there's no difference in what I'm doing vs what we do currently,
01:42:42 dansmith other than we're not serializing the queries
01:43:18 dansmith without my change we issue the cell0 one and then the cell1 one, where now we're doing both at once
01:43:28 dansmith is this a devstack vm on your laptop or something better?
01:43:38 mriedem it's in a vexxhost vm
01:44:43 mriedem the fault stuff is the only major difference i can think of, since we'll be joining on fault all the time, rather than just for instances in ERROR state
01:45:09 mriedem maybe that is equaling things out somehow, idk, like if i had 1000 all in ERROR state before/after your change, that might be different in favor of yours
01:45:31 dansmith hmm, yeah, I guess maybe that might be it
01:46:08 mriedem i do have the numbers from yesterday before your change with 1000 ACTIVE,
01:46:20 mriedem so tomorrow i could run yours through with all active and see if there is a bigger difference because of the fault join
01:46:26 dansmith well, I guess we could go back to the not automatic loading of fault
01:46:48 mriedem i'll run that all active scenario tomorrow to see if it could be the fault stuff,
01:46:54 mriedem it's nearly 9pm so i'm not going to do it tonight
01:47:09 dansmith there was something the API was doing that made it seem way better to do this than what it was doing
01:47:19 dansmith but it's been a while now
01:48:34 dansmith we could also plumb the logic of when to load the fault into the lower layers
01:49:14 mriedem yup i was thinking that too
01:49:34 mriedem another thing that might be causing the microversion bloat, is maybe the microversion to pull the embedded flavor out of the instance
01:49:42 mriedem added in pike
01:50:01 dansmith you could run through each microversion and see where the spike is
01:50:33 dansmith the sorting layer on top of this really has nothing to do with what we're sorting though
01:50:42 dansmith it doesn't make any more copies of things, nor iterate the list more times
01:50:47 mriedem 2.47
01:51:59 dansmith so, the change right before the switchover should do the fault loading but not the sorting, so you could run against that and see if it's more like the earlier or more like the later
01:53:49 mriedem https://review.openstack.org/#/c/506774/ ?
01:53:59 mriedem like, revert that on top of the change that uses the new code in the API
01:54:00 mriedem ?
01:54:14 mriedem oh, nvm,
01:54:15 dansmith oh, I guess you were running on master already?
01:54:23 mriedem yes, new devstack as of today
01:54:38 dansmith yeah, okay
01:55:19 mriedem so 2.16 makes us join on services, 2.26 makes us join on tags, 2.47 returns instance.flavor, and your change always joins on faults
01:55:25 mriedem 2.47 is suspicious
01:55:31 mriedem since that's from instance_extra
01:55:41 dansmith but again, it shouldn't be any different

Earlier   Later