Earlier  
Posted Nick Remark
#openstack-nova - 2020-05-12
16:30:14 dansmith and if everything goes nuts during that, we have to be able to heal away those stale ones when we reboot
16:30:25 dansmith and if images get sideloaded, we have to notice, and if images get locally deleted, we have to notice
16:30:36 dansmith sean-k-mooney: it does work and we (redhat) have prescribed it in a few cases :)
16:31:02 sean-k-mooney .... im glad you have added the image cacahing feature in teh api to have an alternitive noew
16:31:03 dansmith and I know people have in the past manually purged images from that cache before the timer fires
16:31:35 dansmith sean-k-mooney: yep, but that only works (by design) per aggregate, and doesn't let you purge,
16:31:59 sean-k-mooney ya
16:32:06 sean-k-mooney sitll an improvement
16:32:08 dansmith so I think it would be foolish and fragile to not be able to reconcile the state of the disk with the other system
16:32:19 dansmith for sure
16:34:31 sean-k-mooney would you be ok with the compute agent calling placment to make those allocaitons and clean them up. i think it should be fine since its already updating placment in the update_avaiable_resouces funtion just not sure if i missed anything
16:35:01 sean-k-mooney the compute node already needs to be able to reach the placmenet api so it really should not be much of a change in that regard
16:37:01 dansmith the compute node the _only_ thing that _could_ do it...
16:39:49 sean-k-mooney well normally we do the allocation candiate request in the conductor right and we claim the allocation before we get to the compute
16:40:09 sean-k-mooney so the second allcotion for the cache image would have to be done on the ocmpute node and the claim
16:40:33 sean-k-mooney but ya the compute node is really the only thing that could keep them in sync
16:41:03 dansmith but nothing outside the compute node knows about the state of the cache
16:41:17 dansmith so nothing else would have any idea if an image is cached or could be cached
16:41:24 sean-k-mooney yep
16:41:34 sean-k-mooney and even if they did it would be racey
16:41:53 dansmith you mean "even if they tried to guess" :)
16:41:54 sean-k-mooney so that means the healing task could not be part of nova audit
16:42:00 dansmith no
16:42:28 sean-k-mooney it would have to be in the compute manager which is fine
16:43:12 dansmith image cache management is kinda weirdly split between the compute manager and virt driver
16:43:24 dansmith which means we probably need some extra stuff between them I think
16:43:26 dansmith I'd have to go look
16:43:52 dansmith this *would* be a pretty heavyweight new addition, to be clear
16:44:17 sean-k-mooney ya
16:44:17 dansmith it's unfortunate that we'd need to do this, IMHO, given the complexity required
16:44:40 sean-k-mooney do you see another path forward beyound just disabling hte cache
16:44:55 dansmith so we need to make sure we think this is really worth all of that, other than kinda talking our way out of it
16:45:00 sean-k-mooney or puting a size limit on the cache and contionally disabling
16:45:50 dansmith I don't really have any better ideas, no, I just don't like this one enough to be excited about it
16:46:53 dansmith we should consider some other ideas before we pull the trigger on this I mean
16:47:21 dansmith like, we *could* look for a-cs that have $imgsize+$flavor.root available disk
16:47:44 dansmith which may generate some operator confusion, and will definitely avoid being able to schedule the last byte of disk space
16:48:08 dansmith but also makes some sense if you explain it to someone: there has to be enough disk for the image and the root, even if the image might be cached
16:49:09 dansmith and that could be a behavior you enable with a pre-filter, which tries to avoid situations like bfv
16:49:16 sean-k-mooney dansmith: we have 5 slightly different impleentation fo caching in that module
16:49:35 sean-k-mooney each image backend is slightly different but mostly the same
16:50:43 sean-k-mooney dansmith: could we do that and then srink the allocation
16:51:59 dansmith sean-k-mooney: right, we'd not allocate that much, just look for hosts with enough to cover it
16:52:12 sean-k-mooney so do $imgsize+$flavor.root then both the instance and shirnk the disk_gb ot $flavor.root
16:52:24 dansmith except in the most pathological cases, we'd be fine.. you could come up with a race scenario, but it'd be very very targeted
16:52:36 dansmith it's also something we could try and roll back without having to change or migrate data,
16:53:00 sean-k-mooney ya its worth a try
16:53:01 dansmith whereas the new allocations-per-image thing would be something we have to live with and migrate for a while if it doesn't pan out
16:53:12 sean-k-mooney ir at least consiering
16:54:00 sean-k-mooney most of the the time $imgsize+$flavor.root is not going to cause boot failure either
16:54:11 sean-k-mooney as it will only be an issue if the cloud is very full
16:54:32 dansmith right, and if it was, it's because you're trying to schedule the last byte of disk, which is something we say is not in our project scope
16:54:35 sean-k-mooney if its a configurable prefilter then those that want every last gb could opt out
16:54:52 sean-k-mooney well ya that too
16:56:06 sean-k-mooney well it sound liek we have too light weight things. a disabel cache, b look for $imgsize+$flavor.root in the placment query and srhink to $flavor.root
16:56:09 dansmith yeah, so maybe if gibi really has a do-not-cache fix here, and we provide that prefilter, maybe that's good enough for the moment
16:56:12 dansmith yep :)
16:56:25 sean-k-mooney and then we could look at the allcotion per image later if we needed too
16:57:14 sean-k-mooney ya i think he need to copy paste it to the 4 other location i commented on but it might work
16:57:24 dansmith yes, I'd be much happier punting that out to a last-resort type of situation
16:57:31 sean-k-mooney this is a part of the code i would normally ping mdbooth or lyarwood to look at
16:58:50 sean-k-mooney ok well im going to grab food o/
16:59:11 dansmith aye dee ohs
16:59:57 lyarwood oh joy the cache manager
17:00:41 sean-k-mooney lyarwood: the intent of that patch is to have a config optioon to trun it off
17:01:06 sean-k-mooney well to workaround the larger bug
17:01:48 sean-k-mooney i think i commented on the other placces where it also need to be done but i proably missed one and i also dont realy understand the calling code paths so its just a guess
17:03:31 dansmith honestly, as much as I trust you guys, I'd have to test it myself before I believed you
17:03:39 dansmith just because I know how loopy all that code is
17:03:46 dansmith it's partial'd up the butthole
17:05:18 lyarwood yup same I'd have to play around with this, at first glance I'd be worried about this racing with multiple requests to spawn from the same image tbh
17:05:42 lyarwood but I'm likely missing locking somewhere in the imagebackend or driver that stops this
17:06:15 dansmith well, this is trying to avoid ever downloading the image to the base location,
17:06:31 dansmith so it really shouldn't be able to race because only one instance is booting per instance uuid at a time obviously
17:06:47 lyarwood ah so it's not caching and then copying?
17:06:47 dansmith but the call path to know what self.path is here is the critical bit,
17:06:49 lyarwood my bad
17:06:53 dansmith because normally that is the base image path
17:07:07 lyarwood right understood
17:07:11 dansmith lyarwood: that's the assertion, I just don't know how it got to that point here
18:08:34 openstackgerrit Ghanshyam Mann proposed openstack/nova master: Bump hacking min version to 3.0.1 https://review.opendev.org/727347
18:15:19 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Don't delete disks on shared storage during evacuate https://review.opendev.org/578846
18:15:19 openstack bug 1550919 in OpenStack Compute (nova) "[Libvirt]Evacuate fail may cause disk image be deleted" [Medium,In progress] https://launchpad.net/bugs/1550919 - Assigned to Matthew Booth (mbooth-9)
18:15:19 openstackgerrit Lee Yarwood proposed openstack/nova master: Add functional test for bug 1550919 https://review.opendev.org/631294
18:52:40 melwitt gmann: I've got a queens backport that consistently fails the releasenotes job bc of a job timeout. we are wondering if, with the addition of ussuri releasenotes, if it's just too much for the job to get done in a limited time? https://review.opendev.org/722822
18:54:28 gmann melwitt: it takes lot of time for all reno to build but should not be timeout. it pass on master.
18:54:43 gmann let me check job timeout for that
18:54:54 melwitt yeah... I just don't get how my backport could cause this and only cause it on queens? confused
18:55:20 melwitt it got through rocky, stein, train without issue
19:00:25 elod melwitt: when those got merged there weren't stable/ussuri yet, were there?
19:00:46 elod this shows some master fails too: https://zuul.opendev.org/t/openstack/builds?job_name=build-openstack-releasenotes&project=openstack/nova
19:01:13 melwitt elod: no there wasn't, was merged before branching
19:01:56 melwitt oh you mean the stable merges. let me check
19:03:19 melwitt train change merged on Apr 18 and I assume stable/ussuri was cut on Apr 24
19:03:21 gmann reno job timeout is 60 min which should be well enough https://github.com/openstack/openstack-zuul-jobs/blob/13ef0adb415e6296fe5c73d9ff9d1ca557843c54/zuul.d/jobs.yaml#L638
19:03:37 melwitt stein change merged on Apr 24
19:03:55 melwitt rocky change merged on May 11
19:06:19 melwitt the rocky change build does _not_ show processing of stable/ussuri notes https://zuul.opendev.org/t/openstack/build/bd1684e2f6d4460a985d822dfe773b81/log/job-output.txt#580-583
19:06:58 melwitt whereas the failing queens change build _does_ show processing of stable/ussuri notes https://zuul.opendev.org/t/openstack/build/e63085cb0ebf4e0ea8d91e07c94557ae/log/job-output.txt#536-542 so I think elod is right
19:15:23 gmann melwitt: elod it seems job most of time run on the edge time when it pass. >50 for master reno also and 59m 52s in case of https://review.opendev.org/#/c/725146/

Earlier   Later