Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-31
17:33:50 jaypipes dansmith: because the scheduler is the thing that creates that "doubled-up" allocation.
17:34:11 dansmith jaypipes: oh I thought you were talking about how to clean up the doubled-for-same-host allocation on the compute
17:34:50 dansmith jaypipes: why does the scheduler need to probe which is the compute by looking at vpcu? it just needs to add the new allocation to the existing one, and if the RPs are the same then sum the values
17:35:08 jaypipes dansmith: that's part of it I yeah, and that would be unnecessary once we remove the allocations on compute stuff, but there's the first step needed to actually create the doubled-up alloc in the scheduler.
17:35:41 jaypipes dansmith: ack, yeah that's true.
17:36:00 cdent no if the rps are the same, it might be shared
17:36:21 cdent so we need a way to distinguish a same rp that is a "host"
17:36:29 cdent and don't want to double the shared
17:36:34 openstackgerrit Merged openstack/nova master: Add cinder keystone client opts to config reference https://review.openstack.org/488530
17:37:08 dansmith cdent: that's not necessarily true
17:37:12 dansmith cdent: if you have a shared ceph provider,
17:37:26 dansmith you still need a double allocation during the migration, since you copy (even if COW) the disk
17:37:41 cdent the current code isn't doing that
17:37:45 dansmith you don't for a live migration, but I think that's something we can ignore as the scheduler doesn't know
17:38:36 cdent if we're happy to over-allocate for the duration of the move, this gets a lot simpler
17:39:01 cdent but the current code tries rather strenuously to not over-allocate, nor to be resource class conscious
17:39:08 openstackgerrit Merged openstack/nova master: Increase cpu time for image conversion https://review.openstack.org/486642
17:39:24 dansmith cdent: well currently things are wrong, but point me at something specific if you want :)
17:39:32 cdent heh
17:40:07 jmlowe cdent: thank you kindly, you are on a pretty reliable 4k messages between bourbons
17:41:35 cdent dansmith: in https://review.openstack.org/#/c/487589/ we added the doubling code
17:42:01 cdent it doesn't double shared providers
17:42:25 cdent jmlowe: i have no response to that
17:43:14 cdent dansmith: are you suggesting: merge the source and dest into one, if the rps are the same, sum ?
17:44:18 dansmith cdent: ah I see what you mean.. however, that's the broken code
17:44:43 jmlowe cdent: nova is supposed to take over managing network interfaces?
17:44:51 dansmith cdent: it would assume that the whole compute is shared because there are already allocations for that node right?
17:45:17 cdent dansmith: yes
17:45:18 dansmith cdent: the only way around this, aside from looking for compute providers by VCPU, is to have traits, which we don't have
17:45:21 jmlowe cdent: specifically adding and removing them from instances
17:46:03 dansmith cdent: personally, I don't know why we wouldn't just unceremoniously allocate the additional things (summing where there are existing allocations)
17:46:08 dansmith because for cold migration, resize, etc, you *do* need the doubled allocation for the shared provider,
17:46:14 dansmith live migration being the only one where you don't, AfAIK
17:46:28 dansmith jaypipes: right?
17:46:31 cdent dansmith: I'd be happy with that
17:46:40 cdent jmlowe: wat?
17:47:59 diablo_rojo mriedem, any idea how many people will be coming for Nova to the PTG?
17:48:09 jaypipes dansmith: I think you're on to something here...
17:49:15 dansmith diablo_rojo: how would mriedem know that?
17:49:37 jaypipes dansmith: that would definitely mean the cleanup process would be ickier, though. in that confirm_resize() block we'd need to somehow figure out how to also remove the shared provider doubled allocations
17:50:01 dansmith jaypipes: I think it's wrong otherwise
17:50:10 diablo_rojo dansmith, being your elected leader I quessed he might have some notion :)
17:50:33 dansmith diablo_rojo: I don't think his elected status means he tracks individual plans :)
17:50:39 dansmith diablo_rojo: if you're asking because of ticket sales,
17:50:43 diablo_rojo Just looking for an approximation. More or less than last time etc.
17:50:49 dansmith diablo_rojo: we _just_ got approval for the first round of people last week, so..
17:50:54 diablo_rojo dansmith, not for ticket sales. For room allocation.
17:51:05 diablo_rojo I can always put nova in a closet ;)
17:51:37 jmlowe cdent: I thought I remembered something about moving the add and remove of instance nics out of neutron and into nova, could have been imagining that
17:51:48 diablo_rojo dansmith, I know estimates are rough for everyone, but a general idea of how many would be a huge help.
17:52:28 jaypipes diablo_rojo: 25-35 would be my guess.
17:52:56 cfriesen_ question about resource tracking...why do we need the "reserved_host_cpus" config option when we already have "vcpu_pin_set"?
17:53:06 cdent jaypipes: does the confirm resize have access to the flavor/requestspec/whatever against which it could compare the values of a retrieve allocation against?
17:53:11 diablo_rojo jaypipes, that works. Thanks :)
17:53:26 jaypipes cdent: yes.
17:54:47 cdent jaypipes: actually, why compare, just write the allocation that is the result of the spec, it will cleanup the shared doubling there and then, no math required?
17:55:23 dansmith cdent: the allocation that is the result of the old flavor is non-trivial to determine.. size, yes, but RP, not so much
17:55:40 cdent we don't care about the old flavor
17:55:45 cdent we're going to clobber the old flavor
17:55:51 cdent the allocations that resulted from the old flavor
17:55:55 jaypipes cdent: because unfortunately the source host is what runs confirm_resize() and doesn't know the UUID of the dest host so the only way to do it is to look for the source host's UUID in the allocation list and delete those, leaving the others
17:56:07 dansmith cdent: well, if you meant subtract the old one to get the new net, but I don't think we know enough to generate a full new one
17:56:09 cdent argh!
17:56:10 dansmith because yeah, that&
17:56:31 jaypipes though I suppose the source host *could* look up the uUID of the dest host by looking at the Migration object
17:56:45 cdent I know understand why there was so much table flipping last week
17:56:47 dansmith please no
17:56:48 cdent now
17:57:01 jaypipes dansmith: yeah, I don't want to do that either.
17:57:36 jaypipes cdent: ok, so do you have all the answers and ideas you need to work on that bug?
17:57:51 cdent jaypipes: apparently not
17:57:58 cdent as we keep dismissing solutions
17:58:52 dansmith is it time for our daily hangout?
17:58:55 jaypipes cdent: well, all of these patches I pretty much consider just "hey, here's one solution to this problem, can you all check it out".
17:58:58 cdent or should I just go ahead and do the the VCPU introspection
17:59:04 jaypipes dansmith: I'm game
17:59:15 cdent i've got a different hangout now :(
17:59:21 jaypipes cdent: no, I think we're recommending trying the "just sum it" approach.
17:59:30 jaypipes from Mr. dansmith
17:59:35 cdent but we said we can't clean up the sum it approach?
18:00:23 dansmith if we're the only RP in the allocations,
18:00:25 jaypipes cdent: for shared providers, we will heal that on the *destination host* but after the move operation is entirely ended..
18:00:36 dansmith jaypipes: we will?
18:00:39 dansmith I don't think we will
18:01:08 jaypipes dansmith: yeah, because _update_usage_from_instance() will overwrite the allocations to match a single amount of the flavor.
18:01:23 dansmith it can't
18:01:26 dansmith until after confirm
18:01:37 jaypipes right, which is what I said above, no?
18:01:38 dansmith and the destination doesn't know about confirm
18:01:57 jaypipes dansmith: "but after the move operation is entirely ended.."
18:02:22 dansmith oh
18:02:28 dansmith yeah, and you can't
18:02:36 dansmith the destination host does not know when the move has ended
18:02:57 dansmith jaypipes: https://hangouts.google.com/call/2phok3vj6nhipcx6gp62zvaly4u
18:23:59 cdent dansmith, jaypipes still hanging out?
18:24:05 dansmith cdent: yes
18:24:14 dansmith cdent: just getting to the "wtf now?" phase
18:30:53 dansmith mriedem: are you aware of a patch up that adds uuid for migration objects?
18:47:28 openstackgerrit Merged openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/488283
18:50:29 chohoor I have a question, driver.deallocate_networks_on_reschedule(instance) will return True or False when nova do reschedulter(/nova/compute/manager.py:1822), why only ironic driver return True but other drivers return False?

Earlier   Later