Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-08
17:32:28 dansmith yeah, but I looked at it, I don't think it's worth the trouble
17:33:04 mriedem this is also gross https://github.com/openstack/nova/blob/master/nova/scheduler/utils.py#L325
17:33:05 mriedem but meh
17:33:24 mriedem that code probably dies in stein when i can drop the legacy reqspec compat stuff
17:33:41 melwitt heh
17:40:39 openstackgerrit Merged openstack/nova master: conf: Deprecate 'network_manager' https://review.openstack.org/530923
18:00:12 openstackgerrit Merged openstack/python-novaclient master: Use uuidutils of oslo.utils https://review.openstack.org/589717
18:19:23 cdent my stats thank you mriedem
18:23:35 openstackgerrit melanie witt proposed openstack/nova master: Add functional test for affinity with multiple cells https://review.openstack.org/585073
18:23:36 openstackgerrit melanie witt proposed openstack/nova master: Make scheduler.utils.setup_instance_group query all cells https://review.openstack.org/540258
18:28:52 mriedem cdent: heh because i approved 2 changes in 2 days?
18:35:18 openstackgerrit melanie witt proposed openstack/nova master: Use TLSv1.2 for secure VNC access https://review.openstack.org/589992
18:46:52 cdent mriedem: no for assigning that ancient bug t me
18:47:48 mriedem oh heh
18:47:54 mriedem yeah doing some late summer cleaning
19:08:48 dansmith mriedem: melwitt assuming no cells meeting today
19:09:15 openstackgerrit Matt Riedemann proposed openstack/nova master: FakeDriver: adding and removing instances on live migration. https://review.openstack.org/243613
19:09:16 mriedem i don't think i have any major developments since last week,
19:09:44 mriedem as of yesterday, it sounds like tssurya's handling a down cell spec will need to also account for doing server group calculations where members are in a down cell
19:09:48 mriedem and what we do in that case
19:10:41 melwitt yeah, that. and the multi-cell affinity bug fix of mine is finally ready for review again, I got a decent functional test working with it https://review.openstack.org/540258
19:11:03 dansmith ack, since FF I've been assuming we'll pick that up after things open up
19:11:04 dansmith melwitt: yeah but not until after rc1 at least right?
19:11:10 melwitt and got rid of the unit test malarkey
19:11:48 melwitt dansmith: yeah. given the timing and the latentness + multi-cell-only nature, wait until after rc1
19:13:04 melwitt was just bringing it up for celly info
19:14:00 mriedem in cells related news, i pushed up a fix for this over the weekend https://review.openstack.org/#/c/588943/
19:14:16 mriedem follows the same pattern as cleaning up instance mappings and reqspecs for archived instances
19:14:59 dansmith ah yeah
19:17:51 melwitt for the multi-cell affinity, I could also see it being, wait for stein and backport to rocky, queens, pike. whatever you all think is best
19:29:38 mriedem you already said that was the plan above yeah?
19:29:51 mriedem "yeah. given the timing and the latentness + multi-cell-only nature, wait until after rc1"
19:30:01 mriedem maybe you meant rc2?
19:30:13 mriedem tbc, we generally shouldn't assume/expect an rc2
19:30:18 mriedem otherwise we failed at rc1
19:30:19 dansmith I don't think we should do an rc2 for that
19:30:55 mriedem i remember loving explaining how we did rc's upstream to our downstream PMs at ibm since i tried to model the same release process
19:31:00 mriedem "ok so when do we do rc2 and rc3?"
19:31:09 mriedem "we don't, get your fixes into rc1 or don't cut rc1"
19:31:12 mriedem "BUT!!!"
19:31:25 dansmith well, we did have those couple of cycles where we always had an rc2 for translations on the things we got into rc1
19:31:44 mriedem right, but the PMs at ibm planned for multiple candidates because they knew we had a backlog of bugs
19:31:48 mriedem total waterfall
19:31:53 dansmith heh yeah, and,
19:32:01 dansmith why call it rc1 if there's no rc2 amirite/
19:32:10 mriedem that's deep
19:32:21 dansmith it's like an outline with one bullet.. no larger sin.
19:32:24 melwitt yeah, sorry, I meant rc2 vs wait until stein, that I was asking of you both
19:32:38 mriedem i'd say stein
19:32:43 dansmith for sure
19:32:47 melwitt k, cool
19:32:48 mriedem make sure there is no immediate regression,
19:32:49 mriedem then backports
19:32:53 sean-k-mooney mriedem: nova is pretty waterfall too in general that said the runways has made nova mor agile this cycle
19:33:09 mriedem hey man
19:33:10 melwitt so what about the placement perf thing? just try to land it for RC1 and not wait for RC2?
19:33:25 mriedem melwitt: my understanding on that is there are 2 fixes
19:33:41 mriedem cdent's is the more important of the two and is already approved
19:33:50 melwitt that's correct. one is approved, other is up for review
19:33:51 mriedem i.e. cdent's drops the perf by 50%
19:33:57 mriedem jay's drops that another 50%
19:34:14 sean-k-mooney dansmith: :) most of intel liked that word but followed water-scurm-fall developemnt instead
19:34:37 mriedem i'm pretty sure my first 10 years at ibm the dev model was really code-and-fix
19:34:56 mriedem the bestest of models
19:35:35 mriedem melwitt: i'll defer to efried and cdent and the placement boyz on how comfortable they are on jay's fix for rc2
19:35:39 sean-k-mooney mriedem: by drops the perf by 50% do you mean makes it better or worse
19:35:50 mriedem improves perf by 50%
19:35:58 sean-k-mooney :)
19:36:03 mriedem it's the dpdk of patches
19:36:27 sean-k-mooney mriedem: really fast and imposible to debug
19:36:49 mriedem heh that's a pretty good analogy for placement
19:37:05 melwitt mriedem: okay, sounds fair. efried and cdent, let me know what you think of it once you've reviewed
19:38:19 jaypipes melwitt: re: https://bugs.launchpad.net/nova/+bug/1746863, I thought we'd always said server groups were restricted to a single cell. is that not the case?
19:38:19 openstack Launchpad bug 1746863 in OpenStack Compute (nova) "scheduler affinity doesn't work with multiple cells" [High,In progress] - Assigned to melanie witt (melwitt)
19:38:31 cdent melwitt, mriedem : I'm happy to see them both go in
19:39:14 melwitt what is your assessment of the risk of the change somehow making the final release worse?
19:39:57 mriedem jaypipes: it's totally possible to have server group members wind up unintentionally in separate cells
19:40:04 mriedem even if in an affinity group
19:40:41 cdent melwitt: we talking about the placement thing on "final release worse"? the risk in those changes is very very low. the value is very very high.
19:41:33 melwitt jaypipes: they are, in a sense that affinity means same host (cells or not) and anti-affinity means different hosts. but the bug is that if you land on hostA for your first instance, because we don't look for members in all cells, we won't find that a group member is on hostA and therefore we need to co-locate instance2 to hostA for affinity if you want to add another host. I hope that makes sense
19:41:45 sean-k-mooney mriedem: jaypipes i guess mabe you should use dansmith's pre placement filter stuff to avoid that.
19:41:51 melwitt jaypipes: *if you want to add another instance
19:41:58 sean-k-mooney *could use
19:42:13 mriedem sean-k-mooney: how?
19:42:24 melwitt cdent: yeah, exactly. thanks for confirming it is low risk
19:42:24 mriedem placement doesn't know about server groups
19:42:32 mriedem nor cells
19:42:44 dansmith well, if we had a same-resource-provider thing we could kindof hack up a thing to do it via placement
19:43:02 dansmith but agree, it's not easy
19:43:43 dansmith it would be trivial to just fail a boot request for affinity if we can't talk to the cell where the other members are
19:43:53 sean-k-mooney mriedem: well the pre filter is for tenant affinity to a cell. i was thicnking if we had a request in a server group we could have a prefiltr that just picks a cell and only trys to place within that cell for the entire group
19:43:55 dansmith since we clearly can't honor the affinity goal
19:44:16 dansmith sean-k-mooney: that doesn't help us
19:44:23 dansmith sean-k-mooney: you might not be keeping tenants to cells
19:45:12 mriedem dansmith: yeah i think that's what gibi said on mel's patch
19:45:15 sean-k-mooney dansmith: i was not suggesting it was a depency just that if we detected there was an affinity group the only consider 1 cell for the request
19:45:20 dansmith mriedem: ack, haven't looked
19:45:37 sean-k-mooney dansmith: anyway it was just a tought.
19:45:49 mriedem we can still race our way around affinity and wind up in different cells
19:45:58 mriedem if you create the servers at the same time

Earlier   Later