| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-03-14 | |||
| 16:36:47 | dansmith | it's opinionated | |
| 16:37:00 | stephenfin | bauzas: I think the RFC is wrong | |
| 16:37:01 | stephenfin | :) | |
| 16:37:38 | bauzas | it says 16 octets, period. | |
| 16:37:43 | dansmith | bauzas: note it allows removing the dashes :) | |
| 16:37:44 | dansmith | bauzas: because it thinks that's okay :) | |
| 16:37:49 | dansmith | which we could also have done | |
| 16:39:33 | bauzas | we could have done many things | |
| 16:39:34 | stephenfin | bauzas: It didn't matter much before anyway since we were undoing it. This will only start to bite us if/when o.vo decides to make that warning an error https://github.com/openstack/nova/blob/fd59fbd4d1914d2adf35a85435ba4aa433f082cd/nova/cmd/manage.py#L1171 | |
| 16:40:06 | bauzas | but I'm trying to see how we could better coerce in o.vo so that would make both not changing the DB, and make edleafe and stephenfin happy | |
| 16:40:19 | edleafe | bauzas: I really don't care | |
| 16:40:29 | dansmith | stephenfin: that would (a) be a change to our (and others') RPC APIs, but also (b) we could easily handle this in _from_db_obj(), or by doing the migration process with the low-level routines instead of objects | |
| 16:40:33 | edleafe | I was just trying to fix a potential issue | |
| 16:40:37 | bauzas | \o/ | |
| 16:40:55 | bauzas | I'm litterally 20 mins away from a long holiday period | |
| 16:41:07 | edleafe | but spaces in a UUID are not valid. Removing the dashes is fine | |
| 16:41:07 | bauzas | would those 20 mins well spent in fixing that then ? | |
| 16:41:26 | bauzas | c'on | |
| 16:41:32 | bauzas | it's a *string* | |
| 16:41:34 | dansmith | a UUID is a number | |
| 16:41:42 | dansmith | the string representation of it can be many things | |
| 16:41:49 | dansmith | microsoft encloses them in {} to make them stand out | |
| 16:41:55 | bauzas | how the string translates to an hex is something fine even with spaces | |
| 16:42:07 | bauzas | yeah | |
| 16:42:48 | edleafe | hey, I wanted to store all uuids are 128-bit integers, but got out-voted by the human-readable people | |
| 16:42:56 | edleafe | s/are/as | |
| 16:43:17 | bauzas | edleafe: it's comprehensive | |
| 16:43:32 | bauzas | edleafe: but the o.vo coercing method shouldn't care at all about the formatting | |
| 16:44:04 | bauzas | it should just assume Good Faith (c) | |
| 16:46:45 | cfriesen | on a totally different topic...does nova wait to ensure vifs are actually plugged when doing a live migration? I see it calling self.virtapi.wait_for_instance_event() on instance spawn and cold migration, but not for live. I assume the flow is somewhat different? | |
| 16:47:22 | stephenfin | bauzas: Before you go, fancy pushing these two patches through? https://review.openstack.org/#/c/385071 | |
| 16:47:39 | stephenfin | Additional shuffling things around/adding docstring patches | |
| 16:47:48 | bauzas | my review stats are poor this week | |
| 16:47:57 | bauzas | thanks to the NUMA spec | |
| 16:48:18 | bauzas | (and the f*** tax-credit document I had to write) | |
| 16:48:38 | stephenfin | bauzas: No better time, in that case | |
| 16:48:53 | bauzas | https://docs.python.org/2/library/uuid.html#uuid.UUID "When a string of hex digits is given, curly braces, hyphens, and a URN prefix are all optional.' | |
| 16:49:13 | bauzas | so, yeah, really the space carries a lot more, but meh | |
| 16:49:26 | dansmith | tssurya: ah, yeah, kudos for spotting it in the initial patch :) | |
| 16:50:17 | bauzas | dansmith: planning to do a Cells meeting in 10 mins ? | |
| 16:50:21 | dansmith | yar | |
| 16:50:29 | bauzas | cool, I can attend it for once \o/ | |
| 16:50:41 | dansmith | bauzas: I thought you were leaving in 20 minutes? :) | |
| 16:51:09 | bauzas | for the first time since 6 months I can attend a cells meeting, I won't skip it | |
| 16:51:16 | dansmith | heh okay | |
| 16:51:32 | bauzas | if I need to leave, I'll | |
| 16:51:54 | dansmith | bauzas: you will leave only if I dismiss you! | |
| 16:52:17 | bauzas | fine, I'll tell my children I can't go to Disneyland because of you | |
| 16:52:28 | bauzas | they'll understaznd | |
| 16:52:40 | dansmith | dansmith -- keeping children from happiness since 1981 | |
| 16:52:54 | bauzas | heh | |
| 16:55:20 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: Snapshot https://review.openstack.org/543023 | |
| 16:59:08 | openstackgerrit | Dan Smith proposed openstack/nova master: Add --by-service to discover_hosts https://review.openstack.org/552691 | |
| 17:01:02 | dansmith | melwitt: mriedem tssurya bauzas cells meeting? | |
| 17:03:18 | mriedem | melwitt: comments in your etherpad | |
| 17:06:37 | openstackgerrit | Chris Dent proposed openstack/nova master: DNM: Demo code for microversion parse extraction https://review.openstack.org/550265 | |
| 17:06:43 | arvindn05 | mriedem: i've added comments to the review | |
| 17:07:29 | cdent | jaypipes: if you like this https://review.openstack.org/#/c/159382/ (multi workers for scheduler) can you kick it in. Seems like one of those things we'd like to exercise for as long as possible | |
| 17:07:30 | arvindn05 | mriedem: i think the only change would be to remove the workitem for ImageExtraSpecsFilter | |
| 17:08:02 | stephenfin | mriedem: Gerrit keeps scrolling up when I try to reply to comments. Did you have a workaround for that? | |
| 17:08:34 | mriedem | stephenfin: i've noticed that today again too - the fix used to be to change the render setting to 'slow' | |
| 17:08:39 | mriedem | arvindn05: in the cells meeting atm | |
| 17:12:28 | jaypipes | cdent: done | |
| 17:12:36 | cdent | thanks | |
| 17:17:08 | cfriesen | cdent: jaypipes: I'm nervous about defaulting to multiple workers...seems like a good way to hit races even with placement | |
| 17:17:46 | cdent | cfriesen: a) how?, b) that's why we're merging early | |
| 17:19:01 | openstackgerrit | Stephen Finucane proposed openstack/nova-specs master: Add 'numa-aware-vswitches' spec https://review.openstack.org/541290 | |
| 17:19:03 | cfriesen | cdent: numa resources aren't allocated until you actually hit the compute node, so can fail late. Also, do we model server group policies in placement currently? | |
| 17:19:56 | cfriesen | cdent: in order to fix server group policy stuff we had to serialize scheduling of instances in the same server group, otherwise we hit races even with just a single scheduler worker. | |
| 17:21:10 | jaypipes | cfriesen: no, we do not model server group policies in placement. | |
| 17:21:56 | cfriesen | cdent: jaypipes: the problem with server groups is that the group membership isn't updated until the instance actually hits the node, so there's a big window from when the scheduler made the decision until the membership changes | |
| 17:22:02 | cfriesen | (in the DB) | |
| 17:22:35 | cfriesen | sorry, not group membership but the list of compute nodes being used by the group members | |
| 17:22:53 | jaypipes | cfriesen: yes, that's group membership. | |
| 17:23:33 | jaypipes | cfriesen: so because of a crappy affinity implementation and crappy numa resource tracking, we'll continue to slow down the rest of the world... | |
| 17:24:31 | jaypipes | cfriesen: can't we tell folks using that functionality to run a single worker? | |
| 17:24:46 | sean-k-mooney | cfriesen: defulting to muliple works has other issue. like we used to hit the db max conncetion limits because several service in openstack defualted to use multiple works for things and defulted to 1 worker per cpu. | |
| 17:24:47 | cfriesen | jaypipes: just pointing out that this could cause unexpected races if it's enabled by default. We might want to put something in the release notes. | |
| 17:25:05 | cdent | also, if the aforementioned serialization is already present, will it help? | |
| 17:25:28 | jaypipes | sean-k-mooney: that's unrelated. | |
| 17:25:40 | cfriesen | cdent: I don't think that serialization is upstream yet. How would we serialize across multiple workers? | |
| 17:25:51 | sean-k-mooney | jaypipes: proably i am just skimming the scoll back now | |
| 17:26:00 | cdent | cfriesen: ah, sorry, I had misunderstood which "we" you meant :) | |
| 17:26:07 | sean-k-mooney | i was in meeting all day up until this point | |
| 17:27:02 | jaypipes | cfriesen: a single nova boot request is only handled by a single worker, no? the case you're worrying about is multiple clients calling nova boot with the same server group. | |
| 17:27:44 | jaypipes | cfriesen: as for the NUMA issues, I don't think that changing to multiple workers will make a difference in the number of retries that happen. | |
| 17:27:59 | sean-k-mooney | jaypipes: or perhaps the nova multi boot support we you say boot x instance of y flavour on network z ? | |
| 17:28:22 | jaypipes | sean-k-mooney: I don't understand how that's relevant? | |
| 17:28:39 | jaypipes | sean-k-mooney: that would be handled in a single thread. | |
| 17:29:04 | jaypipes | sean-k-mooney: the only situation cfriesen is worried about is when multiple nova boot requests involving the same server group were executed simultaneously. | |
| 17:29:12 | sean-k-mooney | jaypipes: i was wondering if that was a single boot request form the api point of view or x independet ones and a client feature | |
| 17:29:35 | jaypipes | sean-k-mooney: a single nova boot is a single thread of execution. | |
| 17:30:47 | sean-k-mooney | jaypipes: ya that makes sense. and nova boot --min 3 --max 3 --server-group... is considered a singel boot request | |
| 17:31:32 | sean-k-mooney | jaypipes: so the only race would be if two client tried to boot servers in the same group concureently without the server group already having running instances | |
| 17:32:17 | jaypipes | sean-k-mooney: yes, that is considered a single boot request. | |
| 17:33:51 | jaypipes | sean-k-mooney: almost. the only race is two clients concurrently attempting to add instances to the same server group (regardless of whether the server group has members). and that is a situation I find pretty rare and not worth shooting the rest of the world in the foot for. | |
| 17:35:26 | cfriesen | jaypipes: the scenario I'm worried about is where two instances race to schedule and get put onto the same compute node, but then one of them claims a mix of resources that causes the other to fail it's resource allocation. With a single sched worker this is less likely (though still possible). | |
| 17:35:57 | jaypipes | cfriesen: unless those are NUMA resources, it's not possible to do that. | |
| 17:36:01 | mriedem | arvindn05: ok what's up | |
| 17:36:21 | cfriesen | jaypipes: yes, numa resources. CPUs, hugepages, PCI devices with strict affinity | |