| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-10-09 | |||
| 05:28:14 | openstackgerrit | Naichuan Sun proposed openstack/nova master: xenapi(N-R-P): support compute node resource provider update https://review.openstack.org/521041 | |
| 05:28:27 | openstackgerrit | Naichuan Sun proposed openstack/nova master: os-xenapi(n-rp): add traits for vgpu n-rp https://review.openstack.org/604269 | |
| 06:02:11 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: api-ref: Replace non UUID string with UUID https://review.openstack.org/608854 | |
| 06:28:12 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Transform volume.usage notification https://review.openstack.org/580345 | |
| 06:29:08 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Transform compute_task notifications https://review.openstack.org/482629 | |
| 06:30:56 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova stable/rocky: Imported Translations from Zanata https://review.openstack.org/604260 | |
| 07:29:06 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in unit/network/test_neutronv2.py (21) https://review.openstack.org/576709 | |
| 07:29:26 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in unit/network/test_neutronv2.py (22) https://review.openstack.org/576712 | |
| 07:36:52 | bauzas | good morning nova | |
| 07:57:08 | gibi | good morning | |
| 07:58:04 | gibi | bauzas, mriedem, jaypipes, efried: I saw the discussion about force live migrate on the scheduler meeting. I think I will post a mail with a summary of the different pieces and possibilities on the ML | |
| 07:58:13 | bauzas | ok cool | |
| 07:59:45 | openstackgerrit | Rodolfo Alonso Hernandez proposed openstack/os-vif master: Remove IPTools deprecated implementation https://review.openstack.org/605422 | |
| 08:02:50 | openstackgerrit | Jan Gutter proposed openstack/os-vif master: Add support for generic representors https://review.openstack.org/608693 | |
| 08:54:44 | mnaser | fwiw | |
| 08:54:46 | mnaser | http://logs.openstack.org/15/608315/4/check/gpu-test/7243464/job-output.txt.gz | |
| 08:54:50 | mnaser | https://review.openstack.org/#/c/608315/ | |
| 08:54:57 | mnaser | you should be able to do tests with access to a k80 gpu | |
| 08:56:02 | mnaser | with nested virt too | |
| 08:56:15 | mnaser | cc sean-k-mooney ^ | |
| 09:24:50 | openstackgerrit | Zhenyu Zheng proposed openstack/nova-specs master: Detach and attach boot volumes - Stein https://review.openstack.org/600628 | |
| 09:25:18 | openstackgerrit | Brin Zhang proposed openstack/nova master: Add compute API version for when a ``volume_type`` is requested https://review.openstack.org/605573 | |
| 09:36:30 | jaypipes | gibi: cool with me. thank you! | |
| 09:42:17 | gibi | bauzas, mriedem, jaypipes, efried: I've posted the mail. Sorry it turned out as a long one. http://lists.openstack.org/pipermail/openstack-dev/2018-October/135551.html | |
| 09:43:06 | jaypipes | gibi: :) no worries, it's a complicated subject. | |
| 09:46:12 | sean-k-mooney | mnaser: oh cool. bauzas should be pleased. thanks :) now we just need to figure out how to use devstack to deploy with vgpus | |
| 09:46:37 | sean-k-mooney | bauzas: do you have a local.conf/devstack plugin that automates teh setup | |
| 09:47:36 | mnaser | yeah feel free to hack away at it | |
| 09:49:22 | sean-k-mooney | mnaser: actully looking at https://docs.nvidia.com/grid/gpus-supported-by-vgpu.html nvidia may have locked out the k80s... we can still try them | |
| 09:50:34 | mnaser | #justnvidiathings | |
| 09:51:41 | sean-k-mooney | mnaser: you know im really looking forward to the point were someone implement the vgpu support in the nouveau | |
| 09:51:57 | sean-k-mooney | driver so there are not nvida locks on the hardware | |
| 09:52:27 | sean-k-mooney | that said i personally use the nvidia binary driver becase performance | |
| 09:54:10 | mnaser | i'd say this is where things go beyond my knowledge | |
| 09:56:17 | sean-k-mooney | the nouveau is the opensource linux driver for nvidia gpus. the offical binary blob give you about 10-15% better performace in games and i thin its require to use some of the nvida only techs like hair works. for vgpus instead of the normal driver you run there grid driver which is desinged for there datachenter gpus | |
| 09:57:17 | sean-k-mooney | that said im 98% certin that if you could remove the sku check that it would likely work on there desktop and workstation gpus but nvidia want to charge for the privlage of useing virtualisation with there gpus | |
| 10:02:38 | openstackgerrit | Tetsuro Nakamura proposed openstack/nova stable/rocky: Add alloc cands test with nested and aggregates https://review.openstack.org/607454 | |
| 10:02:38 | openstackgerrit | Tetsuro Nakamura proposed openstack/nova stable/rocky: Fix aggregate members in nested alloc candidates https://review.openstack.org/608903 | |
| 10:28:37 | openstackgerrit | Merged openstack/nova master: conf: Gather 'live_migration_scheme', 'live_migration_inbound_addr' https://review.openstack.org/456572 | |
| 10:29:58 | nehaalhat_ | Hi, can any one help me to merge this patch: https://review.openstack.org/#/c/581218/ | |
| 11:32:07 | openstackgerrit | Brin Zhang proposed openstack/nova master: Add microversion 2.67 to support volume_type https://review.openstack.org/606398 | |
| 11:46:22 | openstackgerrit | Merged openstack/nova master: Move test.nested to utils.nested_contexts https://review.openstack.org/608416 | |
| 12:41:19 | pooja_jadhav | Hi team, anyone help me in the https://github.com/openstack/nova/blob/85b36cd2f82ccd740057c1bee08fc722209604ab/nova/tests/functional/api_sample_tests/test_simple_tenant_usage.py#L85-L93.. When we run test name "test_get_tenants_usage" and passed instance_uuid_1 in the query. If we check the instances in simple tenant usages controller. we can see instance-2 object data in the instances list. If we pass instance-2 in query then in the | |
| 12:41:20 | pooja_jadhav | instances we can see instance-3. So how actually its working? | |
| 12:45:32 | openstackgerrit | Lucian Petrut proposed openstack/nova master: Fix os-simple-tenant-usage result order https://review.openstack.org/608685 | |
| 12:49:29 | bauzas | mnaser: sean-k-mooney: sorry was afk | |
| 12:50:13 | bauzas | mnaser: thanks for the proposal, but unfortunately, AFAIK, k80 devices aren't supported by nvidia for vGPUs | |
| 12:51:28 | bauzas | gibi: saw your thread, I need proper time to read it and reply to it | |
| 12:55:37 | gibi | bauzas: sure it needs time. The reason for the mail was to summarize the problem as it is pretty hard for solve it consistently without seeing every corner of it | |
| 13:02:09 | stephenfin | dansmith: What's the by-service approach you refer to here? https://review.openstack.org/#/c/608703/ (link to a spec/commit is good) | |
| 13:32:19 | dansmith | stephenfin: did you see the link to the bug? | |
| 13:32:45 | dansmith | stephenfin: regardless, as I said, I think mentioning the discover step (whether by service or regular) in the devstack docs is the right thing to do | |
| 13:33:18 | stephenfin | Yup, that's the plan. Didn't click the bug link though. Will do now | |
| 13:33:29 | stephenfin | Well, soon as I'm done with downstream fun | |
| 13:47:06 | efried | gibi: Can you help me understand some basics of selecting a destination host during evac/migrate? | |
| 13:51:03 | efried | or dansmith bauzas | |
| 13:51:15 | efried | You can specify a host without the force flag and we'll run GET /a_c, right? | |
| 13:51:18 | dansmith | efried: I'm not sure what you're asking.. it's not really any different | |
| 13:51:33 | dansmith | yeah, IIRC | |
| 13:51:42 | dansmith | only the force flag makes us totally skip I think | |
| 13:51:48 | efried | And then what happens if the GET /a_c returns no candidates for the requested host? | |
| 13:51:58 | efried | Do we fail or do we select a different host? | |
| 13:53:31 | dansmith | we should fail | |
| 13:53:42 | dansmith | lemme find a thread to pull | |
| 13:55:45 | dansmith | https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L983-L995 | |
| 13:56:12 | dansmith | schedule with a single host in the destination field.. if we get back novalidhost, we error the migration | |
| 13:58:08 | efried | So what use was the force flag ever supposed to be? Literally an optimization to avoid running some code, but behaviorally/algorithmically no difference? | |
| 14:00:18 | efried | Because in theory since we're using placement now, which is fast, that optimization buys us almost nothing. So if ^ is true, the force flag is basically obsolete anyway. | |
| 14:02:08 | mriedem | i can field this one... | |
| 14:02:21 | mriedem | efried: before the force flag, specifying a host at all bypassed the scheduler, | |
| 14:02:49 | mriedem | then a microversion was added which made passing a host go through the scheduler for validation, but apparently at least one person though we should preserve the ability to bypass the scheduler, so the force flag was added to do that backdoor | |
| 14:03:11 | efried | What was the motivation to "bypass the scheduler"? Because it was inefficient? | |
| 14:03:53 | mriedem | idk, i'm assuming to oversubscribe a host just to move things around | |
| 14:03:58 | mriedem | at least temporarily | |
| 14:04:10 | efried | is oversubscribe even possible at this point? | |
| 14:04:23 | efried | (Outside of allocation ratio, which doesn't count) | |
| 14:04:37 | mriedem | i don't think so, at least not for vcpu/ram/dis | |
| 14:04:38 | mriedem | *disk | |
| 14:04:52 | mriedem | as i noted in gibi's patch - we broke that in pike when force still goes through claiming resource allocations in conductor | |
| 14:04:54 | sean-k-mooney | efried: if it bypassed the scduer it proably bypassed the placement claim too but i have never check that | |
| 14:05:05 | mriedem | sean-k-mooney: incorrect | |
| 14:05:28 | efried | Yeah, that's the point. Since we started claiming from placement, you can't oversubscribe. | |
| 14:05:35 | mriedem | efried: in the old days, before placement, if you bypass the scheduler, conductor would not send limits down to the RT so it wouldn't fail the limits check on the resource claim | |
| 14:05:46 | sean-k-mooney | mriedem: was that unintentional however as you jsut said we "broke" that in pike | |
| 14:05:59 | efried | has anyone screamed about that breakage? | |
| 14:06:13 | efried | Or do we still not have enough serious operators on pike yet? :P | |
| 14:06:17 | mriedem | there are like 3 people i now on >= pike but no... | |
| 14:06:47 | mriedem | *know | |
| 14:06:51 | edleafe | efried: one reason for the force option was that admins wanted to be able to say "I know what I'm doing, dammit!". It wasn't about code efficiency | |
| 14:07:17 | efried | edleafe: But IIUC, the destination host is observed regardless. | |
| 14:07:32 | mriedem | observed? | |
| 14:07:41 | efried | meaning you either get on the suggested host or you die | |
| 14:07:50 | mriedem | correct | |
| 14:07:51 | efried | I'm not sure how we would "fix" the oversubscribe thing at this point, without adding a placement feature to allow it. | |
| 14:08:22 | mriedem | i don't think adding features to support shooting yourself is something we want to do at this point | |
| 14:08:22 | efried | which I doubt we want to do | |
| 14:08:40 | efried | "Hey placement, you know that one thing you're supposed to be designed to do? Yeah, don't do that." | |
| 14:08:42 | edleafe | efried: it isn't just oversubscription that would cause the scheduler to reject it. Things like affinity, etc., would also cause it to fail | |
| 14:08:52 | mriedem | note that even with forced live migration, conductor still runs some checks that could cause us to reject the host | |
| 14:09:05 | mriedem | edleafe: nope | |