| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-07-19 | |||
| 18:31:10 | efried | Oh, yeah, I've seen him any number of times. Just didn't realize he had a name. | |
| 18:31:16 | mriedem | "that dns guy" | |
| 18:31:48 | efried | Been trying to get the mugsie/penick cage match going. | |
| 18:32:35 | mriedem | battle royale | |
| 18:43:34 | efried | Remind me what are the rules for what releases a conductor and its computes can be at? | |
| 18:43:52 | mriedem | computes n-1 | |
| 18:44:44 | mriedem | the control plane can tolerate n-1 computes for rolling upgrades | |
| 18:45:04 | efried | and GET /allocation_candidates happens on... the controller? | |
| 18:45:09 | mriedem | yeah, scheduler | |
| 18:45:24 | mriedem | controller = api, conductor, scheduler | |
| 18:45:50 | efried | Trying to figure out why this allocation_request_version thing exists | |
| 18:46:40 | efried | Sounds like its purpose is to allow us to "cheat" and use a higher microversion on the n-1 compute than that codebase actually knows about. | |
| 18:47:02 | mriedem | correct | |
| 18:47:05 | mriedem | well, not the compute, | |
| 18:47:07 | mriedem | the cell conductor | |
| 18:47:24 | efried | um | |
| 18:47:29 | efried | because the cell conductor can be n-1 as well? | |
| 18:48:02 | mriedem | api->superconductor->scheduler (get allocation candidates and get alternate hosts from those allocatoin requests at the given microversion used in the scheduler)->pass down to compute to build; fails and reschedule to cell conductor which uses the allocation request, created in the scheduler, to claim resources against the alternate host in placement | |
| 18:48:41 | efried | And the cell conductor can be downlevel from the superconductor? | |
| 18:48:44 | mriedem | long-term i think we want to be able to do rolling upgrades of the cells themselves so yes cell conductor could be n- | |
| 18:48:45 | mriedem | *n-1 | |
| 18:49:01 | mriedem | it's been a long time since i've talked about rolling upgrades for cells with dansmith | |
| 18:49:41 | mriedem | gibi: good news - the delattr stuff for getting alternate hosts was a bugaboo so i'm removing all that code and added a wrinkle to the functional test to assert there are no alternate hosts when this bug is fixed | |
| 18:50:31 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add regression test for bug 1781710 https://review.openstack.org/583339 | |
| 18:50:33 | openstack | bug 1781710 in OpenStack Compute (nova) "ServersOnMultiNodesTest.test_create_server_with_scheduler_hint_group_anti_affinity failing with "Servers are on the same host"" [High,Fix released] https://launchpad.net/bugs/1781710 - Assigned to Matt Riedemann (mriedem) | |
| 18:50:34 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Update RequestSpec.instance_uuid during scheduling https://review.openstack.org/583347 | |
| 18:50:36 | mriedem | melwitt: jaypipes: ^ should be better now | |
| 18:51:12 | mriedem | tl;dr the change is now just make sure and set the in-context correct instance_uuid on the RequestSpec before calling the filters | |
| 18:54:15 | efried | Because of da new rulez, if the conductor is running N, the version of placement that all of its cell conductors and computes will be dealing with is at least the minumum for N, not the minimum for N-1, right? | |
| 18:54:56 | efried | guess the rules aren't new | |
| 18:57:02 | mriedem | new rules being we don't do version negotiation for placement? | |
| 18:57:24 | mriedem | that doesn't apply if the code is on an N-1 service and it thinks the minimum required version is whatever it was for N-1 | |
| 18:57:57 | mriedem | i mean, that's kind of the point of why we send the allocation candidate request version down to the cell, | |
| 18:58:19 | mriedem | because if it's downlevel, and the request body format changed, we need to tell the cell exactly what version to make that request with that particular body | |
| 18:58:32 | mriedem | it's like gd time travel | |
| 18:58:53 | mriedem | efried: do you want to continue talking about this or want me to review https://review.openstack.org/#/c/556669/ ? | |
| 18:58:58 | mriedem | because i need to get in the zone | |
| 18:59:44 | efried | mriedem: zone away. I'm going to write a patch to clean this shit up, based on now being able to assume we're talking to a queens minimum. | |
| 19:08:19 | jaypipes | mriedem: cool, will look shortly. | |
| 19:54:17 | openstackgerrit | Chris Dent proposed openstack/nova master: WIP POC: Use os-resource-classes in placement https://review.openstack.org/584084 | |
| 19:54:18 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] Move resource_class_cache into placement hierarchy https://review.openstack.org/584085 | |
| 19:54:19 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] ensure_rc_cache only at start of process https://review.openstack.org/584086 | |
| 20:20:38 | openstackgerrit | Merged openstack/nova master: Implement migrate_instance_start method for neutron https://review.openstack.org/556334 | |
| 20:24:57 | mriedem | kashyap: what does this mean? " libvirtError: unsupported configuration: Attribute mode is only allowed for guest CPU" | |
| 20:26:01 | mriedem | https://www.redhat.com/archives/libvir-list/2012-January/msg00232.html | |
| 20:26:57 | openstackgerrit | Artom Lifshitz proposed openstack/nova master: DNM: extra logging for 1775947 https://review.openstack.org/584032 | |
| 20:43:56 | cfriesen_ | mriedem: looks like its if you try to parse XML with a <host><cpu> section that then tries to specify a "mode". That's only allowed for XML describing guest cpus. | |
| 20:44:54 | cfriesen_ | https://github.com/libvirt/libvirt/blob/master/src/conf/cpu_conf.c#L317 | |
| 20:51:25 | melwitt | nova meeting in 9 minutes | |
| 20:55:01 | openstackgerrit | Eric Fried proposed openstack/nova master: Check provider generation and retry on conflict https://review.openstack.org/556669 | |
| 21:31:40 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in test_compute_api.py (4) https://review.openstack.org/568462 | |
| 21:31:58 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in unit/network/test_neutronv2.py (3) https://review.openstack.org/574104 | |
| 21:32:17 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in unit/network/test_neutronv2.py (4) https://review.openstack.org/574106 | |
| 21:46:25 | mriedem | gonna rebase and address comments in the handling a down cell series | |
| 21:49:41 | melwitt | tonyb: https://review.openstack.org/#/c/560317/ is failing powerkvm CI and we need help figuring out why | |
| 21:49:54 | tonyb | melwitt: okay I'll look at it | |
| 21:50:20 | melwitt | the powerkvm CI owner is on PTO this week, we found out that's why we didn't get response to pings | |
| 22:05:04 | openstackgerrit | Takashi NATSUME proposed openstack/nova master: Remove mox in virt/test_block_device.py https://review.openstack.org/566153 | |
| 22:07:24 | melwitt | mriedem: os-traits release proposed https://review.openstack.org/584130 | |
| 22:07:28 | melwitt | fyi | |
| 22:08:51 | openstackgerrit | Takashi NATSUME proposed openstack/nova-specs master: Create specs directory for Stein https://review.openstack.org/573602 | |
| 22:12:29 | openstackgerrit | do3meli proposed openstack/nova master: docs: add nova host-evacuate command to evacuate documentation https://review.openstack.org/578040 | |
| 22:15:00 | openstackgerrit | do3meli proposed openstack/nova master: docs: add nova host-evacuate command to evacuate documentation https://review.openstack.org/578040 | |
| 22:15:54 | mriedem | tonyb: well, it was originally failing because the CI is configured with cpu_mode='none' and so i believe the model put into the xml was ppc64le, so that was changed to mode=host-model and model=power8 (per someone from libvirt/qemu that knows about this), | |
| 22:16:02 | mriedem | but now it fails on something else, which cfriesen_ said might be: | |
| 22:16:03 | mriedem | (3:45:04 PM) cfriesen_: https://github.com/libvirt/libvirt/blob/master/src/conf/cpu_conf.c#L317 | |
| 22:16:03 | mriedem | (3:44:06 PM) cfriesen_: mriedem: looks like its if you try to parse XML with a <host><cpu> section that then tries to specify a "mode". That's only allowed for XML describing guest cpus. | |
| 22:16:41 | mriedem | tonyb: i was suggesting the powerkvm ci could just set cpu_mode=host-model and cpu_model=power8, but that's basically what the code is now doing as a workaround | |
| 22:26:40 | cfriesen_ | mriedem: if you want to specify cpu_model=power8, wouldn't you want a cpu_mode of custom? | |
| 22:28:17 | mriedem | if a tree falls in the woods | |
| 22:28:55 | mriedem | cfriesen_: yeah that makes more sense | |
| 22:29:14 | cfriesen_ | mriedem: hmm...https://bugzilla.redhat.com/show_bug.cgi?id=1237025 seems to think mode of "host-model" and "model" of "power8" is valid. weird. | |
| 22:29:15 | openstack | bugzilla.redhat.com bug 1237025 in libvirt "Guest can not start with different combinations of <cpu> mode and <model>" [Medium,Closed: notabug] - Assigned to abologna | |
| 22:29:33 | cfriesen_ | wonder if this is a bizarre powerpc-ism | |
| 22:30:12 | mriedem | jlk: said s390x causes all the problems, | |
| 22:30:15 | mriedem | but power is right up there | |
| 22:32:29 | cfriesen_ | mriedem: https://libvirt.org/formatdomain.html#elementsCPU under the host-model section documents PowerPC weirdness. | |
| 22:33:09 | mriedem | "Specifying CPU model is not supported either" | |
| 22:34:01 | mriedem | Since 1.2.11 PowerISA allows processors to run VMs in binary compatibility mode supporting an older version of ISA. Libvirt on PowerPC architecture uses the host-model to signify a guest mode CPU running in binary compatibility mode | |
| 22:34:10 | cfriesen_ | yeah, that's the interesting bit | |
| 22:34:24 | mriedem | ii libvirt-bin 4.0.0-1ubuntu8.3~cloud0 | |
| 22:34:25 | jlk | of course Power would override that to means omething else | |
| 22:36:08 | cfriesen_ | I don't get why they wouldn't just use a mode of "custom" with a model of "power8" (or whatever) instead. | |
| 22:36:39 | mriedem | could try that | |
| 22:36:48 | mriedem | kind of throwing things at the wall at this point until something works | |
| 22:43:47 | openstackgerrit | Eric Fried proposed openstack/nova master: Check provider generation and retry on conflict https://review.openstack.org/556669 | |
| 22:52:41 | tonyb | mriedem: Yeah setting power8 is wrong long term but shoudl be fine for now (but I wouldn't merge it that way) | |
| 22:53:36 | tonyb | mriedem: How do I debug placement failures? or at least see what allocation_candidates returned | |
| 22:54:34 | tonyb | Oh nm /me finds the 'renew until 27/07/18 08:05:33 | |
| 22:54:40 | tonyb | ' in n-cpu | |
| 22:54:56 | mriedem | so you don't need help debugging placement failures now? | |
| 22:55:20 | tonyb | mriedem: No I don't think so for this thing but in gernal it'd be good to know how to do it | |
| 22:55:22 | mriedem | n-sch logs will tell you for a given request which hosts are being filtered | |
| 22:55:45 | mriedem | i don't think we log the allocation_candidates response since it could be huge | |
| 22:57:28 | tonyb | mriedem: Okay what I'm seeing 'Got no allocation candidates from the Placement API' so I think that measn I the filters don't run which is why I'm not seeing the output I'm expecting in select_destindations? | |
| 22:58:13 | openstackgerrit | Merged openstack/python-novaclient master: Fix inconsistency https://review.openstack.org/572770 | |
| 22:59:03 | mriedem | correct | |
| 22:59:17 | mriedem | tonyb: what's the scenario? normal server create? or a rebuild, or force_hosts? | |
| 22:59:36 | mriedem | could just be there are no hosts with available capacity for the flavor being used | |
| 23:01:52 | tonyb | mriedem: Lots of scenarios: tempest.api.compute.admin.test_auto_allocate_network.AutoAllocateNetworkTest is the one I picked at random | |