| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-02-04 | |||
| 14:33:16 | ralonsoh | thanks | |
| 14:33:41 | sean-k-mooney | oh the nova net removal | |
| 14:33:46 | ralonsoh | yeah! | |
| 14:33:52 | ralonsoh | just a couple of things | |
| 14:34:21 | openstackgerrit | Mykola Yakovliev proposed openstack/nova master: Fix boot_roles in InstanceSystemMetadata https://review.opendev.org/698040 | |
| 14:34:31 | sean-k-mooney | ralonsoh: is this the issue https://review.opendev.org/#/c/697153/16/nova/api/openstack/compute/floating_ips.py@39 | |
| 14:34:52 | ralonsoh | sean-k-mooney, exactly | |
| 14:35:02 | sean-k-mooney | we shoudl be useing floating_ip.get('prot_detail') | |
| 14:35:10 | sean-k-mooney | to not cause a key error if it is not set | |
| 14:35:13 | ralonsoh | sean-k-mooney, and the way the "pool" key is retrieved | |
| 14:35:29 | ralonsoh | the network id is stored in "floating_network_id" | |
| 14:35:37 | ralonsoh | https://github.com/openstack/neutron/blob/master/neutron/db/l3_db.py#L1030-L1037 | |
| 14:35:38 | sean-k-mooney | ah i see | |
| 14:36:03 | sean-k-mooney | im glad we are so consistent | |
| 14:36:18 | ralonsoh | hahahahaha | |
| 14:36:21 | ralonsoh | sorry for that | |
| 14:36:34 | sean-k-mooney | no its fine so that shoudl be an easy fix | |
| 14:36:52 | efried | bauzas, sean-k-mooney: I really don't want the list. I think a boolean toggle is the right thing. As we move forward, for example moving VGPUs from under the root to under the NUMA RPs, things will still work correctly. | |
| 14:37:28 | sean-k-mooney | efried: ya im totally fine with the boolean toggel | |
| 14:37:37 | bauzas | efried: we don't really need a boolean for triggering if so | |
| 14:37:38 | efried | We'll "gain" support for new flavors that express affinity of the VGPUs, but old flavors that don't express such affinity will still work. | |
| 14:37:59 | efried | bauzas: we do; one of the important distinctions here is that we're segregating the data center into numa-aware and not. | |
| 14:38:00 | sean-k-mooney | i think by defualt we shoudl not do the reshape on upgrade but when the config option is set we will do the reshape and it should not be possibel to undo while instance are on the host | |
| 14:38:03 | bauzas | unless we wanna somehow prepare ops to switch when they want | |
| 14:38:32 | bauzas | cool enough, let's then make it a boolean | |
| 14:38:46 | sean-k-mooney | bauzas: we need the config option because we said we would partion the cloud into numa hosts and non numa hosts | |
| 14:38:50 | efried | I agree we should reshape when the config option is set and the service is restarted. We knew the restriction of "reshape only on upgrade" was artificial/temporary when we made it. | |
| 14:39:12 | sean-k-mooney | so that on the non numa hosts you could continue to run floating instance that use more resouces then fit in 1 numa node | |
| 14:39:20 | sean-k-mooney | without haveing to make the multi numa instnces | |
| 14:39:43 | bauzas | that looks seamless indeed | |
| 14:39:43 | stephenfin | ralonsoh: ah, so previously we were making two requests to neutron via neutronclient: one to retrieve all floating IPs and one to retrieve all ports | |
| 14:39:57 | bauzas | but my only concern is that this conf opt is only for Ussuri | |
| 14:40:00 | sean-k-mooney | e.g. the i want to use openstack to run one giant vm per host to run anohter orcastrator on top usecase we all hate | |
| 14:40:11 | bauzas | and then we would turn into it either way in Victoria | |
| 14:40:17 | sean-k-mooney | bauzas: i wont be | |
| 14:40:20 | dansmith | efried: sean-k-mooney bauzas but they will have to reshape eventually, right? | |
| 14:40:22 | stephenfin | ralonsoh: and then we simply matched 'port_id' for each entry in the former to the corresponding port from the latter. I tried to make that cleverer by using 'port_details', but you're saying that's an extension that I can't rely on | |
| 14:40:29 | sean-k-mooney | i will need to be kept as long as we support non numa instnaces | |
| 14:40:30 | bauzas | dansmith is expressing my concern | |
| 14:40:39 | stephenfin | ralonsoh: I can fix that now if you haven't already, but how can I check if that extension is available? | |
| 14:40:41 | efried | dansmith: I'm not sure we ever need to force them to make a host NUMA-aware if they don't want to. | |
| 14:40:56 | ralonsoh | stephenfin, exactly, this is an extension and this key is not mandatory | |
| 14:41:12 | bauzas | efried: if so, we somehow need to make the path clear that we *will* remove filter things in Victoria either way | |
| 14:41:14 | sean-k-mooney | dansmith: they will have to reshape only if we decied that all instance will have a numa toplogy of 1 gust numa node by defualt | |
| 14:41:18 | efried | iow the segregated-on-NUMA-ness is a permanent characteristic of the data center | |
| 14:41:58 | dansmith | if they can stay off forever, then fine, but I imagine that leaves us with a pretty big variable for a long time | |
| 14:42:05 | bauzas | also, | |
| 14:42:23 | bauzas | what I'm not okay is keeping the old filter processing for NUMA placement for a while | |
| 14:42:39 | dansmith | anyway, my point was that 95% of people will not pay attention and flip that flag until they have to, so reshape-on-flag only buys you N cycles until you turn it on permanently | |
| 14:42:48 | ralonsoh | stephenfin, you can use the neutron-client | |
| 14:42:49 | sean-k-mooney | dansmith: ya so i have been advocating that we should make all instnaces numa instance for a long time. but it breaks the "run one giant vm" usecase that some people care about | |
| 14:43:03 | ralonsoh | stephenfin, you can retrieve the extension list | |
| 14:43:29 | ralonsoh | but in this case, do you need it? just checking if this key is in the FIP dict | |
| 14:43:55 | efried | Hm, dansmith won't "people who care about NUMA" (like "edge") flip it right away so they can take advantage? Otherwise NUMA-ness won't work at all. | |
| 14:43:58 | sean-k-mooney | dansmith: if we decided that if you want to do that you have to specify multiple numa nodes in the flavor then we could get rid of the flag | |
| 14:43:58 | dansmith | sean-k-mooney: you mean just in how they specify the one big VM, not that we need to be able to provide numa-ignorant vms right? | |
| 14:44:06 | bauzas | efried: that's my point | |
| 14:44:25 | sean-k-mooney | dansmith: ya they would jsut add hw:numa_nodes=2 or whatever to the giant flavor | |
| 14:44:29 | efried | Like, I thought we were making this dividing line here such that, if you ask for a NUMA topo, and you don't have any computes with that flag switched on, you just won't land. | |
| 14:44:32 | stephenfin | ralonsoh: I need to figure out what instance the floating IP is associated with so I can return 'instance_id' in the API response | |
| 14:44:34 | bauzas | efried: if we want them to flip the conf opt, we somehow need to stop supporting the existing by the filter | |
| 14:44:58 | efried | bauzas: but we don't *want* them to flip the conf opt... unless they *need* NUMA on that host. | |
| 14:44:59 | stephenfin | ralonsoh: which appears to be stored in the 'device_id' field of the floating IP's port | |
| 14:45:06 | dansmith | efried: if it's required for numa at all then don't those people have to switch at upgrade anyway? | |
| 14:45:20 | efried | dansmith: only if they want NUMA-aware VMs on that host. | |
| 14:45:26 | bauzas | efried: correct | |
| 14:45:36 | bauzas | efried: people who don't care a bit about NUMA wouldn't get reshapes | |
| 14:45:54 | sean-k-mooney | dansmith: as it stand we should not be mixing numa vms and non numa vms on the same host | |
| 14:45:59 | dansmith | efried: right so people that have numa guests right now, they do the upgrade and if they don't flip it, they're unable to boot new instances? | |
| 14:46:07 | bauzas | but people who care should opt-in now so that in Victoria we could remove the filter bits responsible for placing instances | |
| 14:46:10 | sean-k-mooney | that is becasue of how we track and affitize the numa vms memory | |
| 14:46:14 | efried | ah, I see what you're saying. Yes dansmith that's right. | |
| 14:46:22 | dansmith | sean-k-mooney: I understand | |
| 14:46:38 | dansmith | efried: right, so anyone else will not flip that flag until they have to | |
| 14:46:46 | efried | why is that a problem? | |
| 14:47:03 | sean-k-mooney | dansmith: yes we could maybe check if the host has numa instnace currenly an for enable it? | |
| 14:47:05 | dansmith | efried: as I said before, if we're going to support numa-aware and numa-ignorant forever, then maybe it's not | |
| 14:47:23 | dansmith | sean-k-mooney: seems a little scary | |
| 14:47:30 | sean-k-mooney | ya i agree | |
| 14:47:34 | efried | Right, that's what I thought we were going to do. If FooHost doesn't ever care about NUMA instances, it can just leave that flag off forever, and never have a flavor with hw:numa* in it, and go on happily forever. | |
| 14:47:48 | ralonsoh | stephenfin, yes, but I'm reviewing the code and if the extension is not enabled, this information won't be there, in the FIP | |
| 14:47:53 | sean-k-mooney | i dont really like basing the behavior off what the current vms are using | |
| 14:48:03 | efried | UNTIL we can design & implement proper granular anti-affinity syntax in placement, which IMO is too ambitious for Ussuri. | |
| 14:48:03 | efried | gibi, bauzas: To close on group_policy: sean-k-mooney and I talked about it a bit yesterday and decided that, in order to support both NUMA and bandwidth, we need to retain the group_policy=none default and do the anti-affinity in the NTF | |
| 14:48:09 | efried | sean-k-mooney: +1 | |
| 14:48:19 | sean-k-mooney | efried: keep in mind that almost all host are numa hosts | |
| 14:48:32 | dansmith | sean-k-mooney: in reality, ALL of them are | |
| 14:48:38 | efried | meaning almost all hosts are *capable* of NUMA. | |
| 14:48:45 | efried | Not that all hosts are hosting NUMA-aware instances. | |
| 14:48:49 | sean-k-mooney | dansmith: yes | |
| 14:48:59 | dansmith | not capable, they are numa and if not configured, are losing performance to that fact | |
| 14:49:00 | sean-k-mooney | efried: correct | |
| 14:49:12 | efried | right, which some VMs don't care about. | |
| 14:49:13 | stephenfin | ralonsoh: WDYM? As I understood it, we won't have the 'port_details' field but the 'device_id' field will still be there in the ports API response, right? | |
| 14:49:41 | dansmith | efried: no, they all care about it, just some operators don't care to take on the burden of configuring it because we make it so difficult | |
| 14:49:41 | sean-k-mooney | efried: no i think dansmit ment at the host level. not the vm level | |
| 14:49:48 | ralonsoh | stephenfin, no no | |
| 14:50:00 | dansmith | nobody is opting out of numa for any reason other than the cost benefit isn't there for configuring it | |
| 14:50:02 | ralonsoh | stephenfin, if you don't have 'port_details', you won't have this info | |
| 14:50:14 | stephenfin | ohhhh, they're tied together | |