| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-04-29 | |||
| 14:16:39 | artom | True | |
| 14:16:49 | belmoreira | our typical use case: | |
| 14:16:52 | belmoreira | select count(*) from compute_nodes nodes; #176 | |
| 14:16:52 | belmoreira | select count(*) from instance_extra where deleted = 0; #711 | |
| 14:16:53 | belmoreira | select * from instance_extra where deleted = 0 and numa_topology is not null and numa_topology like "%nova_object.name%"; #took 38.2 ms | |
| 14:17:10 | sean-k-mooney | ok so 700 ish is not bad | |
| 14:17:27 | sean-k-mooney | its not 2000 but its not 10 | |
| 14:18:07 | artom | So based on those number it might actually be OK to go ahead as is... | |
| 14:18:11 | sean-k-mooney | if we were to assume this was liniar then i think it would be accpetable | |
| 14:18:13 | artom | *numbers | |
| 14:19:09 | dansmith | so, if we just convert to migrate-on-load, we don't take any overhead for string searching at all, migrate only the instances that need it, and are assured that they have migrated all instances by the end of the cycle, right? | |
| 14:19:22 | dansmith | even with nova-manage, we can't make them run it without a blocker migration (which is also expensive) | |
| 14:19:48 | belmoreira | in this case nova_topology is not defined, but may help in the analysis: | |
| 14:19:55 | belmoreira | select count(*) from compute_nodes nodes; #88 | |
| 14:19:55 | belmoreira | select count(*) from instance_extra where deleted = 0; #1573 | |
| 14:19:57 | belmoreira | select * from instance_extra where deleted = 0 and numa_topology is not null and numa_topology like "%nova_object.name%"; #took 106 ms | |
| 14:19:58 | dansmith | if we did that, stephenfin could just remove his nova-manage bits entirely, no impact for anyone that doesn't have old instances | |
| 14:20:35 | stephenfin | Yeah, I can do that too. Let me see if I can wrangle something up | |
| 14:20:39 | artom | dansmith, the end of cycle stuff is complicated by FFUs and operators skipping releases, no? | |
| 14:20:47 | sean-k-mooney | ya if we migrate on load that also works but the only realy implciation is keeping that for a relase or two for FFU | |
| 14:21:00 | dansmith | artom: ah, yeah, good point... | |
| 14:21:34 | artom | So 2 or 3 cycles, I guess | |
| 14:21:39 | dansmith | artom: if we had migrate-on-load for a while, then it'd be fine, but that's the problem trying to do the switch quickly | |
| 14:21:41 | sean-k-mooney | we coudl do both e.g. do the migrate on load and then wen we drop that provide a nova manange command with a blocker migration | |
| 14:22:15 | dansmith | sean-k-mooney: stephenfin has been trying to get this done for a long time so we're trying not to make this a career-long arc | |
| 14:22:46 | artom | dansmith, it's a thing he'll bequeath to his grandchildren | |
| 14:23:08 | sean-k-mooney | yep i know if this was a few weeks ato i woudl have suggeted doing the migrate on load in ussuri and then the blocker/migrate command in victoria | |
| 14:23:16 | stephenfin | 826 days | |
| 14:23:55 | stephenfin | oh my, I appear to have broken zuul https://review.opendev.org/#/c/724332/ | |
| 14:24:24 | artom | Dayyyyuuuuum | |
| 14:24:30 | sean-k-mooney | hehe let me check with infra | |
| 14:27:43 | sean-k-mooney | zuul is being restart | |
| 14:28:00 | sean-k-mooney | it hit an out of memory issue and they are currently trying to fix it | |
| 14:29:17 | sean-k-mooney | so hold rechecks for a few minutes while they sort this out | |
| 14:29:50 | sean-k-mooney | well clearly his patch ate all the memory | |
| 14:35:00 | sean-k-mooney | ok zuul is back up. changes running before 14:00 UTC have been requed anything uploaded or approved bettween 14:00 and 14:30 needs to be rechecked | |
| 14:35:19 | sean-k-mooney | infra are going to send a staus update for the same shortly | |
| 14:52:11 | kashyap | sean-k-mooney: Since you've reviewed an older version (PS-5) and a newer one (PS-7), I'll just address the PS-7 bits here: https://review.opendev.org/#/c/631154/ | |
| 14:52:15 | kashyap | sean-k-mooney: That okay? | |
| 14:55:54 | sean-k-mooney | kashyap: am sure | |
| 14:56:25 | sean-k-mooney | i think i coppied most of the relevent bits although if you read both and just resond on 7 that is fine with me | |
| 14:56:42 | sean-k-mooney | or well update it in version 8 | |
| 14:56:54 | kashyap | sean-k-mooney: While I comment in the spec, on the "increased memory usage" bit -- I knew that thing, but haven't explicitly mentioned it because it requires precise tests, in what scenarios, etc | |
| 14:57:06 | kashyap | You can't just put a generic: "in all cases memory is increased" | |
| 14:57:32 | sean-k-mooney | kashyap: not really even with 1 pci root port dan found it used more memeory the pc | |
| 14:57:34 | kashyap | It requires more testing; so somebody ought to do the "performance testing guy's job"... | |
| 14:57:45 | sean-k-mooney | so i think in all configurtion it has more memory overhead | |
| 14:58:06 | kashyap | sean-k-mooney: Right; I'll mention that, but need to carefully write it in context and with a config example | |
| 14:58:42 | sean-k-mooney | or we can jsut say we expect that q35 will use more memroy as we have never seen a case where it uses less | |
| 14:58:44 | kashyap | sean-k-mooney: Also I don't want us to get sucked into that black-hole and get derailed... | |
| 14:58:56 | kashyap | But that begs the question: "how much more memory than before" | |
| 14:59:13 | sean-k-mooney | sure an that we can leave to the operator | |
| 14:59:14 | kashyap | Which requires clear a example benchmark | |
| 14:59:18 | sean-k-mooney | or perfomance guy | |
| 14:59:42 | kashyap | Okido; I actually first mentioned it locally and then removed it, as I was still thinking of it | |
| 15:00:37 | sean-k-mooney | i mainly just want them to be aware that they should consider it when upgrading so they can factor it in to there host memory reservation and capastity planning | |
| 15:01:55 | kashyap | Yeah, definitely. Thx for the taking time respond. | |
| 15:35:28 | spatel | sean-k-mooney: do you know how much CPU would be enough to reserve for hypervisor? using isolcpus option? | |
| 15:45:11 | sean-k-mooney | i dont adviase using isolcpus | |
| 15:46:11 | sean-k-mooney | generally 1 phsycial core is more then enough on a compute node or 1 per numa nodes if you want to do affintiy of interupts | |
| 15:47:59 | spatel | sean-k-mooney: but mostly for NFV they suggest using isolcpus for isolation | |
| 15:48:15 | sean-k-mooney | spatel: i generally recommend you use the vcpu_pin_set or in train+ cpu_dedicate_set and cpu_shared_set to do the reservation | |
| 15:48:34 | spatel | you are saying 2 cpu core would be more than enough for hypervisor (per NUMA)? | |
| 15:48:40 | sean-k-mooney | spatel: you should only use isolcpus on realtime hosts and only on the cpus used for pinned vms | |
| 15:49:32 | spatel | I am doing all cpu pinning (dedicated) option for all my workload | |
| 15:49:55 | spatel | we need performance not quantity. | |
| 15:51:14 | spatel | I am planning to use isolcpus + vcpu_pin_set (both option to allocate dedicated CPU) | |
| 15:55:21 | openstackgerrit | Jiri Suchomel proposed openstack/nova master: Add ability to download Glance images into the libvirt image cache via RBD https://review.opendev.org/574301 | |
| 15:55:24 | spatel | sean-k-mooney: ^^ | |
| 15:56:13 | sean-k-mooney | 1 phsyical core(2 hyperthreads) is normally enouch for a compute node if you are not using heavy telemetry | |
| 15:56:40 | sean-k-mooney | spatel: you can use isolcpus + vcpu_pin_set but only if the vm is pinned | |
| 15:57:52 | sean-k-mooney | when you use isolcpus it disables the linux kernel shcduler for those cores | |
| 15:58:05 | sean-k-mooney | so if you have floating vms then they wont float | |
| 15:58:30 | spatel | sean-k-mooney: sounds good, yes we do pinned VM (currently i have assigned 8 cores but wanted to see what people mostly recommend ) | |
| 15:58:32 | sean-k-mooney | in generally isolcpus is only a good idea if you are running realtime wrokloads | |
| 15:58:59 | sean-k-mooney | spatel: you might want to look into tuned by the way | |
| 15:59:11 | sean-k-mooney | it supports configuring this via userspace/sysfs | |
| 15:59:25 | spatel | tuned profile? | |
| 15:59:39 | sean-k-mooney | https://github.com/redhat-performance/tuned/tree/master/profiles/cpu-partitioning | |
| 16:00:02 | sean-k-mooney | isolcpus is a deprecated kernel argument | |
| 16:00:04 | sean-k-mooney | https://github.com/redhat-performance/tuned/blob/master/profiles/cpu-partitioning/cpu-partitioning-variables.conf | |
| 16:00:25 | melwitt | gmann: I'm trying to understand a bit about how/why two grenade jobs would run on stable/ussuri (we don't have an example yet). and I looked at the tempest change and realized I don't understand why it ran two grenade jobs on the name change patch https://review.opendev.org/722551 could you please explain why two jobs run on openstack/tempest? I thought it would have been only one | |
| 16:01:01 | sean-k-mooney | tuned uses the the sysfs/cgroups interface to achive the same effect without the drawbacks | |
| 16:01:31 | spatel | sean-k-mooney: ohh good to know :) | |
| 16:01:37 | spatel | will look into that | |
| 16:01:54 | sean-k-mooney | normally i would jsut set isolated_cores=2,4-7 and not set no_balance_cores=5-10 | |
| 16:02:09 | sean-k-mooney | although no_balance_cores=5-10 would be useful for realtime hosts or ovs-dpdk | |
| 16:02:52 | spatel | Do i need to restart machine to set this values ? | |
| 16:03:54 | gmann | melwitt: sure. for nova stable/ussuri, it will be both job running if you recheck any ussuri backport (or testing patch) until 724189 is merged. this is because compute template in Tempest switched to new job (https://review.opendev.org/#/c/722551/3/.zuul.yaml@543) and nova stable/ussuri ./.zuul.yaml have old job also listed for irrelevant file | |
| 16:04:01 | sean-k-mooney | spatel: am i dont think so | |
| 16:04:16 | sean-k-mooney | you would on teh kernel command line but not with tuned | |
| 16:04:23 | gmann | melwitt: https://github.com/openstack/nova/blob/stable/ussuri/.zuul.yaml#L400 | |
| 16:04:39 | spatel | sean-k-mooney: yes kernel does require reboot but lets me test in tuned | |
| 16:05:14 | melwitt | gmann: sorry I mean as an aside, why did two jobs run on https://review.opendev.org/722551 ? I realized I didn't understand that | |
| 16:05:43 | gmann | melwitt: on Tempest side it is still running because Tempest gate run 'integrated-gate-py3' template which is running all service tests. and that template is on openstack-zuul-jobs side so once i update that template then Tempest also will have single new job | |
| 16:06:00 | melwitt | thanks | |
| 16:06:54 | gmann | compute and service specific template are taken care by 722551 but integrated-gate-py3 template is not yet. | |
| 16:07:59 | gmann | best things we did in grenade side is we alias the grenade-py3 to new zuulv3 native job to avoid running legacy + new jobs during this migration. It is same zuulv3 jobs running twice with different name so will not cause issue. | |
| 17:09:53 | openstackgerrit | Stephen Finucane proposed openstack/nova master: WIP: objects: Add migrate-on-load behavior for legacy NUMA objects https://review.opendev.org/724381 | |
| 17:10:18 | stephenfin | dansmith: That's not complete, but when you've a chance can you sanity check to see if that's what you're after? ^ | |