| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-01-16 | |||
| 14:30:49 | sean-k-mooney | dansmith: efried it looks like there are some issues with the arq binding on the cyborg side http://paste.openstack.org/show/788470/ | |
| 14:31:18 | stephenfin | Yup, have spotted the use of 'TimestampMixin' in NovaBase | |
| 14:32:05 | efried | gibi: I'd like to queue up some review time from you for the provider config file bp if possible please. It's pretty straightforward stuff, kind of fun to look at IMO. It's very close to ready, if you want to preview: https://review.opendev.org/#/q/status:open+project:openstack/nova+branch:master+topic:bp/provider-config-file | |
| 14:32:30 | gibi | efried: ack, tomorrow I can start the day with that | |
| 14:32:35 | efried | great, thank you. | |
| 14:32:50 | dansmith | sean-k-mooney: ack | |
| 14:33:31 | belmoreira | Hi, we are hitting https://review.opendev.org/#/c/684375/3 can we get an agreement on how to proceed? kashyap stephenfin | |
| 14:33:49 | efried | sean-k-mooney: are other nova api comms working in that env? Could it just be a creds-in-config thing? | |
| 14:34:09 | sean-k-mooney | ill try booting a normal vm | |
| 14:34:17 | sean-k-mooney | that shoudl trigger the neutron events | |
| 14:35:51 | sean-k-mooney | ya that booted fine | |
| 14:35:52 | efried | sean-k-mooney: point would be cyborg.conf's creds to [nova] | |
| 14:35:55 | sean-k-mooney | so it could be | |
| 14:36:22 | sean-k-mooney | nova] | |
| 14:36:24 | sean-k-mooney | project_domain_name = Default | |
| 14:36:26 | sean-k-mooney | project_name = service | |
| 14:36:28 | sean-k-mooney | user_domain_name = Default | |
| 14:36:30 | sean-k-mooney | password = password | |
| 14:36:32 | sean-k-mooney | username = nova | |
| 14:36:34 | sean-k-mooney | auth_url = http://192.168.135.77/identity | |
| 14:36:36 | sean-k-mooney | auth_type = password | |
| 14:36:38 | sean-k-mooney | that looks fine to me | |
| 14:37:30 | sean-k-mooney | well fine for devstack | |
| 14:37:32 | stephenfin | belmoreira: Should have been more exact about that. I'm not happy with changing the default for everyone to work around a temporary blip with a buggy kernel (which is what that appears to be) | |
| 14:38:34 | sean-k-mooney | we should not proceed with that patch unless we really really reallly have too | |
| 14:38:49 | stephenfin | sean-k-mooney: are you referring to the hugepages one or something else? | |
| 14:39:10 | efried | stephenfin: would you be okay if we implemented the [workaround] opt as suggested? | |
| 14:39:13 | sean-k-mooney | the mempages mode=preferred change | |
| 14:40:07 | stephenfin | efried: That depends on when the bug was introduced and when it was fixed. Latter seems to be 4.19 but kashyap's notes don't appear to capture the former | |
| 14:40:19 | belmoreira | sean-k-mooney what's your main concern? | |
| 14:40:51 | stephenfin | not sean-k-mooney but I'm going to guess potentially huge performance penalties if you haven't correctly configured your host | |
| 14:42:20 | sean-k-mooney | so i originally suggested perfered to kashyap as a potentil solution but im more conserned about non hugepage guest that use hw:numa_nodes=* | |
| 14:42:34 | kashyap | belmoreira: Hi | |
| 14:42:45 | sean-k-mooney | for hugepage we pin the pages seperately so they will be allocated form the numa node we expect | |
| 14:43:29 | sean-k-mooney | however setting perfered by default means we can be sure that the memory of a gues that has a numa doploty but is not backed by hugepage will come form the numa node we expect | |
| 14:43:58 | sean-k-mooney | which will cause issue if we ever model numa in placment | |
| 14:45:01 | sean-k-mooney | i suspect that most of the time it would likely be fine but im not sure what all the sideeffect will be | |
| 14:45:17 | sean-k-mooney | belmoreira: have you tested the patch and conrimed that it fixes the issue? | |
| 14:45:41 | sean-k-mooney | belmoreira: we were not able to get our customer to do that and we did not have a plathform where we could repoduce the issue | |
| 14:45:59 | sean-k-mooney | belmoreira: so we dont acatully know that will fix it although we suspect it will | |
| 14:46:08 | kashyap | stephenfin: I don't know when exactly it was introduced; I need to do kernel commit message digging (if it's not there, then worse, do bisection myself or ask a KVM dev) | |
| 14:47:10 | sean-k-mooney | stephenfin: belmoreira i would be more ok with his as a workaround config option and keeping the default to strict with the expection we might remvoe it in the future. | |
| 14:48:07 | kashyap | stephenfin: On forcing distros to backport — normally, that would be my approach such kernel cases. But here, we need to bear in mind all the existing OpenStack deployments in the wild. | |
| 14:49:28 | sean-k-mooney | i think this issue is only going to show up on centos/rhel maybe suse | |
| 14:49:56 | kashyap | sean-k-mooney: I think they'd be the distros that are most likely to _backport_ those | |
| 14:50:06 | sean-k-mooney | if you are using the hwe kernel on ubunutu 18.04 you will be on 5.x | |
| 14:50:12 | kashyap | Right | |
| 14:50:55 | sean-k-mooney | i supose if you stuck with the orginal ubuntu 18.04 kernel that was 4.15 | |
| 14:52:06 | kashyap | belmoreira: I take it that rebooting the host kernel is not an option in your deployments? | |
| 14:52:27 | sean-k-mooney | belmoreira: o i see you noted that the patch did solve the problem | |
| 14:52:53 | kashyap | Yeah, he tested it | |
| 14:53:33 | kashyap | sean-k-mooney: And in general, 'preferred' is what we recommend _anyway_, no? | |
| 14:53:57 | sean-k-mooney | no | |
| 14:54:07 | sean-k-mooney | we recommend strict | |
| 14:54:09 | kashyap | I have to remind myself of this everytime. | |
| 14:54:43 | sean-k-mooney | prefered makes the memory allcoation of non hugepage memory a hint not a requirement | |
| 14:54:46 | kashyap | request is turned to a "hint", i.e. "if possible, allocate from the | |
| 14:54:46 | kashyap | "Where, 'preferred' means: disable NUMA affinity and so the | |
| 14:54:48 | kashyap | given NUMA node; otherwise, fallback to other nodes" | |
| 14:54:54 | kashyap | Yes | |
| 14:55:10 | sean-k-mooney | which will break our numa aware memory tracking | |
| 14:55:24 | sean-k-mooney | and will not work if we model memory in plamcent in the future per numa node | |
| 14:55:36 | sean-k-mooney | since we can rely on the memory being allcoated form the numa node we said | |
| 14:55:58 | sean-k-mooney | for hugepage memeory the option will be ignored so hugepage backed vms will be fine | |
| 14:56:36 | kashyap | Yeah, then I don't want to muck with defaults. With all these future caveats (which are not properly documented.) | |
| 14:56:40 | sean-k-mooney | but vm with a numa toplogy but no hugepages could have there memory allcoated from any numa node if the one we select was full | |
| 14:58:21 | sean-k-mooney | stephenfin: kashyap belmoreira i left a comment on the patch | |
| 15:14:40 | belmoreira | sorry, needed to leave the office | |
| 15:15:50 | belmoreira | sean-k-mooney yes, I can reproduce with centos 7.6 kernel: Linux XXX 3.10.0-957.1.3.el7.x86_64 #1 SMP Thu Nov 29 14:49:43 UTC 2018 x86_64 x86_64 x86_64 GNU/Linux | |
| 15:15:59 | belmoreira | libvirtd 4.5.0-23 | |
| 15:16:05 | belmoreira | qemu 2.12.0-33 | |
| 15:16:40 | belmoreira | changing to "preferred" fix it (can create new instances) | |
| 15:16:58 | sean-k-mooney | ya as far as i know this is not fixed in 7.6 | |
| 15:17:07 | sean-k-mooney | i think the fix should be in 7.8 | |
| 15:17:20 | belmoreira | kashyap we upgraded to centos 7.7 but didn't reboot the compute nodes. It's a very painful process | |
| 15:17:32 | sean-k-mooney | it might be in 7.7 | |
| 15:18:02 | sean-k-mooney | belmoreira: are you ok with it if we make it a workaround config option and default to strict | |
| 15:18:10 | belmoreira | sean-k-mooney for now I can't reproduce it with 7.7 | |
| 15:18:29 | sean-k-mooney | ok cool then that likely already has the backport | |
| 15:18:41 | sean-k-mooney | although this is hardware specific | |
| 15:18:54 | sean-k-mooney | so if its a differen model fo server then you migh not see it | |
| 15:19:21 | belmoreira | sean-k-mooney yes, took me some time to isolate this | |
| 15:20:08 | sean-k-mooney | ya it took us a day or two to figure this out downstream too when it was first reported to us | |
| 15:20:18 | sean-k-mooney | its a very rare edgecase | |
| 15:20:28 | sean-k-mooney | normlaly you will have dma regions per socket | |
| 15:21:03 | sean-k-mooney | or numa node | |
| 15:21:14 | belmoreira | I think that the workaround config will be useful for everyone that is running NUMA aware VMs. Then needs to be backported to older releases | |
| 15:21:29 | sean-k-mooney | so you need a server that does not + requst a numa toploy + have a kernel with the kvm bug | |
| 15:22:08 | sean-k-mooney | well no it will be useful for the minority that have ^ | |
| 15:22:38 | sean-k-mooney | but i would be ok with proceeding on that basis | |
| 15:22:43 | belmoreira | sean-k-mooney ok, true :) | |
| 15:24:27 | belmoreira | but also shows that the nova community is aware of the issue and has a work around | |
| 15:24:32 | belmoreira | sean-k-mooney thanks | |
| 15:25:18 | kashyap | belmoreira: We're very much aware, FWIW, see this response on 'openstack-discuss' list: http://lists.openstack.org/pipermail/openstack-discuss/2019-September/009770.html | |
| 15:26:08 | belmoreira | sean-k-mooney but to clarify, "preferred" would break the future placement tracking and this should be disable as soon the new fixed kernel is loaded | |
| 15:26:37 | sean-k-mooney | that placement tracking proably wont be a thing until V | |
| 15:26:50 | sean-k-mooney | i could mess up our tracking in the resouce tracker today | |
| 15:27:18 | sean-k-mooney | but if you have the choice between not being able to use an entire socket | |
| 15:27:35 | sean-k-mooney | or slightly reduced performace im guess your ok with the slight perfromace hit | |
| 15:27:48 | sean-k-mooney | and slightly in accurate tracking | |