| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-02-07 | |||
| 19:04:50 | Sundar | gibi: Any pointers to how you handled reschedule for bandwidth provider? | |
| 19:04:56 | spatel | Do you think i need to do something with -> hw:cpu_thread_policy ? | |
| 19:05:28 | sean-k-mooney | spatel: it woudl appear the aplication is not correctly optimising for the numa toplogy. | |
| 19:05:41 | sean-k-mooney | hw:cpu_thread_polciy wont help | |
| 19:06:23 | sean-k-mooney | you can alther the cpu thread and socket toplogy in the guest which might but this does not seeam to be a nova/libvirt issue | |
| 19:07:12 | sean-k-mooney | spatel: my guess is that you have not toled the guest it is using hyper treads so it is miss allocationg threads internally to cores it think are independed but are not | |
| 19:07:33 | spatel | I am just trying to understand if i run application on bare metal then performance is great but when i create VM with all CPU then performance is bad.. | |
| 19:07:36 | sean-k-mooney | hw:cpu_thread_policy=isolate woudl help but it woudl reduce the available core count | |
| 19:08:06 | sean-k-mooney | you shoudl set hw:cpu_treads=2 and hw:cpu_sockets=2 | |
| 19:08:11 | spatel | how may core count it will reduce ? | |
| 19:08:16 | sean-k-mooney | but you cant do that with 30 cpus | |
| 19:08:26 | sean-k-mooney | you need either 28 or 32 | |
| 19:08:35 | spatel | i will go with 28.. | |
| 19:08:39 | sean-k-mooney | so that you have an even number of cpus per numa node/socket | |
| 19:08:39 | Sundar | gibi: "does the series gracefully handle that move operations are not supported yet with accelerators?" There was an earlier attempt to explicitly block unsupported ops: https://review.opendev.org/#/c/674726/ But it was decided that it is not needed, and we should document the supported ops in Cyborg. | |
| 19:09:20 | spatel | you are saying i should set flavor -> hw:cpu_treads=2 and hw:cpu_sockets=2 and run benchmark right? | |
| 19:09:35 | sean-k-mooney | spatel: set hw:cpu_treads=2 and hw:cpu_sockets=2 and do not set hw:cpu_thread_policy=isolate and see if that fixes the issue | |
| 19:09:48 | spatel | sean-k-mooney: sounds good | |
| 19:10:12 | sean-k-mooney | spatel: actuly i misspelled the first one check the flavor docs to confrim they are right | |
| 19:10:27 | sean-k-mooney | spatel: https://docs.openstack.org/nova/latest/user/flavors.html | |
| 19:10:37 | spatel | I will check doc don't worry all i need clue to try something :) | |
| 19:11:17 | Sundar | gibi: "nova needs to grow support for migration with arqs" Since Cyborg does PCI passthrough, live migration is not supported. | |
| 19:11:40 | spatel | sean-k-mooney: you are saying that option will tell my VM you have two socket.. | |
| 19:11:41 | sean-k-mooney | by default openstack emulates each gust cpu as a sperate socket which is incorrect if you have HT enabled on the host | |
| 19:11:54 | spatel | I do have HT enabled | |
| 19:12:17 | sean-k-mooney | yep so you want to tell the vm it has 2 threads per cpu and in this case 2 socket 1 per numa node | |
| 19:12:34 | sean-k-mooney | that will help the guest kerenl make correct schduling decisions | |
| 19:13:03 | spatel | sean-k-mooney: i think you are correct.. that could be the issue why erlang getting confused | |
| 19:13:19 | spatel | because erlang also run inside VM and it should have own scheduler | |
| 19:14:40 | spatel | sean-k-mooney: I will keep you posted about my testing.. (today isn't possible but Monday i will have some result. | |
| 19:15:44 | spatel | sean-k-mooney: let me create i should set hw:cpu_treads=2 and hw:cpu_sockets=2 and hw:numa_node=2 right? | |
| 19:16:06 | spatel | otherwise it won't let me run my VM on two numa nodes | |
| 19:16:52 | sean-k-mooney | the socket and numa nodes dont have to match but it general works better | |
| 19:17:13 | Sundar | sean-k-mooney: Catching up on earlier discussion about instance ops with Cyborg. Explicit blocking of requests with a device profile name in the extra specs ( https://review.opendev.org/#/c/674726/) was not considered the way to go. Is that still your thinking? | |
| 19:18:11 | spatel | sean-k-mooney: i will set all three option and give it a try.. thank you.. | |
| 19:18:38 | sean-k-mooney | Sundar: am honestly i dont rember the full context of the discussin. i know we discussed this at lenght but cant recal what the decision was | |
| 19:19:32 | sean-k-mooney | i think we said skip the explcit check but document what works? but we can aslo block it if we want | |
| 19:20:18 | sean-k-mooney | So the consensus is Option 2: | |
| 19:20:19 | sean-k-mooney | > Gradually phase in the support for the server operations and document the limitations in the meantime but don't actively block them in the API like this change does. | |
| 19:20:21 | sean-k-mooney | We will state the limitations, if any, in Cyborg documentation. | |
| 19:20:30 | sean-k-mooney | that was the last comment on that | |
| 19:22:11 | sean-k-mooney | option 2 was "2. Gradually phase in the support for the server operations and document the limitations in the meantime but don't actively block them in the API like this change does. They either work (by chance) or they don't, but they aren't officially supported. Once they are supported, we patch them in without a new microversion as bug fixes (or just claim test support so they are no longer | |
| 19:22:13 | sean-k-mooney | considered experimental)." | |
| 19:23:18 | sean-k-mooney | gibi: dansmith ^ are ye still ok with that regarding the livemigation check | |
| 19:23:44 | sean-k-mooney | so no check for now and document. and we can add one if we want in the future | |
| 19:25:37 | dansmith | sean-k-mooney: where was that? in the spec? | |
| 19:25:46 | dansmith | it really depends on what the result is | |
| 19:25:55 | sean-k-mooney | https://review.opendev.org/#/c/674726/ | |
| 19:26:01 | sean-k-mooney | its was a nova patch | |
| 19:26:11 | dansmith | if it's data corruption, state intervention required, etc then it needs a check to be graceful | |
| 19:26:22 | dansmith | if it fails in some reasonable way then I'm not so concerned | |
| 19:26:42 | sean-k-mooney | dansmith: i think libvirt will raise an error | |
| 19:27:03 | sean-k-mooney | qemu will reject a migration if the domain has a hostdev that is not of type usb | |
| 19:27:20 | sean-k-mooney | so we will get to the migrate call and it will fail | |
| 19:27:20 | dansmith | for live migration I assume, but based on the changes we've made to the flow, I'm not sure what will happen on resize | |
| 19:27:24 | sean-k-mooney | yes | |
| 19:27:28 | dansmith | we might migrate the instance and ignore the fact that it's missing an accelerator | |
| 19:27:59 | sean-k-mooney | ya i dont know what will happen for resize evacuate or shelve | |
| 19:28:01 | sean-k-mooney | ill find out | |
| 19:28:05 | dansmith | if live fails in a predictable and recoverable way, then I'm fine without a check on that one, | |
| 19:28:22 | dansmith | which is what I said earlier that virt-specific limitations shouldn't be enforced in the api when we can help it | |
| 19:29:45 | Sundar | dansmith, sean-k-mooney: FWIW, IMHO, an explicit check for device profiles in extra specs, as in https://review.opendev.org/#/c/674726/, is probably the safest and clearest to the user. | |
| 19:30:26 | Sundar | We did say that it should be documentation only. But I don;t know if folks will read Cborg dics, or any docs, before kicking off an op. | |
| 19:31:24 | sean-k-mooney | Sundar: the main issue with that patch was it was too agressive in what it blocked | |
| 19:31:27 | Sundar | *Cyborg docs | |
| 19:31:39 | sean-k-mooney | and users never read docs until it breaks | |
| 19:31:58 | Sundar | sean-k-mooney: The details of the patch can be adjusted. | |
| 19:32:38 | Sundar | dansmith, efried, gibi, sean-k-mooney: Would you all recommend to bring back https://review.opendev.org/#/c/674726/? | |
| 19:36:02 | sean-k-mooney | i would not restore it as is. if we add code to block it we shoudl only block the operation we know do not work. | |
| 19:38:21 | sean-k-mooney | that change blocks all snapshoting and backups, and interface and volume atache/removal and other operations like rescue and lock. | |
| 19:39:53 | sean-k-mooney | blocking resize,live-migrate and evacuate might make sense. the rest i think are questionable | |
| 19:40:38 | sean-k-mooney | also blocking shelve might make sense but again we said document and fix as bugfixes so we could adress them one by one without api changes for each | |
| 19:42:44 | Sundar | "Without API changes" -- are you thinking of microversion changes for blocking now and end every unblock in the future, if we support more ops? | |
| 19:43:39 | Sundar | *and every | |
| 19:43:48 | sean-k-mooney | yes im saying we didnt want to do a microversion bump for evey one | |
| 19:47:18 | sean-k-mooney | speaking of op i just booted a vm with your latest revision so ill go test some of them | |
| 19:48:23 | Sundar | Sure. Thanks. | |
| 19:50:08 | sean-k-mooney | my isp broke my home network so its takeing longer then i hoped to test this. i might set up the multi node setup on my laptop instead of my home openstack on monday but ill do what i can on a singel node first | |
| 20:09:13 | efried | Sundar: I've always been in favor of that idea (blocking unsupported operations with a useful message rather than letting them fail "organically" and mysteriously), but I know others disagree. | |
| 20:11:25 | sean-k-mooney | for what its worth i have just done boot,stop,start,reboot,add/remove volume, rescue/unrescue and so far no errors | |
| 20:11:56 | sean-k-mooney | i am not seeing any real interaction with cyborg during those operations | |
| 20:12:42 | sean-k-mooney | which is more or less expected but im not sure if we would loss the acclerator when we regenerate the xml | |
| 20:13:26 | sean-k-mooney | i will look at this more closely next week but right now im just check to make sure the operation complette correctly | |
| 20:29:01 | Sundar | sean-k-mooney: Great. FWIW, I do most of these with FPGAs. Except rescue/unrescue with different images. The list of ops i have checked are in: https://review.opendev.org/673735 | |
| 20:29:44 | sean-k-mooney | im still makeing my way through the list. | |
| 20:30:11 | sean-k-mooney | im hoping to get access to a real server with a rush creek or vista creek next week or the week after | |
| 20:30:20 | sean-k-mooney | which one are you using again? | |
| 20:31:09 | sean-k-mooney | i have jsut done add/remove network interface, pause,unpause, suspend,resume, lock,unlock and rebuild | |
| 20:31:13 | Sundar | Rush Creek, DCP 1.2 | |
| 20:31:40 | sean-k-mooney | cool we have 1 server i think with one of each so ill ask for the rush creek system | |
| 20:31:56 | Sundar | OPAE version 1.1.2-1 | |
| 20:32:26 | Sundar | The OPAE packages come with some sample bitstreams, sp. NLB modes 0 and 3. I use both. | |
| 20:33:58 | Sundar | I can also be reached at ns1.sundar AT gmail DOT com if there is a need for more detailed responses or file transfers. | |
| 20:35:29 | Sundar | I have some utility functions to create/delete device profiles, ARQs using curl. ould be happy to share them if you prefer. | |
| 20:37:01 | sean-k-mooney | im just using the openstack client and i have a test script | |
| 20:38:26 | Sundar | The openstack client is WIP. There are patches to recast it to use openstacksdk etc. https://review.opendev.org/#/c/681391/ | |
| 20:38:48 | sean-k-mooney | yep im using that | |
| 20:39:02 | sean-k-mooney | with the openstacksdk patch too | |
| 20:39:26 | sean-k-mooney | there are a bunch of design issue with it that should be adressed but its kind of useable | |
| 20:40:11 | sean-k-mooney | like when you create a device profile you can only spify its name not the uuid but it only allows you to show a device profile by uuid not name | |