Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-24
13:22:52 sean-k-mooney hw_arcitecture is menat to defien the arch of the vm.
13:23:37 sean-k-mooney mnaser: your relying on the fact that today kvm acclerate does not work when you are emulating a different architecture
13:23:41 openstackgerrit Chris Dent proposed openstack/nova master: DNM: Make nova tests pass with externalized placement https://review.openstack.org/596291
13:24:07 openstackgerrit Balazs Gibizer proposed openstack/nova-specs master: Add subtree filter for GET /resource_providers https://review.openstack.org/595236
13:24:31 mnaser sean-k-mooney: i don't think that's likely going to change anytime soon™
13:24:55 sean-k-mooney mnaser: well my first respocne to that patch is we should revert it with out a spec
13:25:12 mnaser and that would mean we can't deliver multiarchitecture in our cloud
13:25:17 mnaser so i'm not sure how that benefits anyone
13:25:33 mnaser the patch doesn't hurt anyone except help users deliver a better openstack experience
13:25:34 sean-k-mooney mnaser: you can you just need a weigher.
13:26:00 mnaser doesnt the scheduler filter out N number of items to weight
13:26:08 sean-k-mooney mnaser: my issue is you did this in a filter so its a hard requirement instead of a perference
13:26:15 mnaser it is not
13:26:32 mnaser if the hw_architecture field is empty
13:26:41 mnaser or there is no default configured value
13:26:43 mnaser there is no enforcement
13:27:03 sean-k-mooney yes but i cant set hw_architecture and hv_type anymore
13:27:09 mnaser yes you can
13:27:19 mnaser this patch only uses the configured default IF none is specified in the image.
13:27:32 mnaser you are more than welcome to set hw_arch to aarc64 and hv_type to qemu
13:27:40 mnaser and you'll get software virtualized arm64 vm
13:27:59 mnaser this is just the default values if *none* are specified, which the current openstack nova behavior = put it anywhere.
13:28:22 sean-k-mooney so i can set hw_architecture=x86 and hv_type=kvm and land on a powerkvm node?
13:28:49 mnaser sean-k-mooney: no, if no values are set, you can end up anywhere yes
13:28:57 mnaser but the thing is, 99% of our customers *dont* send hw_arch and hv_type
13:29:11 mnaser because they assume that the cloud is x86_64/kvm which is a reasonable assumption.
13:29:36 sean-k-mooney mnaser: its not how rackspace used to run there public cloud
13:29:38 mnaser multiarchitecture clouds are rare, so asking every single customer/user to tag their images with an architecture is not really smooth user experience, especially in a transition
13:30:21 mnaser sure. so you're saying that for someone to upload an image into openstack, we need to start asking for 3 more parameters
13:30:23 sean-k-mooney they had a large amount of powerpc nodes at one point with xen as the hypervior running with x86 guests
13:31:15 mnaser well, i'd say that's more of an outlier. i don't see how this helps solve things and make them easier for users overall in a multiarch environment
13:31:34 mnaser because the rax env you're describing sounds still like a single arch one for the user facing stuff
13:32:45 sean-k-mooney mnaser: well my point is if you implemented this as a weigher you would land on the same acrh in 99% of cases and fall back to hardware acllerated emulaton if not native nodes were available.
13:33:06 mnaser sean-k-mooney: but don't weighers run on a subset of compute nodes?
13:33:19 sean-k-mooney the chage is merged in anycase but i really dont think this should have been a bug.
13:33:30 sean-k-mooney mnaser: am the run after all the filters
13:33:49 sean-k-mooney so you go through placment then the filters then the weighers
13:33:57 mnaser i was pretty sure that weighers ran on a subset of hosts
13:33:58 sean-k-mooney they just reorder what the filters pass
13:34:13 mnaser hence the values that would let you pick subsets of hosts when scheduling
13:35:44 sean-k-mooney sorry i missed that last point?
13:36:03 mnaser i thought weighers only ran on a subset of filtered hosts
13:36:10 sean-k-mooney yes
13:37:19 mdbooth mriedem: Briefly back to the topic of evacuation tests, am I right in thinking we probably won't be able to write a full job which does failure injection? If not, I don't think I don't think I'm blocked on that. If we *can* I'm happy to write it if you point me at the right repo and an example to crib from.
13:37:56 mnaser sean-k-mooney: was rax doing x86 on ppc via software or hardware emulation?
13:38:18 mdbooth 'example' wouldn't have to be that similar, just enough to orient myself. I haven't touched the top-level CI jobs before.
13:38:20 sean-k-mooney mnaser: i belive xen allows cross arch hardware emulation
13:40:17 mdbooth mriedem: The ability to assert success in the non-failure case might still be interesting, I guess. It would essentially be a regression test.
13:40:27 mnaser sean-k-mooney: i am not too familiar with xen but i would assume that wouldn't be hardware assisted virt (i.e. using VT-x in intel, etc)
13:40:36 mnaser but that's something i dunno much about so i could be entirely wrong
13:41:33 sean-k-mooney mnaser: well xen does use vt-x normally if you are running x86 on x86
13:41:40 sean-k-mooney as does kvm
13:42:04 mnaser right, so this would be the same as software virtualization which can be really slow, which i'm trying to avoid having anyone end up with
13:42:22 mnaser and all of this is optional, if you set none of those settings, it won't affect you. but if you run a cloud where you want to make sure all VMs are hardware virt'd, then you can do that
13:47:11 sean-k-mooney mnaser: its not so much that this change was done that i object to its that the semantics of the existing api were changed without a spec.
13:47:49 mnaser sean-k-mooney: but if you don't configure this very optional value, there will zero behavior change
13:48:36 sean-k-mooney mnaser: yes and if you had those configure before an you upgrade then your cloud may now be broken if they do not have the same semantics
13:48:56 mnaser sean-k-mooney: you couldn't possible have them configured because they don't even exist before?
13:49:13 mnaser s/possible/possibly
13:49:25 sean-k-mooney the image properties
13:50:35 openstackgerrit Stephen Finucane proposed openstack/nova master: api: Remove unnecessary default parameter https://review.openstack.org/564451
13:58:48 mriedem mdbooth: what type of failure needs to be injected? like delete the disk before starting the evacuatoin?
13:59:06 mdbooth mriedem: Any failure in spawn on the destination
13:59:08 mriedem can't really do that from tempest, but could if run from a script
13:59:21 mdbooth I think I chose plug_vifs in my functional test
13:59:36 dansmith mriedem: so on your gate job,
14:00:02 dansmith mriedem: just force-downing the service isn't going to be enough I think, since it will still be running, potentially mucking with allocations and state correction, etc right?
14:00:07 mdbooth Just need to assert that if we fail in spawn in the dest, we only cleanup what we're supposed to and don't delete shared data.
14:00:18 mdbooth Whether that's shared instance dir or shared block.
14:00:20 dansmith you have to actually kill the n-cpu process and force-down it to avoid the wait
14:00:54 dansmith (well, not mucking with allocations until it's deleted maybe anymore, but screaming at least)
14:01:06 dansmith but I think it's more legit if you stop it
14:01:25 mriedem dansmith: force down is what we do in functional tests
14:01:30 mriedem that's all the API checks for
14:01:38 dansmith mriedem: I know it will make evac work,
14:02:05 dansmith but the n-cpu for the "down" one will still be running and doing stuff, which isn't how we describe evac as working properly in real life
14:02:13 dansmith we say you have to fence it
14:02:50 mriedem in a multinode job the post-test script runs on the control node which has an n-cpu process on it which we could kill
14:03:06 mriedem we'd have to force the server to that host so it's the source
14:03:37 stephenfin mriedem, melwitt, dansmith: btw, given that sahid immediately replied yesterday, I figured it wasn't necessary to investigate the queue size thing any further myself. Sorry if I picked that up wrong
14:03:46 mriedem stephenfin: it's fine
14:04:11 mriedem after looking at the existing release note and config option help for the rx/tx queue size stuff i think they were good enough to say that it was very limited to certain types of vifs
14:04:27 mriedem and i think what moshe is trying to do is make it work for something else
14:05:37 stephenfin mriedem: aight, good to hear
14:06:06 melwitt .
14:06:35 mriedem mdbooth: btw yes i'm also interested in simply have regression testing for evac in a real env
14:06:40 mriedem *having
14:07:28 mriedem melwitt: that last rocky change is in the gate
14:09:51 melwitt mriedem: I saw, thank you for getting that. the master change didn't land until late last night after I went to sleep
14:13:59 mriedem mdbooth: btw do you still care about this? https://review.openstack.org/#/c/572790/
14:38:15 mriedem dansmith: a couple of thoughts on batch size floor in https://review.openstack.org/#/c/593131/
14:41:48 dansmith mriedem: thanks will look shortly
14:57:22 mriedem and i'm done with that stack \o/
15:02:35 mdbooth mriedem: I do. Came up yesterday, in fact.
15:02:54 mdbooth mriedem: I don't think I had any strong objections there?
15:03:47 mriedem i'm pretty sure you wanted to +1 it
15:04:28 mdbooth Just re-reading my last review. Sounds like I was prevaricating rather than objecting.
15:06:12 mdbooth Ah... the attach_mode thing.
15:13:46 mdbooth mriedem: So there's attach_mode and access_mode
15:14:23 mdbooth It seems like we only *use* access_mode, but that code is checking attach_mode

Earlier   Later