| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-05-08 | |||
| 12:16:14 | mriedem | we do have a grenade job specifically that tests just live migration back and forth from n-1 to n and back | |
| 12:17:07 | artom | mriedem, I suppose I could just test numa-aware live migration with cpu pinning since we don't need special hardware for that, but there'd be no way to assert anything about the instance XML, which is kinda what we want | |
| 12:17:28 | sean-k-mooney | mriedem: ah you are right looking at http://52.27.155.124/portland/2018-04-27/553072/6/check/tempest-dsvm-multinode-ovsdpdk-nfv-networking-xenial/943bdb9/logs/tempest_conf.txt.gz livemigration is turned on the quest then is why is it not running those tests | |
| 12:18:08 | mriedem | sean-k-mooney: blacklist? | |
| 12:18:22 | sean-k-mooney | mriedem: proably ill have to look into that | |
| 12:18:47 | mriedem | t_live_migration.*)|(?:tempest\.api\.network.*)).*$ != '' ]] | |
| 12:18:47 | mriedem | ver_connectivity_cold_migration.*)|(?:.*\.TestNetworkAdvancedServerOps.test_server_connectivity_resize.*)|(?:.*\.TestNetworkBasicOps\.test_update_router_admin_state.*)|(?:.*\.TestNetworkBasicOps\.test_network_basic_op.*)|(?:.*\.TestNetworkBasicOps\.test_update_instance_port_admin_state.*))((?:tempest\.scenario\.test_network_basic_ops.*)|(?:tempest\.scenario\.test_network_advanced_server_ops.*)|(?:tempest\.api\.compute\.admin\ | |
| 12:18:47 | mriedem | [[ ^(?!.*(?:tempest\.scenario\.test_network_advanced_server_ops\.TestNetworkAdvancedServerOps\.test_server_connectivity_suspend_resume.*)|(?:tempest\.scenario\.test_network_advanced_server_ops\.TestNetworkAdvancedServerOps\.test_server_connectivity_pause_unpause.*)|(?:.*\.admin\.test_live_migration.*)|(?:.*\.TestNetworkAdvancedServerOps.test_server_connectivity_cold_migration_revert.*)|(?:.*\.TestNetworkAdvancedServerOps.test | |
| 12:19:02 | mriedem | yes it's blacklisted live migration | |
| 12:19:08 | sean-k-mooney | artom: you need a multi numa host vm as we have a restiction in the libvirt drivver that maps each guest numa node to a different host numa node | |
| 12:19:12 | artom | Which makes sense, because it's broken | |
| 12:19:37 | sean-k-mooney | artom: actully the reson the multinode job was created was to test livemigration | |
| 12:19:53 | artom | But... we know it doens't work :) | |
| 12:19:58 | mriedem | artom: what do you need to get out of the instance xml? | |
| 12:20:09 | mriedem | the numa config? | |
| 12:20:10 | sean-k-mooney | artom: it was turned off temporaly feburay last year and then the ci moved team and i guess it never got turned back on | |
| 12:20:12 | artom | mriedem, yeah | |
| 12:20:20 | artom | sean-k-mooney, how did it ever pass before? | |
| 12:20:26 | mriedem | artom: we have the instance diagnostics api... | |
| 12:20:41 | sean-k-mooney | artom: livemigration works it just does not respect pinning after | |
| 12:20:53 | artom | sean-k-mooney, ah, indeed | |
| 12:21:06 | sean-k-mooney | e.g. the vm will actully livemigrate we just broke all SLAs | |
| 12:21:36 | mriedem | artom: you could potentially add a numa_details field to the instance diagnostics response and use that for validating things in tempest (likely a tempest plugin) | |
| 12:21:38 | artom | mriedem, I don't think those are enough | |
| 12:22:01 | sean-k-mooney | mriedem: the issue with multi numa testing in the upstream ci is we cannot spawn a 2 numa node guest on a host with 1 numa node | |
| 12:22:10 | sean-k-mooney | using the libvirt driver | |
| 12:22:39 | sean-k-mooney | that is a limitaion that i personally think should never have been there but that is a different issue | |
| 12:23:30 | artom | sean-k-mooney, so is the intel CI multi-physical-node? Or multinode on VMs, so we can't have VMs with more than 1 node? | |
| 12:24:14 | artom | mriedem, would it really be a plugin at that point though? If it's 100% through the API, it's legit tempest | |
| 12:24:27 | sean-k-mooney | artom: mriedem: all the host vms for the intel nfv ci have 2 numa nodes to allow numa testing, the phyical hosts have 2-4 numa nodes dending on the server the nodepool vm lands on | |
| 12:24:45 | mriedem | artom: meaning, it's only compute api | |
| 12:25:01 | mriedem | there isn't a real need to make all other projects gate on a numa test in the common tempest repo | |
| 12:25:39 | artom | mriedem, so how does it work currently for nova-only tempest tests? | |
| 12:26:05 | mriedem | those are legacy | |
| 12:26:13 | mriedem | and some are interop tests | |
| 12:26:21 | artom | mriedem, oh, they're not accepting new tests that are single-service? | |
| 12:26:21 | mriedem | the real basic stuff is interop | |
| 12:26:31 | mriedem | depends | |
| 12:26:46 | artom | Christ, it's getting less useful by the minute | |
| 12:26:54 | artom | (Sorry tempest folks!) | |
| 12:27:03 | mriedem | i just know that over time, single service, non-interop tests were supposed to be moved into the project tree or a tempest plugin for that repo | |
| 12:27:27 | mriedem | most other projects already have tempest plugins, i know cinder and neutron have had their own for a long time | |
| 12:27:50 | artom | I suppose it kinda make sense... | |
| 12:28:11 | artom | Keep the "useful to everyone" stuff in-tree, the rest can be out of scope in plugins | |
| 12:28:16 | artom | Anyways | |
| 12:28:32 | artom | So, I think first step for me is to get live migration re-enabled in the intel NFV CI | |
| 12:28:35 | mriedem | i'm no QA gate keeper, but just don't be surprised if that's what they tell you | |
| 12:29:04 | artom | They'll pass, sortof | |
| 12:29:28 | sean-k-mooney | artom: do you want to test multi numa guests or just guest with a numa topology | |
| 12:29:46 | sean-k-mooney | artom: hw:numa_nodes=1 should work in the upstream ci | |
| 12:29:46 | artom | sean-k-mooney, uh, there's a difference? | |
| 12:29:54 | artom | Ah, in that sense | |
| 12:29:55 | artom | Hrmm | |
| 12:29:59 | artom | True, true | |
| 12:30:37 | artom | mriedem, well, in the short term at least, nothing would stop me from proposing a patch to show the test passing | |
| 12:30:39 | sean-k-mooney | artom: cpu pinning will not work in the upstream ci however which is that the feature you really want to test? | |
| 12:30:53 | artom | And it it gets -2, then we can think about plugins | |
| 12:31:07 | artom | sean-k-mooney, well, everything, ideally | |
| 12:31:10 | artom | Even hugepages | |
| 12:31:17 | artom | I'm not writing any new NUMA code | |
| 12:31:23 | artom | Just calling the old one when live migrating | |
| 12:31:42 | artom | So technically just showing that it gets called for 1 NUMA-ish thing (and does the right thing) would be enough | |
| 12:31:49 | artom | But... the more coverage the better | |
| 12:31:50 | mriedem | melwitt: fyi i've marked https://blueprints.launchpad.net/nova/+spec/convert-consoles-to-objects complete | |
| 12:31:58 | mriedem | artom: sure | |
| 12:32:21 | mriedem | artom: also, a new test would get by the intel 3rd party ci blacklist which is currently based on test names | |
| 12:32:53 | artom | mriedem, oh, hah, ineed. Sneaky :D | |
| 12:33:20 | mriedem | so you'll probably have to do something like have your nova series, and then have a DNM nova patch on top that depends on the tempest change, | |
| 12:33:29 | mriedem | because the intel CI runs on nova changes, but probably not tempest changes | |
| 12:33:52 | sean-k-mooney | i have to run to a meeting but ill be back in an hour or 2 | |
| 12:34:13 | artom | mriedem, yep, and a patch to the intel CI plugin that does stuff like check instance XML | |
| 12:34:29 | artom | mriedem, Or. Or! A patch to nova that adds what I need to the diagnostics API | |
| 12:34:43 | artom | Not sure what would be simpler. | |
| 12:34:56 | mriedem | hypervisor-specific stuff in tempest sucks, | |
| 12:35:04 | mriedem | which is why i suggested adding a new field to the diagnostics api | |
| 12:35:29 | mriedem | alternatively, | |
| 12:35:37 | mriedem | does any of this numa stuff for the guest get modeled in placement? | |
| 12:35:45 | mriedem | as a consumed resource? | |
| 12:35:52 | alex_xu | + | |
| 12:35:57 | artom | mriedem, some of it, I think? But allocations are still on the compute via resource tracker, I believe | |
| 12:35:59 | mriedem | alex likes it | |
| 12:36:16 | mriedem | the numa resource allocations would be on numa resource providers in the compute node provider tree | |
| 12:36:40 | mriedem | but given an instance (consumer) uuid, you can get it's resource class allocations against which providers in placement | |
| 12:36:53 | mriedem | so your test could assert that the instance has NUMA resource class allocations | |
| 12:36:56 | alex_xu | mriedem: my daugther just smash my keyboard... | |
| 12:37:04 | mriedem | ha, lot of that going on today | |
| 12:37:04 | efried | nice find tetsuro | |
| 12:37:22 | artom | alex_xu, she clearly didn't smash hard enough since we can all read what you're typing | |
| 12:37:46 | mriedem | artom: but i don't think the numa stuff is done, or close(?) | |
| 12:37:54 | artom | mriedem, I'm not sure placement would be enough, since we would ideally check specific CPUs, not just quantities | |
| 12:38:03 | artom | And pinning can't be checked at all | |
| 12:38:36 | mriedem | artom: i'm not sure if there is a placement solution in the works for that yet or not, but in that case you could just hack up the diagnostics api | |
| 12:38:41 | artom | mriedem, placement is a pool I swim in, but I still breath through a snorkel, so I don't know what liquid surrounds me | |
| 12:39:12 | artom | At some point I will need to grow gills to breath through the placement pool fluid | |
| 12:39:15 | alex_xu | artom: yea, a little hulk | |
| 12:44:47 | openstackgerrit | Takahito Hirose proposed openstack/python-novaclient master: api_version decorator becomes an error in Python 3.5.0. https://review.openstack.org/564702 | |
| 13:16:59 | openstackgerrit | Kashyap Chamarthy proposed openstack/nova master: libvirt: Deprecate support for monitoring Intel CMT `perf` events https://review.openstack.org/565242 | |
| 13:17:56 | kashyap | mriedem: When you get a minute, I read the scrollback from yesterday here, and went with the: "deprecate in Rocky and hard-fail in Stein" | |
| 13:19:22 | kashyap | I don't think I got the "assert_called_once_with" quite right here: https://review.openstack.org/#/c/565242/5/nova/tests/unit/virt/libvirt/test_driver.py@6623 | |
| 13:20:16 | wznoinsk | mriedem, hi | |