Earlier  
Posted Nick Remark
#openstack-nova - 2018-05-08
12:26:46 artom Christ, it's getting less useful by the minute
12:26:54 artom (Sorry tempest folks!)
12:27:03 mriedem i just know that over time, single service, non-interop tests were supposed to be moved into the project tree or a tempest plugin for that repo
12:27:27 mriedem most other projects already have tempest plugins, i know cinder and neutron have had their own for a long time
12:27:50 artom I suppose it kinda make sense...
12:28:11 artom Keep the "useful to everyone" stuff in-tree, the rest can be out of scope in plugins
12:28:16 artom Anyways
12:28:32 artom So, I think first step for me is to get live migration re-enabled in the intel NFV CI
12:28:35 mriedem i'm no QA gate keeper, but just don't be surprised if that's what they tell you
12:29:04 artom They'll pass, sortof
12:29:28 sean-k-mooney artom: do you want to test multi numa guests or just guest with a numa topology
12:29:46 artom sean-k-mooney, uh, there's a difference?
12:29:46 sean-k-mooney artom: hw:numa_nodes=1 should work in the upstream ci
12:29:54 artom Ah, in that sense
12:29:55 artom Hrmm
12:29:59 artom True, true
12:30:37 artom mriedem, well, in the short term at least, nothing would stop me from proposing a patch to show the test passing
12:30:39 sean-k-mooney artom: cpu pinning will not work in the upstream ci however which is that the feature you really want to test?
12:30:53 artom And it it gets -2, then we can think about plugins
12:31:07 artom sean-k-mooney, well, everything, ideally
12:31:10 artom Even hugepages
12:31:17 artom I'm not writing any new NUMA code
12:31:23 artom Just calling the old one when live migrating
12:31:42 artom So technically just showing that it gets called for 1 NUMA-ish thing (and does the right thing) would be enough
12:31:49 artom But... the more coverage the better
12:31:50 mriedem melwitt: fyi i've marked https://blueprints.launchpad.net/nova/+spec/convert-consoles-to-objects complete
12:31:58 mriedem artom: sure
12:32:21 mriedem artom: also, a new test would get by the intel 3rd party ci blacklist which is currently based on test names
12:32:53 artom mriedem, oh, hah, ineed. Sneaky :D
12:33:20 mriedem so you'll probably have to do something like have your nova series, and then have a DNM nova patch on top that depends on the tempest change,
12:33:29 mriedem because the intel CI runs on nova changes, but probably not tempest changes
12:33:52 sean-k-mooney i have to run to a meeting but ill be back in an hour or 2
12:34:13 artom mriedem, yep, and a patch to the intel CI plugin that does stuff like check instance XML
12:34:29 artom mriedem, Or. Or! A patch to nova that adds what I need to the diagnostics API
12:34:43 artom Not sure what would be simpler.
12:34:56 mriedem hypervisor-specific stuff in tempest sucks,
12:35:04 mriedem which is why i suggested adding a new field to the diagnostics api
12:35:29 mriedem alternatively,
12:35:37 mriedem does any of this numa stuff for the guest get modeled in placement?
12:35:45 mriedem as a consumed resource?
12:35:52 alex_xu +
12:35:57 artom mriedem, some of it, I think? But allocations are still on the compute via resource tracker, I believe
12:35:59 mriedem alex likes it
12:36:16 mriedem the numa resource allocations would be on numa resource providers in the compute node provider tree
12:36:40 mriedem but given an instance (consumer) uuid, you can get it's resource class allocations against which providers in placement
12:36:53 mriedem so your test could assert that the instance has NUMA resource class allocations
12:36:56 alex_xu mriedem: my daugther just smash my keyboard...
12:37:04 efried nice find tetsuro
12:37:04 mriedem ha, lot of that going on today
12:37:22 artom alex_xu, she clearly didn't smash hard enough since we can all read what you're typing
12:37:46 mriedem artom: but i don't think the numa stuff is done, or close(?)
12:37:54 artom mriedem, I'm not sure placement would be enough, since we would ideally check specific CPUs, not just quantities
12:38:03 artom And pinning can't be checked at all
12:38:36 mriedem artom: i'm not sure if there is a placement solution in the works for that yet or not, but in that case you could just hack up the diagnostics api
12:38:41 artom mriedem, placement is a pool I swim in, but I still breath through a snorkel, so I don't know what liquid surrounds me
12:39:12 artom At some point I will need to grow gills to breath through the placement pool fluid
12:39:15 alex_xu artom: yea, a little hulk
12:44:47 openstackgerrit Takahito Hirose proposed openstack/python-novaclient master: api_version decorator becomes an error in Python 3.5.0. https://review.openstack.org/564702
13:16:59 openstackgerrit Kashyap Chamarthy proposed openstack/nova master: libvirt: Deprecate support for monitoring Intel CMT `perf` events https://review.openstack.org/565242
13:17:56 kashyap mriedem: When you get a minute, I read the scrollback from yesterday here, and went with the: "deprecate in Rocky and hard-fail in Stein"
13:19:22 kashyap I don't think I got the "assert_called_once_with" quite right here: https://review.openstack.org/#/c/565242/5/nova/tests/unit/virt/libvirt/test_driver.py@6623
13:20:16 wznoinsk mriedem, hi
13:20:42 zzzeek jaypipes: what would cause lock wait timeout exceeded for an INSERT?
13:21:29 jaypipes zzzeek: another thread executing LOCK TABLES <table>?
13:21:50 zzzeek jaypipes: just that? nothing more subtle? ceilometer is doing it
13:22:32 jaypipes zzzeek: got a log output or something more for me? :)
13:22:58 zzzeek jaypipes: i have the error message and the query id have to spend time looking for the logs.
13:23:10 zzzeek jaypipes: but there's nothing like, the "auto increment" feature or somethign locks
13:23:14 jaypipes zzzeek: the only other thing I can think of would be threads attempting to execute huge transactions.
13:23:26 jaypipes zzzeek: all concurrently
13:23:37 zzzeek jaypipes: right and then innodb locks ...a set of potential rows?
13:24:31 jaypipes zzzeek: no, autoinc won't produce that lock wait timeout generally, unless like I said, you have multiple threads simultaneously attempting to commit huge transactions (with thousands or tens of thousands of data modifications in each trx)
13:25:05 jaypipes zzzeek: yes, innodb will do its gap locks if the PK isn't autoinc.
13:25:11 zzzeek jaypipes: ok but in that csae, what is the lock that the INSERT is waiting for? OK gap locks. got it
13:25:25 jaypipes zzzeek: but again... you need some serious concurrency and huge trx to see this impact IME
13:25:38 zzzeek jaypipes: this is a load test
13:25:39 BlackDex Hello there. Does queens support active/active rw in multiple instance using ceph storage and the correct kvm version
13:25:41 BlackDex ?
13:26:18 jaypipes zzzeek: my guess would be ceilometer is attempting to commit batches of record changes. maybe try reducing the length of time between those commits?
13:26:39 zzzeek jaypipes: I dont even know wehre ceilometer's database code is
13:26:50 jaypipes zzzeek: what version?
13:26:55 zzzeek master
13:27:09 jaypipes zzzeek: lemme grep and see.
13:27:21 jaypipes zzzeek: been a very long time since I looked at ceilometer.
13:27:34 zzzeek jaypipes: $ find ceilometer/ -name "*.py" -exec grep -l sql {} \;
13:27:34 zzzeek [classic@photon2 ceilometer]$
13:27:36 zzzeek zero
13:27:45 zzzeek they've hidden it
13:28:16 jaypipes zzzeek: gnocchi is now the backend data storage for meters, though, right?
13:28:19 zzzeek that's pretty impressive the string "sql" does not appear in their source base at all
13:28:24 jaypipes zzzeek: ceilometer is just the polling thing right?
13:28:39 zzzeek jaypipes: right. but the log is the "ceilometer agent-notification"
13:28:41 jaypipes zzzeek: https://github.com/openstack/ceilometer/blob/master/ceilometer/gnocchi_client.py
13:30:11 zzzeek jaypipes: table name is "event"
13:30:19 zzzeek jaypipes: isn't that the old mysql driver?
13:30:29 jaypipes zzzeek: no idea :(
13:30:32 zzzeek jaypipes: ok
13:33:07 jaypipes zzzeek: is this happening in like a tempest run or something? or is this in a prod env?
13:33:31 jaypipes zzzeek: https://github.com/openstack/ceilometer/blob/master/ceilometer/polling/manager.py#L46 <-- maybe try setting that to False and seeing if lock wait timeouts go down (due to smaller trx sizes)
13:33:43 zzzeek jaypipes: top seekrit :)

Earlier   Later