Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-27
12:39:58 mriedem so we ended last night knowing that conductor and scheduler are working but when super conductor should cast to n-cpu, everything stops there
12:40:06 mriedem super-conductor logs stop, rabbitmq logs stop
12:40:22 mriedem and the request isn't going to like the cell1 conductor, there is nothing in it's logs
12:40:52 mriedem so i've got the d-g patch to try and dump the rabbitmqctl report when cleaning up the host and collecting logs
12:41:08 mriedem but so far i haven't had a run with that work yet
12:41:17 mriedem https://review.openstack.org/#/c/487664/
12:43:56 vdrok mriedem: sdague also uploaded this one https://review.openstack.org/487809
12:44:28 sdague mriedem: the fix was definitely wrong last night, it was still spawning 2 conductors
12:47:18 vdrok sdague: mriedem: regarding that bug, it seems like it's the same thing again, http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-partition-redfish-tinyipa-ubuntu-xenial/7aa067c/logs/apache/error.txt.gz, apache was being restarted at that moment
12:47:57 vdrok we did work around it already a couple of times
12:48:15 sdague vdrok: why apache getting restarted?
12:49:14 vdrok sdague: previously it was because of some configuration changes, not sure why now, need to look into logs
12:49:31 sdague anyway... the main issue
12:50:03 sdague mriedem: so with this nova.conf - http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-partition-redfish-tinyipa-ubuntu-xenial/7aa067c/logs/etc/nova/nova.conf.txt.gz
12:50:08 mriedem sdague: yeah the 2 conductors was the culprit, but dan and i couldn't see or think of anything about why that would be causing issues
12:50:11 sdague how do we know where cell0 db is?
12:50:32 mriedem sdague: conductor gets the cell mappings from the api db,
12:50:38 mriedem and the cell mapping has the cell0 mq and db urls in it
12:51:44 mriedem cdent: replied some and asked a question about another test in here https://review.openstack.org/#/c/487589/
12:51:57 cdent yeah, just reading/responding to that. thanks
12:52:21 mriedem cdent: the *best* way to tell if an instance is being migrated would be to check if it has a migration_context attribute in it
12:52:36 sdague mriedem: ok, so which config file should the single conductor be running off of?
12:52:54 mriedem sdague: nova.conf
12:53:31 sdague
12:53:31 sdague which gives it a transport url of - transport_url = rabbit://stackrabbit:secretrabbit@10.16.80.100:5672/
12:54:13 mriedem cdent: i share concerns about relying on the existing allocations to know if we're doing a migration or not
12:54:26 mriedem because of (1) timing and (2) weird operatoins like soft delete and shelve
12:54:40 mriedem and the wonkiness that is the RT
12:55:19 mriedem sdague: so in the singleconductor case, i think the only nova config that matters is nova.conf http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-partition-redfish-tinyipa-ubuntu-xenial/7aa067c/logs/etc/nova/
12:55:25 mriedem nova-cpu.conf and nova_cell1.conf aren't used
12:55:47 mriedem nova-cpu.conf wouldn't even work b/c it doesn't have any information in it about which compute driver to use, or how to talk to placement/cinder/neutron/etc
12:56:35 mriedem http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-partition-redfish-tinyipa-ubuntu-xenial/7aa067c/logs/devstacklog.txt.gz#_2017-07-27_12_04_56_099
12:56:35 mriedem the cell1 mapping is created here:
12:56:42 mriedem nova-manage --config-file /etc/nova/nova.conf --config-file /etc/nova/nova_cell1.conf cell_v2 create_cell --name cell1
12:56:50 mriedem note that is using nova_cell1.conf rather than nova.conf
12:56:52 sdague mriedem: sure, where would the message queue get set
12:57:17 sdague maybe that's the missing piece, dumping the cell mappings
12:57:32 mriedem well we know the mq and db for cell1, it's taken from http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-partition-redfish-tinyipa-ubuntu-xenial/7aa067c/logs/etc/nova/nova_cell1.conf.txt.gz
12:57:43 sdague right
12:57:50 sdague but that's not used anywhere
12:58:01 mriedem it's used when creating the cell1 mapping
12:58:02 sdague so that's stating that the cell1 mq is going to be on a vhost
12:58:06 mriedem nova-manage --config-file /etc/nova/nova.conf --config-file /etc/nova/nova_cell1.conf cell_v2 create_cell --name cell1
12:58:13 sdague but nova-compute is started not listening to that vhost
12:59:32 mriedem right nova-compute is listening on rabbit://stackrabbit:secretrabbit@10.16.80.100:5672/
12:59:42 mriedem the cell1 mapping is sending to rabbit://stackrabbit:secretrabbit@10.16.80.100:5672/nova_cell1
12:59:48 sdague but conductor isn't sending messages there, right?
13:00:09 mriedem because https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L1042
13:00:21 mriedem conductor does an mq switch when it casts to compute
13:00:21 sdague https://github.com/openstack-dev/devstack/blob/9596fdddccd04c26aa5adb923b9bd8e64c6593ec/lib/nova#L709-L712
13:01:02 sdague mriedem: ok, so are you agreeing or disagreeing with me that nova-cond and nova-compute aren't talking on the same mq :)
13:01:37 mriedem can i phone a friend?
13:01:50 mriedem i think i'm agreeing,
13:01:57 mriedem and that would explain why we're not seeing the message get to n-cpu
13:02:26 sdague yeh
13:02:37 sdague is there a nova-manage command to dump the cell mappings?
13:02:46 sdague I think that's kind of critical to see the mismatch
13:03:48 mriedem yes,
13:03:52 mriedem nova-manage cell_v2 list_cells
13:04:17 mriedem with --verbose
13:04:28 mriedem --verbose dumps the db and mq urls
13:05:25 mriedem fwiw, this is a run before the fleetify patch where we create cell1
13:05:26 mriedem http://logs.openstack.org/23/485823/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/206b79e/logs/devstacklog.txt.gz#_2017-07-21_01_52_17_794
13:05:48 mriedem nova-manage cell_v2 create_cell --transport-url rabbit://stackrabbit:secretrabbit@10.0.1.31:5672/ --name cell1
13:05:51 mriedem there is a single nova.conf
13:06:13 mriedem and it's using the same transport_url http://logs.openstack.org/23/485823/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/206b79e/logs/etc/nova/nova.conf.txt.gz
13:06:32 mriedem so yeah, i think that's the problem, the cell1 mapping is using the nova_cell1 mq and nova-compute is using the main mq
13:08:44 sdague ok, fix proposed
13:08:57 sdague including the dump of the cell mapping
13:10:50 mriedem i left a comment in ps2,
13:10:55 mriedem but it looks like you addressed it in ps4
13:11:07 mriedem with https://review.openstack.org/#/c/487809/4/lib/nova@592
13:14:19 dansmith mriedem: oh snap.. just woke up, but excellent call
13:15:09 mriedem sdague connected the dots for me,
13:15:12 mriedem plus sleep helps
13:15:25 dansmith I should have thought of that
13:15:39 mriedem hard to think of anything at the end of a day like yesterday
13:16:06 sdague mriedem: oh, sorry, I wasn't even looking at comments, I was just reading logs trying to understand
13:17:34 mriedem cdent: regarding functional testing, i was thinking the same yesterday, but didn't have time,
13:17:52 vdrok morning dansmith , missed all the fun :)
13:18:23 dansmith vdrok: me too, but.. good morning vdrok :P
13:18:29 mriedem but was thinking it could be relatively simple to write a functional test that starts 2 compute services, creates a server, checks allocations are just on the source host, does a resize to the 2nd compute, checks allocations are retained for both nodes, then does a resize confirm (and another test that does a resize revert), and then validates allocations after that is done
13:19:15 cdent relatively
13:19:34 mriedem cdent: but me working on something like that probably can't happen until tomorrow at this rate
13:19:51 mriedem i actually have *gasp* family commitments tonight
13:20:11 cdent I’m sicker than sick at the moment, but bored enough to still be hanging out
13:21:26 sean-k-mooney mriedem: hi i know your pretty busy with the ff today but any chance of inlcuding the 1 patch for mellanox's ovs offload and the 2 required for our nic feature based scheudling blueprint in pike or will they be pushed to queens?
13:22:08 kashyap mdbooth: When you get a moment, hope I addressed all of your remarks here: "libvirt: Post-migration, set cache value for Cinder volume(s)" -- https://review.openstack.org/#/c/485752/
13:22:22 gibi mriedem: do we already have functional test that boot VMs which really uses the placement service. If there is such then I can try to put together the above suggested resize test
13:22:31 mriedem sean-k-mooney: i think the ovs offload from moshele is probably doable - i think i'd like a bp in nova for that though, just something simple to mirrow the neutron RFE
13:22:38 mriedem since nova doesn't do RFE bugs for blueprints
13:22:53 mriedem gibi: we do
13:23:09 mriedem gibi: the PlacementFixture is used in the _IntegratedHelpersMixin
13:23:13 mriedem which several functional tests use
13:23:26 gibi mriedem: cool, then I will put something together
13:23:35 mriedem there are other functional regression tests which don't use that mixin but still use the placement fixture to avoid warnings from the compute services in the logs
13:23:39 gibi mriedem: I anyhow wanted to test VM moving with placement
13:23:55 mriedem gibi: cool, that would be something that builds on https://review.openstack.org/#/c/483566/
13:24:07 mriedem the other tricky thing though is the test has to use the filter scheduler,
13:24:08 sean-k-mooney mriedem: ok good to know. i saw that netronomes patchs seem to have merged so im sure jangutter is happy with that it would be nice to finish netronomes version too. im sure it will make moshele equally happy

Earlier   Later