Earlier  
Posted Nick Remark
#openstack-nova - 2017-07-27
14:47:24 vdrok mriedem: yup, saw that. seems like I need to add the same variable to grenade settings
14:47:37 mriedem so, running list_cells can only happen on the primary, which is done before the subnodes, so list_cells is pretty useless there
14:47:59 mriedem what we'd really want is to call back into the primary from devstack-gate to run list_cells at the very end
14:48:09 mriedem but i'd make that a separate effort from what's being fixed in https://review.openstack.org/#/c/487809/
14:48:12 mriedem sdague: ^ agree?
14:49:29 sdague mriedem: why is list_cells useless there?
14:49:50 sdague the cells are all configured on the primary, right?
14:49:54 mriedem yeah, good point
14:50:02 sdague I do understand the issue around not running on subnode
14:50:42 mriedem ok so just conditional on n-api and we're good there
14:50:49 sdague though, it seemed to work fine
14:51:08 sdague the grenade failure for ironic is because they are using the old param to start_compute
14:51:24 sdague that got broken with our introduction of CELLSV2_SETUP
14:52:06 mriedem yeah vdrok just pushed a change for that
14:52:15 mriedem sdague: ok so you want to update the devstack patch or i can
14:52:27 sdague mriedem: so... this did pass on the subnodes
14:52:42 vdrok I'll remove the nomulticell flag that we did introduce in a later patch
14:52:42 mriedem huh?
14:52:44 mriedem http://logs.openstack.org/58/487458/3/check/gate-tempest-dsvm-ironic-ipa-wholedisk-agent_ipmitool-tinyipa-multinode-ubuntu-xenial/6cf92e0/logs/subnode-2/devstacklog.txt.gz#_2017-07-27_14_10_41_647
14:52:51 mriedem sdague: ^ is the stack on the subnode
14:53:56 sdague oh, sorry, it was an nv job
14:54:36 sdague mriedem: so, wrap it in n-api enabled?
14:55:08 mriedem sdague: i just updated it
14:55:47 sdague mriedem: ++
14:58:02 sdague mriedem / dansmith - https://review.openstack.org/#/c/487246/ my proposed solution on the nova-compute wait for ready
14:58:14 sdague it seems to have worked on the 3 node job correctly
15:02:21 mriedem oh that's devstack, i thought that was nova
15:02:35 openstackgerrit Takashi NATSUME proposed openstack/nova master: List/show all server migration types (1/2) https://review.openstack.org/430608
15:02:57 openstackgerrit Takashi NATSUME proposed openstack/nova master: List/show all server migration types (2/2) https://review.openstack.org/459483
15:03:00 sdague mriedem: right, it's a devstack change
15:03:21 sdague but trying to put it in devstack directly instead of more orchestration in d-g that gets messy
15:03:47 mriedem jaypipes: sdague: went through https://review.openstack.org/#/c/357726/ - i guess if you want to put it in then ok
15:03:53 mriedem we should fix the extension name in the reno
15:04:00 jaypipes k
15:04:09 mriedem it's one of those issues that we're not going to hear about for 18 months
15:04:17 mriedem when pike is the oldest stable branch
15:05:10 dansmith sdague: that won't reliably wait for the third node, right?
15:05:32 ys__ Hi, All. After I changed cpu_allocation_ratio in controller, Should I restart all nova services or just nova scheduler?
15:06:10 dansmith ys__: see topic please
15:06:32 mriedem ys__: that's used in both the scheduler and the compute service
15:07:26 mriedem actually the option is only used in the compute service
15:07:30 mriedem to update the compute node recored,
15:07:32 mriedem *record
15:07:38 mriedem which is used by the scheduler
15:09:03 sdague dansmith: yes, it will
15:09:14 dansmith sdague: how?
15:09:23 sdague each node is waiting for it's own hostname to show up
15:09:33 sdague stack.sh doesn't complete until it has
15:10:10 dansmith and something else waits for stack.sh on all the nodes before we run tempest/
15:10:37 sdague yes, stack.sh executions are linear
15:10:43 dansmith okay
15:10:50 sdague otherwise tempest would run before services were setup
15:13:50 mriedem dansmith: devstack-gate is waiting for the subnode stacks to be done
15:13:56 mriedem before calling discover_hosts
15:14:00 mriedem so yeah this looks ok
15:14:03 dansmith ack
15:14:07 mriedem http://logs.openstack.org/46/487246/2/experimental/gate-tempest-dsvm-neutron-dvr-ha-multinode-full-ubuntu-xenial-nv/6cd2a5b/logs/devstack-gate-discover-hosts.txt.gz
15:14:11 mriedem ^ is the 3 node job on that change
15:14:21 dansmith I hadn't scrolled right enough to see it was querying its own record, so assumed this was waiting for _a_ compute
15:15:06 mriedem sdague: dansmith: btw, would like to get rid of those ugly ass DEBUG outputs for oslo.concurrency from nova-manage https://review.openstack.org/#/c/487179/
15:15:22 mriedem ^ is probably backportable if you want me to open a bug
15:18:07 openstackgerrit Matt Riedemann proposed openstack/nova master: Add oslo_concurrency=INFO to default log levels for nova-manage https://review.openstack.org/487179
15:21:53 s-dean Hi, I got past the Cells issue yesterday and managed to list hypervisors and display nodes and services, then i cocked up on install neutron, and now im back to square one. cant list hypervisors
15:22:13 s-dean <class 'oslo_messaging.exceptions.MessagingTimeout'> __call__ /usr/lib/python2.7/dist-packages/nova/api/openstack/wsgi.py:1039
15:22:32 mriedem installing neutron shouldn't do anything with nova
15:22:47 s-dean i swear all services are connected to rabbit so why cant I list hypervisors and services
15:22:54 s-dean MessagingTimeout: Timed out waiting for a reply to message ID e65f06c471cc4e75858f936cf4dff041
15:23:06 s-dean is there a database that i can check to see the transport URL
15:23:31 s-dean i had to start a fresh again
15:23:46 mriedem s-dean: nova-manage cell_v2 list_cells --verbose
15:23:55 mriedem will dump the db and mq urls for the cell mappings
15:24:00 mriedem which should just be cell0 and cell1
15:24:04 s-dean ive ran that command
15:24:09 s-dean transport is fine
15:24:16 s-dean to my eye
15:24:28 s-dean hell conductor can even see my compute node
15:24:36 mriedem and the cell1 transport url is the same as the transport url in nova.conf?
15:24:43 s-dean yes
15:24:57 s-dean =INFO REPORT==== 27-Jul-2017::16:17:08 ===
15:24:57 s-dean Connection <0.1768.0> (10.30.0.2:32936 -> 10.30.0.2:5672) has a client-provided name: nova-api:4929:dec90db5-2cda-44ea-8ce7-e007fe5280a2
15:24:58 mriedem are the computes in the nova_api.host_mappings table?
15:25:08 s-dean not yet
15:25:18 s-dean tried to list the hypervisor
15:25:20 mriedem that's probably your problem
15:25:29 mriedem run: nova-manage cell_v2 discover_hosts
15:25:54 mriedem --verbose option on that too
15:26:09 s-dean Found 1 computes in cell: f2852cfb-ae7f-47d0-bef2-d00069bb57ca
15:26:46 s-dean and again ERROR 500 on nova compute api
15:26:57 s-dean just times out
15:27:10 s-dean i can see it connecting to rabbit
15:28:01 s-dean ive had this issue at least 4 times now
15:28:33 s-dean only yesterday, dont know what i did different but it was happy and let me list hypervisors and services
15:30:33 s-dean this is the rabbit mq driver
15:30:55 s-dean oslo_messaging _drivers amqpdriver.py
15:33:46 mriedem does the cell_mapping in the host_mappings table entries match cell1 in the cell_mappings table?
15:33:54 mriedem i.e. are the computes mapped to cell1 properly?
15:34:21 sdague mriedem: was just talking with dhellman, the sitemap fix isn't going to fix the 404s
15:34:30 sdague we're going to need to build a redirect list up ourselves
15:36:20 s-dean in that table, this is what i have | 2017-07-27 14:35:05 | NULL | 1 | 2 | compute01 |
15:36:22 s-dean | 2017-07-27 14:35:05 | NULL | 1 | 2 | compute01 |

Earlier   Later