| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-07-27 | |||
| 20:56:17 | sdague | so... actually, why isn't the host set that way on powervm and ironic? | |
| 20:56:32 | openstackgerrit | Ed Leafe proposed openstack/nova master: Handle hash ring rebalancing in ironic flavor migration https://review.openstack.org/487954 | |
| 20:56:48 | sdague | efried: you have a powervm setup somewhere that you can query? | |
| 20:57:01 | efried | esberglu needs to be involved here. | |
| 20:59:17 | edleafe | dansmith: ^^ incorporated rloo's suggestions | |
| 20:59:47 | efried | That powervm failure *might* be unrelated. We shouldn't be trying to connect to localhost. | |
| 21:00:01 | efried | sdague Did that change *when* the compute service gets started? | |
| 21:00:02 | sdague | mriedem: I'm actually not sure why hostname wouldn't match in the db | |
| 21:00:15 | sdague | efried: ?? | |
| 21:00:23 | mriedem | efried: no | |
| 21:00:27 | efried | sdague Yeah, I wouldn't have thought so. | |
| 21:00:29 | mriedem | efried: it's polling for the compute node to show up | |
| 21:00:31 | mriedem | by the hostname | |
| 21:00:35 | efried | So the net is, we're looking into it. | |
| 21:01:00 | sdague | efried: I'm ok with a revert atm because it broke ironic, and we had enough breaks on them this week | |
| 21:01:10 | sdague | but I am curious why those don't seem to line up | |
| 21:01:55 | thorst | I think the main thing for powervm is it shouldn't be taking that long to start up...so that's what we're looking into :-/ | |
| 21:03:02 | mriedem | oh right i forgot it takes 10 minutes for the powervm node to register | |
| 21:03:04 | mriedem | in init_host | |
| 21:03:46 | tonyb | mikal, sdague, melwitt: I don't knwo if this email was wider distrubuted but you know how we moved last_bytes recently .. it seems it was used by nova-lxd | |
| 21:04:15 | melwitt | I think I saw that email | |
| 21:04:17 | sdague | tonyb: they are out of tree, kind of don't care | |
| 21:04:31 | tonyb | mikal, sdague, melwitt: having the (out of tree) nova-lxd driver call into the libvirt code isn't cool :( so shoudl I revert it? | |
| 21:04:42 | tonyb | sdague: Well that was my initial response | |
| 21:04:53 | edmondsw | mriedem I started a change to get the powervm driver up faster but put it aside when we closed things down for pike | |
| 21:05:04 | tonyb | sdague: especially as when we get to queens we move it again | |
| 21:05:09 | sdague | tonyb: no, show up and interact in the community if you are using internal functions in the rest of the tree | |
| 21:05:10 | melwitt | tonyb: didn't they say they're already copy-pasting it somewhere? | |
| 21:05:27 | mriedem | edmondsw: i don't see how nova has control over how fast the backend node comes up | |
| 21:05:32 | tonyb | melwitt: not in the email I have but there may be more | |
| 21:05:40 | mriedem | edmondsw: you were just working on auto-enable the service i thought | |
| 21:05:52 | melwitt | okay, lemme check. maybe I misunderstood it | |
| 21:05:55 | edmondsw | mriedem https://review.openstack.org/#/c/471773/ | |
| 21:07:00 | edmondsw | should complete init_host much faster | |
| 21:07:25 | sdague | tonyb: I wasn't on any such email, but my patience is low for out of tree driver that's not working in the community | |
| 21:08:12 | melwitt | tonyb: okay, I just saw it was the last comment in the review https://review.openstack.org/#/c/472228/ and he's saying that they've now copy-pasted it, not that they had been until now. so I misread it | |
| 21:08:20 | tonyb | okay so we more or less want to say, "Sorry. this is part of the move to privsep so until xxx merged you'll just need to work around it in your driver' | |
| 21:09:01 | tonyb | melwitt: Thanks. | |
| 21:09:15 | tonyb | I'll contact them after breakfast. | |
| 21:09:30 | melwitt | yeah, if it's going to be available again in privsep after the work is done, that seems not so bad to me | |
| 21:09:31 | tonyb | I just wanted to make check we're on the same page | |
| 21:09:37 | tonyb | cool | |
| 21:13:35 | sdague | tonyb: yes, also, work upstream :P | |
| 21:23:25 | efried | sdague mriedem Okay, we've figured out how that change affected us. We need to do some stuff to the systemctl service file and restart compute. As of now, we're doing that after devstack finishes, not caring that compute wasn't coming up all the way during the stack. With this change, it started to matter that compute wasn't starting. | |
| 21:23:52 | efried | So we're going to figure out how to tweak the service file before stacking starts. Should make compute come right up during the stack process. | |
| 21:24:28 | mriedem | live migration job is done on my sanity check change http://logs.openstack.org/87/488187/1/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/482b6f4/ | |
| 21:27:20 | mriedem | oh i've got a bug in there | |
| 21:33:34 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Sanity check delete_allocation_for_instance https://review.openstack.org/488187 | |
| 21:35:45 | dansmith | jaypipes: https://bugs.launchpad.net/nova/+bug/1707071 | |
| 21:35:46 | openstack | Launchpad bug 1707071 in OpenStack Compute (nova) "Compute nodes will fight over allocations during migration" [Undecided,New] | |
| 21:35:52 | dansmith | dude how awesome is that bug number? | |
| 21:36:03 | dansmith | palindromic and all primes | |
| 21:37:49 | dansmith | well, I guess zero isn't a prime.. damn | |
| 21:40:37 | mriedem | hmm, so on rebuild to the same host | |
| 21:40:54 | mriedem | conductor calls select_destinations | |
| 21:41:11 | mriedem | https://github.com/openstack/nova/blob/master/nova/conductor/manager.py#L770 | |
| 21:41:23 | mriedem | will we double allocate that host then? | |
| 21:43:27 | mriedem | oh nvm | |
| 21:43:29 | mriedem | that's if not host | |
| 21:44:51 | mriedem | whew | |
| 21:47:44 | smcginnis | dansmith: You'll have to try to catch 1737371 | |
| 21:52:15 | dansmith | smcginnis: yeah | |
| 21:56:34 | melwitt | time to buy a lotto ticket | |
| 22:07:58 | mriedem | jaypipes: dansmith: ok, https://review.openstack.org/#/c/483566/ | |
| 22:08:06 | mriedem | there is one thing in there that worries me | |
| 22:08:19 | mriedem | https://review.openstack.org/#/c/483566/20/nova/scheduler/filter_scheduler.py@171 | |
| 22:09:31 | dansmith | hrm | |
| 22:10:13 | mriedem | oh L86 | |
| 22:10:13 | mriedem | if len(selected_hosts) < num_instances: | |
| 22:10:42 | dansmith | what happens if we return less than enough hosts to whatever calls us? | |
| 22:11:00 | dansmith | heh yeah that | |
| 22:11:30 | dansmith | and min/max_instances is only a quota check, so no problem there | |
| 22:12:29 | mriedem | ok +2 | |
| 22:12:35 | mriedem | i tried my hardest | |
| 22:12:39 | mriedem | to find fault | |
| 22:14:08 | dansmith | ack | |
| 22:14:31 | melwitt | mriedem: docstring doesn't match params here https://review.openstack.org/#/c/483566/20/nova/scheduler/filter_scheduler.py@252 if you wanted something :P | |
| 22:14:37 | mriedem | there might still be something to https://review.openstack.org/#/c/483566/20/nova/scheduler/filter_scheduler.py@208 | |
| 22:14:59 | mriedem | where if we know allocation requests are constantly failing for a host, we should stop trying it | |
| 22:15:17 | mriedem | omg | |
| 22:15:19 | mriedem | -10 | |
| 22:15:36 | melwitt | hehe | |
| 22:15:38 | mriedem | you know i think i was looking for where cn_uuid was used in there too | |
| 22:15:39 | mriedem | b/c of the docstring | |
| 22:18:53 | mriedem | pike-3 tag https://review.openstack.org/#/c/488218/ | |
| 22:21:09 | mriedem | jaypipes: thanks for hanging in there | |
| 22:21:24 | mriedem | ooo just in time as laura rolls back up to the house with the kid | |
| 22:21:37 | mriedem | time to get my county fair food eating clothes on | |
| 22:21:43 | dansmith | heh | |
| 22:21:52 | mriedem | over sized drawers, bib, etc | |
| 22:24:41 | mriedem | sdague: the compute host wait thing also blew up the cellsv1 job http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-cells-ubuntu-xenial/8ac25f1/logs/devstacklog.txt.gz | |
| 22:28:03 | mriedem | dansmith: jaypipes: here is the sanity check in action on the live migration job http://logs.openstack.org/87/488187/2/check/gate-tempest-dsvm-multinode-live-migration-ubuntu-xenial/107a810/logs/subnode-2/screen-n-cpu.txt.gz#_Jul_27_22_21_14_766229 | |
| 22:28:26 | mriedem | i didn't see any cases where the source node deleting the allocation isn't in the list of current allocations | |
| 22:28:36 | mriedem | but it's definitely stomping over the 'double up' allocatoins | |
| 22:29:04 | mriedem | Removing allocations for instance which are currently against more than one compute node resource provider. Current allocations: {u'13b1e5e0-66ef-4533-9a07-b1a3220d6b00': {u'generation': 8, u'resources': {u'VCPU': 1, u'MEMORY_MB': 64}}, u'7aa9619d-db83-4da9-b822-f4d66e7143f8': {u'generation': 6, u'resources': {u'VCPU': 1, u'MEMORY_MB': 64}}} | |
| 22:33:20 | dansmith | okay I think that's just getting lucky, | |
| 22:33:29 | dansmith | as I think it could either win or lose that race, but cool | |
| 22:37:45 | mriedem | oh no you didn't | |
| 22:37:45 | mriedem | https://www.youtube.com/watch?v=mBluR6cLxJ8 | |
| 22:44:32 | melwitt | dunno | |