| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-21 | |||
| 16:25:14 | efried | sdague Without a session param, it's pretty lightweight: https://github.com/openstack/os-service-types/blob/master/os_service_types/service_types.py#L51 | |
| 16:25:15 | mordred | sdague: yes- os-service-types is - os-service-types didn't make it in to the keystoneauth release for pike, so we need to add that now that queens is open | |
| 16:25:37 | efried | Almost all of the work is done at import time, actually. | |
| 16:25:46 | efried | https://github.com/openstack/os-service-types/blob/master/os_service_types/service_types.py#L22 | |
| 16:26:06 | mordred | efried: yah, I think it's actually fine for this cycle - the contents really do not change frequently - and CERTAINLY not for the services that nova cares about | |
| 16:26:19 | efried | The race would be harmless. I can take the lock out. | |
| 16:26:20 | mordred | so we can get live-update landed as a follow on | |
| 16:26:24 | sdague | um... https://github.com/openstack/os-service-types/blob/master/os_service_types/service_types.py#L59 ? closet network call? | |
| 16:26:53 | edmondsw | efried sdague mordred to be clear, long-term I think we'll need auth conf options for everything. But we don't today | |
| 16:26:56 | sdague | that seems like something people should have to opt into | |
| 16:26:58 | mordred | sdague: yah - it's there - but the keystoneauth consumption of the library at least at first will not pass a session | |
| 16:27:01 | mordred | sdague: yup | |
| 16:27:36 | efried | sdague The opt-in is by passing a `session` param. | |
| 16:27:43 | mordred | sdague: so when we land the next patch to ksa to consume that, nova can just switch to always passing the correct/official type to keystoneauth and keystoneauth will dtrt | |
| 16:28:16 | mordred | we'll make sure subsequent turning on of remote access/ network calls is appropriately opt-in when we add it | |
| 16:28:30 | sdague | I'm actually not super clear why the remote part is there | |
| 16:28:48 | efried | to get a fresh copy of the service-types-authority data | |
| 16:29:05 | efried | os-service-types ships with a cached copy | |
| 16:29:18 | openstackgerrit | Merged openstack/os-vif master: Update reno for stable/pike https://review.openstack.org/488671 | |
| 16:29:19 | sdague | I thought the point of this was a no requirements version which we rev every time there is a service types update | |
| 16:29:29 | sdague | so people just replace it with the new one | |
| 16:29:42 | efried | Yup. Unless they don't. | |
| 16:29:50 | sdague | if they don't, then they don't | |
| 16:30:05 | openstackgerrit | Merged openstack/os-traits master: Update reno for stable/pike https://review.openstack.org/488669 | |
| 16:30:09 | sdague | closet network calls that could end up with very interesting results or random network hangs | |
| 16:30:15 | sdague | don't seem useful | |
| 16:30:46 | sdague | like https://github.com/openstack/os-service-types/blob/master/os_service_types/service_types.py#L59 under some circumstances can hang forever | |
| 16:31:06 | mordred | sdague: yah - I don't actually think it's useful for nova to opt-in to the fetch behavior | |
| 16:31:53 | efried | Hence the comment in the nova change | |
| 16:32:03 | sdague | mordred: sure, but I'm not super clear where it's a good idea to put a potentially arbitrary delay into anyone's code path | |
| 16:32:35 | sdague | I get that it's clever, but the potential failure domains get huge | |
| 16:32:57 | openstackgerrit | Eric Fried proposed openstack/nova master: nova.utils.get_ksa_adapter() https://review.openstack.org/488137 | |
| 16:33:07 | efried | sdague mordred ^ updated to remove lock | |
| 16:33:08 | mordred | sdague: totally. this is why it will always be opt-in | |
| 16:33:09 | openstackgerrit | Merged openstack/nova master: Restore '[vnc] vnc_*' option support https://review.openstack.org/505831 | |
| 16:33:58 | sdague | mordred: ok. I'm still not sure why it's useful, but I've said my piece :) | |
| 16:34:07 | mordred | efried: yah- I think honestly you could just initialize the _SERVICE_TYPES at the top and have it done at import | |
| 16:34:38 | efried | mordred Considering that ost loads the local file at import time (which I didn't realize) I think you're right. | |
| 16:34:38 | sdague | having spent a couple of months discovering requests can hang forever on simple get calls depending on what's happening on the network, which was locking up 1/3 of my smart home devices, I'm super twitchy on it | |
| 16:36:03 | openstackgerrit | Eric Fried proposed openstack/nova master: nova.utils.get_ksa_adapter() https://review.openstack.org/488137 | |
| 16:36:10 | efried | mordred Did that ^ | |
| 16:37:57 | sdague | efried: +2 | |
| 16:38:03 | efried | thanks! | |
| 16:38:26 | openstackgerrit | Sean Dague proposed openstack/nova master: Support qemu >= 2.10 https://review.openstack.org/505673 | |
| 16:39:19 | mriedem | efried: stephenfin: you guys did that fancy scheduler call flow diagram. i will pay a shiny nickel to whoever can make this live migration call flow into a docs diagram https://photos.app.goo.gl/Q8JdpjM0PZhAzsv32 | |
| 16:39:51 | sdague | cburgess: https://review.openstack.org/#/c/505673 - live snapshot by default | |
| 16:41:40 | dansmith | mriedem: lol | |
| 16:41:48 | dansmith | how ... analog of you | |
| 16:42:03 | mriedem | dansmith: i think i'm just going to slap that into our official docs | |
| 16:42:17 | mriedem | "Technical reference deep dive > live migration > THIS" | |
| 16:42:27 | dansmith | heh | |
| 16:43:05 | openstackgerrit | Elod Illes proposed openstack/nova master: Add instance.interface_detach notification https://review.openstack.org/506284 | |
| 16:43:49 | mriedem | if only i were on the twitters | |
| 16:44:17 | sdague | mriedem: that is a solvable problem | |
| 16:45:18 | efried | mriedem What does "cast" mean? | |
| 16:45:44 | efried | async invocation? | |
| 16:46:28 | mriedem | rpc cast | |
| 16:46:29 | mriedem | vs call | |
| 16:46:42 | mriedem | which is important to understand the hot potato between the computes during live migration | |
| 16:46:51 | mriedem | especially if you love rpc timeouts | |
| 16:46:59 | mriedem | because your instance has 20 ports and 20 volumes attached to it | |
| 16:47:09 | mriedem | and token timeouts | |
| 16:47:33 | dansmith | mmmm, rpc timeout | |
| 16:49:07 | sdague | mriedem: service tokens fix that | |
| 16:49:22 | sdague | well, some of it | |
| 16:50:54 | mriedem | yeah i know | |
| 16:50:59 | mriedem | wonder if anyone is using those yet | |
| 17:04:25 | bauzas | dansmith: mriedem: question, when I was testing https://review.openstack.org/#/c/506093/2 I discovered a weird stack http://paste.openstack.org/show/621641/ | |
| 17:05:35 | bauzas | dansmith: mriedem: when looking at https://github.com/openstack/nova/blob/master/nova/compute/resource_tracker.py#L1227 it lookups the in-memory dict of all compute nodes for finding the right node | |
| 17:06:59 | bauzas | dansmith: mriedem: but AFAICS, we're passing the source node as attribute to the destination RT https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5735 | |
| 17:07:22 | bauzas | dansmith: mriedem: so since the destination RT doesn't know the source node, it fails with a KeyError | |
| 17:07:26 | bauzas | amirite? | |
| 17:07:44 | bauzas | unless post_live_mig runs on the source node | |
| 17:07:52 | bauzas | that's confusing | |
| 17:07:54 | mriedem | post live migrate runs on the source onde | |
| 17:07:56 | mriedem | *node | |
| 17:08:11 | mriedem | see my awesome call flow diagram above | |
| 17:08:21 | mriedem | gd this is already paying for itself | |
| 17:08:27 | mriedem | https://photos.app.goo.gl/Q8JdpjM0PZhAzsv32 | |
| 17:08:28 | bauzas | grmblbl | |
| 17:08:36 | openstackgerrit | Dan Smith proposed openstack/nova master: Fix CellDatabases fixture swallowing exceptions https://review.openstack.org/506312 | |
| 17:09:08 | bauzas | mriedem: any reason you would see why I'm getting that KeyError ? I guess it's because I don't run the periodics ? | |
| 17:09:36 | openstackgerrit | Merged openstack/nova master: Updated from global requirements https://review.openstack.org/505839 | |
| 17:10:19 | openstackgerrit | Merged openstack/nova master: Fix hyperlinks in document https://review.openstack.org/506101 | |
| 17:12:25 | mriedem | bauzas: that should get set in memory when the compute service starts up | |
| 17:12:29 | mriedem | when it calls pre_start_hook | |
| 17:12:46 | bauzas | mmm | |
| 17:12:50 | mriedem | that will call update_available_resource_for_node | |
| 17:13:01 | bauzas | looking at the func tests, jay needed to explicitly start them | |
| 17:13:09 | mriedem | start what? | |
| 17:13:10 | mriedem | the computes? | |
| 17:13:14 | bauzas | the RT update things | |
| 17:13:14 | jaypipes | hmm? | |
| 17:13:14 | mriedem | the fixture starts those | |
| 17:13:43 | bauzas | here, I'm getting a trace because self.compute_nodes[my_node] isn't set yet | |
| 17:13:44 | mriedem | when the compute service starts, it will get the available nodes from the driver, and use those to call rt.update_available_resource | |
| 17:13:54 | mriedem | passing in the node name which gets populated in the rt.compute_nodes dict | |
| 17:13:55 | bauzas | which is populated AFAIK by running the RT update call | |
| 17:15:23 | bauzas | mmm, I'm introspecting | |
| 17:15:32 | mriedem | what you have in the test setup looks fine to me | |