| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-19 | |||
| 15:06:09 | openstackgerrit | Brianna Poulos proposed openstack/nova master: Add trusted_certs to instance_extra https://review.openstack.org/457711 | |
| 15:06:09 | openstackgerrit | Brianna Poulos proposed openstack/nova master: Add trusted_certs to Instance object https://review.openstack.org/489408 | |
| 15:06:10 | openstackgerrit | Brianna Poulos proposed openstack/nova master: Add trusted_certificates to REST API https://review.openstack.org/486204 | |
| 15:06:42 | jianghuaw | sahid, I thought the problem to be resolved by this option is specific for vGPU. One pGPU can support different sized vGPUs. different capacity for different size. If created one size/type of vGPU on a pGPU, it can't create other size/type of vGPUs. | |
| 15:06:55 | jianghuaw | For other devices, is there similar issue? | |
| 15:11:10 | sahid | jianghuaw: hum... operators are going to creates some kind (size/type) vGPUs, but then that is not dynamic so the option is just about to parse which vCPUs we want to expose to Nova, no? | |
| 15:12:42 | jianghuaw | yes, the option is to make the vGPU type be static on one compute node. | |
| 15:12:50 | sahid | so even if a pGPU support different kinds, we want first to "allocate" the kind we want and bzsically create a pool of media which will be exposed to Nova | |
| 15:13:07 | sahid | jianghuaw: yes so basically it's what we have for sriov | |
| 15:15:04 | sahid | jianghuaw: just put your thinking on the review and let see what other contributors think... my point is just to implement one thing which could work for mdev/sriov | |
| 15:15:19 | jianghuaw | It will pre-restrict the type; and only create a resource provider for this specific type. | |
| 15:15:35 | mriedem | gibi: ack | |
| 15:16:03 | jianghuaw | Sure. if that's true a common thing, we can consider to make it general. Thanks sahid. | |
| 15:17:00 | openstackgerrit | Merged openstack/nova master: use unicode in tests to avoid SQLA warning https://review.openstack.org/505198 | |
| 15:21:48 | tasker | mriedem: it looks like https://bugs.launchpad.net/nova/pike/+bug/1717365 affects several different code paths in addition to just the one identified previously. | |
| 15:21:49 | openstack | Launchpad bug 1717365 in OpenStack Compute (nova) "binding:profile is None breaks migration" [Medium,In progress] - Assigned to Matt Riedemann (mriedem) | |
| 15:22:23 | mriedem | tasker: want to point those out in the review then? | |
| 15:22:30 | tasker | sure can. | |
| 15:22:54 | tasker | I'm going to be a bit blind ( not understanding all of nova ), but I'll identify all of the lines that could result in the same problem. | |
| 15:23:04 | mriedem | there was one other place i saw that it was not checked but i didn't think we could have a failure in the other location if we fixed the one spot you hit properly | |
| 15:23:55 | tasker | I just hit it in _update_port_binding_for_instance() during post-migration tasks. | |
| 15:26:20 | mriedem | tasker: here? https://review.openstack.org/#/c/504260/2/nova/network/neutronv2/api.py@2509 | |
| 15:29:47 | ratailor | melwitt, Could you please have a look on https://review.openstack.org/#/c/504885/ | |
| 15:31:57 | melwitt | ratailor: yeah, that's an old issue and I'm not sure where things landed as far as how to fix it. I don't remember us wanting to change the collation type of the column or if there was a reason we shouldn't. sdague might be able to recall past discussion on that | |
| 15:32:37 | ratailor | melwitt, Thanks, I will try to catch him tomorrow. | |
| 15:51:12 | gibi | mriedem: Is it OK to have the following small notification addition as a specless BP? https://blueprints.launchpad.net/nova/+spec/emit-service-create-notification | |
| 15:55:35 | gibi | cdent: have you seen my question in https://review.openstack.org/#/c/502155/5/nova/tests/functional/db/test_resource_provider.py@1522 ? I'm eager to approve your patch series | |
| 16:00:12 | tasker | mriedem: yes. | |
| 16:00:53 | ratailor | sdague, Could you please have a look on https://review.openstack.org/#/c/504885/ | |
| 16:03:10 | openstackgerrit | Merged openstack/nova stable/pike: Set error state after failed evacuation https://review.openstack.org/504979 | |
| 16:05:01 | openstackgerrit | Eric Fried proposed openstack/nova master: Nix warning about protocol-less glance api_servers https://review.openstack.org/505317 | |
| 16:05:04 | efried | sdague ^ | |
| 16:09:03 | tasker | mriedem: I'm having some browser troubles intercepting the "ctrl+s" to save the comment. is there another way to commit my comment? | |
| 16:10:26 | tasker | mriedem: I think that https://review.openstack.org/#/c/504260/2/nova/network/neutronv2/api.py@263 needs to be guarded as well, but I can't get the comment to apply on the page. | |
| 16:12:48 | efried | tasker What's your gerrit ID? | |
| 16:13:43 | tasker | efried: my launchpad ID is egrh3. is that the same thing? | |
| 16:15:06 | efried | tasker Was trying to find other things you've commented on in the past. Go to the front page of that change set, click the "Add Reviewer" button (little person icon to the right on the "Reviewers" line) and punch "Add Me". | |
| 16:16:15 | efried | tasker gotcha. So what's happening when you try to submit a comment? | |
| 16:17:02 | tasker | I highlight the line, press 'c', write my comment, press [save] or "ctrl+s" and the box collapses and says "draft". | |
| 16:17:08 | tasker | but I don't know what to do beyond that. | |
| 16:18:11 | efried | tasker Ah, okay. You go back to the front page of the change and hit the Reply button. Vote if you like (or not) and hit the Post button. | |
| 16:18:36 | efried | tasker That'll accumulate all comments you've made at once into a review note, and make them visible to others. | |
| 16:18:40 | tasker | ahhh | |
| 16:19:08 | tasker | got it! | |
| 16:19:20 | efried | Yup, I can see it. | |
| 16:23:06 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404 | |
| 16:27:41 | mriedem | gibi: i'm going to miss the notifications meeting today | |
| 16:27:45 | mriedem | so might as well cancel | |
| 16:34:05 | gibi | mriedem: thanks for the heads up. I will open it to see if somebody else will join or not but won't wait for long | |
| 16:34:12 | openstackgerrit | Elod Illes proposed openstack/nova master: Add instance.interface_attach notification https://review.openstack.org/503089 | |
| 16:37:28 | sdague | efried: for https://review.openstack.org/505317, can you remove the if condition too (the 2 lines before it) | |
| 16:37:49 | sdague | basically it's been deprecated since newton to have non urls in there | |
| 16:38:06 | efried | sdague Was wondering about that. So we just let it fail hard in the request if they've omitted the protocol? | |
| 16:39:08 | sdague | yes | |
| 16:39:21 | efried | sdague Roger wilco. | |
| 16:39:33 | sdague | if they don't provide the url then the whole uwsgi thing can't work | |
| 16:42:12 | openstackgerrit | Eric Fried proposed openstack/nova master: Don't fix protocol-less glance api_servers anymore https://review.openstack.org/505317 | |
| 16:42:17 | efried | sdague ^ Done. | |
| 16:43:13 | sdague | efried: ok, cool, last thing, I think we probably need a reno to tell people that they will hard fail if they only use IPs | |
| 16:43:40 | efried | sdague Okay. | |
| 16:49:56 | openstackgerrit | Eric Fried proposed openstack/nova master: Don't fix protocol-less glance api_servers anymore https://review.openstack.org/505317 | |
| 16:49:57 | efried | sdague ^ Not real experienced with release notes, that okay? | |
| 16:51:39 | sdague | efried: wfm | |
| 16:52:13 | dansmith | jaypipes: you around? | |
| 16:58:50 | openstackgerrit | Eric Fried proposed openstack/nova-specs master: Spec: Use keystoneauth1 Adapter for endpoints https://review.openstack.org/500190 | |
| 16:59:03 | efried | sdague ^ Added barbican. That sucker should be ready for review, I think. | |
| 16:59:17 | efried | mriedem FYI ^ | |
| 17:31:45 | cfriesen | anyone know if region names are case sensitive? | |
| 17:32:13 | openstackgerrit | Ed Leafe proposed openstack/nova-specs master: Re-propose nested resource providers spec https://review.openstack.org/505209 | |
| 17:32:45 | edleafe | cfriesen: I know that when I worked with RAX regions, they were | |
| 17:55:35 | tasker | the live-migration of the instance that was interrupted during the "post-migrtaion" tasks and the BINDING_PROFILE=None exception, the instance moved to the target host, but nova was interrupted and didn't update the database. nova thinks the instance is still on the original host. how can I get nova to update itself? | |
| 17:56:16 | tasker | restarting the `nova-compute` service did not do what I thought it might -- look to see what instances were active on the compute-host and update the tables accordingly. | |
| 17:58:09 | tasker | restarting results in "While synchronizing instance power states, found 2 instances in the database and 1 instances on the hypervisor." followed by "Instance is unexpectedly not found. Ignore." | |
| 17:58:43 | tasker | that's on the original host. the destination host has "While synchronizing instance power states, found 3 instances in the database and 4 instances on the hypervisor." and nothing else. | |
| 18:02:05 | tasker | i found https://bugs.launchpad.net/nova/+bug/1288958 on the same "Instance is unexpectedly not found. Igonre." message, but that one expired in April `16. the steps disclosed there are different but the state of Nova is similar. | |
| 18:02:10 | openstack | Launchpad bug 1288958 in OpenStack Compute (nova) "nova out of sync with the hypervisor " [Medium,Expired] | |
| 18:02:52 | tasker | nova list shows my particular instance with "power state = NOSTATE" | |
| 18:12:35 | melwitt | tasker: was it the 'setup_networks_on_host' that failed during post_live_migration_at_destination? because if so, that means the network failed setup on the destination and probably isn't working properly on the target | |
| 18:13:32 | melwitt | (I would think, I'm not a live migration expert) | |
| 18:20:11 | tasker | I don't see `setup_networks_on_host` in the traceback. it's "post_live_migration_at_destination -> migrate_instance_finish -> _update_port_binding_for_instance" | |
| 18:20:29 | tasker | do you have a pastie site that you prefer to use? I can flip you the traceback. | |
| 18:22:19 | melwitt | okay, I see the network_api.migrate_instance_finish call (just looking at the code). it's part of the network setup AFAICT | |
| 18:23:11 | melwitt | we use paste.openstack.org and pastebin.com | |
| 18:24:20 | melwitt | but just looking at the code, you're seeing it fail during the network setup at the destination, which would mean it's not going to work properly there, so it's not updating it to the target | |
| 18:25:57 | melwitt | yeah, the migrate_instance_finish is what updates the port binding for the instance (see nova/network/neutronv2/api.py) | |
| 18:27:46 | melwitt | if that failed, I don't think it would be correct to update the instance host to the target host. so AFAICT, the instance is still on the original host and there's a half-baked instance on the destination which needs to be cleaned up | |
| 18:28:16 | tasker | virsh list shows no instance on the original host. | |
| 18:28:49 | tasker | it does show it at the target host, but I cannot tell what state it's in. | |
| 18:28:52 | melwitt | *looking at more code* yeah, I was trying to see what has happened to the libvirt domain by this point | |
| 18:31:33 | tasker | melwitt: http://paste.openstack.org/show/621469/ | |
| 18:31:38 | melwitt | so it moved the domain, failed to update the port binding, so it's not in a working state. so far I'm not seeing how you could recover from that | |
| 18:32:13 | tasker | deletion is recovery, right? <G> | |
| 18:33:45 | melwitt | heh | |
| 18:34:06 | openstackgerrit | Claudiu Belu proposed openstack/nova master: hyperv: report disk_available_least field https://review.openstack.org/504904 | |
| 18:38:03 | tasker | this is a test cluster so no real worries. | |
| 18:42:00 | melwitt | okay, well that's good. the live migration flow is known to be gnarly and if anything fails in the middle of it, manual intervention is needed to fix things. I'm not familiar with more detail about it | |
| 18:42:35 | melwitt | from the trace you posted though, it looks like there's a bug in assuming binding_profile can't be None | |
| 18:43:18 | melwitt | er, actually it should be able to be assumed because: binding_profile = p.get(BINDING_PROFILE, {}) | |
| 18:45:26 | tasker | yeah, mriedem and I have addressed that with https://review.openstack.org/#/c/504260 | |
| 18:46:19 | melwitt | oh, I see. it's set but it's set to None. got it | |