Earlier  
Posted Nick Remark
#openstack-nova - 2019-01-29
13:02:23 kashyap gibi: Pointed out by jangutter earlier; where the 'run-devstack' task was taking 30 minutes
13:02:33 kashyap jaosorior: Hi, you can you ask, I can try :-)
13:03:18 jaosorior kashyap: I'm trying to reload libvirtd, given that it's certificate got rotated. in the man pages it says SIGHUP should do that... but it seems it doesn't... wanted to know if this is a bug, or if there's some other mechanism I could/should use instead.
13:06:54 kashyap jaosorior: Hmm, interesting problem
13:07:26 kashyap jaosorior: Can you afford to do an explicit stop+start of the daemon?
13:07:58 kashyap If so, does that pick up the new certs?
13:08:48 jaosorior kashyap: would that disrupt instances running on the host?
13:08:57 kashyap jaosorior: Good news: no.
13:09:07 kashyap jaosorior: Hang on, though.
13:09:13 jaosorior kashyap: what about connections from nova?
13:09:16 kashyap First, is that a production instance?
13:09:26 jaosorior yes
13:09:35 jaosorior well... not where I'm testing all this stuff :D
13:09:37 kashyap jaosorior: But know that: restarting `libvirtd` will _not_ bring your instances down.
13:09:40 jaosorior but that's the goal, yes
13:11:24 kashyap jaosorior: BTW, the SIGHUP config is ill-defined in the `libvirtd` man page :-(
13:11:39 kashyap It says: "On receipt of SIGHUP libvirtd will reload its configuration."
13:11:46 jaosorior right
13:11:50 kashyap But doesn't expand what is "its configuration".
13:12:02 kashyap I'll get that sorted w/ the libvirt upstream to expand and elaborate.
13:12:03 jaosorior TLS certs/keys are configuration IMO :D
13:12:59 kashyap jaosorior: Okay, so you're using TLS.
13:13:42 kashyap And you want to pick up renewed certificates. Then problem is not libvirtd, I'm afraid — QEMU is the real blocker
13:14:27 jaosorior what are my options?
13:14:31 kashyap Because it (QEMU) can't reload certificates on the fly. (And also QEMU can't just be "restarted")
13:14:51 kashyap Let me check with libvirt folks.
13:15:37 jaosorior kashyap: thanks!
13:16:20 kashyap jaosorior: So, your only viable option is a bit heavy: live migration of the QEMU instance(s).
13:16:21 kashyap (I.e. Nova instances)
13:16:48 ioni where the certs are used?
13:19:21 kashyap ioni: For TLS for various Nova services, I presume
13:19:47 kashyap jaosorior: Related, recently we merged this: http://git.openstack.org/cgit/openstack/nova/tree/doc/source/admin/secure-live-migration-with-qemu-native-tls.rst
13:20:04 kashyap There is a nicer rendered version
13:20:53 kashyap jaosorior: http://logs.openstack.org/83/631283/3/check/openstack-tox-docs/615979c/html/admin/secure-live-migration-with-qemu-native-tls.html
13:22:08 jaosorior kashyap: I'm a little confused though. the TLS implementation for connections to libvirtd is not the same one then as the one for qemu, right?
13:22:27 kashyap (But the above URL doesn't solve the existing problem, though. It is "future goodness", when you upgrade.)
13:23:20 kashyap jaosorior: Are you using the 'live_migration_tunnelled' Nova config?
13:24:05 kashyap jaosorior: They're somewhat connected: when you use TLS with libvirt, it is passing the TLS creds under the hood to QEMU.
13:24:18 jaosorior kashyap: we are, yeah
13:24:19 kashyap jaosorior: I can tell more details, if you post your QEMU command-line of your Nova instance (from the TLS env)
13:24:30 jaosorior kashyap: sure
13:24:55 kashyap I need to grab a quick bite (not going to eat in front of the computer, as a rule). Be back in 15 mins
13:25:13 jaosorior kashyap: sure. Good time for me to make some coffee
13:32:24 openstackgerrit Matt Riedemann proposed openstack/nova stable/rocky: Skip checking of target_dev for vhostuser https://review.openstack.org/633734
13:43:24 kashyap jaosorior: Let me know when you have the output.
13:50:20 kashyap jaosorior: But regardless, this requires support in QEMU and libvirt. I'll file RFEs w/ relevant upstreams
13:50:35 jaosorior kashyap: that would be great! (sorry, I'm in a meeting)
15:06:51 gibi jaypipes, mriedem: thanks for the reviews on the bandwidth series. I will propose the fup patch based on your comments
15:07:31 jaypipes gibi: ack. will try to get through remaining patches later.
15:08:13 gibi jaypipes: I saw your -1 in the middle. I will need to get back to dansmith about the second microversion
15:09:54 mriedem gibi: yw, glad to help move this along
15:10:20 mriedem also hopefully we can maybe merge some code yet this month https://review.openstack.org/#/c/633566/
15:15:09 openstackgerrit Surya Seetharaman proposed openstack/nova master: API microversion 2.68: Handles Down Cells https://review.openstack.org/591657
15:17:44 gibi mriedem: yeah, the gate is not on my side :)
15:22:29 openstackgerrit Matt Riedemann proposed openstack/nova master: Add configuration of maximum disk devices to attach https://review.openstack.org/616777
15:23:19 mriedem i'm +2 on ^ now
15:23:23 mriedem and the one below
15:24:33 jaypipes mriedem: amen on the merging code comment.
15:26:21 kashyap mriedem: I'm only 3 months late on your question here: https://review.openstack.org/#/c/506720/
15:26:39 kashyap (In short, I'm planning to revive that Secure Boot spec for Train.)
15:29:26 kashyap I posted a few interim updates there based on the completed work in OVMF, QEMU, libvirt & related tooling. I didn't post one since last August due to the libvirt bug still being in ASSIGNED state: https://bugzilla.redhat.com/show_bug.cgi?id=1605127
15:29:27 openstack bugzilla.redhat.com bug 1605127 in libvirt "RFE: QEMU firmware metadata format - libvirt support" [Unspecified,Assigned] - Assigned to phrdina
15:33:30 mriedem jangutter: i looked at the evacuate network-vif-plugged timeout errors in https://review.openstack.org/#/c/571325/ - tldr is it looks like we get the event before we're ready for it, which is weird
15:34:20 jangutter mriedem: think it's something transient or still surviving in master branches?
15:34:39 mriedem i assume it's transient otherwise that job would be 100% failure
15:34:59 mriedem would be useful if we logged when we register the event callbacks
15:37:04 mriedem oh we do
15:37:05 mriedem Jan 29 01:11:55.567189 ubuntu-xenial-rax-ord-0002201755 nova-compute[15903]: DEBUG nova.compute.manager [None req-252304a1-6eff-4ff3-aa4d-b4e0ab87601c demo admin] [instance: 2e82ddcd-75b8-4a41-8ecd-ce175adbdc67] Preparing to wait for external event network-vif-plugged-e241f79f-fb0d-4961-b0c8-aea9de2755bf {{(pid=15903) prepare_for_instance_event /opt/stack/new/nova/nova/compute/manager.py:325}}
15:38:35 mriedem which is too late
15:39:41 mriedem Jan 29 01:11:49.707004 ubuntu-xenial-rax-ord-0002201755 devstack@n-api.service[22319]: DEBUG nova.api.openstack.wsgi [req-3ed0ada1-7328-4d5a-a3ea-da34dcdb252d req-1f5aeede-83c6-44d5-afd4-b5435e71d61f service nova] Action: 'create', calling method: <bound method ServerExternalEventsController.create of <nova.api.openstack.compute.server_external_events.ServerExternalEventsController object at 0x7fb0133aa050>>, body: {"e
15:39:41 mriedem s": [{"status": "completed", "tag": "e241f79f-fb0d-4961-b0c8-aea9de2755bf", "name": "network-vif-plugged", "server_uuid": "2e82ddcd-75b8-4a41-8ecd-ce175adbdc67"}]} {{(pid=22322) _process_stack /opt/stack/new/nova/nova/api/openstack/wsgi.py:520}}
15:41:25 jangutter mriedem: nicely reversed causality there.
15:43:29 mriedem http://logstash.openstack.org/#dashboard/file/logstash.json?query=message%3A%5C%22Timeout%20waiting%20for%20%5B('network-vif-plugged'%5C%22%20AND%20message%3A%5C%22for%20instance%20with%20vm_state%20error%20and%20task_state%20rebuild_spawning.%5C%22%20AND%20tags%3A%5C%22screen-n-cpu.txt%5C%22&from=7d
15:46:36 mnaser is there any stable cores around to +w this? https://review.openstack.org/#/c/619352/
15:47:07 mriedem jangutter: https://bugs.launchpad.net/nova/+bug/1813789
15:47:08 openstack Launchpad bug 1813789 in OpenStack Compute (nova) "Evacuate test intermittently fails with network-vif-plugged timeout exception" [Medium,Confirmed]
16:08:10 mriedem jaypipes: since melwitt is out this week i replied on https://review.openstack.org/#/c/632904/ for that TooManyDiskDevices API error check
16:08:28 mriedem tl;dr there was already a lot of discussion between myself, melwitt, cdent and edleafe on that 403
16:09:16 mriedem i personally don't care if we change to 409, but i don't really think 400 is correct
16:11:57 cdent yeah, 409 is most correct, but 403 is consistent
16:17:40 efried mriedem: Where's the code that converts ComputeDriver.capabilities into traits on the compute node RP??
16:18:20 mriedem efried: https://review.openstack.org/#/c/538498/
16:18:28 openstackgerrit Balazs Gibizer proposed openstack/nova master: Fup for the bandwidth resource provider series https://review.openstack.org/633776
16:18:32 efried o
16:18:50 efried Why hasn't this merged yet <== because some a-hole -1'd it.
16:19:00 mriedem because i haven't updated it
16:19:05 gibi jaypipes, mriedem: here is the followup fixing comments in patches that are approved https://review.openstack.org/633776
16:19:26 mriedem gibi: ack
16:19:46 efried mriedem: Cool beans, just knew I had seen said code somewhere, but couldn't find it in the master branch. Thanks.
16:21:35 kashyap jangutter: Where do you see the 'run-devstack' taking 30 minutes here: http://logs.openstack.org/80/630980/7/gate/tempest-full/5b6ebba/ara-report/
16:22:12 kashyap jangutter: The one I found actually successfully "FINISHED"
16:22:28 kashyap (Although taking 30 mins & completing are not mutually exclusive.)
16:22:57 jangutter kashyap: in the devstack-tempest.yaml line, click on the "16 Tasks" button to pull them down. The run-devstack : Run devstack task went over 30 minutes.
16:23:20 kashyap jangutter: Right, there the status of 'run-devstack' task is "CHANGED".
16:23:29 jangutter kashyap: if you click on the parameters for the plays, you'll also see "timeout 30"
16:24:27 kashyap jangutter: Right, there I see it's 30 mins. And when I click on its 'Status' ("CHANGED" - what does it even mean), the job finished.
16:24:39 kashyap So I don't see any action item there.
16:24:42 jangutter kashyap: the playbook was interrupted, so the ARA results are inconsistent
16:24:47 kashyap Hmm
16:25:08 kashyap Just want to put this change out of its misery; it's been in time-out/recheck hell for 9 days

Earlier   Later