Earlier  
Posted Nick Remark
#openstack-nova - 2017-12-19
00:27:02 rybridges also not working
00:27:06 rybridges now trying 3.7.0
00:27:09 rybridges and it seems to be working
00:27:19 rybridges i have suspended 40 instances without error
00:27:58 rybridges doh
00:28:01 rybridges take that back
00:28:06 rybridges tried 20 in parallel
00:28:09 rybridges still got a few errors
00:31:13 clarkb rybridges: libvirt-python is supposed to be compatible with any libvirt that is the same release as it or an older release. so libvirt-python 3.0 can tlak to libvirt 2.5 but libvirt-python 2.5 cn't talk to libvirt 3.0
00:31:21 clarkb this is why rhel 7.4 broke the 2.5.0 cap
00:31:26 clarkb (they did a major upgrade of libvirt
00:32:22 rybridges right
00:32:24 rybridges yes
00:32:34 rybridges 2.5 breaks on later versions of libvirt
00:33:01 rybridges so we had to switch up
01:31:09 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/528881
01:44:36 mriedem alex_xu: here is a question about something from long ago https://review.openstack.org/#/c/97727/
01:44:38 openstackgerrit OpenStack Proposal Bot proposed openstack/python-novaclient master: Updated from global requirements https://review.openstack.org/528911
01:44:59 mriedem alex_xu: why does populate_retry not check for MaxRetriesExceeded if max_attempts = 1?
01:45:23 mriedem i realize that means reschedules are disabled, but why wouldn't we compare num_attempts > max_attempts?
01:46:57 mriedem i guess that's what the code always did...
01:48:20 mriedem goes way back to https://review.openstack.org/#/c/9540/
01:51:21 mriedem oh nvm, i know why
01:51:39 mriedem if max_attempts == 1, we never set the retry key in the filter properties passed to compute
01:52:08 mriedem https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L1855
01:52:13 mriedem and then we don't reschedule
01:53:17 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/528881
01:53:44 openstackgerrit Matt Riedemann proposed openstack/nova master: Don't try to delete build requests on MaxRetriesExceeded https://review.openstack.org/528835
01:53:45 mriedem melwitt: %
01:53:46 mriedem ^
01:54:36 mriedem gah, that also goes back to newton
02:16:44 openstackgerrit Matt Riedemann proposed openstack/nova master: Don't try to delete build request during a reschedule https://review.openstack.org/528835
02:17:32 rybridges so I think i found the root problem with suspend
02:17:44 rybridges when i run suspend like this: openstack server suspend <uuid1> <uuid2> <uuid3> <uuid4> <uuid5>..... <uuid40>
02:17:51 rybridges almost all of the instances go to error state
02:17:53 rybridges but
02:18:30 rybridges when I run suspend in a simple for loop
02:18:33 rybridges like this:
02:18:43 rybridges for i in {1..20}
02:18:45 rybridges do
02:18:56 rybridges openstack server suspend ryan-rhel68-$i &
02:19:03 rybridges done
02:19:08 rybridges i get no errors
02:19:27 rybridges all of the instances go to suspended state (and NOT error state like the first command)
02:19:59 mriedem the compute api only takes a single instance for suspend,
02:20:03 mriedem so not sure what osc cli is doing
02:20:33 mriedem looks like it should be doing the same thing as you are, in a loop https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131
02:21:03 rybridges right
02:21:06 rybridges i was just looking at that
02:21:15 rybridges it looks like it should do essentially the same thing as the loop
02:21:17 rybridges but its not
02:21:26 mriedem what is the actual error in the nova logs?
02:21:27 rybridges because 80% of the instances go to error state
02:21:46 rybridges https://pastebin.com/jTedyZVJ
02:21:56 rybridges weird libvirt error
02:22:15 rybridges but i dont get that at all when i call suspend in a loop from a shell scrip
02:22:18 rybridges script*
02:22:33 mriedem huh, shouldn't make any difference
02:22:42 mriedem definitely looks like you're killing libvirt
02:22:56 mriedem seeing libvirt crash in the libvirtd logs or syslog?
02:23:07 rybridges i checked the libvirtd log
02:23:14 rybridges and did not see anything useful at all
02:23:20 rybridges not really any errors that seem meaningful
02:23:32 rybridges even if that was the case
02:23:46 mriedem i don't know why it would be any different
02:23:48 rybridges why would running the command in one way crash it and running the command in another way be just fine
02:23:49 mriedem either way you're running it
02:23:50 rybridges yea
02:24:00 mriedem unless there is some timing difference
02:24:00 rybridges it is though, i have 4 ocata clusters
02:24:05 rybridges all of them the behavior is like this
02:24:12 rybridges we have an ntp server
02:24:20 rybridges it also cant be timing
02:24:23 rybridges because if it was
02:24:31 rybridges it would be reproducible with both commands
02:24:33 rybridges right?
02:24:36 rybridges unless
02:24:46 rybridges one command is doing something different than the other
02:24:54 mriedem well,
02:24:57 rybridges do you know if that .suspend() call is asynch?
02:25:00 mriedem there is overhead to simply issuing an osc command
02:25:03 mriedem it is
02:25:20 mriedem https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131
02:25:22 mriedem oops
02:25:26 mriedem https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131
02:25:28 mriedem damn
02:25:35 mriedem anyway yeah it's an rpc cast from api to compute
02:25:45 rybridges right
02:25:50 mriedem so i'm wondering if your script is hitting the osc overhead just enough that each iteration is slow enough
02:26:02 rybridges hmm could be
02:26:05 mriedem but when doing them in batch via osc itself, it doesn't have the per-issue overhead
02:26:22 rybridges in theory, you would think that running the script would actually be calling that .suspend() method slower than passing all the uuids
02:26:26 mriedem try running both using timeit?
02:26:38 mriedem that's what i'm saying,
02:26:41 mriedem i think the script way is slower
02:26:49 mriedem and you're slowing it down, effectively load balancing :)
02:26:53 rybridges yea that makes sense
02:26:57 mriedem so you don't DoS libvirt
02:28:10 mriedem i didn't know osc actually let you specify a list of uuids to perform some action
02:28:28 rybridges well

Earlier   Later