| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-12-19 | |||
| 01:51:39 | mriedem | if max_attempts == 1, we never set the retry key in the filter properties passed to compute | |
| 01:52:08 | mriedem | https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L1855 | |
| 01:52:13 | mriedem | and then we don't reschedule | |
| 01:53:17 | openstackgerrit | OpenStack Proposal Bot proposed openstack/nova master: Updated from global requirements https://review.openstack.org/528881 | |
| 01:53:44 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Don't try to delete build requests on MaxRetriesExceeded https://review.openstack.org/528835 | |
| 01:53:45 | mriedem | melwitt: % | |
| 01:53:46 | mriedem | ^ | |
| 01:54:36 | mriedem | gah, that also goes back to newton | |
| 02:16:44 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Don't try to delete build request during a reschedule https://review.openstack.org/528835 | |
| 02:17:32 | rybridges | so I think i found the root problem with suspend | |
| 02:17:44 | rybridges | when i run suspend like this: openstack server suspend <uuid1> <uuid2> <uuid3> <uuid4> <uuid5>..... <uuid40> | |
| 02:17:51 | rybridges | almost all of the instances go to error state | |
| 02:17:53 | rybridges | but | |
| 02:18:30 | rybridges | when I run suspend in a simple for loop | |
| 02:18:33 | rybridges | like this: | |
| 02:18:43 | rybridges | for i in {1..20} | |
| 02:18:45 | rybridges | do | |
| 02:18:56 | rybridges | openstack server suspend ryan-rhel68-$i & | |
| 02:19:03 | rybridges | done | |
| 02:19:08 | rybridges | i get no errors | |
| 02:19:27 | rybridges | all of the instances go to suspended state (and NOT error state like the first command) | |
| 02:19:59 | mriedem | the compute api only takes a single instance for suspend, | |
| 02:20:03 | mriedem | so not sure what osc cli is doing | |
| 02:20:33 | mriedem | looks like it should be doing the same thing as you are, in a loop https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131 | |
| 02:21:03 | rybridges | right | |
| 02:21:06 | rybridges | i was just looking at that | |
| 02:21:15 | rybridges | it looks like it should do essentially the same thing as the loop | |
| 02:21:17 | rybridges | but its not | |
| 02:21:26 | mriedem | what is the actual error in the nova logs? | |
| 02:21:27 | rybridges | because 80% of the instances go to error state | |
| 02:21:46 | rybridges | https://pastebin.com/jTedyZVJ | |
| 02:21:56 | rybridges | weird libvirt error | |
| 02:22:15 | rybridges | but i dont get that at all when i call suspend in a loop from a shell scrip | |
| 02:22:18 | rybridges | script* | |
| 02:22:33 | mriedem | huh, shouldn't make any difference | |
| 02:22:42 | mriedem | definitely looks like you're killing libvirt | |
| 02:22:56 | mriedem | seeing libvirt crash in the libvirtd logs or syslog? | |
| 02:23:07 | rybridges | i checked the libvirtd log | |
| 02:23:14 | rybridges | and did not see anything useful at all | |
| 02:23:20 | rybridges | not really any errors that seem meaningful | |
| 02:23:32 | rybridges | even if that was the case | |
| 02:23:46 | mriedem | i don't know why it would be any different | |
| 02:23:48 | rybridges | why would running the command in one way crash it and running the command in another way be just fine | |
| 02:23:49 | mriedem | either way you're running it | |
| 02:23:50 | rybridges | yea | |
| 02:24:00 | mriedem | unless there is some timing difference | |
| 02:24:00 | rybridges | it is though, i have 4 ocata clusters | |
| 02:24:05 | rybridges | all of them the behavior is like this | |
| 02:24:12 | rybridges | we have an ntp server | |
| 02:24:20 | rybridges | it also cant be timing | |
| 02:24:23 | rybridges | because if it was | |
| 02:24:31 | rybridges | it would be reproducible with both commands | |
| 02:24:33 | rybridges | right? | |
| 02:24:36 | rybridges | unless | |
| 02:24:46 | rybridges | one command is doing something different than the other | |
| 02:24:54 | mriedem | well, | |
| 02:24:57 | rybridges | do you know if that .suspend() call is asynch? | |
| 02:25:00 | mriedem | there is overhead to simply issuing an osc command | |
| 02:25:03 | mriedem | it is | |
| 02:25:20 | mriedem | https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131 | |
| 02:25:22 | mriedem | oops | |
| 02:25:26 | mriedem | https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131 | |
| 02:25:28 | mriedem | damn | |
| 02:25:35 | mriedem | anyway yeah it's an rpc cast from api to compute | |
| 02:25:45 | rybridges | right | |
| 02:25:50 | mriedem | so i'm wondering if your script is hitting the osc overhead just enough that each iteration is slow enough | |
| 02:26:02 | rybridges | hmm could be | |
| 02:26:05 | mriedem | but when doing them in batch via osc itself, it doesn't have the per-issue overhead | |
| 02:26:22 | rybridges | in theory, you would think that running the script would actually be calling that .suspend() method slower than passing all the uuids | |
| 02:26:26 | mriedem | try running both using timeit? | |
| 02:26:38 | mriedem | that's what i'm saying, | |
| 02:26:41 | mriedem | i think the script way is slower | |
| 02:26:49 | mriedem | and you're slowing it down, effectively load balancing :) | |
| 02:26:53 | rybridges | yea that makes sense | |
| 02:26:57 | mriedem | so you don't DoS libvirt | |
| 02:28:10 | mriedem | i didn't know osc actually let you specify a list of uuids to perform some action | |
| 02:28:28 | rybridges | well | |
| 02:28:31 | rybridges | it wasnt always like that | |
| 02:28:43 | rybridges | in juno we could not do that for the suspend command | |
| 02:28:55 | mriedem | yeah but now you guys are all upgraded to ocata | |
| 02:28:59 | mriedem | and have shiny new ways to kill yourselves | |
| 02:29:18 | rybridges | lololol | |
| 02:31:16 | lbragstad | mriedem: responded with more context/questions, hopefully it's clearer https://review.openstack.org/#/c/525772/1 | |
| 02:32:37 | mriedem | lbragstad: i think v1 of this thing needs to probably default to allowing whatever we support today, | |
| 02:32:42 | mriedem | which is admin == god | |
| 02:32:52 | mriedem | so in this thing, god == system scope | |
| 02:32:53 | mriedem | yes? | |
| 02:32:55 | lbragstad | so - ['system', 'project'] | |
| 02:33:00 | mriedem | yeah, | |
| 02:33:07 | lbragstad | because right now if you're admin you're god | |
| 02:33:18 | mriedem | and then for deployments that are doing a god -> project admin -> sheep setup, they can tweak their policy | |
| 02:33:20 | lbragstad | and can do anything everywhere | |
| 02:33:29 | mriedem | cburgess: ^ | |
| 02:33:43 | mriedem | cburgess would be a good person to ask because i think he's in the god role | |
| 02:33:56 | mriedem | i.e. the hosting company operator | |
| 02:34:00 | lbragstad | right | |
| 02:34:24 | lbragstad | so the big question is, how much power do i want to give customers without giving them the power to hose my deployment | |
| 02:37:20 | mriedem | today by default its all or none right? | |
| 02:37:23 | mriedem | admin or not admin | |
| 02:37:39 | lbragstad | pretty much | |