| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-12-19 | |||
| 02:23:14 | rybridges | and did not see anything useful at all | |
| 02:23:20 | rybridges | not really any errors that seem meaningful | |
| 02:23:32 | rybridges | even if that was the case | |
| 02:23:46 | mriedem | i don't know why it would be any different | |
| 02:23:48 | rybridges | why would running the command in one way crash it and running the command in another way be just fine | |
| 02:23:49 | mriedem | either way you're running it | |
| 02:23:50 | rybridges | yea | |
| 02:24:00 | mriedem | unless there is some timing difference | |
| 02:24:00 | rybridges | it is though, i have 4 ocata clusters | |
| 02:24:05 | rybridges | all of them the behavior is like this | |
| 02:24:12 | rybridges | we have an ntp server | |
| 02:24:20 | rybridges | it also cant be timing | |
| 02:24:23 | rybridges | because if it was | |
| 02:24:31 | rybridges | it would be reproducible with both commands | |
| 02:24:33 | rybridges | right? | |
| 02:24:36 | rybridges | unless | |
| 02:24:46 | rybridges | one command is doing something different than the other | |
| 02:24:54 | mriedem | well, | |
| 02:24:57 | rybridges | do you know if that .suspend() call is asynch? | |
| 02:25:00 | mriedem | there is overhead to simply issuing an osc command | |
| 02:25:03 | mriedem | it is | |
| 02:25:20 | mriedem | https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131 | |
| 02:25:22 | mriedem | oops | |
| 02:25:26 | mriedem | https://github.com/openstack/python-openstackclient/blob/master/openstackclient/compute/v2/server.py#L2131 | |
| 02:25:28 | mriedem | damn | |
| 02:25:35 | mriedem | anyway yeah it's an rpc cast from api to compute | |
| 02:25:45 | rybridges | right | |
| 02:25:50 | mriedem | so i'm wondering if your script is hitting the osc overhead just enough that each iteration is slow enough | |
| 02:26:02 | rybridges | hmm could be | |
| 02:26:05 | mriedem | but when doing them in batch via osc itself, it doesn't have the per-issue overhead | |
| 02:26:22 | rybridges | in theory, you would think that running the script would actually be calling that .suspend() method slower than passing all the uuids | |
| 02:26:26 | mriedem | try running both using timeit? | |
| 02:26:38 | mriedem | that's what i'm saying, | |
| 02:26:41 | mriedem | i think the script way is slower | |
| 02:26:49 | mriedem | and you're slowing it down, effectively load balancing :) | |
| 02:26:53 | rybridges | yea that makes sense | |
| 02:26:57 | mriedem | so you don't DoS libvirt | |
| 02:28:10 | mriedem | i didn't know osc actually let you specify a list of uuids to perform some action | |
| 02:28:28 | rybridges | well | |
| 02:28:31 | rybridges | it wasnt always like that | |
| 02:28:43 | rybridges | in juno we could not do that for the suspend command | |
| 02:28:55 | mriedem | yeah but now you guys are all upgraded to ocata | |
| 02:28:59 | mriedem | and have shiny new ways to kill yourselves | |
| 02:29:18 | rybridges | lololol | |
| 02:31:16 | lbragstad | mriedem: responded with more context/questions, hopefully it's clearer https://review.openstack.org/#/c/525772/1 | |
| 02:32:37 | mriedem | lbragstad: i think v1 of this thing needs to probably default to allowing whatever we support today, | |
| 02:32:42 | mriedem | which is admin == god | |
| 02:32:52 | mriedem | so in this thing, god == system scope | |
| 02:32:53 | mriedem | yes? | |
| 02:32:55 | lbragstad | so - ['system', 'project'] | |
| 02:33:00 | mriedem | yeah, | |
| 02:33:07 | lbragstad | because right now if you're admin you're god | |
| 02:33:18 | mriedem | and then for deployments that are doing a god -> project admin -> sheep setup, they can tweak their policy | |
| 02:33:20 | lbragstad | and can do anything everywhere | |
| 02:33:29 | mriedem | cburgess: ^ | |
| 02:33:43 | mriedem | cburgess would be a good person to ask because i think he's in the god role | |
| 02:33:56 | mriedem | i.e. the hosting company operator | |
| 02:34:00 | lbragstad | right | |
| 02:34:24 | lbragstad | so the big question is, how much power do i want to give customers without giving them the power to hose my deployment | |
| 02:37:20 | mriedem | today by default its all or none right? | |
| 02:37:23 | mriedem | admin or not admin | |
| 02:37:39 | lbragstad | pretty much | |
| 02:38:02 | mriedem | ok so i would think in queens, anything that's an admin rule by default today, would be system and project scopes | |
| 02:38:08 | mriedem | for compat | |
| 02:38:21 | mriedem | then over time you could start restricting the defaults from system to just project with release notes | |
| 02:38:35 | mriedem | these are just defaults in the code, and can be overridden | |
| 02:38:42 | lbragstad | so - i kinda tried to go about doing that here: https://review.openstack.org/#/c/528847/1 | |
| 02:38:57 | lbragstad | and i'd be super curious to get cburgess' feedback on that | |
| 02:39:13 | mriedem | oh so you have a global switch | |
| 02:39:22 | lbragstad | where an operator can go through and flip that switch once they have the right role infrastructure in place | |
| 02:39:31 | lbragstad | and they have audited their users to have the right roles | |
| 02:39:40 | rybridges | so the whole reason why i was asking about suspend originally is because snapshots were failing | |
| 02:39:55 | rybridges | and the snapshot flow (to my knowledge) is suspend > snapshot > resume | |
| 02:39:57 | lbragstad | (e.g. bob had the admin role but based on good faith, he didn't hose my deployment) | |
| 02:40:02 | rybridges | and it was always failing on suspend | |
| 02:40:22 | rybridges | and they still fail most of the time on suspend | |
| 02:40:25 | rybridges | with the same error above | |
| 02:40:42 | mriedem | rybridges: what libvirt calls suspend is likely != the compute api suspend | |
| 02:40:44 | rybridges | even though i cannot reproduce the error with suspending on the cli with the loop | |
| 02:41:31 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L1786 | |
| 02:41:48 | mriedem | https://github.com/openstack/nova/blob/master/nova/virt/libvirt/driver.py#L2686 | |
| 02:41:57 | mriedem | former is what libvirt calls on the guest during a snapshot | |
| 02:42:11 | mriedem | latter is what you get with 'openstack server suspend' | |
| 02:42:35 | rybridges | oh | |
| 02:42:38 | rybridges | ok that is interesting | |
| 02:43:00 | mriedem | oh jeez, nvm | |
| 02:43:01 | mriedem | self.suspend(context, instance) | |
| 02:43:03 | mriedem | derp | |
| 02:43:05 | mriedem | you're right | |
| 02:43:18 | mriedem | i was thinking of this https://github.com/openstack/nova/blob/master/nova/virt/libvirt/guest.py#L600 | |
| 02:43:54 | mriedem | rybridges: did you see where/why the snapshot was actually failing? have you tried doing live snapshots? | |
| 02:44:19 | rybridges | the snapshots are failing with the exact same error as i posted in the pastebin above | |
| 02:44:28 | mriedem | you might want to try live snapshot if libvirt / qemu on the host is new enough | |
| 02:44:36 | rybridges | it looks like it is just failing on the suspend | |
| 02:44:45 | mriedem | we don't call suspend if you do a live snapshot | |
| 02:44:57 | rybridges | we are running the latest libvirt / qemu that is available for rhel7 | |
| 02:45:06 | mriedem | which is what? | |
| 02:45:16 | mriedem | https://github.com/openstack/nova/blob/stable/ocata/nova/conf/workarounds.py#L68 | |
| 02:45:44 | rybridges | can you do live snapshot from horizon? | |
| 02:45:57 | mriedem | live vs cold is a config option in nova-compute in this case | |