| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-09-26 | |||
| 15:42:44 | efried | "typically"? | |
| 15:42:50 | efried | not really, no. | |
| 15:42:59 | efried | I couldn't really even tell you how ours is configured :) | |
| 15:43:09 | efried | edmondsw: Any ideas ^ ? | |
| 15:43:49 | edmondsw | not sure what gate/test_evacuate.sh is | |
| 15:43:51 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add attach kwarg to base/nova-net allocate_for_instance methods https://review.openstack.org/605464 | |
| 15:43:56 | jaypipes | gibi: question for you on https://review.openstack.org/#/c/591811/. I'm sure I'm just missing something silly... | |
| 15:43:57 | mriedem | mdbooth: they wouldn't | |
| 15:43:59 | efried | mdbooth: It looks to me like this is only going to change the legacy nova-live-migration job. Does that even get triggered anymore? | |
| 15:44:06 | mdbooth | edmondsw: It's new in https://review.openstack.org/#/c/602174/ | |
| 15:44:07 | mriedem | efried: yes | |
| 15:44:08 | gibi | jaypipes: looking | |
| 15:44:16 | mriedem | devstack-gate can run post-test hook scripts | |
| 15:44:39 | mdbooth | Basically if it's 'opt-in' then failing if libvirt isn't configured is the correct behaviour | |
| 15:44:40 | mriedem | some 3rd party CI still uses devstack-gate, some are moving to zuul v3 which doesn't use devstack-gate (unless you define a legacy-style job, like nova-live-migration) | |
| 15:44:48 | mriedem | it's definitely opt-in | |
| 15:44:53 | mdbooth | If we always run it we'd probably want it to just skip | |
| 15:44:59 | mdbooth | mriedem: Thanks | |
| 15:45:52 | edmondsw | mdbooth our CI is currently using devstack all-in-ones for each run, so it can't do things like evacuate that require multiple nodes yet... working on that | |
| 15:46:10 | mriedem | tempest doesn't test evacuate anyway | |
| 15:46:20 | efried | and we're using zuulv3, right? | |
| 15:46:23 | mriedem | that's why this in a separate script, when tempest isn't running | |
| 15:46:29 | openstackgerrit | Merged openstack/nova stable/queens: Follow devstack-plugin-ceph job rename https://review.openstack.org/602019 | |
| 15:46:36 | openstackgerrit | Merged openstack/nova stable/queens: nova-status - don't count deleted compute_nodes https://review.openstack.org/604786 | |
| 15:47:17 | gibi | jaypipes: you are right we are not catching AllocationDeleteFailed explicitly above in the call stack. Fortunately there are generic exception handling in place alreasy that puts the instance in ERROR state. | |
| 15:47:56 | gibi | jaypipes: the move operations are async on the API so when the fault happens there is no way to return that back to the API user anyhow | |
| 15:49:34 | edmondsw | efried we are not using zuulv3 in PowerVM CI yet, if that's what you were asking | |
| 15:49:51 | efried | dah, okay, thought we were | |
| 15:50:24 | edmondsw | last I heard, zuulv3 wasn't really ready for 3rd party CI usage yet | |
| 15:51:10 | melwitt | mriedem: I was looking at whether I should add the vmware live migration patch (in the queue) to a runway but saw it's failing vmware CI, and I see you've been discussing it with rado https://review.openstack.org/#/c/270116 | |
| 15:53:05 | mriedem | i haven't looked at it since my last comments | |
| 15:53:48 | mriedem | it's also failing unit test | |
| 15:54:20 | melwitt | ok. I'll make a note next to it in the queue | |
| 15:57:42 | mriedem | a shiny donkey to whoever can bring me the head of https://bugs.launchpad.net/nova/+bug/1789998 | |
| 15:57:42 | openstack | Launchpad bug 1789998 in OpenStack Compute (nova) "ResourceProviderAllocationRetrievalFailed ERROR log message on fresh n-cpu startup" [Low,Triaged] | |
| 15:58:13 | efried | F, I forgot *again* to collect my shiny nickel in Denver. | |
| 15:58:20 | mriedem | it's still in my backpack | |
| 15:58:53 | efried | That should probably be my bug. But I'm not likely to have time to look at it today. | |
| 15:59:55 | efried | also, /me wonders what "shiny donkey" means. Sounds like a euphemism for something. | |
| 16:00:00 | efried | Will it also fit in your backpack? | |
| 16:01:19 | mdbooth | mriedem: Speaking of common gate bugs: https://review.openstack.org/#/c/605436/ | |
| 16:01:41 | mriedem | efried: https://www.youtube.com/watch?v=UNV44oqUF6k | |
| 16:01:56 | mdbooth | Although I didn't to a full test run on it locally first, so I won't be surprised if there's a kink to work out. | |
| 16:02:30 | mriedem | evacuate + affinity + locks = my head will explode | |
| 16:03:59 | mdbooth | mriedem: Add in a context manager which is a closure and some tail recursion ;) | |
| 16:05:13 | melwitt | I added cfriesen to the review | |
| 16:05:18 | mdbooth | cfriesen: https://review.openstack.org/#/c/605436/ | |
| 16:05:32 | mdbooth | melwitt: Yeah, I was going to ping him earlier but he wasn't around | |
| 16:05:57 | mdbooth | I saw on the bug he looked at it before, and I assume there's some alternative solution in StarlingX | |
| 16:06:30 | melwitt | yeah | |
| 16:07:17 | cfriesen | for the "validate flavor extra-specs and image properties" work item, do we need a spec since it'll presumably result in an error message to the user? or are we allowed to return new error messages? | |
| 16:07:24 | mdbooth | Like I said it's not central to anything on my plate right now, though, so if somebody else wants to do a better job I'm cool with that. I probably won't spend a huge amount of time on it myself, though. | |
| 16:08:07 | mdbooth | I just fixed it because I saw it. | |
| 16:08:46 | cfriesen | mdbooth: taking a look | |
| 16:11:02 | openstackgerrit | Matthew Booth proposed openstack/nova master: Add volume-backed evacuate test https://review.openstack.org/604397 | |
| 16:11:03 | openstackgerrit | Matthew Booth proposed openstack/nova master: Run evacuate tests with local/lvm and shared/rbd storage https://review.openstack.org/604400 | |
| 16:11:04 | cfriesen | second question, for the "vcpu model extension" change where we'd allow specifying a list of CPU models in nova.conf instead of a single model, would we need a spec even though we're not changing the API? | |
| 16:11:32 | mgariepy | hello, I am upgrading from Pike to Queens but when running nova-manage db online_data_migrations, i get Some instances are still missing keypair information. Unable to run keypair migration at this time | |
| 16:14:13 | mgariepy | i found a few bug in lp concerning a workaround for kilo > liberty upgrade but the fix doesn't work for me as i don't have missing instance in my db. | |
| 16:14:16 | mgariepy | https://bugs.launchpad.net/nova/+bug/1684861 | |
| 16:14:16 | openstack | Launchpad bug 1684861 in OpenStack Compute (nova) newton "Mitaka -> Newton: Database online_data_migrations in newton fail due to missing keypairs" [Medium,In progress] - Assigned to Lee Yarwood (lyarwood) | |
| 16:15:08 | mgariepy | I have 845 entry for select count(instance_uuid) from instance_extra where keypairs is NULL; | |
| 16:15:13 | mdbooth | mgariepy: See #topic. You should probably try #openstack instead | |
| 16:15:51 | mgariepy | well it's a nova issue. | |
| 16:16:12 | mgariepy | i've been upgrading to N o p q. and it fails a Q. | |
| 16:16:16 | openstackgerrit | Matthew Booth proposed openstack/nova master: Raise error on timeout in wait_for_versioned_notifications https://review.openstack.org/604859 | |
| 16:16:17 | openstackgerrit | Matthew Booth proposed openstack/nova master: Add regression test for bug 1550919 https://review.openstack.org/591733 | |
| 16:16:17 | openstack | bug 1550919 in OpenStack Compute (nova) "[Libvirt]Evacuate fail may cause disk image be deleted" [Medium,In progress] https://launchpad.net/bugs/1550919 - Assigned to Matthew Booth (mbooth-9) | |
| 16:16:17 | openstackgerrit | Matthew Booth proposed openstack/nova master: Don't delete disks on shared storage during evacuate https://review.openstack.org/578846 | |
| 16:17:13 | melwitt | cfriesen: for extra spec and image properties validation, I think we would do a spec for it because it's an API change. for the cpu model list, from the ptg notes it looks like we thought we'd need a spec, maybe just to capture all of the related information. any opinion on either of these, mriedem? | |
| 16:17:31 | mdbooth | mgariepy: Indeed, but kilo and liberty are long out of support. Perhaps try your vendor? | |
| 16:18:42 | mdbooth | mriedem: We don't have any NFS CI jobs, do we? | |
| 16:18:44 | mgariepy | i'm upgrading from pike to queens. | |
| 16:19:11 | mdbooth | mriedem: I probably asked this before: my memory is terrible. | |
| 16:23:38 | imacdonn | mgariepy: you should at least try in #openstack ... "how do I...?" questions should start there. This channel is about development, not deployment .. if it's determined that there's a current bug, it could be brought here | |
| 16:24:34 | mriedem | mdbooth: we do, | |
| 16:24:37 | mriedem | it's in the experimental queue | |
| 16:24:47 | mriedem | melwitt: yes for api spec for extra spec validation in the api | |
| 16:24:59 | mriedem | as for cpu models stuff in nova.conf, idk, wasn't paying attention to that at the ptg | |
| 16:25:35 | mriedem | mdbooth: pro tip: look in nova's .zuul.yaml file | |
| 16:25:46 | mriedem | ye shall behold legacy-tempest-dsvm-full-devstack-plugin-nfs | |
| 16:26:01 | cfriesen | mdbooth: the starlingx server group validation stuff is here: https://github.com/starlingx-staging/stx-nova/blob/master/nova/compute/manager.py#L1408-L1436 and the check against "older" instances is here: https://github.com/starlingx-staging/stx-nova/blob/master/nova/objects/instance_group.py#L550-L572 | |
| 16:26:30 | cfriesen | melwitt: okay, specs it is. | |
| 16:26:36 | mriedem | cfriesen: i think i might have mentioned this to you before, but you know how starlingx has a patched/upgraded flag it sets on the HostState object in the scheduler and then has a weigher for those? | |
| 16:26:45 | melwitt | thanks. cfriesen ^ you could try the cpu models as a specless bp and when we ask for approval during the nova meeting, someone might point out why it should be a spec, so you might have to write one at that point | |
| 16:26:49 | mriedem | i think the idea being, send new requests to patched/upgraded hosts? | |
| 16:27:19 | mriedem | cfriesen: any reason to not just check the compute's service version to see if it's the latest? | |
| 16:27:22 | mriedem | that would tell you if it's upgraded | |
| 16:28:11 | melwitt | mgariepy: you said earlier that you have not manually deleted any instances from the database? | |
| 16:28:14 | cfriesen | mriedem: one issue was that patching for a bugfix might not affect the service version | |
| 16:28:28 | cfriesen | mriedem: but that would work for the upgrade case | |
| 16:29:02 | mgariepy | melwitt, nop i didn't | |
| 16:29:11 | mriedem | weighing based on bug fix patches seems excessive | |
| 16:29:26 | mgariepy | i think the upgrade db didn't updated the deleted instances attrbutes. | |
| 16:29:35 | mriedem | but i guess i get it | |
| 16:30:32 | cfriesen | mriedem: so in the original model due to limitations we had to reboot the compute node when patching, so when rolling out a patch we really didn't want to have to migrate instances multiple times if we could avoid it. Probably less of an issue now. | |
| 16:31:16 | mgariepy | melwitt, updated the keypairs fields in the db.. | |
| 16:34:33 | imacdonn | efried: ping me if you want to discuss https://review.openstack.org/#/c/605329/ - there's probably a sexier way to do it | |
| 16:35:15 | mriedem | imacdonn: gonna need tests | |
| 16:35:32 | mriedem | b/c clearly we weren't testing this before which is why it's a bug | |