Earlier  
Posted Nick Remark
#openstack-nova - 2018-09-26
15:47:17 gibi jaypipes: you are right we are not catching AllocationDeleteFailed explicitly above in the call stack. Fortunately there are generic exception handling in place alreasy that puts the instance in ERROR state.
15:47:56 gibi jaypipes: the move operations are async on the API so when the fault happens there is no way to return that back to the API user anyhow
15:49:34 edmondsw efried we are not using zuulv3 in PowerVM CI yet, if that's what you were asking
15:49:51 efried dah, okay, thought we were
15:50:24 edmondsw last I heard, zuulv3 wasn't really ready for 3rd party CI usage yet
15:51:10 melwitt mriedem: I was looking at whether I should add the vmware live migration patch (in the queue) to a runway but saw it's failing vmware CI, and I see you've been discussing it with rado https://review.openstack.org/#/c/270116
15:53:05 mriedem i haven't looked at it since my last comments
15:53:48 mriedem it's also failing unit test
15:54:20 melwitt ok. I'll make a note next to it in the queue
15:57:42 openstack Launchpad bug 1789998 in OpenStack Compute (nova) "ResourceProviderAllocationRetrievalFailed ERROR log message on fresh n-cpu startup" [Low,Triaged]
15:57:42 mriedem a shiny donkey to whoever can bring me the head of https://bugs.launchpad.net/nova/+bug/1789998
15:58:13 efried F, I forgot *again* to collect my shiny nickel in Denver.
15:58:20 mriedem it's still in my backpack
15:58:53 efried That should probably be my bug. But I'm not likely to have time to look at it today.
15:59:55 efried also, /me wonders what "shiny donkey" means. Sounds like a euphemism for something.
16:00:00 efried Will it also fit in your backpack?
16:01:19 mdbooth mriedem: Speaking of common gate bugs: https://review.openstack.org/#/c/605436/
16:01:41 mriedem efried: https://www.youtube.com/watch?v=UNV44oqUF6k
16:01:56 mdbooth Although I didn't to a full test run on it locally first, so I won't be surprised if there's a kink to work out.
16:02:30 mriedem evacuate + affinity + locks = my head will explode
16:03:59 mdbooth mriedem: Add in a context manager which is a closure and some tail recursion ;)
16:05:13 melwitt I added cfriesen to the review
16:05:18 mdbooth cfriesen: https://review.openstack.org/#/c/605436/
16:05:32 mdbooth melwitt: Yeah, I was going to ping him earlier but he wasn't around
16:05:57 mdbooth I saw on the bug he looked at it before, and I assume there's some alternative solution in StarlingX
16:06:30 melwitt yeah
16:07:17 cfriesen for the "validate flavor extra-specs and image properties" work item, do we need a spec since it'll presumably result in an error message to the user? or are we allowed to return new error messages?
16:07:24 mdbooth Like I said it's not central to anything on my plate right now, though, so if somebody else wants to do a better job I'm cool with that. I probably won't spend a huge amount of time on it myself, though.
16:08:07 mdbooth I just fixed it because I saw it.
16:08:46 cfriesen mdbooth: taking a look
16:11:02 openstackgerrit Matthew Booth proposed openstack/nova master: Add volume-backed evacuate test https://review.openstack.org/604397
16:11:03 openstackgerrit Matthew Booth proposed openstack/nova master: Run evacuate tests with local/lvm and shared/rbd storage https://review.openstack.org/604400
16:11:04 cfriesen second question, for the "vcpu model extension" change where we'd allow specifying a list of CPU models in nova.conf instead of a single model, would we need a spec even though we're not changing the API?
16:11:32 mgariepy hello, I am upgrading from Pike to Queens but when running nova-manage db online_data_migrations, i get Some instances are still missing keypair information. Unable to run keypair migration at this time
16:14:13 mgariepy i found a few bug in lp concerning a workaround for kilo > liberty upgrade but the fix doesn't work for me as i don't have missing instance in my db.
16:14:16 openstack Launchpad bug 1684861 in OpenStack Compute (nova) newton "Mitaka -> Newton: Database online_data_migrations in newton fail due to missing keypairs" [Medium,In progress] - Assigned to Lee Yarwood (lyarwood)
16:14:16 mgariepy https://bugs.launchpad.net/nova/+bug/1684861
16:15:08 mgariepy I have 845 entry for select count(instance_uuid) from instance_extra where keypairs is NULL;
16:15:13 mdbooth mgariepy: See #topic. You should probably try #openstack instead
16:15:51 mgariepy well it's a nova issue.
16:16:12 mgariepy i've been upgrading to N o p q. and it fails a Q.
16:16:16 openstackgerrit Matthew Booth proposed openstack/nova master: Raise error on timeout in wait_for_versioned_notifications https://review.openstack.org/604859
16:16:17 openstackgerrit Matthew Booth proposed openstack/nova master: Don't delete disks on shared storage during evacuate https://review.openstack.org/578846
16:16:17 openstack bug 1550919 in OpenStack Compute (nova) "[Libvirt]Evacuate fail may cause disk image be deleted" [Medium,In progress] https://launchpad.net/bugs/1550919 - Assigned to Matthew Booth (mbooth-9)
16:16:17 openstackgerrit Matthew Booth proposed openstack/nova master: Add regression test for bug 1550919 https://review.openstack.org/591733
16:17:13 melwitt cfriesen: for extra spec and image properties validation, I think we would do a spec for it because it's an API change. for the cpu model list, from the ptg notes it looks like we thought we'd need a spec, maybe just to capture all of the related information. any opinion on either of these, mriedem?
16:17:31 mdbooth mgariepy: Indeed, but kilo and liberty are long out of support. Perhaps try your vendor?
16:18:42 mdbooth mriedem: We don't have any NFS CI jobs, do we?
16:18:44 mgariepy i'm upgrading from pike to queens.
16:19:11 mdbooth mriedem: I probably asked this before: my memory is terrible.
16:23:38 imacdonn mgariepy: you should at least try in #openstack ... "how do I...?" questions should start there. This channel is about development, not deployment .. if it's determined that there's a current bug, it could be brought here
16:24:34 mriedem mdbooth: we do,
16:24:37 mriedem it's in the experimental queue
16:24:47 mriedem melwitt: yes for api spec for extra spec validation in the api
16:24:59 mriedem as for cpu models stuff in nova.conf, idk, wasn't paying attention to that at the ptg
16:25:35 mriedem mdbooth: pro tip: look in nova's .zuul.yaml file
16:25:46 mriedem ye shall behold legacy-tempest-dsvm-full-devstack-plugin-nfs
16:26:01 cfriesen mdbooth: the starlingx server group validation stuff is here: https://github.com/starlingx-staging/stx-nova/blob/master/nova/compute/manager.py#L1408-L1436 and the check against "older" instances is here: https://github.com/starlingx-staging/stx-nova/blob/master/nova/objects/instance_group.py#L550-L572
16:26:30 cfriesen melwitt: okay, specs it is.
16:26:36 mriedem cfriesen: i think i might have mentioned this to you before, but you know how starlingx has a patched/upgraded flag it sets on the HostState object in the scheduler and then has a weigher for those?
16:26:45 melwitt thanks. cfriesen ^ you could try the cpu models as a specless bp and when we ask for approval during the nova meeting, someone might point out why it should be a spec, so you might have to write one at that point
16:26:49 mriedem i think the idea being, send new requests to patched/upgraded hosts?
16:27:19 mriedem cfriesen: any reason to not just check the compute's service version to see if it's the latest?
16:27:22 mriedem that would tell you if it's upgraded
16:28:11 melwitt mgariepy: you said earlier that you have not manually deleted any instances from the database?
16:28:14 cfriesen mriedem: one issue was that patching for a bugfix might not affect the service version
16:28:28 cfriesen mriedem: but that would work for the upgrade case
16:29:02 mgariepy melwitt, nop i didn't
16:29:11 mriedem weighing based on bug fix patches seems excessive
16:29:26 mgariepy i think the upgrade db didn't updated the deleted instances attrbutes.
16:29:35 mriedem but i guess i get it
16:30:32 cfriesen mriedem: so in the original model due to limitations we had to reboot the compute node when patching, so when rolling out a patch we really didn't want to have to migrate instances multiple times if we could avoid it. Probably less of an issue now.
16:31:16 mgariepy melwitt, updated the keypairs fields in the db..
16:34:33 imacdonn efried: ping me if you want to discuss https://review.openstack.org/#/c/605329/ - there's probably a sexier way to do it
16:35:15 mriedem imacdonn: gonna need tests
16:35:32 mriedem b/c clearly we weren't testing this before which is why it's a bug
16:35:41 imacdonn mriedem: yeah, I figured that'd come up ... but there aren't any tests at all for this stuff currently, or I failed to find them
16:36:28 imacdonn Trying to verify the contents of a PrettyTable didn't seem appealing :/
16:36:39 mriedem https://github.com/openstack/nova/blob/41ac87f812f239e38ad3b72c2f2108523ad3444d/nova/tests/unit/test_nova_manage.py#L749
16:37:08 mriedem there are tests for each individual migration method but you're right that they aren't tested together in aggregate
16:37:09 imacdonn oh. Why isn't that under "cmd" ?
16:37:16 mriedem b/c history
16:37:19 imacdonn mmkay
16:37:28 mriedem no reason we couldn't move it
16:37:35 mriedem except it makes backports a pain in the ass
16:37:37 melwitt mgariepy: oh, hm. ok, we have had online data migrations before that skip deleted records, so that sounds legit but I need to look and double check. it's surprising though that we haven't had this reported before, I'd think it would be common
16:38:10 imacdonn k, I'll look into why the test didn't catch the bug
16:38:40 imacdonn oh, I guess because it didn't have more than 50 of any migration type
16:40:44 melwitt tssurya, mriedem: I've got a call at the top of the hour, so I won't be able to participate much in a cells meeting if we have one
16:43:21 mgariepy melwitt, the cloud i'm upgarding lives from some times. first install was kilo and upgraded
16:43:43 mgariepy but depending on the when it's upgraded, it can cause inconsistency some times..
16:46:07 melwitt mgariepy: yeah, so I'm looking at this and the migration is that it's moving keypairs from the "main" database to the api database. and it's counting only the non-deleted InstanceExtra. InstanceExtra are marked as deleted when instances are deleted through the nova API https://github.com/openstack/nova/blob/stable/queens/nova/objects/keypair.py#L245
16:50:20 melwitt I don't yet understand how this works, I see the migration creates a keypair in the api db, deletes the keypair from the cell db, but I'm not yet seeing the connection with instance_extra
16:51:48 melwitt it looks like the instance_extra part is a separate migration that should have happened before this cell db => api db one that I'm looking at
16:52:43 melwitt ok yeah, instance_obj.migrate_instance_keypairs
16:52:49 melwitt from newton
16:53:21 mgariepy i'm currently using : nova_git_install_branch: d7864fbb9c2c558c409559e1d5989f84c7403832 # HEAD of "stable/queens" as of 23.07.2018
16:53:54 mgariepy melwitt, can you point me to the commit ?
16:54:51 melwitt mgariepy: here's the code https://github.com/openstack/nova/blob/stable/ocata/nova/objects/instance.py#L1398 that should have run when you upgraded to newton, when you ran the nova-manage db online_data_migrations. did you do that for newton?
16:55:53 melwitt the code no longer exists in queens, so if you missed running that online data migration in newton, you can't run it with the queens code

Earlier   Later