| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-09-26 | |||
| 20:30:19 | melwitt | mriedem: https://bugs.launchpad.net/nova/+bug/1719730 | |
| 20:30:21 | openstack | Launchpad bug 1719730 in OpenStack Compute (nova) "Reschedule after the late affinity check fails with "'NoneType' object is not iterable"" [Undecided,New] | |
| 20:32:00 | efried | mriedem sdague https://review.openstack.org/#/c/488137/ should be ready again | |
| 20:32:49 | melwitt | heh | |
| 21:07:39 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Add recreate test for live migrate rollback not cleaning up dest allocs https://review.openstack.org/507677 | |
| 21:07:50 | mriedem | dansmith: ^ thus begins another round of these | |
| 21:11:50 | openstackgerrit | Eric Berglund proposed openstack/nova master: PowerVM Driver: config drive https://review.openstack.org/409404 | |
| 21:36:24 | pino | Hi Folks, I'm just getting started on a project that would provide an alternative to using key-pairs for instances: ssh certificates. This requires injecting into the instance (before startup) a host certificate, a user CA public key, and authorized principals file(s); then modifying sshd_config to use them. What's the right way to hook into the co | |
| 21:36:25 | pino | mpute instance lifecycle? | |
| 21:37:40 | pino | I'm just experimenting, but I'm wondering if this would be best built as part of Nova itself, or separately hook into the lifecycle. | |
| 21:38:34 | jaypipes | pino: definitely not part of Nova itself, no. apart from writing files to a config drive, Nova doesn't mess with the VM. | |
| 21:39:14 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove dest node allocations during live migration rollback https://review.openstack.org/507687 | |
| 21:39:47 | pino | jaypipes: ok, fair enough... but shouldn't support for ssh certificates be modelled similar to keypair support? | |
| 21:41:31 | jaypipes | pino: honestly, I'm not sure what the diff is between a key pair, with the private part of the pair downloaded to the user and the public part laid down on the VM config drive, and the SSH certificates thing you're describing. | |
| 21:41:44 | pino | And in terms of doing it outside of Nova, do you agree Nova notifications are not the right mechanism? The injection of various files, plus modification of the sshd_config must be done before first boot. Any advice about where, and how to do the hook? | |
| 21:41:47 | jaypipes | pino: I'm not an expert in ssh stuff, apologies. | |
| 21:43:00 | pino | jaypipes: I probably gave too much detail. I'm just looking for some hints about how I can hook into the startup workflow and block it until I've configured the VMs SSH the way I want it. | |
| 21:44:13 | jaypipes | pino: I think cloud-init is more what you are looking for? | |
| 21:45:22 | mriedem | pino: https://docs.openstack.org/nova/latest/user/vendordata.html | |
| 21:45:45 | mriedem | setup an external rest service that provides metadata to the guest when it's created | |
| 21:46:24 | mriedem | example https://github.com/openstack/novajoin | |
| 21:47:14 | penick | pino: I use SSH CA in my environment, maybe I can help? | |
| 21:47:21 | pino | mriedem: I saw that but wasn't sure it was the right approach. I'll take a closer look, thanks for the example. | |
| 21:48:02 | penick | I think I see what you're trying to do, and I think what you're probably going to want is to build a small webservice to create and sign SSH certificates, then tie that in with the nova vendordata stuff to get injected into the instance on boot | |
| 21:48:19 | mriedem | it's a wild penick | |
| 21:48:25 | mriedem | i wonder what the keyword is here | |
| 21:48:29 | penick | I identify as feral | |
| 21:48:47 | pino | jaypipes: I'm looking at cloud-init too... but I want my setup script to run without the user's help (they shouldn't have to do any setup). | |
| 21:49:11 | pino | penick: that makes perfect sense. | |
| 21:50:20 | pino | Ok, so I have a few topics/approaches to study. Thanks! | |
| 21:51:40 | penick | np :) | |
| 22:10:23 | openstackgerrit | Eric Fried proposed openstack/nova master: Use ksa adapter for keystone conf & requests https://review.openstack.org/507693 | |
| 22:26:36 | rybridges | Hey guys, I have a question. I am trying to inject some default user data into every instance while it is provisioning at this location -> https://github.com/openstack/nova/blob/stable/ocata/nova/compute/api.py#L1011 (I am adding an internal patch for this) In my patch, I create the user data, merge it with any existing user data on the instance, and then try to write the new user data to the | |
| 22:26:38 | rybridges | database. I am stuck getting it to write into the database. instance.save( | |
| 22:27:18 | rybridges | instance.save() is throwing stack traces. so I am trying to use the update_instance() method defined in the API class defined here https://github.com/openstack/nova/blob/stable/ocata/nova/compute/api.py#L2622 | |
| 22:27:31 | rybridges | but that does not seem to actually be saving the user data in the instance for some reason | |
| 22:27:51 | rybridges | it calls build_req.save() in the update_instance() method | |
| 22:32:42 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Remove dest node allocations during live migration rollback https://review.openstack.org/507687 | |
| 22:32:48 | melwitt | rybridges: doing it that way is a bad idea IMHO. if you're looking to have default data injected into every instance, you should look into the vendordata stuff that was linked earlier | |
| 22:36:44 | rybridges | so its not default perse, it will actually change based on some parameters. i just figured it would be easier for people to understand my problem if i said default | |
| 22:37:05 | rybridges | i dont like the vendordata stuff because it requires us to write an external webservice which complicates our deployment | |
| 22:37:19 | rybridges | i would rather just make a small patch which hits an entry point and injects the user data into the instance | |
| 22:37:25 | rybridges | it is much simpler and easier to debug/work with | |
| 22:38:14 | rybridges | i feel like it should not be this difficult.. | |
| 22:38:27 | rybridges | to just save the instance and get the user data written to the db | |
| 22:39:39 | mriedem | you know what's going to complicate your deployment? | |
| 22:39:53 | mriedem | constantly rebasing your fork, and when we change the internals that it depends on | |
| 22:40:50 | rybridges | we plan on adding a vendordata service eventually | |
| 22:41:08 | rybridges | also, the rebase is basically nothing | |
| 22:41:11 | rybridges | my patch is 3 lines | |
| 22:41:23 | melwitt | +1. I've been there before (patching nova) and would not recommend it | |
| 22:41:29 | rybridges | because i use an entry point that just passes locals to another function defined in a separate package | |
| 22:42:08 | rybridges | what i am trying to do is very simple | |
| 22:42:22 | rybridges | i just want to get a poc working right now | |
| 22:42:24 | mriedem | dansmith: i joked about this, but somone made it a reality http://forumtopics.openstack.org/cfp/details/6 | |
| 22:45:47 | mriedem | new devstack setup, still can't create 500 vms at once, they all go to NoValidHost, so maybe i didn't try to create this many at once yesterday - i wonder if i'm making placement bomb out, and the scheduler is rolling everything back | |
| 22:47:48 | rybridges | is user data immutable now? similar to how provision_updated_at is immutable in ironic | |
| 22:48:26 | openstackgerrit | Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947 | |
| 22:48:27 | openstackgerrit | Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950 | |
| 22:48:27 | openstackgerrit | Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948 | |
| 22:48:28 | openstackgerrit | Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419 | |
| 22:48:28 | openstackgerrit | Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949 | |
| 22:48:29 | openstackgerrit | Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638 | |
| 22:48:29 | openstackgerrit | Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420 | |
| 22:50:26 | melwitt | mriedem: did the scheduler logs offer any clues? | |
| 22:51:20 | mriedem | i think those have wrapped by now | |
| 22:51:36 | mriedem | i can create in chunks of 100 just fine | |
| 22:52:06 | melwitt | okay. was just curious | |
| 22:52:31 | mriedem | now i've got 100 ACTIVE instances, with 100 consumers in the api db and 300 allocations, | |
| 22:52:44 | mriedem | which makes sense b/c 1 cpu, 1 ram, 1 disk allocation per instance | |
| 22:52:49 | mriedem | 500 in the nova_cell0 db | |
| 22:53:48 | mriedem | aha | |
| 22:53:49 | mriedem | Sep 26 22:28:37 devstack nova-scheduler[2951]: WARNING nova.scheduler.client.report [None req-af92d5f2-4c99-4231-966e-939e1da04239 demo admin] Unable to submit allocation for instance 5f9f4f7d-8a2f-4fb8-b30a-024ed2e8e49d (409 {"errors": [{"status": 409, "request_id": "req-cda80554-6083-45b0-87bf-9e9c9924213f", "detail": "There was a conflict when trying to complete your request.\n\n Inventory changed while attempting to alloc | |
| 22:53:49 | mriedem | ubuntu@devstack:~$ sudo journalctl -a -u devstack@n-sch.service | grep Unable | |
| 22:53:50 | mriedem | ubuntu@devstack:~$ | |
| 22:53:50 | mriedem | Sep 26 22:28:37 devstack nova-scheduler[2951]: DEBUG nova.scheduler.filter_scheduler [None req-af92d5f2-4c99-4231-966e-939e1da04239 demo admin] Unable to successfully claim against any host. {{(pid=2951) _schedule /opt/stack/nova/nova/scheduler/filter_scheduler.py:221}} | |
| 22:53:50 | mriedem | Another thread concurrently updated the data. Please retry your update ", "title": "Conflict"}]}) | |
| 22:54:13 | mriedem | and then that removes all allocations for all instances | |
| 22:54:19 | mriedem | dansmith: melwitt: ^ | |
| 22:54:27 | mriedem | so yeah that's my failure here | |
| 22:55:05 | melwitt | um, so is that a new scheduling race condition that has to be resolved with reschedules? to replace the old claim race? | |
| 22:55:09 | mriedem | it does retry http://paste.openstack.org/show/621996/ | |
| 22:55:21 | mriedem | we do a retry in the scheduler | |
| 22:56:00 | melwitt | yeah, but I thought after "claims in the scheduler" we don't have concurrent request race problems that get kicked out to be retried | |
| 22:56:03 | mriedem | i don't know why that's logged 6 times | |
| 22:56:06 | dansmith | I do | |
| 22:56:23 | mriedem | heh | |
| 22:56:24 | dansmith | because you're scheduling so many things to one compute, and it only retries a certain number of times | |
| 22:56:25 | mriedem | do tell :) | |
| 22:56:59 | mriedem | i'm not surprised it's hitting a conflict | |
| 22:57:04 | mriedem | but why is that logged 3 times? | |
| 22:57:05 | dansmith | melwitt: we still have concurrent updates that we have to retry | |
| 22:57:07 | mriedem | 6 i mean | |
| 22:57:08 | mriedem | https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L1007 | |
| 22:57:12 | mriedem | because ^ we retry 3 times | |
| 22:57:37 | mriedem | you know, hitting a conflict that we have to retry once in 500 instances with a single compute, is pretty good | |
| 22:57:46 | mriedem | although i'm not sure which of the 500 this is that falied | |
| 22:57:54 | melwitt | so this is like the claim race except worse in that the things should have succeeded but can't. maybe it won't happen in real life because by the time that retry limit would be hit, the compute host would already be rejecting claims | |
| 22:58:00 | dansmith | well, actually.. are you doing one boot there or is there anything else going on? | |
| 22:58:18 | mriedem | single boot, --min-count 500 | |