Earlier  
Posted Nick Remark
#openstack-nova - 2017-09-26
21:37:40 pino I'm just experimenting, but I'm wondering if this would be best built as part of Nova itself, or separately hook into the lifecycle.
21:38:34 jaypipes pino: definitely not part of Nova itself, no. apart from writing files to a config drive, Nova doesn't mess with the VM.
21:39:14 openstackgerrit Matt Riedemann proposed openstack/nova master: Remove dest node allocations during live migration rollback https://review.openstack.org/507687
21:39:47 pino jaypipes: ok, fair enough... but shouldn't support for ssh certificates be modelled similar to keypair support?
21:41:31 jaypipes pino: honestly, I'm not sure what the diff is between a key pair, with the private part of the pair downloaded to the user and the public part laid down on the VM config drive, and the SSH certificates thing you're describing.
21:41:44 pino And in terms of doing it outside of Nova, do you agree Nova notifications are not the right mechanism? The injection of various files, plus modification of the sshd_config must be done before first boot. Any advice about where, and how to do the hook?
21:41:47 jaypipes pino: I'm not an expert in ssh stuff, apologies.
21:43:00 pino jaypipes: I probably gave too much detail. I'm just looking for some hints about how I can hook into the startup workflow and block it until I've configured the VMs SSH the way I want it.
21:44:13 jaypipes pino: I think cloud-init is more what you are looking for?
21:45:22 mriedem pino: https://docs.openstack.org/nova/latest/user/vendordata.html
21:45:45 mriedem setup an external rest service that provides metadata to the guest when it's created
21:46:24 mriedem example https://github.com/openstack/novajoin
21:47:14 penick pino: I use SSH CA in my environment, maybe I can help?
21:47:21 pino mriedem: I saw that but wasn't sure it was the right approach. I'll take a closer look, thanks for the example.
21:48:02 penick I think I see what you're trying to do, and I think what you're probably going to want is to build a small webservice to create and sign SSH certificates, then tie that in with the nova vendordata stuff to get injected into the instance on boot
21:48:19 mriedem it's a wild penick
21:48:25 mriedem i wonder what the keyword is here
21:48:29 penick I identify as feral
21:48:47 pino jaypipes: I'm looking at cloud-init too... but I want my setup script to run without the user's help (they shouldn't have to do any setup).
21:49:11 pino penick: that makes perfect sense.
21:50:20 pino Ok, so I have a few topics/approaches to study. Thanks!
21:51:40 penick np :)
22:10:23 openstackgerrit Eric Fried proposed openstack/nova master: Use ksa adapter for keystone conf & requests https://review.openstack.org/507693
22:26:36 rybridges Hey guys, I have a question. I am trying to inject some default user data into every instance while it is provisioning at this location -> https://github.com/openstack/nova/blob/stable/ocata/nova/compute/api.py#L1011 (I am adding an internal patch for this) In my patch, I create the user data, merge it with any existing user data on the instance, and then try to write the new user data to the
22:26:38 rybridges database. I am stuck getting it to write into the database. instance.save(
22:27:18 rybridges instance.save() is throwing stack traces. so I am trying to use the update_instance() method defined in the API class defined here https://github.com/openstack/nova/blob/stable/ocata/nova/compute/api.py#L2622
22:27:31 rybridges but that does not seem to actually be saving the user data in the instance for some reason
22:27:51 rybridges it calls build_req.save() in the update_instance() method
22:32:42 openstackgerrit Matt Riedemann proposed openstack/nova master: Remove dest node allocations during live migration rollback https://review.openstack.org/507687
22:32:48 melwitt rybridges: doing it that way is a bad idea IMHO. if you're looking to have default data injected into every instance, you should look into the vendordata stuff that was linked earlier
22:36:44 rybridges so its not default perse, it will actually change based on some parameters. i just figured it would be easier for people to understand my problem if i said default
22:37:05 rybridges i dont like the vendordata stuff because it requires us to write an external webservice which complicates our deployment
22:37:19 rybridges i would rather just make a small patch which hits an entry point and injects the user data into the instance
22:37:25 rybridges it is much simpler and easier to debug/work with
22:38:14 rybridges i feel like it should not be this difficult..
22:38:27 rybridges to just save the instance and get the user data written to the db
22:39:39 mriedem you know what's going to complicate your deployment?
22:39:53 mriedem constantly rebasing your fork, and when we change the internals that it depends on
22:40:50 rybridges we plan on adding a vendordata service eventually
22:41:08 rybridges also, the rebase is basically nothing
22:41:11 rybridges my patch is 3 lines
22:41:23 melwitt +1. I've been there before (patching nova) and would not recommend it
22:41:29 rybridges because i use an entry point that just passes locals to another function defined in a separate package
22:42:08 rybridges what i am trying to do is very simple
22:42:22 rybridges i just want to get a poc working right now
22:42:24 mriedem dansmith: i joked about this, but somone made it a reality http://forumtopics.openstack.org/cfp/details/6
22:45:47 mriedem new devstack setup, still can't create 500 vms at once, they all go to NoValidHost, so maybe i didn't try to create this many at once yesterday - i wonder if i'm making placement bomb out, and the scheduler is rolling everything back
22:47:48 rybridges is user data immutable now? similar to how provision_updated_at is immutable in ironic
22:48:26 openstackgerrit Dan Smith proposed openstack/nova master: Move allocation manipulation out of drop_move_claim() https://review.openstack.org/498947
22:48:27 openstackgerrit Dan Smith proposed openstack/nova master: Pre-create migration object https://review.openstack.org/498950
22:48:27 openstackgerrit Dan Smith proposed openstack/nova master: Make allocation cleanup honor new by-migration rules https://review.openstack.org/498948
22:48:28 openstackgerrit Dan Smith proposed openstack/nova master: Refactor resource tracker to account for migration allocations https://review.openstack.org/506419
22:48:28 openstackgerrit Dan Smith proposed openstack/nova master: Revert allocations by migration uuid https://review.openstack.org/498949
22:48:29 openstackgerrit Dan Smith proposed openstack/nova master: Make live migration hold resources with a migration allocation https://review.openstack.org/507638
22:48:29 openstackgerrit Dan Smith proposed openstack/nova master: Make migration uuid hold allocations for migrating instances https://review.openstack.org/506420
22:50:26 melwitt mriedem: did the scheduler logs offer any clues?
22:51:20 mriedem i think those have wrapped by now
22:51:36 mriedem i can create in chunks of 100 just fine
22:52:06 melwitt okay. was just curious
22:52:31 mriedem now i've got 100 ACTIVE instances, with 100 consumers in the api db and 300 allocations,
22:52:44 mriedem which makes sense b/c 1 cpu, 1 ram, 1 disk allocation per instance
22:52:49 mriedem 500 in the nova_cell0 db
22:53:48 mriedem aha
22:53:49 mriedem Sep 26 22:28:37 devstack nova-scheduler[2951]: WARNING nova.scheduler.client.report [None req-af92d5f2-4c99-4231-966e-939e1da04239 demo admin] Unable to submit allocation for instance 5f9f4f7d-8a2f-4fb8-b30a-024ed2e8e49d (409 {"errors": [{"status": 409, "request_id": "req-cda80554-6083-45b0-87bf-9e9c9924213f", "detail": "There was a conflict when trying to complete your request.\n\n Inventory changed while attempting to alloc
22:53:49 mriedem ubuntu@devstack:~$ sudo journalctl -a -u devstack@n-sch.service | grep Unable
22:53:50 mriedem ubuntu@devstack:~$
22:53:50 mriedem Sep 26 22:28:37 devstack nova-scheduler[2951]: DEBUG nova.scheduler.filter_scheduler [None req-af92d5f2-4c99-4231-966e-939e1da04239 demo admin] Unable to successfully claim against any host. {{(pid=2951) _schedule /opt/stack/nova/nova/scheduler/filter_scheduler.py:221}}
22:53:50 mriedem Another thread concurrently updated the data. Please retry your update ", "title": "Conflict"}]})
22:54:13 mriedem and then that removes all allocations for all instances
22:54:19 mriedem dansmith: melwitt: ^
22:54:27 mriedem so yeah that's my failure here
22:55:05 melwitt um, so is that a new scheduling race condition that has to be resolved with reschedules? to replace the old claim race?
22:55:09 mriedem it does retry http://paste.openstack.org/show/621996/
22:55:21 mriedem we do a retry in the scheduler
22:56:00 melwitt yeah, but I thought after "claims in the scheduler" we don't have concurrent request race problems that get kicked out to be retried
22:56:03 mriedem i don't know why that's logged 6 times
22:56:06 dansmith I do
22:56:23 mriedem heh
22:56:24 dansmith because you're scheduling so many things to one compute, and it only retries a certain number of times
22:56:25 mriedem do tell :)
22:56:59 mriedem i'm not surprised it's hitting a conflict
22:57:04 mriedem but why is that logged 3 times?
22:57:05 dansmith melwitt: we still have concurrent updates that we have to retry
22:57:07 mriedem 6 i mean
22:57:08 mriedem https://github.com/openstack/nova/blob/master/nova/scheduler/client/report.py#L1007
22:57:12 mriedem because ^ we retry 3 times
22:57:37 mriedem you know, hitting a conflict that we have to retry once in 500 instances with a single compute, is pretty good
22:57:46 mriedem although i'm not sure which of the 500 this is that falied
22:57:54 melwitt so this is like the claim race except worse in that the things should have succeeded but can't. maybe it won't happen in real life because by the time that retry limit would be hit, the compute host would already be rejecting claims
22:58:00 dansmith well, actually.. are you doing one boot there or is there anything else going on?
22:58:18 mriedem single boot, --min-count 500
22:58:37 mriedem only thing that could be changing inventory is the RT?
22:58:46 melwitt oh, okay so all or none. so nvm what I said
22:58:48 dansmith so we should be only making one call to scheduler I guess
23:00:20 mriedem hmm, so before claims in the scheduler, if you multi-create, wouldn't we only fail some of these after claims in the compute and reschedules?
23:00:31 mriedem so chances are you'd have some/most active, but others in error after reschedules?
23:00:46 mriedem now it's all or none
23:00:53 dansmith all or none for the claim process
23:00:56 dansmith you can still fail for other reasons
23:01:10 mriedem filters i suppose yeah

Earlier   Later