Earlier  
Posted Nick Remark
#openstack-nova - 2017-12-19
20:18:50 openstack Launchpad bug 1381153 in OpenStack Compute (nova) "Cannot create instance live snapshots in Centos7 (icehouse)" [Undecided,Invalid]
20:28:19 cfriesen mriedem: thanks for the quota answer earlier...in objects.Quotas we're still looking at both the api DB and the main DB. is that left over from Pike? Presumably now we could remove the main DB access?
20:30:07 mriedem cfriesen: that's for the online data migration
20:30:17 alee hi - does anyone know how to force nova to re-fetch images from glance?
20:30:22 mriedem so we can't remove that until people have run through the online data migrations,
20:30:35 mriedem and we have a blocker schema migration in place to enforce that there are no quota limits/classes in the main cell db
20:30:35 cfriesen mriedem: that should have happened in Pike though, right? so we could remove it in Q?
20:30:49 mriedem cfriesen: you'd have to add a blocker migration
20:31:05 cfriesen mriedem: okay
21:04:45 openstackgerrit Ed Leafe proposed openstack/nova master: Make conductor pass and use host_lists https://review.openstack.org/511358
21:04:46 openstackgerrit Ed Leafe proposed openstack/nova master: Change compute RPC to use alternates for resize https://review.openstack.org/526436
21:04:52 edleafe mriedem: ^^ should address the functional test failures
21:06:11 ludovic Hi, I 'd like to ask few questions to nova experts: is it exists tips to prioritize evacuation order for instances
21:06:53 openstackgerrit Merged openstack/nova master: Some nit fix in multi_cell_list https://review.openstack.org/527597
21:07:15 openstackgerrit Merged openstack/nova master: doc: add note about fixing admin-only APIs without a microversion https://review.openstack.org/527421
21:07:33 ludovic or is it possible to reserve spare nodes ti ensure host-evacuation
21:08:15 mriedem ludovic: no and no,
21:08:24 mriedem you can specify a target host during evacuate, but that doesn't reserve it
21:08:35 mriedem and you evacuate one instance at a time,
21:08:42 mriedem so if there is priority, the caller handles that
21:10:56 ludovic ok, so we can't ensure a good SLA
21:11:37 ludovic of we can't reserve resources to securely evacuate all workload
21:11:54 ludovic the SLA is impacted
21:15:59 ludovic mriedem: does the scheduler propose as a trick to overcome this?
21:19:07 mriedem ludovic: how is this any different than the scheduler correctly picking a host during the initial server create?
21:19:28 mriedem besides reschedules
21:20:14 ludovic excuse-me , maybe can we have a private discussion about the Filter scheduler VS host-evacuate process with HA compute instances ?
21:20:28 mriedem why private?
21:21:01 ludovic in order to not disturb the room
21:21:11 ludovic np
21:21:48 ludovic I'm testing the HA compute instances with tripleo deployment
21:22:13 ludovic so the nova-evacuate pacemaker process is working well
21:22:35 ludovic but i was suprised by the Ram Filter scheduler
21:24:37 ludovic to explain the case , I have two compute Nodes. I have 2 Computes Nodes. The first compute Node with 31 instances ( total RAM allocated = 177152 M) , the second one with 2 instances ( Total RAM allocated = 24576 )
21:24:55 ludovic RAM of computes Nodes is 196483 (memory_mb)
21:26:07 ludovic The Host-evacuate worked but 5 VMs was on ERROR with insufficient memory (nova-compute log) and on was with NO STATE (No Host found by RAM Filter)
21:26:24 ludovic one
21:26:37 ludovic I expected the same ERROR on these 6 instances
21:27:04 ludovic HOST NOT FOUND or ERROR with insufficient memory
21:27:44 mriedem so you're evacuating the 31 instances on the one compute node to the other compute node with 2 instances?
21:27:48 ludovic So i don't understand the result
21:28:02 ludovic yes
21:28:08 mriedem and 5 of 31 fail
21:28:13 ludovic 6
21:28:14 mriedem which release?
21:28:40 ludovic 5 with insufficient memory and 1 with no valid host found
21:28:44 mriedem if you're in pike+ and using the FilterScheduler, you should remove the RamFilter
21:28:46 ludovic ocata
21:29:06 mriedem so the problem is,
21:29:22 mriedem the scheduler has a point in time snapshot of the resources from the compute, per request,
21:29:37 ludovic i use ocata for the moment
21:29:53 mriedem and the ram usage doesn't change until one of the evacuations makes it to the compute and claims those resources, and updates the compute node record in the db, which the scheduler will read in the next scheduling attempt
21:30:21 mriedem so the problem is if you're sending all 31 evacuate requests in a for loop, for example, with no time in between for the scheduler to catch up to the changes in the computes,
21:30:47 mriedem the scheduler thinks the compute is fine and sends the instance there, but the claim on the compute might fail because another request claimed those resources in the meantime
21:31:01 ludovic yes exactly
21:31:18 mriedem this should be fixed in pike,
21:31:35 mriedem because in pike, the RamFilter can be removed and the FilterScheduler uses the Placement service to claim the resources during scheduling
21:32:05 mriedem any claim collisions on the compute node in pike due to concurrent requests will be retried up to 3 times
21:32:35 mriedem otherwise what you're hitting is latent behavior
21:32:36 ludovic ah ok so with ocata the process is not enough efficient ?
21:32:59 mriedem correct; the late claim on the compute has always been a known issue with scheduling
21:33:10 mriedem with server create, if you hit this, we would reschedule to another compute
21:33:16 mriedem we don't do reschedules with evacuate though
21:33:45 ludovic But is it possible to influence the evacuate with max_concurrent_build ?
21:33:49 ludovic in nova.conf ?
21:34:27 ludovic by default this parameter is 10 , if we reduce to one ?
21:36:34 mriedem no max_concurrent_build is for server create, not evacuate
21:36:56 ludovic ah ok i understand .
21:36:58 mriedem there is no option like that for limiting evacuates
21:37:34 mriedem that's not to say one couldn't be added, but adding that in queens or pike doesn't make sense when we've solved this part of the problem in the scheduler
21:38:58 mriedem still, it seems reasonable to allow limiting the number of concurrent evacuates on a given compute, since we do that for spawn and live migrate
21:39:03 mriedem i wouldn't be opposed to adding something like that
21:40:54 ludovic ok thank you for your explanations, it's precious for me because I'm testing openstack for a big French Company you know
21:41:40 ludovic And the goal is to see if we can in the future replace the massive usage of WMare with OpenStack you know
21:41:59 mriedem oolala
21:42:00 ludovic near future
21:42:21 mriedem ok. would be cool if you could test this out on a pike deployment.
21:42:32 mriedem to make sure the filter scheduler + placement is correctly handling this for you
21:42:42 ludovic That's why i asked if reservation and /or prioritizing exist under OpenStack
21:42:56 mriedem remember to remove the RamFilter in pike if you're using the FilterScheduler since (1) it's redundant and (2) it will remove the memory_mb claim in the compute
21:43:39 ludovic ok
21:44:40 ludovic don't you think it will be interesting to add the possibility to have spare Compute Nodes ?
21:45:02 ludovic with aggregate host spare for example
21:45:10 mriedem you mean build something into nova to mark specific computes as only used for evacuate?
21:45:37 ludovic or in a aggregat to propose the prioritinzing of important workload when evacuate
21:45:46 ludovic yes for example
21:46:35 mriedem yeah idk, maybe. i wouldn't want to change the evacuate api to pass through scheduler hints probably.
21:46:36 ludovic That will be ensure a very good SLA because the evacuate processus would be securized
21:46:54 mriedem i don't know how many deployments just have compute nodes lying around as spares for evacuate
21:47:20 mriedem you can also control which host is used client-side
21:47:21 mriedem as noted
21:48:00 mriedem so as a client, if you have a special "evacuate" host aggregate, you could round robin through those hosts and send it with the evacuate request
21:48:38 mriedem but, as your evacuate aggregate starts to fill it, it is no longer spare capacity
21:48:39 mriedem so...
21:48:46 mriedem *fill up
21:49:55 openstackgerrit Ed Leafe proposed openstack/nova master: Make conductor pass and use host_lists https://review.openstack.org/511358
21:49:56 openstackgerrit Ed Leafe proposed openstack/nova master: Change compute RPC to use alternates for resize https://review.openstack.org/526436
21:50:01 edleafe ^^ fixed pep8 booboo
21:50:33 ludovic it seems to not be easy to design ...
21:50:49 ludovic evacuate is temporary

Earlier   Later