Earlier  
Posted Nick Remark
#openstack-nova - 2021-03-12
13:25:25 sean-k-mooney but unless we can come up with a better way then pass it to every function i would prefer not to
13:25:51 lyarwood I'll have to read the oslo.log code again to recall what it actually looks for
13:25:58 lyarwood it might look for other things aside from context
13:25:58 artom lyarwood, I thought the request ID was stashed in a thread global or something? At least that's what I was told...
13:26:19 sean-k-mooney artom: i dont think so
13:26:20 lyarwood hmm I thought oslo.log pulled it out of context tbh
13:26:27 sean-k-mooney that would not work well with eventlests
13:26:51 artom *shrug* I have to do the dad taxi thing anyways
13:26:55 sean-k-mooney unless you mean a green thread local variable of some kind
13:27:10 sean-k-mooney bu it i still would assume its not form that
13:27:21 lyarwood oh it's both
13:27:30 lyarwood https://github.com/openstack/oslo.log/blob/51324b276a8a5f69847d3a17fa89924dabad4759/oslo_log/formatters.py#L61-L78
13:29:10 lyarwood ah right cool
13:29:19 lyarwood okay so passing context isn't required
13:30:20 openstackgerrit Lee Yarwood proposed openstack/nova master: libvirt: Ensure all volume drivers log the instance whenever possible https://review.opendev.org/c/openstack/nova/+/780260
13:30:59 sean-k-mooney https://github.com/openstack/oslo.context/blob/0d02866365bc8b779aef9ebd0b79a52c96ae40e5/oslo_context/context.py#L508-L513
13:31:08 sean-k-mooney so it can retrun None
13:31:34 sean-k-mooney _request_store is a thread local
13:31:36 sean-k-mooney https://github.com/openstack/oslo.context/blob/0d02866365bc8b779aef9ebd0b79a52c96ae40e5/oslo_context/context.py#L40
13:32:10 sean-k-mooney so if the green thread resoumes on a different thread becaue we are usign a thread pool it might be empty
13:32:36 lyarwood kk makes sense, both connect and disconnect wait behind a lock
13:32:45 lyarwood so I assume that's where we end up losing it
13:32:49 lyarwood ew
13:33:07 sean-k-mooney do i want to know what you just saw
13:33:37 lyarwood that was more at the logging changing because of thread local storage
13:34:03 lyarwood I'm looking at yet another example failure of the detach thing
13:34:11 lyarwood another test attaches a volume
13:34:19 lyarwood the failing test attaches a volume
13:34:23 lyarwood the other test detaches
13:34:31 lyarwood and then our failing test attempts to detach
13:34:45 sean-k-mooney thats also ew :)
13:35:08 lyarwood well it's all locked and serialised so that's fine
13:35:18 lyarwood but the failing test just times out trying to detach the device
13:35:19 sean-k-mooney can we just chnage this locally in nova to workaround the issue
13:35:22 lyarwood there's some stuff in syslog
13:35:33 lyarwood but I can't get my head around it tbh
13:36:16 sean-k-mooney you still think this is due to conflicting WWNs right
13:37:14 lyarwood yeah I don't think that's helping but I can't find a smoking gun here
13:47:17 belmoreira hi, we are live migrating thousands of instances because hardware retirement and hit an old issue :) this was discussed in the past for instance creation, but now I'm facing it at live migration, when the migration allocation is created. https://bugs.launchpad.net/nova/+bug/1918419
13:47:19 openstack Launchpad bug 1918419 in OpenStack Compute (nova) "vCPU resource max_unit is hardcoded" [Undecided,New]
13:48:48 sean-k-mooney belmoreira: it would be incorrect to factor in allocation ration when setting max_unit
13:48:49 belmoreira I feel that for the vcpu particular case the operator could take the responsibility to defined the max_unit value.
13:49:27 sean-k-mooney i dont think i agree
13:49:40 sean-k-mooney you might be abel to use provider.yaml to cahnge it
13:49:53 sean-k-mooney but its not valid for a vm to oversubsribe against itself
13:50:06 sean-k-mooney so the max_unit should never exceed total
13:50:36 sean-k-mooney if you are disabling hyperthread thats havlfing your total cpus so the max_unit should also be reduced
13:50:48 belmoreira "its not valid for a vm to oversubsribe against itself" I agree with that
13:51:12 sean-k-mooney max_unit in placment is there to prevent that
13:51:42 sean-k-mooney belmoreira: the correct approch in your case would be to resize the vms before the migration
13:51:53 sean-k-mooney form the 32 core flavor to 16 i assume
13:52:03 sean-k-mooney or 32 but across 2 numa nodes
13:52:11 sean-k-mooney likely depending on the vm
13:52:50 belmoreira of course in the general I agree with you. And it's a good default. What I'm talking about is to give the operator the freedom to change this sensible default for very particular operations.
13:53:08 sean-k-mooney that the thing its all or nothing
13:53:23 sean-k-mooney if you change the max_unit in the placment inventory
13:53:29 sean-k-mooney it will apply to all operations
13:53:41 sean-k-mooney which could allow vms to boot that were over subscibing againt themselves
13:54:47 sean-k-mooney we dont suport setting max_unit for indivigual resouce in the allocation candiates request
13:55:33 sean-k-mooney belmoreira: if you realy need too you can use https://specs.openstack.org/openstack/nova-specs/specs/ussuri/approved/provider-config-file.html
13:55:43 sean-k-mooney to orverride the max_unit
13:55:58 sean-k-mooney assuming you have at least ussuri
13:57:27 belmoreira I'm not familiar with that spec. Let me have a look. (still in Stein for Nova)
13:57:31 sean-k-mooney you would basically do
13:57:33 sean-k-mooney meta:
13:57:35 sean-k-mooney schema_version: 1.0
13:57:37 sean-k-mooney providers:
13:57:39 sean-k-mooney # List of dicts
13:57:41 sean-k-mooney - identification:
13:57:43 sean-k-mooney uuid: $COMPUTE_NODE
13:57:45 sean-k-mooney inventories:
13:57:47 sean-k-mooney additional:
13:57:49 sean-k-mooney VCPU:
13:57:51 sean-k-mooney max_unit: 32
13:59:04 belmoreira If we can do that in Ussuri it looks good enough to me for any special case
13:59:23 belmoreira thanks sean-k-mooney
13:59:25 sean-k-mooney actully it wont work for your usecase
13:59:41 sean-k-mooney i forgot we can only use this for CUSTOM_ resouces
14:00:09 sean-k-mooney we intentionally blocked overwriding sthe standard ones form the virt driver
14:05:25 belmoreira my thinking on all of this is how an operator should move forward when (for whatever reason) the resource inventory changes.
14:06:12 sean-k-mooney well form a downstream and i thnk upstream point of vew we do not supprot changing the number of hyper treads on a host with vms
14:06:13 belmoreira For my particular case, vCPUs were cut in half, but we need to keep the oversubcribed instances running. Now to live migrate then we need "patch" the compute nodes
14:06:40 sean-k-mooney the same aslo gos for numa nodes e.g. by enable cluster on die ot changeing numa per socket in the bios
14:07:00 sean-k-mooney so today the only supported way it to do a resize to a differnt flavor
14:07:46 sean-k-mooney belmoreira: yep i understand unfortuetly the senario your are attempting to do is not currenlty supported by nova
14:07:52 belmoreira sean-k-mooney :) fair enough. I didn't think we would ever disable SMT in production nodes.
14:08:34 sean-k-mooney the "keep vms alive requirement is really the toughest"
14:09:16 sean-k-mooney belmoreira: for live migration we cant really change the toplogy of the guest
14:10:07 sean-k-mooney we can change teh mapping the host but putting that 32 core vm on a 16 core host would not be a valid schudlign decision based on our current oversubscpiton rules
14:10:33 sean-k-mooney belmoreira: actully i hate to say this but did you attempt a force live migration
14:10:57 sean-k-mooney i assume that still fails because placment will block it
14:11:11 sean-k-mooney when we try to update the allcoations
14:11:21 belmoreira sean-k-mooney it's fine. And in most part I agree with you, I just wanted to raise this issue here because other operators may suffer from the same thing. And these are the "use cases" that we didn't think about... Like I said I never thought we would need to disable SMT in production nodes.
14:12:01 sean-k-mooney it is certenly something we could put in the docs somewhwere
14:12:09 belmoreira sean-k-mooney actually I didn't... but in the case is placement, so it should fail
14:12:55 sean-k-mooney i think the allocation update will fail but i have never tried this
14:13:16 sean-k-mooney nova with the old microversion will skip the schduler fileter if you force it
14:13:27 sean-k-mooney but i think we always do the placment update
14:14:19 sean-k-mooney belmoreira: if you want to hack around it without changing the code there is one thing you could try
14:14:49 sean-k-mooney you could increase the interval taht we update placment at in the config on a node temporally so say once an hour

Earlier   Later