Earlier  
Posted Nick Remark
#openstack-nova - 2018-03-22
19:35:07 dansmith sdeath: nope, and I feel your pain, but still.. normally we run people out of here _before_ answering their questions
19:35:13 sean-k-mooney jaypipes: similar feeling...
19:35:26 cfriesen sean-k-mooney: the numatopologyfilter would still catch stuff like that I think
19:35:27 jaypipes sean-k-mooney: no, in all seriousness, it doesn't impact the accounting of CPU resources.
19:35:38 jaypipes cfriesen: yes, it would indeed.
19:35:57 sean-k-mooney jaypipes: yes that is true i shoudl be placement
19:36:24 sean-k-mooney it does not care becase it doesn't impact the accounting of CPU resources. numa topology filter might
19:37:21 sdeath dansmith: then thanks for the assistance! there any objection to lurking?
19:37:37 dansmith sdeath: nope
19:37:46 sean-k-mooney jaypipes: so jay for the snowflake. it lands on a host with HT off and boots fine. it lands on a host with HT on. what happens. retry?
19:37:49 melwitt sdeath: fwiw the openstack@lists.openstack.org is better than the channel for getting responses about deployment/config issues, IMHO
19:38:01 melwitt mailing list
19:38:14 cfriesen jaypipes: so if we ignore ht in placement we don't get the atomicity of reservation benefits, but it's no worse than it is currently. we might fail ResourceTracker.instance_claim() or rebuild_claim() or resize_claim() once we get to the compute node.
19:38:20 jaypipes sean-k-mooney: NUMATopologyFilter has no impact on resource accounting whatsoever, for two reasons: a) the results of numa_fit_instance_to_host() are immediately discarded and b) it looks at *assignment*, not allocation.
19:38:29 melwitt gah, more POST_FAILURE on zuul >:|
19:38:48 sdeath melwitt: will give that one a shot if/when needed... thanks!
19:39:01 cfriesen sean-k-mooney: if it lands on a host with HT on, but there's enough free host CPUs, then we behave as now and we're good
19:39:18 jaypipes sean-k-mooney: see what cfriesen just said... the NUMATopologyFilter will catch the HT-snowflake stuff in the scheduler before going to the compute host.
19:39:18 sean-k-mooney jaypipes: yes i know. im aske about the usecase boot me a vm with 2 pinned cpus and tread policy isolate. will that still work with out a retry if the host has hypertreads on
19:39:20 cfriesen sean-k-mooney: actually no, we'd need to account in placement for the extra CPUs we consume
19:40:05 cfriesen jaypipes: given that the snowflake instance will actually consume more CPUs, we need to account for those extra ones in placement
19:40:17 sean-k-mooney cfriesen: jaypipes yes so there are two options at that point, 1 we fail to boot and retry, or 2 we ask placement to extend the allocation and fail if it would not fit
19:40:31 jaypipes cfriesen, sean-k-mooney: the fundamental problem with the cpu_thread_policy is that it leads to non-deterministic amounts of requested resources.
19:40:50 cfriesen jaypipes: it's deterministic, but it depends on the host
19:40:55 jaypipes sean-k-mooney: I'm not going to change the allocation request.
19:41:04 jaypipes cfriesen: omg, I'm gonna slap you.
19:41:09 jaypipes :)
19:41:12 cfriesen jaypipes: what about a whole new allocation request
19:41:22 jaypipes cfriesen: it's non-deterministic from the viewpoint of the scheduler
19:41:33 sean-k-mooney jaypipes: ok so we just define it a retry and maybe you could use a weigher to minimies the change it would happen
19:42:17 sean-k-mooney e.g. you land on a host. figure out you need more resouce to run on that host then you asked for and retry on next host
19:42:24 cfriesen sean-k-mooney: so now we're saying that ISOLATE can't possible run on a host with HT enabled. how is this different from an aggregate?
19:42:24 jaypipes cfriesen: for a whole new allocation request, we'd need to re-submit to GET /allocation_candidates with a new requested resource amount, which would give us back a different set of compute hosts, which we would send to the NUMATopologyFilter, which would re-work the allocation request again, and we'd end up in an infinite loop of sadnsees.
19:42:59 cfriesen jaypipes: can we specify a particular compute node when doing the allocation request?
19:43:04 sean-k-mooney cfriesen: actully you can make it work but it doubles the amount of flavors
19:43:14 jaypipes cfriesen: no.
19:43:31 cfriesen jaypipes: how do we handle specifying the compute node when doing a migration?
19:43:33 sean-k-mooney cfriesen: you set flavor.vcpu=4 and resouce[vcpu]=8 and it will work only on ht systems
19:43:56 jaypipes cfriesen: you can ask for an aggregate, an amount of resources, required traits, but not a specific provider (since that would defeat the entire purpose of the GET /allocation_candidates endpoint.
19:43:56 cfriesen sean-k-mooney: that'd work on non-ht as well, technically
19:44:30 jaypipes cfriesen: we don't call GET /allocation_candidates when specifying a compute node (force_host).
19:44:37 sean-k-mooney cfriesen: yes but it really expecive maybe add trait:HT_COUNT_2=required to avoid that
19:45:03 rybridges mlavalle: This is how I am creating the network / segment / subnet: http://paste.openstack.org/show/708997/ As you can see the aggregate host list is empty. I was under the impression that this process creates a routed network. Is there something I am missing?
19:45:05 jaypipes sean-k-mooney: that's a trait that looks suspiciously like a quantity of resources.
19:45:27 sean-k-mooney jaypipes: the trait is based on how we said we would do cpu frequency
19:45:58 jaypipes sean-k-mooney: you mean vGPU display heads?
19:46:05 sean-k-mooney e.g. tag it with multipel tratis so 4GHZ cpu would have 1GHZ,2GHZ,3GHZ and 4GHZ traits applied
19:46:36 mriedem melwitt: that's a known issue http://status.openstack.org/elastic-recheck/#1758054
19:46:40 mriedem everything is blocked until that's merged
19:46:48 sean-k-mooney jaypipes: ya i think that does the same thing but did not look at that that closely
19:46:59 jaypipes sean-k-mooney: but in this case, the amount of resources being *requested* changes depending on which host a workload ends up on :( that's the whole problem with this...
19:47:06 melwitt mriedem: ah, thanks
19:47:09 dansmith is gerrit sucking hard for everyone else?
19:47:42 sean-k-mooney jaypipes: yes so if you wanted an isoleated vm on a host with HT on you would have a flavor like this
19:48:17 cfriesen jaypipes: what do we call when specifying a compute node on a migration?
19:48:41 jaypipes dansmith: I'm too busy wanting to shoot myself in the head with a bazooka to feel any pain from gerrit.
19:48:41 sean-k-mooney flavor.vcpu=4,resouce[vcpu]=8:traits:HT_count_2=forbid thread_policy=isolate
19:48:52 dansmith jaypipes: roger that
19:49:10 cfriesen jaypipes: and why couldn't we do that in the scheduler if we realize after selecting a host that we need to account for some extra PCPU resources?
19:49:30 cfriesen sean-k-mooney: at that point you may as well use a host aggregate
19:49:45 jaypipes cfriesen: we call PUT /allocations/{migration_uuid} to reserve resources on the source host for the migration and PUT /allocations/{instance_uuid} to consume the instance resources on the destination host.
19:49:57 jaypipes cfriesen: we don't go through the shceduler at all when force_host.
19:50:24 cfriesen jaypipes: we do for migrations (when it's really a "suggested host" rather than force)
19:51:03 sean-k-mooney cfriesen: well you can have lavor.vcpu=4,resouce[vcpu]=8:traits:ht_count=require thread_policy=isolate and flavor.vcpu=4,resouce[vcpu]=4:traits:HT_count_2=forbid thread_policy=isolate
19:51:47 sean-k-mooney cfriesen: it should have been HT_count_2=require when resouce[vcpu]=8 not forbid originally
19:52:08 jaypipes sean-k-mooney: that just doesn't seem right to me.
19:52:10 cfriesen I really don't want to have multiple extra-spec keys that depend on the value of other extra-spec keys
19:52:22 jaypipes and certainly isn't very understandable to me.
19:52:28 cfriesen agreed, that's a mess. :)
19:53:13 mriedem we don't go through the scheduler when a host is forced, but conductor does the resource allocation 'claim'
19:53:22 mriedem for live migrate and evacuate
19:53:41 sean-k-mooney so we keep coming back to we deprecate and remove tread_policy=isolate or extend the allocation on the compute host which breaks the workflow
19:54:29 sean-k-mooney or we change isolate to mean "host with HT off"
19:55:35 cfriesen mriedem: for "cold migrate to this compute node" we need to be going through the scheduler to do cpu pinning, pci, etc.
19:57:29 mriedem cfriesen: we do go through the scheduler for cold migrate
19:57:34 mriedem there is no force for cold migrate
19:57:35 mriedem remember
19:57:37 mriedem -5
19:57:41 cfriesen jaypipes: is there an actual API spec somewhere for placement? https://docs.openstack.org/nova/latest/user/placement.html doesn't seem to document the HTTP calls.
19:57:56 mriedem cfriesen: https://developer.openstack.org/api-ref/placement/
19:58:17 mriedem https://docs.openstack.org/nova/latest/user/placement.html#rest-api goes to ^
19:58:19 cfriesen mriedem: we allow an optional host, which is not a "force" but is a suggested destination
19:58:28 mriedem cfriesen: yeah, and
19:58:29 mriedem ?
19:58:32 mriedem it goes through the scheduler filters
19:58:33 mriedem and placement
19:58:58 cfriesen mriedem: right, but for that placement call do we ask for all the possible allocation candidates or do we specify a particular host?
19:59:15 jaypipes sean-k-mooney: at this point, I'd much prefer a single trait called HW_CPU_HYPERTHREADING whose absence indicates that hyperthreads are not enabled on the host.
19:59:30 mriedem cfriesen: placement doesn't know about 'hosts'
19:59:41 jaypipes sean-k-mooney: and using the forbidden traits stuff to find compute hosts that don't have hyperthreading enabled.
19:59:42 mriedem cfriesen: we say, 'hey placement, here is my request, give me your shit'
19:59:51 mriedem and then we restrict to just that requested host for the filtering
20:00:00 jaypipes mriedem: and placement goes 'sorry, I don't speak jive'.
20:00:14 mriedem is that a 418?
20:00:19 cfriesen mriedem: okay, so the call to placement still filters all the RPs, then the scheduler filters narrow it down to the requested host?
20:00:19 jaypipes correct.
20:00:29 mriedem cfriesen: yes
20:00:31 jaypipes cfriesen: yes
20:00:41 sean-k-mooney jaypipes: ya im leaning that way too. just get rid of hw:cpu_thread_policy as technical debt but i know that will piss of several people
20:01:18 jaypipes sean-k-mooney: if by several people you do accurately mean 2-3 people in the world, I'm willing to live with that.

Earlier   Later