Earlier  
Posted Nick Remark
#openstack-nova - 2018-12-05
18:04:17 cfriesen so I guess you'd need to still use the ComputeCapabilitiesFilter for that?
18:04:34 cfriesen combine that with requesting the trait in the flavor, and you'd get both, no?
18:06:00 sean-k-mooney it still misses the point that HW_CPU_X86_AVX is ment to only be used if the host support it in hardware
18:07:22 cfriesen sean-k-mooney: if ComputeCapabilitiesFilter fails hosts that don't have it in hardware, and specifying the trait requests that it's available in the guest, that seems like it'd work.
18:07:49 cfriesen The other option is to simply not specify CPU models with features unsupported by your hardware...which is up to the operator.
18:08:11 sean-k-mooney it would work but i realy dislike that we are not useing the HW e.g. hardware namespace to model only hardware things
18:08:22 cfriesen sean-k-mooney: have you got any specs/docs for the intent of the trait?
18:09:41 sean-k-mooney that was my intent when i asked for namespacing in os traits https://github.com/openstack/os-traits/commit/23d81d4451dd29c23150b29b8fa9d3025ee8878f#diff-f290aedb8b7fdc21b2b04be76222f6c3
18:10:23 sean-k-mooney the HW namespace was for hardware features
18:10:42 cfriesen we use the "hw" namespace for all sorts of virtual hardware stuff
18:11:10 sean-k-mooney we only started doing that this cylce with the vtpm spec
18:11:12 cfriesen number of numa nodes, cpu distribution between numa nodes, etc
18:11:19 cfriesen oh, you mean for trait
18:11:23 sean-k-mooney yes
18:12:10 sean-k-mooney anyway looks like we approved and implmented https://specs.openstack.org/openstack/nova-specs/specs/rocky/implemented/report-cpu-features-as-traits.html#libvirt so we are stuck with it
18:13:04 sean-k-mooney i was hopping not to make the same mistakes we did with flavors in os-traits but if i can rely on traits to modele this correctly i will always need filters
18:13:16 cfriesen I think it'd be worth adding something to the release notes warning operators to consider this when adding CPU models to the list.
18:13:55 cfriesen In practice, I would expect that most operators advertising high-performance would not want emulated features, no?
18:14:16 sean-k-mooney ya or better to the config option so its in the docs for setting the models/extra flags
18:14:27 cfriesen right, that makes sense
18:14:51 sean-k-mooney cfriesen: yes but as a tenat i can no longer rely on it being a hardware feature in a public cloud
18:15:15 sean-k-mooney i get the opertor usecase but this originally came form the mano folks
18:15:51 cfriesen gotta run for lunch
18:15:59 sean-k-mooney o/
18:16:08 sean-k-mooney ill comment back on the spec.
18:16:21 sean-k-mooney since we are already missuing them wwe might as well continue
18:45:17 openstackgerrit Matt Riedemann proposed openstack/nova-specs master: Per aggregate scheduling weight (spec) https://review.openstack.org/599308
18:46:35 mriedem_away johnthetubaguy: bauzas: ^ when you're around, i'm +2 on ^ now - seems like a good compromise from the per-flavor weights it started from
18:46:49 mriedem_away mgagne_: since you said you had something like this downstream already ^ it would be good if you can ack that works for you as well
18:52:31 mgagne_ mriedem: done. thanks
19:32:06 openstackgerrit Jack Ding proposed openstack/nova master: Preserve UEFI NVRAM variable store https://review.openstack.org/621646
21:09:26 efried jaypipes, mriedem: Are you waiting for CERN to deploy on ironic nodes before reviewing series at https://review.openstack.org/#/c/615677/ ?
21:10:56 mriedem efried: not really
21:11:02 mriedem i'm not intentionally avoiding it
21:11:34 efried Cool, I know you're stretched pretty thin.
21:11:46 mriedem but my waist line continues to grow
21:13:57 mriedem here is an easy gate related fix https://review.openstack.org/#/c/623011/
21:15:40 mriedem efried: i've been ignoring some other stuff for awhile so trying to burn that list down first
21:15:49 mriedem reviewing your series will be my xmas gift to you
21:15:52 efried mriedem: +2 on the tempest timeout
21:15:54 mriedem thanks
21:15:56 efried Thanks :)
21:38:38 efried mriedem: reserved=total for which resource? And how do you stop the virt driver from overwriting that? (Re ML post about CERN workaround for low alloc candidates limit)
21:40:24 mriedem VCPU? all of them?
21:40:27 dansmith for the compute node
21:40:35 dansmith for all of them yeah
21:40:57 mriedem the compute would probably need to know if it's service is disabled and if so, not ovewrite it
21:41:04 mriedem *its
21:41:04 dansmith mriedem: I think we discussed just having an rpc cast to compute to have it do it, vs. a periodic
21:41:22 mriedem dansmith: sure, but the next update_available_resource periodic would overwrite it
21:41:30 mriedem when reporting inventory
21:41:34 dansmith the compute can stop calling the virt driver's update method if it's disabled I would think
21:41:36 sean-k-mooney efried: the virt dirver really should not be touching the reserved value excpet when its first creating the RP
21:41:46 mriedem sean-k-mooney: the ironic driver does all the time
21:41:49 mriedem when the node is being cleaned
21:41:49 dansmith sean-k-mooney: uh, why?
21:41:57 dansmith the reserved amount is owned by the virt driver, IMHO
21:42:02 mriedem that's exactly why the placement api change was made so that reserved can equal total
21:42:04 dansmith only the virt driver knows what it should be
21:42:25 sean-k-mooney dansmith: well for the same reason as teh have the cpu allocation ratios vs inital cpu allocation ratios spec
21:42:28 efried agreed, all the inventory values ought to be owned by the virt driver, period. We start making exceptions, we end up with messes like allocation ratio... and reserved.
21:42:36 sean-k-mooney controling via api or config
21:42:57 efried but making that contract stick for this kind of workaround is going to be tricksy.
21:42:58 dansmith sean-k-mooney: that's not an example that helps your case I think :)
21:43:24 mriedem the alternative was a trait i think
21:43:32 efried unless we do it with a handy-dandy provider config yaml file https://review.openstack.org/#/c/612497/
21:43:32 mriedem and a pre-request placement filter in nova-scheduler
21:43:35 mriedem or something like that
21:43:51 mriedem this isn't blues clues
21:43:55 dansmith mriedem: yeah, that's also an option
21:44:28 mriedem so api sets a trait (or removes it), virt doesn't overwrite it, and scheduler filters on it (essentially it becomes the ComputeFilter)
21:44:31 efried omg, now you totally remind me of Steve from blues clues.
21:44:35 sean-k-mooney efried: provider config yaml for what?
21:44:43 mriedem efried: there were a couple of steves
21:44:48 mriedem or was it one and then another guy?
21:44:50 dansmith mriedem: not scheduler filters, request filter
21:44:51 efried no, there was only one steve.
21:44:55 sean-k-mooney how does that allw you to set it via the api and not have the compute overriede
21:44:56 efried the other guy was...
21:44:59 efried joe?
21:45:07 mriedem yes joe
21:45:17 mriedem http://www.gstatic.com/tv/thumb/persons/250022/250022_v9_ba.jpg
21:45:48 mriedem sean-k-mooney: b/c the compute/virt driver isn't supposed to overwrite externally set traits
21:46:01 mriedem and the api can set traits just like it mirrors aggregates
21:46:06 sean-k-mooney anyway if we say only the entity that created the RP may set the reserved value on the invetories im fine with that if we document it
21:46:58 dansmith I thought there was some other reason for doing the rpc cast on disable, but I can't remember what it was
21:47:46 dansmith maybe so ironic driver can do something?
21:47:52 sean-k-mooney on disableing a compute node. i think some dirver can take actions on disable and ironic might be one of them but i dont remember eiter
21:48:14 efried we can ask each driver's upt to do
21:48:14 efried if self.disabled:
21:48:14 efried <do a right thing to disable the inventory>
21:48:39 dansmith efried: we still have to do it as soon as the api is called though,
21:48:44 dansmith so waiting for the next periodic is not good enough
21:48:51 efried mm
21:49:08 dansmith oh yeah, so..
21:49:11 efried do we have a ComputeDriver hook that we call on disable?
21:49:11 mriedem right i could disable an entire cell,
21:49:15 dansmith disable is on the service not the node, right?
21:49:23 efried or could/should we implement one?
21:49:25 mriedem er half a cell, and then live migrate stuff within the cell

Earlier   Later