Earlier  
Posted Nick Remark
#openstack-nova - 2018-08-01
08:39:31 tetsuro yikun: Looks good to me. Thanks!
08:39:50 yikun thanks~
08:41:25 openstackgerrit Zhenyu Zheng proposed openstack/nova master: Expand aggregate_metadata table index to cover 'value' as well https://review.openstack.org/587704
09:15:50 openstackgerrit Pranab proposed openstack/os-vif master: Support for OVS DB TCP socket communication. https://review.openstack.org/587378
09:30:10 openstackgerrit Merged openstack/nova master: compute node local_gb_used include swap disks https://review.openstack.org/585928
09:33:52 openstackgerrit Merged openstack/nova master: Define common variables for irrelevant-files https://review.openstack.org/578882
09:53:04 openstackgerrit Merged openstack/nova master: Change deprecated policies to policy https://review.openstack.org/583434
09:59:05 kosamara Hi bauzas, I was checking your spec numa-topology-with-rps. What's the status on this, are you actively reworking it?
09:59:36 bauzas kosamara: yup, I'm planning to rework on it at the end of August
10:00:20 bauzas I'll be on PTO for 3 weeks starting next Wed
10:00:31 bauzas gibi: stephenfin: cdent: ^
10:00:57 cdent
10:05:29 kosamara bauzas: you mention RPs for PCI devices (L139 https://review.openstack.org/#/c/552924/9/specs/rocky/approved/numa-topology-with-rps.rst ). This is what I'm interested in. Nova currently doesn't create them, so I think it makes sense to do so either as part of this or as a new spec (maybe better, since your spec is already complex?)
10:05:56 gibi bauzas: have a nice time off!
10:09:51 bauzas kosamara: the spec is about discussing the topology, not really how to provide those devices
10:10:10 bauzas kosamara: so, yeah maybe a new spec
10:10:30 bauzas kosamara: at least about the upgrade between the existing and what you want
10:12:33 gibi bauzas: FYI, we talked about the PCI device handling with kosamara in openstack-neutron as well http://eavesdrop.openstack.org/irclogs/%23openstack-neutron/%23openstack-neutron.2018-08-01.log.html#t2018-08-01T08:23:33
10:14:59 bauzas gibi: ack
10:16:19 bauzas oh, fun, GPUs look like a good driver for PCI devices :p
10:17:07 gibi bauzas: fun indeed
10:22:54 bauzas mdbooth: others, I'm torn with https://review.openstack.org/#/c/478330/
10:23:56 lyarwood oh it's that time in the cycle again? :)
10:24:11 lyarwood I'm sure this has been nack'd in the past
10:25:46 bauzas nvm, found the spec
10:25:48 bauzas https://blueprints.launchpad.net/nova/+spec/volume-backed-server-rebuild
10:25:58 bauzas I'll accordingly -2 it
10:26:08 bauzas and refer to this series
10:26:13 lyarwood https://review.openstack.org/#/c/532407/ yeah
10:34:49 openstackgerrit Zengzhi Wang proposed openstack/nova master: Fix using last request-id for all log when multiple request calls. https://review.openstack.org/587754
10:35:09 mdbooth bauzas: On a call. Still want me to look?
10:35:39 bauzas mdbooth: nope, I gently -2'd the change :)
10:35:52 mdbooth 'gently' ;)
10:36:11 bauzas I'm French but I can be a gentleman
10:36:13 mdbooth bauzas: You used the fluffy ban hammer
10:36:38 bauzas mdbooth: there is one spec and already an open series
10:36:45 bauzas I just asked the proposer to collaborate
10:37:03 bauzas -2 just means "please stop what you're doing and discuss"
10:37:32 bauzas it's not a "I don't like you, and for god's sake, you'll suffer pain'
10:37:45 mdbooth bauzas: Although I recall the other discussion. I'm not in favour of the approach in this patch, tbh.
10:38:08 mdbooth To me, rebuild on BFV means you keep the original volume and re-image it.
10:39:01 sean-k-mooney mdbooth: is that not want was originally propsoed?
10:39:25 mdbooth sean-k-mooney: I think that was in the other proposal I recall.
10:39:34 mdbooth sean-k-mooney: My memory is notoriously shonky, though.
10:41:15 sean-k-mooney sorry im thinking of something else i was vagly rembering the rescure rescure bfv changes and mixing it up with rebuild
10:41:18 bauzas anyway, this is a behavioural change proposal, hence a spec
10:41:36 bauzas for people wanting to get the last proposal, the spec is thrre
10:44:06 openstackgerrit Pranab proposed openstack/os-vif master: Support for OVS DB TCP socket communication. https://review.openstack.org/587378
10:58:33 openstackgerrit Radoslav Gerganov proposed openstack/nova master: Reload oslo_context after calling monkey_patch() https://review.openstack.org/587772
11:30:18 kosamara bauzas: thanks! As I said in openstack-neutron, I'm examining how to create RPs for GPUs from the pci devices that nova tracks, in order to move the scheduling of PCI passthrough to placement. This should in the future allow quotas on them too, as resource classes.
11:32:14 kosamara Btw, could someone please take another look at https://review.openstack.org/#/c/579897 ?
11:34:11 sean-k-mooney kosamara: i belive that was planned to be started next cycle. not jsut for gpus but for pci devices in general
11:38:32 sean-k-mooney kosamara: can you explain what the vendorid in https://review.openstack.org/#/c/579897/2/nova/virt/libvirt/config.py
11:39:03 kosamara sean-k-mooney: makes sense not to do it just for GPUs. I don't know about the plans though. I was researching for any relevant specs and found the numa-rps and the net-bandwidth-rps. Can you point me to anything else?
11:41:10 kosamara sean-k-mooney: the vendorid can apparently be a random alphanumeric. I've tried it with the one proposed. My reference: https://wiki.archlinux.org/index.php/PCI_passthrough_via_OVMF#Troubleshooting
11:47:52 sean-k-mooney kosamara: the numa aware resouce providers was a prequisit for moving pci tracking to placement
11:50:04 kosamara sean-k-mooney: I'm not sure of that. PCI passthrough now doesn't require information about the NUMA topology. The PCI RPs could initially be children of the compute node RP, and then NUMA topology information can be added as extra layer in between, no?
11:51:40 sean-k-mooney kosamara: that would require the reshpaer work to be completed. but no today in openstack we track numa affinity of all pcidevces
11:52:47 kosamara sean-k-mooney: what's reshpaer ?
11:52:57 sean-k-mooney kosamara: what im guessing confused about with https://review.openstack.org/#/c/579897/2/nova/virt/libvirt/config.py is why you are adding a vendior id element to the hyperv features elememnt i nthe first place if you are trying to hide the fact you are being virutalised.
11:53:31 sean-k-mooney kosamara: https://specs.openstack.org/openstack/nova-specs/specs/rocky/approved/reshape-provider-tree.html
11:54:31 sean-k-mooney kosamara: its the ablitiy to reshap an exisiting nested resouce provider tree and move resouce providers form the root node to child nodes
11:54:32 kosamara sean-k-mooney: Maybe I should word this more precisely. We are trying to block detection of hyperv specifically. Making the vendorid anything other than the genuine one should work.
11:55:03 sean-k-mooney kosamara: are you using libvirt to manage hyperv nodes?
11:57:13 sean-k-mooney kosamara: the code you are modifying is changing the hyperv element in https://libvirt.org/formatdomain.html#elementsFeatures
12:06:05 lpetrut kosamara: I guess you're overriding the signature returned through cpuid 0x40000000
12:06:43 lpetrut sean-k-mooney: qemu can mimic some of the hyper-v enlightments
12:07:30 sean-k-mooney lpetrut: yep but im wondering if we shoudl entirely remove the hyperv element so that if nvidia make there check smarter in the future we dont need to change the code again
12:09:59 lpetrut but then I guess you'd lose all those hyper-v enlighments, which would probably affect Windows guests performance. I think the idea is to keep those, hiding the signature only. AFAIK, those guest features don't rely on the cpuid signature
12:10:25 kosamara sean-k-mooney: This is for kvm nodes. The <hyperv> element appears also on kvm nodes, when creating VMs from images with os_type='windows'
12:11:29 sean-k-mooney kosamara: yes i am aware. but im wondering if setting the vendor id is sufficent to hide everythin from the windows guest kernel.
12:11:44 sean-k-mooney or at least the driver layer
12:12:33 kosamara sean-k-mooney: I've tested inside the guest that I could install the nvidia drivers and run applications on the GPU.
12:12:48 sean-k-mooney lpetrut: i agree it would be nice to keep some of the performance aspects enabled if they are not observable within the guest.
12:13:42 sean-k-mooney kosamara: yes but is that because nvidia are just not current checking for other cpu flags that the enablments are setting or are you shure the vendor_id is enought to ensure we dont have to change this code again in the future
12:14:18 kosamara sean-k-mooney: unfortunately, I believe it's the 1st.
12:15:30 sean-k-mooney kosamara: in that case i would like at least a note to that effect in the code so that if we do have to revisit we understand exactly why the current workaround is not sufficent
12:16:20 kosamara The more explicit solution would be to remove the hyperv element as you say, but I've read that indeed this would reduce guest performance, like lpetrut says.
12:16:45 sean-k-mooney kosamara: well the kvm hidden flag also reduces the guest performance
12:18:36 sean-k-mooney it force the linux kernel to assume its running on phyical hardware so it cannont performa any optimisation that are posible when it knows its running on kvm
12:19:29 kosamara sean-k-mooney: since right now we don't need to cause another perf degradation though, I believe it's better to only do that if need be. I'll write a note on the code, thanks!
12:19:50 openstackgerrit Merged openstack/nova master: Add placement.concurrent_udpate to generation pre-checks https://review.openstack.org/581771
12:35:19 stephenfin bauzas: I'm gone for another week and a half next Wednesday too
13:31:39 openstackgerrit Merged openstack/nova master: [placement] Use oslotest CaptureOutput fixture https://review.openstack.org/587129
13:32:32 openstackgerrit Merged openstack/nova master: [placement] Use a non-nova log capture fixture https://review.openstack.org/587130
13:32:40 openstackgerrit Merged openstack/nova master: [placement] Use a simplified WarningsFixture https://review.openstack.org/587131
13:45:29 openstackgerrit Merged openstack/nova master: Fix comments in _anchors_for_sharing_providers and related test https://review.openstack.org/587679
13:45:38 openstackgerrit Merged openstack/nova master: Convert 'placement_api_docs' into a Sphinx extension https://review.openstack.org/578826
13:48:30 openstackgerrit Matt Riedemann proposed openstack/nova master: Add recreate test for RT.stats bug 1784705 https://review.openstack.org/587614
13:48:30 openstack bug 1784705 in OpenStack Compute (nova) "ResourceTracker.stats can leak across multiple ironic nodes" [High,In progress] https://launchpad.net/bugs/1784705 - Assigned to Matt Riedemann (mriedem)
13:48:31 openstackgerrit Matt Riedemann proposed openstack/nova master: Make ResourceTracker.stats node-specific https://review.openstack.org/587636
13:59:05 dansmith mriedem: don't you watch the stats leak bug tagged for rc1?
13:59:24 mriedem dansmith: it's not new, so not really an rc thing
13:59:27 mriedem regressed in ocata
13:59:36 efried mode #openstack-nova
13:59:38 dansmith but fairly critical no?
13:59:50 mriedem i would think so....
14:00:04 dansmith so I dunno, seems important enough to go into rc1 to me
14:00:18 mriedem tssurya: do you know if cern is using the ComputeCapabilitiesFilter for baremetal scheduling where the baremetal flavors are have capabilities extra specs to link them to certain bm nodes?
14:00:30 mriedem dansmith: i wouldn't argue against getting it fixed asap of course :)

Earlier   Later