Earlier  
Posted Nick Remark
#openstack-nova - 2018-10-19
15:31:48 mriedem_afk i've triaged https://bugs.launchpad.net/nova/+bug/1798805 but not sure if we would ever do anything about it
15:33:07 finucannot bauzas: If I delete a compute nodes RP, it'll be recreated, right?
15:34:04 SteelyDan mriedem_afk: I think that means disable in the "break its needs" sort of sense
15:34:31 SteelyDan I'm strongly in favor of keeping disable purely for scheduler reasons, but I think it's saying if the compute is actually down, we shouldn't take action on instances
15:34:35 SteelyDan which is probably reasonable
15:35:01 SteelyDan or make our super awesome bug-free state sync periodic power the instance on when the compute comes back
15:36:03 imacdonn I kinda sorta wish there was a "don't schedule anything new" status, which is different from "don't try to interact with me at all"
15:36:26 SteelyDan imacdonn: that's what disable means
15:36:29 SteelyDan it's the only thing it means
15:36:43 imacdonn the former, you mean
15:36:59 SteelyDan it means don't schedule anything there
15:37:04 imacdonn right
15:37:31 SteelyDan you're saying you wish there was a "kneecap this compute"
15:37:33 SteelyDan API?
15:37:45 imacdonn IMO, it should really mean "don't talk that node at all right now", and there should be some other way to tell the scheduler not to pout anything NEW there
15:38:23 finucannot mriedem_afk, awaugama: More investigation needed but I think there's something wrong with those foo_allocation_ratio options. I configured them on the compute node, waited ages aaand...nada. Deleting the RP fixed things
15:38:27 imacdonn (or maybe the opposite, so keep the existing meaning of "disable") ... but I think they are use-cases for each
15:38:37 SteelyDan we're not changing the existing meaning for disable
15:38:57 SteelyDan we could add another state, but it just means for every single operation we do another db hit to check the state of the host
15:39:03 finucannot mriedem_afk, awaugama: At least, assuming my understanding of how that's _supposed_ to work is correct. It it Friday so maybe it's not
15:39:24 imacdonn "hard_disabled" ? :)
15:39:31 SteelyDan no.
15:42:13 imacdonn possible use-case: node is being taken down for hardware maintenance. operator has shut down the instances on it. We don't want to allow the user to star their instances back up until the maintenance is completed
15:42:46 imacdonn (we do this all the time, in a private cloud context)
15:44:37 spatel cfriesen: I have other VMs running on same compute node they are fine.. very strange issue
15:45:31 SteelyDan imacdonn: yeah I'm sure everyone does.. I get the use case, it's unfortunate to need to check the status on every call, but I get it
15:53:16 openstackgerrit Stephen Finucane proposed openstack/nova master: api-ref: 'os-hypervisors' doesn't reflect overcommit ratio https://review.openstack.org/611604
16:07:26 cfriesen imacdonn: so lock the instances?
16:09:01 cfriesen imacdonn: or stop the nova-compute process?
16:09:08 imacdonn cfriesen: I guess, but the owner can unlock it?
16:09:47 cfriesen imacdonn: you could always make lock/unlock admin-only.
16:09:57 imacdonn cfriesen: I guess stopping nova-compute may cause the bug that mriedem_afk was reviewing ... although there are some open questions in the bug
16:16:10 cfriesen I'm not sure it's really a bug...more like a feature to more gracefully handle cases where things get "stuck" in what was supposed to be a transitional state.
16:52:27 openstackgerrit Merged openstack/nova master: Fix deprecated base64.decodestring warning https://review.openstack.org/610401
17:00:27 openstackgerrit Stephen Finucane proposed openstack/nova master: Fail to live migration if instance has a NUMA topology https://review.openstack.org/611088
17:05:51 openstackgerrit Stephen Finucane proposed openstack/nova master: Fail to live migration if instance has a NUMA topology https://review.openstack.org/611088
17:23:29 awaugama finucannot, was at lunch. so you're able to reproduce that setting the value on the compute node doesn't change anything in placement?
17:29:21 sean-k-mooney awaugama: which setting is that?
17:30:12 awaugama cpu_allocation_ratio I believe
17:30:45 awaugama yeah
17:30:46 sean-k-mooney awaugama: there is a spec current for changing how we adress this for stien
17:31:11 awaugama I'm not aware of one
17:31:38 sean-k-mooney awaugama: there are 2 but i belive we have setteled on one ill see if i can find it
17:31:46 awaugama cool
17:32:57 sean-k-mooney awaugama: this is one of the specs https://review.openstack.org/#/c/552105/
17:33:25 sean-k-mooney i think that is the most recent one
17:35:03 sean-k-mooney this is the other one https://review.openstack.org/#/c/544683/ but im not sure if it will still be requried
17:35:11 awaugama checking
17:36:02 awaugama yeah that looks like what finucannot was talking about being messed up
17:36:24 sean-k-mooney well that depends on what you were expecting
17:36:36 sean-k-mooney awaugama: what was your initall issue
17:37:07 awaugama setting cpu_allocation_ratio to 1 in nova.conf changed the value on the compute node (even on compute node startup it showed it was changed) but placement still has it set to 16
17:37:54 sean-k-mooney awaugama: right i belive currently it will only use the value if it is first creating the provider
17:38:16 sean-k-mooney if the provider exists it will not update it i think...
17:38:57 awaugama that seems counter-intuitive
17:39:25 sean-k-mooney awaugama: this will be changed going forward to allow the cpu_allocation_ratio confige option to allow specify the placement value too and cpu_inital_allocation_ratio to specify what it shoudl be for newly creeted resouce providers
17:39:52 sean-k-mooney awaugama: the current behavior i belive is intened to allow you to mainge the allocation ratios via the api instead of the config
17:40:26 awaugama interesting. finucannot and mriedem_afk ^
17:40:39 awaugama i need to step away for a few minutes, brb
17:41:43 sean-k-mooney awaugama: finucannot is likely not around anymore.
17:42:12 sean-k-mooney awaugama: he is usually enjoying his weekend by now :)
17:50:39 awaugama sean-k-mooney, yeah, I figure he'll get the scrollback at somepoint
17:57:41 sean-k-mooney awaugama: the inital allocation ratios spec is likely the one that is most relevent to you https://review.openstack.org/#/c/552105/
18:04:31 spatel sean-k-mooney: hey! how are you doing :)
18:04:54 sean-k-mooney spatel: quite well thank you. how are you :)
18:04:57 spatel after longtime i am seeing you online or may be i was not paying attention
18:05:13 sean-k-mooney i was on company traing most of this week
18:05:21 spatel I am great!! and my openstack also going great
18:05:46 sean-k-mooney so ya i was not online much. that is good to hear.
18:06:01 spatel I have added 80 SR-IOV compute nodes and put them in production :)
18:06:31 sean-k-mooney wow that was fast. are you happy with the feature set/perfromance you have achived?
18:06:51 spatel Yup!! i am not seeing any performance impact for my application so far
18:07:38 sean-k-mooney you decided to go with vnic_type direct with a flavor with hugepages and pinned cpus in the end?
18:07:59 spatel I have create two AZ group as we spoken last time, tor-1 and tor-2 and spreading application across two AZ
18:08:30 spatel Yes vnic_type=direct / hugepages / cpu pinning / numa=2
18:08:46 spatel whatever setting you suggested last time
18:09:03 sean-k-mooney cool that should be a good starting point for any VNF deployed on openstack
18:09:23 spatel I am happy with those setting :)
18:09:27 sean-k-mooney and for the tor config did you split the control and dataplane across the tors or keep them on the same tor
18:11:08 spatel That i didn't tested yet because of deployment urgency!
18:11:18 spatel i need to test that in LAB first
18:11:55 spatel that is in my to-do list
18:12:32 sean-k-mooney well the original config you proposed should work but i was not certen the failure mode was optimal. that said you intended to add a deicated manament switch at a later poit so that will likely be the best longterm solution in anycase
18:12:38 spatel as soon as i get time i want to run test on dpdk (SR-IOV is still painful even its better)
18:14:17 spatel sean-k-mooney: yes i have plan in future to isolate mgmt traffic using extra nic
18:16:48 spatel sean-k-mooney: Quick question, how do you guys monitor instance stats? like CPU/memory from KVM
18:18:09 spatel currently i have snmp agent running on compute node and also inside instance but that won't give you 100% result right? i want to monitor hypervisor level stats
18:18:39 spatel KVM level monitoring in short
18:19:58 sean-k-mooney am you have several options
18:20:17 sean-k-mooney in generall i would recommened collectd for host level perfomance monitoring
18:20:53 sean-k-mooney if you have deployed the openstack ceilometer project you can also confure it to monitro the libvirt isntances
18:21:48 sean-k-mooney i belive the prometious project has good monitoring capablites for the openstack servcies but it does not have good monitoring capablites for the host or the vms as far as i am aware
18:22:21 spatel I didn't configure ceilometer because i wasn't sure i need that component because its private cloud and we don't need billing
18:22:55 sean-k-mooney spatel: that is a common missconceptions ceilometer is for telemetry not billing
18:23:09 sean-k-mooney that said its not nessisarly the best solution for telemetry either
18:24:10 spatel i was worried it will eat my resources
18:24:19 sean-k-mooney if you decide to use collectd you could look into https://github.com/openstack/collectd-openstack-plugins to plublis metrics to celometer or gnocci but collectd can also advertise the stata it monitors over snmp or toher protocols
18:24:54 sean-k-mooney spatel: yes that is a concern with celimiter. it does not scale well which is why i generally dont recomend it as my first choice
18:25:14 spatel Agreed ++
18:25:31 spatel i had same concern so i didn't bother to install

Earlier   Later