| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-10-19 | |||
| 15:42:46 | imacdonn | (we do this all the time, in a private cloud context) | |
| 15:44:37 | spatel | cfriesen: I have other VMs running on same compute node they are fine.. very strange issue | |
| 15:45:31 | SteelyDan | imacdonn: yeah I'm sure everyone does.. I get the use case, it's unfortunate to need to check the status on every call, but I get it | |
| 15:53:16 | openstackgerrit | Stephen Finucane proposed openstack/nova master: api-ref: 'os-hypervisors' doesn't reflect overcommit ratio https://review.openstack.org/611604 | |
| 16:07:26 | cfriesen | imacdonn: so lock the instances? | |
| 16:09:01 | cfriesen | imacdonn: or stop the nova-compute process? | |
| 16:09:08 | imacdonn | cfriesen: I guess, but the owner can unlock it? | |
| 16:09:47 | cfriesen | imacdonn: you could always make lock/unlock admin-only. | |
| 16:09:57 | imacdonn | cfriesen: I guess stopping nova-compute may cause the bug that mriedem_afk was reviewing ... although there are some open questions in the bug | |
| 16:16:10 | cfriesen | I'm not sure it's really a bug...more like a feature to more gracefully handle cases where things get "stuck" in what was supposed to be a transitional state. | |
| 16:52:27 | openstackgerrit | Merged openstack/nova master: Fix deprecated base64.decodestring warning https://review.openstack.org/610401 | |
| 17:00:27 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Fail to live migration if instance has a NUMA topology https://review.openstack.org/611088 | |
| 17:05:51 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Fail to live migration if instance has a NUMA topology https://review.openstack.org/611088 | |
| 17:23:29 | awaugama | finucannot, was at lunch. so you're able to reproduce that setting the value on the compute node doesn't change anything in placement? | |
| 17:29:21 | sean-k-mooney | awaugama: which setting is that? | |
| 17:30:12 | awaugama | cpu_allocation_ratio I believe | |
| 17:30:45 | awaugama | yeah | |
| 17:30:46 | sean-k-mooney | awaugama: there is a spec current for changing how we adress this for stien | |
| 17:31:11 | awaugama | I'm not aware of one | |
| 17:31:38 | sean-k-mooney | awaugama: there are 2 but i belive we have setteled on one ill see if i can find it | |
| 17:31:46 | awaugama | cool | |
| 17:32:57 | sean-k-mooney | awaugama: this is one of the specs https://review.openstack.org/#/c/552105/ | |
| 17:33:25 | sean-k-mooney | i think that is the most recent one | |
| 17:35:03 | sean-k-mooney | this is the other one https://review.openstack.org/#/c/544683/ but im not sure if it will still be requried | |
| 17:35:11 | awaugama | checking | |
| 17:36:02 | awaugama | yeah that looks like what finucannot was talking about being messed up | |
| 17:36:24 | sean-k-mooney | well that depends on what you were expecting | |
| 17:36:36 | sean-k-mooney | awaugama: what was your initall issue | |
| 17:37:07 | awaugama | setting cpu_allocation_ratio to 1 in nova.conf changed the value on the compute node (even on compute node startup it showed it was changed) but placement still has it set to 16 | |
| 17:37:54 | sean-k-mooney | awaugama: right i belive currently it will only use the value if it is first creating the provider | |
| 17:38:16 | sean-k-mooney | if the provider exists it will not update it i think... | |
| 17:38:57 | awaugama | that seems counter-intuitive | |
| 17:39:25 | sean-k-mooney | awaugama: this will be changed going forward to allow the cpu_allocation_ratio confige option to allow specify the placement value too and cpu_inital_allocation_ratio to specify what it shoudl be for newly creeted resouce providers | |
| 17:39:52 | sean-k-mooney | awaugama: the current behavior i belive is intened to allow you to mainge the allocation ratios via the api instead of the config | |
| 17:40:26 | awaugama | interesting. finucannot and mriedem_afk ^ | |
| 17:40:39 | awaugama | i need to step away for a few minutes, brb | |
| 17:41:43 | sean-k-mooney | awaugama: finucannot is likely not around anymore. | |
| 17:42:12 | sean-k-mooney | awaugama: he is usually enjoying his weekend by now :) | |
| 17:50:39 | awaugama | sean-k-mooney, yeah, I figure he'll get the scrollback at somepoint | |
| 17:57:41 | sean-k-mooney | awaugama: the inital allocation ratios spec is likely the one that is most relevent to you https://review.openstack.org/#/c/552105/ | |
| 18:04:31 | spatel | sean-k-mooney: hey! how are you doing :) | |
| 18:04:54 | sean-k-mooney | spatel: quite well thank you. how are you :) | |
| 18:04:57 | spatel | after longtime i am seeing you online or may be i was not paying attention | |
| 18:05:13 | sean-k-mooney | i was on company traing most of this week | |
| 18:05:21 | spatel | I am great!! and my openstack also going great | |
| 18:05:46 | sean-k-mooney | so ya i was not online much. that is good to hear. | |
| 18:06:01 | spatel | I have added 80 SR-IOV compute nodes and put them in production :) | |
| 18:06:31 | sean-k-mooney | wow that was fast. are you happy with the feature set/perfromance you have achived? | |
| 18:06:51 | spatel | Yup!! i am not seeing any performance impact for my application so far | |
| 18:07:38 | sean-k-mooney | you decided to go with vnic_type direct with a flavor with hugepages and pinned cpus in the end? | |
| 18:07:59 | spatel | I have create two AZ group as we spoken last time, tor-1 and tor-2 and spreading application across two AZ | |
| 18:08:30 | spatel | Yes vnic_type=direct / hugepages / cpu pinning / numa=2 | |
| 18:08:46 | spatel | whatever setting you suggested last time | |
| 18:09:03 | sean-k-mooney | cool that should be a good starting point for any VNF deployed on openstack | |
| 18:09:23 | spatel | I am happy with those setting :) | |
| 18:09:27 | sean-k-mooney | and for the tor config did you split the control and dataplane across the tors or keep them on the same tor | |
| 18:11:08 | spatel | That i didn't tested yet because of deployment urgency! | |
| 18:11:18 | spatel | i need to test that in LAB first | |
| 18:11:55 | spatel | that is in my to-do list | |
| 18:12:32 | sean-k-mooney | well the original config you proposed should work but i was not certen the failure mode was optimal. that said you intended to add a deicated manament switch at a later poit so that will likely be the best longterm solution in anycase | |
| 18:12:38 | spatel | as soon as i get time i want to run test on dpdk (SR-IOV is still painful even its better) | |
| 18:14:17 | spatel | sean-k-mooney: yes i have plan in future to isolate mgmt traffic using extra nic | |
| 18:16:48 | spatel | sean-k-mooney: Quick question, how do you guys monitor instance stats? like CPU/memory from KVM | |
| 18:18:09 | spatel | currently i have snmp agent running on compute node and also inside instance but that won't give you 100% result right? i want to monitor hypervisor level stats | |
| 18:18:39 | spatel | KVM level monitoring in short | |
| 18:19:58 | sean-k-mooney | am you have several options | |
| 18:20:17 | sean-k-mooney | in generall i would recommened collectd for host level perfomance monitoring | |
| 18:20:53 | sean-k-mooney | if you have deployed the openstack ceilometer project you can also confure it to monitro the libvirt isntances | |
| 18:21:48 | sean-k-mooney | i belive the prometious project has good monitoring capablites for the openstack servcies but it does not have good monitoring capablites for the host or the vms as far as i am aware | |
| 18:22:21 | spatel | I didn't configure ceilometer because i wasn't sure i need that component because its private cloud and we don't need billing | |
| 18:22:55 | sean-k-mooney | spatel: that is a common missconceptions ceilometer is for telemetry not billing | |
| 18:23:09 | sean-k-mooney | that said its not nessisarly the best solution for telemetry either | |
| 18:24:10 | spatel | i was worried it will eat my resources | |
| 18:24:19 | sean-k-mooney | if you decide to use collectd you could look into https://github.com/openstack/collectd-openstack-plugins to plublis metrics to celometer or gnocci but collectd can also advertise the stata it monitors over snmp or toher protocols | |
| 18:24:54 | sean-k-mooney | spatel: yes that is a concern with celimiter. it does not scale well which is why i generally dont recomend it as my first choice | |
| 18:25:14 | spatel | Agreed ++ | |
| 18:25:31 | spatel | i had same concern so i didn't bother to install | |
| 18:26:09 | sean-k-mooney | ya so in general i would check out collectd and see if it meets your needs | |
| 18:26:26 | spatel | I am going to try collectd sure! | |
| 18:26:49 | sean-k-mooney | my former team enhanced the collectd libvirt plugin to allow reporting stats for vms using the uuid field instead of the id field | |
| 18:27:02 | sean-k-mooney | the uuid filed is set to the nova instance uuid | |
| 18:27:17 | sean-k-mooney | so its easy to do a 1:1 mapping between the stats and the workload | |
| 18:27:40 | spatel | hmm! interesting i think i have to look into collectd now | |
| 18:28:04 | spatel | also i was thinking push collectd data to influx + grafana so i have good dashboard | |
| 18:28:08 | sean-k-mooney | you can also have collectd use its network plugin to send the stats to influxdb and then use graphana to visualise the results | |
| 18:28:15 | sean-k-mooney | jinks | |
| 18:28:32 | sean-k-mooney | :) that is a good solution | |
| 18:28:49 | spatel | yup! | |
| 18:29:15 | sean-k-mooney | are you familar with OPNFV? | |
| 18:29:23 | spatel | no much | |
| 18:29:54 | sean-k-mooney | thats ok its rather vast in scope like openstack | |
| 18:30:12 | sean-k-mooney | spatel: i just wanted to highlight the barometer project https://wiki.opnfv.org/display/fastpath/Barometer+Home | |
| 18:31:01 | spatel | reading.. | |
| 18:32:13 | sean-k-mooney | spatel: they have be working with integrating collect graphana and other tools for openstack monitoring for telco/NFV usecases | |
| 18:32:16 | sean-k-mooney | https://www.youtube.com/watch?v=-82UShFiBBM | |
| 18:32:26 | spatel | sean-k-mooney: very interesting... also i had question about SR-IOV VF nic stats, it doesn't tell you how much data flowing through your VF | |
| 18:32:57 | sean-k-mooney | spatel: that depend on the nic | |
| 18:33:18 | spatel | I have Qlogic so i doubt it has that feature | |
| 18:33:23 | sean-k-mooney | so most nics dont have enough hardware counters to gatar stats on all the VFs from the host side | |
| 18:34:00 | sean-k-mooney | spatel: if you are able to run collectd in the guest however it can monitor the kernel stats and give you that info | |