Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-21
14:08:29 mriedem ^ is the compute node stats before the NoValidHost
14:12:23 mriedem seems odd given there is capacity left on that node and it's also weird that we have CoreFilter enabled in devstack since it was disabled by default in code
14:14:22 alex_xu gibi: thanks, i will read the irc log
14:14:52 mriedem https://goo.gl/g7LGmM
14:15:50 cdent what manner of malware are you sending us to with that link mriedem ?
14:15:56 mriedem logstash
14:16:37 cdent mriedem: you’re being semi-ignored becasue I think anyone who cares is over in the scheduler meeting at the mo
14:16:45 mriedem i know
14:18:41 alex_xu mriedem: fyi, a regression in python-novaclient https://review.openstack.org/#/c/492003/, not sure whether it is worth a release
14:19:46 mriedem alex_xu: so it's been broken since 9.0.0
14:19:52 mriedem if https://review.openstack.org/#/c/429512/ caused it
14:20:03 alex_xu yea, probably for a while
14:20:27 mriedem then we'll probably just fix in a patch release once stable/pike is ok for python-novaclient releases
14:20:54 sean-k-mooney mriedem: it is requsting 8 vcpus it should fit but it would be oversubscribing though not against itself
14:21:17 mriedem sean-k-mooney: it's requesting 1 vcpu
14:21:19 sean-k-mooney mriedem: i assume that the 4 cores already in use are not dedicated?
14:21:19 alex_xu ok
14:21:35 mriedem sean-k-mooney: we don't have cpu pinning in the upstream infra gate jobs no
14:21:53 sean-k-mooney mriedem: ah sorry missread http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_765
14:22:07 sean-k-mooney {"cores": 1, "cells": 1, "threads": 1, "sockets": 8}
14:22:11 sean-k-mooney that is for the host
14:22:16 sean-k-mooney not the vm right
14:22:20 mriedem correct
14:25:19 mriedem looking at the CoreFilter code this doesn't make any sense
14:26:03 mriedem we have 8 total, and cpu_allocation_ratio is 16.0, so we end up with vcpus_total = 24, 24-vcpus_used (4) = 20, so we have 20 free vcpus and we're requesting 1
14:27:32 cdent mriedem: is that a new vm or move/migration/whatever?
14:27:37 mriedem new vm
14:28:03 sean-k-mooney mriedem: well you should be able to use 128 vcps with 8 cpus and an allocation ratios of 16. anyway did you notice this error 'disabled_reason': u'AUTO: Connection to libvirt lost:
14:28:42 sean-k-mooney at http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_766
14:29:31 mriedem hmm seeing that now yeah
14:30:40 mriedem oh duh,
14:30:48 mriedem it's not the CoreFilter, it's the ComputeFilter
14:30:50 mriedem god
14:31:19 sean-k-mooney there are a load of tracebacks in the n-cpu log
14:31:46 sean-k-mooney NeutronAdminCredentialConfigurationInvalid: Networking client is experiencing an unauthorized exception.
14:31:52 openstackgerrit Balazs Gibizer proposed openstack/nova master: WIP: test allocation handling during scheduler retry https://review.openstack.org/495891
14:32:11 mriedem sean-k-mooney: that one is a known issue,
14:32:16 mriedem the real problem is libvirt crashed
14:32:25 mriedem so the libvirt driver auto-disables the compute service
14:35:16 sean-k-mooney mriedem: there does not seam to be any usfull info in the libvirtd log other then a bunch of virNetlinkEventCallback spam
14:35:33 guimaluf Could you guys help me to debug this error? http://paste.openstack.org/show/618918/ I'm trying to launch an instance within my testing environment, which consist of 4 vms, keystone/api-controllers, mysql/ceph/rabbitmq, neutron and nova.
14:36:04 guimaluf my concern is mainly focused on HTTPMultipleChoices: HTTPMultipleChoices (HTTP 300) Requested version of OpenStack Images API is not available. I expect nova to contact glance v2 api, but it is trying to reach v1 even with only v2 available
14:37:46 mriedem sean-k-mooney: i won't pursue any further - libvirt crashing is not a new thing in the gate
14:41:01 sean-k-mooney mriedem: well strangly the libvirtd log seams to still be active when the connection drops on the nova side but ya this does not really look like it s novas problem
14:41:28 sean-k-mooney mriedem: the compute agent will reconnect and reactivate itself anyway correct
14:41:33 mriedem yes
14:42:03 mriedem in a prod cloud with hundreds/thousands of nodes, that's probably good enough
14:42:09 mriedem for single node gate jobs it's a killer
14:43:32 sean-k-mooney ya tempest could maybe retry the failed jobs automatically to a certenlimit at the end fo the run in the future but honestly that will likey just slow down the gate when code is actully broken
14:49:57 maciejjozefczyk guimaluf: check nova-scheduler and rabbitmq
14:50:32 maciejjozefczyk guimaluf: s/nova-scheduler/nova-conductor/
14:51:17 mriedem maciejjozefczyk: anything glance related is likely blowing up in the api
14:51:21 mriedem sounds like misconfiguration
14:51:32 mriedem [glance]api_servers is where i'd look
14:51:38 mriedem in nova.conf
14:51:56 mriedem and make sure that matches what's in the service catalog for the image service
14:57:34 maciejjozefczyk mriedem: Could you check this cherry-pick https://review.openstack.org/#/c/494974/ please :) ?
14:59:03 mriedem maciejjozefczyk: has to go through pike first,
14:59:19 mriedem and i see the cherry pick for pike, but this isn't a regression in pike, so i'm not sure we should include it in the rc2 for pike
15:01:09 maciejjozefczyk mriedem: So it should wait for stable pike?
15:01:25 mriedem maciejjozefczyk: it has to be in stable/pike before we can merge it in stable/ocata, yes
15:01:32 mriedem i'm not sure yet if it should be in rc2 for pike
15:03:08 maciejjozefczyk mriedem: ok
15:05:47 mriedem man i guess you guys were busy last week https://review.openstack.org/#/q/status:merged+project:openstack/nova+branch:stable/pike
15:06:44 dansmith mriedem: anything you see in there that wasn't legit?
15:06:47 dtantsur folks, do you by chance know who is leading https://etherpad.openstack.org/p/queens-PTG-vmbm ? I should have been aware of it, but I'm not :)
15:07:02 cdent dtantsur: probably johnthetubaguy ?
15:07:03 dansmith mriedem: some of the work required some test backports for cleanliness
15:07:20 mriedem dansmith: yeah that looks like the evacuate/shelve offload tests from givi
15:07:22 mriedem *gib
15:07:23 mriedem gdi
15:07:24 mriedem gibi
15:07:26 dansmith yeah
15:07:42 dtantsur johnthetubaguy: is it you? :)
15:09:12 mriedem dansmith: i found one major issue here https://review.openstack.org/#/c/493037/4/nova/compute/resource_tracker.py@1225
15:09:39 dansmith mriedem: oh the humanity
15:10:13 gibi mriedem: welcome back. :)
15:13:04 mriedem thanks
15:16:31 mriedem it's going to be a rough couple of days so give me softballs if you have them
15:17:51 gibi I'm going for vacation this week so I can promise that I will not file yet another resource allocation bugs. :)
15:19:01 mriedem q
15:19:04 mriedem oops
15:22:22 mriedem ocata 15.0.7 release request https://review.openstack.org/495903
15:26:27 guimaluf maciejjozefczyk, the error I mentioned is happening on nova-compute node, no errs on conductor logs. I'll enable debug to see if there is something meaningful
15:28:03 dansmith mriedem: I left this for you to make a call on: https://review.openstack.org/#/c/494973/
15:29:20 openstackgerrit Merged openstack/nova master: replace chance with filter scheduler in func tests https://review.openstack.org/491529
15:29:41 openstackgerrit Merged openstack/nova master: Clarify that vlan feature means nova-network support https://review.openstack.org/478551
15:33:39 mriedem dansmith: yeah saw that, wasn't sure since it's not a regression in pike
15:34:17 dansmith mriedem: yeah thought that was the right call, but figured I'd leave it to you.. can wait until after pike releases and then push the backport like normal
15:45:13 openstackgerrit Stephen Finucane proposed openstack/nova master: Remove plug_ovs_hybrid, unplug_ovs_hybrid https://review.openstack.org/483030
15:48:48 lvdombrkr hello guys, im trieng deploy tripleO , and after trying deploy overcloud get error: http://paste.openstack.org/raw/618917/ , in nova scheduler logs i see that tha node is filtred out http://paste.openstack.org/raw/618930/
15:48:54 lvdombrkr but i dont uderstand why
15:49:12 lvdombrkr any ideas?
15:50:18 openstackgerrit Stephen Finucane proposed openstack/nova master: tests: Remove useless test https://review.openstack.org/483031
15:55:19 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce basic framework for security proxying https://review.openstack.org/345396
15:55:20 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce framework for RFB authentication https://review.openstack.org/345397
15:55:20 openstackgerrit Stephen Finucane proposed openstack/nova master: console: introduce the VeNCrypt RFB authentication scheme https://review.openstack.org/345398
15:55:21 openstackgerrit Stephen Finucane proposed openstack/nova master: console: provide an RFB security proxy implementation https://review.openstack.org/345399
15:57:08 mriedem disabling RamFilter by default hits it's first victim https://bugs.launchpad.net/nova/+bug/1712057
15:57:09 openstack Launchpad bug 1712057 in OpenStack Compute (nova) "When the specified destination host deploys the virtual machine, the allocation ratio is not valid" [Undecided,Invalid] - Assigned to wangyicheng (wang-yicheng)

Earlier   Later