Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-21
13:44:08 edleafe Partial
13:45:09 cdent like your soul
13:45:33 dansmith I'm at 99.6%
13:45:42 dansmith don't expect me around much until after :)
13:46:08 cdent dansmith: are you overrun nearby by tourists and such?
13:46:38 dansmith cdent: no, I live in a heavily fortified mountaintop stronghold remember?
13:46:58 cdent dansmith: understood, but I wondered if there might be people at the walls
13:46:59 edleafe cdent: this is as good as it will get for me: http://imgur.com/a/z6FfF
13:47:01 dansmith but in reality, so far the traffic apocolypse hasn't been a thing
13:47:08 dansmith cdent: heh, not yet :)
13:47:33 edleafe dansmith: my sister-in-law lives in Albany. She says traffic has been hell there for the past week
13:48:16 dansmith edleafe: yeah, I think the places without road capacity are feeling it, but the bigger places haven't had the troubles they expected
13:48:38 dansmith we were over in eastern/central this weekend and it was a ghost town
13:48:53 dansmith gas stations have been stockpiling toilet paper and bullets for weeks and it was all for naught
13:49:11 edleafe heh
13:49:31 edleafe Well, for once Albany has something to attract tourists
13:49:43 dansmith heh
13:55:05 edleafe Scheduler subteam meeting in 5 minutes in #openstack-meeting-alt
14:07:37 mriedem hmm, so i heard stable gate sucked last week
14:07:39 mriedem and continues to suck
14:07:53 mriedem looking at a failure, the ComputeFilter kicked something out because the single-node host job has 4 VMs on it already
14:08:12 mriedem vcpus=8,vcpus_used=4
14:08:29 mriedem ^ is the compute node stats before the NoValidHost
14:12:23 mriedem seems odd given there is capacity left on that node and it's also weird that we have CoreFilter enabled in devstack since it was disabled by default in code
14:14:22 alex_xu gibi: thanks, i will read the irc log
14:14:52 mriedem https://goo.gl/g7LGmM
14:15:50 cdent what manner of malware are you sending us to with that link mriedem ?
14:15:56 mriedem logstash
14:16:37 cdent mriedem: you’re being semi-ignored becasue I think anyone who cares is over in the scheduler meeting at the mo
14:16:45 mriedem i know
14:18:41 alex_xu mriedem: fyi, a regression in python-novaclient https://review.openstack.org/#/c/492003/, not sure whether it is worth a release
14:19:46 mriedem alex_xu: so it's been broken since 9.0.0
14:19:52 mriedem if https://review.openstack.org/#/c/429512/ caused it
14:20:03 alex_xu yea, probably for a while
14:20:27 mriedem then we'll probably just fix in a patch release once stable/pike is ok for python-novaclient releases
14:20:54 sean-k-mooney mriedem: it is requsting 8 vcpus it should fit but it would be oversubscribing though not against itself
14:21:17 mriedem sean-k-mooney: it's requesting 1 vcpu
14:21:19 sean-k-mooney mriedem: i assume that the 4 cores already in use are not dedicated?
14:21:19 alex_xu ok
14:21:35 mriedem sean-k-mooney: we don't have cpu pinning in the upstream infra gate jobs no
14:21:53 sean-k-mooney mriedem: ah sorry missread http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_765
14:22:07 sean-k-mooney {"cores": 1, "cells": 1, "threads": 1, "sockets": 8}
14:22:11 sean-k-mooney that is for the host
14:22:16 sean-k-mooney not the vm right
14:22:20 mriedem correct
14:25:19 mriedem looking at the CoreFilter code this doesn't make any sense
14:26:03 mriedem we have 8 total, and cpu_allocation_ratio is 16.0, so we end up with vcpus_total = 24, 24-vcpus_used (4) = 20, so we have 20 free vcpus and we're requesting 1
14:27:32 cdent mriedem: is that a new vm or move/migration/whatever?
14:27:37 mriedem new vm
14:28:03 sean-k-mooney mriedem: well you should be able to use 128 vcps with 8 cpus and an allocation ratios of 16. anyway did you notice this error 'disabled_reason': u'AUTO: Connection to libvirt lost:
14:28:42 sean-k-mooney at http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_766
14:29:31 mriedem hmm seeing that now yeah
14:30:40 mriedem oh duh,
14:30:48 mriedem it's not the CoreFilter, it's the ComputeFilter
14:30:50 mriedem god
14:31:19 sean-k-mooney there are a load of tracebacks in the n-cpu log
14:31:46 sean-k-mooney NeutronAdminCredentialConfigurationInvalid: Networking client is experiencing an unauthorized exception.
14:31:52 openstackgerrit Balazs Gibizer proposed openstack/nova master: WIP: test allocation handling during scheduler retry https://review.openstack.org/495891
14:32:11 mriedem sean-k-mooney: that one is a known issue,
14:32:16 mriedem the real problem is libvirt crashed
14:32:25 mriedem so the libvirt driver auto-disables the compute service
14:35:16 sean-k-mooney mriedem: there does not seam to be any usfull info in the libvirtd log other then a bunch of virNetlinkEventCallback spam
14:35:33 guimaluf Could you guys help me to debug this error? http://paste.openstack.org/show/618918/ I'm trying to launch an instance within my testing environment, which consist of 4 vms, keystone/api-controllers, mysql/ceph/rabbitmq, neutron and nova.
14:36:04 guimaluf my concern is mainly focused on HTTPMultipleChoices: HTTPMultipleChoices (HTTP 300) Requested version of OpenStack Images API is not available. I expect nova to contact glance v2 api, but it is trying to reach v1 even with only v2 available
14:37:46 mriedem sean-k-mooney: i won't pursue any further - libvirt crashing is not a new thing in the gate
14:41:01 sean-k-mooney mriedem: well strangly the libvirtd log seams to still be active when the connection drops on the nova side but ya this does not really look like it s novas problem
14:41:28 sean-k-mooney mriedem: the compute agent will reconnect and reactivate itself anyway correct
14:41:33 mriedem yes
14:42:03 mriedem in a prod cloud with hundreds/thousands of nodes, that's probably good enough
14:42:09 mriedem for single node gate jobs it's a killer
14:43:32 sean-k-mooney ya tempest could maybe retry the failed jobs automatically to a certenlimit at the end fo the run in the future but honestly that will likey just slow down the gate when code is actully broken
14:49:57 maciejjozefczyk guimaluf: check nova-scheduler and rabbitmq
14:50:32 maciejjozefczyk guimaluf: s/nova-scheduler/nova-conductor/
14:51:17 mriedem maciejjozefczyk: anything glance related is likely blowing up in the api
14:51:21 mriedem sounds like misconfiguration
14:51:32 mriedem [glance]api_servers is where i'd look
14:51:38 mriedem in nova.conf
14:51:56 mriedem and make sure that matches what's in the service catalog for the image service
14:57:34 maciejjozefczyk mriedem: Could you check this cherry-pick https://review.openstack.org/#/c/494974/ please :) ?
14:59:03 mriedem maciejjozefczyk: has to go through pike first,
14:59:19 mriedem and i see the cherry pick for pike, but this isn't a regression in pike, so i'm not sure we should include it in the rc2 for pike
15:01:09 maciejjozefczyk mriedem: So it should wait for stable pike?
15:01:25 mriedem maciejjozefczyk: it has to be in stable/pike before we can merge it in stable/ocata, yes
15:01:32 mriedem i'm not sure yet if it should be in rc2 for pike
15:03:08 maciejjozefczyk mriedem: ok
15:05:47 mriedem man i guess you guys were busy last week https://review.openstack.org/#/q/status:merged+project:openstack/nova+branch:stable/pike
15:06:44 dansmith mriedem: anything you see in there that wasn't legit?
15:06:47 dtantsur folks, do you by chance know who is leading https://etherpad.openstack.org/p/queens-PTG-vmbm ? I should have been aware of it, but I'm not :)
15:07:02 cdent dtantsur: probably johnthetubaguy ?
15:07:03 dansmith mriedem: some of the work required some test backports for cleanliness
15:07:20 mriedem dansmith: yeah that looks like the evacuate/shelve offload tests from givi
15:07:22 mriedem *gib
15:07:23 mriedem gdi
15:07:24 mriedem gibi
15:07:26 dansmith yeah
15:07:42 dtantsur johnthetubaguy: is it you? :)
15:09:12 mriedem dansmith: i found one major issue here https://review.openstack.org/#/c/493037/4/nova/compute/resource_tracker.py@1225
15:09:39 dansmith mriedem: oh the humanity
15:10:13 gibi mriedem: welcome back. :)
15:13:04 mriedem thanks

Earlier   Later