Earlier  
Posted Nick Remark
#openstack-nova - 2017-08-21
12:13:44 cdent We got any cores around yet today, besides jaypipes ? This would be good to kick home: https://review.openstack.org/#/c/491529/
13:20:49 sean-k-mooney jaypipes: got a link to cdent's shared storage provider patches? ill try to take a look at them if i get a chance. by the way is the master brance still locked down for pike or has it repoend for queens
13:27:20 edleafe jaypipes: jsut saw https://review.openstack.org/#/c/471927/
13:27:40 edleafe jaypipes: I have series to do similar that I was just about to push :(
13:28:12 edleafe I took the series that I had already done before the switch to alloc candidates and updated it
13:28:31 edleafe oh geez - this is just the spec
13:28:34 edleafe nvm
13:33:01 sean-k-mooney edleafe: sould like an easy win then if you add implementes(with the correct spelling): placement-allocation-requests
13:34:19 alex_xu gibi: a question about https://review.openstack.org/#/c/491529/14//COMMIT_MSG@17, is it a bug?
13:36:19 openstackgerrit Ed Leafe proposed openstack/nova master: WIP - add alternate hosts https://review.openstack.org/486215
13:36:20 openstackgerrit Ed Leafe proposed openstack/nova master: WIP - Add allocations to the values returned from the scheduler https://review.openstack.org/495854
13:36:20 openstackgerrit Ed Leafe proposed openstack/nova master: WIP - return alternates along with their allocations https://review.openstack.org/486253
13:39:17 gibi alex_xu: no, it is not a bug. It is the way how the resouce claiming was implemented in the scheduler for the move operations
13:39:36 gibi alex_xu: there was couple of irc discussion about it in the past
13:40:16 gibi alex_xu: here is some irc history http://p.anticdent.org/logs/openstack-nova?dated=2017-08-08%2012:29:21.129291#15Hv
13:40:54 stephenfin cdent: Done
13:40:59 stephenfin https://review.openstack.org/#/c/491529/, that is
13:41:10 cdent thanks stephenfin welcome back
13:41:17 stephenfin cdent: Cheers :)
13:42:19 cdent I’m trying to sit outside in the sun but it may be too squinty
13:43:18 stephenfin The rain is bouncing off the pavement outside. I'm not at all jealous :D
13:43:31 edleafe cdent: I gotta make my eclipse viewing box this morning
13:43:32 gibi alex_xu, stephenfin: as you are already looking at functional tests and scheduling, there are two additional test patch that needs core eyes
13:43:38 gibi alex_xu, stephenfin https://review.openstack.org/#/c/495159/
13:43:47 gibi alex_xu, stephenfin: https://review.openstack.org/#/c/493865/
13:43:58 cdent edleafe: are you in the path? I thought you were too far south
13:44:08 edleafe Partial
13:45:09 cdent like your soul
13:45:33 dansmith I'm at 99.6%
13:45:42 dansmith don't expect me around much until after :)
13:46:08 cdent dansmith: are you overrun nearby by tourists and such?
13:46:38 dansmith cdent: no, I live in a heavily fortified mountaintop stronghold remember?
13:46:58 cdent dansmith: understood, but I wondered if there might be people at the walls
13:46:59 edleafe cdent: this is as good as it will get for me: http://imgur.com/a/z6FfF
13:47:01 dansmith but in reality, so far the traffic apocolypse hasn't been a thing
13:47:08 dansmith cdent: heh, not yet :)
13:47:33 edleafe dansmith: my sister-in-law lives in Albany. She says traffic has been hell there for the past week
13:48:16 dansmith edleafe: yeah, I think the places without road capacity are feeling it, but the bigger places haven't had the troubles they expected
13:48:38 dansmith we were over in eastern/central this weekend and it was a ghost town
13:48:53 dansmith gas stations have been stockpiling toilet paper and bullets for weeks and it was all for naught
13:49:11 edleafe heh
13:49:31 edleafe Well, for once Albany has something to attract tourists
13:49:43 dansmith heh
13:55:05 edleafe Scheduler subteam meeting in 5 minutes in #openstack-meeting-alt
14:07:37 mriedem hmm, so i heard stable gate sucked last week
14:07:39 mriedem and continues to suck
14:07:53 mriedem looking at a failure, the ComputeFilter kicked something out because the single-node host job has 4 VMs on it already
14:08:12 mriedem vcpus=8,vcpus_used=4
14:08:29 mriedem ^ is the compute node stats before the NoValidHost
14:12:23 mriedem seems odd given there is capacity left on that node and it's also weird that we have CoreFilter enabled in devstack since it was disabled by default in code
14:14:22 alex_xu gibi: thanks, i will read the irc log
14:14:52 mriedem https://goo.gl/g7LGmM
14:15:50 cdent what manner of malware are you sending us to with that link mriedem ?
14:15:56 mriedem logstash
14:16:37 cdent mriedem: you’re being semi-ignored becasue I think anyone who cares is over in the scheduler meeting at the mo
14:16:45 mriedem i know
14:18:41 alex_xu mriedem: fyi, a regression in python-novaclient https://review.openstack.org/#/c/492003/, not sure whether it is worth a release
14:19:46 mriedem alex_xu: so it's been broken since 9.0.0
14:19:52 mriedem if https://review.openstack.org/#/c/429512/ caused it
14:20:03 alex_xu yea, probably for a while
14:20:27 mriedem then we'll probably just fix in a patch release once stable/pike is ok for python-novaclient releases
14:20:54 sean-k-mooney mriedem: it is requsting 8 vcpus it should fit but it would be oversubscribing though not against itself
14:21:17 mriedem sean-k-mooney: it's requesting 1 vcpu
14:21:19 alex_xu ok
14:21:19 sean-k-mooney mriedem: i assume that the 4 cores already in use are not dedicated?
14:21:35 mriedem sean-k-mooney: we don't have cpu pinning in the upstream infra gate jobs no
14:21:53 sean-k-mooney mriedem: ah sorry missread http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_765
14:22:07 sean-k-mooney {"cores": 1, "cells": 1, "threads": 1, "sockets": 8}
14:22:11 sean-k-mooney that is for the host
14:22:16 sean-k-mooney not the vm right
14:22:20 mriedem correct
14:25:19 mriedem looking at the CoreFilter code this doesn't make any sense
14:26:03 mriedem we have 8 total, and cpu_allocation_ratio is 16.0, so we end up with vcpus_total = 24, 24-vcpus_used (4) = 20, so we have 20 free vcpus and we're requesting 1
14:27:32 cdent mriedem: is that a new vm or move/migration/whatever?
14:27:37 mriedem new vm
14:28:03 sean-k-mooney mriedem: well you should be able to use 128 vcps with 8 cpus and an allocation ratios of 16. anyway did you notice this error 'disabled_reason': u'AUTO: Connection to libvirt lost:
14:28:42 sean-k-mooney at http://logs.openstack.org/28/468528/2/check/gate-tempest-dsvm-neutron-full-ubuntu-xenial/337af98/logs/screen-n-sch.txt#_2017-08-21_04_15_14_766
14:29:31 mriedem hmm seeing that now yeah
14:30:40 mriedem oh duh,
14:30:48 mriedem it's not the CoreFilter, it's the ComputeFilter
14:30:50 mriedem god
14:31:19 sean-k-mooney there are a load of tracebacks in the n-cpu log
14:31:46 sean-k-mooney NeutronAdminCredentialConfigurationInvalid: Networking client is experiencing an unauthorized exception.
14:31:52 openstackgerrit Balazs Gibizer proposed openstack/nova master: WIP: test allocation handling during scheduler retry https://review.openstack.org/495891
14:32:11 mriedem sean-k-mooney: that one is a known issue,
14:32:16 mriedem the real problem is libvirt crashed
14:32:25 mriedem so the libvirt driver auto-disables the compute service
14:35:16 sean-k-mooney mriedem: there does not seam to be any usfull info in the libvirtd log other then a bunch of virNetlinkEventCallback spam
14:35:33 guimaluf Could you guys help me to debug this error? http://paste.openstack.org/show/618918/ I'm trying to launch an instance within my testing environment, which consist of 4 vms, keystone/api-controllers, mysql/ceph/rabbitmq, neutron and nova.
14:36:04 guimaluf my concern is mainly focused on HTTPMultipleChoices: HTTPMultipleChoices (HTTP 300) Requested version of OpenStack Images API is not available. I expect nova to contact glance v2 api, but it is trying to reach v1 even with only v2 available
14:37:46 mriedem sean-k-mooney: i won't pursue any further - libvirt crashing is not a new thing in the gate
14:41:01 sean-k-mooney mriedem: well strangly the libvirtd log seams to still be active when the connection drops on the nova side but ya this does not really look like it s novas problem
14:41:28 sean-k-mooney mriedem: the compute agent will reconnect and reactivate itself anyway correct
14:41:33 mriedem yes
14:42:03 mriedem in a prod cloud with hundreds/thousands of nodes, that's probably good enough
14:42:09 mriedem for single node gate jobs it's a killer
14:43:32 sean-k-mooney ya tempest could maybe retry the failed jobs automatically to a certenlimit at the end fo the run in the future but honestly that will likey just slow down the gate when code is actully broken
14:49:57 maciejjozefczyk guimaluf: check nova-scheduler and rabbitmq
14:50:32 maciejjozefczyk guimaluf: s/nova-scheduler/nova-conductor/
14:51:17 mriedem maciejjozefczyk: anything glance related is likely blowing up in the api

Earlier   Later