Earlier  
Posted Nick Remark
#openstack-nova - 2019-11-04
15:20:44 efried isn't that the provider that's been choking?
15:20:49 efried http://grafana.openstack.org/d/BhcSH5Iiz/nodepool-ovh?orgId=1
15:20:51 dansmith you said ovn
15:20:55 efried oh, sorry
15:20:58 efried yes, ovh
15:21:01 dansmith which is a network technology, so just making sure :)
15:21:09 mriedem ez bake ovn
15:21:25 efried taking mriedem back to his childhood
15:21:39 mriedem dude i'm making some muffins in the thing right now
15:22:01 efried Takes *way* longer with these stupid CFLs
15:22:14 dansmith pretty sure mriedem never left childhood
15:23:34 mriedem i will say, since maya likes legos i do find myself tempted around xmas to splurge on a $150 lego set just so i can help build it
15:23:52 dansmith nice
15:24:08 efried donnyd: can you think of an easy way $nodepool_provider could switch up configuration to reduce the chance of choking I/O?
15:24:51 donnyd On FN?
15:24:58 mriedem "but dad i only like the friends and disney princess sets" "you're getting medieval castle gdi"
15:25:05 efried specifically ovh seems to be the problem at the moment.
15:25:06 dansmith mriedem: lol
15:25:47 donnyd Oh, well I know I avoid that issue by using local storage on FN. Not sure what ovh has on the backend of their instances
15:26:10 efried not asking for action on your part donnyd, just advice. If a patch has 10 zuul jobs an any one of them lands on ovh and chokes on I/O and times out, the whole patch has to be retried, which is a royal PITA. Been happening a really lot over the last week or two.
15:26:43 efried Yeah, I guess they would need to look into where the actual bottleneck is.
15:26:53 efried not sure if there's any way to tell from here via grafana...
15:26:55 donnyd The reality is we should try and label jobs by what they are bound by
15:27:09 efried heh, we were just talking about your email from Sept.
15:27:25 donnyd If there is a cpu bound job, scheduling on FN would be less optimal
15:27:44 efried I couldn't see anywhere infra had responded to that idea, but I thought I remembered them shutting it down hard for some reason.
15:27:50 cdent efried, mriedem : what I meant by less cpu, was less cpu for all nodes, not per node
15:27:51 donnyd But IO bound jobs will go like stink on FN
15:27:53 cdent so that we can run fewr jobs
15:28:13 cdent because we run too many at once
15:28:22 donnyd I shut down FN when I put in my gen set
15:28:24 cdent because oversubscribers are liars
15:28:48 efried cdent: okay, that's what I was suggesting too.
15:28:53 donnyd I oversub FN by a small margin, but not much
15:29:07 donnyd Usually 100% of the memory is utilized
15:29:08 efried so like a white lie?
15:29:25 cdent donnyd in my experience what you're doing seems to be working much better than some of the other providers, so kudos to you
15:29:31 donnyd But CPU can be oversub like 1.125
15:29:50 donnyd cdent: I need faster CPUs for sure
15:30:01 donnyd Mine are so slow compared to others
15:30:12 mriedem looking dstat when this messaging timeout happened https://zuul.opendev.org/t/openstack/build/ef0196fe84804b44ac106d011c8c29ea/log/controller/logs/screen-n-api.txt.gz?severity=4 around the time of the timeout cpu usage is low, io is basically 0 and load is spiked
15:30:31 donnyd The biggest difference for FN is local NVME storage for all instances
15:30:43 cdent question: isn't this something infra will already have a plan and solution for?
15:30:52 cdent this is what cloud mgt/provisioning is all about, presumably?
15:31:11 mriedem i wonder if these providers live migrate the ci vms around frequently?
15:31:30 donnyd If there was a way to benchmark providers and then prefer them for jobs that are bound by something I think we could optimize the CI
15:32:19 donnyd Butttt.... if that provider breaks or goes away the issue becomes jobs failing because they were dependent on that super fast thing provider_x does
15:32:28 mriedem cdent: infra is just tenants so i'm not sure how much control they have
15:32:48 mriedem besides may requesting minimums in a flavor or something
15:32:49 mriedem ?
15:32:51 cdent mriedem: but presumably they can tweak their consumption?
15:32:53 mriedem *maybe
15:32:54 donnyd mriedem: well a simple benchmark could expose what jobs run optimal
15:34:19 donnyd That was my idea and it got quickly struck down because of the issue listed above efried
15:35:42 donnyd I do all kinds of custom things and so does mnaser at vexxhost to support the CI... not sure how much usage it gets though. I think limestone is on board with it as well
15:35:45 efried ah, thanks donnyd, I remember now
15:35:51 donnyd NP
15:36:01 donnyd Hopefully that helps
15:36:17 efried So really $provider needs to be able to run $job. In this case ovh needs to go figure out what the bottleneck is and fix it.
15:36:31 donnyd Bingo
15:36:44 efried even if it's by reducing the number of nodes they provide
15:36:49 donnyd But they probably don't even know its an issue till someone complains
15:36:59 efried not sure where else to complain
15:37:08 donnyd probably infra
15:37:10 efried openstack-discuss@ has a thread
15:37:44 donnyd You can always schedule the job using maybe the numa label made for sean-k-mooney
15:38:00 efried I guess we wait until after the summit then, infra folks are going to be pretty sparse for the next week I imagine.
15:38:05 donnyd That would ensure the job runs on FN vexxhost or limestone if I am not mistaken
15:38:12 efried donnyd: yeah, but you have to put that in the job def itself, right?
15:38:26 efried You can't do it one off to make a particular patch merge
15:38:29 donnyd Yea, you have to assign it a label
15:38:31 donnyd No
15:38:44 efried that's our problem here, getting approved patches through the gate.
15:38:44 donnyd Sry, that is correct efried
15:39:16 donnyd Well patch the job with that label... then it will run anyways wont it
15:39:20 efried heh, we should just change all our devstack-based job defs to run on the fast providers...
15:39:33 donnyd Well I wouldn't call FN fast
15:39:36 efried wait until somebody notices
15:39:41 efried okay, s/fast/reliable/
15:39:43 donnyd Just better at IO bound jobs
15:40:06 efried Yeah, I would rather my patch take 2h and succeed than 1.5h and fail.
15:40:17 donnyd I'm sure ironic and tripleo hate FN because my CPUs are old
15:40:41 donnyd And they are mostly CPU bound
15:40:54 donnyd Give it a swing with the label
15:41:05 donnyd And if it works, you have a bandaid
15:42:21 openstackgerrit Eric Fried proposed openstack/nova master: Add cyborg tempest job. https://review.opendev.org/670999
16:14:29 efried donnyd: I wasn't serious about the label. That would be pretty publicly greedy of us.
16:19:19 dansmith mriedem: question for you in here while I look at the test: https://review.opendev.org/#/c/635646/48/nova/conductor/tasks/cross_cell_migrate.py
16:19:50 dansmith holy hell, everything in the check/gate are running tests and only minutes old
16:19:58 dansmith we should send the whole community to china more often
16:24:15 donnyd Well not really efried
16:24:45 donnyd If a job has requirements to run and a specific label meets the requirements then I don't see the issue
16:25:02 donnyd But it's your call... just pointing out the options
16:26:44 mriedem dansmith: thanks, replied
16:27:18 mriedem dansmith: is my sarcasm detector picking this up correctly? https://review.opendev.org/#/c/635080/48/nova/tests/unit/compute/test_compute_mgr.py@10545
16:27:32 dansmith mriedem: yes.
16:28:02 mriedem heh, verbose commentage is how i keep from feeling lonely
16:34:55 dansmith mriedem: oops, my "...yes I see" was supposed to go into the fault_clone complaint nit
16:35:13 dansmith went back to note that and picked the wrong comment to edit
16:35:48 dansmith mriedem: one more question

Earlier   Later