Earlier  
Posted Nick Remark
#openstack-nova - 2018-03-19
22:53:06 dansmith that kinda goes against the original supposition of melwitt's patch to client RPC though
22:53:11 melwitt here's the original bug that led to using overwrite=False https://bugs.launchpad.net/nova/+bug/1627838
22:53:12 openstack Launchpad bug 1627838 in OpenStack Compute (nova) newton "Context is reset to by ClientRouter in nova/rpc.py" [High,Fix committed] - Assigned to Matt Riedemann (mriedem)
22:54:59 dansmith hmm, I don't get that
23:00:07 melwitt yeah, I guess that was a different thing. normally we run periodics with get_admin_context which does an overwrite=False and when I added a different unrelated periodic I didn't use get_admin_context and inadvertently made it overwrite=True without realizing it
23:01:26 melwitt and the reported problem was that the periodic was blowing away other requests auth token in-flight
23:01:57 dansmith although ClientRouter does schedule periodics quite differently from service it seems
23:03:28 melwitt yeah, maybe there is something wrong with how it's running periodics?
23:03:36 dansmith well,
23:03:59 dansmith I'm wondering if that's why it needed overwrite=False, but our periodics in compute manager need overwrite=True
23:04:11 melwitt oh, I see
23:04:24 dansmith melwitt: that thing needed periodics to age out clients but that code seems gone I think, so can we remove thatl ine anyway?
23:04:40 melwitt yeah, I was thinking that too
23:06:18 dansmith so we're running periodics in our managers via threadgroup timers, which I assume run in threads of their own, but I should check
23:06:52 dansmith mriedem: yeah, so that _remove_stale_clients() periodic was removed later, so we should nuke that line
23:06:56 dansmith er, melwitt ^
23:07:25 melwitt ack. I can propose that
23:10:12 openstackgerrit Matt Riedemann proposed openstack/nova master: Remove RequestContext.instance_lock_checked https://review.openstack.org/554378
23:10:13 mriedem more cleanup while i was poking around ^
23:11:14 dansmith so I'm not exactly seeing how the way periodics are used in ClientRouter would have worked, other than the single go of them run via __init__
23:11:24 dansmith but my guess is that they're not run the way we run them in our service
23:11:50 dansmith melwitt: do we know for sure that those periodics ever ran more than once?
23:11:56 dansmith that code wasn't in the tree for very long, so..
23:12:52 dansmith it sure looks like we just fired them off once during __init_ and never ran them again
23:13:25 dansmith which is also puzzling as to why that would be a problem needing to be reported, but the bug doesn't have much info
23:13:31 melwitt dansmith: maybe not. ClientRouter inherits from oslo PeriodicTasks, so it would do this https://github.com/openstack/oslo.service/blob/master/oslo_service/periodic_task.py#L193
23:13:59 dansmith melwitt: well, it does that once when you run it in __init__
23:14:00 dansmith but I don't see where that gets scheduled any other place
23:14:02 dansmith service.py will do it, but you're not a service there
23:16:37 dansmith and the tests just fire the handler manually
23:16:49 melwitt hm
23:16:50 dansmith (as expected, we don't really have good coverage over periodics)
23:17:43 dansmith https://www.youtube.com/watch?v=C9MG9b4mEU0
23:18:29 melwitt lol, great movie
23:18:53 dansmith the only thing I can think of,
23:19:07 dansmith would be that we instantiate a new compute rpc class for certain things at runtime,
23:19:20 dansmith and maybe we blow away whatever context was current at that time,
23:19:46 dansmith because anything running ComputeRPCAPI.__init() will run ClientRouter.__init__(), which will instantiate a context with overwrite=True before your patch
23:20:13 dansmith er, rpcapi.ComputeAPI() I mean
23:20:56 melwitt yeah :\
23:22:33 dansmith mriedem: so I think maybe chatting with someone like dhellmann might be a good idea, but I kinda think we should be running with overwrite=True on those periodics' context
23:24:09 dansmith I'd also like someone to tell me why TLS isn't cleared when a new thread is spawned and/or would like to confirm that timers as we use them to start periodics in our services are actually running in their own threads and not as if they were signal handlers
23:26:38 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix message for unexpected external event https://review.openstack.org/554380
23:27:18 mriedem ok, but not tonight
23:31:46 openstackgerrit melanie witt proposed openstack/nova master: Remove useless run_periodic_tasks call in ClientRouter https://review.openstack.org/554381
23:34:02 melwitt dansmith, mriedem ^ I left the related overwrite=False regression test case alone thinking we would address that when we get answers about what we should be doing with overwrite + periodics later
23:35:22 dansmith melwitt: hmm, I'm not sure what that test is testing TBH
23:35:35 dansmith I'm pretty fried and well past punkin time, so I'll look tomorrow with a fresh head
23:35:57 melwitt it's testing that we don't overwrite=True during periodics because at the time we thought that was messing up things
23:36:03 melwitt okay, cool
23:42:15 dansmith I don't think it does though, because we just start the service and then assert our thing,
23:42:23 dansmith I don't think we know that the periodics have run yet
23:43:06 dansmith we don't run service.periodic_tasks() until the first time we fire the timer, which is delayed from startup, and that's where context gets generated for the periodic run
#openstack-nova - 2018-03-20
00:04:14 openstackgerrit Merged openstack/nova master: Remove old flavor_extra_specs_delete db api method https://review.openstack.org/539702
00:04:55 openstackgerrit Merged openstack/nova master: Report client: Remove version discovery comment https://review.openstack.org/554253
00:05:08 openstackgerrit Merged openstack/nova master: Remove version/date from CLI documentation https://review.openstack.org/553903
00:13:31 pooja Hi.. I am seeing an issue with NumInstancesFilter in nova scheduler (Newton release) when provisioning multiple instances in parallel (not in one batch api call)
00:14:09 pooja The scheduler's view of host isn't updated and so multiple instances get placed on a host, which exceeds the max_instances value set for that host.
00:14:31 pooja Is this a known issue and is there a solution for it in ocata/pike release?
00:15:11 pooja Appreciate any pointers or change links. Thanks!
00:16:17 melwitt pooja: are you running a single scheduler? there was a change in pike to do resource claims in the scheduler via placement. I'm not yet familiar with the "max_instances" value you mentioned though
00:16:51 melwitt okay, so it's a config option
00:21:03 pooja melwitt: Yes, I'm running a single instance of nova-scheduler
00:21:53 pooja This is the filter I'm referring to - https://github.com/openstack/nova/blob/master/nova/scheduler/filters/num_instances_filter.py#L28
00:23:34 melwitt I found it too, currently looking through the code. it looks like the problem you described should be fixed as of pike with the claims in the scheduler. let me see if I can find a patch related to that specific area
00:24:08 pooja Great! thanks for your help looking into it, melwitt!
00:25:28 Spaz-Home Morning
00:30:39 melwitt pooja: to be honest, I'm not sure if the issue is fixed as of the new code. to be sure, it would be better to ask someone like bauzas or edleafe. here's a link to where I started tracing, if that might help in the meantime https://github.com/openstack/nova/blob/master/nova/scheduler/host_manager.py#L283-L284
00:33:50 melwitt and the consume_from_request is used in filter_scheduler.py, looking at it more in filter_scheduler.py, it does seem like it would not be resilient to the issue of parallel requests
00:35:34 melwitt it looks like the report of "num_instances" comes from a compute node stat report, which may not be updating in real-time
00:48:56 pooja melwitt: Sure, let me check with bauzas or edleafe too.
00:49:50 pooja Yes, the problem is that stats get updated asynchronously and num_instances value used by scheduler would be inaccurate based on that timing.
00:50:35 melwitt pooja: I see. I think I understand now, and based on that, it's probably still a problem in the current code now
00:51:04 pooja Oh okay.. should I file a bug for it?
00:51:27 pooja Do these filters work the same way with the new Placement API?
00:52:46 melwitt some do, some don't. the [Core|Ram|Disk]Filter became obsolete in the filter scheduler once we started calling placement. because we pre-filter based on answers from placement and the scheduler claims with placement along the way
00:54:09 melwitt but the rest of the filters run after the placement call, as they did before
00:56:17 melwitt again, bauzas, edleafe, and co are the people to chat with about that
00:56:47 melwitt if you want to open a bug, I can point them to it and ask them to comment on it. it's up to you, however you want to do it
01:09:48 pooja Sounds good! I will connect with them and see if I need to file a bug for this. Will let you know if I do that. Thanks again!
01:09:59 melwitt cool, thanks pooja
01:59:43 openstackgerrit Merged openstack/nova master: VMware: fix TypeError while get console log https://review.openstack.org/549182
02:52:27 alex_xu_ jaypipes: what do you think about the preferred_traits, that resovled the cyborg weigher problem
04:28:50 openstackgerrit Merged openstack/nova master: Add placeholder migrations for Queens backports https://review.openstack.org/553831
05:18:25 openstackgerrit Naichuan Sun proposed openstack/nova master: xenapi: Use XAPI pool instead of aggregate pool for shared SR migration https://review.openstack.org/554154
05:27:04 openstackgerrit Naichuan Sun proposed openstack/nova master: xenapi: Use XAPI pool instead of aggregate pool for shared SR migration https://review.openstack.org/554154
06:03:36 openstackgerrit Pranab proposed openstack/os-vif master: Add abstract OVSDB API https://review.openstack.org/476612
06:10:44 openstackgerrit OpenStack Proposal Bot proposed openstack/nova master: Imported Translations from Zanata https://review.openstack.org/548772
06:15:04 openstackgerrit Jianghua Wang proposed openstack/nova master: libvirt: Improve 'qemu-img convert' performance https://review.openstack.org/522067
07:45:51 openstackgerrit Michael Still proposed openstack/nova master: Move xenapi partition copies to privsep. https://review.openstack.org/553605
07:45:51 openstackgerrit Michael Still proposed openstack/nova master: Move xenapi disk resizing to privsep. https://review.openstack.org/552242
07:45:52 openstackgerrit Michael Still proposed openstack/nova master: We no longer need rootwrap. https://review.openstack.org/554438
07:45:52 openstackgerrit Michael Still proposed openstack/nova master: Move image conversion to privsep. https://review.openstack.org/554437
07:45:53 openstackgerrit Michael Still proposed openstack/nova master: We don't need utils.trycmd any more. https://review.openstack.org/554439
08:08:45 openstackgerrit Zhenyu Zheng proposed openstack/nova-specs master: Allow abort live migrations in queued status https://review.openstack.org/536722
08:22:44 openstackgerrit sahid proposed openstack/nova master: Revert "[libvirt] Add _get_vcpu_realtime_scheduler()" https://review.openstack.org/554448
08:22:55 openstackgerrit sahid proposed openstack/nova master: Revert "[libvirt] Add _get_numa_memnode()" https://review.openstack.org/554449
08:30:42 sahid jaypipes: can you have a look at https://review.openstack.org/#/c/511188/
08:30:58 openstackgerrit licanwei proposed openstack/nova master: Make nova-manage capable of syncing all cell databases https://review.openstack.org/519275
08:31:31 sahid i mean it would be nice if you cn cut on it we can make some progress, basically you brought that idea and we never see you again :)

Earlier   Later