| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-12-06 | |||
| 15:15:55 | mriedem | there is only one image field we deal with the schema and that's disk_format | |
| 15:16:20 | mriedem | https://review.openstack.org/#/c/375875/ | |
| 15:19:53 | diliprenkila | mriedem: so we should fix the os_hidden type in nova ? not in glance | |
| 15:20:56 | openstackgerrit | Silvan Kaiser proposed openstack/nova master: Exec systemd-run without --user flag in Quobyte driver https://review.openstack.org/554195 | |
| 15:21:59 | mriedem | diliprenkila: i don't think there is probably anything to change in glance, | |
| 15:22:14 | mriedem | but i'd like to know why tempest isn't failing with this, but i need to know the recreate steps, | |
| 15:22:15 | openstackgerrit | Chris Dent proposed openstack/nova master: Correct lower-constraints.txt and the related tox job https://review.openstack.org/622972 | |
| 15:22:30 | mriedem | because tempest has very basic tests where it creates a server and then creates a snapshot of that server, | |
| 15:22:40 | mriedem | which i would think should cause this failure | |
| 15:26:27 | mriedem | maybe tempest isn't using image api v2.7/ | |
| 15:26:29 | mriedem | ? | |
| 15:27:59 | diliprenkila | mridem: may be | |
| 15:28:06 | artom | mriedem, I think by default it goes to the lowest microversion | |
| 15:28:21 | mriedem | glance doesn't have microversions but... | |
| 15:28:22 | artom | Unless the test specifies min_microversion (or just microversion?) | |
| 15:28:33 | artom | Oh, so just endpoints? | |
| 15:28:56 | mriedem | i need a mordred | |
| 15:29:14 | mordred | I didn't do it | |
| 15:29:19 | mriedem | image api versions, go! | |
| 15:29:23 | mriedem | as in, wtf | |
| 15:29:25 | artom | You're holding out for a mordred 'till the end of the night | |
| 15:29:30 | mriedem | does the user opt into those or you just get what the server has available? | |
| 15:29:31 | diliprenkila | mriedem: i am using nova: 18.0.0 , glance: 2.9.1 | |
| 15:29:54 | mordred | they're silly. there is no selection - you just get the API described by the highest number in that list | |
| 15:30:06 | mordred | so - basically - ignore the thing after the 2. | |
| 15:30:17 | mriedem | so if glance is rocky, i get 2.7 | |
| 15:30:22 | mriedem | https://developer.openstack.org/api-ref/image/versions/index.html#version-history | |
| 15:30:26 | mordred | yeah | |
| 15:30:40 | mriedem | then i don't know why tempest would not fail on https://bugs.launchpad.net/nova/+bug/1806239 | |
| 15:30:40 | openstack | Launchpad bug 1806239 in OpenStack Compute (nova) "nova-api should handle type conversion while creating server snapshots " [Undecided,New] | |
| 15:31:07 | mordred | nova is still using glanceclient right? | |
| 15:31:12 | mriedem | yeah | |
| 15:31:33 | mordred | yeah - tempest uses direct rest calls - so it's possible glanceclient is doing something wrong. or tempest is doing something wrong | |
| 15:32:29 | mriedem | well, tempest would just create a server and tell nova to snapshot it | |
| 15:32:33 | mriedem | and then nova will use glanceclient | |
| 15:32:52 | mriedem | if that's all it takes to tickle this with rocky glance, i'm not sure why tempest wouldn't blow up | |
| 15:33:37 | mriedem | anyway, i'll wait for diliprenkila to provide recreate steps | |
| 15:33:59 | mordred | oh. gotcha | |
| 15:34:05 | mordred | yeah. that's super weird | |
| 15:35:10 | ShilpaSD | Hi All, facing issue while stack, E: Sub-process /usr/bin/dpkg returned an error code (1), any suggestions to resolve this? | |
| 15:37:28 | dansmith | mriedem: so on that rpc logging thing at startup, do we think the actual query is slowing things down, or the logging of the pointless message? | |
| 15:37:47 | mriedem | idk | |
| 15:38:26 | mriedem | looking at logs | |
| 15:39:20 | mriedem | well we take 2 seconds dumping our gd config options :) | |
| 15:39:35 | dansmith | which should be mostly just log traffic, right? | |
| 15:39:40 | mriedem | yeah | |
| 15:39:41 | dansmith | so maybe the logging is the thing | |
| 15:39:42 | dansmith | also | |
| 15:39:44 | mriedem | we start loading extensions at Dec 05 20:14:00.919520 | |
| 15:40:03 | dansmith | you know that if we were to actually start compute first, we would cache the service version the way we want and avoid the multiple lookups | |
| 15:40:24 | mriedem | looks like we're done loading extensions at Dec 05 20:14:27.718587 | |
| 15:40:43 | mriedem | so there is another thing here, | |
| 15:40:49 | mriedem | the rpcapi client does the version query thing, | |
| 15:41:04 | mriedem | but API also constructs a SchedulerReportClient per instance, which apparently uses a lock | |
| 15:41:04 | mriedem | Dec 05 20:14:27.687766 ubuntu-xenial-ovh-bhs1-0000959981 devstack@n-api.service[23459]: DEBUG oslo_concurrency.lockutils [None req-dfdfad07-2ff4-43ed-9f67-2acd59687e0c None None] Lock "placement_client" acquired by "nova.scheduler.client.report._create_client" :: waited 0.000s {{(pid=23462) inner /usr/local/lib/python2.7/dist-packages/oslo_concurrency/lockutils.py:327}} | |
| 15:41:49 | mriedem | so... | |
| 15:41:55 | dansmith | I also just noticed/remembered that the multi-cell version of this does not cache at all | |
| 15:41:57 | mriedem | and that's an in-memory lock | |
| 15:44:28 | mriedem | heh remember how efried_cya_jan removed that lazy-load scheduler report client stuff? https://github.com/openstack/nova/blob/master/nova/compute/api.py#L256 | |
| 15:44:45 | dansmith | yeah | |
| 15:44:51 | efried_cya_jan | uh oh | |
| 15:45:17 | dansmith | mriedem: is there a warn-once pattern I should be using for logs? | |
| 15:46:23 | mriedem | dansmith: i think just a global | |
| 15:46:31 | dansmith | ack | |
| 15:46:48 | mriedem | i'll strip out a separate bug for this report client init thing | |
| 15:51:01 | mriedem | https://bugs.launchpad.net/nova/+bug/1807219 | |
| 15:51:01 | openstack | Launchpad bug 1807219 in OpenStack Compute (nova) "SchedulerReporClient init slows down nova-api startup" [Medium,Triaged] | |
| 15:52:22 | cdent | mdbooth: is that ^ related at all to the slow down you were seeing in your explorations yesterday(?)? | |
| 15:53:20 | efried_cya_jan | mriedem: We ought to singleton that guy. If we're not having caching conflicts, it can only be out of luck because the API obj isn't doing anything that touches the cache. | |
| 15:53:47 | edleafe | efried_cya_jan: more like efried_big_fat_liar | |
| 15:54:45 | efried_cya_jan | I'm here for like another 20 minutes today. This does two things: 1) keeps my inbox down to triple digits; 2) makes you scared to talk about me behind my back. | |
| 15:55:03 | mriedem | i intentionally talked about you to summon you | |
| 15:55:06 | mdbooth | cdent: Only if we also do this during server create. Possible. The number of API objects created was insane. | |
| 15:55:08 | kaisers_ | stephenfin: mdbooth: Updated https://review.openstack.org/#/c/554195/ as discussed yesterday | |
| 15:55:15 | edleafe | Oh, it's more fun to talk about you to your face! | |
| 15:56:23 | efried_cya_jan | mriedem: Is the removal of lazy-load causing (or even related to) that one-reportclient-per-API? | |
| 15:57:59 | mriedem | efried_cya_jan: not sure, we might have always been loading it during API init | |
| 15:58:41 | mriedem | but i think it would have done the lazy-load | |
| 15:58:52 | mriedem | since LazyLoader only creates the client thing until something is accessed on it | |
| 15:59:10 | efried_cya_jan | And it's only expensive because it's serialized, not because the report client is doing anything heavy, right? | |
| 15:59:29 | mriedem | yeah i think so | |
| 16:00:00 | mriedem | https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L260 | |
| 16:00:21 | mriedem | we create 2 provider trees in there for some reason... | |
| 16:00:26 | mriedem | https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L270 | |
| 16:00:31 | mriedem | https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L279 | |
| 16:00:35 | mriedem | oops | |
| 16:00:40 | mriedem | https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L286 | |
| 16:01:15 | mriedem | is there any good reason we need to do that in both places? | |
| 16:02:33 | efried_cya_jan | mriedem: There should never be a need for one process to create two separate provider trees <=> report client instances, period. | |
| 16:02:48 | efried_cya_jan | That could only ever lead to bugs. | |
| 16:02:57 | efried_cya_jan | It will only ever *not* lead to bugs by luck. | |
| 16:03:12 | efried_cya_jan | hold on, looking at your links... | |
| 16:03:24 | mriedem | so, i'm going to remove this https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L286-L287 | |
| 16:03:35 | mriedem | and change https://github.com/openstack/nova/blob/c9dca64fa64005e5bea327f06a7a3f4821ab72b1/nova/scheduler/client/report.py#L270-L272 | |
| 16:03:39 | mriedem | to just call clear_provider_cache | |
| 16:03:45 | mriedem | for great success | |
| 16:03:47 | mriedem | ok?! | |
| 16:03:56 | mriedem | butthole surfers got me all amped up | |
| 16:05:18 | efried_cya_jan | mriedem: That will be fine, it's removing redundant calls, but it's not going to change anything functionally. We were calling ProviderTree init twice before, but not creating two separate provider trees, just overwriting self._provider_tree. | |
| 16:06:06 | efried_cya_jan | mriedem: However, I don't agree with the details. You can't get rid of _create_client without affecting @safe_connect. For now (until @safe_connect has died a fiery death), I would just remove the provider tree and assoc timer inits from __init__. | |