Earlier  
Posted Nick Remark
#openstack-nova - 2020-04-28
14:54:04 gmann melwitt: all setup for devstack and grenade now. we can approve this - https://review.opendev.org/#/c/704364/5
14:54:08 gmann and then backport
14:54:26 dansmith artom: I think I'd rather the string filter in the query than just continually select every instance out of the database on every run, only to find that none are needed
14:54:28 gmann *done now
14:54:41 dansmith artom: I'd have to iterate every instance in the DB every time you asked me "are you done?"
14:55:06 artom dansmith, but isn't that what the string filter in the query is doing anyways, behind the scenes?
14:55:15 artom Looking at every instance in the database?
14:55:40 dansmith artom: no, it's looking at every value in a table
14:56:16 dansmith which is massively less expensive than selecting multiple tables, feeding those to python, constructing tons and tons of objects, for us to go through and do...that same string comparison
14:56:51 artom dansmith, right sorry - but what's what I was suggesting we do in python - select instance_uuid, numa_topology from instance_extra <in steps of 100>, fitler in python, use those instance_uuids for the actual update
14:57:17 artom But maybe ^^^ is impossible with our models/sqlalchemy?
14:57:31 dansmith you mean a direct query and not an ORM one? that's better, but you're still sending the results of all those queries across the network every time
14:57:52 dansmith no, we can do it without an ORM query, we do that in places
14:58:46 artom dansmith, yeah :/
14:59:19 artom I guess we'd need metrics of how long the string filter would take, as a function of # of instances
14:59:47 artom To see if we're making a mountain our of a molehill
14:59:59 dansmith it'll definitely be faster for the db server to do it, the only benefit I can see is you offload the cpu processing to another machine, but you pay the penalty of increasing latency and definitely network traffic
15:00:34 dansmith besides, stephenfin will want to nuke this migration in a cycle or two anyway, so it won't be run forever
15:01:33 gibi away
15:02:29 artom dansmith, I suppose asking CERN to test this for us is not an option ;)
15:03:15 dansmith artom: sure you can ask them for their opinion on the pain
15:03:23 dansmith I dunno what the alternative is though
15:03:31 artom Yeah, valid point
15:03:44 dansmith especially in cern's case, the impact of sending their massive amount of data over the network is not likely better
15:03:52 dansmith and, they have upgrade windows where they can handle this kind of thing
15:04:10 dansmith artom: well, one alternative is to avoid doing this, which is valid
15:04:21 dansmith artom: this is cleanup that we don't need to do, other than to address stephenfin's OCD
15:04:36 artom I tend to like stephenfin OCD ;)
15:04:38 dansmith which while this seems like a good one to do, because it's data at rest, it's not *really* costing us much to maintain the compatibility
15:05:02 artom dansmith, I don't suppose doing it on the fly is an option?
15:05:12 dansmith artom: we're already doing it on the fly, that's my point
15:05:19 artom Like, every time we call those legacy dict methods, update the DB?
15:05:28 artom Seriously?
15:05:40 artom Yeeeaaahhh
15:05:49 artom I have to say, that sways me the other way around
15:05:50 dansmith we're doing it on read, and if we save it for some reason (like a migration) then we update it
15:06:07 dansmith we could do update-on-read too, which we've done in other places
15:06:23 artom Update on read would be nice
15:06:33 dansmith I'd be cool with that
15:06:52 dansmith despite this looking nicer, it really isn't *necessary* to update this.. these compat routines have been here for ages
15:07:11 artom I'd be less worried if the hit to the DB was less
15:07:28 artom If jaypipes was still around he could set us straight
15:07:39 artom Maybe I'm paranoid for nothing
15:07:50 dansmith yeah, I think it's a good insight.. I honestly hadn't really considered it, but I think you've got a good point
15:08:08 dansmith if it was something we really had to do, then I'd be for this approach, let the DB server do the hard bit, but...
15:08:33 dansmith the other thing about this, is for instances created since like 2018, this is just pain for no reason
15:08:41 dansmith or whenever we moved to the new format
15:08:54 dansmith this is really just looking for super ancient instances that have been nursed for a long time,
15:09:00 dansmith which definitely puts it into perspective for me
15:09:12 stephenfin 2014
15:09:17 dansmith hah
15:12:27 artom stephenfin, your call I guess - looks like update on read might be a more acceptable solution, though with it we'd never be 100% sure we got all of those ancient instances
15:12:39 stephenfin And this isn't OCD. I've to maintain the hardware.py module and uglier corners of libvirt driver code, and I'm continuously trying to burn down tech debt inflicted upon me by others to make that job easier
15:13:39 openstackgerrit Merged openstack/nova master: Imported Translations from Zanata https://review.opendev.org/722644
15:14:12 stephenfin artom: personally, I was for deleting all the compat code to handle instances that hadn't been _touched_ since 2014, but dansmith said that wasn't good enough so I did this
15:14:23 stephenfin and have been working on it for two years
15:14:41 artom stephenfin, sorry dude, didn't mean to sabotage it :(
15:14:57 artom But it would have been dishonest to silence my concern
15:15:37 artom belmoreira, around? We're having a DB performance discussion, and figured CERN's scale might shed some light
15:17:21 belmoreira artom I just need to leave now. If you can I will be around tomorrow
15:17:47 artom belmoreira, ok, I'll ping you again earlier tomorrow morning (I'm on EDT)
15:20:27 artom I'm trying to read up on LIKE performance, which I *assume* is what sqlalchemy boils down .contains() to...
15:21:08 dansmith like is a little more expensive than a strstr() I think, although it might collapse to that if mysql notices there are no pattern characters
15:25:01 gibi dansmith, gmann, stephenfin, zigo: OK so I see that the change in the policy generator is not feasible before ussuri GA and might also seemed as a step backward to have that change in V. Then are we suggesting to zigo that his use case is not really valid / supported in Ussuri and in Victoria?
15:25:29 gibi or what is the way forward top of the reno update?
15:26:17 dansmith gibi: no, I think zigo's case is very valid, and common
15:26:55 dansmith gibi: I think we've screwed up here, and I think all I've said is that I don't think we have a lot of good options
15:27:45 gmann but should we consider the "new policy file generated without deprecated rules to switch to new defaults" are not valid ?
15:28:18 gmann because in both cases, policy file can be used that way
15:28:19 gibi dansmith: thanks. I feel that we have no good options
15:28:27 dansmith indeed :/
15:29:19 gibi gmann: do you mean that use the new generated file and configure nova and keystone with scopes?
15:29:36 gibi gmann: I think that is a valid use case
15:29:43 dansmith we can't do that in the upgrade I think
15:29:55 dansmith because people with non-scoped current policy files will break
15:30:04 gmann yes. and in past also if we deprecate any policy rule and operator want to move to new defalt, policy overwrite is the option they had
15:30:06 gibi dansmith: true, not for the upgrade, but for the new deployments
15:30:23 dansmith gibi: only distros have the ability to do something different for upgrade vs. new I think
15:30:33 dansmith gibi: we the nova project have to assume upgrade
15:30:40 gmann yeah
15:30:59 dansmith I think what we can do is set up for the most default case, which is where we are now, and accept that the case where you're using the generation tool is going to break
15:31:12 dansmith it's likely common, but less common than all the other cases I think
15:32:05 dansmith accept, and document/warn about the potential problem I mean
15:32:29 openstackgerrit Merged openstack/nova stable/ussuri: Imported Translations from Zanata https://review.opendev.org/723160
15:32:34 gibi OK, I see. thanks. I will summarize it in the bug
15:32:42 gmann dansmith: i agree it might be common to re-generate policy file and end up no deprecated rule but is not that wrong usage ? and some point we have to tell them its not right one
15:33:26 dansmith gmann: no, I don't think it's wrong.. it's not what we want people to do, but what we want them to do isn't very convenient or user-friendly, which is why I think it's likely common
15:33:48 dansmith gibi: also the other action item is to switch to yaml policy by default going forward I think, and encourage people to move to that
15:33:51 gmann and it is like it was not reported before when few policy default were changed. i am not sure if new default were superset of old so that same problem would occur
15:34:22 gmann dansmith: +1 on switch yaml.
15:34:30 gmann and this bug - https://bugs.launchpad.net/oslo.policy/+bug/1853170
15:34:30 openstack Launchpad bug 1853170 in oslo.policy "Need documentation on recommended operator workflow for deprecated policies" [High,Triaged]
15:34:30 gibi dansmith: good point about yaml, adding that to the PTG etherpad
15:34:30 dansmith gmann: this we can all agree on :P
15:34:50 gmann we need some consistent usage guide also
15:35:10 gmann gibi: its there, in cross project section
15:35:24 gibi gmann: even the yaml usage?
15:36:08 gmann gibi: ah i thought it is written in description but i missed. please add
15:36:29 gibi done :)
15:36:34 gmann thanks

Earlier   Later