Earlier  
Posted Nick Remark
#openstack-nova - 2018-09-05
17:40:14 jroll mgagne: ah, cool, so I guess what I would do is add some sort of "NoNicsAvailable" exception that the ironic driver can return in this case
17:40:29 melwitt thanks, was just looking for that
17:41:43 melwitt indeed, call is synchronous
17:42:23 dansmith this would be a call to compute which does an http call to ironic, yeah?
17:42:28 mgagne jroll: there is already an exception for that (NoFreePhysicalPorts which is mapped to Invalid) https://github.com/openstack/ironic/blob/master/ironic/drivers/modules/network/common.py#L170-L174
17:43:46 melwitt dansmith: I think they'd add something to the already-existing synchronous attach_interface call to call ironic
17:43:56 mgagne compute does a call to virt driver which performs the HTTP call: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1450-L1451
17:44:42 jroll mgagne: oh, I meant add that exception to nova, so we can return it here: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1474
17:44:54 dansmith yeah the unfortunate bit there is that if the second call takes a long time, we'll time out, report to the user that it failed, but eventually it succeeded
17:45:17 dansmith but, maybe this would be a good application of my long-call stuff
17:45:18 jroll or add VirtualInterfacePlugException to the list here: https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/attach_interfaces.py#L134
17:45:19 mgagne jroll: yes! just need a way to detect this kind of error from Ironic API response and map accordingly
17:45:56 melwitt dansmith: ack
17:45:57 jroll mgagne: yep, this is where someone says all API errors should have a "code" :)
17:46:41 mgagne jroll: it's mapped to InterfaceAttachFailed in the compute manager https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5956
17:46:51 mgagne which is handled already here https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/attach_interfaces.py#L149
17:47:29 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix formatting in changes-since guide https://review.openstack.org/600150
17:47:56 jroll mgagne: ah, ok. so we should detect this specific case and make a new exception with better info for the user eh?
17:48:03 mgagne jroll: yep =)
17:48:52 mgagne I'm just not sure how to do it properly without parsing error message which could theoretically be translated to a different language than english
17:49:03 mgagne I have to run right now, brb in ~1h
17:49:11 jroll gotcha
17:49:37 jroll mgagne: might be more of a topic for the ironic channel when you're back
17:53:36 sean-k-mooney jroll: i was away but i did mention having a 400 with and embded code :P
17:54:10 jroll sean-k-mooney: see, someone said it before I even predicted it :P
17:54:23 sean-k-mooney but if my memory is correct there was some converstation about makeing the apis more consitent in vancouver
17:54:38 sean-k-mooney e.g. between different services
17:54:47 sean-k-mooney i dont know hat the outcome of that was
17:55:40 sean-k-mooney the only thing i do remember is never embed html in and error responce
17:55:56 jroll yes, there was
17:56:05 jroll people agree we should do it, nobody has set aside the time afaik
17:56:50 sean-k-mooney is it a topic at the ptg. i assumed it would continue to be discussed in that api working group or in one of the cross project sessions
18:02:52 jroll sean-k-mooney: dunno, haven't been following it closely
18:15:35 mriedem i hope not
18:17:23 mriedem but maybe it'll be on tuesday when i might have to sit in the kata containers qemu room all day because someone from my company has to know what's going on in that "community"
18:21:32 cdent sean-k-mooney: feel free to come to api-sig session on monday if you wanna talk about that. the agenda is pretty light at the moment (although we intend to flesh it out more during tomorrow's meeting): https://etherpad.openstack.org/p/api-sig-stein-ptg
19:04:13 sean-k-mooney cdent: am i might, i think "starardising how we report errors" is a good goal. not sure i know enough about the options to add much to the conversation
19:08:01 cdent sean-k-mooney: there's an existing api guideline for errors: http://specs.openstack.org/openstack/api-wg/guidelines/errors.html but it is not implemented by much (placement does a bit of it, mostly to use 'code' to distinguish different 409 responses
19:13:54 sean-k-mooney oh ok i should read that.
19:14:24 sean-k-mooney by the way im glad ye used status code 418 in the example :)
19:18:10 sean-k-mooney mgagne: you proably should check out the link cdent posted http://specs.openstack.org/openstack/api-wg/guidelines/errors.html
19:18:34 sean-k-mooney anyway i need to go sleep/pack. talk to everyone tomorrow
19:19:01 mgagne sean-k-mooney: looks like a very cool spec =)
20:20:08 openstackgerrit Matt Riedemann proposed openstack/nova master: Fix evacuate logging https://review.openstack.org/593055
20:53:51 mriedem UUIDs or accept ValueErrors for invalid UUIDs. See https://docs.openstack.org/oslo.versionedobjects/latest/reference/fields.html#oslo_versionedobjects.fields.UUIDField for further details
20:53:51 mriedem FutureWarning: ImageMeta(checksum=<?>,container_format=<?>,created_at=<?>,direct_url=<?>,disk_format=<?>,id=<?>,min_disk=<?>,min_ram=<?>,name=<?>,owner=<?>,properties=ImageMetaProps,protected=<?>,size=<?>,status=<?>,tags=<?>,updated_at=<?>,virtual_size=<?>,visibility=<?>) is an invalid UUID. Using UUIDFields with invalid UUIDs is no longer supported, and will be removed in a future release. Please update your code to input va
20:53:51 mriedem ugh i thought we had squashed this
20:54:28 mriedem i guess an ImageMeta object is certainly not a uuid
20:58:30 melwitt yeah... is it an error in a test or something? why is an ImageMeta object being treated as a UUIDField
20:59:19 mriedem missing mock i think
21:06:19 cdent Is "thin provisioning" of disk mostly a vmware thing, or does it also happen when using some other hypervisor+storage things?
21:08:16 dansmith cdent: other things have it
21:08:30 dansmith not everything amazing is made by vmware. JEEZ
21:09:04 dansmith a unix sparse file is kinda thing provisioning and pre-dates a lot of stuff
21:09:10 dansmith like 3.5" floppies
21:09:56 cdent the reason I ask is I'm wondering how appropriate it is to consider dynamically adjusting allocation_ratio on DISK_GB in that sort of setting: compare actual versus perceived usage
21:09:57 dansmith *thin
21:10:21 dansmith cdent: well, it's not really the same thing
21:10:24 dansmith it's similar
21:10:29 cdent also: dansmith you know I hate everything, so I don't think anything from vmware is amazing
21:10:39 dansmith but technically thin provisioning can blow up in your face if you overcommit
21:10:47 cdent right, that's exactly the issue
21:10:50 dansmith I do know that.
21:11:22 dansmith IMHO, allocation_ratio can only be over 1.0 for disk if you like to gamble
21:11:34 dansmith and shouldn't really be related to use of thin provisioning,
21:11:42 dansmith except in that it may be the mechanism by which you gamble
21:12:00 cdent the nearby concrete problem here is: datastore with 50TB total, allocated to 50TB, but in reality has 41TB free because of "thing provisioning"
21:12:19 cdent so the short term fix is bump allocation_ratio and watch real usage "real close like"
21:12:19 dansmith but yep
21:12:27 dansmith yeah, i.e. gambling
21:12:33 cdent that was not mockery, I just can't type
21:12:37 cdent which you also know
21:12:50 dansmith so, every night before the op goes to sleep,
21:13:07 dansmith he decides how far he thinks he can get before 8am the next morning and sets the allocation_ratio accordingly
21:13:13 cdent pretty much
21:13:32 dansmith solid plan
21:14:40 cdent so the questioning is: if you know what placement thinks about usage, and you know what the datastore thinks of its free space, you ought to be able to calculate a dynamic allocation ratio every now and again, and save that op some sleep
21:15:12 dansmith if your usage is very consistent
21:15:23 dansmith the other thing to think about/remember is:
21:15:52 dansmith thin provisioning is somewhat of a space-saving thing, but it's also just to speed up the actual provisioning step
21:16:06 dansmith most filesystems will end up filling out the block device over time,
21:16:29 cdent yeah, which is why it would need to be a failure regular thing
21:16:33 dansmith unless they intentionally compact and fstrim() now and then to let go of things back to the underlying storage
21:16:36 cdent like maybe every update_provider_tree
21:17:00 cdent that was an awesome freudian slide on my part: meant fairly, but failure works too
21:17:16 dansmith but if you go to sleep with it set at 1.1,
21:17:29 dansmith and something happens where a big tenant has a script that runs amok and fills the disk with logs,
21:17:37 dansmith you could end up getting a page
21:17:55 dansmith so it's still a ticking time bomb, even if you calculate a conservative ratio based on recent history
21:18:24 cdent I'm not sure I'm following that logic: at time 1 we set it to 1.1. a few minutes later the actual usage on the store goes up, so at time 2 we set it to 1.0
21:18:48 cdent if 1 and 2 are close together the risk is lessed (but not entirely removed)
21:19:00 dansmith you mean some automated thing that looks for the crash on the horizon and adjusts the ratio before it becomes a problem?
21:21:34 dansmith I'm trying to think about how that helps
21:21:36 cdent in this vmware case the datastore knows how many bytes it says it has "free". this is different from what placement says. The ratio of that difference can be a multiplier for setting allocation_ratio. If you're saying "the datastores sense of 'free' is not real" then yeah, sure, we got a problem.
21:21:46 dansmith until you get to 1.0, it might as well be 1.0, and once you're past 1.0 you can't go back
21:22:10 dansmith it's free to the datastore, but it's not "uncommitted"
21:22:58 cdent I think that may be where the vmware stuff is actually doing something "amazing" that is a bit different from sparse files
21:23:37 dansmith the other two things that it can be doing are compression and dedup
21:23:48 dansmith both of which have an optimal ratio, but that aren't consistent
21:24:23 dansmith like, if your users are into storing usenet archives, you might get 1.5x from compression, but if they're archiving "those" usenet articles, it'll be much closer to 1.0x

Earlier   Later