| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-09-05 | |||
| 17:38:57 | mgagne | https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/attach_interfaces.py#L131-L154 | |
| 17:40:02 | mgagne | https://github.com/openstack/nova/blob/master/nova/compute/rpcapi.py#L465-L474 | |
| 17:40:07 | mgagne | call() is used | |
| 17:40:14 | jroll | mgagne: ah, cool, so I guess what I would do is add some sort of "NoNicsAvailable" exception that the ironic driver can return in this case | |
| 17:40:29 | melwitt | thanks, was just looking for that | |
| 17:41:43 | melwitt | indeed, call is synchronous | |
| 17:42:23 | dansmith | this would be a call to compute which does an http call to ironic, yeah? | |
| 17:42:28 | mgagne | jroll: there is already an exception for that (NoFreePhysicalPorts which is mapped to Invalid) https://github.com/openstack/ironic/blob/master/ironic/drivers/modules/network/common.py#L170-L174 | |
| 17:43:46 | melwitt | dansmith: I think they'd add something to the already-existing synchronous attach_interface call to call ironic | |
| 17:43:56 | mgagne | compute does a call to virt driver which performs the HTTP call: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1450-L1451 | |
| 17:44:42 | jroll | mgagne: oh, I meant add that exception to nova, so we can return it here: https://github.com/openstack/nova/blob/master/nova/virt/ironic/driver.py#L1474 | |
| 17:44:54 | dansmith | yeah the unfortunate bit there is that if the second call takes a long time, we'll time out, report to the user that it failed, but eventually it succeeded | |
| 17:45:17 | dansmith | but, maybe this would be a good application of my long-call stuff | |
| 17:45:18 | jroll | or add VirtualInterfacePlugException to the list here: https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/attach_interfaces.py#L134 | |
| 17:45:19 | mgagne | jroll: yes! just need a way to detect this kind of error from Ironic API response and map accordingly | |
| 17:45:56 | melwitt | dansmith: ack | |
| 17:45:57 | jroll | mgagne: yep, this is where someone says all API errors should have a "code" :) | |
| 17:46:41 | mgagne | jroll: it's mapped to InterfaceAttachFailed in the compute manager https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L5956 | |
| 17:46:51 | mgagne | which is handled already here https://github.com/openstack/nova/blob/master/nova/api/openstack/compute/attach_interfaces.py#L149 | |
| 17:47:29 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Fix formatting in changes-since guide https://review.openstack.org/600150 | |
| 17:47:56 | jroll | mgagne: ah, ok. so we should detect this specific case and make a new exception with better info for the user eh? | |
| 17:48:03 | mgagne | jroll: yep =) | |
| 17:48:52 | mgagne | I'm just not sure how to do it properly without parsing error message which could theoretically be translated to a different language than english | |
| 17:49:03 | mgagne | I have to run right now, brb in ~1h | |
| 17:49:11 | jroll | gotcha | |
| 17:49:37 | jroll | mgagne: might be more of a topic for the ironic channel when you're back | |
| 17:53:36 | sean-k-mooney | jroll: i was away but i did mention having a 400 with and embded code :P | |
| 17:54:10 | jroll | sean-k-mooney: see, someone said it before I even predicted it :P | |
| 17:54:23 | sean-k-mooney | but if my memory is correct there was some converstation about makeing the apis more consitent in vancouver | |
| 17:54:38 | sean-k-mooney | e.g. between different services | |
| 17:54:47 | sean-k-mooney | i dont know hat the outcome of that was | |
| 17:55:40 | sean-k-mooney | the only thing i do remember is never embed html in and error responce | |
| 17:55:56 | jroll | yes, there was | |
| 17:56:05 | jroll | people agree we should do it, nobody has set aside the time afaik | |
| 17:56:50 | sean-k-mooney | is it a topic at the ptg. i assumed it would continue to be discussed in that api working group or in one of the cross project sessions | |
| 18:02:52 | jroll | sean-k-mooney: dunno, haven't been following it closely | |
| 18:15:35 | mriedem | i hope not | |
| 18:17:23 | mriedem | but maybe it'll be on tuesday when i might have to sit in the kata containers qemu room all day because someone from my company has to know what's going on in that "community" | |
| 18:21:32 | cdent | sean-k-mooney: feel free to come to api-sig session on monday if you wanna talk about that. the agenda is pretty light at the moment (although we intend to flesh it out more during tomorrow's meeting): https://etherpad.openstack.org/p/api-sig-stein-ptg | |
| 19:04:13 | sean-k-mooney | cdent: am i might, i think "starardising how we report errors" is a good goal. not sure i know enough about the options to add much to the conversation | |
| 19:08:01 | cdent | sean-k-mooney: there's an existing api guideline for errors: http://specs.openstack.org/openstack/api-wg/guidelines/errors.html but it is not implemented by much (placement does a bit of it, mostly to use 'code' to distinguish different 409 responses | |
| 19:13:54 | sean-k-mooney | oh ok i should read that. | |
| 19:14:24 | sean-k-mooney | by the way im glad ye used status code 418 in the example :) | |
| 19:18:10 | sean-k-mooney | mgagne: you proably should check out the link cdent posted http://specs.openstack.org/openstack/api-wg/guidelines/errors.html | |
| 19:18:34 | sean-k-mooney | anyway i need to go sleep/pack. talk to everyone tomorrow | |
| 19:19:01 | mgagne | sean-k-mooney: looks like a very cool spec =) | |
| 20:20:08 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Fix evacuate logging https://review.openstack.org/593055 | |
| 20:53:51 | mriedem | ugh i thought we had squashed this | |
| 20:53:51 | mriedem | FutureWarning: ImageMeta(checksum=<?>,container_format=<?>,created_at=<?>,direct_url=<?>,disk_format=<?>,id=<?>,min_disk=<?>,min_ram=<?>,name=<?>,owner=<?>,properties=ImageMetaProps,protected=<?>,size=<?>,status=<?>,tags=<?>,updated_at=<?>,virtual_size=<?>,visibility=<?>) is an invalid UUID. Using UUIDFields with invalid UUIDs is no longer supported, and will be removed in a future release. Please update your code to input va | |
| 20:53:51 | mriedem | UUIDs or accept ValueErrors for invalid UUIDs. See https://docs.openstack.org/oslo.versionedobjects/latest/reference/fields.html#oslo_versionedobjects.fields.UUIDField for further details | |
| 20:54:28 | mriedem | i guess an ImageMeta object is certainly not a uuid | |
| 20:58:30 | melwitt | yeah... is it an error in a test or something? why is an ImageMeta object being treated as a UUIDField | |
| 20:59:19 | mriedem | missing mock i think | |
| 21:06:19 | cdent | Is "thin provisioning" of disk mostly a vmware thing, or does it also happen when using some other hypervisor+storage things? | |
| 21:08:16 | dansmith | cdent: other things have it | |
| 21:08:30 | dansmith | not everything amazing is made by vmware. JEEZ | |
| 21:09:04 | dansmith | a unix sparse file is kinda thing provisioning and pre-dates a lot of stuff | |
| 21:09:10 | dansmith | like 3.5" floppies | |
| 21:09:56 | cdent | the reason I ask is I'm wondering how appropriate it is to consider dynamically adjusting allocation_ratio on DISK_GB in that sort of setting: compare actual versus perceived usage | |
| 21:09:57 | dansmith | *thin | |
| 21:10:21 | dansmith | cdent: well, it's not really the same thing | |
| 21:10:24 | dansmith | it's similar | |
| 21:10:29 | cdent | also: dansmith you know I hate everything, so I don't think anything from vmware is amazing | |
| 21:10:39 | dansmith | but technically thin provisioning can blow up in your face if you overcommit | |
| 21:10:47 | cdent | right, that's exactly the issue | |
| 21:10:50 | dansmith | I do know that. | |
| 21:11:22 | dansmith | IMHO, allocation_ratio can only be over 1.0 for disk if you like to gamble | |
| 21:11:34 | dansmith | and shouldn't really be related to use of thin provisioning, | |
| 21:11:42 | dansmith | except in that it may be the mechanism by which you gamble | |
| 21:12:00 | cdent | the nearby concrete problem here is: datastore with 50TB total, allocated to 50TB, but in reality has 41TB free because of "thing provisioning" | |
| 21:12:19 | dansmith | but yep | |
| 21:12:19 | cdent | so the short term fix is bump allocation_ratio and watch real usage "real close like" | |
| 21:12:27 | dansmith | yeah, i.e. gambling | |
| 21:12:33 | cdent | that was not mockery, I just can't type | |
| 21:12:37 | cdent | which you also know | |
| 21:12:50 | dansmith | so, every night before the op goes to sleep, | |
| 21:13:07 | dansmith | he decides how far he thinks he can get before 8am the next morning and sets the allocation_ratio accordingly | |
| 21:13:13 | cdent | pretty much | |
| 21:13:32 | dansmith | solid plan | |
| 21:14:40 | cdent | so the questioning is: if you know what placement thinks about usage, and you know what the datastore thinks of its free space, you ought to be able to calculate a dynamic allocation ratio every now and again, and save that op some sleep | |
| 21:15:12 | dansmith | if your usage is very consistent | |
| 21:15:23 | dansmith | the other thing to think about/remember is: | |
| 21:15:52 | dansmith | thin provisioning is somewhat of a space-saving thing, but it's also just to speed up the actual provisioning step | |
| 21:16:06 | dansmith | most filesystems will end up filling out the block device over time, | |
| 21:16:29 | cdent | yeah, which is why it would need to be a failure regular thing | |
| 21:16:33 | dansmith | unless they intentionally compact and fstrim() now and then to let go of things back to the underlying storage | |
| 21:16:36 | cdent | like maybe every update_provider_tree | |
| 21:17:00 | cdent | that was an awesome freudian slide on my part: meant fairly, but failure works too | |
| 21:17:16 | dansmith | but if you go to sleep with it set at 1.1, | |
| 21:17:29 | dansmith | and something happens where a big tenant has a script that runs amok and fills the disk with logs, | |
| 21:17:37 | dansmith | you could end up getting a page | |
| 21:17:55 | dansmith | so it's still a ticking time bomb, even if you calculate a conservative ratio based on recent history | |
| 21:18:24 | cdent | I'm not sure I'm following that logic: at time 1 we set it to 1.1. a few minutes later the actual usage on the store goes up, so at time 2 we set it to 1.0 | |
| 21:18:48 | cdent | if 1 and 2 are close together the risk is lessed (but not entirely removed) | |
| 21:19:00 | dansmith | you mean some automated thing that looks for the crash on the horizon and adjusts the ratio before it becomes a problem? | |
| 21:21:34 | dansmith | I'm trying to think about how that helps | |
| 21:21:36 | cdent | in this vmware case the datastore knows how many bytes it says it has "free". this is different from what placement says. The ratio of that difference can be a multiplier for setting allocation_ratio. If you're saying "the datastores sense of 'free' is not real" then yeah, sure, we got a problem. | |
| 21:21:46 | dansmith | until you get to 1.0, it might as well be 1.0, and once you're past 1.0 you can't go back | |
| 21:22:10 | dansmith | it's free to the datastore, but it's not "uncommitted" | |
| 21:22:58 | cdent | I think that may be where the vmware stuff is actually doing something "amazing" that is a bit different from sparse files | |