| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-12-12 | |||
| 18:45:12 | mriedem | the empty vol attachment keeps the volume reserved but that vol attachment isn't actually connected to any host | |
| 18:45:16 | ildikov | melwitt: so when you call attachment_delete in terminate_connections that moves to volume back to 'available' | |
| 18:45:33 | mriedem | unless there is another attachment on the volume already | |
| 18:45:34 | ildikov | mriedem: do we call attachment_complete in every bit of shelve, where we need to? | |
| 18:45:42 | mriedem | once the attachment count is 0, then the volume is 'available' again | |
| 18:46:10 | mriedem | ildikov: you mean unshelve? | |
| 18:46:28 | mriedem | ildikov: yes via https://github.com/openstack/nova/blob/master/nova/compute/manager.py#L4783 | |
| 18:46:34 | mriedem | same as normal server create | |
| 18:46:53 | mriedem | happens down in the DriverVolumeBlockDevice.attach magic cauldron | |
| 18:47:15 | mriedem | magic cauldron of attachitude | |
| 18:47:24 | mriedem | henceforth shall it be known | |
| 18:47:35 | ildikov | mriedem: I meant whether we have actions as shelve and attaching a volume to a shelved instance and which call inside the Nova code covers these | |
| 18:47:56 | ildikov | mriedem: as we want the volume to be 'in-use' with the new flow as well so we added the additional attachment_complete call | |
| 18:48:05 | mriedem | attaching to a shelved offloaded instance happens in the api and we complete the attachment there too as you pointed out | |
| 18:48:37 | mriedem | remember that we can complete an empty attachment, | |
| 18:48:41 | mriedem | and update a completed attachment | |
| 18:48:48 | mriedem | weird as that may be | |
| 18:48:53 | mriedem | that's why unshelve works | |
| 18:49:07 | mriedem | *that's why unshelve works for the case that the volume is attached while the instance is shelved offloaded | |
| 18:51:49 | ildikov | mriedem: I understand that part | |
| 18:52:08 | ildikov | mriedem: I was just wondering whether we do all the steps when we shelve offload an instance wit volume attached | |
| 18:52:13 | ildikov | as I guess we can do that too | |
| 18:52:46 | ildikov | which might be the same set of calls under the hood, just wanted to double check after melwitt's confusion :) | |
| 18:54:33 | mriedem | yes, _prep_block_device happens it | |
| 18:54:36 | mriedem | on unshelve | |
| 18:54:40 | mriedem | *handles | |
| 18:56:24 | ildikov | ok, cool | |
| 18:56:50 | ildikov | thanks | |
| 19:04:05 | mriedem | sdague: where is the code for your api-ref burndown chart | |
| 19:04:06 | mriedem | ? | |
| 19:04:16 | mriedem | https://github.com/sdague/rst-burndown ? | |
| 19:05:17 | cdent | mriedem: gibi's sort of fork is at https://github.com/gibizer/nova-versioned-notification-transformation-burndown | |
| 19:05:24 | mriedem | yup, just found that too | |
| 19:05:27 | sdague | mriedem: yeh | |
| 19:05:41 | mriedem | mrhillsman is looking for something similar for sdk gap burndown | |
| 19:06:07 | cdent | I've got some totally 100% non-fancy infrastructure in place to run that stuff for gibi, if mrhillsman wants to get in on that: http://burndown.peermore.com/nova-notification/ | |
| 19:06:19 | cdent | it runs a cron job to update itself | |
| 19:06:43 | mriedem | i'll pass that on thanks | |
| 19:06:47 | cdent | but presumably if this is for openlap, they've already got such stuff | |
| 19:06:56 | cdent | openlap, hmmm | |
| 19:07:02 | cdent | if that's not already started, probably ought to be | |
| 19:07:31 | mriedem | yeah i'm assuming he can host it himself | |
| 19:08:51 | mriedem | er yeah http://openlabtesting.org/ | |
| 19:09:22 | cdent | even comes with uml | |
| 19:30:27 | mriedem | dansmith: want to check out the first 2 changes in the alternate hosts series here? https://review.openstack.org/#/c/516707/ - just needs final +2 | |
| 19:30:48 | dansmith | okay in a bit | |
| 19:39:22 | mriedem | for anyone that cares, the instance action paging series is also ready to go i think https://review.openstack.org/#/q/topic:bp/pagination-add-changes-since-for-instance-action-list+status:open | |
| 19:39:50 | mriedem | i've started moving fully complete series up to the top of https://etherpad.openstack.org/p/nova-queens-blueprint-status | |
| 20:15:49 | openstackgerrit | Matt Riedemann proposed openstack/nova master: [WIP] POC to use neutron port_list when filtering instance by ip https://review.openstack.org/525505 | |
| 20:18:43 | openstackgerrit | Michael Still proposed openstack/nova master: Convert ext filesystem resizes to privsep. https://review.openstack.org/517516 | |
| 20:18:44 | openstackgerrit | Michael Still proposed openstack/nova master: Start moving users of parted to privsep. https://review.openstack.org/519011 | |
| 20:18:44 | openstackgerrit | Michael Still proposed openstack/nova master: Move flushing block devices to privsep. https://review.openstack.org/519010 | |
| 20:18:45 | openstackgerrit | Michael Still proposed openstack/nova master: Convert users of tune2fs to privsep. https://review.openstack.org/519484 | |
| 20:18:45 | openstackgerrit | Michael Still proposed openstack/nova master: Move remaining uses of parted to privsep. https://review.openstack.org/519483 | |
| 20:18:46 | openstackgerrit | Michael Still proposed openstack/nova master: Move makefs to privsep https://review.openstack.org/527510 | |
| 20:51:45 | mnaser | throwing this up for discussion - how would folks feel if image caching used the imagebackend instead of just it's own cache | |
| 20:52:15 | mnaser | it would be beneficial for scenarios when you have a qcow2 image stored, with ceph backend for example, avoiding the import and convert on every iteration | |
| 20:53:01 | mnaser | another interesting scenario is multiple cells with independent ceph clusters (thereby, image is not directly accessible to do a cow new image) | |
| 20:53:55 | mnaser | (we're going to use cells v2 to add another az in an entirely different physical facility with a 10g link between them, fun times ahead :>) | |
| 20:56:26 | melwitt | mnaser: if you're using ceph and its native qcow2 format, you will be using glance images of type RAW and will get fast-clone and no image conversions, right? | |
| 20:56:48 | melwitt | that is, what do you mean by 'import and convert'? | |
| 20:56:55 | mnaser | melwitt: ceph with images stored in raw will get fast clone | |
| 20:57:33 | mnaser | melwitt: so - current scenario - images stored in raw in `images` pool in ceph, vm disks stored in `vms` pool in ceph. when nova boots a new vm, it can access `images` directly so it does a fast-clone and you get a disk very quickly | |
| 20:58:12 | mnaser | melwitt: if the image format is qcow2, ceph wont be able to do fast clone. nova will pull the image down locally, convert it to raw, and upload/import it back to ceph, then boot the machine | |
| 20:58:21 | melwitt | right | |
| 20:59:21 | mnaser | melwitt: now in an env where you have 2 cells, each one has a ceph cluster of its own, cell A which uses the same ceph cluster as glance can do fast clone. cell B will check the location and realize it has no way to reach that ceph cluster, so it will download the raw image locally and import/upload it to ceph | |
| 20:59:56 | mnaser | and for every new VM, it will do this (obviously leveraging local image cache), but that can be a lot of traffic before an image gets cached | |
| 21:00:39 | mnaser | my two ideas to work this around is... a) use the multiple locations feature of glance to mirror the images and add an rbd location in both clusters (nova loops over all locations till it finds one, then falls back to download) | |
| 21:00:46 | melwitt | from the ceph docs "Important Ceph doesn’t support QCOW2 for hosting a virtual machine disk. Thus if you want to boot virtual machines in Ceph (ephemeral backend or boot from volume), the Glance image format must be RAW." http://docs.ceph.com/docs/master/rbd/rbd-openstack/ | |
| 21:01:04 | melwitt | the cells thing is different though, and something to think about | |
| 21:01:23 | mnaser | or b) refactor image cache to use the image backend so that images can be cached in the ceph cluster, avoiding multiple downloads | |
| 21:01:49 | mnaser | melwitt: let me check the code, i am fairly confident qcow2 images will be downloaded by nova and converted before being imported. *me checks* | |
| 21:02:02 | melwitt | how does b) help if each cell is using an independent ceph? | |
| 21:02:43 | melwitt | mnaser: they are, it's just the ceph docs say they don't support qcow2 format | |
| 21:02:58 | mnaser | melwitt: it will create a centralized cache so to speak, so an image is downloaded from cell A to cell B once only | |
| 21:03:12 | mnaser | rather than being downloaded potenitally N times where N is the # of hypervisors | |
| 21:03:18 | mnaser | potentially* | |
| 21:03:24 | melwitt | okay, I'm just missing how it's centralized if each ceph is separate and can't find each other like you said | |
| 21:04:39 | melwitt | if you configure things unlike the ceph docs instruct, then yes it will download and convert. that's probably why it says don't do that, use raw | |
| 21:05:06 | mnaser | melwitt: okay, but if i use raw, it'll still download it across the network and cache it inside every hypervisor | |
| 21:05:27 | mnaser | ideally, i'd imagine it caching it inside the ceph cluster in cell B, so that it is downloaded once only, and nova can do fast clone | |
| 21:05:49 | melwitt | so cache one copy per cell is what you're saying | |
| 21:05:56 | mnaser | correct | |
| 21:06:12 | mnaser | so currently: cell B HVs downloads raw image from glance (which is in another location), caches it locally, imports to create disk | |
| 21:06:34 | mnaser | idea: cell B HV downloads raw image from glance, stores it in ceph, fast-clone to create disk. further hypervisors would fast-clone that same image | |
| 21:06:35 | melwitt | okay. yeah, I guess I didn't know ceph images are cached on disk on hypervisors. I thought they just accessed the ceph pool each time | |
| 21:07:05 | mnaser | melwitt: generally they are not, but because the cell wouldnt be able to reach that cluster, it will fall back to plain o'l downloading image via http and importing it | |
| 21:07:42 | mnaser | and that can be a pretty big bottleneck as things scale | |
| 21:08:31 | melwitt | that seems reasonable to me, mdbooth is the imagebackend guru so he'd be the best person to run the idea by | |
| 21:09:40 | mnaser | melwitt: just drumming up things to hear what it looks like. this whole idea can be worked around by storing things twice in glance (it supports multiple image locations) | |
| 21:10:08 | mnaser | but then it would be a glance thing rather than nova (nova would just try all locations and find that one of them is in a ceph cluster it can access) | |
| 21:10:22 | mnaser | it's been a fun little problem to think about | |
| 21:13:14 | melwitt | yeah, good points. let me dig up our etherpad for the to-resolve things we've been keeping track of | |
| 21:14:49 | mnaser | im sure ill run into plenty of other fun stuff as we'll use segments and routed networks too | |
| 21:16:43 | melwitt | I expect so too | |
| 21:17:31 | melwitt | mriedem: what's the current multi-cells issues/todos/considerations etherpad we have? | |
| 21:17:51 | melwitt | I found https://etherpad.openstack.org/p/nova-pike-cells-v2-todos but wasn't sure if it's the latest thing | |
| 21:18:17 | mriedem | this? https://etherpad.openstack.org/p/cellsv1-to-v2-migration | |
| 21:18:48 | mriedem | i'm fairly sure https://etherpad.openstack.org/p/nova-pike-cells-v2-todos hasn't been touched in a long time | |
| 21:19:20 | mriedem | so yeah https://etherpad.openstack.org/p/nova-pike-cells-v2-todos is probably the most recent since it was the running list of stuff for pike | |