| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-03-15 | |||
| 18:20:48 | jaypipes | efried, edleafe: where are we on your battling microversion changes? | |
| 18:21:20 | efried | jaypipes: It's tied up on the home stretch while zuul unwinds its panties. | |
| 18:21:27 | efried | See topic | |
| 18:21:34 | jaypipes | k | |
| 18:24:51 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Use ksa adapter for cinder client https://review.openstack.org/508345 | |
| 18:24:54 | mriedem | efried: i've tried to rebase this but there are some known broken things in it ^ | |
| 18:25:44 | efried | mriedem: Yeah, I had to put it aside for other "more urgent" things. | |
| 18:26:01 | efried | I swear I had it working at some point in the cycle. | |
| 18:26:34 | efried | but by the end, it was definitely busted and I couldn't figure out how to fix it without spending a big chunk of time. | |
| 18:28:31 | openstackgerrit | Claudiu Belu proposed openstack/nova master: compute: Adds instance live-resize https://review.openstack.org/248581 | |
| 18:28:31 | openstackgerrit | Claudiu Belu proposed openstack/nova master: db: Adds live-resize to Migration model migration_type https://review.openstack.org/185961 | |
| 18:34:20 | openstackgerrit | Surya Seetharaman proposed openstack/nova master: Add disabled field to CellMapping object https://review.openstack.org/550090 | |
| 18:34:20 | openstackgerrit | Surya Seetharaman proposed openstack/nova master: [WIP] Add CellMappingList.get_all_enabled() query method https://review.openstack.org/550188 | |
| 18:40:24 | cdent | nice message efried | |
| 18:40:40 | openstackgerrit | Surya Seetharaman proposed openstack/nova master: Add CellMappingList.get_all_enabled() query method https://review.openstack.org/550188 | |
| 18:40:53 | efried | Thanks cdent. I guess if anyone was gonna bother to read the whole thing, it'd be you :) | |
| 18:41:21 | cdent | yeah, spose so | |
| 18:41:34 | cdent | I think you'll find there's a vast army of silent readers out there | |
| 18:41:48 | cdent | their silence can sometimes be rather disturbing | |
| 18:46:06 | efried | cdent: I'm quite a slow reader, and I often feel that pain when something's important enough that I know I gotta read it, but really long (like dhellman's missive on requirements). I guess for people who read at normal speeds, it's not such an onerous task. | |
| 18:47:25 | cdent | efried: yeah, I've been reminded many times that my attitude towards reading comes from something of a position of privilege. Apparently I read _very_ fast when it comes to email and similar forms. | |
| 18:47:42 | efried | That explains a lot. | |
| 18:48:49 | efried | I envy those who can read fast and not miss stuff. That's why I read slow - because I'm being real thorough (terrified of missing some detail or - gods forbid - failing to catch a typo!) | |
| 18:49:23 | edleafe | heh, just started reading efried's email | |
| 18:50:12 | edleafe | An hourglass would be good enough :) | |
| 18:52:35 | cdent | efried: I'm certain that I miss stuff, but I'm usually grazing for meaning, not details | |
| 18:52:58 | efried | I should develop that skill. FOMO. | |
| 18:53:18 | cdent | maybe not, probably useful to have both styles around | |
| 18:53:23 | mriedem | mgoddard_: stephenfin: dansmith: so maybe the numa topology filter is ok with ironic http://logs.openstack.org/12/553412/1/check/ironic-tempest-dsvm-ipa-wholedisk-bios-agent_ipmitool-tinyipa/2b781af/logs/screen-n-sch.txt.gz#_Mar_15_15_24_11_093172 | |
| 18:53:27 | mriedem | that devstack change didn't blow up | |
| 18:53:33 | dansmith | sweet | |
| 18:53:41 | cdent | for me it's in part a learned skill to intentionally miss out on some stuff | |
| 18:55:03 | mriedem | could have sworn something about the IronicNodeState object was different such that the filter would fail with it | |
| 18:55:12 | mriedem | also, it's not like we actually have pci requests in these CI jobs | |
| 18:55:30 | edleafe | I tried to learn to skim back in college. I always felt like I never got anything out of it, so I went back to reading details | |
| 18:55:58 | edleafe | efried: excellent recap | |
| 18:55:59 | dansmith | mriedem: that's true, I guess a flavor or request with pci or numa could break if there are ironic hosts in there somehow, | |
| 18:56:09 | dansmith | but I dunno what that would be really | |
| 18:56:29 | efried | edleafe: Thanks. | |
| 18:56:36 | mriedem | dansmith: the other thing might have been something to do with allocation ratios, but grasping at straws | |
| 18:57:29 | dansmith | we don't have ratios for those types though | |
| 18:58:04 | mriedem | goes into the numa topology limits | |
| 18:58:08 | mriedem | the cpu and ram allocation ratois | |
| 18:58:55 | cfriesen | mriedem: we have a private patch to enable ironic and regular nodes...had to make some of the scheduler filters check the hypervisor type. | |
| 18:59:12 | mriedem | cfriesen: is there anything you guys don't have a private patch for? | |
| 18:59:33 | cfriesen | mriedem: we try to upstream stuff, but it takes forever | |
| 18:59:40 | mgoddard_ | mriedem: that's good news! | |
| 19:00:01 | cfriesen | mriedem: plus, we only need to worry about one hypervisor | |
| 19:00:16 | mriedem | mgoddard_: well, it's kind of a false positive i think | |
| 19:00:34 | mriedem | cfriesen: it takes even longer when you don't even propose them | |
| 19:00:45 | dansmith | mriedem: it's not a false positive, it's just not comprehensive.. it means something, it just doesn't mean it all works fine :) | |
| 19:01:39 | mriedem | i think host_topology_and_format_from_host might be the thing | |
| 19:02:20 | mriedem | that's just always None for ironic i think | |
| 19:02:22 | cfriesen | mriedem: yeah, I know. I don't control how much upstreaming time I get. Looks like we modified NUMATopologyFilter to check the hypervisor type, and modified AggregateInstanceExtraSpecsFilter to ignore the "baremetal" and "storage" keys. | |
| 19:02:45 | dansmith | mriedem: if that breaks if it's none, that probably also precludes using any two hypervisors together where one is libvirt with numa and the other doesn't support it, right? | |
| 19:03:14 | mriedem | nvm, that would just filter out that ironic node | |
| 19:03:19 | mriedem | / host | |
| 19:03:38 | openstackgerrit | Eric Berglund proposed openstack/nova master: WIP: Resize https://review.openstack.org/553583 | |
| 19:03:59 | mriedem | cfriesen: if you could figure out *why* you "Looks like we modified NUMATopologyFilter to check the hypervisor type" that would be nice | |
| 19:04:13 | mriedem | because at some point i thought these wouldn't work but i can't figure out why now | |
| 19:04:24 | dansmith | yeah at least upstreaming the bug would be worthwhile | |
| 19:04:33 | mriedem | ++ | |
| 19:04:42 | dansmith | although they do that, so if you did in this case, then .. cool :) | |
| 19:04:56 | cfriesen | mriedem: let me check with the author | |
| 19:10:48 | sean-k-mooney | mriedem: cfriesen perhaps because you wanted to avoid qemu hosts when there are numa requests? | |
| 19:11:59 | mgoddard_ | perhaps the bug is just that the NUMA filter doesn't work for ironic, rather than that it rejects all bare metal hosts? | |
| 19:12:11 | cfriesen | sean-k-mooney: don't think so, this was specifically part of allowing one nova-scheduler to handle both libvirt/kvm and ironic compute nodes | |
| 19:12:29 | sean-k-mooney | cfriesen: ah ok | |
| 19:13:41 | sean-k-mooney | mgoddard_: well if i ask for a 2 numa node instance the ironic should be able to select a node with 2 numa nodes however i dont think ironic adds numa info into the compute nodes table for the filter to use | |
| 19:13:53 | mgoddard_ | exactly | |
| 19:14:08 | mgoddard_ | same with cpu pinning | |
| 19:14:19 | sean-k-mooney | mgoddard_: no cpu pinning is different | |
| 19:14:20 | mgoddard_ | and hyperthreading | |
| 19:14:35 | cfriesen | mriedem: back to the nova-compute service delete issue, when deleting a service via the API, nova.db.sqlalchemy.api.service_destroy() will soft-delete both the service and the service and the compute_node entry, but placement is still around. Then we create the compute node again and get new service and compute_node entries with a different uuid but the same hostname. | |
| 19:14:36 | sean-k-mooney | cpu pinning does not make sense in a phyical server context | |
| 19:14:39 | dansmith | it wouldn't matter these days anyway, ironic nodes report a CUSTOM_IRONIC_FOO resource and no cpu/mem | |
| 19:15:36 | mgoddard_ | sean-k-mooney: ok, you're right about pinning. hyperthreading could (but doesn't) work though | |
| 19:15:53 | dansmith | cfriesen: mriedem because we store service_id in the compute node, so we won't find the existing one when we re-create the service | |
| 19:16:15 | dansmith | cfriesen: mriedem I bet that api was never updated when we added the node concept.. it probably needs to delete the node(s) as well when it does that | |
| 19:16:19 | sean-k-mooney | mgoddard_: hyperthreading is a bios config option and should work but nova does not allow you to enable hyperthreading as a flavor extra spec | |
| 19:16:21 | dansmith | (and thus placement( | |
| 19:16:49 | sean-k-mooney | mgoddard_: the closest you have to that is setting the tread count in the cpu topology extra specs | |
| 19:17:24 | mriedem | i didn't realize that service_destroy also deleted the related compute node record | |
| 19:17:40 | cfriesen | me either...had to go read the code. | |
| 19:18:55 | dansmith | cfriesen: I thought you said it didn't? | |
| 19:18:56 | mriedem | ok so if service delete also deletes the compute node record, then yeah we should also remove the compute node RP in placement | |
| 19:19:09 | mriedem | i said deleting the service didn't delete the compute node record | |
| 19:19:19 | dansmith | oh okay | |
| 19:19:31 | mriedem | there was some bug about a year ago we were both talking about this | |
| 19:19:41 | mriedem | but https://github.com/openstack/nova/blob/master/nova/db/sqlalchemy/api.py#L404 | |
| 19:21:31 | cfriesen | I should be able to open a bug and post a WIP fix later today | |
| 19:21:31 | mriedem | https://github.com/openstack/nova/commit/f0d44c5b09f3f3c84038d40b621bb629a1f8110e#diff-3104166b3e802b86db6c5fa92ad08f43 | |
| 19:21:36 | mriedem | ^ is why i thought this | |
| 19:23:28 | mriedem | see the exchange between myself and alex | |
| 19:24:23 | mriedem | so in this case, they deleted the service (and compute node record) but didn't stop the nova-compute service, | |
| 19:24:32 | mriedem | so it re-created the compute node record | |
| 19:25:27 | mriedem | so in that bug, when they listed compute nodes, the api tries to find the related service which was deleted and the api blows up | |
| 19:25:34 | mriedem | because you don't recreate the service until you restart the service | |
| 19:26:16 | melwitt | gibi: I'm working on the nova/neutron session summary, would be helpful if you could fill in any gaps for the bandwidth-based scheduling agreements/decisions when you get a chance https://etherpad.openstack.org/p/nova-ptg-rocky-neutron-summary | |
| 19:26:25 | mriedem | melwitt: he's on vacation | |