| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-03-18 | |||
| 13:21:15 | sean-k-mooney | i think its just "pip install osc-placement" | |
| 13:21:53 | ygk_12345 | sean-k-mooney should I do it in the utility node in openstack ansible ? | |
| 13:21:55 | sean-k-mooney | well "python3 -m pip install --user osc-placement" woudl be more modern way of doing it | |
| 13:21:58 | lyarwood | bauzas: random scheduler question for you, if we fail to schedule an instance and it moves to ERROR can we ever retry launching that instance or can we only delete it? | |
| 13:22:20 | lyarwood | bauzas: working through an issue where failure to schedule an instance with attached volumes leaves these volumes in a reserved state | |
| 13:22:31 | sean-k-mooney | ygk_12345: this really isnt a supprot channel and definetly not for openstack ansible so i dont know. | |
| 13:22:46 | lyarwood | bauzas: but if we are able to schedule and/or go straight to a specific host and then fail the volumes become available | |
| 13:23:17 | lyarwood | I assume we want things to stay reserved *if* we can relaunch the instances somehow | |
| 13:23:38 | ygk_12345 | sean-k-mooney oh ok | |
| 13:35:06 | luyao | stephenfin: Hi, I saw you rechecked my vpmem-live-migration patch. FYI, I submitted a new version, there are a bug in the previous one causing the tempest test failed under mixed-version compute services, I think I addressed it for now, you can look at it if you have time. | |
| 13:37:12 | bauzas | lyarwood: I'm back | |
| 13:37:35 | kashyap | lyarwood: I can't see how your patch is the reason for the 'tempest-integrated-compute' failure | |
| 13:37:42 | kashyap | (Not that you're claiming it :)) | |
| 13:39:19 | lyarwood | kashyap: which one sorry? | |
| 13:39:56 | luyao | artom: I'm afraid I can't response to you promptly in IRC , so I'll appreciate that if you can leave comments under that patch. | |
| 13:39:58 | kashyap | lyarwood: The blockCopy one; sorry - https://zuul.opendev.org/t/openstack/build/1796646647854d128ba128b45d5c7a14 | |
| 13:40:12 | artom | luyao, yeah, that's fine | |
| 13:42:14 | bauzas | lyarwood: to answer your question, you could rebuild I think but not just 'create' again | |
| 13:42:42 | artom | dansmith, when you're about, you don't have to look at the entire patch, but just the https://review.opendev.org/#/c/687856/15/nova/conductor/tasks/live_migrate.py@57 bit, about using compute service versions to judge support for a thing | |
| 13:43:54 | lyarwood | bauzas: so we can't just `nova start $instance` | |
| 13:43:59 | lyarwood | bauzas: we have to rebuild? | |
| 13:44:23 | lyarwood | bauzas: either way, I'd assume that given that's an option the volumes should stay reserved right? | |
| 13:49:21 | bauzas | lyarwood: 'start' recreates the guest XML but doesn't ask for cinder and neutron | |
| 13:49:27 | bauzas | (AFAIR) | |
| 13:50:12 | lyarwood | bauzas: it should update the attachment but it wouldn't recreate it | |
| 13:51:15 | bauzas | yup, this | |
| 13:51:55 | bauzas | either way ,if the instance was set to ERROR on scheduling stage, then the related cell is cell° | |
| 13:51:55 | bauzas | either way ,if the instance was set to ERROR on scheduling stage, then the related cell is cell° | |
| 13:51:59 | bauzas | cell0 sorry | |
| 13:52:09 | bauzas | so you're messed up | |
| 13:52:21 | bauzas | that's why a rebuild or resize is the only option | |
| 13:56:41 | sean-k-mooney | the nova start api action calls power on the virt driver which in the libvirt driver calls hard reboot because reasons | |
| 13:57:24 | sean-k-mooney | so yes it recreates the xml. it might rettrive attachemnt info form cinder but it will not modify the sate of the volume in cinder | |
| 13:57:54 | lyarwood | sean-k-mooney / bauzas ; ack thanks | |
| 13:58:08 | lyarwood | so I'm thinking that either way the volume needs to remain reserved | |
| 13:58:19 | lyarwood | if we ERROR out during scheduling (current behaviour) | |
| 13:58:36 | lyarwood | or if we ERROR out in the virt driver during spawn (currently moves back to available) | |
| 13:58:37 | sean-k-mooney | well if you stop an instance then start it we should not hit the scuduler | |
| 13:58:43 | sean-k-mooney | we dont for a hard reboot | |
| 13:58:59 | lyarwood | right | |
| 13:59:16 | sean-k-mooney | a rebuild need to hit the schduler to assert the new image does not contian metadata that invaldiate the current host | |
| 13:59:26 | sean-k-mooney | but start and reboot dont cause nothing changes | |
| 13:59:38 | lyarwood | right | |
| 13:59:45 | sean-k-mooney | lyarwood: whats the issue you are looking at by the way | |
| 14:01:00 | lyarwood | sean-k-mooney: as above, when we ERROR out during scheduling attached volumes remain in a reserved state (as we don't delete their attachments in cinder), if we fail during spawn the compute layer moves the volume back to available (by deleting the attachments in cinder). | |
| 14:01:20 | nightmare_unreal | hello where can I see list of bugs for openstack CLI client or nova related bugs | |
| 14:01:43 | sean-k-mooney | lyarwood: ah ok | |
| 14:01:48 | lyarwood | sean-k-mooney: so I'm thinking that for the latter we should keep the volume reserved to avoid something or someone else coming along and using it | |
| 14:01:58 | lyarwood | if it isn't multiattach etc etc | |
| 14:02:51 | sean-k-mooney | lyarwood: the later being during spawn when we delete the attachemtn we leave it in a reserved sate | |
| 14:03:20 | lyarwood | sean-k-mooney: when we delete the attachments the volume moves back to an available state | |
| 14:03:37 | lyarwood | sean-k-mooney: I'd like to change this to keep it reserved by not deleting the attachments | |
| 14:04:23 | sean-k-mooney | in the case wehre start fails or boot fails or both | |
| 14:04:42 | sean-k-mooney | for boot it will end up in cell 0 and have host so no attachemts | |
| 14:04:57 | sean-k-mooney | *no host | |
| 14:05:17 | lyarwood | for spawn, so we have a host and a cell already | |
| 14:06:29 | lyarwood | brb | |
| 14:24:31 | sean-k-mooney | nightmare_unreal: nova bugs are here https://bugs.launchpad.net/nova | |
| 14:24:45 | sean-k-mooney | nightmare_unreal: the openstack client uses storyborad | |
| 14:25:52 | sean-k-mooney | so here https://storyboard.openstack.org/#!/project/openstack/python-openstackclient but this is all the releated projects https://storyboard.openstack.org/#!/project_group/80 | |
| 14:26:15 | nightmare_unreal | thanks :) | |
| 14:34:45 | dansmith | artom: your sigh about migration func testing is lack of review I assume? | |
| 14:34:55 | dansmith | I'm sure I'm at least partially to blame there | |
| 14:35:51 | artom | dansmith, no, mostly towards myself | |
| 14:35:57 | dansmith | oh okay | |
| 14:35:58 | artom | For not staying on top of it | |
| 14:36:33 | artom | I wanted to do a push, well, starting now | |
| 14:36:51 | artom | I still have stephenfin's feedback from the previous millenium to address (about the 2/3 NUMA nodes split) | |
| 14:37:16 | artom | But if we manage to land it before the VPMEM live migration patch, it would make things easier | |
| 14:37:25 | dansmith | ack | |
| 14:37:48 | dansmith | I'm pretty un-jazzed about that patch | |
| 14:38:24 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Support unshelve with qos ports https://review.opendev.org/704759 | |
| 14:38:24 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Enable unshelve with qos ports https://review.opendev.org/705475 | |
| 14:38:37 | artom | dansmith, live migration with VPMEM? | |
| 14:38:44 | dansmith | yes | |
| 14:39:10 | artom | In terms of current implementation or the concept in general? | |
| 14:39:34 | artom | I'm just neutral-jazzed - I got involved because my pawprints are all over that code after NUMA live migration | |
| 14:39:56 | dansmith | well, I'm rather negative on the vpmem stuff in general, but that ship has sailed..But yeah, the current patch more acutely | |
| 14:40:09 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Merge qos related renos for Ussuri https://review.opendev.org/706766 | |
| 14:40:17 | dansmith | it smacks of "let me just move around all these critical things so they work the way I need them" | |
| 14:40:44 | artom | Yeah... | |
| 14:40:49 | lyarwood | sean-k-mooney: https://bugs.launchpad.net/neutron/+bug/1867936 seen this before? | |
| 14:40:49 | openstack | Launchpad bug 1867936 in neutron "test_update_delete_extra_route failing due to timeout when creating subnets" [Undecided,New] | |
| 14:41:28 | artom | dansmith, you've essentially put words on a vague feeling I had | |
| 14:41:43 | sean-k-mooney | lyarwood: only in the failure of one of your patches | |
| 14:41:48 | artom | Which is why I would really be more comfortable if we at least had NUMA LM func tests in-tree before merging | |
| 14:42:07 | gibi | dansmith: hi! added the global service version check to https://review.opendev.org/#/c/705475 | |
| 14:42:08 | artom | Those at least exercise those code paths | |
| 14:42:18 | dansmith | artom: well, I certainly don't think that's unreasonable, as long as we can get that moving | |
| 14:42:33 | artom | dansmith, ack, lemme get on that | |
| 14:42:56 | lyarwood | sean-k-mooney: yeah I've seen it a few times, was just wondering if someones already raised a more specific bug report for this | |
| 14:43:04 | dansmith | artom: using your patch to validate what we have in tree so we can more directly assess the changes this one makes would be helpful | |
| 14:43:07 | dansmith | gibi: ack | |
| 14:43:39 | artom | dansmith, the bottom two of https://review.opendev.org/#/c/705545/ would be ready, I'd say | |
| 14:43:41 | sean-k-mooney | lyarwood: there have been some issue with the dhcp agent by the way. e.g. it not seting up the dhcp config in a timely manner | |
| 14:43:52 | dansmith | gibi: I feel like we might be at the point with the cyborg set where someone else needs to take over it.. I'm not fully caught up yet this morning but I think feedback has been pending for a while now | |
| 14:43:58 | sean-k-mooney | this might just be a side effect of that althogh i though that neuton had fixed that | |
| 14:44:07 | artom | dansmith, if you have any -1s to get in on those, do it now ;) | |
| 14:44:28 | lyarwood | sean-k-mooney: kk | |
| 14:44:29 | dansmith | artom: at least the bottom one is -W, should it not be? | |