| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2020-11-09 | |||
| 14:36:52 | stephenfin | I'd like to move pretty much all of OSC first, since we should have no issues doing that once openstacksdk is good enough | |
| 14:37:54 | sean-k-mooney | since you have put this effort in to close the gap i just dont want it to reopen | |
| 14:44:14 | openstackgerrit | sean mooney proposed openstack/nova master: libvirt: delegate ovs plug to os-vif https://review.opendev.org/602432 | |
| 14:44:48 | sean-k-mooney | stephenfin: lyarwood can ye take a look at ^ | |
| 14:44:58 | sean-k-mooney | fixed the pep8 issue otherwise its the same | |
| 14:47:31 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: doc: require openstack client change for every new API microversion https://review.opendev.org/717727 | |
| 14:47:54 | gibi | stephenfin, sean-k-mooney: this ^^ is my contribution to the OSC topic | |
| 14:49:21 | sean-k-mooney | :) | |
| 14:51:50 | gibi | "Delay in Elastic Search: Indexing behind by 141 hours" :( | |
| 14:54:50 | sean-k-mooney | thats only slightly longer then normally its normally 72 hours i think | |
| 14:55:26 | sean-k-mooney | if you are fering to logstash/kibana upstream | |
| 14:55:32 | sean-k-mooney | *refering | |
| 15:01:35 | gibi | it remember when it was close to 0 | |
| 15:01:41 | gibi | last week it hanged around 100 hours | |
| 15:02:34 | sean-k-mooney | it does tend to vary but i think the target was no more the 72 hours | |
| 15:02:45 | sean-k-mooney | it does catch up from time to time | |
| 15:03:45 | gibi | I hope so | |
| 15:13:12 | brinzhang_ | stephenfin: thanks fix that issue from openstack server migration list CLI, I am sure I was tested ti, but the strange thing is that this issue was not found | |
| 15:21:05 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Remove compute service level check for qos ops https://review.opendev.org/735570 | |
| 15:21:52 | gibi | stephenfin: I think you will like this code removal patch ^ | |
| 15:32:38 | openstackgerrit | Balazs Gibizer proposed openstack/nova stable/victoria: Warn when starting services with older than N-1 computes https://review.opendev.org/761923 | |
| 15:42:10 | openstackgerrit | Balazs Gibizer proposed openstack/nova stable/victoria: Add upgrade check about old computes https://review.opendev.org/761924 | |
| 15:45:17 | stephenfin | gibi: Done. I assume we're okay to merge things like that now that we've got the service version check? | |
| 15:45:31 | gibi | stephenfin: that is my idea too | |
| 15:45:56 | dansmith | the per-release service version check should catch anything that merged N-2 releases ago, | |
| 15:46:11 | dansmith | so it should be fine to remove older ones and rely on the macro one yeah | |
| 15:46:50 | dansmith | (assuming the qos one is old enough, I didn't look) | |
| 15:47:05 | gibi | qos move ops are added in Ussuri | |
| 15:47:35 | gibi | I mean that last one | |
| 15:48:15 | gibi | so in Victoria we could have removed the service level check. But we did not for extra safety | |
| 15:49:48 | openstackgerrit | Lee Yarwood proposed openstack/nova master: Migrate nova-grenade-multinode job to zuulv3 native https://review.opendev.org/742056 | |
| 15:52:22 | dansmith | stephenfin: we're not really breaking RPC specifically here, because we're not changing any rpc versions or signatures or anything, but this is one of those "not covered by the rpc versions" behaviors.. it's breaking service-to-service interaction, but only for older computes (not even older RPC versions), but for which we've already said isn't supported | |
| 15:53:03 | stephenfin | that makes sense | |
| 15:53:05 | lyarwood | gmann: ^ I'd like to push ahead with this btw, we are currently hitting https://bugs.launchpad.net/nova/+bug/1901739 with the original bionic based job so I'd rather switch to Focal and add the ceph coverage later. | |
| 15:53:05 | openstack | Launchpad bug 1901739 in OpenStack Compute (nova) " libvirt.libvirtError: internal error: missing block job data for disk 'vda'" [High,Fix released] - Assigned to Lee Yarwood (lyarwood) | |
| 15:53:08 | dansmith | stephenfin: I'll put this in a comment once I review, but just echoing here since you called me out :) | |
| 15:53:18 | stephenfin | ta : | |
| 15:53:19 | stephenfin | :) | |
| 15:53:37 | gibi | stephenfin, dansmith: thanks for the review btw | |
| 16:05:08 | openstackgerrit | Merged openstack/nova master: Remove six.moves https://review.opendev.org/727224 | |
| 16:10:06 | gmann | lyarwood: but we are going to loose ceph coverage right? or we can add ceph coverage as separate job using existing script and move them once ceph greande base job is ready | |
| 16:10:35 | lyarwood | gmann: we can try but I'd take a working gate over missing ceph coverage for a few weeks at the moment | |
| 16:12:32 | gmann | lyarwood: existing zuulv2 grenade jobs is working right or it is failing? | |
| 16:12:34 | jgwentworth | lyarwood: so what's the plan for adding ceph coverage back? seems like a risk to leave it uncovered for an extended period of time. do we have any idea how to do it for v3 jobs? | |
| 16:13:18 | lyarwood | gmann: it's failing pretty often at the moment due to a an issue with libvirt/QEMU on bionic | |
| 16:13:56 | gmann | you mean on victoria gate? on master gate, it should run on Focal | |
| 16:13:56 | lyarwood | melwitt: https://review.opendev.org/#/q/owner:self+topic:native-zuulv3-migration+status:open - wire up a native zuulv3 multinode ceph job | |
| 16:14:13 | lyarwood | gmann: master gate, multinode grenade still uses bionic | |
| 16:14:23 | lyarwood | gmann: and that's the problem here | |
| 16:14:32 | lyarwood | gmann: or we can move it to NV | |
| 16:14:41 | gmann | lyarwood: oh we should move it to Focal | |
| 16:15:04 | gmann | ah legacy job | |
| 16:15:09 | melwitt | lyarwood: sorry, I don't understand what I'm looking at here that's related to ceph | |
| 16:15:41 | gmann | lyarwood: i did not move base legacy job on bionic | |
| 16:17:47 | gmann | lyarwood: I think we can move to zuulv3 using script for now like in PS1 - https://review.opendev.org/#/c/742056/1 and once ceph grenade base is ready then remove the use of script ? | |
| 16:18:06 | lyarwood | gmann: we can try | |
| 16:18:24 | gmann | ok let me update it. | |
| 16:22:03 | lyarwood | melwitt: sorry nova-grenade-multinode currently does some ceph stuff manually via the live migration hook - https://github.com/openstack/nova/blob/80b807a4c590980b9e514042778bd8c277e89e40/playbooks/legacy/nova-grenade-multinode/run.yaml#L58 & https://github.com/openstack/nova/blob/80b807a4c590980b9e514042778bd8c277e89e40/gate/live_migration/hooks/run_tests.sh#L55-L65 | |
| 16:22:41 | lyarwood | melwitt: I wanted to break this out into a native zuulv3 job based on a multinode ceph job but that's taking a while to work out on the topic I shared above | |
| 16:23:37 | melwitt | lyarwood: ok, so the manual stuff could be made "native" somehow. I did not know that | |
| 16:24:25 | lyarwood | melwitt: yeah instead of calling specific bash functions from the plugin I just wanted to have a generic job that would deploy multinode ceph that we'd run the LM tests on in Nova | |
| 16:25:00 | lyarwood | melwitt: I made some progress a while ago with the key sharing etc just became stuck at the end with getting the subnode to actually connect to the ceph cluster on the main node | |
| 16:26:23 | openstackgerrit | Stephen Finucane proposed openstack/nova master: functional: Add live migration tests for PCI, SR-IOV servers https://review.opendev.org/746950 | |
| 16:26:23 | openstackgerrit | Stephen Finucane proposed openstack/nova master: functional: Expand SR-IOV live migration tests with NUMA https://review.opendev.org/749360 | |
| 16:27:12 | melwitt | lyarwood: I see, thanks, that helps. I hoped to be able to help in some way since I am concerned about the coverage loss but didn't know where to look or start | |
| 16:32:17 | lyarwood | gmann: so are you just going to change the base job and hope that works? | |
| 16:32:46 | lyarwood | to tempest-multinode-full-py3 | |
| 16:34:14 | openstackgerrit | Merged openstack/nova master: Remove six.iteritems/itervalues/iterkeys https://review.opendev.org/727757 | |
| 16:34:25 | openstackgerrit | Merged openstack/nova master: Remove six.byte2int/int2byte https://review.opendev.org/727777 | |
| 16:35:24 | gmann | lyarwood: not tempest but grenade-multinode and running run_tests.sh in post phase with disable to run smoke tests on new node (which run as part of grenade-multinode playbooks) | |
| 16:37:17 | lyarwood | gmann: right I'm asking which base job you're going to use | |
| 16:37:53 | gmann | lyarwood: for grenade anyways we need to use grenade-multinode | |
| 16:39:21 | lyarwood | gmann: right sorry, that's zuulv3 based and would just use the scripts to avoid us losing coverage | |
| 16:39:48 | gmann | yeah | |
| 16:40:00 | lyarwood | got ya | |
| 17:09:47 | openstackgerrit | Ghanshyam Mann proposed openstack/nova master: Migrate nova-grenade-multinode job to zuulv3 native https://review.opendev.org/742056 | |
| 17:13:12 | openstackgerrit | Lee Yarwood proposed openstack/nova master: Add os-volume_attachments reference docs https://review.opendev.org/760971 | |
| 17:24:18 | openstackgerrit | Merged openstack/nova stable/victoria: [doc]: Fix glance image_metadata link https://review.opendev.org/761423 | |
| 17:27:41 | openstackgerrit | Dat Le proposed openstack/nova master: Fix unplugging VIF when migrate/resize VM https://review.opendev.org/751642 | |
| 17:44:38 | openstackgerrit | sean mooney proposed openstack/nova master: libvirt: delegate ovs plug to os-vif https://review.opendev.org/602432 | |
| 17:45:08 | sean-k-mooney | stephenfin: lyarwood can you re +w https://review.opendev.org/751642 | |
| 17:45:15 | sean-k-mooney | it was rebased | |
| 17:45:37 | stephenfin | done | |
| 17:45:42 | sean-k-mooney | thanks | |
| 18:07:12 | openstackgerrit | Ghanshyam Mann proposed openstack/nova master: DNM: Testing system scope in tempest https://review.opendev.org/740124 | |
| 18:20:16 | openstackgerrit | Elod Illes proposed openstack/nova stable/ussuri: [doc]: Fix glance image_metadata link https://review.opendev.org/761977 | |
| 18:23:06 | mnaser | sean-k-mooney: have you done some investigation by any chance on any potential ways of speeding up nova startup time in environments with compute nodes that have a large port count | |
| 18:23:28 | mnaser | in this case, it takes almost 5-6 minutes for the agent to go up in an env with ~150-160 ports | |
| 18:23:40 | mnaser | using osv | |
| 18:23:52 | mnaser | i havent tried migrating to the native ovsdb driver, maybe that might help, but a whole lot of plugging happens | |
| 18:26:42 | sean-k-mooney | am not really. we have seen that usign native does indeed help form your previous issues but in general we do need to ensure that all network interfacs are plugged for all vms when the compute agent starts | |
| 18:27:40 | sean-k-mooney | it also depend on the release i know we fixed some issue with ip adress checking on newer release | |
| 18:27:51 | mnaser | sean-k-mooney: wonder if it might make sense to make this happen in a threadpool or something, as it seems to be happening one-by-one | |
| 18:28:07 | sean-k-mooney | not really | |
| 18:28:23 | mnaser | because right now in some of those bigger vms, nova-compute reports down as it goes through the restart that takes ~6m | |
| 18:28:27 | sean-k-mooney | we could but we have to wait for them all to complete | |
| 18:29:09 | sean-k-mooney | so if we dispatch them to a tread pool we will get more paralium but we will the fill the pool and have to wait for it to complete | |
| 18:29:33 | sean-k-mooney | plugging form a nova side shoudl not take that long in general | |
| 18:29:45 | sean-k-mooney | with the native driver its much much faster | |
| 18:30:17 | mnaser | ok i guess thats' probably going to be my next step to see how that improves it | |