| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-06 | |||
| 13:41:12 | andreykurilin | superdan: marker doesn't work in some cases | |
| 13:41:49 | superdan | andreykurilin: heh, any more detail than that? :) | |
| 13:42:01 | mriedem | andreykurilin: have a failed job log? | |
| 13:42:03 | mriedem | novaclient or rally? | |
| 13:42:03 | andreykurilin | superdan: sure, just need to collect links:) | |
| 13:42:14 | andreykurilin | give me a sec | |
| 13:44:38 | andreykurilin | rally gates are failing due to an issue with pagination. we use limit=-1 option from novaclient. It is designed to make an inf loop changing the marker until the response will include an empty list. For some reasons API ignores the market in some cases. I copy-pasted the code from novaclient and added some debug messages | |
| 13:44:53 | andreykurilin | here is a ok execution - http://logs.openstack.org/83/509783/3/check/gate-rally-dsvm-neutron-existing-users-rally/1a480ab/console.html#_2017-10-05_23_02_16_892620 | |
| 13:45:12 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add error notification for instance.interface_attach https://review.openstack.org/506643 | |
| 13:45:27 | andreykurilin | and just after several seconds, there is one more execution (another iteration of the workload) and it stucks | |
| 13:45:37 | andreykurilin | http://logs.openstack.org/83/509783/3/check/gate-rally-dsvm-neutron-existing-users-rally/1a480ab/console.html#_2017-10-05_23_02_19_441231 | |
| 13:45:52 | andreykurilin | here is a code which I'm using for dedbugging purpose - https://review.openstack.org/#/c/509783/3/rally/plugins/openstack/scenarios/nova/utils.py | |
| 13:46:13 | superdan | andreykurilin: is that implying that you get a page with a marker and you get back a page with the marker in it? | |
| 13:46:16 | andreykurilin | it is equal to what we have in novaclient but with some debug message as I already mentione | |
| 13:46:38 | andreykurilin | superdan: yes | |
| 13:46:42 | andreykurilin | i think so | |
| 13:47:01 | superdan | hrm | |
| 13:47:06 | andreykurilin | but it happens not regulary. In some cases the page includes the marker, in others - no | |
| 13:47:36 | superdan | what sort key are you using? | |
| 13:47:51 | andreykurilin | superdan: due to migration to Zull v3 I cannot say when it had happend actually, but can assume that 2 days ago | |
| 13:47:58 | andreykurilin | no sort keys | |
| 13:48:14 | superdan | andreykurilin: yeah I know what changed, so no question there | |
| 13:48:16 | andreykurilin | suyuperdan: here is a query http://logs.openstack.org/83/509783/3/check/gate-rally-dsvm-neutron-existing-users-rally/1a480ab/console.html#_2017-10-05_23_02_19_452283 | |
| 13:48:18 | superdan | andreykurilin: okay so default sort | |
| 13:48:29 | andreykurilin | just marker in query, nothing more | |
| 13:48:56 | superdan | andreykurilin: are these instances from a num_instances=N type create operation? | |
| 13:49:07 | superdan | andreykurilin: such that they probably have very similar create times? | |
| 13:49:11 | andreykurilin | no | |
| 13:49:58 | superdan | andreykurilin: no meaning they were created one at a time in a client loop? | |
| 13:50:00 | andreykurilin | yes | |
| 13:50:08 | andreykurilin | sec | |
| 13:50:10 | superdan | and how many(ish)? | |
| 13:50:47 | andreykurilin | superdan: there are 2 instances which acre created in one time (~1 sec), but from different threads and with different names | |
| 13:51:08 | superdan | andreykurilin: so you're literally paging through two instances? | |
| 13:52:59 | andreykurilin | yes. just need to mention, that there are 2 cases and both failed. first one boot_and_list actions are performed twice in the same time. the second: list action performed once after both vms are booted | |
| 13:53:50 | superdan | andreykurilin: okay so limit=1 then? | |
| 13:54:13 | superdan | andreykurilin: and both instances are ACTIVE right? | |
| 13:57:06 | bauwser | zioproto: sorry, I have a huge internal backlog to do | |
| 13:59:49 | andreykurilin | superdan: so there are 2 cases. The shared logging relates to the first case, but the behaviour of nova the same and for the second case. Let me dedscribe it more details. there are 2 threads which perfroms boot_and_list actions. the listing is performed right after the vm become active. Both threads are using the same user and tenant | |
| 14:00:00 | zioproto | bauwser: no worries ! | |
| 14:00:35 | andreykurilin | superdan: in this case the first thread performs list action successfully (with using limit=-1 option of novaclient) and the second thread fails | |
| 14:00:59 | superdan | andreykurilin: and what does limit=-1 mean to novaclient? | |
| 14:01:15 | mriedem | page until there is nothing returned i think | |
| 14:01:20 | andreykurilin | yes | |
| 14:01:32 | superdan | right but with what limit to the api? | |
| 14:01:37 | superdan | no limit= default? | |
| 14:01:39 | andreykurilin | no limit | |
| 14:01:50 | mriedem | so default limit of 1000 | |
| 14:01:54 | andreykurilin | yes | |
| 14:02:13 | superdan | okay, so this really should get both instances in the first page, | |
| 14:02:17 | superdan | try another with result[-1] and get an empty page, yes? | |
| 14:03:05 | andreykurilin | `marker = result[-1]` gives the same page as previously with marker in it | |
| 14:03:25 | superdan | right, I was describing what _should_ be happening | |
| 14:03:48 | andreykurilin | yes | |
| 14:03:52 | superdan | okay | |
| 14:04:08 | superdan | I might have an idea of what is going on, but I need to do some experimentation | |
| 14:04:20 | superdan | andreykurilin: in the meantime, can you alter that loop a bit just to see if it helps? | |
| 14:04:22 | gibi | cburgess: hi! Is there any next step about https://blueprints.launchpad.net/nova/+spec/libvirt-virtio-set-queue-sizes I can look at / help with? | |
| 14:04:35 | andreykurilin | superdan: sure | |
| 14:04:44 | superdan | andreykurilin: can you set the sort_keys=['uuid'] | |
| 14:05:23 | superdan | although that really shouldn't matter since we're only iterating instances in a single cell db here | |
| 14:05:40 | superdan | andreykurilin: and I can throw up a nova patch you can depends-on right? | |
| 14:06:09 | superdan | andreykurilin: got a bug number for this yet? | |
| 14:06:42 | andreykurilin | superdan: yes, we can do depends-on to check the fix. no, I do not have a bug report | |
| 14:08:13 | andreykurilin | superdan: superdan: made a patch to check sort_keys, but based on the queue of the zuul the results will be in several hours | |
| 14:11:01 | andreykurilin | superdan: should I create a bug report? | |
| 14:11:14 | superdan | andreykurilin: yeah please create a bug and I'll some debugging | |
| 14:11:19 | superdan | sorry, I'm stuck on a call atm | |
| 14:11:24 | andreykurilin | thanks | |
| 14:21:57 | andreykurilin | superdan: https://bugs.launchpad.net/nova/+bug/1721791 | |
| 14:21:58 | openstack | Launchpad bug 1721791 in OpenStack Compute (nova) "Pagination of instances works incorrect" [Undecided,New] | |
| 14:22:34 | superdan | andreykurilin: thanks, I'm not sure how this is happening, but I'm really distracted on this call | |
| 14:22:39 | superdan | andreykurilin: are you going to be around for a while? | |
| 14:23:03 | andreykurilin | np, I'll planning to be there :) | |
| 14:30:20 | openstackgerrit | Merged openstack/nova master: Remove dest node allocations during live migration rollback https://review.openstack.org/507687 | |
| 14:33:38 | superdan | oooh, I might have a recreate | |
| 14:34:09 | cdent | must be because you’re super | |
| 14:35:15 | superdan | andreykurilin: do you do a regular unpaged list after the fail at all? | |
| 14:35:30 | superdan | andreykurilin: I kinda feel like one of the instances has to be in ERROR state, in cell0 to make this happen | |
| 14:38:57 | andreykurilin | superdan: while doing unpaged list, both instances are returned. After making a boot request, we are fetching the status of VM and do not continue until it become ACTIVE. Both VMs returned ACTIVE status | |
| 14:39:18 | andreykurilin | so I'm pretty sure that they are not in ERROR while listing | |
| 14:40:08 | superdan | andreykurilin: okay they should be sorted by created_at,id which is stable if you only have one database. Unless you have multiple cells here, or instances in cell0, I'm not sure how you could end up with unstable sort | |
| 14:41:59 | andreykurilin | superdan: it is dsvm job with a single node. it doesn't have any special configs | |
| 14:42:12 | superdan | yeah | |
| 14:42:26 | openstackgerrit | Dan Smith proposed openstack/nova master: WIP Always put 'uuid' into sort_keys for stabile instance lists https://review.openstack.org/510140 | |
| 14:42:35 | superdan | andreykurilin: can you try with this in place ^ ? | |
| 14:43:02 | superdan | if you revert the functional part of that change, the test added fails in the same way | |
| 14:43:53 | andreykurilin | ok, will make a depends on patch | |
| 14:51:34 | andreykurilin | superdan: btw, performance of list action is quite good. before I added a limit to the loop, debug messages flooded the log file by 10gb of text (until jenkins kicked the job by temout) :D | |
| 14:51:53 | superdan | andreykurilin: hah, cool | |
| 15:04:27 | superdan | andreykurilin: hmm, actually, that test isn't fully stable, so I need to keep working on it | |
| 15:05:08 | mriedem | fried_rice: merry friday https://review.openstack.org/#/c/488137/22 | |
| 15:05:15 | mriedem | i didn't -1, but i'm sort of inlined to | |
| 15:05:24 | fried_rice | mriedem ack, looking. | |
| 15:13:51 | cdent | ronlund: nobody wants to love on your doc fix? https://review.openstack.org/#/c/502168/ | |
| 15:17:28 | cdent | fried_rice: you still have https://review.openstack.org/#/c/499826/ in your mind? what do we need to do to resolve that? | |
| 15:17:43 | fried_rice | ... | |
| 15:18:11 | cdent | i hear that | |
| 15:18:21 | fried_rice | cdent I actually keep forgetting to put it on the "stuck reviews" list for the nova meetings. | |
| 15:18:45 | fried_rice | It really seems like overkill to put up a whole microversion for that change. | |