| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2017-10-06 | |||
| 03:37:23 | mriedem | reproduced that 409 during claim resources in the scheduler, creating 1000 instances at once | |
| 03:40:09 | mriedem | retried 18 times across that 1000 | |
| 03:40:15 | mriedem | one of those poor saps just couldn't hack it | |
| 03:40:26 | mriedem | need to run with https://review.openstack.org/#/c/507705/2/nova/scheduler/client/report.py to find out which one i guess | |
| 03:41:43 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/pike: Log consumer uuid when retrying claims in the scheduler https://review.openstack.org/509961 | |
| 04:43:59 | openstackgerrit | melanie witt proposed openstack/nova master: Improve the CellDatabases test fixture and usage https://review.openstack.org/508432 | |
| 04:44:00 | openstackgerrit | melanie witt proposed openstack/nova master: Target context for build notification in conductor https://review.openstack.org/509967 | |
| 04:44:00 | openstackgerrit | melanie witt proposed openstack/nova master: Elevate existing RequestContext to get bandwidth usage https://review.openstack.org/509968 | |
| 05:11:37 | hanish | one of my compute node is disabled due to 10 vm launch failing, how can i recover that node | |
| 05:20:37 | Tengu | you must re-enable nova agent, hanish | |
| 05:25:43 | hanish | @Tengu: i restarted nova-compute agent on compute node, but still i facing the issue | |
| 05:29:10 | Tengu | hmmm nope, not via systemctl | |
| 05:29:26 | Tengu | there's an openstack command for that, in order to re-enable it at openstack level | |
| 05:30:33 | takashin | ||
| 05:32:09 | Tengu | hanish: there's something like nova service-list - you should see it's disabled for your host. | |
| 05:32:17 | Tengu | then you have nova service-enable | |
| 05:32:30 | Tengu | hanish: I don't remember the "openstack unified" command for those. | |
| 05:38:30 | hanish | tengu: thanks | |
| 06:17:34 | Tengu | hanish: did it do the trick? | |
| 06:20:20 | hanish | Tengu: thanks, yes it worked. | |
| 06:23:26 | Tengu | hanish: good :). | |
| 06:23:51 | Tengu | I had that kind of issue earlier with tripleO. | |
| 06:36:42 | openstackgerrit | Takashi NATSUME proposed openstack/nova-specs master: Abort Cold Migration https://review.openstack.org/334732 | |
| 06:50:21 | takashin | gmann: cdent: I have modified the spec for "Abort cold migration" function. Would you review https://review.openstack.org/#/c/334732/ again? | |
| 06:50:27 | openstackgerrit | zhangyangyang proposed openstack/nova master: Move libvirts qemu-img support to privsep https://review.openstack.org/507848 | |
| 08:24:00 | zioproto | hello :) Working on Newton I have a region in my cloud where openstack usage list goes in stacktrace. Instance not found | |
| 08:24:23 | zioproto | looking at the nova bugs I did not find anything usefull | |
| 08:25:34 | zioproto | tracking my logs it looks like this is broken since I upgraded to Newton | |
| 08:25:57 | zioproto | dansmith: usually you like these database stories :) | |
| 08:26:53 | bauwser | zioproto: stacktrace ? | |
| 08:27:20 | bauwser | zioproto: by Newton, we began to use the API DB | |
| 08:27:45 | zioproto | https://pastebin.com/gtJbutvi | |
| 08:28:01 | zioproto | it is funny because this happens only in 1 production region | |
| 08:28:23 | zioproto | I have the same setup on dev/staging/prod and there it works | |
| 08:28:38 | zioproto | my feeling is that there is a broken database entry, or something specific to that region that breaks it | |
| 08:29:28 | zioproto | of course openstack server show 72afef44-1a4b-46e5-8cbc-7bd4f0eb31ff gives me a instance not found as well | |
| 08:29:39 | zioproto | what other tables should I dig to look for this uuid ? | |
| 08:32:06 | gmann | takashin: thanks. i will check soon. | |
| 08:33:13 | takashin | gmann: Thanks in advance. | |
| 08:39:47 | bauwser | zioproto: are you aware of those commands when you deploy a Newton cloud ? https://docs.openstack.org/nova/latest/cli/nova-manage.html#man-page-cells-v2 | |
| 08:40:25 | zioproto | bauwser: I think I did this when I upgraded to Mitaka | |
| 08:40:38 | zioproto | to split the nova db into two DBs | |
| 08:41:35 | bauwser | zioproto: if I were you, I'd be looking at the nova_api DB for the instance_mappings table | |
| 08:41:51 | bauwser | zioproto: and check which cell is for the instance UUID | |
| 08:41:51 | zioproto | I go have a look | |
| 08:42:23 | bauwser | zioproto: then looking at host_mappings if we have a cell record for each host in it | |
| 08:52:43 | zioproto | bauwser: my host_mappings table is empty, is that a bad sign ? | |
| 08:53:18 | bauwser | zioproto: indeed, how many hosts do you have? | |
| 08:53:31 | zioproto | you mean compute hosts ? | |
| 08:53:44 | zioproto | like hundreds | |
| 08:56:15 | zioproto | looks like we upgraded to newton without doing this cells housekeeping | |
| 08:56:32 | zioproto | this was mandatory at this point ? | |
| 08:56:38 | zioproto | to have a cell0 database ? | |
| 09:01:57 | bauwser | zioproto: I don't exactly remember when we used the api DB for getting the instance list | |
| 09:02:41 | bauwser | zioproto: but you should definitely map the hosts | |
| 09:06:18 | openstackgerrit | zhangyangyang proposed openstack/nova master: Move libvirts qemu-img support to privsep https://review.openstack.org/507848 | |
| 09:07:53 | openstackgerrit | zhangyangyang proposed openstack/nova master: Move libvirts qemu-img support to privsep https://review.openstack.org/507848 | |
| 09:15:11 | openstackgerrit | zhangyangyang proposed openstack/nova master: Move libvirts qemu-img support to privsep https://review.openstack.org/507848 | |
| 10:35:28 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Transform instance.exists notification https://review.openstack.org/403660 | |
| 10:35:29 | openstackgerrit | Balazs Gibizer proposed openstack/nova master: Add sample test for instance audit https://review.openstack.org/480955 | |
| 11:03:52 | openstackgerrit | Merged openstack/nova master: Blacklist test_extend_attached_volume from cells v1 job https://review.openstack.org/509907 | |
| 11:17:57 | gibi | hm https://review.openstack.org/509907 has been merged so I guess it is recheck time | |
| 11:25:22 | zioproto | bauwser: in nova-manage cell_v2 simple_cell_setup [--transport-url <transport_url>] how does <transport_url> look like ? | |
| 11:28:58 | zioproto | ok I found it here https://docs.openstack.org/nova/latest/user/cells.html | |
| 11:43:33 | fried_rice | jaypipes yt? Wanted to brainstorm a couple of edge cases. | |
| 12:06:07 | gibi | could a second core look at this code removal patch? https://review.openstack.org/#/c/505164/ | |
| 12:18:33 | gibi | bauwser: hi! you can still support my embarrassment in https://review.openstack.org/#/c/509750 if you would like to :) | |
| 12:19:14 | fried_rice | Gluten free, if you please. | |
| 12:20:11 | fried_rice | cdent Perhaps you'd be willing to give some feedback on these edge cases that came to me in my sleep (or lack thereof) last night. | |
| 12:21:10 | cdent | I can try, but these cases keep confusing me, but since it will probably help to talk about it, shoot. | |
| 12:21:26 | fried_rice | If I ask for CUSTOM_FOO:4, is it *ever* legal for placement to give me 3 CUSTOM_FOOs from one RP and 1 from a different RP? | |
| 12:22:32 | cdent | In the fundament case of resource providers and inventory, no | |
| 12:22:40 | cdent | the request is a for a chunk of size 4 | |
| 12:22:57 | fried_rice | Here's the thing: it makes total sense for the answer to be "yes" for something like VFs on separate PFs, all other traits being equal. Because e.g. what if I ask for 4 VFs and I've only got 2 on each PF? | |
| 12:23:13 | cdent | right, I was just going to say “vfs make that weird" | |
| 12:23:36 | fried_rice | But even there, if I'm asking for e.g. bandwidth inventory along with my VF count, how do I split that up? | |
| 12:23:47 | cdent | and also why I said “fundament[al]” because I think nested makes some of these decisions less clear | |
| 12:24:05 | cdent | (btw, it is great that you are exploring this stuff) | |
| 12:24:06 | fried_rice | And it makes *no* sense for something like DISK_GB. If I ask for 4, I don't want you giving me 1GB from each of 4 providers. | |
| 12:24:28 | fried_rice | So let's put that to bed and say Nay. | |
| 12:24:45 | cdent | this issue may be why in some conversations VFs have been proposed as resource providers | |
| 12:24:55 | fried_rice | eek | |
| 12:25:02 | cdent | ikr | |
| 12:25:23 | fried_rice | I mean, I guess you could do that, in the "pre-create" case like current VF passthrough does. | |
| 12:25:47 | fried_rice | No good in the "create VF dynamically" case of the future. | |
| 12:25:58 | fried_rice | cdent Okay, so next thing: | |
| 12:26:18 | fried_rice | Back to the VF scenario, if I ask for VF:2,BANDWIDTH:20000 | |
| 12:26:33 | fried_rice | I'm answering my own question. | |
| 12:26:35 | zioproto | bauwser: still around here ? | |
| 12:26:42 | fried_rice | Those are total chunks on the RP. | |
| 12:26:48 | zioproto | it comes out that the uuid from the stacktrace is pretty unique | |
| 12:26:53 | fried_rice | Placement doesn't care how they're going to be split up | |
| 12:27:02 | zioproto | it is the only instance in the cloud that matches this query | |
| 12:27:06 | zioproto | select * from instances where vm_state="shelved_offloaded" | |
| 12:27:13 | fried_rice | That's up to the virt driver once it gets that information (which we still don't have a way of doing yet - discussion Monday) | |
| 12:27:45 | cdent | yup | |
| 12:27:56 | fried_rice | So I guess the op would have to assume each VF will get 10000. And if they want a different split, they can specify them as separate request numbers (per the spec I'm composing). | |
| 12:28:02 | fried_rice | Cool cool. | |
| 12:28:10 | fried_rice | cdent Thanks for sounding-boarding. | |
| 12:28:22 | cdent | you’re welcome | |