| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-05-17 | |||
| 12:33:03 | mriedem | belmoreira: yeah https://docs.openstack.org/osc-placement/latest/index.html | |
| 12:33:11 | mriedem | openstack resource class create - to create your custom resource class | |
| 12:33:17 | gibi | mriedem: and it seems that the fix https://review.openstack.org/#/c/569048 shows a race condition as well | |
| 12:33:22 | mriedem | openstack resource provider list - to find your compute node resource provider | |
| 12:33:58 | mriedem | belmoreira: and then you want to set that custom resource class inventory on the compute node resource provider? use https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-inventory-set | |
| 12:34:11 | mriedem | or actually https://docs.openstack.org/osc-placement/latest/cli/index.html#resource-provider-inventory-class-set is probably easier | |
| 12:34:27 | mriedem | openstack resource provider inventory set - makes you specify the full inventory with the changes, since it's a full overwrite | |
| 12:34:40 | mriedem | openstack resource provider class set - just modifies the inventory for the specified class | |
| 12:34:58 | mriedem | gibi: yeah i was wondering why the osc-placement functional test was still failng | |
| 12:36:27 | mriedem | gibi: are you sure that https://review.openstack.org/#/c/569048/ ran against the merged nova change? | |
| 12:36:55 | belmoreira | mriedem: thanks. This will be always defined with placement API. There isn't the option for the compute_node advertise the custom resource classes that it supports? | |
| 12:37:42 | gibi | mriedem: the nova patch merged yesterday, I pushed the fix for osc_placement today | |
| 12:37:46 | gibi | mriedem: so I assume | |
| 12:38:29 | mriedem | belmoreira: the compute node doesn't make up custom resource classes on the fly randomly :) | |
| 12:38:39 | mriedem | the compute node will report standard resource classes | |
| 12:38:39 | gibi | mriedem: and the osc_placement test fails differently with the fix than before | |
| 12:38:58 | mriedem | gibi: it looks like assertCommandFailed might just be doing a prefix vs assertIn check? | |
| 12:39:01 | mriedem | because you got HTTP 404 back | |
| 12:39:06 | mriedem | http://logs.openstack.org/48/569048/1/check/osc-placement-dsvm-functional/2800b3e/job-output.txt.gz#_2018-05-17_08_20_04_900926 | |
| 12:39:33 | mriedem | belmoreira: we plan on reporting compute driver capabilities as standard traits on the compute node provider, see https://review.openstack.org/#/c/538498/ | |
| 12:39:57 | mriedem | belmoreira: so we can report things like, "does this compute node support device tags, does this compute note support multiattach volumes, etc" | |
| 12:40:29 | mriedem | those are being defined here https://github.com/openstack/os-traits/tree/master/os_traits/compute | |
| 12:40:35 | gibi | mriedem: ahh. I've just realized that it fails in cleanup not in the test itself http://logs.openstack.org/48/569048/1/check/osc-placement-dsvm-functional/2800b3e/job-output.txt.gz#_2018-05-17_08_20_04_880092 | |
| 12:40:44 | mriedem | i just need to cleanup that patch and get on the update_provider_tree() train | |
| 12:40:58 | mriedem | gibi: aha | |
| 12:41:17 | gibi | mriedem: so the test deletes the trait, then fails to clean up as that want to delete the trait again | |
| 12:41:32 | mriedem | gibi: ah yes https://review.openstack.org/#/c/569048/1/osc_placement/tests/functional/base.py@212 | |
| 12:41:43 | mriedem | if 'not found' not in err_message: | |
| 12:41:46 | mriedem | just need to fix that | |
| 12:41:52 | gibi | mriedem: good point | |
| 12:41:59 | mriedem | if if 'not found' not in err_message or 'HTTP 404' in err_message: | |
| 12:42:00 | gibi | mriedem: I missed that | |
| 12:42:02 | mriedem | to support old and new placement | |
| 12:42:18 | mriedem | er or 'HTTP 404' not in err_message: | |
| 12:42:20 | gibi | mriedem: even better point to support old and new placement | |
| 12:42:31 | mriedem | gibi: want to fix that quick and then i'll rebase avolkov's series | |
| 12:42:39 | mriedem | belmoreira: does that help you out? | |
| 12:42:50 | gibi | mriedem: I'm doing it as we speak | |
| 12:44:16 | belmoreira | mriedem traits would do the trick as well. What I was thinking is an option to specify "custom" traits/resource classes per host. | |
| 12:45:02 | belmoreira | report compute_driver capabilities is good but we may want to use different properties as well | |
| 12:48:44 | gibi | mriedem: both the old and the new error message contains HTTP 404, so we dont need the ORed conditional | |
| 12:48:52 | gibi | stack@ubuntu:~/osc-placement$ openstack trait delete ALMA --os-placement-api-version 1.6 | |
| 12:48:55 | gibi | No such trait(s): ALMA. (HTTP 404) | |
| 12:49:06 | mriedem | belmoreira: you can do that yourself | |
| 12:49:22 | mriedem | belmoreira: you mean you want to set a config option for nova to proxy all of that information for you? | |
| 12:50:04 | gibi | stack@ubuntu:~/nova$ openstack trait delete ALMA --os-placement-api-version 1.6 | |
| 12:50:07 | gibi | Not found (HTTP 404) | |
| 12:50:14 | mriedem | belmoreira: the beauty of custom resource classes and traits and placement is that it's external to nova, so you don't need to proxy stuff through nova and add the complexity / bottleneck there | |
| 12:50:27 | mriedem | gibi: cool | |
| 12:51:00 | openstackgerrit | Balazs Gibizer proposed openstack/osc-placement master: Fix error message asserts in functional test https://review.openstack.org/569048 | |
| 12:51:04 | gibi | mriedem: ^^ here it is | |
| 12:51:20 | mriedem | yup, watching zuul | |
| 12:54:42 | mriedem | tssurya: i guess belmiro didn't like that answer? | |
| 12:55:02 | tssurya | mriedem: I am sure that's not the case :D | |
| 12:55:09 | mriedem | he hates me now doesn't he | |
| 12:55:14 | tssurya | haha | |
| 12:55:19 | mriedem | I BLEW IT | |
| 12:55:48 | mriedem | tssurya: btw, got any cells v2 upgrade news? | |
| 12:56:32 | jroll | alex_xu: thanks | |
| 12:57:03 | tssurya | mriedem: its actually okay.. | |
| 12:57:24 | tssurya | we are just trying to understand what is taking time in the scheduler | |
| 12:57:39 | tssurya | I mean what takes most of the time in the path of the scheduling | |
| 12:58:13 | mriedem | my guess would be that it's pulling more compute node records from the db than it used to do with the per-cell scheduler | |
| 12:58:48 | tssurya | and what's more weird is, how do I test exactly this scatter-gather versus normal sequential querying ? | |
| 12:58:52 | gibi | mriedem: the lates ps of the osc_placement fix still failing the functional test locally :/ | |
| 12:58:56 | tssurya | like the time taken | |
| 12:59:02 | mriedem | so if you have 50 cells with 200 computes, before the top level scheduler would pick a cell and that would pull at most 200 compute nodes, right? but now it could be pulling at most 10K compute nodes | |
| 12:59:44 | tssurya | mriedem: we use dan's pre-placement filter to ensure we don't pull that much | |
| 12:59:54 | mriedem | tssurya: hmmm, we could add some debug logging to the scatter gather routine to log how long an operation took per cell... | |
| 13:00:26 | tssurya | mriedem: that is what I did, but it gives weird times, | |
| 13:00:32 | tssurya | not convincing ones | |
| 13:01:20 | mriedem | is the time being spent in pulling the compute nodes from the cells or from the filters? | |
| 13:02:21 | tssurya | mriedem: filters don't seem to take much time, pulling computes (the query directly done on the DB is very fast, however logging the scatter-gather routine gives weird longer times) | |
| 13:02:34 | tssurya | what is taking time is in the select destinations | |
| 13:02:44 | tssurya | I am working on it now to narrow it down further | |
| 13:02:47 | mriedem | also, do you know if you're hitting the same compute nodes for concurrent requests (packing)? just wondering if you're getting 409s from placement when trying to allocate resources against a target node | |
| 13:02:49 | gibi | mriedem: ohh, there is an upper() call on the message somewhere as {{{err_message:no such trait(s): custom_fake_hw_gpu_class_264c6d7df889415ea5d3b428d7a4a2fc. (http 404)}}} | |
| 13:03:32 | mriedem | tssurya: hmm, might be the alternate selection code? | |
| 13:03:51 | tssurya | mriedem: yea I am right at that point now | |
| 13:04:06 | tssurya | will see in prod if that's the one | |
| 13:04:52 | openstackgerrit | Balazs Gibizer proposed openstack/osc-placement master: Fix error message asserts in functional test https://review.openstack.org/569048 | |
| 13:05:11 | mriedem | that will add more time compared to before queens because we are trying to now find up to max_attempts alternate hosts in the same cell, and that means running those alternates through the filters again | |
| 13:05:38 | tssurya | mriedem: oh right! | |
| 13:05:42 | mriedem | and that's per instance | |
| 13:06:08 | mriedem | so if you do server multicreate in a single request with like 10 instances and max_attempts is 3, we are trying to find i think 30 hosts | |
| 13:07:37 | tssurya | mriedem: hmm yea, will know if this is the one taking time, will check the logs | |
| 13:07:45 | tssurya | meanwhile mriedem I have another issue : http://logs.openstack.org/71/568271/2/check/tempest-full-py3/7142c54/testr_results.html.gz , not sure on how to fix this | |
| 13:08:18 | gibi | mriedem: ps 3 passes the functional locally | |
| 13:08:19 | tssurya | it totally depends on if the host runs both cell and super conductors or runs one of them | |
| 13:09:40 | openstackgerrit | Nguyen Hai proposed openstack/nova-specs master: Follow the new PTI for document build https://review.openstack.org/551802 | |
| 13:11:53 | mriedem | tssurya: hmm, the cell conductor should be reporting into the cell1 database | |
| 13:12:22 | mriedem | this is n-super-cond: http://logs.openstack.org/71/568271/2/check/tempest-full-py3/7142c54/controller/logs/screen-n-super-cond.txt.gz#_May_17_08_20_09_133873 | |
| 13:12:35 | mriedem | http://logs.openstack.org/71/568271/2/check/tempest-full-py3/7142c54/controller/logs/etc/nova/nova_conf.txt.gz | |
| 13:12:43 | mriedem | which is in cell0 | |
| 13:12:44 | mriedem | [database] connection = mysql+pymysql://root:secretdatabase@127.0.0.1/nova_cell0?charset=utf8 | |
| 13:13:03 | mriedem | the cell conductor is using the nova cell1 config | |
| 13:13:03 | mriedem | http://logs.openstack.org/71/568271/2/check/tempest-full-py3/7142c54/controller/logs/screen-n-cond-cell1.txt.gz#_May_17_08_20_10_145541 | |
| 13:13:19 | mriedem | http://logs.openstack.org/71/568271/2/check/tempest-full-py3/7142c54/controller/logs/etc/nova/nova_cell1_conf.txt.gz | |
| 13:13:23 | mriedem | [database] connection = mysql+pymysql://root:secretdatabase@127.0.0.1/nova_cell1?charset=utf8 | |
| 13:14:43 | mriedem | tssurya: it's a buggy test http://git.openstack.org/cgit/openstack/tempest/tree/tempest/api/compute/admin/test_services.py#n61 | |