Earlier  
Posted Nick Remark
#openstack-nova - 2017-10-18
12:14:17 janki efried, ya. didnot find anything useful in cpu logs
12:14:37 efried rgerganov Is there some kind of API endpoint that can be queried to get the details on the NFS store?
12:14:41 janki efried, ya. policies is 1 of the reason for 404 as per the api doc
12:14:46 efried rgerganov That would be another way to do it.
12:14:48 sean-k-mooney if we did we coudl store info such as the ip adress as a trait on the resouce provider
12:14:51 efried janki Oh, then maybe that's the whole problem.
12:14:58 efried sean-k-mooney Just so.
12:15:30 efried sean-k-mooney Though I'm not sure that's a great idea. More likely you would want to maintain a mapping of RP UUIDs to configs.
12:15:52 sean-k-mooney efried: ya the down side to that is it is a public api...
12:16:07 rgerganov efried, yeah, this is what I am calling "connect info" for resource provider
12:16:10 sean-k-mooney at least i think it is placement is now admin only correct
12:16:17 rgerganov maybe the right term is just "config", idk
12:16:19 efried A trait like CUSTOM_NFS_SHARE_IP_192_168_0_55 is a pretty hideous thought.
12:17:01 sean-k-mooney efried: ya at least that was not an ipv6 address ...
12:17:12 efried Hah, totally
12:17:59 efried sean-k-mooney rgerganov As with PCI devices, we would rather maintain the specifics (like PCI address) external to placement, and have the driver responsible for mapping RP UUIDs to whatever entry in that external data store.
12:18:28 janki efried, POST /compute/v2.1/os-floating-ips => generated 73 bytes in 180 msecs (HTTP/1.1 404) 9 headers in 378 bytes (1 switches on core 0) - "generated" in this line means HTTP package is generated and not FIP is generated right
12:19:07 efried janki I don't know what FIP is. I think that's just saying the response payload was 73 bytes.
12:19:11 rgerganov efried, the problem with this is that it may go out of sync; why not storing this info in placement itself?
12:19:38 efried rgerganov How does it go out of sync? You would have to keep it up to date in placement somehow too, nah?
12:19:45 janki efried, FIP = Floating IP. sorry for short form
12:20:09 efried janki Oh, yeah, it's talking about the response payload. The API message knows nothing about the internals of what you're doing with the call.
12:22:16 janki efried, ack. "Policy check for os_compute_api:os-extended-server-attributes failed with credentials" doesnot mean that os-floating-ip has policy issues and couldnot find anthing in log which suggests that
12:22:36 rgerganov efried, ok maybe keeping them in sync is not a real problem; for me it just feels natural the config info to be associated with the RP and be available through the placement api
12:22:48 sean-k-mooney im sure we can come up with a clean way to store config info for a resouce provders such as adding a new field to the rp itself or reusing the description field
12:23:06 rgerganov sean-k-mooney, +1
12:23:31 sean-k-mooney that config section though should be admin only
12:23:34 janki efried, but again there is no os-extended-server-attributes API
12:23:54 sean-k-mooney normal users never need to see it only openstack services like the virt diriver
12:24:40 efried janki I really don't understand the test case, but if the policy check failed, is it possible it didn't even get to the API you're concerned about?
12:24:51 sean-k-mooney janki: that api is being moved into the normal server respoce i belive
12:25:00 efried janki The API call in question, if it needed to use a specific microversion, would presumably be coded up to do that.
12:25:48 janki efried, I dont think it needed a specific microversion. I was just checking if that is the reason for failure
12:26:09 efried janki Okay, well, it sounds like the policy is the first thing to sort out.
12:27:19 sean-k-mooney janki: this might be of interest to you https://review.openstack.org/#/c/508101/5/specs/queens/approved/api-extensions-policy-removal.rst
12:28:02 sean-k-mooney you will see on line 75 the proposal is to add the extended attributes to the get server responce
12:29:57 janki sean-k-mooney, ya. looks like os-extended-server-attributes is a collection of REST APIs exposed by Nova. right?
12:30:33 sean-k-mooney janki: it was technically a buch of rest apis exposed by a nova api extention
12:31:04 sean-k-mooney janki: as of pike we nolonger support extending the api like that but the existing extions were more or less kept
12:31:59 sean-k-mooney janki: i think the useful ones will be adopted into the main api and the unsed ones will be drop. i would expect the extended server attribute to be adopted in to the server resouce
12:32:14 janki sean-k-mooney, ohk. so I am getting 404 on os-floating-ips and policy error for os-extended-server-attributes. could this be related?
12:32:51 sean-k-mooney janki: if you query with admin prviliges dose the same happen
12:33:16 sean-k-mooney janki: the reason for the clean up spec is that there are two set of policies applied to the extentions
12:33:36 sean-k-mooney the main api policies and the extention level policies that we are looking to remove
12:34:02 janki sean-k-mooney, I dont have acces to the setup. Its jenkins build. But I think the naswer is no. there is 1 more variable in policy.json "os_compute_api:os-floating-ips": "rule:admin_or_owner" so there are not related
12:36:45 sean-k-mooney dumb question im assumeing the setup is using neutron, im not sure how the extentions were enabled in the past but if it was using nova-network you would have no floating-ips
12:37:37 alex_xu janki: it looks like the floating pool isn't there
12:38:22 janki alex_xu, yes. because the API that creates floating ip is failing
12:39:24 alex_xu janki: ok, great, you already know that, I didn't read the full chat log yet :)
12:40:29 janki alex_xu, I am trying to figure out the reason for the API to fail. I checked, with help of efried, that microversion used is 2.1 (< 2.36).
12:42:11 janki alex_xu, also there is policy check failure for os-extended-server-attributes in log but no such failure for os-floating-ips
12:44:33 alex_xu janki: they are sounds unrelated
12:45:09 janki alex_xu, Ya. initially I thought they are but looking at sample policy file, I do think the same
12:48:22 openstackgerrit Jan Zerebecki proposed openstack/nova master: Only log not correcting allocation once per period https://review.openstack.org/508262
12:50:53 openstackgerrit Jan Zerebecki proposed openstack/nova master: Only log not correcting allocation once per period https://review.openstack.org/508262
12:55:22 alex_xu janki: looks like you needn't worry about that policy check failure, the log will be emitted each time when the user doesn't have the permission
12:55:33 alex_xu and that log looks like annoying
12:56:36 janki alex_xu, well then, the next step is look into tempest.conf
13:00:45 alex_xu nova api meeting is running at #openstack-meeting-4
13:01:34 mriedem sdague: johnthetubaguy: can you take a look at this ocata-only fix? https://review.openstack.org/#/c/512406/ turns out we backported something awhile ago that requires some special handling based on the version of libvirt you're running
13:03:39 gmann janki: what is error actually, i can help from tempest.conf side
13:06:28 janki gmann, so I am running tempest for ODL + Pike setup and https://github.com/openstack/tempest/edit/master/tempest/scenario/test_server_basic_ops.py is failing with error https://logs.opendaylight.org/releng/jenkins092/netvirt-csit-1node-openstack-pike-upstream-stateful-carbon/36/tempest/tempest_results.html.gz
13:07:05 stephenfin mriedem: Any suggestions on who else I can ask to review https://review.openstack.org/#/c/361140/ in jaypipes' absence?
13:08:06 janki gmann, as per https://developer.openstack.org/api-ref/compute/#create-allocate-floating-ip-address, I verified that microverion is 2.1 and there are no logs about policy failure for os-floating-ips
13:08:28 mriedem stephenfin: i think that dansmith guy is pretty smart
13:09:00 janki gmann, next is to check tempest.conf for proper flag settings
13:10:32 gmann janki: its not policy related, what FIP pool you configured on nova>
13:11:29 janki gmann, I didnot manually configure anything. These logs are from jenkins build. Which file would those be in? I can dig that file up
13:11:54 stephenfin mriedem: I agree: dansmith is a super smart guy who'd surely love to review https://review.openstack.org/#/c/361140/
13:11:57 gmann janki: sure in api meeting, ll ping/debug after that
13:12:18 janki gmann, ack.
13:13:08 stephenfin Then again, if jaypipes is back from the dead, he might also like to review https://review.openstack.org/#/c/361140/ ...
13:13:14 stephenfin No rest for the weary :)
13:26:42 efried stephenfin I'm probably just missing something fundamental on https://review.openstack.org/#/c/361140/
13:29:28 lyarwood mdbooth: re https://github.com/openstack/nova/commit/29735973336b1038a2a9fb9072049fe3ed151502 - do you recall why context.auth_token isn't set?
13:29:48 mdbooth lyarwood: looking
13:29:51 lyarwood mdbooth: https://github.com/openstack/nova/commit/5e650e3681d40069dacf1ea2e43b07b362cf1bc3 that introduced the workaround doesn't really say why it isn't there via init_host
13:29:54 cdent efried: going back through the log looking at the conversation you had with sean-k-mooney and rgerganov; where in the powervm universe is the mapping from rp uuid to <other> going to live?
13:30:42 stephenfin efried: The problem I'm trying to get at is, even if we somehow managed to keep non-PCI-needing instances off of PCI-having NUMA nodes, we can still end up in the situation where we have no free PCI-having NUMA nodes for PCI-needing instances
13:31:09 mdbooth lyarwood: Could it be cause there's no request context at that point?
13:31:18 mdbooth i.e. no user, no api call?
13:31:26 efried cdent That's a pretty big TBD. We don't have any persistent data on our "hypervisor" (the NovaLink partition), which is a fairly fundamental point of architecture. You're supposed to be able to trash the partition and recreate it with no loss.
13:31:32 stephenfin This approach reduces the possibility but it doesn't mitigate it entirely. We want to use the two in tandem, which is what you've kind of hinted at (I think)
13:31:45 lyarwood mdbooth: yeah sorry, https://github.com/openstack/nova/blob/master/nova/context.py#L279
13:31:59 sean-k-mooney stephenfin: that was the usecase we created https://github.com/openstack/nfv-filters/blob/master/nfv_filters/nova/scheduler/filters/aggregate_instance_type_filter.py to address
13:32:11 openstackgerrit Chris Dent proposed openstack/nova-specs master: Add spec for symmetric GET and PUT of allocations https://review.openstack.org/508164
13:32:16 efried stephenfin Yes. You'll never get around that limitation. Unless you want to totally deny non-PCI-needing instances from booting on PCI-having NUMA nodes. Which ain't reasonable IMO.
13:32:43 sean-k-mooney it allows you to create aggregates of nodes with scarce resouces and require that they are requested in the flavor to schedule to those nodes
13:32:55 stephenfin efried: But you will with this spec, which allows booting PCI-needing instances from using non-PCI-having NUMA nodes
13:32:58 efried cdent Off the cuff, the idea would be to make some part of the RP relate to some part of whatever thingy we're "mapping" to.
13:33:16 sean-k-mooney this filter will work with traits in request in the flavor too by the way
13:33:19 mdbooth lyarwood: Yeah, that would be required for a glance request, I guess
13:33:24 stephenfin sean-k-mooney: I saw that, but it wasn't too inflexible, hence https://blueprints.launchpad.net/nova/+spec/reserve-numa-with-pci
13:33:41 mdbooth Although there's useful work it can do without that
13:33:50 sean-k-mooney that is your new weigher right
13:34:00 lyarwood mdbooth: it's the cinder encryption metadata lookup that's failing at the moment
13:34:08 sean-k-mooney stephenfin: i tought that merged in pike?
13:34:24 efried stephenfin The "problem" you won't get around is where I boot a hundred non-PCI-needing instances and run out of non-PCI-having NUMA nodes after the first fifty, so I start using my PCI-having NUMA nodes. Once those are full, I can no longer boot "required"-affinity PCI-needing instances.
13:34:28 stephenfin sean-k-mooney: Correct. If you had to divide your cloud into aggregates for PCI devices, it made it very inflexible if you wanted to, say, start using a lot of PCI-having instances or vice versa
13:34:43 stephenfin sean-k-mooney: It did. I'm saying https://review.openstack.org/#/c/361140/ forms a one-two combo :)

Earlier   Later