| Posted | Nick | Remark | |
|---|---|---|---|
| #openstack-nova - 2018-07-11 | |||
| 12:34:38 | gibi | in the notification | |
| 12:35:48 | gibi | yikun: commenting it in the review... | |
| 12:37:47 | gibi | yikun: done | |
| 12:38:00 | gibi | yikun: thanks for your patientes | |
| 12:38:42 | yikun | OK, I see, much thanks for your time and help. :) | |
| 12:38:45 | gibi | s/patientes/patience | |
| 12:39:07 | yikun | :), ha, I know | |
| 12:39:22 | gibi | :) | |
| 12:40:30 | yikun | and it time to leave and back home, have a good day. | |
| 12:40:46 | gibi | good day to you too | |
| 12:41:17 | mriedem | gibi: so we're leaving it as a string in the notification payload? | |
| 12:41:25 | gibi | mriedem: yes | |
| 12:42:32 | gibi | mriedem: because the field type is DictOfString and when we finally generate json schemas for notification classes then that will also created from the field typw | |
| 12:42:35 | gibi | type | |
| 12:44:23 | mriedem | i'm assuming yikun pointed out that for InstanceGroup.rules we coerce the known max_server_per_host field value to an int | |
| 12:47:31 | gibi | mriedem: yes, I saw that. gmann also pointed out that can be lead to some json validation inconvenience | |
| 12:47:41 | gibi | mriedem: if we flatten the API | |
| 12:49:03 | openstackgerrit | Matt Riedemann proposed openstack/nova master: Skip ServerShowV257Test.test_rebuild_server for cells v1 job https://review.openstack.org/581717 | |
| 12:51:04 | mriedem | for the rest api, we accept strings or ints apparently | |
| 12:51:07 | mriedem | which was news to me | |
| 12:51:13 | mriedem | for the positive_integer type | |
| 12:51:50 | mriedem | https://github.com/openstack/nova/blob/master/nova/api/validation/parameter_types.py#L238 | |
| 12:51:51 | gibi | mriedem: and we patternmatch, I see | |
| 12:52:16 | mriedem | so i'm not sure why {'max_server_per_host': 3} vs {'max_server_per_host': '3'} | |
| 12:52:20 | mriedem | will make a difference | |
| 12:52:39 | mriedem | tbc, the InstanceGroup.rules doing the cast to int for that value is for internal convenience more than anything, | |
| 12:52:54 | mriedem | so we don't have to remember to cast the value in all of the code that uses it, like the view builder in the api, the scheduler filter and the late affinity check in the compute | |
| 12:54:45 | gibi | mriedem: I think the api accepts strings for historical reasons | |
| 12:55:33 | gibi | mriedem: I sure I'm OK to cast it to int internally for counting and comparison | |
| 12:57:08 | mriedem | ok. i just don't understand the json schema validation concern | |
| 12:57:19 | mriedem | is that a concern for the consumer of the notification? | |
| 12:59:01 | gibi | mriedem: no it is not. I refer to gmann's comment in https://review.openstack.org/#/c/563401/28/nova/notifications/objects/server_group.py@41 about the need to move the validation from json to code | |
| 13:01:19 | mriedem | oy | |
| 13:01:36 | mriedem | there is only 1 policy per group, and the rules are applied to that policy on that group, | |
| 13:01:46 | mriedem | we've already named the InstanceGroup.rules field | |
| 13:03:58 | gibi | yeah that is gmann's point | |
| 13:03:59 | mriedem | which i think is ok, it's like a 2 line validator | |
| 13:04:12 | gibi | I agree | |
| 13:04:15 | mriedem | if group.policy == 'anti-affinity' and group.rules: raise HTTPBadRequest | |
| 13:04:23 | mriedem | oops | |
| 13:04:23 | openstackgerrit | Stephen Finucane proposed openstack/nova master: Revert "docs: Disable smartquotes" https://review.openstack.org/578841 | |
| 13:04:29 | mriedem | != 'anti-affinity' | |
| 13:05:04 | gibi | sure | |
| 13:05:30 | gibi | policy_rule feels a better name for me but I accept rule with a good api doc | |
| 13:06:51 | openstackgerrit | sahid proposed openstack/nova stable/queens: hardware: fix hugepages memory usage per intances https://review.openstack.org/581736 | |
| 13:08:14 | openstack | mriedem: Error: "=" is not a valid command. | |
| 13:17:08 | openstackgerrit | Matt Riedemann proposed openstack/nova stable/queens: Fix TypeError in prep_resize allocation cleanup https://review.openstack.org/581741 | |
| 13:17:20 | mriedem | gibi: i think you might have officially made yourself the other core to review https://review.openstack.org/#/c/563375/ | |
| 13:17:32 | mriedem | dansmith would +2 it but he made some non-trivial changes to it | |
| 13:19:56 | openstackgerrit | Chris Dent proposed openstack/nova master: [placement] add error.code on a ConcurrentUpdateDetected https://review.openstack.org/581742 | |
| 13:21:25 | mriedem | gibi: btw, you've been quite about reviews on the bw resource provider series - what's the status on that? are you looking for reviews on incremental things? is it done end to end or still working on changes for the full stack in nova (and maybe deps in neutron)? | |
| 13:21:55 | mriedem | my giant port binding live migration series is basically blocked on neutron dependencies | |
| 13:24:29 | openstackgerrit | Chris Dent proposed openstack/nova master: Test for unsanitized consumer UUID https://review.openstack.org/581137 | |
| 13:28:53 | gibi | mriedem: I will go through the remainings of https://review.openstack.org/#/c/563375/ soon | |
| 13:29:23 | gibi | mriedem: the status of bandwidth is that the code up on review stopped working after the official support for nested allocation_candidate support is merged | |
| 13:29:38 | gibi | mriedem: the solution is to support the new placement API versions from the scheduler report client | |
| 13:30:28 | gibi | mriedem: for that I started with a failing functional test here https://review.openstack.org/#/c/527728/ | |
| 13:31:05 | mriedem | yikun: gmann: alex_xu: so i see there must have been some discussion on https://review.openstack.org/#/c/567534/ and the request format wasn't changed but the response fields were changed...but no comments in the review about that decision? | |
| 13:37:32 | alex_xu | mriedem: no decision yet, I will try to go through the review | |
| 13:38:15 | alex_xu | mriedem: do you remember whether we discuss what should we do if the image traits check failed for rebuild https://review.openstack.org/#/c/569498/ ? | |
| 13:43:09 | mriedem | alex_xu: glancing at the comments, the question is if the instance should be in ACTIVE or ERROR state? | |
| 13:43:23 | alex_xu | mriedem: yea, that is the question | |
| 13:43:56 | mriedem | if we hit the scheduler during rebuild today because the image changes, and we get NoValidHost, we put the instance into ERROR state | |
| 13:44:22 | mriedem | so i'd expect we do the same if the new image traits are invalid for the instance.host - which is essentially the same thing as a NoValidHost we'd get from the scheduler | |
| 13:45:31 | mriedem | you can rebuild an instance in ERROR state btw | |
| 13:45:57 | alex_xu | oh, we can rebuild an instance in error, that sounds better | |
| 13:46:11 | mriedem | https://review.openstack.org/#/q/Ibb7bee15a3d4ee6f0ef53ba12e8b41f65a1fe999 | |
| 13:46:23 | mriedem | setting the instance to ERROR is a relatively new change | |
| 13:46:30 | alex_xu | otherwise I feel the user just chocie a wrong image, then his instance get into error, he only can get help from admin | |
| 13:47:09 | mriedem | they can rebuild with a good image | |
| 13:47:21 | mriedem | i remember this change and brought it up during the ptg in dublin | |
| 13:47:28 | mriedem | discussing maybe re-architecting how rebuild works, | |
| 13:47:37 | mriedem | the reason we set it to ERROR state is not just because scheduling failed, | |
| 13:47:47 | mriedem | but also because we've changed values on the instance itself in the api before casting to conductor, | |
| 13:47:54 | mriedem | so at that point, we can't rollback those changes | |
| 13:48:06 | mriedem | see my comments in https://review.openstack.org/#/c/536268/ | |
| 13:50:34 | alex_xu | mriedem: ah, i see now, that problem is out of the scope https://review.openstack.org/#/c/569498/ | |
| 13:51:09 | mriedem | yup | |
| 13:51:17 | mriedem | i need to review that rebuild + image traits change too | |
| 13:51:28 | mriedem | since i'm the one that got it stuck in committee during spec review | |
| 13:52:34 | mriedem | dansmith: so you don't see a need to name the new fields in the server group API policy_name and policy_rules right? | |
| 13:52:39 | alex_xu | mriedem: yea | |
| 13:52:55 | dansmith | mriedem: I don't, but gmann said he thought it was important | |
| 13:52:55 | mriedem | because yikun re-wrote the api change to do that, which makes it inconsistent with both the notification payload and the internal object modeling | |
| 13:53:06 | mriedem | i don't think it's important or really confusing | |
| 13:53:11 | dansmith | me either | |
| 13:53:18 | mriedem | i think it's more important that we have consistency up and down the stack | |
| 13:53:20 | mriedem | internal and external | |
| 13:57:26 | openstackgerrit | Eric Fried proposed openstack/nova master: Delete orphan compute nodes before updating resources https://review.openstack.org/579922 | |
| 14:03:22 | mriedem | gibi: just ping me if / when you need reviews on the bw provider changes | |
| 14:12:33 | openstackgerrit | Merged openstack/nova master: Remove irrelevant comment https://review.openstack.org/578821 | |
| 14:16:11 | openstackgerrit | Surya Seetharaman proposed openstack/nova master: Update queued-for-delete from the ComputeAPI during deletion/restoration https://review.openstack.org/566813 | |
| 14:33:00 | gibi | mriedem: thanks. | |
| 14:34:42 | efried | alex_xu: Are you still -1 on that change? Given that we've concluded NoValidHost is the right thing, having those logs be ERROR seems appropriate, 对吧 | |
| 14:37:30 | mriedem | efried: we don't ERROR for NoValidHost in the scheduler | |
| 14:37:41 | mriedem | ERROR means the operator needs to investigate b/c there is a problem in the system, | |
| 14:37:48 | mriedem | in this case, it's really a user error like a 400 | |
| 14:37:51 | mriedem | we don't log errors for that | |
| 14:37:56 | mriedem | debug at most | |
| 14:38:34 | efried | mriedem: So in this case, how would the user know what went wrong? | |