Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1342219 > unrolled thread

Re: [PATCH 1/1] intel_pstate: Increase hold-off time before busyness is scaled

Started byStephane Gasparini <stephane.gasparini@linux.intel.com>
First post2016-02-24 18:10 +0100
Last post2016-02-25 21:00 +0100
Articles 2 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH 1/1] intel_pstate: Increase hold-off time before busyness is scaled Stephane Gasparini <stephane.gasparini@linux.intel.com> - 2016-02-24 18:10 +0100
    RE: [PATCH 1/1] intel_pstate: Increase hold-off time before busyness is scaled "Doug Smythies" <dsmythies@telus.net> - 2016-02-25 21:00 +0100

#1342219 — Re: [PATCH 1/1] intel_pstate: Increase hold-off time before busyness is scaled

FromStephane Gasparini <stephane.gasparini@linux.intel.com>
Date2016-02-24 18:10 +0100
SubjectRe: [PATCH 1/1] intel_pstate: Increase hold-off time before busyness is scaled
Message-ID<r5L6O-7D3-15@gated-at.bofh.it>
Hi Doug


> On Feb 19, 2016, at 5:38 PM, Doug Smythies <dsmythies@telus.net> wrote:
> 
> Hi Steph,
> 
> On 2016.02.19 03:12 Stephane Gasparini wrote:
>> 
>> The issue you are reporting looks like one we improved on android by using 
>> the average pstate instead of using the last requested pstate
>> 
>> We know that this is improving the ffmpeg encoding performance when using the
>> load algorithm.
>> 
>> see patch attached
>> 
>> This patch is only applied on get_target_pstate_use_cpu_load however you can give
>> it a try on get_target_pstate_use_performance
> 
> Yes, that type of patch works on the load based approach.
> I’m not talking about using average p-state in the scaled_busy computation.

I’m talking adding the output of the PID (the number of pstate to ad or subtract)
to the average pstate rather than adding this to the current p-sate.

The current p-state is in some situation not reflecting the reality as the 
current p-state can be imposed by a "linked CPU". This is the case when you have a
thread migration on "linked CPU" that was not loaded. Its current P-State will be low
while its average p-state will reflect the activity of the "linked CPU".

I will not claim this is a perfect solution, but this combined to the topology 
awareness of the scheduler is helping to take better decision.

> However, I do not think it works on the performance based approach. Why not?
> Well, and if I understand correctly, follow the math and you end up with:
> 
> scaled_busy = 100%
> 
> scaled_busy = (aperf * 100% / mperf) * (max_pstate / * ((aperf * max_pstate) / mperf))
> 
> ... Doug
> 
> 
—
Steph

[toc] | [next] | [standalone]


#1343441

From"Doug Smythies" <dsmythies@telus.net>
Date2016-02-25 21:00 +0100
Message-ID<r6aeS-g0-1@gated-at.bofh.it>
In reply to#1342219
Hi Steph,

On 2016.02.24 08:20 Stephane Gasparini wrote:
>> On Feb 19, 2016, at 5:38 PM, Doug Smythies <dsmythies@telus.net> wrote: 
>>> On 2016.02.19 03:12 Stephane Gasparini wrote:
>>> 
>>> The issue you are reporting looks like one we improved on android by using 
>>> the average pstate instead of using the last requested pstate
>>> 
>>> We know that this is improving the ffmpeg encoding performance when using the
>>> load algorithm.
>>> 
>>> see patch attached
>>> 
>>> This patch is only applied on get_target_pstate_use_cpu_load however you can give
>>> it a try on get_target_pstate_use_performance
>> 
>> Yes, that type of patch works on the load based approach.
>
> I’m not talking about using average p-state in the scaled_busy computation.
> I’m talking adding the output of the PID (the number of pstate to ad or subtract)
> to the average pstate rather than adding this to the current p-sate.

For the situation we are dealing with here, that would actually make it worse,
wouldn't it?

Let's work through a real very low load example from the Mel V2 patch where
the target pstate is increased whereas it should have been decreased:

Mel patch version 2 (12X hold off added to rjw 3 patch v10 set added to kernel 4.5-rc4):

CPU: 3
Core busy: 105
Scaled busy: 143
Old pstate: 25
New pstate: 34
mperf: 52039
aperf: 55097
tsc: 335265689
freq: 3599750 KHz
Load: 0.02%
Duration (mS): 98.293

New pstate = old pstate + (scaled_busy-setpoint) * p_gain
           = 25 + (143 - 97) * 0.2
           = 34 (as above)

Ave pstate = max_pstate * aperf / mperf
           = 34 * 55097 / 52039
           = 36

Steph average pstate method added to the above:
New pstate = ave pstate + (scaled_busy-setpoint) * p_gain
           = 36 + (143 - 97) * 0.2
           = 45 (before clamping)

Now, just for completeness show the no Mel patch math:
Scaled busy = Core busy * max_pstate / old pstate * sample time / duration
            = 105 * 34 / 25 * 10 / 98.293
            = 14.53
New pstate = old pstate + (scaled_busy-setpoint) * p_gain
            = 25 + (14.53 - 97) * .2
            = 8.5
            = 16 clamped minimum

Regardless, I coded the average pstate method and observe little
difference between it and the Mel V2 patch with limited testing.

... Doug

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web