Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1330619 > unrolled thread

Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

Started by"Rafael J. Wysocki" <rjw@rjwysocki.net>
First post2016-02-09 21:10 +0100
Last post2016-02-11 21:50 +0100
Articles 20 on this page of 40 — 8 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-09 21:10 +0100
    Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Steve Muckle <steve.muckle@linaro.org> - 2016-02-10 02:10 +0100
      Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 03:00 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 04:10 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Steve Muckle <steve.muckle@linaro.org> - 2016-02-10 20:50 +0100
            Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 22:50 +0100
              Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Steve Muckle <steve.muckle@linaro.org> - 2016-02-10 23:10 +0100
                Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 23:20 +0100
      Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 13:10 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Juri Lelli <juri.lelli@arm.com> - 2016-02-11 13:30 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 16:30 +0100
            Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks Vincent Guittot <vincent.guittot@linaro.org> - 2016-02-11 19:30 +0100
              Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-12 15:10 +0100
                Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks Vincent Guittot <vincent.guittot@linaro.org> - 2016-02-12 15:50 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Steve Muckle <steve.muckle@linaro.org> - 2016-02-11 18:10 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-11 18:40 +0100
            Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 18:40 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 18:40 +0100
            Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Steve Muckle <steve.muckle@linaro.org> - 2016-02-11 20:00 +0100
              Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-11 20:10 +0100
                Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-12 14:50 +0100
              Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-12 15:20 +0100
                Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-12 17:10 +0100
                  Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-12 17:20 +0100
                    Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks Ashwin Chaugule <ashwin.chaugule@linaro.org> - 2016-02-12 18:00 +0100
                      Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-13 00:20 +0100
                  RE: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Doug Smythies" <dsmythies@telus.net> - 2016-02-12 18:10 +0100
                    Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-13 00:20 +0100
    Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Juri Lelli <juri.lelli@arm.com> - 2016-02-10 13:40 +0100
      Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 14:30 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Juri Lelli <juri.lelli@arm.com> - 2016-02-10 15:10 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 15:30 +0100
            Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Juri Lelli <juri.lelli@arm.com> - 2016-02-10 15:50 +0100
              Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-10 16:50 +0100
                Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Juri Lelli <juri.lelli@arm.com> - 2016-02-10 17:10 +0100
    Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 13:00 +0100
      Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-11 13:10 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-02-11 16:30 +0100
          Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-11 17:00 +0100
        Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-11 21:50 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1332748

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-12 14:50 +0100
Message-ID<r1mgG-7Hp-19@gated-at.bofh.it>
In reply to#1332323
On Thu, Feb 11, 2016 at 8:04 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Thu, Feb 11, 2016 at 7:52 PM, Steve Muckle <steve.muckle@linaro.org> wrote:
>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>>> My concern above is that pokes are guaranteed to keep occurring when
>>>> > there is only RT or DL activity so nothing breaks.
>>>
>>> The hook in their respective tick handler should ensure stuff is called
>>> sporadically and isn't stalled.
>>
>> But that's only true if the RT/DL tasks happen to be running when the
>> tick arrives right?
>>
>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>> the cpufreq update. This could go on for an arbitrarily long time
>> depending on the periodicity of the work.
>
> I'm thinking that two additional hooks in enqueue_task_rt/dl() might
> help here.  Then, we will hit either the tick or enqueue and that
> should do the trick.
>
> Peter, what do you think?

In any case I posted a v9 with those changes
(https://patchwork.kernel.org/patch/8290791/).

Again, it doesn't appear to break things.

If the enqueue hooks are bad (unwanted at all or in wrong places),
please let me know.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1332775 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromPeter Zijlstra <peterz@infradead.org>
Date2016-02-12 15:20 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r1mJI-87O-9@gated-at.bofh.it>
In reply to#1332320
On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
> >> My concern above is that pokes are guaranteed to keep occurring when
> >> > there is only RT or DL activity so nothing breaks.
> >
> > The hook in their respective tick handler should ensure stuff is called
> > sporadically and isn't stalled.
> 
> But that's only true if the RT/DL tasks happen to be running when the
> tick arrives right?
> 
> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
> if no CFS tasks happen to be executing on that CPU, we'll never trigger
> the cpufreq update. This could go on for an arbitrarily long time
> depending on the periodicity of the work.

Possible yes, but why do we care? Such a CPU would be so much idle that
cpufreq doesn't matter one way or another, right?

[toc] | [prev] | [next] | [standalone]


#1332846

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-12 17:10 +0100
Message-ID<r1osa-Pl-25@gated-at.bofh.it>
In reply to#1332775
On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>> >> My concern above is that pokes are guaranteed to keep occurring when
>> >> > there is only RT or DL activity so nothing breaks.
>> >
>> > The hook in their respective tick handler should ensure stuff is called
>> > sporadically and isn't stalled.
>>
>> But that's only true if the RT/DL tasks happen to be running when the
>> tick arrives right?
>>
>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>> the cpufreq update. This could go on for an arbitrarily long time
>> depending on the periodicity of the work.
>
> Possible yes, but why do we care? Such a CPU would be so much idle that
> cpufreq doesn't matter one way or another, right?

Well, in theory you can get 50% or so of the time active in bursts
that happen to fit between ticks.  If we happen to do those in the
lowest P-state, we may burn more energy than necessary on platforms
where more idle is preferred.

[toc] | [prev] | [next] | [standalone]


#1332854

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-12 17:20 +0100
Message-ID<r1oBQ-SL-15@gated-at.bofh.it>
In reply to#1332846
On Fri, Feb 12, 2016 at 5:01 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>> >> My concern above is that pokes are guaranteed to keep occurring when
>>> >> > there is only RT or DL activity so nothing breaks.
>>> >
>>> > The hook in their respective tick handler should ensure stuff is called
>>> > sporadically and isn't stalled.
>>>
>>> But that's only true if the RT/DL tasks happen to be running when the
>>> tick arrives right?
>>>
>>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>>> the cpufreq update. This could go on for an arbitrarily long time
>>> depending on the periodicity of the work.
>>
>> Possible yes, but why do we care? Such a CPU would be so much idle that
>> cpufreq doesn't matter one way or another, right?
>
> Well, in theory you can get 50% or so of the time active in bursts
> that happen to fit between ticks.  If we happen to do those in the
> lowest P-state, we may burn more energy than necessary on platforms
> where more idle is preferred.

At least intel_pstate should be able to figure out which P-state to
use then on the APERF/MPERF basis.

[toc] | [prev] | [next] | [standalone]


#1332902

FromAshwin Chaugule <ashwin.chaugule@linaro.org>
Date2016-02-12 18:00 +0100
Message-ID<r1pey-17H-31@gated-at.bofh.it>
In reply to#1332854
On 12 February 2016 at 11:15, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Fri, Feb 12, 2016 at 5:01 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>>> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>>>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>>> >> My concern above is that pokes are guaranteed to keep occurring when
>>>> >> > there is only RT or DL activity so nothing breaks.
>>>> >
>>>> > The hook in their respective tick handler should ensure stuff is called
>>>> > sporadically and isn't stalled.
>>>>
>>>> But that's only true if the RT/DL tasks happen to be running when the
>>>> tick arrives right?
>>>>
>>>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>>>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>>>> the cpufreq update. This could go on for an arbitrarily long time
>>>> depending on the periodicity of the work.
>>>
>>> Possible yes, but why do we care? Such a CPU would be so much idle that
>>> cpufreq doesn't matter one way or another, right?
>>
>> Well, in theory you can get 50% or so of the time active in bursts
>> that happen to fit between ticks.  If we happen to do those in the
>> lowest P-state, we may burn more energy than necessary on platforms
>> where more idle is preferred.
>
> At least intel_pstate should be able to figure out which P-state to
> use then on the APERF/MPERF basis.

Speaking for the generic case, it would be great to make use of such
feedback counters for selecting the next freq request. Use (num of
cycles used/total cycles) to figure out %ON time for the CPU. I
understand its not the goal for this patch series, but in the future
if we can do this in your callbacks where possible, then I think we
will do better than Ondemand.

Regards,
Ashwin.

> --
> To unsubscribe from this list: send the line "unsubscribe linux-pm" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

[toc] | [prev] | [next] | [standalone]


#1333198

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-13 00:20 +0100
Message-ID<r1vah-5dV-1@gated-at.bofh.it>
In reply to#1332902
On Fri, Feb 12, 2016 at 5:53 PM, Ashwin Chaugule
<ashwin.chaugule@linaro.org> wrote:
> On 12 February 2016 at 11:15, Rafael J. Wysocki <rafael@kernel.org> wrote:
>> On Fri, Feb 12, 2016 at 5:01 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
>>> On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>>>> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>>>>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>>>> >> My concern above is that pokes are guaranteed to keep occurring when
>>>>> >> > there is only RT or DL activity so nothing breaks.
>>>>> >
>>>>> > The hook in their respective tick handler should ensure stuff is called
>>>>> > sporadically and isn't stalled.
>>>>>
>>>>> But that's only true if the RT/DL tasks happen to be running when the
>>>>> tick arrives right?
>>>>>
>>>>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>>>>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>>>>> the cpufreq update. This could go on for an arbitrarily long time
>>>>> depending on the periodicity of the work.
>>>>
>>>> Possible yes, but why do we care? Such a CPU would be so much idle that
>>>> cpufreq doesn't matter one way or another, right?
>>>
>>> Well, in theory you can get 50% or so of the time active in bursts
>>> that happen to fit between ticks.  If we happen to do those in the
>>> lowest P-state, we may burn more energy than necessary on platforms
>>> where more idle is preferred.
>>
>> At least intel_pstate should be able to figure out which P-state to
>> use then on the APERF/MPERF basis.
>
> Speaking for the generic case, it would be great to make use of such
> feedback counters for selecting the next freq request. Use (num of
> cycles used/total cycles) to figure out %ON time for the CPU. I
> understand its not the goal for this patch series, but in the future
> if we can do this in your callbacks where possible, then I think we
> will do better than Ondemand.

Yes, we can do that at least in principle.  intel_pstate is a proof of that.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1332916

From"Doug Smythies" <dsmythies@telus.net>
Date2016-02-12 18:10 +0100
Message-ID<r1poe-1rA-19@gated-at.bofh.it>
In reply to#1332846
On 2016.02.12 08:01 Rafael J. Wysocki wrote:
> On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>>>> My concern above is that pokes are guaranteed to keep occurring when
>>>>> there is only RT or DL activity so nothing breaks.
>>>>
>>>> The hook in their respective tick handler should ensure stuff is called
>>>> sporadically and isn't stalled.
>>>
>>> But that's only true if the RT/DL tasks happen to be running when the
>>> tick arrives right?
>>>
>>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>>> the cpufreq update. This could go on for an arbitrarily long time
>>> depending on the periodicity of the work.
>>
>> Possible yes, but why do we care? Such a CPU would be so much idle that
>> cpufreq doesn't matter one way or another, right?

> Well, in theory you can get 50% or so of the time active in bursts
> that happen to fit between ticks.  If we happen to do those in the
> lowest P-state, we may burn more energy than necessary on platforms
> where more idle is preferred.

I believe this happens considerably more often than is commonly thought,
and is the exact reason I was opposed to the introduction of the
"duration" method into the intel_pstate driver in the first
place. The probability of occurrence (of a relatively busy CPU being idle
on jiffy boundaries) is very use dependant, occurring more on desktops than
servers, and sometime more with video frame rate based tasks. Data to support
my claim is a couple of years old and not very complete, but I see the issue
often on trace data acquired from desktop users on bugzilla reports.

Disclaimer: I fully admit that my related tests on the other thread have
been rigged to exaggerate the issue.

[toc] | [prev] | [next] | [standalone]


#1333199

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-13 00:20 +0100
Message-ID<r1vai-5dV-3@gated-at.bofh.it>
In reply to#1332916
On Fri, Feb 12, 2016 at 6:02 PM, Doug Smythies <dsmythies@telus.net> wrote:
> On 2016.02.12 08:01 Rafael J. Wysocki wrote:
>> On Fri, Feb 12, 2016 at 3:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>>> On Thu, Feb 11, 2016 at 10:52:20AM -0800, Steve Muckle wrote:
>>>> On 02/11/2016 09:30 AM, Peter Zijlstra wrote:
>>>>>> My concern above is that pokes are guaranteed to keep occurring when
>>>>>> there is only RT or DL activity so nothing breaks.
>>>>>
>>>>> The hook in their respective tick handler should ensure stuff is called
>>>>> sporadically and isn't stalled.
>>>>
>>>> But that's only true if the RT/DL tasks happen to be running when the
>>>> tick arrives right?
>>>>
>>>> Couldn't we have RT/DL activity which doesn't overlap with the tick? And
>>>> if no CFS tasks happen to be executing on that CPU, we'll never trigger
>>>> the cpufreq update. This could go on for an arbitrarily long time
>>>> depending on the periodicity of the work.
>>>
>>> Possible yes, but why do we care? Such a CPU would be so much idle that
>>> cpufreq doesn't matter one way or another, right?
>
>> Well, in theory you can get 50% or so of the time active in bursts
>> that happen to fit between ticks.  If we happen to do those in the
>> lowest P-state, we may burn more energy than necessary on platforms
>> where more idle is preferred.
>
> I believe this happens considerably more often than is commonly thought,
> and is the exact reason I was opposed to the introduction of the
> "duration" method into the intel_pstate driver in the first
> place. The probability of occurrence (of a relatively busy CPU being idle
> on jiffy boundaries) is very use dependant, occurring more on desktops than
> servers, and sometime more with video frame rate based tasks. Data to support
> my claim is a couple of years old and not very complete, but I see the issue
> often on trace data acquired from desktop users on bugzilla reports.

The approach with update callbacks from the scheduler should not be
affected by this, because it takes updates not only at the tick time,
but also on other scheduler events.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1331118 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromJuri Lelli <juri.lelli@arm.com>
Date2016-02-10 13:40 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r0CdR-2bY-11@gated-at.bofh.it>
In reply to#1330619
Hi Rafael,

On 09/02/16 21:05, Rafael J. Wysocki wrote:

[...]

> +/**
> + * cpufreq_update_util - Take a note about CPU utilization changes.
> + * @util: Current utilization.
> + * @max: Utilization ceiling.
> + *
> + * This function is called by the scheduler on every invocation of
> + * update_load_avg() on the CPU whose utilization is being updated.
> + */
> +void cpufreq_update_util(unsigned long util, unsigned long max)
> +{
> +	struct update_util_data *data;
> +
> +	rcu_read_lock();
> +
> +	data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
> +	if (data && data->func)
> +		data->func(data, cpu_clock(smp_processor_id()), util, max);

Are util and max used anywhere? It seems to me that cpu_clock is used by
the callbacks to check if the sampling period is elapsed, but I couldn't
yet find who is using util and max.

Thanks,

- Juri

[toc] | [prev] | [next] | [standalone]


#1331159

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-10 14:30 +0100
Message-ID<r0D0f-2IY-13@gated-at.bofh.it>
In reply to#1331118
On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> Hi Rafael,
>
> On 09/02/16 21:05, Rafael J. Wysocki wrote:
>
> [...]
>
>> +/**
>> + * cpufreq_update_util - Take a note about CPU utilization changes.
>> + * @util: Current utilization.
>> + * @max: Utilization ceiling.
>> + *
>> + * This function is called by the scheduler on every invocation of
>> + * update_load_avg() on the CPU whose utilization is being updated.
>> + */
>> +void cpufreq_update_util(unsigned long util, unsigned long max)
>> +{
>> +     struct update_util_data *data;
>> +
>> +     rcu_read_lock();
>> +
>> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
>> +     if (data && data->func)
>> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
>
> Are util and max used anywhere?

They aren't yet, but they will be.

Maybe not in this cycle (it it takes too much time to integrate the
preliminary changes), but we definitely are going to use those
numbers.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1331189 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromJuri Lelli <juri.lelli@arm.com>
Date2016-02-10 15:10 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r0DCY-3cV-59@gated-at.bofh.it>
In reply to#1331159
On 10/02/16 14:23, Rafael J. Wysocki wrote:
> On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> > Hi Rafael,
> >
> > On 09/02/16 21:05, Rafael J. Wysocki wrote:
> >
> > [...]
> >
> >> +/**
> >> + * cpufreq_update_util - Take a note about CPU utilization changes.
> >> + * @util: Current utilization.
> >> + * @max: Utilization ceiling.
> >> + *
> >> + * This function is called by the scheduler on every invocation of
> >> + * update_load_avg() on the CPU whose utilization is being updated.
> >> + */
> >> +void cpufreq_update_util(unsigned long util, unsigned long max)
> >> +{
> >> +     struct update_util_data *data;
> >> +
> >> +     rcu_read_lock();
> >> +
> >> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
> >> +     if (data && data->func)
> >> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
> >
> > Are util and max used anywhere?
> 
> They aren't yet, but they will be.
> 
> Maybe not in this cycle (it it takes too much time to integrate the
> preliminary changes), but we definitely are going to use those
> numbers.
> 

Oh OK. However, I was under the impression that this set was only
proposing a way to get rid of timers and use the scheduler as heartbeat
for cpufreq governors. The governors' sample based approach wouldn't
change, though. Am I wrong in assuming this?

Also, is linux-pm/bleeding-edge the one I want to fetch to try this set
out?

Thanks,

- Juri

[toc] | [prev] | [next] | [standalone]


#1331208

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-10 15:30 +0100
Message-ID<r0DWi-3jE-21@gated-at.bofh.it>
In reply to#1331189
On Wed, Feb 10, 2016 at 3:03 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> On 10/02/16 14:23, Rafael J. Wysocki wrote:
>> On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
>> > Hi Rafael,
>> >
>> > On 09/02/16 21:05, Rafael J. Wysocki wrote:
>> >
>> > [...]
>> >
>> >> +/**
>> >> + * cpufreq_update_util - Take a note about CPU utilization changes.
>> >> + * @util: Current utilization.
>> >> + * @max: Utilization ceiling.
>> >> + *
>> >> + * This function is called by the scheduler on every invocation of
>> >> + * update_load_avg() on the CPU whose utilization is being updated.
>> >> + */
>> >> +void cpufreq_update_util(unsigned long util, unsigned long max)
>> >> +{
>> >> +     struct update_util_data *data;
>> >> +
>> >> +     rcu_read_lock();
>> >> +
>> >> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
>> >> +     if (data && data->func)
>> >> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
>> >
>> > Are util and max used anywhere?
>>
>> They aren't yet, but they will be.
>>
>> Maybe not in this cycle (it it takes too much time to integrate the
>> preliminary changes), but we definitely are going to use those
>> numbers.
>>
>
> Oh OK. However, I was under the impression that this set was only
> proposing a way to get rid of timers and use the scheduler as heartbeat
> for cpufreq governors. The governors' sample based approach wouldn't
> change, though. Am I wrong in assuming this?

Your assumption is correct.

The sample-based approach doesn't change at this time, simply to avoid
making too many changes in one go.

The next step, as I'm seeing it, would be to use the
scheduler-provided utilization in the governor computations instead of
the load estimation made by governors themselves.

> Also, is linux-pm/bleeding-edge the one I want to fetch to try this set out?

You can get it from there, but possibly with some changes unrelated to cpufreq.

You can also pull from the pm-cpufreq-test branch to get the cpufreq
changes only.

Apart from that, I'm going resend the $subject set with updated patch
[1/3] for completeness.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1331233 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromJuri Lelli <juri.lelli@arm.com>
Date2016-02-10 15:50 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r0EfE-3qm-31@gated-at.bofh.it>
In reply to#1331208
On 10/02/16 15:26, Rafael J. Wysocki wrote:
> On Wed, Feb 10, 2016 at 3:03 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> > On 10/02/16 14:23, Rafael J. Wysocki wrote:
> >> On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> >> > Hi Rafael,
> >> >
> >> > On 09/02/16 21:05, Rafael J. Wysocki wrote:
> >> >
> >> > [...]
> >> >
> >> >> +/**
> >> >> + * cpufreq_update_util - Take a note about CPU utilization changes.
> >> >> + * @util: Current utilization.
> >> >> + * @max: Utilization ceiling.
> >> >> + *
> >> >> + * This function is called by the scheduler on every invocation of
> >> >> + * update_load_avg() on the CPU whose utilization is being updated.
> >> >> + */
> >> >> +void cpufreq_update_util(unsigned long util, unsigned long max)
> >> >> +{
> >> >> +     struct update_util_data *data;
> >> >> +
> >> >> +     rcu_read_lock();
> >> >> +
> >> >> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
> >> >> +     if (data && data->func)
> >> >> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
> >> >
> >> > Are util and max used anywhere?
> >>
> >> They aren't yet, but they will be.
> >>
> >> Maybe not in this cycle (it it takes too much time to integrate the
> >> preliminary changes), but we definitely are going to use those
> >> numbers.
> >>
> >
> > Oh OK. However, I was under the impression that this set was only
> > proposing a way to get rid of timers and use the scheduler as heartbeat
> > for cpufreq governors. The governors' sample based approach wouldn't
> > change, though. Am I wrong in assuming this?
> 
> Your assumption is correct.
> 

In this case. Wouldn't be possible to simply put the kicks in
sched/core.c? scheduler_tick() seems a good candidate for that, and you
could complement that with enqueue/dequeue/etc., if needed.

I'm actually wondering if a slow CONFIG_HZ might affect governors'
sampling rate. We might have scheduler tick firing every 40ms and
sampling rate set to 10 or 20ms, don't we?

> The sample-based approach doesn't change at this time, simply to avoid
> making too many changes in one go.
> 
> The next step, as I'm seeing it, would be to use the
> scheduler-provided utilization in the governor computations instead of
> the load estimation made by governors themselves.
> 

OK. But, I'm not sure what does this buy us. If the end goal is still to
do sampling, aren't we better off using the (1 - idle) estimation as
today?

> > Also, is linux-pm/bleeding-edge the one I want to fetch to try this set out?
> 
> You can get it from there, but possibly with some changes unrelated to cpufreq.
> 
> You can also pull from the pm-cpufreq-test branch to get the cpufreq
> changes only.
> 
> Apart from that, I'm going resend the $subject set with updated patch
> [1/3] for completeness.
> 

Great, thanks! Let's see if I can finally find time to run some tests
this time :).

Best,

- Juri

[toc] | [prev] | [next] | [standalone]


#1331284

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-10 16:50 +0100
Message-ID<r0FbI-42v-11@gated-at.bofh.it>
In reply to#1331233
On Wed, Feb 10, 2016 at 3:46 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> On 10/02/16 15:26, Rafael J. Wysocki wrote:
>> On Wed, Feb 10, 2016 at 3:03 PM, Juri Lelli <juri.lelli@arm.com> wrote:
>> > On 10/02/16 14:23, Rafael J. Wysocki wrote:
>> >> On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
>> >> > Hi Rafael,
>> >> >
>> >> > On 09/02/16 21:05, Rafael J. Wysocki wrote:
>> >> >
>> >> > [...]
>> >> >
>> >> >> +/**
>> >> >> + * cpufreq_update_util - Take a note about CPU utilization changes.
>> >> >> + * @util: Current utilization.
>> >> >> + * @max: Utilization ceiling.
>> >> >> + *
>> >> >> + * This function is called by the scheduler on every invocation of
>> >> >> + * update_load_avg() on the CPU whose utilization is being updated.
>> >> >> + */
>> >> >> +void cpufreq_update_util(unsigned long util, unsigned long max)
>> >> >> +{
>> >> >> +     struct update_util_data *data;
>> >> >> +
>> >> >> +     rcu_read_lock();
>> >> >> +
>> >> >> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
>> >> >> +     if (data && data->func)
>> >> >> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
>> >> >
>> >> > Are util and max used anywhere?
>> >>
>> >> They aren't yet, but they will be.
>> >>
>> >> Maybe not in this cycle (it it takes too much time to integrate the
>> >> preliminary changes), but we definitely are going to use those
>> >> numbers.
>> >>
>> >
>> > Oh OK. However, I was under the impression that this set was only
>> > proposing a way to get rid of timers and use the scheduler as heartbeat
>> > for cpufreq governors. The governors' sample based approach wouldn't
>> > change, though. Am I wrong in assuming this?
>>
>> Your assumption is correct.
>>
>
> In this case. Wouldn't be possible to simply put the kicks in
> sched/core.c? scheduler_tick() seems a good candidate for that, and you
> could complement that with enqueue/dequeue/etc., if needed.

That can be done, but they are not needed for things like idle and
stop, are they?

> I'm actually wondering if a slow CONFIG_HZ might affect governors'
> sampling rate. We might have scheduler tick firing every 40ms and
> sampling rate set to 10 or 20ms, don't we?

The smallest HZ you can get from the standard config is 100.  That
would translate to an update every 10ms roughly if my understanding of
things is correct.

Also I think that the scheduler and cpufreq should really work at the
same pace as they affect each other in any case.

>> The sample-based approach doesn't change at this time, simply to avoid
>> making too many changes in one go.
>>
>> The next step, as I'm seeing it, would be to use the
>> scheduler-provided utilization in the governor computations instead of
>> the load estimation made by governors themselves.
>>
>
> OK. But, I'm not sure what does this buy us. If the end goal is still to
> do sampling, aren't we better off using the (1 - idle) estimation as
> today?

First of all, we can avoid the need to compute this number entirely if
we use the scheduler-provided one.

Second, what if we come up with a different idea about the CPU
utilization than the scheduler has?  Who's right then?

Finally, the way this number is currently computed by cpufreq is based
on some questionable heuristics (and not just in one place), so maybe
it's better to stop doing that?

Also I didn't say that the *final* goal would be to do sampling.  I
was talking about the next step. :-)

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1331290 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromJuri Lelli <juri.lelli@arm.com>
Date2016-02-10 17:10 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r0Fv4-4sd-7@gated-at.bofh.it>
In reply to#1331284
On 10/02/16 16:46, Rafael J. Wysocki wrote:
> On Wed, Feb 10, 2016 at 3:46 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> > On 10/02/16 15:26, Rafael J. Wysocki wrote:
> >> On Wed, Feb 10, 2016 at 3:03 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> >> > On 10/02/16 14:23, Rafael J. Wysocki wrote:
> >> >> On Wed, Feb 10, 2016 at 1:33 PM, Juri Lelli <juri.lelli@arm.com> wrote:
> >> >> > Hi Rafael,
> >> >> >
> >> >> > On 09/02/16 21:05, Rafael J. Wysocki wrote:
> >> >> >
> >> >> > [...]
> >> >> >
> >> >> >> +/**
> >> >> >> + * cpufreq_update_util - Take a note about CPU utilization changes.
> >> >> >> + * @util: Current utilization.
> >> >> >> + * @max: Utilization ceiling.
> >> >> >> + *
> >> >> >> + * This function is called by the scheduler on every invocation of
> >> >> >> + * update_load_avg() on the CPU whose utilization is being updated.
> >> >> >> + */
> >> >> >> +void cpufreq_update_util(unsigned long util, unsigned long max)
> >> >> >> +{
> >> >> >> +     struct update_util_data *data;
> >> >> >> +
> >> >> >> +     rcu_read_lock();
> >> >> >> +
> >> >> >> +     data = rcu_dereference(*this_cpu_ptr(&cpufreq_update_util_data));
> >> >> >> +     if (data && data->func)
> >> >> >> +             data->func(data, cpu_clock(smp_processor_id()), util, max);
> >> >> >
> >> >> > Are util and max used anywhere?
> >> >>
> >> >> They aren't yet, but they will be.
> >> >>
> >> >> Maybe not in this cycle (it it takes too much time to integrate the
> >> >> preliminary changes), but we definitely are going to use those
> >> >> numbers.
> >> >>
> >> >
> >> > Oh OK. However, I was under the impression that this set was only
> >> > proposing a way to get rid of timers and use the scheduler as heartbeat
> >> > for cpufreq governors. The governors' sample based approach wouldn't
> >> > change, though. Am I wrong in assuming this?
> >>
> >> Your assumption is correct.
> >>
> >
> > In this case. Wouldn't be possible to simply put the kicks in
> > sched/core.c? scheduler_tick() seems a good candidate for that, and you
> > could complement that with enqueue/dequeue/etc., if needed.
> 
> That can be done, but they are not needed for things like idle and
> stop, are they?
> 

Sorry, I'm not sure I understand you here. In a NO_HZ system tick will
be stopped when idle.

> > I'm actually wondering if a slow CONFIG_HZ might affect governors'
> > sampling rate. We might have scheduler tick firing every 40ms and
> > sampling rate set to 10 or 20ms, don't we?
> 
> The smallest HZ you can get from the standard config is 100.  That
> would translate to an update every 10ms roughly if my understanding of
> things is correct.
> 

Right. Please, forget my question above :).

> Also I think that the scheduler and cpufreq should really work at the
> same pace as they affect each other in any case.
> 

Makes sense yes.

> >> The sample-based approach doesn't change at this time, simply to avoid
> >> making too many changes in one go.
> >>
> >> The next step, as I'm seeing it, would be to use the
> >> scheduler-provided utilization in the governor computations instead of
> >> the load estimation made by governors themselves.
> >>
> >
> > OK. But, I'm not sure what does this buy us. If the end goal is still to
> > do sampling, aren't we better off using the (1 - idle) estimation as
> > today?
> 
> First of all, we can avoid the need to compute this number entirely if
> we use the scheduler-provided one.
> 
> Second, what if we come up with a different idea about the CPU
> utilization than the scheduler has?  Who's right then?
> 
> Finally, the way this number is currently computed by cpufreq is based
> on some questionable heuristics (and not just in one place), so maybe
> it's better to stop doing that?
> 
> Also I didn't say that the *final* goal would be to do sampling.  I
> was talking about the next step. :-)
> 

Oh, this changes things indeed. :)

Thanks,

- Juri

[toc] | [prev] | [next] | [standalone]


#1331852 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromPeter Zijlstra <peterz@infradead.org>
Date2016-02-11 13:00 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r0Y4G-8bX-5@gated-at.bofh.it>
In reply to#1330619
On Tue, Feb 09, 2016 at 09:05:05PM +0100, Rafael J. Wysocki wrote:
> > > One concern I had was, given that the lone scheduler update hook is in
> > > CFS, is it possible for governor updates to be stalled due to RT or DL
> > > task activity?
> > 
> > I don't think they may be completely stalled, but I'd prefer Peter to
> > answer that as he suggested to do it this way.
> 
> In any case, if that concern turns out to be significant in practice, it may
> be addressed like in the appended modification of patch [1/3] from the $subject
> series.
> 
> With that things look like before from the cpufreq side, but the other sched
> classes also get a chance to trigger a cpufreq update.  The drawback is the
> cpu_clock() call instead of passing the time value from update_load_avg(), but
> I guess we can live with that if necessary.
> 
> FWIW, this modification doesn't seem to break things on my test machine.

Not really pretty though. It blows a bit that you require this callback
to be periodic (in order to replace a timer).

Ideally we'd not have to call this if state doesn't change.


> +++ linux-pm/include/linux/sched.h
> @@ -3207,4 +3207,11 @@ static inline unsigned long rlimit_max(u
>  	return task_rlimit_max(current, limit);
>  }
>  
> +void cpufreq_update_util(unsigned long util, unsigned long max);

Didn't you have a timestamp in there?

> +
> +static inline void cpufreq_kick(void)
> +{
> +	cpufreq_update_util(ULONG_MAX, ULONG_MAX);
> +}
> +
>  #endif

[toc] | [prev] | [next] | [standalone]


#1331877

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-11 13:10 +0100
Message-ID<r0Yeo-8w1-65@gated-at.bofh.it>
In reply to#1331852
On Thu, Feb 11, 2016 at 12:51 PM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Tue, Feb 09, 2016 at 09:05:05PM +0100, Rafael J. Wysocki wrote:
>> > > One concern I had was, given that the lone scheduler update hook is in
>> > > CFS, is it possible for governor updates to be stalled due to RT or DL
>> > > task activity?
>> >
>> > I don't think they may be completely stalled, but I'd prefer Peter to
>> > answer that as he suggested to do it this way.
>>
>> In any case, if that concern turns out to be significant in practice, it may
>> be addressed like in the appended modification of patch [1/3] from the $subject
>> series.
>>
>> With that things look like before from the cpufreq side, but the other sched
>> classes also get a chance to trigger a cpufreq update.  The drawback is the
>> cpu_clock() call instead of passing the time value from update_load_avg(), but
>> I guess we can live with that if necessary.
>>
>> FWIW, this modification doesn't seem to break things on my test machine.
>
> Not really pretty though. It blows a bit that you require this callback
> to be periodic (in order to replace a timer).

We need it for now, but that's because of how things work on the cpufreq side.

> Ideally we'd not have to call this if state doesn't change.

When cpufreq starts to use the util numbers, things will work like
that pretty much automatically.

We'll need to avoid thrashing if there are too many state changes over
a short time, but that's a different problem.

>> +++ linux-pm/include/linux/sched.h
>> @@ -3207,4 +3207,11 @@ static inline unsigned long rlimit_max(u
>>       return task_rlimit_max(current, limit);
>>  }
>>
>> +void cpufreq_update_util(unsigned long util, unsigned long max);
>
> Didn't you have a timestamp in there?

I did and I still do in fact.

The last version is here:

https://patchwork.kernel.org/patch/8275271/

but it has the additional hooks for RT/DL which you seem to be
thinking are a mistake.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1332149 — Re: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks

FromPeter Zijlstra <peterz@infradead.org>
Date2016-02-11 16:30 +0100
SubjectRe: [PATCH 0/3] cpufreq: Replace timers with utilization update callbacks
Message-ID<r11lU-2aN-15@gated-at.bofh.it>
In reply to#1331877
On Thu, Feb 11, 2016 at 01:08:28PM +0100, Rafael J. Wysocki wrote:
> > Not really pretty though. It blows a bit that you require this callback
> > to be periodic (in order to replace a timer).
> 
> We need it for now, but that's because of how things work on the cpufreq side.

Right, maybe stick a big comment on cpufreq_trigger_update() noting its
a big ugly hack and will go away 'soon'.

> The last version is here:
> 
> https://patchwork.kernel.org/patch/8275271/
> 
> but it has the additional hooks for RT/DL which you seem to be
> thinking are a mistake.

As long as we make sure everbody knows they're a band-aid and will be
taken out back and shot that should be fine for a little while I
suppose.

[toc] | [prev] | [next] | [standalone]


#1332180

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-11 17:00 +0100
Message-ID<r11OW-2po-23@gated-at.bofh.it>
In reply to#1332149
On Thu, Feb 11, 2016 at 4:29 PM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Thu, Feb 11, 2016 at 01:08:28PM +0100, Rafael J. Wysocki wrote:
>> > Not really pretty though. It blows a bit that you require this callback
>> > to be periodic (in order to replace a timer).
>>
>> We need it for now, but that's because of how things work on the cpufreq side.
>
> Right, maybe stick a big comment on cpufreq_trigger_update() noting its
> a big ugly hack and will go away 'soon'.

I will.

>> The last version is here:
>>
>> https://patchwork.kernel.org/patch/8275271/
>>
>> but it has the additional hooks for RT/DL which you seem to be
>> thinking are a mistake.
>
> As long as we make sure everbody knows they're a band-aid and will be
> taken out back and shot that should be fine for a little while I
> suppose.

Great, thanks!

Yes, I'm treating those as a band-aid for replacement.

Let me update the patch with a comment to explain that.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1332358

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-02-11 21:50 +0100
Message-ID<r16lA-5v2-17@gated-at.bofh.it>
In reply to#1331877
On Thu, Feb 11, 2016 at 1:08 PM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Thu, Feb 11, 2016 at 12:51 PM, Peter Zijlstra <peterz@infradead.org> wrote:
>> On Tue, Feb 09, 2016 at 09:05:05PM +0100, Rafael J. Wysocki wrote:
>>> > > One concern I had was, given that the lone scheduler update hook is in
>>> > > CFS, is it possible for governor updates to be stalled due to RT or DL
>>> > > task activity?
>>> >
>>> > I don't think they may be completely stalled, but I'd prefer Peter to
>>> > answer that as he suggested to do it this way.
>>>
>>> In any case, if that concern turns out to be significant in practice, it may
>>> be addressed like in the appended modification of patch [1/3] from the $subject
>>> series.
>>>
>>> With that things look like before from the cpufreq side, but the other sched
>>> classes also get a chance to trigger a cpufreq update.  The drawback is the
>>> cpu_clock() call instead of passing the time value from update_load_avg(), but
>>> I guess we can live with that if necessary.
>>>
>>> FWIW, this modification doesn't seem to break things on my test machine.
>>
>> Not really pretty though. It blows a bit that you require this callback
>> to be periodic (in order to replace a timer).
>
> We need it for now, but that's because of how things work on the cpufreq side.

In fact, I don't need the new callback to be invoked periodically.  I
only need it to be called often enough, where "enough" means at least
once in every sampling interval (for the lack of a better name) on the
rough average.  Less often than that may be kind of OK too depending
on the case.

I guess I need to explain that in more detail, though, at least for
the record if not anything else, so let me do that.

To start with let me note that things in cpufreq don't happen
periodically even today with timers, because all of those timers are
deferrable, so you never know when you'll get the next update
realistically.  We try to compensate for that in a kind of poor man's
way (which may be a source of problems by itself as mentioned by
Doug), but that's a band-aid rather.

With that in mind, there are two cases, the intel_pstate case and the
ondemand/conservative governor case.

intel_pstate is simpler, because it can do everything it needs in the
new callback (or in a timer function previously).  Periodicity might
matter to it, but it only uses two last points in its computations,
the current one and the previous one.  Thus it is not that important
how long the particular interval is.  Of course, if it is way too
long, we may miss some intermediate peaks and valleys and if the peaks
are intermittent enough, people may see poor performance.  In
practice, though, it turns out that the new callback is invoked (even
from CFS alone) much more frequently than we need on the average, so
we apply a "sample delay" rate limit to it.

In turn, the ondemand/conservative governor case is outright
ridiculous, because they don't even compute anything in the callback
(or a timer function previously).  They simply use it to spawn a work
item in process context that will estimate the "utilization" and
possibly change the P-state.  That may be delayed by the scheduling
interval, then pushed back by RT tasks and so on, so the time between
the moment they decide to take a "sample" and the moment that actually
happens may be, well, arbitrary.  So really timers are used here to
poke at things on a regular basis rather than for any actually
periodic stuff.

That may be improved in two ways in principle.  First, by moving as
much as we can into the utilization update callback without adding too
much overhead to the scheduler path.  Governor computations are the
primary candidate for that.  They need to take all of the tunables
accessible from user space into account, but that shouldn't be a big
problem.  We may be able to call at least some drivers from there too
(even the ACPI driver may be able to switch P-states via register
writes in some cases).  The second way would be to use the utilization
numbers provided by the scheduler for making governor decisions.

If we can do both, we should be much better off than we are today
already, even without the EAS stuff.

Thanks,
Rafael

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web