Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1356939 > unrolled thread
| Started by | Michael Turquette <mturquette@baylibre.com> |
|---|---|
| First post | 2016-03-14 06:30 +0100 |
| Last post | 2016-03-16 01:10 +0100 |
| Articles | 20 on this page of 46 — 9 participants |
Back to article view | Back to linux.kernel
[PATCH 0/8] schedutil enhancements Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
[PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:30 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:50 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Steve Muckle <steve.muckle@linaro.org> - 2016-03-16 04:40 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Peter Zijlstra <peterz@infradead.org> - 2016-03-16 09:10 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Juri Lelli <Juri.Lelli@arm.com> - 2016-03-16 11:10 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Steve Muckle <steve.muckle@linaro.org> - 2016-03-16 19:00 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Michael Turquette <mturquette@baylibre.com> - 2016-03-16 23:10 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Juri Lelli <Juri.Lelli@arm.com> - 2016-03-17 10:40 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Steve Muckle <steve.muckle@linaro.org> - 2016-03-17 15:00 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Patrick Bellasi <patrick.bellasi@arm.com> - 2016-03-17 17:00 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-16 13:50 +0100
Re: [PATCH 4/8] cpufreq/schedutil: sysfs capacity margin tunable Michael Turquette <mturquette@baylibre.com> - 2016-03-16 23:10 +0100
[PATCH 1/8] sched/cpufreq: remove cpufreq_trigger_update() Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 1/8] sched/cpufreq: remove cpufreq_trigger_update() Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:20 +0100
Re: [PATCH 1/8] sched/cpufreq: remove cpufreq_trigger_update() Peter Zijlstra <peterz@infradead.org> - 2016-03-15 23:00 +0100
Re: [PATCH 1/8] sched/cpufreq: remove cpufreq_trigger_update() Peter Zijlstra <peterz@infradead.org> - 2016-03-16 09:10 +0100
[PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:30 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Steve Muckle <steve.muckle@linaro.org> - 2016-03-16 05:00 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Peter Zijlstra <peterz@infradead.org> - 2016-03-16 08:50 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Vincent Guittot <vincent.guittot@linaro.org> - 2016-03-16 09:40 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Peter Zijlstra <peterz@infradead.org> - 2016-03-16 10:00 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Vincent Guittot <vincent.guittot@linaro.org> - 2016-03-16 10:20 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-16 13:40 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Peter Zijlstra <peterz@infradead.org> - 2016-03-16 14:20 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-16 14:30 +0100
Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util Peter Zijlstra <peterz@infradead.org> - 2016-03-16 14:50 +0100
[PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Dietmar Eggemann <dietmar.eggemann@arm.com> - 2016-03-15 20:20 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Michael Turquette <mturquette@baylibre.com> - 2016-03-15 21:50 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Dietmar Eggemann <dietmar.eggemann@arm.com> - 2016-03-16 20:50 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Peter Zijlstra <peterz@infradead.org> - 2016-03-16 21:10 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-03-16 22:40 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:50 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Peter Zijlstra <peterz@infradead.org> - 2016-03-16 08:50 +0100
Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity Peter Zijlstra <peterz@infradead.org> - 2016-03-16 13:50 +0100
[PATCH 2/8] sched/fair: add margin to utilization update Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 2/8] sched/fair: add margin to utilization update Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:20 +0100
Re: [PATCH 2/8] sched/fair: add margin to utilization update Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:50 +0100
Re: [PATCH 2/8] sched/fair: add margin to utilization update Steve Muckle <steve.muckle@linaro.org> - 2016-03-16 04:00 +0100
Re: [PATCH 2/8] sched/fair: add margin to utilization update Michael Turquette <mturquette@baylibre.com> - 2016-03-16 23:20 +0100
[PATCH 3/8] sched/cpufreq: new cfs capacity margin helpers Michael Turquette <mturquette@baylibre.com> - 2016-03-14 06:30 +0100
Re: [PATCH 3/8] sched/cpufreq: new cfs capacity margin helpers Peter Zijlstra <peterz@infradead.org> - 2016-03-15 22:20 +0100
Re: [PATCH 0/8] schedutil enhancements "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-03-16 01:10 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-03-16 05:00 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdaMO-3gF-5@gated-at.bofh.it> |
| In reply to | #1358254 |
Hi Mike, On 03/15/2016 03:06 PM, Michael Turquette wrote: >>> > > void cpufreq_set_freq_update_hook(int cpu, struct freq_update_hook *hook, >>> > > + void (*func)(struct freq_update_hook *hook, >>> > > + enum sched_class_util sched_class, >>> > > + u64 time, unsigned long util, >>> > > + unsigned long max)); >> > >> > Have you looked at the asm that generated? At some point you'll start >> > spilling on the stack and it'll be a god awful mess. >> > > Is your complaint about the enum that I added to the existing function > signature, or do you not like the original function signature[0] from > Rafael's patch, sans enum? The ARM procedure call standard has the first four arguments in registers so the addition of the enum above will start using the stack.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 08:50 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdeno-5MN-13@gated-at.bofh.it> |
| In reply to | #1358254 |
On Tue, Mar 15, 2016 at 03:06:09PM -0700, Michael Turquette wrote:
> Quoting Peter Zijlstra (2016-03-15 14:25:20)
> > On Sun, Mar 13, 2016 at 10:22:09PM -0700, Michael Turquette wrote:
> > > +++ b/include/linux/sched.h
> > > @@ -2362,15 +2362,25 @@ extern u64 scheduler_tick_max_deferment(void);
> > > static inline bool sched_can_stop_tick(void) { return false; }
> > > #endif
> > >
> > > +enum sched_class_util {
> > > + cfs_util,
> > > + rt_util,
> > > + dl_util,
> > > + nr_util_types,
> > > +};
> > > +
> > > #ifdef CONFIG_CPU_FREQ
> > > struct freq_update_hook {
> > > + void (*func)(struct freq_update_hook *hook,
> > > + enum sched_class_util sched_class, u64 time,
> > > unsigned long util, unsigned long max);
> > > };
> > >
> > Have you looked at the asm that generated? At some point you'll start
> > spilling on the stack and it'll be a god awful mess.
> >
>
> Is your complaint about the enum that I added to the existing function
> signature, or do you not like the original function signature[0] from
> Rafael's patch, sans enum?
No, my complaint is more about the call ABI for all our platforms, at
some point we start passing arguments on the stack instead of through
registers.
I'm not sure where that starts hurting, but its always a concern when
adding arguments to functions.
[toc] | [prev] | [next] | [standalone]
| From | Vincent Guittot <vincent.guittot@linaro.org> |
|---|---|
| Date | 2016-03-16 09:40 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdf9M-6p0-7@gated-at.bofh.it> |
| In reply to | #1358595 |
On 16 March 2016 at 08:41, Peter Zijlstra <peterz@infradead.org> wrote:
> On Tue, Mar 15, 2016 at 03:06:09PM -0700, Michael Turquette wrote:
>> Quoting Peter Zijlstra (2016-03-15 14:25:20)
>> > On Sun, Mar 13, 2016 at 10:22:09PM -0700, Michael Turquette wrote:
>> > > +++ b/include/linux/sched.h
>> > > @@ -2362,15 +2362,25 @@ extern u64 scheduler_tick_max_deferment(void);
>> > > static inline bool sched_can_stop_tick(void) { return false; }
>> > > #endif
>> > >
>> > > +enum sched_class_util {
>> > > + cfs_util,
>> > > + rt_util,
>> > > + dl_util,
>> > > + nr_util_types,
>> > > +};
>> > > +
>> > > #ifdef CONFIG_CPU_FREQ
>> > > struct freq_update_hook {
>> > > + void (*func)(struct freq_update_hook *hook,
>> > > + enum sched_class_util sched_class, u64 time,
>> > > unsigned long util, unsigned long max);
>> > > };
>> > >
>> > Have you looked at the asm that generated? At some point you'll start
>> > spilling on the stack and it'll be a god awful mess.
>> >
>>
>> Is your complaint about the enum that I added to the existing function
>> signature, or do you not like the original function signature[0] from
>> Rafael's patch, sans enum?
>
> No, my complaint is more about the call ABI for all our platforms, at
> some point we start passing arguments on the stack instead of through
> registers.
>
> I'm not sure where that starts hurting, but its always a concern when
> adding arguments to functions.
I wonder if it's really worth passing per sched_class request to
sched_util ? sched_util is about selecting a frequency based on the
utilization of the CPU, it only needs a value that reflect the whole
utilization. Can't we sum (or whatever the formula we want to apply)
utilizations before calling cpufreq_update_util
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 10:00 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdft8-6zO-11@gated-at.bofh.it> |
| In reply to | #1358695 |
On Wed, Mar 16, 2016 at 09:29:59AM +0100, Vincent Guittot wrote: > I wonder if it's really worth passing per sched_class request to > sched_util ? sched_util is about selecting a frequency based on the > utilization of the CPU, it only needs a value that reflect the whole > utilization. Can't we sum (or whatever the formula we want to apply) > utilizations before calling cpufreq_update_util So I've thought the same; but I'm conflicted, its a shame to compute anything if the call then doesn't do anything with it. And keeping a structure of all the various numbers to pass in also has cost of yet another cacheline to touch.
[toc] | [prev] | [next] | [standalone]
| From | Vincent Guittot <vincent.guittot@linaro.org> |
|---|---|
| Date | 2016-03-16 10:20 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdfMu-6Yq-5@gated-at.bofh.it> |
| In reply to | #1358744 |
On 16 March 2016 at 09:53, Peter Zijlstra <peterz@infradead.org> wrote: > On Wed, Mar 16, 2016 at 09:29:59AM +0100, Vincent Guittot wrote: >> I wonder if it's really worth passing per sched_class request to >> sched_util ? sched_util is about selecting a frequency based on the >> utilization of the CPU, it only needs a value that reflect the whole >> utilization. Can't we sum (or whatever the formula we want to apply) >> utilizations before calling cpufreq_update_util > > So I've thought the same; but I'm conflicted, its a shame to compute > anything if the call then doesn't do anything with it. yes, at least we shoud skip all that stuff (including adding a margin) if no hook has been set in cpufreq_update_util. I also see potential optimization of updating the value only if the utilization has been decayed > > And keeping a structure of all the various numbers to pass in also has > cost of yet another cacheline to touch.
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rafael@kernel.org> |
|---|---|
| Date | 2016-03-16 13:40 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdiU2-vS-9@gated-at.bofh.it> |
| In reply to | #1358744 |
On Wed, Mar 16, 2016 at 9:53 AM, Peter Zijlstra <peterz@infradead.org> wrote: > On Wed, Mar 16, 2016 at 09:29:59AM +0100, Vincent Guittot wrote: >> I wonder if it's really worth passing per sched_class request to >> sched_util ? sched_util is about selecting a frequency based on the >> utilization of the CPU, it only needs a value that reflect the whole >> utilization. Can't we sum (or whatever the formula we want to apply) >> utilizations before calling cpufreq_update_util > > So I've thought the same; but I'm conflicted, its a shame to compute > anything if the call then doesn't do anything with it. > > And keeping a structure of all the various numbers to pass in also has > cost of yet another cacheline to touch. In principle we can use high-order bits of util and max to encode the information on where they come from. Of course, that translates to additional ifs in the governor, but I guess they are unavoidable anyway.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 14:20 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdjwK-YM-17@gated-at.bofh.it> |
| In reply to | #1358952 |
On Wed, Mar 16, 2016 at 01:39:10PM +0100, Rafael J. Wysocki wrote:
> On Wed, Mar 16, 2016 at 9:53 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> > On Wed, Mar 16, 2016 at 09:29:59AM +0100, Vincent Guittot wrote:
> >> I wonder if it's really worth passing per sched_class request to
> >> sched_util ? sched_util is about selecting a frequency based on the
> >> utilization of the CPU, it only needs a value that reflect the whole
> >> utilization. Can't we sum (or whatever the formula we want to apply)
> >> utilizations before calling cpufreq_update_util
> >
> > So I've thought the same; but I'm conflicted, its a shame to compute
> > anything if the call then doesn't do anything with it.
> >
> > And keeping a structure of all the various numbers to pass in also has
> > cost of yet another cacheline to touch.
>
> In principle we can use high-order bits of util and max to encode the
> information on where they come from.
>
> Of course, that translates to additional ifs in the governor, but I
> guess they are unavoidable anyway.
Another thing we can do, for as long as we have the indirect function
call anyway, is stuff extra pointers in that same cacheline we pull the
function from.
Something like the below; there's room for 8 pointers (including the
function pointer) in a cacheline.
That would allow the callback to fetch whatever data it feels is
required (could be all of it).
We could also put a u64 *now = &rq->clock in, which would leave another
4 pointers for DL/RT support.
And since we're then back to 1-2 arguments on the function, we can add a
flags/mask field to indicate what changed (and if the function
throttles, it can keep a mask of all that changed since last time it
actually did something, or allow punching through the throttle if our
minimum guarantee changes or whatnot).
(this would of course require we allocate struct update_util_data with
the proper alignment thingies etc..)
Then again, maybe this is somewhat overboard :-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index ba49c9efd0b2..d34d75c5cc93 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -3236,8 +3236,10 @@ static inline unsigned long rlimit_max(unsigned int limit)
#ifdef CONFIG_CPU_FREQ
struct update_util_data {
- void (*func)(struct update_util_data *data,
- u64 time, unsigned long util, unsigned long max);
+ unsigned long *cfs_util_avg;
+ unsigned long *cfs_util_max;
+
+ void (*func)(struct update_util_data *data, u64 time);
};
void cpufreq_set_update_util_data(int cpu, struct update_util_data *data);
diff --git a/kernel/sched/cpufreq.c b/kernel/sched/cpufreq.c
index 928c4ba32f68..de5b20b11de3 100644
--- a/kernel/sched/cpufreq.c
+++ b/kernel/sched/cpufreq.c
@@ -32,6 +32,9 @@ void cpufreq_set_update_util_data(int cpu, struct update_util_data *data)
if (WARN_ON(data && !data->func))
return;
+ data->cfs_util_avg = &cpu_rq(cpu)->cfs.avg.util_avg;
+ data->cfs_util_max = &cpu_rq(cpu)->cpu_capacity_orig;
+
rcu_assign_pointer(per_cpu(cpufreq_update_util_data, cpu), data);
}
EXPORT_SYMBOL_GPL(cpufreq_set_update_util_data);
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rafael@kernel.org> |
|---|---|
| Date | 2016-03-16 14:30 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdjGp-136-9@gated-at.bofh.it> |
| In reply to | #1358980 |
On Wed, Mar 16, 2016 at 2:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
> On Wed, Mar 16, 2016 at 01:39:10PM +0100, Rafael J. Wysocki wrote:
>> On Wed, Mar 16, 2016 at 9:53 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> > On Wed, Mar 16, 2016 at 09:29:59AM +0100, Vincent Guittot wrote:
>> >> I wonder if it's really worth passing per sched_class request to
>> >> sched_util ? sched_util is about selecting a frequency based on the
>> >> utilization of the CPU, it only needs a value that reflect the whole
>> >> utilization. Can't we sum (or whatever the formula we want to apply)
>> >> utilizations before calling cpufreq_update_util
>> >
>> > So I've thought the same; but I'm conflicted, its a shame to compute
>> > anything if the call then doesn't do anything with it.
>> >
>> > And keeping a structure of all the various numbers to pass in also has
>> > cost of yet another cacheline to touch.
>>
>> In principle we can use high-order bits of util and max to encode the
>> information on where they come from.
>>
>> Of course, that translates to additional ifs in the governor, but I
>> guess they are unavoidable anyway.
>
> Another thing we can do, for as long as we have the indirect function
> call anyway, is stuff extra pointers in that same cacheline we pull the
> function from.
>
> Something like the below; there's room for 8 pointers (including the
> function pointer) in a cacheline.
>
> That would allow the callback to fetch whatever data it feels is
> required (could be all of it).
>
> We could also put a u64 *now = &rq->clock in, which would leave another
> 4 pointers for DL/RT support.
>
> And since we're then back to 1-2 arguments on the function, we can add a
> flags/mask field to indicate what changed (and if the function
> throttles, it can keep a mask of all that changed since last time it
> actually did something, or allow punching through the throttle if our
> minimum guarantee changes or whatnot).
>
> (this would of course require we allocate struct update_util_data with
> the proper alignment thingies etc..)
>
> Then again, maybe this is somewhat overboard :-)
I was thinking about something along these lines, but then I thought
that passing in registers would be more efficient.
One advantage I can see here is that we don't pass arguments that may
not be used by the callee.
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index ba49c9efd0b2..d34d75c5cc93 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -3236,8 +3236,10 @@ static inline unsigned long rlimit_max(unsigned int limit)
>
> #ifdef CONFIG_CPU_FREQ
> struct update_util_data {
> - void (*func)(struct update_util_data *data,
> - u64 time, unsigned long util, unsigned long max);
> + unsigned long *cfs_util_avg;
> + unsigned long *cfs_util_max;
> +
> + void (*func)(struct update_util_data *data, u64 time);
> };
How do we ensure proper alignment?
> void cpufreq_set_update_util_data(int cpu, struct update_util_data *data);
> diff --git a/kernel/sched/cpufreq.c b/kernel/sched/cpufreq.c
> index 928c4ba32f68..de5b20b11de3 100644
> --- a/kernel/sched/cpufreq.c
> +++ b/kernel/sched/cpufreq.c
> @@ -32,6 +32,9 @@ void cpufreq_set_update_util_data(int cpu, struct update_util_data *data)
> if (WARN_ON(data && !data->func))
> return;
>
> + data->cfs_util_avg = &cpu_rq(cpu)->cfs.avg.util_avg;
> + data->cfs_util_max = &cpu_rq(cpu)->cpu_capacity_orig;
> +
> rcu_assign_pointer(per_cpu(cpufreq_update_util_data, cpu), data);
> }
> EXPORT_SYMBOL_GPL(cpufreq_set_update_util_data);
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 14:50 +0100 |
| Subject | Re: [PATCH 5/8] sched/cpufreq: pass sched class into cpufreq_update_util |
| Message-ID | <rdjZN-1bV-27@gated-at.bofh.it> |
| In reply to | #1358986 |
On Wed, Mar 16, 2016 at 02:23:21PM +0100, Rafael J. Wysocki wrote:
> On Wed, Mar 16, 2016 at 2:10 PM, Peter Zijlstra <peterz@infradead.org> wrote:
> > (this would of course require we allocate struct update_util_data with
> > the proper alignment thingies etc..)
> > diff --git a/include/linux/sched.h b/include/linux/sched.h
> > index ba49c9efd0b2..d34d75c5cc93 100644
> > --- a/include/linux/sched.h
> > +++ b/include/linux/sched.h
> > @@ -3236,8 +3236,10 @@ static inline unsigned long rlimit_max(unsigned int limit)
> >
> > #ifdef CONFIG_CPU_FREQ
> > struct update_util_data {
> > - void (*func)(struct update_util_data *data,
> > - u64 time, unsigned long util, unsigned long max);
> > + unsigned long *cfs_util_avg;
> > + unsigned long *cfs_util_max;
> > +
> > + void (*func)(struct update_util_data *data, u64 time);
> > };
we should add: ____cacheline_aligned here
> How do we ensure proper alignment?
Depends on the allocator; not all of them respect the struct alignment
attribute.
kernel/sched/cpufreq.c:DEFINE_PER_CPU(struct update_util_data *, cpufreq_update_util_data);
That one could use:
DEFINE_PER_CPU_ALIGNED() instead
as would this one:
drivers/cpufreq/cpufreq_governor.c:static DEFINE_PER_CPU(struct cpu_dbs_info, cpu_dbs);
Because when you cacheline align dbs_info, its update_util_data member
will also get the correct alignment because of the structure attribute.
[toc] | [prev] | [next] | [standalone]
| From | Michael Turquette <mturquette@baylibre.com> |
|---|---|
| Date | 2016-03-14 06:30 +0100 |
| Subject | [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rcteO-7Q2-17@gated-at.bofh.it> |
| In reply to | #1356939 |
arch_scale_freq_capacity is weird. It specifies an arch hook for an
implementation that could easily vary within an architecture or even a
chip family.
This patch helps to mitigate this weirdness by defaulting to the
cpufreq-provided implementation, which should work for all cases where
CONFIG_CPU_FREQ is set.
If CONFIG_CPU_FREQ is not set, then try to use an implementation
provided by the architecture. Failing that, fall back to
SCHED_CAPACITY_SCALE.
It may be desirable for cpufreq drivers to specify their own
implementation of arch_scale_freq_capacity in the future. The same is
true for platform code within an architecture. In both cases an
efficient implementation selector will need to be created and this patch
adds a comment to that effect.
Signed-off-by: Michael Turquette <mturquette+renesas@baylibre.com>
---
kernel/sched/sched.h | 16 +++++++++++++++-
1 file changed, 15 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 469d11d..37502ea 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1368,7 +1368,21 @@ static inline int hrtick_enabled(struct rq *rq)
#ifdef CONFIG_SMP
extern void sched_avg_update(struct rq *rq);
-#ifndef arch_scale_freq_capacity
+/*
+ * arch_scale_freq_capacity can be implemented by cpufreq, platform code or
+ * arch code. We select the cpufreq-provided implementation first. If it
+ * doesn't exist then we default to any other implementation provided from
+ * platform/arch code. If those do not exist then we use the default
+ * SCHED_CAPACITY_SCALE value below.
+ *
+ * Note that if cpufreq drivers or platform/arch code have competing
+ * implementations it is up to those subsystems to select one at runtime with
+ * an efficient solution, as we cannot tolerate the overhead of indirect
+ * functions (e.g. function pointers) in the scheduler fast path
+ */
+#ifdef CONFIG_CPU_FREQ
+#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
+#elif !defined(arch_scale_freq_capacity)
static __always_inline
unsigned long arch_scale_freq_capacity(struct sched_domain *sd, int cpu)
{
--
2.1.4
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2016-03-15 20:20 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rd2FA-6of-13@gated-at.bofh.it> |
| In reply to | #1356945 |
On 14/03/16 05:22, Michael Turquette wrote:
> arch_scale_freq_capacity is weird. It specifies an arch hook for an
> implementation that could easily vary within an architecture or even a
> chip family.
>
> This patch helps to mitigate this weirdness by defaulting to the
> cpufreq-provided implementation, which should work for all cases where
> CONFIG_CPU_FREQ is set.
>
> If CONFIG_CPU_FREQ is not set, then try to use an implementation
> provided by the architecture. Failing that, fall back to
> SCHED_CAPACITY_SCALE.
>
> It may be desirable for cpufreq drivers to specify their own
> implementation of arch_scale_freq_capacity in the future. The same is
> true for platform code within an architecture. In both cases an
> efficient implementation selector will need to be created and this patch
> adds a comment to that effect.
For me this independence of the scheduler code towards the actual
implementation of the Frequency Invariant Engine (FEI) was actually a
feature.
In EAS RFC5.2 (linux-arm.org/linux-power.git energy_model_rfc_v5.2 ,
which hasn't been posted to LKML) we establish the link in the ARCH code
(arch/arm64/include/asm/topology.h).
#ifdef CONFIG_CPU_FREQ
#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
...
+#endif
>
> Signed-off-by: Michael Turquette <mturquette+renesas@baylibre.com>
> ---
> kernel/sched/sched.h | 16 +++++++++++++++-
> 1 file changed, 15 insertions(+), 1 deletion(-)
>
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index 469d11d..37502ea 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -1368,7 +1368,21 @@ static inline int hrtick_enabled(struct rq *rq)
> #ifdef CONFIG_SMP
> extern void sched_avg_update(struct rq *rq);
>
> -#ifndef arch_scale_freq_capacity
> +/*
> + * arch_scale_freq_capacity can be implemented by cpufreq, platform code or
> + * arch code. We select the cpufreq-provided implementation first. If it
> + * doesn't exist then we default to any other implementation provided from
> + * platform/arch code. If those do not exist then we use the default
> + * SCHED_CAPACITY_SCALE value below.
> + *
> + * Note that if cpufreq drivers or platform/arch code have competing
> + * implementations it is up to those subsystems to select one at runtime with
> + * an efficient solution, as we cannot tolerate the overhead of indirect
> + * functions (e.g. function pointers) in the scheduler fast path
> + */
> +#ifdef CONFIG_CPU_FREQ
> +#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
> +#elif !defined(arch_scale_freq_capacity)
> static __always_inline
> unsigned long arch_scale_freq_capacity(struct sched_domain *sd, int cpu)
> {
>
[toc] | [prev] | [next] | [standalone]
| From | Michael Turquette <mturquette@baylibre.com> |
|---|---|
| Date | 2016-03-15 21:50 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rd44G-7jA-3@gated-at.bofh.it> |
| In reply to | #1358199 |
Quoting Dietmar Eggemann (2016-03-15 12:13:58)
> On 14/03/16 05:22, Michael Turquette wrote:
> > arch_scale_freq_capacity is weird. It specifies an arch hook for an
> > implementation that could easily vary within an architecture or even a
> > chip family.
> >
> > This patch helps to mitigate this weirdness by defaulting to the
> > cpufreq-provided implementation, which should work for all cases where
> > CONFIG_CPU_FREQ is set.
> >
> > If CONFIG_CPU_FREQ is not set, then try to use an implementation
> > provided by the architecture. Failing that, fall back to
> > SCHED_CAPACITY_SCALE.
> >
> > It may be desirable for cpufreq drivers to specify their own
> > implementation of arch_scale_freq_capacity in the future. The same is
> > true for platform code within an architecture. In both cases an
> > efficient implementation selector will need to be created and this patch
> > adds a comment to that effect.
>
> For me this independence of the scheduler code towards the actual
> implementation of the Frequency Invariant Engine (FEI) was actually a
> feature.
I do not agree that it is a strength; I think it is confusing. My
opinion is that cpufreq drivers should implement
arch_scale_freq_capacity. Having a sane fallback
(cpufreq_scale_freq_capacity) simply means that you can remove the
boilerplate from the arm32 and arm64 code, which is a win.
Furthermore, if we have multiple competing implementations of
arch_scale_freq_invariance, wouldn't it be better for all of them to
live in cpufreq drivers? This means we would only need to implement a
single run-time "selector".
On the other hand, if the implementation lives in arch code and we have
various implementations of arch_scale_freq_capacity within an
architecture, then each arch would need to implement this selector
function. Even worse then if we have a split where some implementations
live in drivers/cpufreq (e.g. intel_pstate) and others in arch/arm and
others in arch/arm64 ... now we have three selectors.
Note that this has nothing to do with cpu microarch invariance. I'm
happy for that to stay in arch code because we can have heterogeneous
cpus that do not scale frequency, and thus would not enable cpufreq.
But if your platform scales cpu frequency, then really cpufreq should be
in the loop.
>
> In EAS RFC5.2 (linux-arm.org/linux-power.git energy_model_rfc_v5.2 ,
> which hasn't been posted to LKML) we establish the link in the ARCH code
> (arch/arm64/include/asm/topology.h).
Right, sorry again about preemptively posting the patch. Total brainfart
on my part.
>
> #ifdef CONFIG_CPU_FREQ
> #define arch_scale_freq_capacity cpufreq_scale_freq_capacity
> ...
> +#endif
The above is no longer necessary with this patch. Same question as
above: why insist on the arch boilerplate?
Regards,
Mike
>
> >
> > Signed-off-by: Michael Turquette <mturquette+renesas@baylibre.com>
> > ---
> > kernel/sched/sched.h | 16 +++++++++++++++-
> > 1 file changed, 15 insertions(+), 1 deletion(-)
> >
> > diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> > index 469d11d..37502ea 100644
> > --- a/kernel/sched/sched.h
> > +++ b/kernel/sched/sched.h
> > @@ -1368,7 +1368,21 @@ static inline int hrtick_enabled(struct rq *rq)
> > #ifdef CONFIG_SMP
> > extern void sched_avg_update(struct rq *rq);
> >
> > -#ifndef arch_scale_freq_capacity
> > +/*
> > + * arch_scale_freq_capacity can be implemented by cpufreq, platform code or
> > + * arch code. We select the cpufreq-provided implementation first. If it
> > + * doesn't exist then we default to any other implementation provided from
> > + * platform/arch code. If those do not exist then we use the default
> > + * SCHED_CAPACITY_SCALE value below.
> > + *
> > + * Note that if cpufreq drivers or platform/arch code have competing
> > + * implementations it is up to those subsystems to select one at runtime with
> > + * an efficient solution, as we cannot tolerate the overhead of indirect
> > + * functions (e.g. function pointers) in the scheduler fast path
> > + */
> > +#ifdef CONFIG_CPU_FREQ
> > +#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
> > +#elif !defined(arch_scale_freq_capacity)
> > static __always_inline
> > unsigned long arch_scale_freq_capacity(struct sched_domain *sd, int cpu)
> > {
> >
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2016-03-16 20:50 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rdpCa-4V9-15@gated-at.bofh.it> |
| In reply to | #1358233 |
On 15/03/16 20:46, Michael Turquette wrote: > Quoting Dietmar Eggemann (2016-03-15 12:13:58) >> On 14/03/16 05:22, Michael Turquette wrote: [...] >> For me this independence of the scheduler code towards the actual >> implementation of the Frequency Invariant Engine (FEI) was actually a >> feature. > > I do not agree that it is a strength; I think it is confusing. My > opinion is that cpufreq drivers should implement > arch_scale_freq_capacity. Having a sane fallback > (cpufreq_scale_freq_capacity) simply means that you can remove the > boilerplate from the arm32 and arm64 code, which is a win. > > Furthermore, if we have multiple competing implementations of > arch_scale_freq_invariance, wouldn't it be better for all of them to > live in cpufreq drivers? This means we would only need to implement a > single run-time "selector". > > On the other hand, if the implementation lives in arch code and we have > various implementations of arch_scale_freq_capacity within an > architecture, then each arch would need to implement this selector > function. Even worse then if we have a split where some implementations > live in drivers/cpufreq (e.g. intel_pstate) and others in arch/arm and > others in arch/arm64 ... now we have three selectors. OK, now I see your point. What I don't understand is the fact why you want different foo_scale_freq_capacity() implementations per cpufreq drivers. IMHO we want to do the cpufreq.c based implementation to abstract from that (at least for target_index() cpufreq drivers). intel_pstate (setpolicy()) is an exception but my humble guess is that systems with intel_pstate driver have X86_FEATURE_APERFMPERF support. > Note that this has nothing to do with cpu microarch invariance. I'm > happy for that to stay in arch code because we can have heterogeneous > cpus that do not scale frequency, and thus would not enable cpufreq. > But if your platform scales cpu frequency, then really cpufreq should be > in the loop. Agreed. > >> >> In EAS RFC5.2 (linux-arm.org/linux-power.git energy_model_rfc_v5.2 , >> which hasn't been posted to LKML) we establish the link in the ARCH code >> (arch/arm64/include/asm/topology.h). > > Right, sorry again about preemptively posting the patch. Total brainfart > on my part. > >> >> #ifdef CONFIG_CPU_FREQ >> #define arch_scale_freq_capacity cpufreq_scale_freq_capacity >> ... >> +#endif > > The above is no longer necessary with this patch. Same question as > above: why insist on the arch boilerplate? OK. [...]
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 21:10 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rdpVw-5jX-41@gated-at.bofh.it> |
| In reply to | #1359258 |
On Wed, Mar 16, 2016 at 07:44:33PM +0000, Dietmar Eggemann wrote: > intel_pstate (setpolicy()) is an exception but my humble guess is that > systems with intel_pstate driver have X86_FEATURE_APERFMPERF support. A quick browse of the Intel SDM says you're right. It looks like everything after Pentium-M; so Core-Solo/Core-Duo and onwards have APERF/MPERF. And it looks like P6 class systems didn't have DVFS support at all, which basically leaves P4 and Pentium-M as the only chips to have DVFS support lacking APERF/MPERF. And while I haven't got any P4 based space heaters left, I might still have a Pentium-M class laptop somewhere (if it still boots).
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rjw@rjwysocki.net> |
|---|---|
| Date | 2016-03-16 22:40 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rdrkC-6j4-27@gated-at.bofh.it> |
| In reply to | #1359280 |
On Wednesday, March 16, 2016 09:07:52 PM Peter Zijlstra wrote: > On Wed, Mar 16, 2016 at 07:44:33PM +0000, Dietmar Eggemann wrote: > > intel_pstate (setpolicy()) is an exception but my humble guess is that > > systems with intel_pstate driver have X86_FEATURE_APERFMPERF support. > > A quick browse of the Intel SDM says you're right. It looks like > everything after Pentium-M; so Core-Solo/Core-Duo and onwards have > APERF/MPERF. > > And it looks like P6 class systems didn't have DVFS support at all, > which basically leaves P4 and Pentium-M as the only chips to have DVFS > support lacking APERF/MPERF. > > And while I haven't got any P4 based space heaters left, I might still > have a Pentium-M class laptop somewhere (if it still boots). intel_pstate depends on APERF/MPERF, so it won't work with anything without them.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-15 22:50 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rd50L-7UJ-27@gated-at.bofh.it> |
| In reply to | #1356945 |
On Sun, Mar 13, 2016 at 10:22:12PM -0700, Michael Turquette wrote:
> +++ b/kernel/sched/sched.h
> @@ -1368,7 +1368,21 @@ static inline int hrtick_enabled(struct rq *rq)
> #ifdef CONFIG_SMP
> extern void sched_avg_update(struct rq *rq);
>
> -#ifndef arch_scale_freq_capacity
> +#ifdef CONFIG_CPU_FREQ
> +#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
> +#elif !defined(arch_scale_freq_capacity)
> static __always_inline
> unsigned long arch_scale_freq_capacity(struct sched_domain *sd, int cpu)
> {
This could not allow x86 to use the APERF/MPERF thing, so no, can't be
right.
Maybe something like
#ifndef arch_scale_freq_capacity
#ifdef CONFIG_CPU_FREQ
#define arch_scale_freq_capacity cpufreq_scale_freq_capacity
#else
static __always_inline
unsigned long arch_scale_freq_capacity(..)
{
return SCHED_CAPACITY_SCALE;
}
#endif
#endif
Which will let the arch override and only falls back to cpufreq if
the arch doesn't do anything.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 08:50 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rdenp-5MN-23@gated-at.bofh.it> |
| In reply to | #1358276 |
On Tue, Mar 15, 2016 at 03:27:21PM -0700, Michael Turquette wrote:
> That solution scales for the case where architectures have different
> methods. It doesn't scale for the case where cpufreq drivers or platform
> code within the same arch have competing implementations.
Sure it does; no matter what interface we use on x86 to set the DVFS
hints (ACPI, intel_p_state, whatever), using APERF/MPERF is the only
actual way of telling WTH the actual frequency was.
> I'm happy with it as a stop-gap, because it will initially work for
> arm{64} and x86, but we'll still need run-time selection of
> arch_scale_freq_capacity some day. Once we have that, I think that we
> should favor a run-time provided implementation over the arch-provided
> one.
Also, I'm thinking we don't need any of this. Your
cpufreq_scale_freq_capacity() is completely and utterly pointless. Since
its implementation simply provides whatever frequency we selected its
identical to not using frequency invariant load metrics and having
cpufreq use the !inv formula.
See:
lkml.kernel.org/r/20160309163930.GP6356@twins.programming.kicks-ass.net
Now, something else (power aware scheduling etc..) might need the freq
invariant stuff, but cpufreq (which we're concerned with here) does not
unless arch_scale_freq_capacity() does something else than simply return
the value we've set earlier.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-16 13:50 +0100 |
| Subject | Re: [PATCH 8/8] sched: prefer cpufreq_scale_freq_capacity |
| Message-ID | <rdj3I-zc-11@gated-at.bofh.it> |
| In reply to | #1358600 |
On Wed, Mar 16, 2016 at 08:47:52AM +0100, Peter Zijlstra wrote:
> On Tue, Mar 15, 2016 at 03:27:21PM -0700, Michael Turquette wrote:
> > I'm happy with it as a stop-gap, because it will initially work for
> > arm{64} and x86, but we'll still need run-time selection of
> > arch_scale_freq_capacity some day. Once we have that, I think that we
> > should favor a run-time provided implementation over the arch-provided
> > one.
>
> Also, I'm thinking we don't need any of this. Your
> cpufreq_scale_freq_capacity() is completely and utterly pointless.
Scrap that, while true to instant utilization, this isn't true for our
case, since our utilization numbers carry history, and any frequency
change in that window is relevant.
[toc] | [prev] | [next] | [standalone]
| From | Michael Turquette <mturquette@baylibre.com> |
|---|---|
| Date | 2016-03-14 06:30 +0100 |
| Subject | [PATCH 2/8] sched/fair: add margin to utilization update |
| Message-ID | <rcteO-7Q2-21@gated-at.bofh.it> |
| In reply to | #1356939 |
Utilization contributions to cfs_rq->avg.util_avg are scaled for both
microarchitecture-invariance as well as frequency-invariance. This means
that any given utilization contribution will be scaled against the
current cpu capacity (cpu frequency). Contributions from long running
tasks, whose utilization grows larger over time, will asymptotically
approach the current capacity.
This causes a problem when using this utilization signal to select a
target cpu capacity (cpu frequency), as our signal will never exceed the
current capacity, which would otherwise be our signal to increase
frequency.
Solve this by introducing a default capacity margin that is added to the
utilization signal when requesting a change to capacity (cpu frequency).
The margin is 1280, or 1.25 x SCHED_CAPACITY_SCALE (1024). This is
equivalent to similar margins such as the default 125 value assigned to
struct sched_domain.imbalance_pct for load balancing, and to the 80%
up_threshold used by the legacy cpufreq ondemand governor.
Signed-off-by: Michael Turquette <mturquette+renesas@baylibre.com>
---
kernel/sched/fair.c | 18 ++++++++++++++++--
kernel/sched/sched.h | 3 +++
2 files changed, 19 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index a32f281..29e8bae 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -100,6 +100,19 @@ const_debug unsigned int sysctl_sched_migration_cost = 500000UL;
*/
unsigned int __read_mostly sysctl_sched_shares_window = 10000000UL;
+/*
+ * Add a 25% margin globally to all capacity requests from cfs. This is
+ * equivalent to an 80% up_threshold in legacy governors like ondemand.
+ *
+ * This is required as task utilization increases. The frequency-invariant
+ * utilization will asymptotically approach the current capacity of the cpu and
+ * the additional margin will cross the threshold into the next capacity state.
+ *
+ * XXX someday expand to separate, per-call site margins? e.g. enqueue, fork,
+ * task_tick, load_balance, etc
+ */
+unsigned long cfs_capacity_margin = CAPACITY_MARGIN_DEFAULT;
+
#ifdef CONFIG_CFS_BANDWIDTH
/*
* Amount of runtime to allocate from global (tg) to local (per-cfs_rq) pool
@@ -2840,6 +2853,8 @@ static inline void update_load_avg(struct sched_entity *se, int update_tg)
if (cpu == smp_processor_id() && &rq->cfs == cfs_rq) {
unsigned long max = rq->cpu_capacity_orig;
+ unsigned long cap = cfs_rq->avg.util_avg *
+ cfs_capacity_margin / max;
/*
* There are a few boundary cases this might miss but it should
@@ -2852,8 +2867,7 @@ static inline void update_load_avg(struct sched_entity *se, int update_tg)
* thread is a different class (!fair), nor will the utilization
* number include things like RT tasks.
*/
- cpufreq_update_util(rq_clock(rq),
- min(cfs_rq->avg.util_avg, max), max);
+ cpufreq_update_util(rq_clock(rq), min(cap, max), max);
}
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index f06dfca..8c93ed2 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -27,6 +27,9 @@ extern __read_mostly int scheduler_running;
extern unsigned long calc_load_update;
extern atomic_long_t calc_load_tasks;
+#define CAPACITY_MARGIN_DEFAULT 1280;
+extern unsigned long cfs_capacity_margin;
+
extern void calc_global_load_tick(struct rq *this_rq);
extern long calc_load_fold_active(struct rq *this_rq);
--
2.1.4
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-03-15 22:20 +0100 |
| Subject | Re: [PATCH 2/8] sched/fair: add margin to utilization update |
| Message-ID | <rd4xI-7J5-15@gated-at.bofh.it> |
| In reply to | #1356946 |
On Sun, Mar 13, 2016 at 10:22:06PM -0700, Michael Turquette wrote:
> @@ -2840,6 +2853,8 @@ static inline void update_load_avg(struct sched_entity *se, int update_tg)
>
> if (cpu == smp_processor_id() && &rq->cfs == cfs_rq) {
> unsigned long max = rq->cpu_capacity_orig;
> + unsigned long cap = cfs_rq->avg.util_avg *
> + cfs_capacity_margin / max;
>
> /*
> * There are a few boundary cases this might miss but it should
> @@ -2852,8 +2867,7 @@ static inline void update_load_avg(struct sched_entity *se, int update_tg)
> * thread is a different class (!fair), nor will the utilization
> * number include things like RT tasks.
> */
> - cpufreq_update_util(rq_clock(rq),
> - min(cfs_rq->avg.util_avg, max), max);
> + cpufreq_update_util(rq_clock(rq), min(cap, max), max);
> }
> }
I really don't see why that is here, and not inside whatever uses
cpufreq_update_util().
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web