Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1220267 > unrolled thread
| Started by | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| First post | 2015-09-07 17:40 +0200 |
| Last post | 2015-09-13 13:10 +0200 |
| Articles | 18 on this page of 58 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-07 17:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-07 18:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-07 21:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-07 21:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 14:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 09:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 14:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 16:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 16:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 17:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-10 00:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-10 13:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-10 13:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-10 14:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-11 10:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-10 19:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 18:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-09 11:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-09 11:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-09 13:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 19:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-17 12:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-17 12:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-21 11:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-21 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-22 09:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Leo Yan <leo.yan@linaro.org> - 2015-09-11 09:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 12:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Leo Yan <leo.yan@linaro.org> - 2015-09-11 16:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-10 05:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-10 12:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 15:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 16:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 17:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 16:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-10 06:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-10 12:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-11 10:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 12:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-11 19:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-12 04:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-14 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-14 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-14 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-15 08:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-15 19:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-16 04:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-16 19:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-17 12:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-15 10:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-16 17:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 13:50 +0200
[tip:sched/core] sched/fair: Get rid of scaling utilization by capacity_orig tip-bot for Dietmar Eggemann <tipbot@zytor.com> - 2015-09-13 13:10 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-10 06:10 +0200 |
| Message-ID | <q71ln-7gf-1@gated-at.bofh.it> |
| In reply to | #1220760 |
On Tue, Sep 08, 2015 at 01:50:38PM +0100, Dietmar Eggemann wrote: > > It's both a unit and a SCALE/SHIFT problem, SCHED_LOAD_SHIFT and > > SCHED_CAPACITY_SHIFT are defined separately so we must be sure to > > scale the value in the right range. In the case of cpu_usage which > > returns sa->util_avg , it's the capacity range not the load range. > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and > CAPACITY have no unit. To be more accurate, probably, LOAD can be thought of as having unit, but UTIL has no unit. Anyway, those are my definitions: 1) unit, only for LOAD, and SCHED_LOAD_X is the unit (but SCHED_LOAD_RESOLUTION make it also some 2, see below) 2) range, aka, resolution or fix-point percentage (as Ben said) 3) timing ratio, LOAD_AVG_MAX etc, unralated with SCHED_LOAD_X > >> I always thought that scale_load_down() takes care of that. > > > > AFAIU, scale_load_down is a way to increase the resolution of the > > load not to move from load to capacity > > I tried to figure out why we have this issue when comparing UTIL w/ > CAPACITY and not LOAD w/ CAPACITY: > > Both are initialized like that: > > sa->load_avg = scale_load_down(se->load.weight); > sa->load_sum = sa->load_avg * LOAD_AVG_MAX; > sa->util_avg = scale_load_down(SCHED_LOAD_SCALE); > sa->util_sum = LOAD_AVG_MAX; > > and we use 'se->on_rq * scale_load_down(se->load.weight)' as 'unsigned > long weight' argument to call __update_load_avg() making sure the > scaling differences between LOAD and CAPACITY are respected while > updating sa->load_sum (and sa->load_avg). Yes, because we used SCHED_LOAD_X as both unit and range for LOAD. > OTAH, we don't apply a scale_load_down for sa->util_[sum/avg] only a '<< > SCHED_LOAD_SHIFT) / LOAD_AVG_MAX' on sa->util_avg. > So changing '<< SCHED_LOAD_SHIFT' to '* > scale_load_down(SCHED_LOAD_SCALE)' would be the logical thing to do. Actually, for UTIL, we only need range, so don't conflate with LOAD, what about we get all these clarified by redefining SCHED_LOAD_RESOLUTION as the resolution/range generic macro for LOAD, UTIL, and CAPACITY: #define SCHED_RESOLUTION_SHIFT 10 #define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT) #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */ # define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT) # define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT) # define SCHED_LOAD_SHIFT (10 + SCHED_RESOLUTION_SHIFT) #else # define scale_load(w) (w) # define scale_load_down(w) (w) # define SCHED_LOAD_SHIFT (10) #endif #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT) For UTIL, e.g., it will be initiated as: sa->util_avg = SCHED_RESOLUTION_SCALE; And for capacity, we just use SCHED_RESOLUTION_SHIFT (so SCHED_CAPACITY_SHIFT is not needed). Thanks, Yuyang -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-10 12:10 +0200 |
| Message-ID | <q76XN-6I3-39@gated-at.bofh.it> |
| In reply to | #1221872 |
On Thu, Sep 10, 2015 at 04:15:20AM +0800, Yuyang Du wrote: > On Tue, Sep 08, 2015 at 01:50:38PM +0100, Dietmar Eggemann wrote: > > > It's both a unit and a SCALE/SHIFT problem, SCHED_LOAD_SHIFT and > > > SCHED_CAPACITY_SHIFT are defined separately so we must be sure to > > > scale the value in the right range. In the case of cpu_usage which > > > returns sa->util_avg , it's the capacity range not the load range. > > > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and > > CAPACITY have no unit. > > To be more accurate, probably, LOAD can be thought of as having unit, > but UTIL has no unit. But I'm thinking that is wrong; it should have one, esp. if we go scale the thing. Giving it the same fixed point unit as load simplifies the code. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-11 10:20 +0200 |
| Message-ID | <q7rIS-4kB-21@gated-at.bofh.it> |
| In reply to | #1222076 |
On Thu, Sep 10, 2015 at 12:07:27PM +0200, Peter Zijlstra wrote:
> > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and
> > > CAPACITY have no unit.
> >
> > To be more accurate, probably, LOAD can be thought of as having unit,
> > but UTIL has no unit.
>
> But I'm thinking that is wrong; it should have one, esp. if we go scale
> the thing. Giving it the same fixed point unit as load simplifies the
> code.
I think we probably are saying the same thing with different terms. Anyway,
let me reiterate what I said and make it a little more formalized.
UTIL has no unit because it is pure ratio, the cpu_running%, which is in the
range of [0, 100%], and we increase the resolution, because we don't want
to lose many (due to integer rounding) by multiplying a number (say 1024), then
the range becomes [0, 1024].
CAPACITY is also a ratio of ACTUAL_PERF/MAX_PERF, from (0, 1]. Even LOAD
is the same, a ratio of NICE_X/NICE_0, from [15/1024=0.015, 88761/1024=86.68],
as it only has relativity meaning (i.e., when comparing to each other).
I said it has unit, it is in the sense that it looks like currency (for instance,
Yuan), you can use to buy CPU fair share. But it is just how you look at it and
there are certainly many other ways.
So, I still propose to generalize all these with the following patch, in the
belief that this makes it simple and clear, and error-reducing.
--
Subject: [PATCH] sched/fair: Generalize the load/util averages resolution
definition
A integer metric needs certain resolution to allow how much detail we
can look into (not losing detail by integer rounding), which also
determines the range of the metrics.
For instance, to increase the resolution of [0, 1] (two levels), one
can multiply 1024 and get [0, 1024] (1025 levels).
In sched/fair, a few metrics depend on the resolution: load/load_avg,
util_avg, and capacity (frequency adjustment). In order to reduce the
risks of making mistakes relating to resolution/range, we therefore
generalize the resolution by defining a basic resolution constant
number, and then formalize all metrics to depend on the basic
resolution. The basic resolution is 1024 or (1 << 10). Further, one
can recursively apply another basic resolution to increase the final
resolution (e.g., 1048676=1<<20).
Signed-off-by: Yuyang Du <yuyang.du@intel.com>
---
include/linux/sched.h | 2 +-
kernel/sched/sched.h | 12 +++++++-----
2 files changed, 8 insertions(+), 6 deletions(-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 119823d..55a7b93 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -912,7 +912,7 @@ enum cpu_idle_type {
/*
* Increase resolution of cpu_capacity calculations
*/
-#define SCHED_CAPACITY_SHIFT 10
+#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
#define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
/*
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 68cda11..d27cdd8 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
*/
#define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
+# define SCHED_RESOLUTION_SHIFT 10
+# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
+
/*
* Increase resolution of nice-level calculations for 64-bit architectures.
* The extra resolution improves shares distribution and load balancing of
@@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
* increased costs.
*/
#if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
-# define SCHED_LOAD_RESOLUTION 10
-# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
-# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
+# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
+# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
+# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
#else
-# define SCHED_LOAD_RESOLUTION 0
+# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
# define scale_load(w) (w)
# define scale_load_down(w) (w)
#endif
-#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
#define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
#define NICE_0_LOAD SCHED_LOAD_SCALE
--
1.9.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-11 12:30 +0200 |
| Message-ID | <q7tKF-7d0-1@gated-at.bofh.it> |
| In reply to | #1222609 |
On Fri, Sep 11, 2015 at 08:28:25AM +0800, Yuyang Du wrote:
> On Thu, Sep 10, 2015 at 12:07:27PM +0200, Peter Zijlstra wrote:
> > > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and
> > > > CAPACITY have no unit.
> > >
> > > To be more accurate, probably, LOAD can be thought of as having unit,
> > > but UTIL has no unit.
> >
> > But I'm thinking that is wrong; it should have one, esp. if we go scale
> > the thing. Giving it the same fixed point unit as load simplifies the
> > code.
>
> I think we probably are saying the same thing with different terms. Anyway,
> let me reiterate what I said and make it a little more formalized.
>
> UTIL has no unit because it is pure ratio, the cpu_running%, which is in the
> range of [0, 100%], and we increase the resolution, because we don't want
> to lose many (due to integer rounding) by multiplying a number (say 1024), then
> the range becomes [0, 1024].
Fully agree, and with frequency invariance we basically scale running
time to take into account that the cpu might be running slower that it
is capable of at the highest frequency. With cpu invariance also scale
by any difference their might be in max frequency and/or cpu
micro-archiecture so utilization becomes comparable between cpus. One
can also see it as we slow down or speed up time depending the current
compute capacity of the cpu relative to the max capacity.
> CAPACITY is also a ratio of ACTUAL_PERF/MAX_PERF, from (0, 1]. Even LOAD
> is the same, a ratio of NICE_X/NICE_0, from [15/1024=0.015, 88761/1024=86.68],
> as it only has relativity meaning (i.e., when comparing to each other).
Fully agree. Though 'LOAD' is a somewhat overloaded term in the
scheduler. Just to be clear, you refer to load.weight, load_avg is the
multiplication of load.weight and the task runnable time ratio.
> I said it has unit, it is in the sense that it looks like currency (for instance,
> Yuan), you can use to buy CPU fair share. But it is just how you look at it and
> there are certainly many other ways.
>
> So, I still propose to generalize all these with the following patch, in the
> belief that this makes it simple and clear, and error-reducing.
>
> --
>
> Subject: [PATCH] sched/fair: Generalize the load/util averages resolution
> definition
>
> A integer metric needs certain resolution to allow how much detail we
> can look into (not losing detail by integer rounding), which also
> determines the range of the metrics.
>
> For instance, to increase the resolution of [0, 1] (two levels), one
> can multiply 1024 and get [0, 1024] (1025 levels).
>
> In sched/fair, a few metrics depend on the resolution: load/load_avg,
> util_avg, and capacity (frequency adjustment). In order to reduce the
> risks of making mistakes relating to resolution/range, we therefore
> generalize the resolution by defining a basic resolution constant
> number, and then formalize all metrics to depend on the basic
> resolution. The basic resolution is 1024 or (1 << 10). Further, one
> can recursively apply another basic resolution to increase the final
> resolution (e.g., 1048676=1<<20).
>
> Signed-off-by: Yuyang Du <yuyang.du@intel.com>
> ---
> include/linux/sched.h | 2 +-
> kernel/sched/sched.h | 12 +++++++-----
> 2 files changed, 8 insertions(+), 6 deletions(-)
>
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index 119823d..55a7b93 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -912,7 +912,7 @@ enum cpu_idle_type {
> /*
> * Increase resolution of cpu_capacity calculations
> */
> -#define SCHED_CAPACITY_SHIFT 10
> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
>
> /*
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index 68cda11..d27cdd8 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> */
> #define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
>
> +# define SCHED_RESOLUTION_SHIFT 10
> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
> +
> /*
> * Increase resolution of nice-level calculations for 64-bit architectures.
> * The extra resolution improves shares distribution and load balancing of
> @@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> * increased costs.
> */
> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
> -# define SCHED_LOAD_RESOLUTION 10
> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
> #else
> -# define SCHED_LOAD_RESOLUTION 0
> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
> # define scale_load(w) (w)
> # define scale_load_down(w) (w)
> #endif
>
> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
>
> #define NICE_0_LOAD SCHED_LOAD_SCALE
I think this is pretty much the required relationship between all the
SHIFTs and SCALEs that Peter checked for in his #if-#error thing
earlier, so no disagreements from my side :-)
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-11 19:10 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q7zZM-7Wu-9@gated-at.bofh.it> |
| In reply to | #1222690 |
Morten Rasmussen <morten.rasmussen@arm.com> writes:
> On Fri, Sep 11, 2015 at 08:28:25AM +0800, Yuyang Du wrote:
>> On Thu, Sep 10, 2015 at 12:07:27PM +0200, Peter Zijlstra wrote:
>> > > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and
>> > > > CAPACITY have no unit.
>> > >
>> > > To be more accurate, probably, LOAD can be thought of as having unit,
>> > > but UTIL has no unit.
>> >
>> > But I'm thinking that is wrong; it should have one, esp. if we go scale
>> > the thing. Giving it the same fixed point unit as load simplifies the
>> > code.
>>
>> I think we probably are saying the same thing with different terms. Anyway,
>> let me reiterate what I said and make it a little more formalized.
>>
>> UTIL has no unit because it is pure ratio, the cpu_running%, which is in the
>> range of [0, 100%], and we increase the resolution, because we don't want
>> to lose many (due to integer rounding) by multiplying a number (say 1024), then
>> the range becomes [0, 1024].
>
> Fully agree, and with frequency invariance we basically scale running
> time to take into account that the cpu might be running slower that it
> is capable of at the highest frequency. With cpu invariance also scale
> by any difference their might be in max frequency and/or cpu
> micro-archiecture so utilization becomes comparable between cpus. One
> can also see it as we slow down or speed up time depending the current
> compute capacity of the cpu relative to the max capacity.
>
>> CAPACITY is also a ratio of ACTUAL_PERF/MAX_PERF, from (0, 1]. Even LOAD
>> is the same, a ratio of NICE_X/NICE_0, from [15/1024=0.015, 88761/1024=86.68],
>> as it only has relativity meaning (i.e., when comparing to each other).
>
> Fully agree. Though 'LOAD' is a somewhat overloaded term in the
> scheduler. Just to be clear, you refer to load.weight, load_avg is the
> multiplication of load.weight and the task runnable time ratio.
>
>> I said it has unit, it is in the sense that it looks like currency (for instance,
>> Yuan), you can use to buy CPU fair share. But it is just how you look at it and
>> there are certainly many other ways.
>>
>> So, I still propose to generalize all these with the following patch, in the
>> belief that this makes it simple and clear, and error-reducing.
>>
>> --
>>
>> Subject: [PATCH] sched/fair: Generalize the load/util averages resolution
>> definition
>>
>> A integer metric needs certain resolution to allow how much detail we
>> can look into (not losing detail by integer rounding), which also
>> determines the range of the metrics.
>>
>> For instance, to increase the resolution of [0, 1] (two levels), one
>> can multiply 1024 and get [0, 1024] (1025 levels).
>>
>> In sched/fair, a few metrics depend on the resolution: load/load_avg,
>> util_avg, and capacity (frequency adjustment). In order to reduce the
>> risks of making mistakes relating to resolution/range, we therefore
>> generalize the resolution by defining a basic resolution constant
>> number, and then formalize all metrics to depend on the basic
>> resolution. The basic resolution is 1024 or (1 << 10). Further, one
>> can recursively apply another basic resolution to increase the final
>> resolution (e.g., 1048676=1<<20).
>>
>> Signed-off-by: Yuyang Du <yuyang.du@intel.com>
>> ---
>> include/linux/sched.h | 2 +-
>> kernel/sched/sched.h | 12 +++++++-----
>> 2 files changed, 8 insertions(+), 6 deletions(-)
>>
>> diff --git a/include/linux/sched.h b/include/linux/sched.h
>> index 119823d..55a7b93 100644
>> --- a/include/linux/sched.h
>> +++ b/include/linux/sched.h
>> @@ -912,7 +912,7 @@ enum cpu_idle_type {
>> /*
>> * Increase resolution of cpu_capacity calculations
>> */
>> -#define SCHED_CAPACITY_SHIFT 10
>> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
>> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
>>
>> /*
>> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
>> index 68cda11..d27cdd8 100644
>> --- a/kernel/sched/sched.h
>> +++ b/kernel/sched/sched.h
>> @@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
>> */
>> #define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
>>
>> +# define SCHED_RESOLUTION_SHIFT 10
>> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
>> +
>> /*
>> * Increase resolution of nice-level calculations for 64-bit architectures.
>> * The extra resolution improves shares distribution and load balancing of
>> @@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
>> * increased costs.
>> */
>> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
>> -# define SCHED_LOAD_RESOLUTION 10
>> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
>> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
>> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
>> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
>> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
>> #else
>> -# define SCHED_LOAD_RESOLUTION 0
>> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
>> # define scale_load(w) (w)
>> # define scale_load_down(w) (w)
>> #endif
>>
>> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
>> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
>>
>> #define NICE_0_LOAD SCHED_LOAD_SCALE
>
> I think this is pretty much the required relationship between all the
> SHIFTs and SCALEs that Peter checked for in his #if-#error thing
> earlier, so no disagreements from my side :-)
> --
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at http://www.tux.org/lkml/
SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not
required to be the same value and should not be conflated.
In particular, since cgroups are on the same timeline as tasks and their
shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so
that SCHED_LOAD_RESOLUTION is invisible), changing that part of
SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0
= 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is
an internal value to the kernel.
In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require
recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-12 04:20 +0200 |
| Message-ID | <q7IA1-3A6-1@gated-at.bofh.it> |
| In reply to | #1222978 |
On Fri, Sep 11, 2015 at 10:05:53AM -0700, bsegall@google.com wrote: > > SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not > required to be the same value and should not be conflated. > In particular, since cgroups are on the same timeline as tasks and their > shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so > that SCHED_LOAD_RESOLUTION is invisible), changing that part of > SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0 > = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is > an internal value to the kernel. > > In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require > recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD. Not fully looked into the concerns, but the new SCHED_RESOLUTION_SHIFT is intended to formalize all the integer metrics that need better resolution. It is not special to any metric, so actually it is to de-conflate whoever is conflated. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-14 19:40 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q8FTs-4AL-25@gated-at.bofh.it> |
| In reply to | #1223287 |
Yuyang Du <yuyang.du@intel.com> writes: > On Fri, Sep 11, 2015 at 10:05:53AM -0700, bsegall@google.com wrote: >> >> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not >> required to be the same value and should not be conflated. > >> In particular, since cgroups are on the same timeline as tasks and their >> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so >> that SCHED_LOAD_RESOLUTION is invisible), changing that part of >> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0 >> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is >> an internal value to the kernel. >> >> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require >> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD. > > Not fully looked into the concerns, but the new SCHED_RESOLUTION_SHIFT > is intended to formalize all the integer metrics that need better resolution. > It is not special to any metric, so actually it is to de-conflate whoever is > conflated. It conflates the userspace-invisible SCHED_LOAD_RESOLUTION with the userspace-visible value of scale_load_down(NICE_0_LOAD). Increasing SCHED_LOAD_RESOLUTION must not change scale_load_down(NICE_0_LOAD). -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-14 15:00 +0200 |
| Message-ID | <q8Bwu-6Es-29@gated-at.bofh.it> |
| In reply to | #1222978 |
On Fri, Sep 11, 2015 at 10:05:53AM -0700, bsegall@google.com wrote:
> Morten Rasmussen <morten.rasmussen@arm.com> writes:
>
> > On Fri, Sep 11, 2015 at 08:28:25AM +0800, Yuyang Du wrote:
> >> diff --git a/include/linux/sched.h b/include/linux/sched.h
> >> index 119823d..55a7b93 100644
> >> --- a/include/linux/sched.h
> >> +++ b/include/linux/sched.h
> >> @@ -912,7 +912,7 @@ enum cpu_idle_type {
> >> /*
> >> * Increase resolution of cpu_capacity calculations
> >> */
> >> -#define SCHED_CAPACITY_SHIFT 10
> >> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
> >> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
> >>
> >> /*
> >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> >> index 68cda11..d27cdd8 100644
> >> --- a/kernel/sched/sched.h
> >> +++ b/kernel/sched/sched.h
> >> @@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> >> */
> >> #define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
> >>
> >> +# define SCHED_RESOLUTION_SHIFT 10
> >> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
> >> +
> >> /*
> >> * Increase resolution of nice-level calculations for 64-bit architectures.
> >> * The extra resolution improves shares distribution and load balancing of
> >> @@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> >> * increased costs.
> >> */
> >> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
> >> -# define SCHED_LOAD_RESOLUTION 10
> >> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
> >> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
> >> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
> >> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
> >> #else
> >> -# define SCHED_LOAD_RESOLUTION 0
> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
> >> # define scale_load(w) (w)
> >> # define scale_load_down(w) (w)
> >> #endif
> >>
> >> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> >> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
> >>
> >> #define NICE_0_LOAD SCHED_LOAD_SCALE
> >
> > I think this is pretty much the required relationship between all the
> > SHIFTs and SCALEs that Peter checked for in his #if-#error thing
> > earlier, so no disagreements from my side :-)
> > --
> > To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> > the body of a message to majordomo@vger.kernel.org
> > More majordomo info at http://vger.kernel.org/majordomo-info.html
> > Please read the FAQ at http://www.tux.org/lkml/
>
> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not
> required to be the same value and should not be conflated.
>
> In particular, since cgroups are on the same timeline as tasks and their
> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so
> that SCHED_LOAD_RESOLUTION is invisible), changing that part of
> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0
> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is
> an internal value to the kernel.
>
> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require
> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD.
I think I follow, but doesn't that mean that the current code is broken
too? NICE_0_LOAD changes if you change SCHED_LOAD_RESOLUTION:
#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
#define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
#define NICE_0_LOAD SCHED_LOAD_SCALE
#define NICE_0_SHIFT SCHED_LOAD_SHIFT
To me it sounds like we need to define it the other way around:
#define NICE_0_SHIFT 10
#define NICE_0_LOAD (1L << NICE_0_SHIFT)
and then add any additional resolution bits from there to ensure that
NICE_0_LOAD and the prio_to_weight/wmult tables are unchanged.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-14 19:40 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q8FTs-4AL-23@gated-at.bofh.it> |
| In reply to | #1224114 |
Morten Rasmussen <morten.rasmussen@arm.com> writes:
> On Fri, Sep 11, 2015 at 10:05:53AM -0700, bsegall@google.com wrote:
>> Morten Rasmussen <morten.rasmussen@arm.com> writes:
>>
>> > On Fri, Sep 11, 2015 at 08:28:25AM +0800, Yuyang Du wrote:
>> >> diff --git a/include/linux/sched.h b/include/linux/sched.h
>> >> index 119823d..55a7b93 100644
>> >> --- a/include/linux/sched.h
>> >> +++ b/include/linux/sched.h
>> >> @@ -912,7 +912,7 @@ enum cpu_idle_type {
>> >> /*
>> >> * Increase resolution of cpu_capacity calculations
>> >> */
>> >> -#define SCHED_CAPACITY_SHIFT 10
>> >> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
>> >> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
>> >>
>> >> /*
>> >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
>> >> index 68cda11..d27cdd8 100644
>> >> --- a/kernel/sched/sched.h
>> >> +++ b/kernel/sched/sched.h
>> >> @@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
>> >> */
>> >> #define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
>> >>
>> >> +# define SCHED_RESOLUTION_SHIFT 10
>> >> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
>> >> +
>> >> /*
>> >> * Increase resolution of nice-level calculations for 64-bit architectures.
>> >> * The extra resolution improves shares distribution and load balancing of
>> >> @@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
>> >> * increased costs.
>> >> */
>> >> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
>> >> -# define SCHED_LOAD_RESOLUTION 10
>> >> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
>> >> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
>> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
>> >> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
>> >> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
>> >> #else
>> >> -# define SCHED_LOAD_RESOLUTION 0
>> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
>> >> # define scale_load(w) (w)
>> >> # define scale_load_down(w) (w)
>> >> #endif
>> >>
>> >> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
>> >> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
>> >>
>> >> #define NICE_0_LOAD SCHED_LOAD_SCALE
>> >
>> > I think this is pretty much the required relationship between all the
>> > SHIFTs and SCALEs that Peter checked for in his #if-#error thing
>> > earlier, so no disagreements from my side :-)
>> > --
>> > To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
>> > the body of a message to majordomo@vger.kernel.org
>> > More majordomo info at http://vger.kernel.org/majordomo-info.html
>> > Please read the FAQ at http://www.tux.org/lkml/
>>
>> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not
>> required to be the same value and should not be conflated.
>>
>> In particular, since cgroups are on the same timeline as tasks and their
>> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so
>> that SCHED_LOAD_RESOLUTION is invisible), changing that part of
>> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0
>> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is
>> an internal value to the kernel.
>>
>> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require
>> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD.
>
> I think I follow, but doesn't that mean that the current code is broken
> too? NICE_0_LOAD changes if you change SCHED_LOAD_RESOLUTION:
>
> #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
>
> #define NICE_0_LOAD SCHED_LOAD_SCALE
> #define NICE_0_SHIFT SCHED_LOAD_SHIFT
>
> To me it sounds like we need to define it the other way around:
>
> #define NICE_0_SHIFT 10
> #define NICE_0_LOAD (1L << NICE_0_SHIFT)
>
> and then add any additional resolution bits from there to ensure that
> NICE_0_LOAD and the prio_to_weight/wmult tables are unchanged.
No, NICE_0_LOAD is supposed to be scale_load(prio_to_weight[nice_0]),
ie including SLR. It has never been clear to me what
SCHED_LOAD_SCALE/SCHED_LOAD_SHIFT were for as opposed to NICE_0_LOAD,
and the new utilization uses of it are entirely unlinked to 1024 == NICE_0
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-15 08:50 +0200 |
| Message-ID | <q8SdX-5qT-5@gated-at.bofh.it> |
| In reply to | #1224365 |
On Mon, Sep 14, 2015 at 10:34:00AM -0700, bsegall@google.com wrote: > >> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not > >> required to be the same value and should not be conflated. > >> > >> In particular, since cgroups are on the same timeline as tasks and their > >> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so > >> that SCHED_LOAD_RESOLUTION is invisible), changing that part of > >> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0 > >> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is > >> an internal value to the kernel. > >> > >> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require > >> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD. > > > > I think I follow, but doesn't that mean that the current code is broken > > too? NICE_0_LOAD changes if you change SCHED_LOAD_RESOLUTION: > > > > #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION) > > #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT) > > > > #define NICE_0_LOAD SCHED_LOAD_SCALE > > #define NICE_0_SHIFT SCHED_LOAD_SHIFT > > > > To me it sounds like we need to define it the other way around: > > > > #define NICE_0_SHIFT 10 > > #define NICE_0_LOAD (1L << NICE_0_SHIFT) > > > > and then add any additional resolution bits from there to ensure that > > NICE_0_LOAD and the prio_to_weight/wmult tables are unchanged. > > No, NICE_0_LOAD is supposed to be scale_load(prio_to_weight[nice_0]), > ie including SLR. It has never been clear to me what > SCHED_LOAD_SCALE/SCHED_LOAD_SHIFT were for as opposed to NICE_0_LOAD, > and the new utilization uses of it are entirely unlinked to 1024 == NICE_0 Presume your SLR means SCHED_LOAD_RESOLUTION: 1) The introduction of (not redefinition of) SCHED_RESOLUTION_SHIFT does not change anything after macro expansion. 2) The constants in prio_to_weight[] and prio_to_wmult[] are tied to a resolution of 10bits NICE_0, i.e., 1024, I guest it is the user visible part you mentioned, so is the cgroup share. To me, it is all ok. With the SCHED_RESOLUTION_SHIFT, the basic resolution unit, it is just for us to state clearly, the NICE_0's weight has a fixed resolution of SCHED_RESOLUTION_SHIFT, or even add this: #if prio_to_weight[20] != 1 << SCHED_RESOLUTION_SHIFT error "NICE_0 weight not calibrated" #endif /* I can learn, Peter */ I guess you are saying we are conflating NICE_0 with NICE_0_LOAD. But to me, they are just integer metrics, needing a resolution respectively. That is it. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-15 19:20 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q923D-2S5-15@gated-at.bofh.it> |
| In reply to | #1224666 |
Yuyang Du <yuyang.du@intel.com> writes: > On Mon, Sep 14, 2015 at 10:34:00AM -0700, bsegall@google.com wrote: >> >> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not >> >> required to be the same value and should not be conflated. >> >> >> >> In particular, since cgroups are on the same timeline as tasks and their >> >> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so >> >> that SCHED_LOAD_RESOLUTION is invisible), changing that part of >> >> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0 >> >> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is >> >> an internal value to the kernel. >> >> >> >> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require >> >> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD. >> > >> > I think I follow, but doesn't that mean that the current code is broken >> > too? NICE_0_LOAD changes if you change SCHED_LOAD_RESOLUTION: >> > >> > #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION) >> > #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT) >> > >> > #define NICE_0_LOAD SCHED_LOAD_SCALE >> > #define NICE_0_SHIFT SCHED_LOAD_SHIFT >> > >> > To me it sounds like we need to define it the other way around: >> > >> > #define NICE_0_SHIFT 10 >> > #define NICE_0_LOAD (1L << NICE_0_SHIFT) >> > >> > and then add any additional resolution bits from there to ensure that >> > NICE_0_LOAD and the prio_to_weight/wmult tables are unchanged. >> >> No, NICE_0_LOAD is supposed to be scale_load(prio_to_weight[nice_0]), >> ie including SLR. It has never been clear to me what >> SCHED_LOAD_SCALE/SCHED_LOAD_SHIFT were for as opposed to NICE_0_LOAD, >> and the new utilization uses of it are entirely unlinked to 1024 == NICE_0 > > Presume your SLR means SCHED_LOAD_RESOLUTION: > > 1) The introduction of (not redefinition of) SCHED_RESOLUTION_SHIFT does not > change anything after macro expansion. > > 2) The constants in prio_to_weight[] and prio_to_wmult[] are tied to a > resolution of 10bits NICE_0, i.e., 1024, I guest it is the user visible > part you mentioned, so is the cgroup share. > > To me, it is all ok. With the SCHED_RESOLUTION_SHIFT, the basic resolution > unit, it is just for us to state clearly, the NICE_0's weight has a fixed > resolution of SCHED_RESOLUTION_SHIFT, or even add this: > > #if prio_to_weight[20] != 1 << SCHED_RESOLUTION_SHIFT > error "NICE_0 weight not calibrated" > #endif > /* I can learn, Peter */ > > I guess you are saying we are conflating NICE_0 with NICE_0_LOAD. But to me, > they are just integer metrics, needing a resolution respectively. That is it. Yes this would change nothing at the moment post-expansion, that's not the point. SLR being 10 bits and the nice-0 being 1024 are completely and utterly unrelated and the headers should not pretend they need to be the same value, any more than there should be a #define that is shared with every other use of 1024 in the kernel. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-16 04:30 +0200 |
| Message-ID | <q9aDT-7JY-11@gated-at.bofh.it> |
| In reply to | #1225458 |
On Tue, Sep 15, 2015 at 10:11:41AM -0700, bsegall@google.com wrote: > > > > I guess you are saying we are conflating NICE_0 with NICE_0_LOAD. But to me, > > they are just integer metrics, needing a resolution respectively. That is it. > > Yes this would change nothing at the moment post-expansion, that's not > the point. SLR being 10 bits and the nice-0 being 1024 are completely > and utterly unrelated and the headers should not pretend they need to be > the same value, I never said they are related, why should they be related. And they need or need not to be the same value, fine. However, the SLR has to be a value. It is because it mighe be 10 or 20 (LOAD), therefore I make SCHED_RESOLUTION_SHIFT 10 (kind of a denominator). Not the other way around. We can define SCHED_RESOLUTION_SHIFT 1, and then define SLR = x * SCHED_RESOLUTION_SHIFT with x being a random number, if you must. And by the way, with SCHED_RESOLUTION_SHIFT, there will not be SLR anymore, we only need SCHED_LOAD_SHIFT, which has a low resolution 1*SCHED_RESOLUTION_SHIFT or a high one 2*SCHED_RESOLUTION_SHIFT. The scale_load*() is the conversion between the resolutions of NICE_0 and NICE_0_LOAD. > any more than there should be a #define that is shared > with every other use of 1024 in the kernel. The point really is, metrics (if not many ) need resolution, not just NICE_0_LOAD does. You can choose to either hardcode a number, like SCHED_CAPACITY_SHIFT now, or you can use SCHED_RESOLUTION_SHIFT, which is even as simple as a sign to say what the defined is (the scaled one with a better resolution vs. the original one). I guess this is to say we now have a (no-big-deal) resolution system. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-16 19:10 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q9onv-2h2-9@gated-at.bofh.it> |
| In reply to | #1225678 |
Yuyang Du <yuyang.du@intel.com> writes: > On Tue, Sep 15, 2015 at 10:11:41AM -0700, bsegall@google.com wrote: >> > >> > I guess you are saying we are conflating NICE_0 with NICE_0_LOAD. But to me, >> > they are just integer metrics, needing a resolution respectively. That is it. >> >> Yes this would change nothing at the moment post-expansion, that's not >> the point. SLR being 10 bits and the nice-0 being 1024 are completely >> and utterly unrelated and the headers should not pretend they need to be >> the same value, > > I never said they are related, why should they be related. And they need or > need not to be the same value, fine. > > However, the SLR has to be a value. It is because it mighe be 10 or 20 (LOAD), > therefore I make SCHED_RESOLUTION_SHIFT 10 (kind of a denominator). Not the > other way around. > > We can define SCHED_RESOLUTION_SHIFT 1, and then define SLR = x * SCHED_RESOLUTION_SHIFT > with x being a random number, if you must. That's sorta the point - you could do this and it would be just as (non-)sensical. > >> any more than there should be a #define that is shared >> with every other use of 1024 in the kernel. > > The point really is, metrics (if not many ) need resolution, not just NICE_0_LOAD does. > You can choose to either hardcode a number, like SCHED_CAPACITY_SHIFT now, > or you can use SCHED_RESOLUTION_SHIFT, which is even as simple as a sign to say what > the defined is (the scaled one with a better resolution vs. the original one). > I guess this is to say we now have a (no-big-deal) resolution system. Yes they were chosen for similar reasons, but they are not conceptually related, and you couldn't decide to just bump up all the resolutions by changing SCHED_RESOLUTION_SHIFT, so doing this would just be misleading. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-17 12:30 +0200 |
| Message-ID | <q9EBY-Ih-29@gated-at.bofh.it> |
| In reply to | #1226293 |
On Wed, Sep 16, 2015 at 10:06:24AM -0700, bsegall@google.com wrote: > > The point really is, metrics (if not many ) need resolution, not just NICE_0_LOAD does. > > You can choose to either hardcode a number, like SCHED_CAPACITY_SHIFT now, > > or you can use SCHED_RESOLUTION_SHIFT, which is even as simple as a sign to say what > > the defined is (the scaled one with a better resolution vs. the original one). > > I guess this is to say we now have a (no-big-deal) resolution system. > > Yes they were chosen for similar reasons, but they are not conceptually > related, and you couldn't decide to just bump up all the resolutions by > changing SCHED_RESOLUTION_SHIFT, so doing this would just be misleading. Yes, it appears they are made seemingly conceptually related. But probably it isn't worth a concern, if one knows it is just a scaled integer metric. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-15 10:40 +0200 |
| Message-ID | <q8TWq-7VB-17@gated-at.bofh.it> |
| In reply to | #1224365 |
On Mon, Sep 14, 2015 at 10:34:00AM -0700, bsegall@google.com wrote:
> Morten Rasmussen <morten.rasmussen@arm.com> writes:
>
> > On Fri, Sep 11, 2015 at 10:05:53AM -0700, bsegall@google.com wrote:
> >> Morten Rasmussen <morten.rasmussen@arm.com> writes:
> >>
> >> > On Fri, Sep 11, 2015 at 08:28:25AM +0800, Yuyang Du wrote:
> >> >> diff --git a/include/linux/sched.h b/include/linux/sched.h
> >> >> index 119823d..55a7b93 100644
> >> >> --- a/include/linux/sched.h
> >> >> +++ b/include/linux/sched.h
> >> >> @@ -912,7 +912,7 @@ enum cpu_idle_type {
> >> >> /*
> >> >> * Increase resolution of cpu_capacity calculations
> >> >> */
> >> >> -#define SCHED_CAPACITY_SHIFT 10
> >> >> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
> >> >> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
> >> >>
> >> >> /*
> >> >> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> >> >> index 68cda11..d27cdd8 100644
> >> >> --- a/kernel/sched/sched.h
> >> >> +++ b/kernel/sched/sched.h
> >> >> @@ -40,6 +40,9 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> >> >> */
> >> >> #define NS_TO_JIFFIES(TIME) ((unsigned long)(TIME) / (NSEC_PER_SEC / HZ))
> >> >>
> >> >> +# define SCHED_RESOLUTION_SHIFT 10
> >> >> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
> >> >> +
> >> >> /*
> >> >> * Increase resolution of nice-level calculations for 64-bit architectures.
> >> >> * The extra resolution improves shares distribution and load balancing of
> >> >> @@ -53,16 +56,15 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> >> >> * increased costs.
> >> >> */
> >> >> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
> >> >> -# define SCHED_LOAD_RESOLUTION 10
> >> >> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
> >> >> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
> >> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
> >> >> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
> >> >> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
> >> >> #else
> >> >> -# define SCHED_LOAD_RESOLUTION 0
> >> >> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
> >> >> # define scale_load(w) (w)
> >> >> # define scale_load_down(w) (w)
> >> >> #endif
> >> >>
> >> >> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> >> >> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
> >> >>
> >> >> #define NICE_0_LOAD SCHED_LOAD_SCALE
> >> >
> >> > I think this is pretty much the required relationship between all the
> >> > SHIFTs and SCALEs that Peter checked for in his #if-#error thing
> >> > earlier, so no disagreements from my side :-)
> >> > --
> >> > To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> >> > the body of a message to majordomo@vger.kernel.org
> >> > More majordomo info at http://vger.kernel.org/majordomo-info.html
> >> > Please read the FAQ at http://www.tux.org/lkml/
> >>
> >> SCHED_LOAD_RESOLUTION and the non-SLR part of SCHED_LOAD_SHIFT are not
> >> required to be the same value and should not be conflated.
> >>
> >> In particular, since cgroups are on the same timeline as tasks and their
> >> shares are not scaled by SCHED_LOAD_SHIFT in any way (but are scaled so
> >> that SCHED_LOAD_RESOLUTION is invisible), changing that part of
> >> SCHED_LOAD_SHIFT would cause issues, since things can assume that nice-0
> >> = 1024. However changing SCHED_LOAD_RESOLUTION would be fine, as that is
> >> an internal value to the kernel.
> >>
> >> In addition, changing the non-SLR part of SCHED_LOAD_SHIFT would require
> >> recomputing all of prio_to_weight/wmult for the new NICE_0_LOAD.
> >
> > I think I follow, but doesn't that mean that the current code is broken
> > too? NICE_0_LOAD changes if you change SCHED_LOAD_RESOLUTION:
> >
> > #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> > #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
> >
> > #define NICE_0_LOAD SCHED_LOAD_SCALE
> > #define NICE_0_SHIFT SCHED_LOAD_SHIFT
> >
> > To me it sounds like we need to define it the other way around:
> >
> > #define NICE_0_SHIFT 10
> > #define NICE_0_LOAD (1L << NICE_0_SHIFT)
> >
> > and then add any additional resolution bits from there to ensure that
> > NICE_0_LOAD and the prio_to_weight/wmult tables are unchanged.
>
> No, NICE_0_LOAD is supposed to be scale_load(prio_to_weight[nice_0]),
> ie including SLR. It has never been clear to me what
> SCHED_LOAD_SCALE/SCHED_LOAD_SHIFT were for as opposed to NICE_0_LOAD,
I see, I wasn't sure if NICE_0_LOAD is being used in the code somewhere
with the assumption that NICE_0_LOAD = load.weight = 1024. The
scale_(down_)_load() conversion between base load (nice_0 = 1024) and
hi-res load makes makes sense.
> and the new utilization uses of it are entirely unlinked to 1024 == NICE_0
Yes, agreed. For utilization we just need to define some fixed point
resolution (as Yuyang said). That resolution is independent of the hi-res
load additional bits and should remain so. The same fixed point
resolution has to be used for capacity as well unless we want to
introduce scale_(down_)_capacity() functions to allow utilization to be
compared to capacity.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-16 17:50 +0200 |
| Message-ID | <q9n86-jT-7@gated-at.bofh.it> |
| In reply to | #1224365 |
On Mon, Sep 14, 2015 at 10:34:00AM -0700, bsegall@google.com wrote: > It has never been clear to me what > SCHED_LOAD_SCALE/SCHED_LOAD_SHIFT were for as opposed to NICE_0_LOAD, SCHED_LOAD_SCALE/SHIFT are the fixed point mult/shift, and NICE_0_LOAD is the load of a nice-0 task. They happen to be the same by the choice that nice-0 has a load of 1 (the only natural choice given proportional weight and hyperboles etc..). But for the fixed point math we use SCHED_LOAD_* and only when we want to explicitly use the unit load we use NICE_0_LOAD (its only used 4-5 times or so). -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-08 13:50 +0200 |
| Message-ID | <q6pzs-39Q-5@gated-at.bofh.it> |
| In reply to | #1220360 |
On Mon, Sep 07, 2015 at 07:54:18PM +0100, Dietmar Eggemann wrote: > I see the point but IMHO this will only be necessary if the SCHED_LOAD_RESOLUTION > stuff gets re-enabled again. Paul, Ben, gentle reminder to look at re-enabling this. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | tip-bot for Dietmar Eggemann <tipbot@zytor.com> |
|---|---|
| Date | 2015-09-13 13:10 +0200 |
| Subject | [tip:sched/core] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q8dkv-5Qt-39@gated-at.bofh.it> |
| In reply to | #1220267 |
Commit-ID: 231678b768da07d19ab5683a39eeb0c250631d02
Gitweb: http://git.kernel.org/tip/231678b768da07d19ab5683a39eeb0c250631d02
Author: Dietmar Eggemann <dietmar.eggemann@arm.com>
AuthorDate: Fri, 14 Aug 2015 17:23:13 +0100
Committer: Ingo Molnar <mingo@kernel.org>
CommitDate: Sun, 13 Sep 2015 09:52:57 +0200
sched/fair: Get rid of scaling utilization by capacity_orig
Utilization is currently scaled by capacity_orig, but since we now have
frequency and cpu invariant cfs_rq.avg.util_avg, frequency and cpu scaling
now happens as part of the utilization tracking itself.
So cfs_rq.avg.util_avg should no longer be scaled in cpu_util().
Signed-off-by: Dietmar Eggemann <dietmar.eggemann@arm.com>
Signed-off-by: Morten Rasmussen <morten.rasmussen@arm.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: Juri Lelli <Juri.Lelli@arm.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Mike Galbraith <efault@gmx.de>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Steve Muckle <steve.muckle@linaro.org>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: daniel.lezcano@linaro.org <daniel.lezcano@linaro.org>
Cc: mturquette@baylibre.com <mturquette@baylibre.com>
Cc: pang.xunlei@zte.com.cn <pang.xunlei@zte.com.cn>
Cc: rjw@rjwysocki.net <rjw@rjwysocki.net>
Cc: sgurrappadi@nvidia.com <sgurrappadi@nvidia.com>
Cc: vincent.guittot@linaro.org <vincent.guittot@linaro.org>
Cc: yuyang.du@intel.com <yuyang.du@intel.com>
Link: http://lkml.kernel.org/r/55EDAF43.30500@arm.com
Signed-off-by: Ingo Molnar <mingo@kernel.org>
---
kernel/sched/fair.c | 38 ++++++++++++++++++++++----------------
1 file changed, 22 insertions(+), 16 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 1b56d63..047fd1c 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4862,33 +4862,39 @@ next:
done:
return target;
}
+
/*
* cpu_util returns the amount of capacity of a CPU that is used by CFS
* tasks. The unit of the return value must be the one of capacity so we can
* compare the utilization with the capacity of the CPU that is available for
* CFS task (ie cpu_capacity).
- * cfs.avg.util_avg is the sum of running time of runnable tasks on a
- * CPU. It represents the amount of utilization of a CPU in the range
- * [0..SCHED_LOAD_SCALE]. The utilization of a CPU can't be higher than the
- * full capacity of the CPU because it's about the running time on this CPU.
- * Nevertheless, cfs.avg.util_avg can be higher than SCHED_LOAD_SCALE
- * because of unfortunate rounding in util_avg or just
- * after migrating tasks until the average stabilizes with the new running
- * time. So we need to check that the utilization stays into the range
- * [0..cpu_capacity_orig] and cap if necessary.
- * Without capping the utilization, a group could be seen as overloaded (CPU0
- * utilization at 121% + CPU1 utilization at 80%) whereas CPU1 has 20% of
- * available capacity.
+ *
+ * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the
+ * recent utilization of currently non-runnable tasks on a CPU. It represents
+ * the amount of utilization of a CPU in the range [0..capacity_orig] where
+ * capacity_orig is the cpu_capacity available at the highest frequency
+ * (arch_scale_freq_capacity()).
+ * The utilization of a CPU converges towards a sum equal to or less than the
+ * current capacity (capacity_curr <= capacity_orig) of the CPU because it is
+ * the running time on this CPU scaled by capacity_curr.
+ *
+ * Nevertheless, cfs_rq.avg.util_avg can be higher than capacity_curr or even
+ * higher than capacity_orig because of unfortunate rounding in
+ * cfs.avg.util_avg or just after migrating tasks and new task wakeups until
+ * the average stabilizes with the new running time. We need to check that the
+ * utilization stays within the range of [0..capacity_orig] and cap it if
+ * necessary. Without utilization capping, a group could be seen as overloaded
+ * (CPU0 utilization at 121% + CPU1 utilization at 80%) whereas CPU1 has 20% of
+ * available capacity. We allow utilization to overshoot capacity_curr (but not
+ * capacity_orig) as it useful for predicting the capacity required after task
+ * migrations (scheduler-driven DVFS).
*/
static int cpu_util(int cpu)
{
unsigned long util = cpu_rq(cpu)->cfs.avg.util_avg;
unsigned long capacity = capacity_orig_of(cpu);
- if (util >= SCHED_LOAD_SCALE)
- return capacity;
-
- return (util * capacity) >> SCHED_LOAD_SHIFT;
+ return (util >= capacity) ? capacity : util;
}
/*
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web