Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1220267 > unrolled thread
| Started by | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| First post | 2015-09-07 17:40 +0200 |
| Last post | 2015-09-13 13:10 +0200 |
| Articles | 20 on this page of 58 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-07 17:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-07 18:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-07 21:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-07 21:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 14:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 09:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 14:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 16:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 16:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 17:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-10 00:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-10 13:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-10 13:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-10 14:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-11 10:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-10 19:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-08 18:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-09 11:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-09 11:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-09 13:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 19:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-17 12:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-17 12:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-21 11:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-21 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-22 09:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Leo Yan <leo.yan@linaro.org> - 2015-09-11 09:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 12:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Leo Yan <leo.yan@linaro.org> - 2015-09-11 16:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-10 05:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-10 12:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 15:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 16:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 17:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Vincent Guittot <vincent.guittot@linaro.org> - 2015-09-08 16:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-08 16:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-10 06:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-10 12:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-11 10:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-11 12:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-11 19:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-12 04:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-14 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-14 15:00 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-14 19:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-15 08:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-15 19:20 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-16 04:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig bsegall@google.com - 2015-09-16 19:10 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Yuyang Du <yuyang.du@intel.com> - 2015-09-17 12:30 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Morten Rasmussen <morten.rasmussen@arm.com> - 2015-09-15 10:40 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-16 17:50 +0200
Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig Peter Zijlstra <peterz@infradead.org> - 2015-09-08 13:50 +0200
[tip:sched/core] sched/fair: Get rid of scaling utilization by capacity_orig tip-bot for Dietmar Eggemann <tipbot@zytor.com> - 2015-09-13 13:10 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-09 11:50 +0200 |
| Message-ID | <q6KaS-7Lb-11@gated-at.bofh.it> |
| In reply to | #1220985 |
On Wed, Sep 09, 2015 at 11:43:05AM +0200, Peter Zijlstra wrote:
> Sadly that makes the code worse; I get 14 mul instructions where
> previously I had 11.
FWIW I count like:
objdump -d defconfig-build/kernel/sched/fair.o |
awk '/<[^>]*>:/ { p=0 }
/<update_blocked_averages>:/ { p=1 }
{ if (p) print $0 }' |
cut -d\: -f2- | grep mul | wc -l
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-09 11:50 +0200 |
| Message-ID | <q6KaS-7Lb-13@gated-at.bofh.it> |
| In reply to | #1220985 |
On Tue, Sep 08, 2015 at 05:53:31PM +0100, Morten Rasmussen wrote:
> On Tue, Sep 08, 2015 at 03:31:58PM +0100, Morten Rasmussen wrote:
> > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote:
> > But if we apply the scaling to the weight instead of time, we would only
> > have to apply it once and not three times like it is now? So maybe we
> > can end up with almost the same number of multiplications.
> >
> > We might be loosing bits for low priority task running on cpus at a low
> > frequency though.
>
> Something like the below. We should be saving one multiplication.
> @@ -2577,8 +2575,13 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> return 0;
> sa->last_update_time = now;
>
> - scale_freq = arch_scale_freq_capacity(NULL, cpu);
> - scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
> + if (weight || running)
> + scale_freq = arch_scale_freq_capacity(NULL, cpu);
> + if (weight)
> + scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
> + if (running)
> + scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
> + >> SCHED_CAPACITY_SHIFT;
>
> /* delta_w is the amount already accumulated against our next period */
> delta_w = sa->period_contrib;
> @@ -2594,16 +2597,15 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> * period and accrue it.
> */
> delta_w = 1024 - delta_w;
> - scaled_delta_w = cap_scale(delta_w, scale_freq);
> if (weight) {
> - sa->load_sum += weight * scaled_delta_w;
> + sa->load_sum += scaled_weight * delta_w;
> if (cfs_rq) {
> cfs_rq->runnable_load_sum +=
> - weight * scaled_delta_w;
> + scaled_weight * delta_w;
> }
> }
> if (running)
> - sa->util_sum += scaled_delta_w * scale_cpu;
> + sa->util_sum += delta_w * scale_freq_cpu;
>
> delta -= delta_w;
>
Sadly that makes the code worse; I get 14 mul instructions where
previously I had 11.
What happens is that GCC gets confused and cannot constant propagate the
new variables, so what used to be shifts now end up being actual
multiplications.
With this, I get back to 11. Can you see what happens on ARM where you
have both functions defined to non constants?
---
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2551,10 +2551,10 @@ static __always_inline int
__update_load_avg(u64 now, int cpu, struct sched_avg *sa,
unsigned long weight, int running, struct cfs_rq *cfs_rq)
{
+ unsigned long scaled_weight, scale_freq, scale_freq_cpu;
+ unsigned int delta_w, decayed = 0;
u64 delta, periods;
u32 contrib;
- unsigned int delta_w, decayed = 0;
- unsigned long scaled_weight = 0, scale_freq, scale_freq_cpu = 0;
delta = now - sa->last_update_time;
/*
@@ -2575,13 +2575,10 @@ __update_load_avg(u64 now, int cpu, stru
return 0;
sa->last_update_time = now;
- if (weight || running)
- scale_freq = arch_scale_freq_capacity(NULL, cpu);
- if (weight)
- scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
- if (running)
- scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
- >> SCHED_CAPACITY_SHIFT;
+ scale_freq = arch_scale_freq_capacity(NULL, cpu);
+
+ scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
+ scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu) >> SCHED_CAPACITY_SHIFT;
/* delta_w is the amount already accumulated against our next period */
delta_w = sa->period_contrib;
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-09 13:10 +0200 |
| Message-ID | <q6Lqh-1gs-7@gated-at.bofh.it> |
| In reply to | #1221348 |
On Wed, Sep 09, 2015 at 11:43:05AM +0200, Peter Zijlstra wrote:
> On Tue, Sep 08, 2015 at 05:53:31PM +0100, Morten Rasmussen wrote:
> > On Tue, Sep 08, 2015 at 03:31:58PM +0100, Morten Rasmussen wrote:
>
> > > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote:
> > > But if we apply the scaling to the weight instead of time, we would only
> > > have to apply it once and not three times like it is now? So maybe we
> > > can end up with almost the same number of multiplications.
> > >
> > > We might be loosing bits for low priority task running on cpus at a low
> > > frequency though.
> >
> > Something like the below. We should be saving one multiplication.
>
> > @@ -2577,8 +2575,13 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > return 0;
> > sa->last_update_time = now;
> >
> > - scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > - scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
> > + if (weight || running)
> > + scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > + if (weight)
> > + scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
> > + if (running)
> > + scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
> > + >> SCHED_CAPACITY_SHIFT;
> >
> > /* delta_w is the amount already accumulated against our next period */
> > delta_w = sa->period_contrib;
> > @@ -2594,16 +2597,15 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > * period and accrue it.
> > */
> > delta_w = 1024 - delta_w;
> > - scaled_delta_w = cap_scale(delta_w, scale_freq);
> > if (weight) {
> > - sa->load_sum += weight * scaled_delta_w;
> > + sa->load_sum += scaled_weight * delta_w;
> > if (cfs_rq) {
> > cfs_rq->runnable_load_sum +=
> > - weight * scaled_delta_w;
> > + scaled_weight * delta_w;
> > }
> > }
> > if (running)
> > - sa->util_sum += scaled_delta_w * scale_cpu;
> > + sa->util_sum += delta_w * scale_freq_cpu;
> >
> > delta -= delta_w;
> >
>
> Sadly that makes the code worse; I get 14 mul instructions where
> previously I had 11.
>
> What happens is that GCC gets confused and cannot constant propagate the
> new variables, so what used to be shifts now end up being actual
> multiplications.
>
> With this, I get back to 11. Can you see what happens on ARM where you
> have both functions defined to non constants?
We repeated the experiment on arm and arm64 but still with functions
defined to constant to compare with your results. The mul instruction
count seems to be somewhat compiler version dependent, but consistently
show no effect of the patch:
arm before after
gcc4.9 12 12
gcc4.8 10 10
arm64 before after
gcc4.9 11 11
I will get numbers with the arch-functions implemented as well and do
hackbench runs to see what happens in terms of performance.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-11 19:20 +0200 |
| Message-ID | <q7A9s-89P-3@gated-at.bofh.it> |
| In reply to | #1221367 |
On Wed, Sep 09, 2015 at 12:13:10PM +0100, Morten Rasmussen wrote: > On Wed, Sep 09, 2015 at 11:43:05AM +0200, Peter Zijlstra wrote: > > Sadly that makes the code worse; I get 14 mul instructions where > > previously I had 11. > > > > What happens is that GCC gets confused and cannot constant propagate the > > new variables, so what used to be shifts now end up being actual > > multiplications. > > > > With this, I get back to 11. Can you see what happens on ARM where you > > have both functions defined to non constants? > > We repeated the experiment on arm and arm64 but still with functions > defined to constant to compare with your results. The mul instruction > count seems to be somewhat compiler version dependent, but consistently > show no effect of the patch: > > arm before after > gcc4.9 12 12 > gcc4.8 10 10 > > arm64 before after > gcc4.9 11 11 > > I will get numbers with the arch-functions implemented as well and do > hackbench runs to see what happens in terms of performance. I have done some runs with the proposed fixes added: 1. PeterZ's util_sum shift fix (change util_sum). 2. Morten's scaling of weight instead of time (reduce bit loss). 3. PeterZ's unconditional calls to arch*() functions (compiler opt). To be clear: 2 includes 1, and 3 includes 1 and 2. Runs where done with the default (#define) implementation of the arch-functions and with arch specific implementation for ARM. I realized that just looking for 'mul' instructions in update_blocked_averages() is probably not a fair comparison on ARM as it turned out that it has quite a few multiply-accumulate instructions. So I have included the total count including those too. Test platforms: ARM TC2 (A7x3 only) perf stat --null --repeat 10 -- perf bench sched messaging -g 50 -l 200 #mul: grep -e mul (in update_blocked_averages()) #mul_all: grep -e mul -e mla -e mls -e mia (in update_blocked_averages()) gcc: 4.9.3 Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz perf stat --null --repeat 10 -- perf bench sched messaging -g 50 -l 15000 #mul: grep -e mul (in update_blocked_averages()) gcc: 4.9.2 Results: perf numbers are average of three (x10) runs. Raw data is available further down. ARM TC2 #mul #mul_all perf bench arch*() default arm default arm default arm 1 shift_fix 10 16 22 36 13.401 13.288 2 scaled_weight 12 14 30 32 13.282 13.238 3 unconditional 12 14 26 32 13.296 13.427 Intel E5-2690 #mul #mul_all perf bench arch*() default default default 1 shift_fix 13 14.786 2 scaled_weight 18 15.078 3 unconditional 14 15.195 Overall it appears that fewer 'mul' instructions doesn't necessarily mean better perf bench score. For ARM, 2 seems the best choice overall. While 1 is better for Intel. If we want to try avoid the bit loss by scaling weight instead of time, 2 is best for both. However, all that said, looking at the raw numbers there is a significant difference between runs of perf --repeat, so we can't really draw any strong conclusions. It all appears to be in the noise. I suggest that I spin a v2 of this series and go with scaled_weight to reduce bit loss. Any objections? While at it, should I include Yuyang's patch redefining the SCALE/SHIFT mess? Raw numbers: ARM TC2 shift_fix default_arch gcc4.9.3 #mul 10 #mul+mla+mls+mia 22 13.384416727 seconds time elapsed ( +- 0.17% ) 13.431014702 seconds time elapsed ( +- 0.18% ) 13.387434890 seconds time elapsed ( +- 0.15% ) shift_fix arm_arch gcc4.9.3 #mul 16 #mul+mla+mls+mia 36 13.271044081 seconds time elapsed ( +- 0.11% ) 13.310189123 seconds time elapsed ( +- 0.19% ) 13.283594740 seconds time elapsed ( +- 0.12% ) scaled_weight default_arch gcc4.9.3 #mul 12 #mul+mla+mls+mia 30 13.295649553 seconds time elapsed ( +- 0.20% ) 13.271634654 seconds time elapsed ( +- 0.19% ) 13.280081329 seconds time elapsed ( +- 0.14% ) scaled_weight arm_arch gcc4.9.3 #mul 14 #mul+mla+mls+mia 32 13.230659223 seconds time elapsed ( +- 0.15% ) 13.222276527 seconds time elapsed ( +- 0.15% ) 13.260275081 seconds time elapsed ( +- 0.21% ) unconditional default_arch gcc4.9.3 #mul 12 #mul+mla+mls+mia 26 13.274904460 seconds time elapsed ( +- 0.13% ) 13.307853511 seconds time elapsed ( +- 0.15% ) 13.304084844 seconds time elapsed ( +- 0.22% ) unconditional arm_arch gcc4.9.3 #mul 14 #mul+mla+mls+mia 32 13.432878577 seconds time elapsed ( +- 0.13% ) 13.417950552 seconds time elapsed ( +- 0.12% ) 13.431682719 seconds time elapsed ( +- 0.18% ) Intel shift_fix default_arch gcc4.9.2 #mul 13 14.905815416 seconds time elapsed ( +- 0.61% ) 14.811113694 seconds time elapsed ( +- 0.84% ) 14.639739309 seconds time elapsed ( +- 0.76% ) scaled_weight default_arch gcc4.9.2 #mul 18 15.113275474 seconds time elapsed ( +- 0.64% ) 15.056777680 seconds time elapsed ( +- 0.44% ) 15.064074416 seconds time elapsed ( +- 0.71% ) unconditional default_arch gcc4.9.2 #mul 14 15.105152500 seconds time elapsed ( +- 0.71% ) 15.346405473 seconds time elapsed ( +- 0.81% ) 15.132933523 seconds time elapsed ( +- 0.82% ) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-17 12:00 +0200 |
| Message-ID | <q9E8V-8lc-7@gated-at.bofh.it> |
| In reply to | #1222996 |
On Fri, Sep 11, 2015 at 06:22:47PM +0100, Morten Rasmussen wrote:
> I have done some runs with the proposed fixes added:
>
> 1. PeterZ's util_sum shift fix (change util_sum).
> 2. Morten's scaling of weight instead of time (reduce bit loss).
> 3. PeterZ's unconditional calls to arch*() functions (compiler opt).
>
> To be clear: 2 includes 1, and 3 includes 1 and 2.
>
> Runs where done with the default (#define) implementation of the
> arch-functions and with arch specific implementation for ARM.
> Results:
>
> perf numbers are average of three (x10) runs. Raw data is available
> further down.
>
> ARM TC2 #mul #mul_all perf bench
> arch*() default arm default arm default arm
>
> 1 shift_fix 10 16 22 36 13.401 13.288
> 2 scaled_weight 12 14 30 32 13.282 13.238
> 3 unconditional 12 14 26 32 13.296 13.427
>
> Intel E5-2690 #mul #mul_all perf bench
> arch*() default default default
>
> 1 shift_fix 13 14.786
> 2 scaled_weight 18 15.078
> 3 unconditional 14 15.195
>
>
> Overall it appears that fewer 'mul' instructions doesn't necessarily
> mean better perf bench score. For ARM, 2 seems the best choice overall.
I suspect you're paying for having to do an actual load which can miss
there. So that makes sense.
> While 1 is better for Intel.
Right, because GCC shits itself with those conditionals. Weirdly though;
the below version does not seem so affected.
> I suggest that I spin a v2 of this series and go with scaled_weight to
> reduce bit loss. Any objections?
Just playing devils advocate to myself; how about cgroups? Will not a
per-cpu share of the cgroup weight often be very small?
So I had a little play, and I'm not at all convinced we want to do this
(I've not actually ran any numbers on it, but I can well imagine the
extra condition to hurt on branch miss predict) but it does show GCC
need not always get confused.
---
kernel/sched/fair.c | 58 +++++++++++++++++++++++++++++++++++------------------
1 file changed, 38 insertions(+), 20 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 9176f7c588a8..1b60fbe3b86c 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2519,7 +2519,25 @@ static u32 __compute_runnable_contrib(u64 n)
#error "load tracking assumes 2^10 as unit"
#endif
-#define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
+static __always_inline unsigned long fp_mult2(unsigned long x, unsigned long y)
+{
+ y *= x;
+ y >>= 10;
+
+ return y;
+}
+
+static __always_inline unsigned long fp_mult3(unsigned long x, unsigned long y, unsigned long z)
+{
+ if (x > y)
+ swap(x,y);
+
+ z *= y;
+ z >>= 10;
+ z *= x;
+
+ return z;
+}
/*
* We can represent the historical contribution to runnable average as the
@@ -2553,9 +2571,9 @@ static __always_inline int
__update_load_avg(u64 now, int cpu, struct sched_avg *sa,
unsigned long weight, int running, struct cfs_rq *cfs_rq)
{
- u64 delta, scaled_delta, periods;
+ u64 delta, periods;
u32 contrib;
- unsigned int delta_w, scaled_delta_w, decayed = 0;
+ unsigned int delta_w, decayed = 0;
unsigned long scale_freq, scale_cpu;
delta = now - sa->last_update_time;
@@ -2577,8 +2595,10 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
return 0;
sa->last_update_time = now;
- scale_freq = arch_scale_freq_capacity(NULL, cpu);
- scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
+ if (weight)
+ scale_freq = arch_scale_freq_capacity(NULL, cpu);
+ if (running)
+ scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
/* delta_w is the amount already accumulated against our next period */
delta_w = sa->period_contrib;
@@ -2594,16 +2614,14 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
* period and accrue it.
*/
delta_w = 1024 - delta_w;
- scaled_delta_w = cap_scale(delta_w, scale_freq);
if (weight) {
- sa->load_sum += weight * scaled_delta_w;
- if (cfs_rq) {
- cfs_rq->runnable_load_sum +=
- weight * scaled_delta_w;
- }
+ unsigned long t = fp_mult3(delta_w, weight, scale_freq);
+ sa->load_sum += t;
+ if (cfs_rq)
+ cfs_rq->runnable_load_sum += t;
}
if (running)
- sa->util_sum += scaled_delta_w * scale_cpu;
+ sa->util_sum += delta_w * fp_mult2(scale_cpu, scale_freq);
delta -= delta_w;
@@ -2620,25 +2638,25 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
/* Efficiently calculate \sum (1..n_period) 1024*y^i */
contrib = __compute_runnable_contrib(periods);
- contrib = cap_scale(contrib, scale_freq);
if (weight) {
- sa->load_sum += weight * contrib;
+ unsigned long t = fp_mult3(contrib, weight, scale_freq);
+ sa->load_sum += t;
if (cfs_rq)
- cfs_rq->runnable_load_sum += weight * contrib;
+ cfs_rq->runnable_load_sum += t;
}
if (running)
- sa->util_sum += contrib * scale_cpu;
+ sa->util_sum += contrib * fp_mult2(scale_cpu, scale_freq);
}
/* Remainder of delta accrued against u_0` */
- scaled_delta = cap_scale(delta, scale_freq);
if (weight) {
- sa->load_sum += weight * scaled_delta;
+ unsigned long t = fp_mult3(delta, weight, scale_freq);
+ sa->load_sum += t;
if (cfs_rq)
- cfs_rq->runnable_load_sum += weight * scaled_delta;
+ cfs_rq->runnable_load_sum += t;
}
if (running)
- sa->util_sum += scaled_delta * scale_cpu;
+ sa->util_sum += delta * fp_mult2(scale_cpu, scale_freq);
sa->period_contrib += delta;
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-17 12:50 +0200 |
| Message-ID | <q9EVj-15r-9@gated-at.bofh.it> |
| In reply to | #1222996 |
On Fri, Sep 11, 2015 at 06:22:47PM +0100, Morten Rasmussen wrote: > While at it, should I include Yuyang's patch redefining the SCALE/SHIFT > mess? I suspect his patch will fail to compile on ARM which uses SCHED_CAPACITY_* outside of kernel/sched/*. But if you all (Ben, Yuyang, you) can agree on a patch simplifying these things I'm not opposed to it. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-21 11:10 +0200 |
| Message-ID | <qb5gK-297-7@gated-at.bofh.it> |
| In reply to | #1226905 |
On Thu, Sep 17, 2015 at 12:38:25PM +0200, Peter Zijlstra wrote:
> On Fri, Sep 11, 2015 at 06:22:47PM +0100, Morten Rasmussen wrote:
>
> > While at it, should I include Yuyang's patch redefining the SCALE/SHIFT
> > mess?
>
> I suspect his patch will fail to compile on ARM which uses
> SCHED_CAPACITY_* outside of kernel/sched/*.
>
> But if you all (Ben, Yuyang, you) can agree on a patch simplifying these
> things I'm not opposed to it.
Yes, indeed. So SCHED_RESOLUTION_SHIFT has to be defined in include/linux/sched.h.
With this, I think the codes still need some cleanup, and importantly
documentation.
But first, I think as load_sum and load_avg can afford NICE_0_LOAD with either high
or low resolution. So we have no reason to have low resolution (10bits) load_avg
when NICE_0_LOAD has high resolution (20bits), because load_avg = runnable% * load,
as opposed to now we have load_avg = runnable% * scale_load_down(load).
We get rid of all scale_load_down() for runnable load average?
--
Subject: [PATCH] sched/fair: Generalize the load/util averages resolution
definition
The metric needs certain resolution to determine how much detail we
can look into (or not losing detail by integer rounding), which also
determines the range of the metrics.
For instance, to increase the resolution of [0, 1] (two levels), one
can multiply 1024 and get [0, 1024] (1025 levels).
In sched/fair, a few metrics depend on the resolution: load/load_avg,
util_avg, and capacity (frequency adjustment). In order to reduce the
risks to make mistakes relating to resolution/range, we therefore
generalize the resolution by defining a basic resolution constant
number, and then formalize all metrics by depending on the basic
resolution. The basic resolution is 1024 or (1 << 10). Further, one
can recursively apply the basic resolution to increase the final
resolution.
Pointed out by Ben Segall, NICE_0's weight (visible to user) and load
have independent resolution, but they must be well calibrated.
Signed-off-by: Yuyang Du <yuyang.du@intel.com>
---
include/linux/sched.h | 9 ++++++---
kernel/sched/fair.c | 4 ----
kernel/sched/sched.h | 15 ++++++++++-----
3 files changed, 16 insertions(+), 12 deletions(-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index bd38b3e..9b86f79 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -909,10 +909,13 @@ enum cpu_idle_type {
CPU_MAX_IDLE_TYPES
};
+# define SCHED_RESOLUTION_SHIFT 10
+# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
+
/*
* Increase resolution of cpu_capacity calculations
*/
-#define SCHED_CAPACITY_SHIFT 10
+#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
#define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
/*
@@ -1180,8 +1183,8 @@ struct load_weight {
* 1) load_avg factors frequency scaling into the amount of time that a
* sched_entity is runnable on a rq into its weight. For cfs_rq, it is the
* aggregated such weights of all runnable and blocked sched_entities.
- * 2) util_avg factors frequency and cpu scaling into the amount of time
- * that a sched_entity is running on a CPU, in the range [0..SCHED_LOAD_SCALE].
+ * 2) util_avg factors frequency and cpu capacity scaling into the amount of time
+ * that a sched_entity is running on a CPU, in the range [0..SCHED_CAPACITY_SCALE].
* For cfs_rq, it is the aggregated such times of all runnable and
* blocked sched_entities.
* The 64 bit load_sum can:
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 4df37a4..c61fd8e 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2522,10 +2522,6 @@ static u32 __compute_runnable_contrib(u64 n)
return contrib + runnable_avg_yN_sum[n];
}
-#if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != 10 || SCHED_CAPACITY_SHIFT != 10
-#error "load tracking assumes 2^10 as unit"
-#endif
-
#define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
/*
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 3845a71..31b4022 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -53,18 +53,23 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
* increased costs.
*/
#if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
-# define SCHED_LOAD_RESOLUTION 10
-# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
-# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
+# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
+# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
+# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
#else
-# define SCHED_LOAD_RESOLUTION 0
+# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
# define scale_load(w) (w)
# define scale_load_down(w) (w)
#endif
-#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
#define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
+/*
+ * NICE_0's weight (visible to user) and its load (invisible to user) have
+ * independent resolution, but they should be well calibrated. We use scale_load()
+ * and scale_load_down(w) to convert between them, the following must be true:
+ * scale_load(prio_to_weight[20]) == NICE_0_LOAD
+ */
#define NICE_0_LOAD SCHED_LOAD_SCALE
#define NICE_0_SHIFT SCHED_LOAD_SHIFT
--
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | bsegall@google.com |
|---|---|
| Date | 2015-09-21 19:40 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <qbdei-51N-17@gated-at.bofh.it> |
| In reply to | #1229130 |
Yuyang Du <yuyang.du@intel.com> writes:
> On Thu, Sep 17, 2015 at 12:38:25PM +0200, Peter Zijlstra wrote:
>> On Fri, Sep 11, 2015 at 06:22:47PM +0100, Morten Rasmussen wrote:
>>
>> > While at it, should I include Yuyang's patch redefining the SCALE/SHIFT
>> > mess?
>>
>> I suspect his patch will fail to compile on ARM which uses
>> SCHED_CAPACITY_* outside of kernel/sched/*.
>>
>> But if you all (Ben, Yuyang, you) can agree on a patch simplifying these
>> things I'm not opposed to it.
>
> Yes, indeed. So SCHED_RESOLUTION_SHIFT has to be defined in include/linux/sched.h.
>
> With this, I think the codes still need some cleanup, and importantly
> documentation.
>
> But first, I think as load_sum and load_avg can afford NICE_0_LOAD with either high
> or low resolution. So we have no reason to have low resolution (10bits) load_avg
> when NICE_0_LOAD has high resolution (20bits), because load_avg = runnable% * load,
> as opposed to now we have load_avg = runnable% * scale_load_down(load).
>
> We get rid of all scale_load_down() for runnable load average?
Hmm, LOAD_AVG_MAX * prio_to_weight[0] is 4237627662, ie barely within a
32-bit unsigned long, but in fact LOAD_AVG_MAX * MAX_SHARES is already
going to give errors on 32-bit (even with the old code in fact). This
should probably be fixed... somehow (dividing by 4 for load_sum on
32-bit would work, though be ugly. Reducing MAX_SHARES by 2 bits on
32-bit might have made sense but would be a weird difference between 32
and 64, and could break userspace anyway, so it's presumably too late
for that).
64-bit has ~30 bits free, so this would be fine so long as SLR is 0 on
32-bit.
>
> --
>
> Subject: [PATCH] sched/fair: Generalize the load/util averages resolution
> definition
>
> The metric needs certain resolution to determine how much detail we
> can look into (or not losing detail by integer rounding), which also
> determines the range of the metrics.
>
> For instance, to increase the resolution of [0, 1] (two levels), one
> can multiply 1024 and get [0, 1024] (1025 levels).
>
> In sched/fair, a few metrics depend on the resolution: load/load_avg,
> util_avg, and capacity (frequency adjustment). In order to reduce the
> risks to make mistakes relating to resolution/range, we therefore
> generalize the resolution by defining a basic resolution constant
> number, and then formalize all metrics by depending on the basic
> resolution. The basic resolution is 1024 or (1 << 10). Further, one
> can recursively apply the basic resolution to increase the final
> resolution.
>
> Pointed out by Ben Segall, NICE_0's weight (visible to user) and load
> have independent resolution, but they must be well calibrated.
>
> Signed-off-by: Yuyang Du <yuyang.du@intel.com>
> ---
> include/linux/sched.h | 9 ++++++---
> kernel/sched/fair.c | 4 ----
> kernel/sched/sched.h | 15 ++++++++++-----
> 3 files changed, 16 insertions(+), 12 deletions(-)
>
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index bd38b3e..9b86f79 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -909,10 +909,13 @@ enum cpu_idle_type {
> CPU_MAX_IDLE_TYPES
> };
>
> +# define SCHED_RESOLUTION_SHIFT 10
> +# define SCHED_RESOLUTION_SCALE (1L << SCHED_RESOLUTION_SHIFT)
> +
> /*
> * Increase resolution of cpu_capacity calculations
> */
> -#define SCHED_CAPACITY_SHIFT 10
> +#define SCHED_CAPACITY_SHIFT SCHED_RESOLUTION_SHIFT
> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
>
> /*
> @@ -1180,8 +1183,8 @@ struct load_weight {
> * 1) load_avg factors frequency scaling into the amount of time that a
> * sched_entity is runnable on a rq into its weight. For cfs_rq, it is the
> * aggregated such weights of all runnable and blocked sched_entities.
> - * 2) util_avg factors frequency and cpu scaling into the amount of time
> - * that a sched_entity is running on a CPU, in the range [0..SCHED_LOAD_SCALE].
> + * 2) util_avg factors frequency and cpu capacity scaling into the amount of time
> + * that a sched_entity is running on a CPU, in the range [0..SCHED_CAPACITY_SCALE].
> * For cfs_rq, it is the aggregated such times of all runnable and
> * blocked sched_entities.
> * The 64 bit load_sum can:
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 4df37a4..c61fd8e 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -2522,10 +2522,6 @@ static u32 __compute_runnable_contrib(u64 n)
> return contrib + runnable_avg_yN_sum[n];
> }
>
> -#if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != 10 || SCHED_CAPACITY_SHIFT != 10
> -#error "load tracking assumes 2^10 as unit"
> -#endif
> -
> #define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
>
> /*
> diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
> index 3845a71..31b4022 100644
> --- a/kernel/sched/sched.h
> +++ b/kernel/sched/sched.h
> @@ -53,18 +53,23 @@ static inline void update_cpu_load_active(struct rq *this_rq) { }
> * increased costs.
> */
> #if 0 /* BITS_PER_LONG > 32 -- currently broken: it increases power usage under light load */
> -# define SCHED_LOAD_RESOLUTION 10
> -# define scale_load(w) ((w) << SCHED_LOAD_RESOLUTION)
> -# define scale_load_down(w) ((w) >> SCHED_LOAD_RESOLUTION)
> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT + SCHED_RESOLUTION_SHIFT)
> +# define scale_load(w) ((w) << SCHED_RESOLUTION_SHIFT)
> +# define scale_load_down(w) ((w) >> SCHED_RESOLUTION_SHIFT)
> #else
> -# define SCHED_LOAD_RESOLUTION 0
> +# define SCHED_LOAD_SHIFT (SCHED_RESOLUTION_SHIFT)
> # define scale_load(w) (w)
> # define scale_load_down(w) (w)
> #endif
>
> -#define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION)
> #define SCHED_LOAD_SCALE (1L << SCHED_LOAD_SHIFT)
>
> +/*
> + * NICE_0's weight (visible to user) and its load (invisible to user) have
> + * independent resolution, but they should be well calibrated. We use scale_load()
> + * and scale_load_down(w) to convert between them, the following must be true:
> + * scale_load(prio_to_weight[20]) == NICE_0_LOAD
> + */
> #define NICE_0_LOAD SCHED_LOAD_SCALE
> #define NICE_0_SHIFT SCHED_LOAD_SHIFT
I still think tying the scale_load shift to be the same as the
SCHED_CAPACITY/etc shift is silly, and tying the NICE_0_LOAD/SHIFT in is
worse. Honestly if I was going to change anything it would be to define
NICE_0_LOAD/SHIFT entirely separately from SCHED_LOAD_SCALE/SHIFT.
However I'm not sure if calculate_imbalance's use of SCHED_LOAD_SCALE is
actually a separate use of 1024*SLR-as-percentage or is basically
assuming most tasks are nice-0 or what. It sure /looks/ like it's
comparing values with different units - it's doing (nr_running * CONST -
group_capacity) and comparing to load, so it looks like both (ie
increasing load.weight of everything on your system by X% would change
load balancer behavior here).
Given that it might make sense to make it clear that capacity units and
nice-0-task units have to be the same thing due to load balancer
approximations (though they are still entirely separate from the
SCHED_LOAD_RESOLUTION multiplier).
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-22 09:30 +0200 |
| Message-ID | <qbqbw-6ZU-9@gated-at.bofh.it> |
| In reply to | #1229597 |
On Mon, Sep 21, 2015 at 10:30:04AM -0700, bsegall@google.com wrote: > > But first, I think as load_sum and load_avg can afford NICE_0_LOAD with either high > > or low resolution. So we have no reason to have low resolution (10bits) load_avg > > when NICE_0_LOAD has high resolution (20bits), because load_avg = runnable% * load, > > as opposed to now we have load_avg = runnable% * scale_load_down(load). > > > > We get rid of all scale_load_down() for runnable load average? > > Hmm, LOAD_AVG_MAX * prio_to_weight[0] is 4237627662, ie barely within a > 32-bit unsigned long, but in fact LOAD_AVG_MAX * MAX_SHARES is already > going to give errors on 32-bit (even with the old code in fact). This > should probably be fixed... somehow (dividing by 4 for load_sum on > 32-bit would work, though be ugly. Reducing MAX_SHARES by 2 bits on > 32-bit might have made sense but would be a weird difference between 32 > and 64, and could break userspace anyway, so it's presumably too late > for that). > > 64-bit has ~30 bits free, so this would be fine so long as SLR is 0 on > 32-bit. > load_avg has no LOAD_AVG_MAX term in it, it is runnable% * load, IOW, load_avg <= load. So, on 32bit, cfs_rq's load_avg can host 2^32/prio_to_weight[0]/1024 = 47, with 20bits load resolution. This is ok, because struct load_weight's load is also unsigned long. If overflown, cfs_rq->load.weight will be overflown in the first place. However, after a second thought, this is not quite right. Because load_avg is not necessarily no greater than load, since load_avg has blocked load in it. Although, load_avg is still at the same level as load (converging to be <= load), we may not want the risk to overflow on 32bit. > > +/* > > + * NICE_0's weight (visible to user) and its load (invisible to user) have > > + * independent resolution, but they should be well calibrated. We use scale_load() > > + * and scale_load_down(w) to convert between them, the following must be true: > > + * scale_load(prio_to_weight[20]) == NICE_0_LOAD > > + */ > > #define NICE_0_LOAD SCHED_LOAD_SCALE > > #define NICE_0_SHIFT SCHED_LOAD_SHIFT > > I still think tying the scale_load shift to be the same as the > SCHED_CAPACITY/etc shift is silly, and tying the NICE_0_LOAD/SHIFT in is > worse. Honestly if I was going to change anything it would be to define > NICE_0_LOAD/SHIFT entirely separately from SCHED_LOAD_SCALE/SHIFT. If NICE_0_LOAD is nice-0's load, and if SCHED_LOAD_SHIFT is to say how to get nice-0's load, I don't understand why you want to separate them. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Leo Yan <leo.yan@linaro.org> |
|---|---|
| Date | 2015-09-11 09:50 +0200 |
| Message-ID | <q7rfQ-3xd-5@gated-at.bofh.it> |
| In reply to | #1220985 |
Hi Morten,
On Tue, Sep 08, 2015 at 05:53:31PM +0100, Morten Rasmussen wrote:
> On Tue, Sep 08, 2015 at 03:31:58PM +0100, Morten Rasmussen wrote:
> > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote:
> > >
> > > Something like teh below..
> > >
> > > Another thing to ponder; the downside of scaled_delta_w is that its
> > > fairly likely delta is small and you loose all bits, whereas the weight
> > > is likely to be large can could loose a fwe bits without issue.
> >
> > That issue applies both to load and util.
> >
> > >
> > > That is, in fixed point scaling like this, you want to start with the
> > > biggest numbers, not the smallest, otherwise you loose too much.
> > >
> > > The flip side is of course that now you can share a multiplcation.
> >
> > But if we apply the scaling to the weight instead of time, we would only
> > have to apply it once and not three times like it is now? So maybe we
> > can end up with almost the same number of multiplications.
> >
> > We might be loosing bits for low priority task running on cpus at a low
> > frequency though.
>
> Something like the below. We should be saving one multiplication.
>
> --- 8< ---
>
> From: Morten Rasmussen <morten.rasmussen@arm.com>
> Date: Tue, 8 Sep 2015 17:15:40 +0100
> Subject: [PATCH] sched/fair: Scale load/util contribution rather than time
>
> When updating load/util tracking the time delta might be very small (1)
> in many cases, scaling it futher down with frequency and cpu invariance
> scaling might cause us to loose precision. Instead of scaling time we
> can scale the weight of the task for load and the capacity for
> utilization. Both weight (>=15) and capacity should be significantly
> bigger in most cases. Low priority tasks might still suffer a bit but
> worst should be improved, as weight is at least 15 before invariance
> scaling.
>
> Signed-off-by: Morten Rasmussen <morten.rasmussen@arm.com>
> ---
> kernel/sched/fair.c | 38 +++++++++++++++++++-------------------
> 1 file changed, 19 insertions(+), 19 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 9301291..d5ee72a 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -2519,8 +2519,6 @@ static u32 __compute_runnable_contrib(u64 n)
> #error "load tracking assumes 2^10 as unit"
> #endif
>
> -#define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
> -
> /*
> * We can represent the historical contribution to runnable average as the
> * coefficients of a geometric series. To do this we sub-divide our runnable
> @@ -2553,10 +2551,10 @@ static __always_inline int
> __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> unsigned long weight, int running, struct cfs_rq *cfs_rq)
> {
> - u64 delta, scaled_delta, periods;
> + u64 delta, periods;
> u32 contrib;
> - unsigned int delta_w, scaled_delta_w, decayed = 0;
> - unsigned long scale_freq, scale_cpu;
> + unsigned int delta_w, decayed = 0;
> + unsigned long scaled_weight = 0, scale_freq, scale_freq_cpu = 0;
>
> delta = now - sa->last_update_time;
> /*
> @@ -2577,8 +2575,13 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> return 0;
> sa->last_update_time = now;
>
> - scale_freq = arch_scale_freq_capacity(NULL, cpu);
> - scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
> + if (weight || running)
> + scale_freq = arch_scale_freq_capacity(NULL, cpu);
> + if (weight)
> + scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
> + if (running)
> + scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
> + >> SCHED_CAPACITY_SHIFT;
maybe below question is stupid :)
Why not calculate the scaled_weight depend on cpu's capacity as well?
So like: scaled_weight = weight * scale_freq_cpu.
> /* delta_w is the amount already accumulated against our next period */
> delta_w = sa->period_contrib;
> @@ -2594,16 +2597,15 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> * period and accrue it.
> */
> delta_w = 1024 - delta_w;
> - scaled_delta_w = cap_scale(delta_w, scale_freq);
> if (weight) {
> - sa->load_sum += weight * scaled_delta_w;
> + sa->load_sum += scaled_weight * delta_w;
> if (cfs_rq) {
> cfs_rq->runnable_load_sum +=
> - weight * scaled_delta_w;
> + scaled_weight * delta_w;
> }
> }
> if (running)
> - sa->util_sum += scaled_delta_w * scale_cpu;
> + sa->util_sum += delta_w * scale_freq_cpu;
>
> delta -= delta_w;
>
> @@ -2620,25 +2622,23 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
>
> /* Efficiently calculate \sum (1..n_period) 1024*y^i */
> contrib = __compute_runnable_contrib(periods);
> - contrib = cap_scale(contrib, scale_freq);
> if (weight) {
> - sa->load_sum += weight * contrib;
> + sa->load_sum += scaled_weight * contrib;
> if (cfs_rq)
> - cfs_rq->runnable_load_sum += weight * contrib;
> + cfs_rq->runnable_load_sum += scaled_weight * contrib;
> }
> if (running)
> - sa->util_sum += contrib * scale_cpu;
> + sa->util_sum += contrib * scale_freq_cpu;
> }
>
> /* Remainder of delta accrued against u_0` */
> - scaled_delta = cap_scale(delta, scale_freq);
> if (weight) {
> - sa->load_sum += weight * scaled_delta;
> + sa->load_sum += scaled_weight * delta;
> if (cfs_rq)
> - cfs_rq->runnable_load_sum += weight * scaled_delta;
> + cfs_rq->runnable_load_sum += scaled_weight * delta;
> }
> if (running)
> - sa->util_sum += scaled_delta * scale_cpu;
> + sa->util_sum += delta * scale_freq_cpu;
>
> sa->period_contrib += delta;
>
> --
> 1.9.1
>
> --
> To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
> Please read the FAQ at http://www.tux.org/lkml/
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Morten Rasmussen <morten.rasmussen@arm.com> |
|---|---|
| Date | 2015-09-11 12:00 +0200 |
| Message-ID | <q7thE-6pR-25@gated-at.bofh.it> |
| In reply to | #1222592 |
On Fri, Sep 11, 2015 at 03:46:51PM +0800, Leo Yan wrote:
> Hi Morten,
>
> On Tue, Sep 08, 2015 at 05:53:31PM +0100, Morten Rasmussen wrote:
> > On Tue, Sep 08, 2015 at 03:31:58PM +0100, Morten Rasmussen wrote:
> > > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote:
> > > >
> > > > Something like teh below..
> > > >
> > > > Another thing to ponder; the downside of scaled_delta_w is that its
> > > > fairly likely delta is small and you loose all bits, whereas the weight
> > > > is likely to be large can could loose a fwe bits without issue.
> > >
> > > That issue applies both to load and util.
> > >
> > > >
> > > > That is, in fixed point scaling like this, you want to start with the
> > > > biggest numbers, not the smallest, otherwise you loose too much.
> > > >
> > > > The flip side is of course that now you can share a multiplcation.
> > >
> > > But if we apply the scaling to the weight instead of time, we would only
> > > have to apply it once and not three times like it is now? So maybe we
> > > can end up with almost the same number of multiplications.
> > >
> > > We might be loosing bits for low priority task running on cpus at a low
> > > frequency though.
> >
> > Something like the below. We should be saving one multiplication.
> >
> > --- 8< ---
> >
> > From: Morten Rasmussen <morten.rasmussen@arm.com>
> > Date: Tue, 8 Sep 2015 17:15:40 +0100
> > Subject: [PATCH] sched/fair: Scale load/util contribution rather than time
> >
> > When updating load/util tracking the time delta might be very small (1)
> > in many cases, scaling it futher down with frequency and cpu invariance
> > scaling might cause us to loose precision. Instead of scaling time we
> > can scale the weight of the task for load and the capacity for
> > utilization. Both weight (>=15) and capacity should be significantly
> > bigger in most cases. Low priority tasks might still suffer a bit but
> > worst should be improved, as weight is at least 15 before invariance
> > scaling.
> >
> > Signed-off-by: Morten Rasmussen <morten.rasmussen@arm.com>
> > ---
> > kernel/sched/fair.c | 38 +++++++++++++++++++-------------------
> > 1 file changed, 19 insertions(+), 19 deletions(-)
> >
> > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> > index 9301291..d5ee72a 100644
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -2519,8 +2519,6 @@ static u32 __compute_runnable_contrib(u64 n)
> > #error "load tracking assumes 2^10 as unit"
> > #endif
> >
> > -#define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
> > -
> > /*
> > * We can represent the historical contribution to runnable average as the
> > * coefficients of a geometric series. To do this we sub-divide our runnable
> > @@ -2553,10 +2551,10 @@ static __always_inline int
> > __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > unsigned long weight, int running, struct cfs_rq *cfs_rq)
> > {
> > - u64 delta, scaled_delta, periods;
> > + u64 delta, periods;
> > u32 contrib;
> > - unsigned int delta_w, scaled_delta_w, decayed = 0;
> > - unsigned long scale_freq, scale_cpu;
> > + unsigned int delta_w, decayed = 0;
> > + unsigned long scaled_weight = 0, scale_freq, scale_freq_cpu = 0;
> >
> > delta = now - sa->last_update_time;
> > /*
> > @@ -2577,8 +2575,13 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > return 0;
> > sa->last_update_time = now;
> >
> > - scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > - scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
> > + if (weight || running)
> > + scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > + if (weight)
> > + scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
> > + if (running)
> > + scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
> > + >> SCHED_CAPACITY_SHIFT;
>
> maybe below question is stupid :)
>
> Why not calculate the scaled_weight depend on cpu's capacity as well?
> So like: scaled_weight = weight * scale_freq_cpu.
IMHO, we should not scale load by cpu capacity since load isn't really
comparable to capacity. It is runnable time based (not running time like
utilization) and the idea is to used it for balancing when when the
system is fully utilized. When the system is fully utilized we can't say
anything about the true compute demands of a task, it may get exactly
the cpu time it needs or it may need much more. Hence it doesn't really
make sense to scale the demand by the capacity of the cpu. Two busy
loops on cpus with different cpu capacities should have the load as they
have the same compute demands.
I mentioned this briefly in the commit message of patch 3 in this
series.
Makes sense?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Leo Yan <leo.yan@linaro.org> |
|---|---|
| Date | 2015-09-11 16:20 +0200 |
| Message-ID | <q7xlf-42z-5@gated-at.bofh.it> |
| In reply to | #1222671 |
On Fri, Sep 11, 2015 at 11:02:33AM +0100, Morten Rasmussen wrote:
> On Fri, Sep 11, 2015 at 03:46:51PM +0800, Leo Yan wrote:
> > Hi Morten,
> >
> > On Tue, Sep 08, 2015 at 05:53:31PM +0100, Morten Rasmussen wrote:
> > > On Tue, Sep 08, 2015 at 03:31:58PM +0100, Morten Rasmussen wrote:
> > > > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote:
> > > > >
> > > > > Something like teh below..
> > > > >
> > > > > Another thing to ponder; the downside of scaled_delta_w is that its
> > > > > fairly likely delta is small and you loose all bits, whereas the weight
> > > > > is likely to be large can could loose a fwe bits without issue.
> > > >
> > > > That issue applies both to load and util.
> > > >
> > > > >
> > > > > That is, in fixed point scaling like this, you want to start with the
> > > > > biggest numbers, not the smallest, otherwise you loose too much.
> > > > >
> > > > > The flip side is of course that now you can share a multiplcation.
> > > >
> > > > But if we apply the scaling to the weight instead of time, we would only
> > > > have to apply it once and not three times like it is now? So maybe we
> > > > can end up with almost the same number of multiplications.
> > > >
> > > > We might be loosing bits for low priority task running on cpus at a low
> > > > frequency though.
> > >
> > > Something like the below. We should be saving one multiplication.
> > >
> > > --- 8< ---
> > >
> > > From: Morten Rasmussen <morten.rasmussen@arm.com>
> > > Date: Tue, 8 Sep 2015 17:15:40 +0100
> > > Subject: [PATCH] sched/fair: Scale load/util contribution rather than time
> > >
> > > When updating load/util tracking the time delta might be very small (1)
> > > in many cases, scaling it futher down with frequency and cpu invariance
> > > scaling might cause us to loose precision. Instead of scaling time we
> > > can scale the weight of the task for load and the capacity for
> > > utilization. Both weight (>=15) and capacity should be significantly
> > > bigger in most cases. Low priority tasks might still suffer a bit but
> > > worst should be improved, as weight is at least 15 before invariance
> > > scaling.
> > >
> > > Signed-off-by: Morten Rasmussen <morten.rasmussen@arm.com>
> > > ---
> > > kernel/sched/fair.c | 38 +++++++++++++++++++-------------------
> > > 1 file changed, 19 insertions(+), 19 deletions(-)
> > >
> > > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> > > index 9301291..d5ee72a 100644
> > > --- a/kernel/sched/fair.c
> > > +++ b/kernel/sched/fair.c
> > > @@ -2519,8 +2519,6 @@ static u32 __compute_runnable_contrib(u64 n)
> > > #error "load tracking assumes 2^10 as unit"
> > > #endif
> > >
> > > -#define cap_scale(v, s) ((v)*(s) >> SCHED_CAPACITY_SHIFT)
> > > -
> > > /*
> > > * We can represent the historical contribution to runnable average as the
> > > * coefficients of a geometric series. To do this we sub-divide our runnable
> > > @@ -2553,10 +2551,10 @@ static __always_inline int
> > > __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > > unsigned long weight, int running, struct cfs_rq *cfs_rq)
> > > {
> > > - u64 delta, scaled_delta, periods;
> > > + u64 delta, periods;
> > > u32 contrib;
> > > - unsigned int delta_w, scaled_delta_w, decayed = 0;
> > > - unsigned long scale_freq, scale_cpu;
> > > + unsigned int delta_w, decayed = 0;
> > > + unsigned long scaled_weight = 0, scale_freq, scale_freq_cpu = 0;
> > >
> > > delta = now - sa->last_update_time;
> > > /*
> > > @@ -2577,8 +2575,13 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa,
> > > return 0;
> > > sa->last_update_time = now;
> > >
> > > - scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > > - scale_cpu = arch_scale_cpu_capacity(NULL, cpu);
> > > + if (weight || running)
> > > + scale_freq = arch_scale_freq_capacity(NULL, cpu);
> > > + if (weight)
> > > + scaled_weight = weight * scale_freq >> SCHED_CAPACITY_SHIFT;
> > > + if (running)
> > > + scale_freq_cpu = scale_freq * arch_scale_cpu_capacity(NULL, cpu)
> > > + >> SCHED_CAPACITY_SHIFT;
> >
> > maybe below question is stupid :)
> >
> > Why not calculate the scaled_weight depend on cpu's capacity as well?
> > So like: scaled_weight = weight * scale_freq_cpu.
>
> IMHO, we should not scale load by cpu capacity since load isn't really
> comparable to capacity. It is runnable time based (not running time like
> utilization) and the idea is to used it for balancing when when the
> system is fully utilized. When the system is fully utilized we can't say
> anything about the true compute demands of a task, it may get exactly
> the cpu time it needs or it may need much more. Hence it doesn't really
> make sense to scale the demand by the capacity of the cpu. Two busy
> loops on cpus with different cpu capacities should have the load as they
> have the same compute demands.
>
> I mentioned this briefly in the commit message of patch 3 in this
> series.
>
> Makes sense?
Yeah, after your reminding, i recognise load only includes runnable
time on rq but not include running time.
Thanks,
Leo Yan
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-09-10 05:00 +0200 |
| Message-ID | <q70fD-5vL-1@gated-at.bofh.it> |
| In reply to | #1220762 |
On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote: > > +#if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != 10 || SCHED_CAPACITY_SHIFT != 10 > +#error "load tracking assumes 2^10 as unit" > +#endif > + Sorry for late response. I might already missed somthing. But I got a bit lost here, with: #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION) #define SCHED_CAPACITY_SHIFT 10 the #if is certainly false. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-10 12:10 +0200 |
| Message-ID | <q76XM-6I3-11@gated-at.bofh.it> |
| In reply to | #1221862 |
On Thu, Sep 10, 2015 at 03:07:48AM +0800, Yuyang Du wrote: > On Tue, Sep 08, 2015 at 02:52:05PM +0200, Peter Zijlstra wrote: > > > > +#if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != 10 || SCHED_CAPACITY_SHIFT != 10 > > +#error "load tracking assumes 2^10 as unit" > > +#endif > > + > > Sorry for late response. I might already missed somthing. > > But I got a bit lost here, with: > > #define SCHED_LOAD_SHIFT (10 + SCHED_LOAD_RESOLUTION) > #define SCHED_CAPACITY_SHIFT 10 > > the #if is certainly false. That is intended, triggering #error would be 'bad'. The reason for this bit is to raise a stink if someone 'accidentally' changes one of these values and expects things to just work. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vincent Guittot <vincent.guittot@linaro.org> |
|---|---|
| Date | 2015-09-08 15:50 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q6rrA-5QO-7@gated-at.bofh.it> |
| In reply to | #1220741 |
On 8 September 2015 at 14:26, Peter Zijlstra <peterz@infradead.org> wrote: > On Tue, Sep 08, 2015 at 09:22:05AM +0200, Vincent Guittot wrote: >> No, but >> sa->util_avg = (sa->util_sum << SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; >> will fix the unit issue. > > Tricky that, LOAD_AVG_MAX very much relies on the unit being 1<<10. > > And where load_sum already gets a factor 1024 from the weight > multiplication, util_sum does not get such a factor, and all the scaling > we do on it loose bits. fair point > > So at the moment we go compute the util_avg value, we need to inflate > util_sum with an extra factor 1024 in order to make it work. > > And seeing that we do the shift up on sa->util_sum without consideration > of overflow, would it not make sense to add that factor before the > scaling and into the addition? Yes this should save 1 left shift and 1 right shift >> > > Now, given all that, units are a complete mess here, and I'd not mind > something like: > > #if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != SCHED_CAPACITY_SHIFT > #error "something usefull" > #endif In this case why not simply doing #define SCHED_CAPACITY_SHIFT SCHED_LOAD_SHIFT or the opposite ? > > somewhere near here. > > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-08 16:20 +0200 |
| Message-ID | <q6rUC-6Ec-29@gated-at.bofh.it> |
| In reply to | #1220797 |
On Tue, Sep 08, 2015 at 03:39:37PM +0200, Vincent Guittot wrote: > > Now, given all that, units are a complete mess here, and I'd not mind > > something like: > > > > #if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != SCHED_CAPACITY_SHIFT > > #error "something usefull" > > #endif > > In this case why not simply doing > #define SCHED_CAPACITY_SHIFT SCHED_LOAD_SHIFT > or the opposite ? Sadly not enough; aside from the fact that we really should do !0 LOAD_RESOLUTION on 64bit, the whole magic tables (runnable_avg_yN_*[]) and LOAD_AVG_MAX* values rely on the unit being 1<<10. So regardless of defining one in terms of the other, we should check both are in fact 10 and error out otherwise. Changing them must involve recomputing these numbers or otherwise mucking about with shifts to ensure its back to 10 when we do this load muck. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vincent Guittot <vincent.guittot@linaro.org> |
|---|---|
| Date | 2015-09-08 17:20 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q6sQF-80Q-3@gated-at.bofh.it> |
| In reply to | #1220827 |
On 8 September 2015 at 16:10, Peter Zijlstra <peterz@infradead.org> wrote: > On Tue, Sep 08, 2015 at 03:39:37PM +0200, Vincent Guittot wrote: >> > Now, given all that, units are a complete mess here, and I'd not mind >> > something like: >> > >> > #if (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) != SCHED_CAPACITY_SHIFT >> > #error "something usefull" >> > #endif >> >> In this case why not simply doing >> #define SCHED_CAPACITY_SHIFT SCHED_LOAD_SHIFT >> or the opposite ? > > Sadly not enough; aside from the fact that we really should do !0 > LOAD_RESOLUTION on 64bit, the whole magic tables (runnable_avg_yN_*[]) > and LOAD_AVG_MAX* values rely on the unit being 1<<10. ah yes, i forgot to take into account the LOAD_RESOLUTION. So after some more thinking, i finally don't see where in the code, we will have a issue if SCHED_CAPACITY_SHIFT is not equal to (SCHED_LOAD_SHIFT - SCHED_LOAD_RESOLUTION) or not equal to 10 with the respect of using a value that doesn't overflow the variables Regards, Vincent > > So regardless of defining one in terms of the other, we should check > both are in fact 10 and error out otherwise. > > Changing them must involve recomputing these numbers or otherwise > mucking about with shifts to ensure its back to 10 when we do this load > muck. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2015-09-08 15:00 +0200 |
| Message-ID | <q6qFc-4GE-15@gated-at.bofh.it> |
| In reply to | #1220545 |
On 08/09/15 08:22, Vincent Guittot wrote: > On 7 September 2015 at 20:54, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >> On 07/09/15 17:21, Vincent Guittot wrote: >>> On 7 September 2015 at 17:37, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >>>> On 04/09/15 00:51, Steve Muckle wrote: >>>>> Hi Morten, Dietmar, >>>>> >>>>> On 08/14/2015 09:23 AM, Morten Rasmussen wrote: >>>>> ... >>>>>> + * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the >>>>>> + * recent utilization of currently non-runnable tasks on a CPU. It represents >>>>>> + * the amount of utilization of a CPU in the range [0..capacity_orig] where >>>>> >>>>> I see util_sum is scaled by SCHED_LOAD_SHIFT at the end of >>>>> __update_load_avg(). If there is now an assumption that util_avg may be >>>>> used directly as a capacity value, should it be changed to >>>>> SCHED_CAPACITY_SHIFT? These are equal right now, not sure if they will >>>>> always be or if they can be combined. >>>> >>>> You're referring to the code line >>>> >>>> 2647 sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX; >>>> >>>> in __update_load_avg()? >>>> >>>> Here we actually scale by 'SCHED_LOAD_SCALE/LOAD_AVG_MAX' so both values are >>>> load related. >>> >>> I agree with Steve that there is an issue from a unit point of view >>> >>> sa->util_sum and LOAD_AVG_MAX have the same unit so sa->util_avg is a >>> load because of << SCHED_LOAD_SHIFT) >>> >>> Before this patch , the translation from load to capacity unit was >>> done in get_cpu_usage with "* capacity) >> SCHED_LOAD_SHIFT" >>> >>> So you still have to change the unit from load to capacity with a "/ >>> SCHED_LOAD_SCALE * SCHED_CAPACITY_SCALE" somewhere. >>> >>> sa->util_avg = ((sa->util_sum << SCHED_LOAD_SHIFT) /SCHED_LOAD_SCALE * >>> SCHED_CAPACITY_SCALE / LOAD_AVG_MAX = (sa->util_sum << >>> SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; >> >> I see the point but IMHO this will only be necessary if the SCHED_LOAD_RESOLUTION >> stuff gets re-enabled again. >> >> It's not really about utilization or capacity units but rather about using the same >> SCALE/SHIFT values for both sides, right? > > It's both a unit and a SCALE/SHIFT problem, SCHED_LOAD_SHIFT and > SCHED_CAPACITY_SHIFT are defined separately so we must be sure to > scale the value in the right range. In the case of cpu_usage which > returns sa->util_avg , it's the capacity range not the load range. Still don't understand why it's a unit problem. IMHO LOAD/UTIL and CAPACITY have no unit. I agree that with the current patch-set we have a SHIFT/SCALE problem once SCHED_LOAD_RESOLUTION is set to != 0. > >> >> I always thought that scale_load_down() takes care of that. > > AFAIU, scale_load_down is a way to increase the resolution of the > load not to move from load to capacity IMHO, increasing the resolution of the load is done by re-enabling this define SCHED_LOAD_RESOLUTION 10 thing (or by setting SCHED_LOAD_RESOLUTION to something else than 0). I tried to figure out why we have this issue when comparing UTIL w/ CAPACITY and not LOAD w/ CAPACITY: Both are initialized like that: sa->load_avg = scale_load_down(se->load.weight); sa->load_sum = sa->load_avg * LOAD_AVG_MAX; sa->util_avg = scale_load_down(SCHED_LOAD_SCALE); sa->util_sum = LOAD_AVG_MAX; and we use 'se->on_rq * scale_load_down(se->load.weight)' as 'unsigned long weight' argument to call __update_load_avg() making sure the scaling differences between LOAD and CAPACITY are respected while updating sa->load_sum (and sa->load_avg). OTAH, we don't apply a scale_load_down for sa->util_[sum/avg] only a '<< SCHED_LOAD_SHIFT) / LOAD_AVG_MAX' on sa->util_avg. So changing '<< SCHED_LOAD_SHIFT' to '* scale_load_down(SCHED_LOAD_SCALE)' would be the logical thing to do. I agree that '<< SCHED_CAPACITY_SHIFT' would have the same effect but why using a CAPACITY related thing on the LOAD/UTIL side? The only reason would be the unit problem which I don't understand. > >> >> So shouldn't: >> >> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >> index 3445d2fb38f4..b80f799aface 100644 >> --- a/kernel/sched/fair.c >> +++ b/kernel/sched/fair.c >> @@ -2644,7 +2644,7 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa, >> cfs_rq->runnable_load_avg = >> div_u64(cfs_rq->runnable_load_sum, LOAD_AVG_MAX); >> } >> - sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX; >> + sa->util_avg = (sa->util_sum * scale_load_down(SCHED_LOAD_SCALE)) / LOAD_AVG_MAX; >> } >> >> return decayed; >> >> fix that issue in case SCHED_LOAD_RESOLUTION != 0 ? > > > No, but > sa->util_avg = (sa->util_sum << SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; > will fix the unit issue. > I agree that i don't change the result because both SCHED_LOAD_SHIFT > and SCHED_CAPACITY_SHIFT are set to 10 but as mentioned above, they > are set separately so it can make the difference if someone change one > SHIFT value. SCHED_LOAD_SHIFT and SCHED_CAPACITY_SHIFT can be set separately but the way to change SCHED_LOAD_SHIFT is by re-enabling the define SCHED_LOAD_RESOLUTION 10 in kernel/sched/sched.h. I guess nobody wants to change SCHED_CAPACITY_[SHIFT/SCALE]. Cheers, -- Dietmar [...] -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vincent Guittot <vincent.guittot@linaro.org> |
|---|---|
| Date | 2015-09-08 16:10 +0200 |
| Subject | Re: [PATCH 5/6] sched/fair: Get rid of scaling utilization by capacity_orig |
| Message-ID | <q6rKW-6sO-17@gated-at.bofh.it> |
| In reply to | #1220760 |
On 8 September 2015 at 14:50, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: > On 08/09/15 08:22, Vincent Guittot wrote: >> On 7 September 2015 at 20:54, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >>> On 07/09/15 17:21, Vincent Guittot wrote: >>>> On 7 September 2015 at 17:37, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >>>>> On 04/09/15 00:51, Steve Muckle wrote: >>>>>> Hi Morten, Dietmar, >>>>>> >>>>>> On 08/14/2015 09:23 AM, Morten Rasmussen wrote: >>>>>> ... >>>>>>> + * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the >>>>>>> + * recent utilization of currently non-runnable tasks on a CPU. It represents >>>>>>> + * the amount of utilization of a CPU in the range [0..capacity_orig] where >>>>>> >>>>>> I see util_sum is scaled by SCHED_LOAD_SHIFT at the end of >>>>>> __update_load_avg(). If there is now an assumption that util_avg may be >>>>>> used directly as a capacity value, should it be changed to >>>>>> SCHED_CAPACITY_SHIFT? These are equal right now, not sure if they will >>>>>> always be or if they can be combined. >>>>> >>>>> You're referring to the code line >>>>> >>>>> 2647 sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX; >>>>> >>>>> in __update_load_avg()? >>>>> >>>>> Here we actually scale by 'SCHED_LOAD_SCALE/LOAD_AVG_MAX' so both values are >>>>> load related. >>>> >>>> I agree with Steve that there is an issue from a unit point of view >>>> >>>> sa->util_sum and LOAD_AVG_MAX have the same unit so sa->util_avg is a >>>> load because of << SCHED_LOAD_SHIFT) >>>> >>>> Before this patch , the translation from load to capacity unit was >>>> done in get_cpu_usage with "* capacity) >> SCHED_LOAD_SHIFT" >>>> >>>> So you still have to change the unit from load to capacity with a "/ >>>> SCHED_LOAD_SCALE * SCHED_CAPACITY_SCALE" somewhere. >>>> >>>> sa->util_avg = ((sa->util_sum << SCHED_LOAD_SHIFT) /SCHED_LOAD_SCALE * >>>> SCHED_CAPACITY_SCALE / LOAD_AVG_MAX = (sa->util_sum << >>>> SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; >>> >>> I see the point but IMHO this will only be necessary if the SCHED_LOAD_RESOLUTION >>> stuff gets re-enabled again. >>> >>> It's not really about utilization or capacity units but rather about using the same >>> SCALE/SHIFT values for both sides, right? >> >> It's both a unit and a SCALE/SHIFT problem, SCHED_LOAD_SHIFT and >> SCHED_CAPACITY_SHIFT are defined separately so we must be sure to >> scale the value in the right range. In the case of cpu_usage which >> returns sa->util_avg , it's the capacity range not the load range. > > Still don't understand why it's a unit problem. IMHO LOAD/UTIL and > CAPACITY have no unit. If you set 2 different values to SCHED_LOAD_SHIFT and SCHED_CAPACITY_SHIFT for test purpose, you will see that util_avg will not use the right range of value If we don't take into account freq and cpu invariance in a 1st step sa->util_sum is a load in the range [0..LOAD_AVG_MAX]. I say load because of the max value the current implementation of util_avg is sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX so sa->util_avg is a load in the range [0..SCHED_LOAD_SCALE] the current implementation of get_cpu_usage is return (sa->util_avg * capacity_orig_of(cpu)) >> SCHED_LOAD_SHIFT so the usage has the same unit and range as capacity of the cpu and can be compared with another capacity value Your patchset returns directly sa->util_avg which is a load to compare it with capacity value So you have to convert sa->util_avg from load to capacity so if you have sa->util_avg = (sa->util_sum << SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX sa->util_avg is now a capacity with the same range as you cpu thanks to the cpu invariance factor that the patch 3 has added. the << SCHED_CAPACITY_SHIFT above can be optimized with the >> SCHED_CAPACITY_SHIFT included in sa->util_sum += scale(contrib, scale_cpu); as mentioned by Peter At now, SCHED_CAPACITY_SHIFT is set to 10 as well as SCHED_LOAD_SHIFT so using one instead of the other doesn't change the result but if it's no more the case, we need to take care of the range/unit that we use Regards, Vincent > > I agree that with the current patch-set we have a SHIFT/SCALE problem > once SCHED_LOAD_RESOLUTION is set to != 0. > >> >>> >>> I always thought that scale_load_down() takes care of that. >> >> AFAIU, scale_load_down is a way to increase the resolution of the >> load not to move from load to capacity > > IMHO, increasing the resolution of the load is done by re-enabling this > define SCHED_LOAD_RESOLUTION 10 thing (or by setting > SCHED_LOAD_RESOLUTION to something else than 0). > > I tried to figure out why we have this issue when comparing UTIL w/ > CAPACITY and not LOAD w/ CAPACITY: > > Both are initialized like that: > > sa->load_avg = scale_load_down(se->load.weight); > sa->load_sum = sa->load_avg * LOAD_AVG_MAX; > sa->util_avg = scale_load_down(SCHED_LOAD_SCALE); > sa->util_sum = LOAD_AVG_MAX; > > and we use 'se->on_rq * scale_load_down(se->load.weight)' as 'unsigned > long weight' argument to call __update_load_avg() making sure the > scaling differences between LOAD and CAPACITY are respected while > updating sa->load_sum (and sa->load_avg). > > OTAH, we don't apply a scale_load_down for sa->util_[sum/avg] only a '<< > SCHED_LOAD_SHIFT) / LOAD_AVG_MAX' on sa->util_avg. > So changing '<< SCHED_LOAD_SHIFT' to '* > scale_load_down(SCHED_LOAD_SCALE)' would be the logical thing to do. > > I agree that '<< SCHED_CAPACITY_SHIFT' would have the same effect but > why using a CAPACITY related thing on the LOAD/UTIL side? The only > reason would be the unit problem which I don't understand. > >> >>> >>> So shouldn't: >>> >>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c >>> index 3445d2fb38f4..b80f799aface 100644 >>> --- a/kernel/sched/fair.c >>> +++ b/kernel/sched/fair.c >>> @@ -2644,7 +2644,7 @@ __update_load_avg(u64 now, int cpu, struct sched_avg *sa, >>> cfs_rq->runnable_load_avg = >>> div_u64(cfs_rq->runnable_load_sum, LOAD_AVG_MAX); >>> } >>> - sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX; >>> + sa->util_avg = (sa->util_sum * scale_load_down(SCHED_LOAD_SCALE)) / LOAD_AVG_MAX; >>> } >>> >>> return decayed; >>> >>> fix that issue in case SCHED_LOAD_RESOLUTION != 0 ? >> >> >> No, but >> sa->util_avg = (sa->util_sum << SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; >> will fix the unit issue. >> I agree that i don't change the result because both SCHED_LOAD_SHIFT >> and SCHED_CAPACITY_SHIFT are set to 10 but as mentioned above, they >> are set separately so it can make the difference if someone change one >> SHIFT value. > > SCHED_LOAD_SHIFT and SCHED_CAPACITY_SHIFT can be set separately but the > way to change SCHED_LOAD_SHIFT is by re-enabling the define > SCHED_LOAD_RESOLUTION 10 in kernel/sched/sched.h. I guess nobody wants > to change SCHED_CAPACITY_[SHIFT/SCALE]. > > Cheers, > > -- Dietmar > > [...] > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2015-09-08 16:30 +0200 |
| Message-ID | <q6s4i-6PU-3@gated-at.bofh.it> |
| In reply to | #1220813 |
On 08/09/15 15:01, Vincent Guittot wrote: > On 8 September 2015 at 14:50, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >> On 08/09/15 08:22, Vincent Guittot wrote: >>> On 7 September 2015 at 20:54, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >>>> On 07/09/15 17:21, Vincent Guittot wrote: >>>>> On 7 September 2015 at 17:37, Dietmar Eggemann <dietmar.eggemann@arm.com> wrote: >>>>>> On 04/09/15 00:51, Steve Muckle wrote: >>>>>>> Hi Morten, Dietmar, >>>>>>> >>>>>>> On 08/14/2015 09:23 AM, Morten Rasmussen wrote: >>>>>>> ... >>>>>>>> + * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the >>>>>>>> + * recent utilization of currently non-runnable tasks on a CPU. It represents >>>>>>>> + * the amount of utilization of a CPU in the range [0..capacity_orig] where >>>>>>> >>>>>>> I see util_sum is scaled by SCHED_LOAD_SHIFT at the end of >>>>>>> __update_load_avg(). If there is now an assumption that util_avg may be >>>>>>> used directly as a capacity value, should it be changed to >>>>>>> SCHED_CAPACITY_SHIFT? These are equal right now, not sure if they will >>>>>>> always be or if they can be combined. >>>>>> >>>>>> You're referring to the code line >>>>>> >>>>>> 2647 sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX; >>>>>> >>>>>> in __update_load_avg()? >>>>>> >>>>>> Here we actually scale by 'SCHED_LOAD_SCALE/LOAD_AVG_MAX' so both values are >>>>>> load related. >>>>> >>>>> I agree with Steve that there is an issue from a unit point of view >>>>> >>>>> sa->util_sum and LOAD_AVG_MAX have the same unit so sa->util_avg is a >>>>> load because of << SCHED_LOAD_SHIFT) >>>>> >>>>> Before this patch , the translation from load to capacity unit was >>>>> done in get_cpu_usage with "* capacity) >> SCHED_LOAD_SHIFT" >>>>> >>>>> So you still have to change the unit from load to capacity with a "/ >>>>> SCHED_LOAD_SCALE * SCHED_CAPACITY_SCALE" somewhere. >>>>> >>>>> sa->util_avg = ((sa->util_sum << SCHED_LOAD_SHIFT) /SCHED_LOAD_SCALE * >>>>> SCHED_CAPACITY_SCALE / LOAD_AVG_MAX = (sa->util_sum << >>>>> SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX; >>>> >>>> I see the point but IMHO this will only be necessary if the SCHED_LOAD_RESOLUTION >>>> stuff gets re-enabled again. >>>> >>>> It's not really about utilization or capacity units but rather about using the same >>>> SCALE/SHIFT values for both sides, right? >>> >>> It's both a unit and a SCALE/SHIFT problem, SCHED_LOAD_SHIFT and >>> SCHED_CAPACITY_SHIFT are defined separately so we must be sure to >>> scale the value in the right range. In the case of cpu_usage which >>> returns sa->util_avg , it's the capacity range not the load range. >> >> Still don't understand why it's a unit problem. IMHO LOAD/UTIL and >> CAPACITY have no unit. > > If you set 2 different values to SCHED_LOAD_SHIFT and > SCHED_CAPACITY_SHIFT for test purpose, you will see that util_avg will > not use the right range of value > > If we don't take into account freq and cpu invariance in a 1st step > > sa->util_sum is a load in the range [0..LOAD_AVG_MAX]. I say load > because of the max value > > the current implementation of util_avg is > sa->util_avg = (sa->util_sum << SCHED_LOAD_SHIFT) / LOAD_AVG_MAX > > so sa->util_avg is a load in the range [0..SCHED_LOAD_SCALE] > > the current implementation of get_cpu_usage is > return (sa->util_avg * capacity_orig_of(cpu)) >> SCHED_LOAD_SHIFT > > so the usage has the same unit and range as capacity of the cpu and > can be compared with another capacity value > > Your patchset returns directly sa->util_avg which is a load to compare > it with capacity value > > So you have to convert sa->util_avg from load to capacity so if you have > sa->util_avg = (sa->util_sum << SCHED_CAPACITY_SHIFT) / LOAD_AVG_MAX > > sa->util_avg is now a capacity with the same range as you cpu thanks > to the cpu invariance factor that the patch 3 has added. > > the << SCHED_CAPACITY_SHIFT above can be optimized with the >> > SCHED_CAPACITY_SHIFT included in > sa->util_sum += scale(contrib, scale_cpu); > as mentioned by Peter > > At now, SCHED_CAPACITY_SHIFT is set to 10 as well as SCHED_LOAD_SHIFT > so using one instead of the other doesn't change the result but if > it's no more the case, we need to take care of the range/unit that we > use No arguing here, I just called this a SHIFT/SCALE problem. [...] -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web