Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1442382 > unrolled thread

Re: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support

Started byVincent Guittot <vincent.guittot@linaro.org>
First post2016-07-13 14:10 +0200
Last post2016-07-13 18:00 +0200
Articles 2 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support Vincent Guittot <vincent.guittot@linaro.org> - 2016-07-13 14:10 +0200
    Re: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity  support Morten Rasmussen <morten.rasmussen@arm.com> - 2016-07-13 18:00 +0200

#1442382 — Re: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support

FromVincent Guittot <vincent.guittot@linaro.org>
Date2016-07-13 14:10 +0200
SubjectRe: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support
Message-ID<rUr9f-2Tk-11@gated-at.bofh.it>
Hi Morten,

On 22 June 2016 at 19:03, Morten Rasmussen <morten.rasmussen@arm.com> wrote:
> Hi,
>
> The scheduler is currently not doing much to help performance on systems with
> asymmetric compute capacities (read ARM big.LITTLE). This series improves the
> situation with a few tweaks mainly to the task wake-up path that considers
> compute capacity at wake-up and not just whether a cpu is idle for these
> systems. This gives us consistent, and potentially higher, throughput in
> partially utilized scenarios. SMP behaviour and performance should be
> unaffected.
>
> Test 0:
>         for i in `seq 1 10`; \
>                do sysbench --test=cpu --max-time=3 --num-threads=1 run; \
>                done \
>         | awk '{if ($4=="events:") {print $5; sum +=$5; runs +=1}} \
>                END {print "Average events: " sum/runs}'
>
> Target: ARM TC2 (2xA15+3xA7)
>
>         (Higher is better)
> tip:    Average events: 146.9
> patch:  Average events: 217.9
>
> Test 1:
>         perf stat --null --repeat 10 -- \
>         perf bench sched messaging -g 50 -l 5000
>
> Target: Intel IVB-EP (2*10*2)
>
> tip:    4.861970420 seconds time elapsed ( +-  1.39% )
> patch:  4.886204224 seconds time elapsed ( +-  0.75% )
>
> Target: ARM TC2 A7-only (3xA7) (-l 1000)
>
> tip:    61.485682596 seconds time elapsed ( +-  0.07% )
> patch:  62.667950130 seconds time elapsed ( +-  0.36% )
>
> More analysis:
>
> Statistics from mixed periodic task workload (rt-app) containing both
> big and little task, single run on ARM TC2:
>
> tu   = Task utilization big/little
> pcpu = Previous cpu big/little
> tcpu = This (waker) cpu big/little
> dl   = New cpu is little
> db   = New cpu is big
> sis  = New cpu chosen by select_idle_sibling()
> figc = New cpu chosen by find_idlest_*()
> ww   = wake_wide(task) count for figc wakeups
> bw   = sd_flag & SD_BALANCE_WAKE (non-fork/exec wake)
>        for figc wakeups
>
> case tu   pcpu tcpu   dl   db  sis figc   ww   bw
> 1    l    l    l     122   68   28  162  161  161
> 2    l    l    b      11    4    0   15   15   15
> 3    l    b    l       0  252    8  244  244  244
> 4    l    b    b      36 1928  711 1253 1016 1016
> 5    b    l    l       5   19    0   24   22   24
> 6    b    l    b       5    1    0    6    0    6
> 7    b    b    l       0   31    0   31   31   31
> 8    b    b    b       1  194  109   86   59   59
> --------------------------------------------------
>                      180 2497  856 1821

I'm not sure to know how to interpret all these statistics

>
> Cases 1-4 + 8 are fine to be served by select_idle_sibling() as both
> this_cpu and prev_cpu are suitable cpus for the task. However, as the
> figc column reveals, those cases are often served by find_idlest_*()
> anyway due to wake_wide() sending the wakeup that way when
> SD_BALANCE_WAKE is set on the sched_domains.
>
> Pulling in the wakee_flip patch (dropped in v2) from v1 shifts a
> significant share of the wakeups to sis from figc:
>
> case tu   pcpu tcpu   dl   db  sis figc   ww   bw
> 1    l    l    l     537    8  537    8    6    6
> 2    l    l    b      49   11   32   28   28   28
> 3    l    b    l       4  323  322    5    5    5
> 4    l    b    b       1 1910 1209  702  458  456
> 5    b    l    l       0    5    0    5    1    5
> 6    b    l    b       0    0    0    0    0    0
> 7    b    b    l       0   32    0   32    2   32
> 8    b    b    b       0  198  168   30   13   13
> --------------------------------------------------
>                      591 2487 2268  810
>
> Notes:
>
> Active migration of tasks away from small capacity cpus isn't addressed
> in this set although it is necessary for consistent throughput in other
> scenarios on asymmetric cpu capacity systems.
>
> The infrastructure to enable capacity awareness for arm64 is not provided here
> but will be based on Juri's DT bindings patch set [1]. A combined preview
> branch is available [2].
>
> [1] https://lkml.org/lkml/2016/6/15/291
> [2] git://linux-arm.org/linux-power.git capacity_awareness_v2_arm64_v1
>
> Patch   1-3: Generic fixes and clean-ups.
> Patch  4-11: Improve capacity awareness.
> Patch 11-12: Arch features for arm to enable asymmetric capacity support.
>
> v2:
>
> - Dropped patch ignoring wakee_flips for pid=0 for now as we can not
>   distinguish cpu time processing irqs from idle time.
>
> - Dropped disabling WAKE_AFFINE as suggested by Vincent to allow more
>   scenarios to use fast-path (select_idle_sibling()). Asymmetric wake
>   conditions adjusted accordingly.
>
> - Changed use of new SD_ASYM_CPUCAPACITY slightly. Now enables
>   SD_BALANCE_WAKE.
>
> - Minor clean-ups and rebased to more recent tip/sched/core.
>
> v1: https://lkml.org/lkml/2014/5/23/621
>
> Dietmar Eggemann (1):
>   sched: Store maximum per-cpu capacity in root domain
>
> Morten Rasmussen (12):
>   sched: Fix power to capacity renaming in comment
>   sched/fair: Consistent use of prev_cpu in wakeup path
>   sched/fair: Optimize find_idlest_cpu() when there is no choice
>   sched: Introduce SD_ASYM_CPUCAPACITY sched_domain topology flag
>   sched: Enable SD_BALANCE_WAKE for asymmetric capacity systems
>   sched/fair: Let asymmetric cpu configurations balance at wake-up
>   sched/fair: Compute task/cpu utilization at wake-up more correctly
>   sched/fair: Consider spare capacity in find_idlest_group()
>   sched: Add per-cpu max capacity to sched_group_capacity
>   sched/fair: Avoid pulling tasks from non-overloaded higher capacity
>     groups
>   arm: Set SD_ASYM_CPUCAPACITY for big.LITTLE platforms
>   arm: Update arch_scale_cpu_capacity() to reflect change to define
>
>  arch/arm/include/asm/topology.h |   5 +
>  arch/arm/kernel/topology.c      |  25 ++++-
>  include/linux/sched.h           |   3 +-
>  kernel/sched/core.c             |  21 +++-
>  kernel/sched/fair.c             | 212 +++++++++++++++++++++++++++++++++++-----
>  kernel/sched/sched.h            |   5 +-
>  6 files changed, 241 insertions(+), 30 deletions(-)
>
> --
> 1.9.1
>

[toc] | [next] | [standalone]


#1442570 — Re: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support

FromMorten Rasmussen <morten.rasmussen@arm.com>
Date2016-07-13 18:00 +0200
SubjectRe: [PATCH v2 00/13] sched: Clean-ups and asymmetric cpu capacity support
Message-ID<rUuJQ-53V-37@gated-at.bofh.it>
In reply to#1442382
On Wed, Jul 13, 2016 at 02:06:17PM +0200, Vincent Guittot wrote:
> Hi Morten,
> 
> On 22 June 2016 at 19:03, Morten Rasmussen <morten.rasmussen@arm.com> wrote:
> > Hi,
> >
> > The scheduler is currently not doing much to help performance on systems with
> > asymmetric compute capacities (read ARM big.LITTLE). This series improves the
> > situation with a few tweaks mainly to the task wake-up path that considers
> > compute capacity at wake-up and not just whether a cpu is idle for these
> > systems. This gives us consistent, and potentially higher, throughput in
> > partially utilized scenarios. SMP behaviour and performance should be
> > unaffected.
> >
> > Test 0:
> >         for i in `seq 1 10`; \
> >                do sysbench --test=cpu --max-time=3 --num-threads=1 run; \
> >                done \
> >         | awk '{if ($4=="events:") {print $5; sum +=$5; runs +=1}} \
> >                END {print "Average events: " sum/runs}'
> >
> > Target: ARM TC2 (2xA15+3xA7)
> >
> >         (Higher is better)
> > tip:    Average events: 146.9
> > patch:  Average events: 217.9
> >
> > Test 1:
> >         perf stat --null --repeat 10 -- \
> >         perf bench sched messaging -g 50 -l 5000
> >
> > Target: Intel IVB-EP (2*10*2)
> >
> > tip:    4.861970420 seconds time elapsed ( +-  1.39% )
> > patch:  4.886204224 seconds time elapsed ( +-  0.75% )
> >
> > Target: ARM TC2 A7-only (3xA7) (-l 1000)
> >
> > tip:    61.485682596 seconds time elapsed ( +-  0.07% )
> > patch:  62.667950130 seconds time elapsed ( +-  0.36% )
> >
> > More analysis:
> >
> > Statistics from mixed periodic task workload (rt-app) containing both
> > big and little task, single run on ARM TC2:
> >
> > tu   = Task utilization big/little
> > pcpu = Previous cpu big/little
> > tcpu = This (waker) cpu big/little
> > dl   = New cpu is little
> > db   = New cpu is big
> > sis  = New cpu chosen by select_idle_sibling()
> > figc = New cpu chosen by find_idlest_*()
> > ww   = wake_wide(task) count for figc wakeups
> > bw   = sd_flag & SD_BALANCE_WAKE (non-fork/exec wake)
> >        for figc wakeups
> >
> > case tu   pcpu tcpu   dl   db  sis figc   ww   bw
> > 1    l    l    l     122   68   28  162  161  161
> > 2    l    l    b      11    4    0   15   15   15
> > 3    l    b    l       0  252    8  244  244  244
> > 4    l    b    b      36 1928  711 1253 1016 1016
> > 5    b    l    l       5   19    0   24   22   24
> > 6    b    l    b       5    1    0    6    0    6
> > 7    b    b    l       0   31    0   31   31   31
> > 8    b    b    b       1  194  109   86   59   59
> > --------------------------------------------------
> >                      180 2497  856 1821
> 
> I'm not sure to know how to interpret all these statistics

Thanks for looking into the details. Let me provide a bit more context.

After our discussion around v1 I wanted to understand how the patches
works with different combinations of task utilization, prev_cpu, and
waking cpu. IIRC, the outcome our discussion was that tasks with
utilization too high to fit little cpus should go on big cpus, tasks
small enough to fit anywhere can go anywhere. For the latter we don't
want to spend too much time placing them as they essentially don't care
so they can be placed using select_idle_sibling().

So, I created a workload with rt-app with a number of different periodic
tasks with different periods and busy times. I traced all wake-ups and
put them into eight categories depending on the wake-up scenario, i.e.
task utilization, prev_cpu, and waking cpu (tu, pcpu, and tcpu).

The next two columns (dl, db) show the number of wake-ups that ended up
on a little or big cpu. If we take case 1 as an example, we had 190
wake-ups in total where a little task last ran on a little cpu and was
woken up by a little cpu. 122 of the wake-ups ended up selecting a
little cpu again, while in 68 cases it went to a big cpu. That is fine
according to our scheduling policy above.

The sis and figc columns show the split between wake-ups being handled
by select_idle_sibling() versus find_idlest_*(). Coming back to case 1,
28 wake-ups where handled by the former and 162 by the latter. We can't
really say exactly how many of select_idle_sibling() wake-ups that ended
up on big or little, but based on the numbers it is clear that in some
cases find_idlest_*() chose a big cpu for a little task (68 > 28).

The last two columns, ww and bw, try to explain why we have so many
wake-ups handled by find_idlest_*() for cases (1-4, and 8) where we
could have used select_idle_sibling(). The bw number is the number of
find_idlest_*() wake-ups that were passed the SD_BALANCE_WAKE flag (i.e.
non-FORK and non-EXEC wake-ups). FORK and EXEC wake-ups always take the
find_idlest_*() route, so we should ignore those. For case 1 it turned
out that only one of the 162 figc wake-ups was a FORK/EXEC wake-up. So
something else mush have caused those wake-ups to not go via
select_idle_sibling(). The ww column explains why as it shows how many
of the figc wake-ups where wake_wide() returned true and therefore
disabled want_affine. Because we have enabled SD_BALANCE_WAKE on the
sched_domains, !wake_affine wake-ups no longer end up using
select_idle_sibling() anyway, but end up using find_idlest_*().

Thinking more about it, should we force those tasks to use
select_idle_sibling() anyway? Something like the below could do it I
think:

@@ -5444,7 +5444,7 @@ select_task_rq_fair(struct task_struct *p, int prev_cpu, int sd_flag, int wake_f
                        new_cpu = cpu;
        }
 
-       if (!sd) {
+       if (!sd || (!wake_cap(p, cpu, prev_cpu) && sd_flags & SD_BALANCE_WAKE)) {
                if (sd_flag & SD_BALANCE_WAKE) /* XXX always ? */
                        new_cpu = select_idle_sibling(p, prev_cpu, new_cpu);

Ideally, cases 5-7 should be handled by find_idlest_*(), which seems to
hold true in the table above, and cases 1-4, and 8 should be handled by
select_idle_sibling(), which isn't always the case due to wake_wide().

Thoughts?

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web