Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1227754 > unrolled thread
| Started by | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| First post | 2015-09-18 12:40 +0200 |
| Last post | 2015-09-21 00:10 +0200 |
| Articles | 3 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [RFCv5 PATCH 32/46] sched: Energy-aware wake-up task placement Dietmar Eggemann <dietmar.eggemann@arm.com> - 2015-09-18 12:40 +0200
Re: [RFCv5 PATCH 32/46] sched: Energy-aware wake-up task placement Steve Muckle <steve.muckle@linaro.org> - 2015-09-20 20:40 +0200
Re: [RFCv5 PATCH 32/46] sched: Energy-aware wake-up task placement Leo Yan <leo.yan@linaro.org> - 2015-09-21 00:10 +0200
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2015-09-18 12:40 +0200 |
| Subject | Re: [RFCv5 PATCH 32/46] sched: Energy-aware wake-up task placement |
| Message-ID | <qa1fc-8td-3@gated-at.bofh.it> |
On 02/09/15 18:11, Leo Yan wrote:
> On Tue, Jul 07, 2015 at 07:24:15PM +0100, Morten Rasmussen wrote:
>> Let available compute capacity and estimated energy impact select
>> wake-up target cpu when energy-aware scheduling is enabled and the
>> system in not over-utilized (above the tipping point).
>>
>> energy_aware_wake_cpu() attempts to find group of cpus with sufficient
>> compute capacity to accommodate the task and find a cpu with enough spare
>> capacity to handle the task within that group. Preference is given to
>> cpus with enough spare capacity at the current OPP. Finally, the energy
>> impact of the new target and the previous task cpu is compared to select
>> the wake-up target cpu.
>>
>> cc: Ingo Molnar <mingo@redhat.com>
>> cc: Peter Zijlstra <peterz@infradead.org>
>>
>> Signed-off-by: Morten Rasmussen <morten.rasmussen@arm.com>
>> ---
>> kernel/sched/fair.c | 85 ++++++++++++++++++++++++++++++++++++++++++++++++++++-
>> 1 file changed, 84 insertions(+), 1 deletion(-)
>>
>> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
>> index 0f7dbda4..01f7337 100644
>> --- a/kernel/sched/fair.c
>> +++ b/kernel/sched/fair.c
>> @@ -5427,6 +5427,86 @@ static int select_idle_sibling(struct task_struct *p, int target)
>> return target;
>> }
>>
>> +static int energy_aware_wake_cpu(struct task_struct *p, int target)
>> +{
>> + struct sched_domain *sd;
>> + struct sched_group *sg, *sg_target;
>> + int target_max_cap = INT_MAX;
>> + int target_cpu = task_cpu(p);
>> + int i;
>> +
>> + sd = rcu_dereference(per_cpu(sd_ea, task_cpu(p)));
>> +
>> + if (!sd)
>> + return target;
>> +
>> + sg = sd->groups;
>> + sg_target = sg;
>> +
>> + /*
>> + * Find group with sufficient capacity. We only get here if no cpu is
>> + * overutilized. We may end up overutilizing a cpu by adding the task,
>> + * but that should not be any worse than select_idle_sibling().
>> + * load_balance() should sort it out later as we get above the tipping
>> + * point.
>> + */
>> + do {
>> + /* Assuming all cpus are the same in group */
>> + int max_cap_cpu = group_first_cpu(sg);
>> +
>> + /*
>> + * Assume smaller max capacity means more energy-efficient.
>> + * Ideally we should query the energy model for the right
>> + * answer but it easily ends up in an exhaustive search.
>> + */
>> + if (capacity_of(max_cap_cpu) < target_max_cap &&
>> + task_fits_capacity(p, max_cap_cpu)) {
>> + sg_target = sg;
>> + target_max_cap = capacity_of(max_cap_cpu);
>> + }
>
> Here should consider scenario for two groups have same capacity?
> This will benefit for the case LITTLE.LITTLE. So the code will be
> looks like below:
>
> int target_sg_cpu = INT_MAX;
>
> if (capacity_of(max_cap_cpu) <= target_max_cap &&
> task_fits_capacity(p, max_cap_cpu)) {
>
> if ((capacity_of(max_cap_cpu) == target_max_cap) &&
> (target_sg_cpu < max_cap_cpu))
> continue;
>
> target_sg_cpu = max_cap_cpu;
> sg_target = sg;
> target_max_cap = capacity_of(max_cap_cpu);
> }
>
It's true that on your SMP system the target sched_group 'sg_target'
depends only on 'task_cpu(p)' because this determines sched_domain 'sd'
(and so the order of sched_groups for the iteration).
So the current do-while loop to select 'sg_target' for an SMP system
makes little sense.
But why should we favour the first sched_group (cluster) (the one w/ the
lower max_cap_cpu number) in this situation?
[...]
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2015-09-20 20:40 +0200 |
| Message-ID | <qaRGN-7Ww-7@gated-at.bofh.it> |
| In reply to | #1227754 |
On 09/18/2015 03:34 AM, Dietmar Eggemann wrote:
>> Here should consider scenario for two groups have same capacity?
>> This will benefit for the case LITTLE.LITTLE. So the code will be
>> looks like below:
>>
>> int target_sg_cpu = INT_MAX;
>>
>> if (capacity_of(max_cap_cpu) <= target_max_cap &&
>> task_fits_capacity(p, max_cap_cpu)) {
>>
>> if ((capacity_of(max_cap_cpu) == target_max_cap) &&
>> (target_sg_cpu < max_cap_cpu))
>> continue;
>>
>> target_sg_cpu = max_cap_cpu;
>> sg_target = sg;
>> target_max_cap = capacity_of(max_cap_cpu);
>> }
>>
>
> It's true that on your SMP system the target sched_group 'sg_target'
> depends only on 'task_cpu(p)' because this determines sched_domain 'sd'
> (and so the order of sched_groups for the iteration).
>
> So the current do-while loop to select 'sg_target' for an SMP system
> makes little sense.
>
> But why should we favour the first sched_group (cluster) (the one w/ the
> lower max_cap_cpu number) in this situation?
Running the originally proposed code on a system with two identical
clusters, it looks like we'll always end up doing an energy-aware search
in the task's prev_cpu cluster (sched_group). If you had small tasks
scattered across both clusters, energy_aware_wake_cpu() would not
consider condensing them on a single cluster. Leo was this the issue you
were seeing?
However I think there may be negative side effects with the proposed
policy above as well - won't this cause us to pack the first cluster
until it's 100% full (running at fmax) before using the second cluster?
That would also be bad for power.
thanks,
Steve
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Leo Yan <leo.yan@linaro.org> |
|---|---|
| Date | 2015-09-21 00:10 +0200 |
| Message-ID | <qaUY2-4jZ-5@gated-at.bofh.it> |
| In reply to | #1228936 |
On Sun, Sep 20, 2015 at 11:39:16AM -0700, Steve Muckle wrote:
> On 09/18/2015 03:34 AM, Dietmar Eggemann wrote:
> >> Here should consider scenario for two groups have same capacity?
> >> This will benefit for the case LITTLE.LITTLE. So the code will be
> >> looks like below:
> >>
> >> int target_sg_cpu = INT_MAX;
> >>
> >> if (capacity_of(max_cap_cpu) <= target_max_cap &&
> >> task_fits_capacity(p, max_cap_cpu)) {
> >>
> >> if ((capacity_of(max_cap_cpu) == target_max_cap) &&
> >> (target_sg_cpu < max_cap_cpu))
> >> continue;
> >>
> >> target_sg_cpu = max_cap_cpu;
> >> sg_target = sg;
> >> target_max_cap = capacity_of(max_cap_cpu);
> >> }
> >>
> >
> > It's true that on your SMP system the target sched_group 'sg_target'
> > depends only on 'task_cpu(p)' because this determines sched_domain 'sd'
> > (and so the order of sched_groups for the iteration).
> >
> > So the current do-while loop to select 'sg_target' for an SMP system
> > makes little sense.
> >
> > But why should we favour the first sched_group (cluster) (the one w/ the
> > lower max_cap_cpu number) in this situation?
>
> Running the originally proposed code on a system with two identical
> clusters, it looks like we'll always end up doing an energy-aware search
> in the task's prev_cpu cluster (sched_group). If you had small tasks
> scattered across both clusters, energy_aware_wake_cpu() would not
> consider condensing them on a single cluster. Leo was this the issue you
> were seeing?
Exactly.
> However I think there may be negative side effects with the proposed
> policy above as well - won't this cause us to pack the first cluster
> until it's 100% full (running at fmax) before using the second cluster?
> That would also be bad for power.
In this case of CPU is running at fmax, it's true that
task_fits_capacity() will return true. But here i think
cpu_overutilized() also will return true, so that means scheduler will
go back to use CFS's old way for loading balance. Finally tasks also
will be spread into two clusters.
Also reviewed the profiling result on Hikey with this modification
[1], rt-app 6%/13%/19%/25% place 8 tasks into one cluster as
possible, but rt-app 31%/38%/44%/50% also will place tasks to second
cluster. NOTE, I get this conclusion from CPU idle's duty cycle, but
not from real power data.
[1] https://lists.linaro.org/pipermail/eas-dev/2015-September/000218.html
Thanks,
Leo Yan
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web