Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1340127 > unrolled thread
| Started by | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| First post | 2016-02-23 02:30 +0100 |
| Last post | 2016-02-23 02:40 +0100 |
| Articles | 20 on this page of 47 — 8 participants |
Back to article view | Back to linux.kernel
[RFCv7 PATCH 00/10] sched: scheduler-driven CPU frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 07/10] sched/fair: jump to max OPP when crossing UP threshold Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 09/10] sched/deadline: split rt_avg in 2 distincts metrics Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 10/10] sched: rt scheduler sets capacity requirement Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 08/10] sched: remove call of sched_avg_update from sched_rt_avg_update Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-23 02:40 +0100
Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow Michael Turquette <mturquette@baylibre.com> - 2016-02-26 02:10 +0100
Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-26 02:20 +0100
Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-26 22:10 +0100
Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow Steve Muckle <steve.muckle@linaro.org> - 2016-02-26 02:20 +0100
[RFCv7 PATCH 06/10] sched/fair: cpufreq_sched triggers for load balancing Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
Re: [RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests Ricky Liang <jcliang@chromium.org> - 2016-03-01 08:00 +0100
Re: [RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests Steve Muckle <steve.muckle@linaro.org> - 2016-03-03 05:00 +0100
[RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-25 05:00 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-02-25 10:30 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-25 22:10 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-02-25 10:30 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-25 22:10 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-02-26 10:20 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-27 01:10 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-03-01 14:00 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-01 20:50 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-25 12:10 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-02-26 01:40 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-27 03:40 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-02-27 05:20 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-28 03:30 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-03-01 15:40 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-01 21:40 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-03-01 14:30 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-03-01 14:20 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Michael Turquette <mturquette@baylibre.com> - 2016-03-02 08:50 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection "Rafael J. Wysocki" <rafael@kernel.org> - 2016-03-03 03:50 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-03-03 05:00 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Juri Lelli <Juri.Lelli@arm.com> - 2016-03-03 10:40 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Peter Zijlstra <peterz@infradead.org> - 2016-03-03 14:10 +0100
Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection Ingo Molnar <mingo@kernel.org> - 2016-03-03 15:30 +0100
[RFCv7 PATCH 05/10] sched/{core,fair}: trigger OPP change request on fork() Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
[RFCv7 PATCH 01/10] sched: Compute cpu capacity available at current frequency Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:30 +0100
Re: [RFCv7 PATCH 01/10] sched: Compute cpu capacity available at current frequency "Rafael J. Wysocki" <rafael@kernel.org> - 2016-02-23 02:50 +0100
Re: [RFCv7 PATCH 01/10] sched: Compute cpu capacity available at current frequency Peter Zijlstra <peterz@infradead.org> - 2016-02-23 10:20 +0100
Re: [RFCv7 PATCH 01/10] sched: Compute cpu capacity available at current frequency "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-02-26 02:40 +0100
Re: [RFCv7 PATCH 01/10] sched: Compute cpu capacity available at current frequency Peter Zijlstra <peterz@infradead.org> - 2016-02-26 10:20 +0100
Re: [RFCv7 PATCH 00/10] sched: scheduler-driven CPU frequency selection Steve Muckle <steve.muckle@linaro.org> - 2016-02-23 02:40 +0100
Page 1 of 3 [1] 2 3 Next page →
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 00/10] sched: scheduler-driven CPU frequency selection |
| Message-ID | <r59Xz-6rQ-1@gated-at.bofh.it> |
Scheduler-driven CPU frequency selection hopes to exploit both
per-task and global information in the scheduler to improve frequency
selection policy and achieve lower power consumption, improved
responsiveness/performance, and less reliance on heuristics and
tunables. For further discussion of this integration see [0].
This patch series implements a cpufreq governor which collects CPU
capacity requests from the fair, realtime, and deadline scheduling
classes. The fair and realtime scheduling classes are modified to make
these requests. The deadline class is not yet modified to make CPU
capacity requests.
Changes in this series since RFCv6 [1], posted December 9, 2015:
Patch 3, sched: scheduler-driven cpu frequency selection
- Added Kconfig dependency on IRQ_WORK.
- Reworked locking.
- Make throttling optional - it is not required in order to ensure that
the previous frequency transition is complete.
- Some fixes in cpufreq_sched_thread related to the task state.
- Changes to support mixed fast and slow path operation.
Patch 7: sched/fair: jump to max OPP when crossing UP threshold
- move sched_freq_tick() call so rq lock is still held
Patch 9: sched/deadline: split rt_avg in 2 distincts metrics
- RFCv6 calculated DL capacity from DL task parameters, RFCv7 restores
the original method of calculation but keeps DL capacity separate
Patch 10: sched: rt scheduler sets capacity requirement
- change #ifdef from CONFIG_SMP, trivial cleanup
Profiling results:
Performance profiling has been done by using rt-app [2] to generate
various periodic workloads with a particular duty cycle. The time to
complete the busy portion of the duty cycle is measured and overhead
is calculated as
overhead = (busy_duration_test_gov - busy_duration_perf_gov)/
(busy_duration_pwrsave_gov - busy_duration_perf_gov)
This shows as a percentage how close the governor is to running the
workload at fmin (100%) or fmax (0%). The number of times the busy
duration exceeds the period of the periodic workload (an "overrun") is
also recorded. In the table below the performance of the ondemand
(sampling_rate = 20ms), interactive (default tunables), and
scheduler-driven governors are evaluated using these metrics. The test
platform is a Samsung Chromebook 2 ("Peach Pi"). The workload is
affined to CPU0, an A15 with an fmin of 200MHz and an fmax of
1.8GHz. The interactive governor was incorporated/adapted from [3]. A
branch with the interactive governor and a few required dependency
patches for ARM is available at [4].
More detailed explanation of the columns below:
run: duration at fmax of the busy portion of the periodic workload in msec
period: duration of the entire period of the periodic workload in msec
loops: number of iterations of the periodic workload tested
OR: number of instances of overrun as described above
OH: overhead as calculated above
SCHED_OTHER workload:
wload parameters ondemand interactive sched
run period loops OR OH OR OH OR OH
1 100 100 0 62.07% 0 100.02% 0 78.49%
10 1000 10 0 21.80% 0 22.74% 0 72.56%
1 10 1000 0 21.72% 0 63.08% 0 52.40%
10 100 100 0 8.09% 0 15.53% 0 17.33%
100 1000 10 0 1.83% 0 1.77% 0 0.29%
6 33 300 0 15.32% 0 8.60% 0 17.34%
66 333 30 0 0.79% 0 3.18% 0 12.26%
4 10 1000 0 5.87% 0 10.21% 0 6.15%
40 100 100 0 0.41% 0 0.04% 0 2.68%
400 1000 10 0 0.42% 0 0.50% 0 1.22%
5 9 1000 2 3.82% 1 6.10% 0 2.51%
50 90 100 0 0.19% 0 0.05% 0 1.71%
500 900 10 0 0.37% 0 0.38% 0 1.82%
9 12 1000 6 1.79% 1 0.77% 0 0.26%
90 120 100 0 0.16% 1 0.05% 0 0.49%
900 1200 10 0 0.09% 0 0.26% 0 0.62%
SCHED_FIFO workload:
wload parameters ondemand interactive sched
run period loops OR OH OR OH OR OH
1 100 100 0 39.61% 0 100.49% 0 99.57%
10 1000 10 0 73.51% 0 21.09% 0 96.66%
1 10 1000 0 18.01% 0 61.46% 0 67.68%
10 100 100 0 31.31% 0 18.62% 0 77.01%
100 1000 10 0 58.80% 0 1.90% 0 15.40%
6 33 300 251 85.99% 0 9.20% 1 30.09%
66 333 30 24 84.03% 0 3.38% 0 33.23%
4 10 1000 0 6.23% 0 12.21% 10 11.54%
40 100 100 100 62.08% 0 0.11% 1 11.85%
400 1000 10 10 62.09% 0 0.51% 0 7.00%
5 9 1000 999 12.29% 1 6.03% 0 0.04%
50 90 100 99 61.47% 0 0.05% 2 6.53%
500 900 10 10 43.37% 0 0.39% 0 6.30%
9 12 1000 999 9.83% 0 0.01% 14 1.69%
90 120 100 99 61.47% 0 0.01% 28 2.29%
900 1200 10 10 43.31% 0 0.22% 0 2.15%
Note that at this point RT CPU capacity is measured via rt_avg. For
the above results sched_time_avg_ms has been set to 50ms.
Known issues:
- More testing with real world type workloads, such as UI workloads and
benchmarks, is required.
- The power side of the characterization is in progress.
- Deadline scheduling class does not yet make CPU capacity requests.
- Not sure what's going on yet with the ondemand numbers above, it seems like
there may a regression with ondemand and RT tasks.
Dependencies:
Frequency invariant load tracking is required. For heterogeneous
systems such as big.Little, CPU invariant load tracking is required as
well. The required support for ARM platforms along with a patch
creating tracepoints for cpufreq_sched is located in [5].
References:
[0] http://article.gmane.org/gmane.linux.kernel/1499836
[1] http://thread.gmane.org/gmane.linux.power-management.general/69176
[2] https://git.linaro.org/power/rt-app.git
[3] https://lkml.org/lkml/2015/10/28/782
[4] https://git.linaro.org/people/steve.muckle/kernel.git/shortlog/refs/heads/interactive
[5] https://git.linaro.org/people/steve.muckle/kernel.git/shortlog/refs/heads/sched-freq-rfcv7
Juri Lelli (3):
sched/fair: add triggers for OPP change requests
sched/{core,fair}: trigger OPP change request on fork()
sched/fair: cpufreq_sched triggers for load balancing
Michael Turquette (2):
cpufreq: introduce cpufreq_driver_is_slow
sched: scheduler-driven cpu frequency selection
Morten Rasmussen (1):
sched: Compute cpu capacity available at current frequency
Steve Muckle (1):
sched/fair: jump to max OPP when crossing UP threshold
Vincent Guittot (3):
sched: remove call of sched_avg_update from sched_rt_avg_update
sched/deadline: split rt_avg in 2 distincts metrics
sched: rt scheduler sets capacity requirement
drivers/cpufreq/Kconfig | 21 ++
drivers/cpufreq/cpufreq.c | 6 +
include/linux/cpufreq.h | 12 ++
include/linux/sched.h | 8 +
kernel/sched/Makefile | 1 +
kernel/sched/core.c | 43 +++-
kernel/sched/cpufreq_sched.c | 459 +++++++++++++++++++++++++++++++++++++++++++
kernel/sched/deadline.c | 2 +-
kernel/sched/fair.c | 108 +++++-----
kernel/sched/rt.c | 48 ++++-
kernel/sched/sched.h | 120 ++++++++++-
11 files changed, 777 insertions(+), 51 deletions(-)
create mode 100644 kernel/sched/cpufreq_sched.c
--
2.4.10
[toc] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 07/10] sched/fair: jump to max OPP when crossing UP threshold |
| Message-ID | <r59Xz-6rQ-7@gated-at.bofh.it> |
| In reply to | #1340127 |
Since the true utilization of a long running task is not detectable
while it is running and might be bigger than the current cpu capacity,
create the maximum cpu capacity head room by requesting the maximum
cpu capacity once the cpu usage plus the capacity margin exceeds the
current capacity. This is also done to try to harm the performance of
a task the least.
Original fair-class only version authored by Juri Lelli
<juri.lelli@arm.com>.
cc: Ingo Molnar <mingo@redhat.com>
cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Juri Lelli <juri.lelli@arm.com>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/core.c | 40 +++++++++++++++++++++++++++++++++++
kernel/sched/fair.c | 57 --------------------------------------------------
kernel/sched/sched.h | 59 ++++++++++++++++++++++++++++++++++++++++++++++++++++
3 files changed, 99 insertions(+), 57 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 86297a2..747a7af 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -3020,6 +3020,45 @@ unsigned long long task_sched_runtime(struct task_struct *p)
return ns;
}
+#ifdef CONFIG_CPU_FREQ_GOV_SCHED
+static unsigned long sum_capacity_reqs(unsigned long cfs_cap,
+ struct sched_capacity_reqs *scr)
+{
+ unsigned long total = cfs_cap + scr->rt;
+
+ total = total * capacity_margin;
+ total /= SCHED_CAPACITY_SCALE;
+ total += scr->dl;
+ return total;
+}
+
+static void sched_freq_tick(int cpu)
+{
+ struct sched_capacity_reqs *scr;
+ unsigned long capacity_orig, capacity_curr;
+
+ if (!sched_freq())
+ return;
+
+ capacity_orig = capacity_orig_of(cpu);
+ capacity_curr = capacity_curr_of(cpu);
+ if (capacity_curr == capacity_orig)
+ return;
+
+ /*
+ * To make free room for a task that is building up its "real"
+ * utilization and to harm its performance the least, request
+ * a jump to max OPP as soon as the margin of free capacity is
+ * impacted (specified by capacity_margin).
+ */
+ scr = &per_cpu(cpu_sched_capacity_reqs, cpu);
+ if (capacity_curr < sum_capacity_reqs(cpu_util(cpu), scr))
+ set_cfs_cpu_capacity(cpu, true, capacity_max);
+}
+#else
+static inline void sched_freq_tick(int cpu) { }
+#endif
+
/*
* This function gets called by the timer code, with HZ frequency.
* We call it with interrupts disabled.
@@ -3037,6 +3076,7 @@ void scheduler_tick(void)
curr->sched_class->task_tick(rq, curr, 0);
update_cpu_load_active(rq);
calc_global_load_tick(rq);
+ sched_freq_tick(cpu);
raw_spin_unlock(&rq->lock);
perf_event_task_tick();
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 5531513..cf7ae0a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4283,9 +4283,6 @@ static inline void hrtick_update(struct rq *rq)
}
#endif
-static unsigned long capacity_orig_of(int cpu);
-static int cpu_util(int cpu);
-
static void update_capacity_of(int cpu)
{
unsigned long req_cap;
@@ -4685,15 +4682,6 @@ static unsigned long target_load(int cpu, int type)
return max(rq->cpu_load[type-1], total);
}
-static unsigned long capacity_of(int cpu)
-{
- return cpu_rq(cpu)->cpu_capacity;
-}
-
-static unsigned long capacity_orig_of(int cpu)
-{
- return cpu_rq(cpu)->cpu_capacity_orig;
-}
static unsigned long cpu_avg_load_per_task(int cpu)
{
@@ -4863,17 +4851,6 @@ static long effective_load(struct task_group *tg, int cpu, long wl, long wg)
#endif
/*
- * Returns the current capacity of cpu after applying both
- * cpu and freq scaling.
- */
-static unsigned long capacity_curr_of(int cpu)
-{
- return cpu_rq(cpu)->cpu_capacity_orig *
- arch_scale_freq_capacity(NULL, cpu)
- >> SCHED_CAPACITY_SHIFT;
-}
-
-/*
* Detect M:N waker/wakee relationships via a switching-frequency heuristic.
* A waker of many should wake a different task than the one last awakened
* at a frequency roughly N times higher than one of its wakees. In order
@@ -5117,40 +5094,6 @@ done:
}
/*
- * cpu_util returns the amount of capacity of a CPU that is used by CFS
- * tasks. The unit of the return value must be the one of capacity so we can
- * compare the utilization with the capacity of the CPU that is available for
- * CFS task (ie cpu_capacity).
- *
- * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the
- * recent utilization of currently non-runnable tasks on a CPU. It represents
- * the amount of utilization of a CPU in the range [0..capacity_orig] where
- * capacity_orig is the cpu_capacity available at the highest frequency
- * (arch_scale_freq_capacity()).
- * The utilization of a CPU converges towards a sum equal to or less than the
- * current capacity (capacity_curr <= capacity_orig) of the CPU because it is
- * the running time on this CPU scaled by capacity_curr.
- *
- * Nevertheless, cfs_rq.avg.util_avg can be higher than capacity_curr or even
- * higher than capacity_orig because of unfortunate rounding in
- * cfs.avg.util_avg or just after migrating tasks and new task wakeups until
- * the average stabilizes with the new running time. We need to check that the
- * utilization stays within the range of [0..capacity_orig] and cap it if
- * necessary. Without utilization capping, a group could be seen as overloaded
- * (CPU0 utilization at 121% + CPU1 utilization at 80%) whereas CPU1 has 20% of
- * available capacity. We allow utilization to overshoot capacity_curr (but not
- * capacity_orig) as it useful for predicting the capacity required after task
- * migrations (scheduler-driven DVFS).
- */
-static int cpu_util(int cpu)
-{
- unsigned long util = cpu_rq(cpu)->cfs.avg.util_avg;
- unsigned long capacity = capacity_orig_of(cpu);
-
- return (util >= capacity) ? capacity : util;
-}
-
-/*
* select_task_rq_fair: Select target runqueue for the waking task in domains
* that have the 'sd_flag' flag set. In practice, this is SD_BALANCE_WAKE,
* SD_BALANCE_FORK, or SD_BALANCE_EXEC.
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 9c26be2..59747d8 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1385,7 +1385,66 @@ unsigned long arch_scale_cpu_capacity(struct sched_domain *sd, int cpu)
}
#endif
+#ifdef CONFIG_SMP
+static inline unsigned long capacity_of(int cpu)
+{
+ return cpu_rq(cpu)->cpu_capacity;
+}
+
+static inline unsigned long capacity_orig_of(int cpu)
+{
+ return cpu_rq(cpu)->cpu_capacity_orig;
+}
+
+/*
+ * cpu_util returns the amount of capacity of a CPU that is used by CFS
+ * tasks. The unit of the return value must be the one of capacity so we can
+ * compare the utilization with the capacity of the CPU that is available for
+ * CFS task (ie cpu_capacity).
+ *
+ * cfs_rq.avg.util_avg is the sum of running time of runnable tasks plus the
+ * recent utilization of currently non-runnable tasks on a CPU. It represents
+ * the amount of utilization of a CPU in the range [0..capacity_orig] where
+ * capacity_orig is the cpu_capacity available at the highest frequency
+ * (arch_scale_freq_capacity()).
+ * The utilization of a CPU converges towards a sum equal to or less than the
+ * current capacity (capacity_curr <= capacity_orig) of the CPU because it is
+ * the running time on this CPU scaled by capacity_curr.
+ *
+ * Nevertheless, cfs_rq.avg.util_avg can be higher than capacity_curr or even
+ * higher than capacity_orig because of unfortunate rounding in
+ * cfs.avg.util_avg or just after migrating tasks and new task wakeups until
+ * the average stabilizes with the new running time. We need to check that the
+ * utilization stays within the range of [0..capacity_orig] and cap it if
+ * necessary. Without utilization capping, a group could be seen as overloaded
+ * (CPU0 utilization at 121% + CPU1 utilization at 80%) whereas CPU1 has 20% of
+ * available capacity. We allow utilization to overshoot capacity_curr (but not
+ * capacity_orig) as it useful for predicting the capacity required after task
+ * migrations (scheduler-driven DVFS).
+ */
+static inline int cpu_util(int cpu)
+{
+ unsigned long util = cpu_rq(cpu)->cfs.avg.util_avg;
+ unsigned long capacity = capacity_orig_of(cpu);
+
+ return (util >= capacity) ? capacity : util;
+}
+
+/*
+ * Returns the current capacity of cpu after applying both
+ * cpu and freq scaling.
+ */
+static inline unsigned long capacity_curr_of(int cpu)
+{
+ return cpu_rq(cpu)->cpu_capacity_orig *
+ arch_scale_freq_capacity(NULL, cpu)
+ >> SCHED_CAPACITY_SHIFT;
+}
+
+#endif
+
#ifdef CONFIG_CPU_FREQ_GOV_SCHED
+#define capacity_max SCHED_CAPACITY_SCALE
extern unsigned int capacity_margin;
extern struct static_key __sched_freq;
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 09/10] sched/deadline: split rt_avg in 2 distincts metrics |
| Message-ID | <r59Xz-6rQ-9@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Vincent Guittot <vincent.guittot@linaro.org>
rt_avg monitors the average load of rt tasks, deadline tasks and
interruptions, when enabled. It's used to calculate the remaining
capacity for CFS tasks. We split rt_avg in 2 metrics, one for rt and
interruptions that keeps the name rt_avg and another one for deadline
tasks that will be named dl_avg.
Both values are still used to calculate the remaining capacity for cfs
task. But rt_avg is now also used to request capacity to the sched-freq
for the rt tasks.
As the irq time is accounted with rt tasks, it will be taken into account
in the request of capacity.
Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/core.c | 1 +
kernel/sched/deadline.c | 2 +-
kernel/sched/fair.c | 1 +
kernel/sched/sched.h | 8 +++++++-
4 files changed, 10 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 747a7af..12a4a3a 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -759,6 +759,7 @@ void sched_avg_update(struct rq *rq)
asm("" : "+rm" (rq->age_stamp));
rq->age_stamp += period;
rq->rt_avg /= 2;
+ rq->dl_avg /= 2;
}
}
diff --git a/kernel/sched/deadline.c b/kernel/sched/deadline.c
index cd64c97..87dcee3 100644
--- a/kernel/sched/deadline.c
+++ b/kernel/sched/deadline.c
@@ -747,7 +747,7 @@ static void update_curr_dl(struct rq *rq)
curr->se.exec_start = rq_clock_task(rq);
cpuacct_charge(curr, delta_exec);
- sched_rt_avg_update(rq, delta_exec);
+ sched_dl_avg_update(rq, delta_exec);
dl_se->runtime -= dl_se->dl_yielded ? 0 : delta_exec;
if (dl_runtime_exceeded(dl_se)) {
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index cf7ae0a..3a812fa 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -6278,6 +6278,7 @@ static unsigned long scale_rt_capacity(int cpu)
*/
age_stamp = READ_ONCE(rq->age_stamp);
avg = READ_ONCE(rq->rt_avg);
+ avg += READ_ONCE(rq->dl_avg);
delta = __rq_clock_broken(rq) - age_stamp;
if (unlikely(delta < 0))
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 3df21f2..ad6cc8b 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -644,7 +644,7 @@ struct rq {
struct list_head cfs_tasks;
- u64 rt_avg;
+ u64 rt_avg, dl_avg;
u64 age_stamp;
u64 idle_stamp;
u64 avg_idle;
@@ -1499,8 +1499,14 @@ static inline void sched_rt_avg_update(struct rq *rq, u64 rt_delta)
{
rq->rt_avg += rt_delta * arch_scale_freq_capacity(NULL, cpu_of(rq));
}
+
+static inline void sched_dl_avg_update(struct rq *rq, u64 dl_delta)
+{
+ rq->dl_avg += dl_delta * arch_scale_freq_capacity(NULL, cpu_of(rq));
+}
#else
static inline void sched_rt_avg_update(struct rq *rq, u64 rt_delta) { }
+static inline void sched_dl_avg_update(struct rq *rq, u64 dl_delta) { }
static inline void sched_avg_update(struct rq *rq) { }
#endif
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 10/10] sched: rt scheduler sets capacity requirement |
| Message-ID | <r59XA-6rQ-13@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Vincent Guittot <vincent.guittot@linaro.org>
RT tasks don't provide any running constraints like deadline ones
except their running priority. The only current usable input to
estimate the capacity needed by RT tasks is the rt_avg metric. We use
it to estimate the CPU capacity needed for the RT scheduler class.
In order to monitor the evolution for RT task load, we must
peridiocally check it during the tick.
Then, we use the estimated capacity of the last activity to estimate
the next one which can not be that accurate but is a good starting
point without any impact on the wake up path of RT tasks.
Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/rt.c | 48 +++++++++++++++++++++++++++++++++++++++++++++++-
1 file changed, 47 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/rt.c b/kernel/sched/rt.c
index 8ec86ab..da9086c 100644
--- a/kernel/sched/rt.c
+++ b/kernel/sched/rt.c
@@ -1426,6 +1426,41 @@ static void check_preempt_curr_rt(struct rq *rq, struct task_struct *p, int flag
#endif
}
+#ifdef CONFIG_CPU_FREQ_GOV_SCHED
+static void sched_rt_update_capacity_req(struct rq *rq)
+{
+ u64 total, used, age_stamp, avg;
+ s64 delta;
+
+ if (!sched_freq())
+ return;
+
+ sched_avg_update(rq);
+ /*
+ * Since we're reading these variables without serialization make sure
+ * we read them once before doing sanity checks on them.
+ */
+ age_stamp = READ_ONCE(rq->age_stamp);
+ avg = READ_ONCE(rq->rt_avg);
+ delta = rq_clock(rq) - age_stamp;
+
+ if (unlikely(delta < 0))
+ delta = 0;
+
+ total = sched_avg_period() + delta;
+
+ used = div_u64(avg, total);
+ if (unlikely(used > SCHED_CAPACITY_SCALE))
+ used = SCHED_CAPACITY_SCALE;
+
+ set_rt_cpu_capacity(rq->cpu, true, (unsigned long)(used));
+}
+#else
+static inline void sched_rt_update_capacity_req(struct rq *rq)
+{ }
+
+#endif
+
static struct sched_rt_entity *pick_next_rt_entity(struct rq *rq,
struct rt_rq *rt_rq)
{
@@ -1494,8 +1529,17 @@ pick_next_task_rt(struct rq *rq, struct task_struct *prev)
if (prev->sched_class == &rt_sched_class)
update_curr_rt(rq);
- if (!rt_rq->rt_queued)
+ if (!rt_rq->rt_queued) {
+ /*
+ * The next task to be picked on this rq will have a lower
+ * priority than rt tasks so we can spend some time to update
+ * the capacity used by rt tasks based on the last activity.
+ * This value will be the used as an estimation of the next
+ * activity.
+ */
+ sched_rt_update_capacity_req(rq);
return NULL;
+ }
put_prev_task(rq, prev);
@@ -2212,6 +2256,8 @@ static void task_tick_rt(struct rq *rq, struct task_struct *p, int queued)
update_curr_rt(rq);
+ sched_rt_update_capacity_req(rq);
+
watchdog(rq, p);
/*
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 08/10] sched: remove call of sched_avg_update from sched_rt_avg_update |
| Message-ID | <r59XA-6rQ-15@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Vincent Guittot <vincent.guittot@linaro.org>
rt_avg is only used to scale the available CPU's capacity for CFS
tasks. As the update of this scaling is done during periodic load
balance, we only have to ensure that sched_avg_update has been called
before any periodic load balancing. This requirement is already
fulfilled by __update_cpu_load so the call in sched_rt_avg_update,
which is part of the hotpath, is useless.
Signed-off-by: Vincent Guittot <vincent.guittot@linaro.org>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/sched.h | 1 -
1 file changed, 1 deletion(-)
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 59747d8..3df21f2 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1498,7 +1498,6 @@ static inline void set_dl_cpu_capacity(int cpu, bool request,
static inline void sched_rt_avg_update(struct rq *rq, u64 rt_delta)
{
rq->rt_avg += rt_delta * arch_scale_freq_capacity(NULL, cpu_of(rq));
- sched_avg_update(rq);
}
#else
static inline void sched_rt_avg_update(struct rq *rq, u64 rt_delta) { }
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r59XA-6rQ-21@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Michael Turquette <mturquette@baylibre.com>
Some architectures and platforms perform CPU frequency transitions
through a non-blocking method, while some might block or sleep. Even
when frequency transitions do not block or sleep they may be very slow.
This distinction is important when trying to change frequency from
a non-interruptible context in a scheduler hot path.
Describe this distinction with a cpufreq driver flag,
CPUFREQ_DRIVER_FAST. The default is to not have this flag set,
thus erring on the side of caution.
cpufreq_driver_is_slow() is also introduced in this patch. Setting
the above flag will allow this function to return false.
[smuckle@linaro.org: change flag/API to include drivers that are too
slow for scheduler hot paths, in addition to those that block/sleep]
Cc: Rafael J. Wysocki <rafael@kernel.org>
Cc: Viresh Kumar <viresh.kumar@linaro.org>
Signed-off-by: Michael Turquette <mturquette@baylibre.com>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
drivers/cpufreq/cpufreq.c | 6 ++++++
include/linux/cpufreq.h | 9 +++++++++
2 files changed, 15 insertions(+)
diff --git a/drivers/cpufreq/cpufreq.c b/drivers/cpufreq/cpufreq.c
index e979ec7..88e63ca 100644
--- a/drivers/cpufreq/cpufreq.c
+++ b/drivers/cpufreq/cpufreq.c
@@ -154,6 +154,12 @@ bool have_governor_per_policy(void)
}
EXPORT_SYMBOL_GPL(have_governor_per_policy);
+bool cpufreq_driver_is_slow(void)
+{
+ return !(cpufreq_driver->flags & CPUFREQ_DRIVER_FAST);
+}
+EXPORT_SYMBOL_GPL(cpufreq_driver_is_slow);
+
struct kobject *get_governor_parent_kobj(struct cpufreq_policy *policy)
{
if (have_governor_per_policy())
diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
index 88a4215..93e1c1c 100644
--- a/include/linux/cpufreq.h
+++ b/include/linux/cpufreq.h
@@ -160,6 +160,7 @@ u64 get_cpu_idle_time(unsigned int cpu, u64 *wall, int io_busy);
int cpufreq_get_policy(struct cpufreq_policy *policy, unsigned int cpu);
int cpufreq_update_policy(unsigned int cpu);
bool have_governor_per_policy(void);
+bool cpufreq_driver_is_slow(void);
struct kobject *get_governor_parent_kobj(struct cpufreq_policy *policy);
#else
static inline unsigned int cpufreq_get(unsigned int cpu)
@@ -316,6 +317,14 @@ struct cpufreq_driver {
*/
#define CPUFREQ_NEED_INITIAL_FREQ_CHECK (1 << 5)
+/*
+ * Indicates that it is safe to call cpufreq_driver_target from
+ * non-interruptable context in scheduler hot paths. Drivers must
+ * opt-in to this flag, as the safe default is that they might sleep
+ * or be too slow for hot path use.
+ */
+#define CPUFREQ_DRIVER_FAST (1 << 6)
+
int cpufreq_register_driver(struct cpufreq_driver *driver_data);
int cpufreq_unregister_driver(struct cpufreq_driver *driver_data);
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rafael@kernel.org> |
|---|---|
| Date | 2016-02-23 02:40 +0100 |
| Subject | Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r5a7i-6wg-29@gated-at.bofh.it> |
| In reply to | #1340136 |
On Tue, Feb 23, 2016 at 2:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote: > From: Michael Turquette <mturquette@baylibre.com> > > Some architectures and platforms perform CPU frequency transitions > through a non-blocking method, while some might block or sleep. Even > when frequency transitions do not block or sleep they may be very slow. > This distinction is important when trying to change frequency from > a non-interruptible context in a scheduler hot path. > > Describe this distinction with a cpufreq driver flag, > CPUFREQ_DRIVER_FAST. The default is to not have this flag set, > thus erring on the side of caution. > > cpufreq_driver_is_slow() is also introduced in this patch. Setting > the above flag will allow this function to return false. > > [smuckle@linaro.org: change flag/API to include drivers that are too > slow for scheduler hot paths, in addition to those that block/sleep] > > Cc: Rafael J. Wysocki <rafael@kernel.org> > Cc: Viresh Kumar <viresh.kumar@linaro.org> > Signed-off-by: Michael Turquette <mturquette@baylibre.com> > Signed-off-by: Steve Muckle <smuckle@linaro.org> Something more sophisticated than this is needed, because one driver may actually be able to do "fast" switching in some cases and may not be able to do that in other cases. For example, in the acpi-cpufreq case all depends on what's there in the ACPI tables. Thanks, Rafael
[toc] | [prev] | [next] | [standalone]
| From | Michael Turquette <mturquette@baylibre.com> |
|---|---|
| Date | 2016-02-26 02:10 +0100 |
| Subject | Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r6f4T-3Xh-47@gated-at.bofh.it> |
| In reply to | #1340158 |
Quoting Rafael J. Wysocki (2016-02-22 17:31:09) > On Tue, Feb 23, 2016 at 2:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote: > > From: Michael Turquette <mturquette@baylibre.com> > > > > Some architectures and platforms perform CPU frequency transitions > > through a non-blocking method, while some might block or sleep. Even > > when frequency transitions do not block or sleep they may be very slow. > > This distinction is important when trying to change frequency from > > a non-interruptible context in a scheduler hot path. > > > > Describe this distinction with a cpufreq driver flag, > > CPUFREQ_DRIVER_FAST. The default is to not have this flag set, > > thus erring on the side of caution. > > > > cpufreq_driver_is_slow() is also introduced in this patch. Setting > > the above flag will allow this function to return false. > > > > [smuckle@linaro.org: change flag/API to include drivers that are too > > slow for scheduler hot paths, in addition to those that block/sleep] > > > > Cc: Rafael J. Wysocki <rafael@kernel.org> > > Cc: Viresh Kumar <viresh.kumar@linaro.org> > > Signed-off-by: Michael Turquette <mturquette@baylibre.com> > > Signed-off-by: Steve Muckle <smuckle@linaro.org> > > Something more sophisticated than this is needed, because one driver > may actually be able to do "fast" switching in some cases and may not > be able to do that in other cases. Those drivers can set the flag dynamically when they probe based on their ACPI tables. > > For example, in the acpi-cpufreq case all depends on what's there in > the ACPI tables. It's all a moot point until the locking in cpufreq is changed. Until those changes are made it is a bad idea to call cpufreq_driver_target() from schedule() context, regardless of the underlying hardware, and all platforms should kick that work out to the kthread. Regards, Mike > > Thanks, > Rafael
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rjw@rjwysocki.net> |
|---|---|
| Date | 2016-02-26 02:20 +0100 |
| Subject | Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r6fey-40K-7@gated-at.bofh.it> |
| In reply to | #1343662 |
On Thursday, February 25, 2016 04:50:29 PM Michael Turquette wrote: > Quoting Rafael J. Wysocki (2016-02-22 17:31:09) > > On Tue, Feb 23, 2016 at 2:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote: > > > From: Michael Turquette <mturquette@baylibre.com> > > > > > > Some architectures and platforms perform CPU frequency transitions > > > through a non-blocking method, while some might block or sleep. Even > > > when frequency transitions do not block or sleep they may be very slow. > > > This distinction is important when trying to change frequency from > > > a non-interruptible context in a scheduler hot path. > > > > > > Describe this distinction with a cpufreq driver flag, > > > CPUFREQ_DRIVER_FAST. The default is to not have this flag set, > > > thus erring on the side of caution. > > > > > > cpufreq_driver_is_slow() is also introduced in this patch. Setting > > > the above flag will allow this function to return false. > > > > > > [smuckle@linaro.org: change flag/API to include drivers that are too > > > slow for scheduler hot paths, in addition to those that block/sleep] > > > > > > Cc: Rafael J. Wysocki <rafael@kernel.org> > > > Cc: Viresh Kumar <viresh.kumar@linaro.org> > > > Signed-off-by: Michael Turquette <mturquette@baylibre.com> > > > Signed-off-by: Steve Muckle <smuckle@linaro.org> > > > > Something more sophisticated than this is needed, because one driver > > may actually be able to do "fast" switching in some cases and may not > > be able to do that in other cases. > > Those drivers can set the flag dynamically when they probe based on > their ACPI tables. No, they can't. Being able to to the "fast" switching is a property of the policy and the driver together and it may change with CPU going online/offline. > > > > For example, in the acpi-cpufreq case all depends on what's there in > > the ACPI tables. > > It's all a moot point until the locking in cpufreq is changed. No, it isn't. Look at this, for example: https://patchwork.kernel.org/patch/8426741/ > Until those changes are made it is a bad idea to call cpufreq_driver_target() > from schedule() context, regardless of the underlying hardware, and all > platforms should kick that work out to the kthread. Calling cpufreq_driver_target() from the scheduler is a bad idea overall, not just because of the locking. But there are other ways to switch frequencies from scheduler paths. I run such code on my test box daily without any problems. Thanks, Rafael
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rafael@kernel.org> |
|---|---|
| Date | 2016-02-26 22:10 +0100 |
| Subject | Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r6xOa-G1-11@gated-at.bofh.it> |
| In reply to | #1343686 |
On Fri, Feb 26, 2016 at 7:55 PM, Michael Turquette <mturquette@baylibre.com> wrote: > Quoting Rafael J. Wysocki (2016-02-25 17:16:17) >> On Thursday, February 25, 2016 04:50:29 PM Michael Turquette wrote: >> > Quoting Rafael J. Wysocki (2016-02-22 17:31:09) >> > > On Tue, Feb 23, 2016 at 2:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote: >> > > For example, in the acpi-cpufreq case all depends on what's there in >> > > the ACPI tables. >> > >> > It's all a moot point until the locking in cpufreq is changed. >> >> No, it isn't. Look at this, for example: https://patchwork.kernel.org/patch/8426741/ > > Thanks for the pointer. Do you mind Cc'ing me on future versions? I'll do that. Thanks, Rafael
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-26 02:20 +0100 |
| Subject | Re: [RFCv7 PATCH 02/10] cpufreq: introduce cpufreq_driver_is_slow |
| Message-ID | <r6fez-40K-29@gated-at.bofh.it> |
| In reply to | #1343662 |
On 02/25/2016 04:50 PM, Michael Turquette wrote: >> > Something more sophisticated than this is needed, because one driver >> > may actually be able to do "fast" switching in some cases and may not >> > be able to do that in other cases. > > Those drivers can set the flag dynamically when they probe based on > their ACPI tables. I was thinking that the reference here was to a driver that may be able to do fast switching for some transitions and not for others, say perhaps depending on the current and target frequencies, or the state of the regulators, or other system conditions. Rafael has proposed a fast_switch() addition to the cpufreq API which currently returns void. Perhaps that could be extended to return success or failure from the driver. The driver aborts if it cannot complete the request atomically and quickly. The scheduler-driven governor could attempt a fast switch if the callback is installed (and the other criteria for the fast switch are met, such as not throttled etc, no request already in flight etc). If the fast switch aborts, fall back to the slow path. I suppose the governor could also just see if policy->cur has changed as opposed to checking cpufreq_driver_fast_switch's return value. But then we can't tell the difference between the fast transition failing because it must be re-attempted in the slow path, and the fast transition failing because of some other more serious reason. In the latter case the request should probably just be dropped rather than retried in the slow path.
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 06/10] sched/fair: cpufreq_sched triggers for load balancing |
| Message-ID | <r59XA-6rQ-27@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Juri Lelli <juri.lelli@arm.com>
As we don't trigger freq changes from {en,de}queue_task_fair() during load
balancing, we need to do explicitly so on load balancing paths.
[smuckle@linaro.org: move update_capacity_of calls so rq lock is held]
cc: Ingo Molnar <mingo@redhat.com>
cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Juri Lelli <juri.lelli@arm.com>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/fair.c | 21 ++++++++++++++++++++-
1 file changed, 20 insertions(+), 1 deletion(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index e7fab8f..5531513 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -6107,6 +6107,10 @@ static void attach_one_task(struct rq *rq, struct task_struct *p)
{
raw_spin_lock(&rq->lock);
attach_task(rq, p);
+ /*
+ * We want to potentially raise target_cpu's OPP.
+ */
+ update_capacity_of(cpu_of(rq));
raw_spin_unlock(&rq->lock);
}
@@ -6128,6 +6132,11 @@ static void attach_tasks(struct lb_env *env)
attach_task(env->dst_rq, p);
}
+ /*
+ * We want to potentially raise env.dst_cpu's OPP.
+ */
+ update_capacity_of(env->dst_cpu);
+
raw_spin_unlock(&env->dst_rq->lock);
}
@@ -7267,6 +7276,11 @@ more_balance:
* ld_moved - cumulative load moved across iterations
*/
cur_ld_moved = detach_tasks(&env);
+ /*
+ * We want to potentially lower env.src_cpu's OPP.
+ */
+ if (cur_ld_moved)
+ update_capacity_of(env.src_cpu);
/*
* We've detached some tasks from busiest_rq. Every
@@ -7631,8 +7645,13 @@ static int active_load_balance_cpu_stop(void *data)
schedstat_inc(sd, alb_count);
p = detach_one_task(&env);
- if (p)
+ if (p) {
schedstat_inc(sd, alb_pushed);
+ /*
+ * We want to potentially lower env.src_cpu's OPP.
+ */
+ update_capacity_of(env.src_cpu);
+ }
else
schedstat_inc(sd, alb_failed);
}
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests |
| Message-ID | <r59XA-6rQ-25@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Juri Lelli <juri.lelli@arm.com>
Each time a task is {en,de}queued we might need to adapt the current
frequency to the new usage. Add triggers on {en,de}queue_task_fair() for
this purpose. Only trigger a freq request if we are effectively waking up
or going to sleep. Filter out load balancing related calls to reduce the
number of triggers.
[smuckle@linaro.org: resolve merge conflicts, define task_new,
use renamed static key sched_freq]
cc: Ingo Molnar <mingo@redhat.com>
cc: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Juri Lelli <juri.lelli@arm.com>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
kernel/sched/fair.c | 49 +++++++++++++++++++++++++++++++++++++++++++++++--
1 file changed, 47 insertions(+), 2 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 3437e01..f1f00a4 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -4283,6 +4283,21 @@ static inline void hrtick_update(struct rq *rq)
}
#endif
+static unsigned long capacity_orig_of(int cpu);
+static int cpu_util(int cpu);
+
+static void update_capacity_of(int cpu)
+{
+ unsigned long req_cap;
+
+ if (!sched_freq())
+ return;
+
+ /* Convert scale-invariant capacity to cpu. */
+ req_cap = cpu_util(cpu) * SCHED_CAPACITY_SCALE / capacity_orig_of(cpu);
+ set_cfs_cpu_capacity(cpu, true, req_cap);
+}
+
/*
* The enqueue_task method is called before nr_running is
* increased. Here we update the fair scheduling stats and
@@ -4293,6 +4308,7 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
{
struct cfs_rq *cfs_rq;
struct sched_entity *se = &p->se;
+ int task_new = !(flags & ENQUEUE_WAKEUP);
for_each_sched_entity(se) {
if (se->on_rq)
@@ -4324,9 +4340,23 @@ enqueue_task_fair(struct rq *rq, struct task_struct *p, int flags)
update_cfs_shares(cfs_rq);
}
- if (!se)
+ if (!se) {
add_nr_running(rq, 1);
+ /*
+ * We want to potentially trigger a freq switch
+ * request only for tasks that are waking up; this is
+ * because we get here also during load balancing, but
+ * in these cases it seems wise to trigger as single
+ * request after load balancing is done.
+ *
+ * XXX: how about fork()? Do we need a special
+ * flag/something to tell if we are here after a
+ * fork() (wakeup_task_new)?
+ */
+ if (!task_new)
+ update_capacity_of(cpu_of(rq));
+ }
hrtick_update(rq);
}
@@ -4384,9 +4414,24 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
update_cfs_shares(cfs_rq);
}
- if (!se)
+ if (!se) {
sub_nr_running(rq, 1);
+ /*
+ * We want to potentially trigger a freq switch
+ * request only for tasks that are going to sleep;
+ * this is because we get here also during load
+ * balancing, but in these cases it seems wise to
+ * trigger as single request after load balancing is
+ * done.
+ */
+ if (task_sleep) {
+ if (rq->cfs.nr_running)
+ update_capacity_of(cpu_of(rq));
+ else if (sched_freq())
+ set_cfs_cpu_capacity(cpu_of(rq), false, 0);
+ }
+ }
hrtick_update(rq);
}
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | Ricky Liang <jcliang@chromium.org> |
|---|---|
| Date | 2016-03-01 08:00 +0100 |
| Subject | Re: [RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests |
| Message-ID | <r7MrM-5z7-3@gated-at.bofh.it> |
| In reply to | #1340138 |
Hi Steve,
On Tue, Feb 23, 2016 at 9:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> From: Juri Lelli <juri.lelli@arm.com>
>
> Each time a task is {en,de}queued we might need to adapt the current
> frequency to the new usage. Add triggers on {en,de}queue_task_fair() for
> this purpose. Only trigger a freq request if we are effectively waking up
> or going to sleep. Filter out load balancing related calls to reduce the
> number of triggers.
>
> [smuckle@linaro.org: resolve merge conflicts, define task_new,
> use renamed static key sched_freq]
>
> cc: Ingo Molnar <mingo@redhat.com>
> cc: Peter Zijlstra <peterz@infradead.org>
> Signed-off-by: Juri Lelli <juri.lelli@arm.com>
> Signed-off-by: Steve Muckle <smuckle@linaro.org>
> ---
> kernel/sched/fair.c | 49 +++++++++++++++++++++++++++++++++++++++++++++++--
> 1 file changed, 47 insertions(+), 2 deletions(-)
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 3437e01..f1f00a4 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -4283,6 +4283,21 @@ static inline void hrtick_update(struct rq *rq)
> }
> #endif
>
> +static unsigned long capacity_orig_of(int cpu);
> +static int cpu_util(int cpu);
> +
> +static void update_capacity_of(int cpu)
> +{
> + unsigned long req_cap;
> +
> + if (!sched_freq())
> + return;
> +
> + /* Convert scale-invariant capacity to cpu. */
> + req_cap = cpu_util(cpu) * SCHED_CAPACITY_SCALE / capacity_orig_of(cpu);
> + set_cfs_cpu_capacity(cpu, true, req_cap);
> +}
> +
The change hunks of this patch should probably all depend on
CONFIG_SMP as capacity_orig_of() and cpu_util() are only available
when CONFIG_SMP is enabled.
[snip...]
Thanks,
Ricky
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-03-03 05:00 +0100 |
| Subject | Re: [RFCv7 PATCH 04/10] sched/fair: add triggers for OPP change requests |
| Message-ID | <r8sAH-1aq-37@gated-at.bofh.it> |
| In reply to | #1346408 |
Hi Ricky, On 02/29/2016 10:51 PM, Ricky Liang wrote: > The change hunks of this patch should probably all depend on > CONFIG_SMP as capacity_orig_of() and cpu_util() are only available > when CONFIG_SMP is enabled. Yeah, I was deferring cleaning that up until there was more buy in on the overall solution. But it looks like we will be moving forward using Rafael's schedutil governor. The most recent posting of that is here: http://thread.gmane.org/gmane.linux.kernel/2166378 thanks, Steve
[toc] | [prev] | [next] | [standalone]
| From | Steve Muckle <steve.muckle@linaro.org> |
|---|---|
| Date | 2016-02-23 02:30 +0100 |
| Subject | [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection |
| Message-ID | <r59XB-6rQ-35@gated-at.bofh.it> |
| In reply to | #1340127 |
From: Michael Turquette <mturquette@baylibre.com>
Scheduler-driven CPU frequency selection hopes to exploit both
per-task and global information in the scheduler to improve frequency
selection policy, achieving lower power consumption, improved
responsiveness/performance, and less reliance on heuristics and
tunables. For further discussion on the motivation of this integration
see [0].
This patch implements a shim layer between the Linux scheduler and the
cpufreq subsystem. The interface accepts capacity requests from the
CFS, RT and deadline sched classes. The requests from each sched class
are summed on each CPU with a margin applied to the CFS and RT
capacity requests to provide some headroom. Deadline requests are
expected to be precise enough given their nature to not require
headroom. The maximum total capacity request for a CPU in a frequency
domain drives the requested frequency for that domain.
Policy is determined by both the sched classes and this shim layer.
Note that this algorithm is event-driven. There is no polling loop to
check cpu idle time nor any other method which is unsynchronized with
the scheduler, aside from an optional throttling mechanism.
Thanks to Juri Lelli <juri.lelli@arm.com> for contributing design ideas,
code and test results, and to Ricky Liang <jcliang@chromium.org>
for initialization and static key inc/dec fixes.
[0] http://article.gmane.org/gmane.linux.kernel/1499836
[smuckle@linaro.org: various additions and fixes, revised commit text]
CC: Ricky Liang <jcliang@chromium.org>
Signed-off-by: Michael Turquette <mturquette@baylibre.com>
Signed-off-by: Juri Lelli <juri.lelli@arm.com>
Signed-off-by: Steve Muckle <smuckle@linaro.org>
---
drivers/cpufreq/Kconfig | 21 ++
include/linux/cpufreq.h | 3 +
include/linux/sched.h | 8 +
kernel/sched/Makefile | 1 +
kernel/sched/cpufreq_sched.c | 459 +++++++++++++++++++++++++++++++++++++++++++
kernel/sched/sched.h | 51 +++++
6 files changed, 543 insertions(+)
create mode 100644 kernel/sched/cpufreq_sched.c
diff --git a/drivers/cpufreq/Kconfig b/drivers/cpufreq/Kconfig
index 659879a..82d1548 100644
--- a/drivers/cpufreq/Kconfig
+++ b/drivers/cpufreq/Kconfig
@@ -102,6 +102,14 @@ config CPU_FREQ_DEFAULT_GOV_CONSERVATIVE
Be aware that not all cpufreq drivers support the conservative
governor. If unsure have a look at the help section of the
driver. Fallback governor will be the performance governor.
+
+config CPU_FREQ_DEFAULT_GOV_SCHED
+ bool "sched"
+ select CPU_FREQ_GOV_SCHED
+ help
+ Use the CPUfreq governor 'sched' as default. This scales
+ cpu frequency using CPU utilization estimates from the
+ scheduler.
endchoice
config CPU_FREQ_GOV_PERFORMANCE
@@ -183,6 +191,19 @@ config CPU_FREQ_GOV_CONSERVATIVE
If in doubt, say N.
+config CPU_FREQ_GOV_SCHED
+ bool "'sched' cpufreq governor"
+ depends on CPU_FREQ
+ select CPU_FREQ_GOV_COMMON
+ select IRQ_WORK
+ help
+ 'sched' - this governor scales cpu frequency from the
+ scheduler as a function of cpu capacity utilization. It does
+ not evaluate utilization on a periodic basis (as ondemand
+ does) but instead is event-driven by the scheduler.
+
+ If in doubt, say N.
+
comment "CPU frequency scaling drivers"
config CPUFREQ_DT
diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
index 93e1c1c..ce8b895 100644
--- a/include/linux/cpufreq.h
+++ b/include/linux/cpufreq.h
@@ -495,6 +495,9 @@ extern struct cpufreq_governor cpufreq_gov_ondemand;
#elif defined(CONFIG_CPU_FREQ_DEFAULT_GOV_CONSERVATIVE)
extern struct cpufreq_governor cpufreq_gov_conservative;
#define CPUFREQ_DEFAULT_GOVERNOR (&cpufreq_gov_conservative)
+#elif defined(CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED)
+extern struct cpufreq_governor cpufreq_gov_sched;
+#define CPUFREQ_DEFAULT_GOVERNOR (&cpufreq_gov_sched)
#endif
/*********************************************************************
diff --git a/include/linux/sched.h b/include/linux/sched.h
index a292c4b..27a6cd8 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -937,6 +937,14 @@ enum cpu_idle_type {
#define SCHED_CAPACITY_SHIFT 10
#define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
+struct sched_capacity_reqs {
+ unsigned long cfs;
+ unsigned long rt;
+ unsigned long dl;
+
+ unsigned long total;
+};
+
/*
* Wake-queues are lists of tasks with a pending wakeup, whose
* callers have already marked the task as woken internally,
diff --git a/kernel/sched/Makefile b/kernel/sched/Makefile
index 6768797..90ed832 100644
--- a/kernel/sched/Makefile
+++ b/kernel/sched/Makefile
@@ -19,3 +19,4 @@ obj-$(CONFIG_SCHED_AUTOGROUP) += auto_group.o
obj-$(CONFIG_SCHEDSTATS) += stats.o
obj-$(CONFIG_SCHED_DEBUG) += debug.o
obj-$(CONFIG_CGROUP_CPUACCT) += cpuacct.o
+obj-$(CONFIG_CPU_FREQ_GOV_SCHED) += cpufreq_sched.o
diff --git a/kernel/sched/cpufreq_sched.c b/kernel/sched/cpufreq_sched.c
new file mode 100644
index 0000000..a113e4e
--- /dev/null
+++ b/kernel/sched/cpufreq_sched.c
@@ -0,0 +1,459 @@
+/*
+ * Copyright (C) 2015 Michael Turquette <mturquette@linaro.org>
+ * Copyright (C) 2015-2016 Steve Muckle <smuckle@linaro.org>
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License version 2 as
+ * published by the Free Software Foundation.
+ */
+
+#include <linux/cpufreq.h>
+#include <linux/module.h>
+#include <linux/kthread.h>
+#include <linux/percpu.h>
+#include <linux/irq_work.h>
+#include <linux/delay.h>
+#include <linux/string.h>
+
+#include "sched.h"
+
+struct static_key __read_mostly __sched_freq = STATIC_KEY_INIT_FALSE;
+static bool __read_mostly cpufreq_driver_slow;
+
+/*
+ * The number of enabled schedfreq policies is modified during GOV_START/STOP.
+ * It, along with whether the schedfreq static key is enabled, is protected by
+ * the gov_enable_lock.
+ */
+static int enabled_policies;
+static DEFINE_MUTEX(gov_enable_lock);
+
+#ifndef CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED
+static struct cpufreq_governor cpufreq_gov_sched;
+#endif
+
+/*
+ * Capacity margin added to CFS and RT capacity requests to provide
+ * some head room if task utilization further increases.
+ */
+unsigned int capacity_margin = 1280;
+
+static DEFINE_PER_CPU(struct gov_data *, cpu_gov_data);
+DEFINE_PER_CPU(struct sched_capacity_reqs, cpu_sched_capacity_reqs);
+
+/**
+ * gov_data - per-policy data internal to the governor
+ * @throttle: next throttling period expiry. Derived from throttle_nsec
+ * @throttle_nsec: throttle period length in nanoseconds
+ * @task: worker thread for dvfs transition that may block/sleep
+ * @irq_work: callback used to wake up worker thread
+ * @policy: pointer to cpufreq policy associated with this governor data
+ * @fastpath_lock: prevents multiple CPUs in a frequency domain from racing
+ * with each other in fast path during calculation of domain frequency
+ * @slowpath_lock: mutex used to synchronize with slow path - ensure policy
+ * remains enabled, and eliminate racing between slow and fast path
+ * @enabled: boolean value indicating that the policy is started, protected
+ * by the slowpath_lock
+ * @requested_freq: last frequency requested by the sched governor
+ *
+ * struct gov_data is the per-policy cpufreq_sched-specific data
+ * structure. A per-policy instance of it is created when the
+ * cpufreq_sched governor receives the CPUFREQ_GOV_POLICY_INIT
+ * condition and a pointer to it exists in the gov_data member of
+ * struct cpufreq_policy.
+ */
+struct gov_data {
+ ktime_t throttle;
+ unsigned int throttle_nsec;
+ struct task_struct *task;
+ struct irq_work irq_work;
+ struct cpufreq_policy *policy;
+ raw_spinlock_t fastpath_lock;
+ struct mutex slowpath_lock;
+ unsigned int enabled;
+ unsigned int requested_freq;
+};
+
+static void cpufreq_sched_try_driver_target(struct cpufreq_policy *policy,
+ unsigned int freq)
+{
+ struct gov_data *gd = policy->governor_data;
+
+ __cpufreq_driver_target(policy, freq, CPUFREQ_RELATION_L);
+ gd->throttle = ktime_add_ns(ktime_get(), gd->throttle_nsec);
+}
+
+static bool finish_last_request(struct gov_data *gd)
+{
+ ktime_t now = ktime_get();
+
+ if (ktime_after(now, gd->throttle))
+ return false;
+
+ while (1) {
+ int usec_left = ktime_to_ns(ktime_sub(gd->throttle, now));
+
+ usec_left /= NSEC_PER_USEC;
+ usleep_range(usec_left, usec_left + 100);
+ now = ktime_get();
+ if (ktime_after(now, gd->throttle))
+ return true;
+ }
+}
+
+static int cpufreq_sched_thread(void *data)
+{
+ struct sched_param param;
+ struct gov_data *gd = (struct gov_data*) data;
+ struct cpufreq_policy *policy = gd->policy;
+ unsigned int new_request = 0;
+ unsigned int last_request = 0;
+ int ret;
+
+ param.sched_priority = 50;
+ ret = sched_setscheduler_nocheck(gd->task, SCHED_FIFO, ¶m);
+ if (ret) {
+ pr_warn("%s: failed to set SCHED_FIFO\n", __func__);
+ do_exit(-EINVAL);
+ } else {
+ pr_debug("%s: kthread (%d) set to SCHED_FIFO\n",
+ __func__, gd->task->pid);
+ }
+
+ mutex_lock(&gd->slowpath_lock);
+
+ while (true) {
+ set_current_state(TASK_INTERRUPTIBLE);
+ if (kthread_should_stop()) {
+ set_current_state(TASK_RUNNING);
+ break;
+ }
+ new_request = gd->requested_freq;
+ if (!gd->enabled || new_request == last_request) {
+ mutex_unlock(&gd->slowpath_lock);
+ schedule();
+ mutex_lock(&gd->slowpath_lock);
+ } else {
+ set_current_state(TASK_RUNNING);
+ /*
+ * if the frequency thread sleeps while waiting to be
+ * unthrottled, start over to check for a newer request
+ */
+ if (finish_last_request(gd))
+ continue;
+ last_request = new_request;
+ cpufreq_sched_try_driver_target(policy, new_request);
+ }
+ }
+
+ mutex_unlock(&gd->slowpath_lock);
+
+ return 0;
+}
+
+static void cpufreq_sched_irq_work(struct irq_work *irq_work)
+{
+ struct gov_data *gd;
+
+ gd = container_of(irq_work, struct gov_data, irq_work);
+ if (!gd)
+ return;
+
+ wake_up_process(gd->task);
+}
+
+static void update_fdomain_capacity_request(int cpu)
+{
+ unsigned int freq_new, index_new, cpu_tmp;
+ struct cpufreq_policy *policy;
+ struct gov_data *gd = per_cpu(cpu_gov_data, cpu);
+ unsigned long capacity = 0;
+
+ if (!gd)
+ return;
+
+ /* interrupts already disabled here via rq locked */
+ raw_spin_lock(&gd->fastpath_lock);
+
+ policy = gd->policy;
+
+ for_each_cpu(cpu_tmp, policy->cpus) {
+ struct sched_capacity_reqs *scr;
+
+ scr = &per_cpu(cpu_sched_capacity_reqs, cpu_tmp);
+ capacity = max(capacity, scr->total);
+ }
+
+ freq_new = capacity * policy->max >> SCHED_CAPACITY_SHIFT;
+
+ /*
+ * Calling this without locking policy->rwsem means we race
+ * against changes with policy->min and policy->max. This should
+ * be okay though.
+ */
+ if (cpufreq_frequency_table_target(policy, policy->freq_table,
+ freq_new, CPUFREQ_RELATION_L,
+ &index_new))
+ goto out;
+ freq_new = policy->freq_table[index_new].frequency;
+
+ if (freq_new == gd->requested_freq)
+ goto out;
+
+ gd->requested_freq = freq_new;
+
+ if (cpufreq_driver_slow || !mutex_trylock(&gd->slowpath_lock)) {
+ irq_work_queue_on(&gd->irq_work, cpu);
+ } else if (policy->transition_ongoing ||
+ ktime_before(ktime_get(), gd->throttle)) {
+ mutex_unlock(&gd->slowpath_lock);
+ irq_work_queue_on(&gd->irq_work, cpu);
+ } else {
+ cpufreq_sched_try_driver_target(policy, freq_new);
+ mutex_unlock(&gd->slowpath_lock);
+ }
+
+out:
+ raw_spin_unlock(&gd->fastpath_lock);
+}
+
+void update_cpu_capacity_request(int cpu, bool request)
+{
+ unsigned long new_capacity;
+ struct sched_capacity_reqs *scr;
+
+ /* The rq lock serializes access to the CPU's sched_capacity_reqs. */
+ lockdep_assert_held(&cpu_rq(cpu)->lock);
+
+ scr = &per_cpu(cpu_sched_capacity_reqs, cpu);
+
+ new_capacity = scr->cfs + scr->rt;
+ new_capacity = new_capacity * capacity_margin
+ / SCHED_CAPACITY_SCALE;
+ new_capacity += scr->dl;
+
+ if (new_capacity == scr->total)
+ return;
+
+ scr->total = new_capacity;
+ if (request)
+ update_fdomain_capacity_request(cpu);
+}
+
+static ssize_t show_throttle_nsec(struct cpufreq_policy *policy, char *buf)
+{
+ struct gov_data *gd = policy->governor_data;
+ return sprintf(buf, "%u\n", gd->throttle_nsec);
+}
+
+static ssize_t store_throttle_nsec(struct cpufreq_policy *policy,
+ const char *buf, size_t count)
+{
+ struct gov_data *gd = policy->governor_data;
+ unsigned int input;
+ int ret;
+
+ ret = sscanf(buf, "%u", &input);
+
+ if (ret != 1)
+ return -EINVAL;
+
+ gd->throttle_nsec = input;
+ return count;
+}
+
+static struct freq_attr sched_freq_throttle_nsec_attr =
+ __ATTR(throttle_nsec, 0644, show_throttle_nsec, store_throttle_nsec);
+
+static struct attribute *sched_freq_sysfs_attribs[] = {
+ &sched_freq_throttle_nsec_attr.attr,
+ NULL
+};
+
+static struct attribute_group sched_freq_sysfs_group = {
+ .attrs = sched_freq_sysfs_attribs,
+ .name = "sched_freq",
+};
+
+static int cpufreq_sched_policy_init(struct cpufreq_policy *policy)
+{
+ struct gov_data *gd;
+ int ret;
+
+ gd = kzalloc(sizeof(*gd), GFP_KERNEL);
+ if (!gd)
+ return -ENOMEM;
+ policy->governor_data = gd;
+ gd->policy = policy;
+ raw_spin_lock_init(&gd->fastpath_lock);
+ mutex_init(&gd->slowpath_lock);
+
+ ret = sysfs_create_group(&policy->kobj, &sched_freq_sysfs_group);
+ if (ret)
+ goto err_mem;
+
+ /*
+ * Set up schedfreq thread for slow path freq transitions if
+ * required by the driver.
+ */
+ if (cpufreq_driver_is_slow()) {
+ cpufreq_driver_slow = true;
+ gd->task = kthread_create(cpufreq_sched_thread, gd,
+ "kschedfreq:%d",
+ cpumask_first(policy->related_cpus));
+ if (IS_ERR_OR_NULL(gd->task)) {
+ pr_err("%s: failed to create kschedfreq thread\n",
+ __func__);
+ goto err_sysfs;
+ }
+ get_task_struct(gd->task);
+ kthread_bind_mask(gd->task, policy->related_cpus);
+ wake_up_process(gd->task);
+ init_irq_work(&gd->irq_work, cpufreq_sched_irq_work);
+ }
+ return 0;
+
+err_sysfs:
+ sysfs_remove_group(&policy->kobj, &sched_freq_sysfs_group);
+err_mem:
+ policy->governor_data = NULL;
+ kfree(gd);
+ return -ENOMEM;
+}
+
+static int cpufreq_sched_policy_exit(struct cpufreq_policy *policy)
+{
+ struct gov_data *gd = policy->governor_data;
+
+ /* Stop the schedfreq thread associated with this policy. */
+ if (cpufreq_driver_slow) {
+ kthread_stop(gd->task);
+ put_task_struct(gd->task);
+ }
+ sysfs_remove_group(&policy->kobj, &sched_freq_sysfs_group);
+ policy->governor_data = NULL;
+ kfree(gd);
+ return 0;
+}
+
+static int cpufreq_sched_start(struct cpufreq_policy *policy)
+{
+ struct gov_data *gd = policy->governor_data;
+ int cpu;
+
+ /*
+ * The schedfreq static key is managed here so the global schedfreq
+ * lock must be taken - a per-policy lock such as policy->rwsem is
+ * not sufficient.
+ */
+ mutex_lock(&gov_enable_lock);
+
+ gd->enabled = 1;
+
+ /*
+ * Set up percpu information. Writing the percpu gd pointer will
+ * enable the fast path if the static key is already enabled.
+ */
+ for_each_cpu(cpu, policy->cpus) {
+ memset(&per_cpu(cpu_sched_capacity_reqs, cpu), 0,
+ sizeof(struct sched_capacity_reqs));
+ per_cpu(cpu_gov_data, cpu) = gd;
+ }
+
+ if (enabled_policies == 0)
+ static_key_slow_inc(&__sched_freq);
+ enabled_policies++;
+ mutex_unlock(&gov_enable_lock);
+
+ return 0;
+}
+
+static void dummy(void *info) {}
+
+static int cpufreq_sched_stop(struct cpufreq_policy *policy)
+{
+ struct gov_data *gd = policy->governor_data;
+ int cpu;
+
+ /*
+ * The schedfreq static key is managed here so the global schedfreq
+ * lock must be taken - a per-policy lock such as policy->rwsem is
+ * not sufficient.
+ */
+ mutex_lock(&gov_enable_lock);
+
+ /*
+ * The governor stop path may or may not hold policy->rwsem. There
+ * must be synchronization with the slow path however.
+ */
+ mutex_lock(&gd->slowpath_lock);
+
+ /*
+ * Stop new entries into the hot path for all CPUs. This will
+ * potentially affect other policies which are still running but
+ * this is an infrequent operation.
+ */
+ static_key_slow_dec(&__sched_freq);
+ enabled_policies--;
+
+ /*
+ * Ensure that all CPUs currently part of this policy are out
+ * of the hot path so that if this policy exits we can free gd.
+ */
+ preempt_disable();
+ smp_call_function_many(policy->cpus, dummy, NULL, true);
+ preempt_enable();
+
+ /*
+ * Other CPUs in other policies may still have the schedfreq
+ * static key enabled. The percpu gd is used to signal which
+ * CPUs are enabled in the sched gov during the hot path.
+ */
+ for_each_cpu(cpu, policy->cpus)
+ per_cpu(cpu_gov_data, cpu) = NULL;
+
+ /* Pause the slow path for this policy. */
+ gd->enabled = 0;
+
+ if (enabled_policies)
+ static_key_slow_inc(&__sched_freq);
+ mutex_unlock(&gd->slowpath_lock);
+ mutex_unlock(&gov_enable_lock);
+
+ return 0;
+}
+
+static int cpufreq_sched_setup(struct cpufreq_policy *policy,
+ unsigned int event)
+{
+ switch (event) {
+ case CPUFREQ_GOV_POLICY_INIT:
+ return cpufreq_sched_policy_init(policy);
+ case CPUFREQ_GOV_POLICY_EXIT:
+ return cpufreq_sched_policy_exit(policy);
+ case CPUFREQ_GOV_START:
+ return cpufreq_sched_start(policy);
+ case CPUFREQ_GOV_STOP:
+ return cpufreq_sched_stop(policy);
+ case CPUFREQ_GOV_LIMITS:
+ break;
+ }
+ return 0;
+}
+
+#ifndef CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED
+static
+#endif
+struct cpufreq_governor cpufreq_gov_sched = {
+ .name = "sched",
+ .governor = cpufreq_sched_setup,
+ .owner = THIS_MODULE,
+};
+
+static int __init cpufreq_sched_init(void)
+{
+ return cpufreq_register_governor(&cpufreq_gov_sched);
+}
+
+/* Try to make this the default governor */
+fs_initcall(cpufreq_sched_init);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 1d58387..17908dd 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1384,6 +1384,57 @@ unsigned long arch_scale_cpu_capacity(struct sched_domain *sd, int cpu)
}
#endif
+#ifdef CONFIG_CPU_FREQ_GOV_SCHED
+extern unsigned int capacity_margin;
+extern struct static_key __sched_freq;
+
+static inline bool sched_freq(void)
+{
+ return static_key_false(&__sched_freq);
+}
+
+DECLARE_PER_CPU(struct sched_capacity_reqs, cpu_sched_capacity_reqs);
+void update_cpu_capacity_request(int cpu, bool request);
+
+static inline void set_cfs_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{
+ if (per_cpu(cpu_sched_capacity_reqs, cpu).cfs != capacity) {
+ per_cpu(cpu_sched_capacity_reqs, cpu).cfs = capacity;
+ update_cpu_capacity_request(cpu, request);
+ }
+}
+
+static inline void set_rt_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{
+ if (per_cpu(cpu_sched_capacity_reqs, cpu).rt != capacity) {
+ per_cpu(cpu_sched_capacity_reqs, cpu).rt = capacity;
+ update_cpu_capacity_request(cpu, request);
+ }
+}
+
+static inline void set_dl_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{
+ if (per_cpu(cpu_sched_capacity_reqs, cpu).dl != capacity) {
+ per_cpu(cpu_sched_capacity_reqs, cpu).dl = capacity;
+ update_cpu_capacity_request(cpu, request);
+ }
+}
+#else
+static inline bool sched_freq(void) { return false; }
+static inline void set_cfs_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{ }
+static inline void set_rt_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{ }
+static inline void set_dl_cpu_capacity(int cpu, bool request,
+ unsigned long capacity)
+{ }
+#endif
+
static inline void sched_rt_avg_update(struct rq *rq, u64 rt_delta)
{
rq->rt_avg += rt_delta * arch_scale_freq_capacity(NULL, cpu_of(rq));
--
2.4.10
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rjw@rjwysocki.net> |
|---|---|
| Date | 2016-02-25 05:00 +0100 |
| Subject | Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection |
| Message-ID | <r5VfQ-6qt-5@gated-at.bofh.it> |
| In reply to | #1340141 |
Hi,
I promised a review and here it goes.
Let me focus on this one as the rest seems to depend on it.
On Monday, February 22, 2016 05:22:43 PM Steve Muckle wrote:
> From: Michael Turquette <mturquette@baylibre.com>
>
> Scheduler-driven CPU frequency selection hopes to exploit both
> per-task and global information in the scheduler to improve frequency
> selection policy, achieving lower power consumption, improved
> responsiveness/performance, and less reliance on heuristics and
> tunables. For further discussion on the motivation of this integration
> see [0].
>
> This patch implements a shim layer between the Linux scheduler and the
> cpufreq subsystem. The interface accepts capacity requests from the
> CFS, RT and deadline sched classes. The requests from each sched class
> are summed on each CPU with a margin applied to the CFS and RT
> capacity requests to provide some headroom. Deadline requests are
> expected to be precise enough given their nature to not require
> headroom. The maximum total capacity request for a CPU in a frequency
> domain drives the requested frequency for that domain.
>
> Policy is determined by both the sched classes and this shim layer.
>
> Note that this algorithm is event-driven. There is no polling loop to
> check cpu idle time nor any other method which is unsynchronized with
> the scheduler, aside from an optional throttling mechanism.
>
> Thanks to Juri Lelli <juri.lelli@arm.com> for contributing design ideas,
> code and test results, and to Ricky Liang <jcliang@chromium.org>
> for initialization and static key inc/dec fixes.
>
> [0] http://article.gmane.org/gmane.linux.kernel/1499836
>
> [smuckle@linaro.org: various additions and fixes, revised commit text]
Well, the changelog is still a bit terse in my view. It should at least
describe the design somewhat (mention the static keys and how they are
used etc) end explain why the things are done this way.
> CC: Ricky Liang <jcliang@chromium.org>
> Signed-off-by: Michael Turquette <mturquette@baylibre.com>
> Signed-off-by: Juri Lelli <juri.lelli@arm.com>
> Signed-off-by: Steve Muckle <smuckle@linaro.org>
> ---
> drivers/cpufreq/Kconfig | 21 ++
> include/linux/cpufreq.h | 3 +
> include/linux/sched.h | 8 +
> kernel/sched/Makefile | 1 +
> kernel/sched/cpufreq_sched.c | 459 +++++++++++++++++++++++++++++++++++++++++++
> kernel/sched/sched.h | 51 +++++
> 6 files changed, 543 insertions(+)
> create mode 100644 kernel/sched/cpufreq_sched.c
>
[cut]
> diff --git a/include/linux/cpufreq.h b/include/linux/cpufreq.h
> index 93e1c1c..ce8b895 100644
> --- a/include/linux/cpufreq.h
> +++ b/include/linux/cpufreq.h
> @@ -495,6 +495,9 @@ extern struct cpufreq_governor cpufreq_gov_ondemand;
> #elif defined(CONFIG_CPU_FREQ_DEFAULT_GOV_CONSERVATIVE)
> extern struct cpufreq_governor cpufreq_gov_conservative;
> #define CPUFREQ_DEFAULT_GOVERNOR (&cpufreq_gov_conservative)
> +#elif defined(CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED)
> +extern struct cpufreq_governor cpufreq_gov_sched;
> +#define CPUFREQ_DEFAULT_GOVERNOR (&cpufreq_gov_sched)
> #endif
>
> /*********************************************************************
> diff --git a/include/linux/sched.h b/include/linux/sched.h
> index a292c4b..27a6cd8 100644
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -937,6 +937,14 @@ enum cpu_idle_type {
> #define SCHED_CAPACITY_SHIFT 10
> #define SCHED_CAPACITY_SCALE (1L << SCHED_CAPACITY_SHIFT)
>
> +struct sched_capacity_reqs {
> + unsigned long cfs;
> + unsigned long rt;
> + unsigned long dl;
> +
> + unsigned long total;
> +};
Without a comment explaining what this represents it is quite hard to
decode it.
> +
> /*
> * Wake-queues are lists of tasks with a pending wakeup, whose
> * callers have already marked the task as woken internally,
> diff --git a/kernel/sched/Makefile b/kernel/sched/Makefile
> index 6768797..90ed832 100644
> --- a/kernel/sched/Makefile
> +++ b/kernel/sched/Makefile
> @@ -19,3 +19,4 @@ obj-$(CONFIG_SCHED_AUTOGROUP) += auto_group.o
> obj-$(CONFIG_SCHEDSTATS) += stats.o
> obj-$(CONFIG_SCHED_DEBUG) += debug.o
> obj-$(CONFIG_CGROUP_CPUACCT) += cpuacct.o
> +obj-$(CONFIG_CPU_FREQ_GOV_SCHED) += cpufreq_sched.o
> diff --git a/kernel/sched/cpufreq_sched.c b/kernel/sched/cpufreq_sched.c
> new file mode 100644
> index 0000000..a113e4e
> --- /dev/null
> +++ b/kernel/sched/cpufreq_sched.c
> @@ -0,0 +1,459 @@
> +/*
Any chance to add one sentence about what's in the file?
Besides, governors traditionally go to drivers/cpufreq. Why is this different?
> + * Copyright (C) 2015 Michael Turquette <mturquette@linaro.org>
> + * Copyright (C) 2015-2016 Steve Muckle <smuckle@linaro.org>
> + *
> + * This program is free software; you can redistribute it and/or modify
> + * it under the terms of the GNU General Public License version 2 as
> + * published by the Free Software Foundation.
> + */
> +
> +#include <linux/cpufreq.h>
> +#include <linux/module.h>
> +#include <linux/kthread.h>
> +#include <linux/percpu.h>
> +#include <linux/irq_work.h>
> +#include <linux/delay.h>
> +#include <linux/string.h>
> +
> +#include "sched.h"
> +
> +struct static_key __read_mostly __sched_freq = STATIC_KEY_INIT_FALSE;
Well, I'm not familiar with static keys and how they work, so you'll need to
explain this part to me. I'm assuming that there is some magic related to
the value of the key that changes set_*_cpu_capacity() into no-ops when the
key is 0, presumably by modifying kernel code.
So this is clever but tricky and my question here is why it is necessary.
For example, compared to the RCU-based approach I'm using, how much better
this is? Yes, there is some tiny overhead related to the checking if callbacks
are present and invoking them, but is it really so bad? Can you actually
measure the different in any realistic workload?
One thing I personally like in the RCU-based approach is its universality. The
callbacks may be installed by different entities in a uniform way: intel_pstate
can do that, the old governors can do that, my experimental schedutil code can
do that and your code could have done that too in principle. And this is very
nice, because it is a common underlying mechanism that can be used by everybody
regardless of their particular implementations on the other side.
Why would I want to use something different, then?
> +static bool __read_mostly cpufreq_driver_slow;
> +
> +/*
> + * The number of enabled schedfreq policies is modified during GOV_START/STOP.
> + * It, along with whether the schedfreq static key is enabled, is protected by
> + * the gov_enable_lock.
> + */
Well, it would be good to explain what the role of the number of enabled_policies is
at least briefly.
> +static int enabled_policies;
> +static DEFINE_MUTEX(gov_enable_lock);
> +
> +#ifndef CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED
> +static struct cpufreq_governor cpufreq_gov_sched;
> +#endif
I don't think you need the #ifndef any more after recent changes in linux-next.
> +
> +/*
> + * Capacity margin added to CFS and RT capacity requests to provide
> + * some head room if task utilization further increases.
> + */
OK, where does this number come from?
> +unsigned int capacity_margin = 1280;
> +
> +static DEFINE_PER_CPU(struct gov_data *, cpu_gov_data);
> +DEFINE_PER_CPU(struct sched_capacity_reqs, cpu_sched_capacity_reqs);
> +
> +/**
> + * gov_data - per-policy data internal to the governor
> + * @throttle: next throttling period expiry. Derived from throttle_nsec
> + * @throttle_nsec: throttle period length in nanoseconds
> + * @task: worker thread for dvfs transition that may block/sleep
> + * @irq_work: callback used to wake up worker thread
> + * @policy: pointer to cpufreq policy associated with this governor data
> + * @fastpath_lock: prevents multiple CPUs in a frequency domain from racing
> + * with each other in fast path during calculation of domain frequency
> + * @slowpath_lock: mutex used to synchronize with slow path - ensure policy
> + * remains enabled, and eliminate racing between slow and fast path
> + * @enabled: boolean value indicating that the policy is started, protected
> + * by the slowpath_lock
> + * @requested_freq: last frequency requested by the sched governor
> + *
> + * struct gov_data is the per-policy cpufreq_sched-specific data
> + * structure. A per-policy instance of it is created when the
> + * cpufreq_sched governor receives the CPUFREQ_GOV_POLICY_INIT
> + * condition and a pointer to it exists in the gov_data member of
> + * struct cpufreq_policy.
> + */
> +struct gov_data {
> + ktime_t throttle;
> + unsigned int throttle_nsec;
> + struct task_struct *task;
> + struct irq_work irq_work;
> + struct cpufreq_policy *policy;
> + raw_spinlock_t fastpath_lock;
> + struct mutex slowpath_lock;
> + unsigned int enabled;
> + unsigned int requested_freq;
> +};
> +
> +static void cpufreq_sched_try_driver_target(struct cpufreq_policy *policy,
> + unsigned int freq)
> +{
> + struct gov_data *gd = policy->governor_data;
> +
> + __cpufreq_driver_target(policy, freq, CPUFREQ_RELATION_L);
> + gd->throttle = ktime_add_ns(ktime_get(), gd->throttle_nsec);
> +}
> +
> +static bool finish_last_request(struct gov_data *gd)
> +{
> + ktime_t now = ktime_get();
> +
> + if (ktime_after(now, gd->throttle))
> + return false;
> +
> + while (1) {
I would write this as
do {
> + int usec_left = ktime_to_ns(ktime_sub(gd->throttle, now));
> +
> + usec_left /= NSEC_PER_USEC;
> + usleep_range(usec_left, usec_left + 100);
> + now = ktime_get();
} while (ktime_before(now, gd->throttle));
return true;
But maybe that's just me. :-)
> + if (ktime_after(now, gd->throttle))
> + return true;
> + }
> +}
I'm not a big fan of this throttling mechanism overall, but more about that later.
> +
> +static int cpufreq_sched_thread(void *data)
> +{
Now, what really is the advantage of having those extra threads vs using
workqueues?
I guess the underlying concern is that RT tasks may stall workqueues indefinitely
in theory and then the frequency won't be updated, but there's much more kernel
stuff run from workqueues and if that is starved, you won't get very far anyway.
If you take special measures to prevent frequency change requests from being
stalled by RT tasks, question is why are they so special? Aren't there any
other kernel activities that also should be protected from that and may be
more important than CPU frequency changes?
Plus if this really is the problem here, then it also affects the other cpufreq
governors, so maybe it should be solved for everybody in some common way?
> + struct sched_param param;
> + struct gov_data *gd = (struct gov_data*) data;
> + struct cpufreq_policy *policy = gd->policy;
> + unsigned int new_request = 0;
> + unsigned int last_request = 0;
> + int ret;
> +
> + param.sched_priority = 50;
> + ret = sched_setscheduler_nocheck(gd->task, SCHED_FIFO, ¶m);
> + if (ret) {
> + pr_warn("%s: failed to set SCHED_FIFO\n", __func__);
> + do_exit(-EINVAL);
> + } else {
> + pr_debug("%s: kthread (%d) set to SCHED_FIFO\n",
> + __func__, gd->task->pid);
> + }
> +
> + mutex_lock(&gd->slowpath_lock);
> +
> + while (true) {
> + set_current_state(TASK_INTERRUPTIBLE);
> + if (kthread_should_stop()) {
> + set_current_state(TASK_RUNNING);
> + break;
> + }
> + new_request = gd->requested_freq;
> + if (!gd->enabled || new_request == last_request) {
This generally is a mistake.
You can't assume that if you had requested a CPU to enter a specific P-state,
that state was actually entered. In the platform-coordinated case the hardware
(or firmware) may choose to ignore your request and you won't be told about
that.
For this reason, you generally need to make the request every time even if it
is identical to the previous one. Even in the one-CPU-per-policy case there
may be HW coordination you don't know about.
> + mutex_unlock(&gd->slowpath_lock);
> + schedule();
> + mutex_lock(&gd->slowpath_lock);
> + } else {
> + set_current_state(TASK_RUNNING);
> + /*
> + * if the frequency thread sleeps while waiting to be
> + * unthrottled, start over to check for a newer request
> + */
> + if (finish_last_request(gd))
> + continue;
> + last_request = new_request;
> + cpufreq_sched_try_driver_target(policy, new_request);
> + }
> + }
> +
> + mutex_unlock(&gd->slowpath_lock);
> +
> + return 0;
> +}
> +
> +static void cpufreq_sched_irq_work(struct irq_work *irq_work)
> +{
> + struct gov_data *gd;
> +
> + gd = container_of(irq_work, struct gov_data, irq_work);
> + if (!gd)
> + return;
> +
> + wake_up_process(gd->task);
I'm wondering what would be wrong with writing it as
if (gd)
wake_up_process(gd->task);
And can gd turn out to be NULL here in any case?
> +}
> +
> +static void update_fdomain_capacity_request(int cpu)
> +{
> + unsigned int freq_new, index_new, cpu_tmp;
> + struct cpufreq_policy *policy;
> + struct gov_data *gd = per_cpu(cpu_gov_data, cpu);
> + unsigned long capacity = 0;
> +
> + if (!gd)
> + return;
> +
Why is this check necessary?
> + /* interrupts already disabled here via rq locked */
> + raw_spin_lock(&gd->fastpath_lock);
Well, if you compare this with the one-CPU-per-policy path in my experimental
schedutil governor code with the "fast switch" patch on top, you'll notice that
it doesn't use any locks and/or atomic ops. That's very much on purpose and
here's where your whole gain from using static keys practically goes away.
> +
> + policy = gd->policy;
> +
> + for_each_cpu(cpu_tmp, policy->cpus) {
> + struct sched_capacity_reqs *scr;
> +
> + scr = &per_cpu(cpu_sched_capacity_reqs, cpu_tmp);
> + capacity = max(capacity, scr->total);
> + }
You could save a few cycles from this in the case when the policy is not
shared.
> +
> + freq_new = capacity * policy->max >> SCHED_CAPACITY_SHIFT;
Where does this formula come from?
> +
> + /*
> + * Calling this without locking policy->rwsem means we race
> + * against changes with policy->min and policy->max. This should
> + * be okay though.
> + */
> + if (cpufreq_frequency_table_target(policy, policy->freq_table,
> + freq_new, CPUFREQ_RELATION_L,
> + &index_new))
> + goto out;
__cpufreq_driver_target() will call this again, so isn't calling it here
a bit wasteful?
> + freq_new = policy->freq_table[index_new].frequency;
> +
> + if (freq_new == gd->requested_freq)
> + goto out;
> +
Again, the above generally is a mistake for reasons explained earlier.
> + gd->requested_freq = freq_new;
> +
> + if (cpufreq_driver_slow || !mutex_trylock(&gd->slowpath_lock)) {
This really doesn't look good to me.
Why is the mutex needed here in the first place? cpufreq_sched_stop() should
be able to make sure that this function won't be run again for the policy
without using this lock.
> + irq_work_queue_on(&gd->irq_work, cpu);
I hope that you are aware of the fact that irq_work_queue_on() explodes
on uniprocessor ARM32 if you run an SMP kernel on it?
And what happens in the !cpufreq_driver_is_slow() case when we don't
initialize the irq_work?
> + } else if (policy->transition_ongoing ||
> + ktime_before(ktime_get(), gd->throttle)) {
If this really runs in the scheduler paths, you don't want to have ktime_get()
here.
> + mutex_unlock(&gd->slowpath_lock);
> + irq_work_queue_on(&gd->irq_work, cpu);
Allright.
I think I have figured out how this is arranged, but I may be wrong. :-)
Here's my understanding of it. If we are throttled, we don't just skip the
request. Instead, we wake up the gd thread kind of in the hope that the
throttling may end when it actually wakes up. So in fact we poke at the
gd thread on a regular basis asking it "Are you still throttled?" and that
happens on every call from the scheduler until the throttling is over if
I'm not mistaken. This means that during throttling every call from the
scheduler generates an irq_work that wakes up the gd thread just to make it
check if it still is throttled and go to sleep again. Please tell me
that I haven't understood this correctly.
The above aside, I personally think that rate limitting should happen at the source
and not at the worker thread level. So if you're throttled, you should just
return immediately from this function without generating any more work. That,
BTW, is what the sampling rate in my code is for.
> + } else {
> + cpufreq_sched_try_driver_target(policy, freq_new);
Well, this is supposed to be the fast path AFAICS.
Did you actually look at what __cpufreq_driver_target() does in general?
Including the wait_event() in cpufreq_freq_transition_begin() to mention just
one suspicious thing? And how much overhead it generates in the most general
case?
No, running *that* from the fast path is not a good idea. Quite honestly,
you'd need a new driver callback and a new way to run it from the cpufreq core
to implement this in a reasonably efficient way.
> + mutex_unlock(&gd->slowpath_lock);
> + }
> +
> +out:
> + raw_spin_unlock(&gd->fastpath_lock);
> +}
> +
> +void update_cpu_capacity_request(int cpu, bool request)
> +{
> + unsigned long new_capacity;
> + struct sched_capacity_reqs *scr;
> +
> + /* The rq lock serializes access to the CPU's sched_capacity_reqs. */
> + lockdep_assert_held(&cpu_rq(cpu)->lock);
> +
> + scr = &per_cpu(cpu_sched_capacity_reqs, cpu);
> +
> + new_capacity = scr->cfs + scr->rt;
> + new_capacity = new_capacity * capacity_margin
> + / SCHED_CAPACITY_SCALE;
> + new_capacity += scr->dl;
Can you please explain the formula here?
> +
> + if (new_capacity == scr->total)
> + return;
> +
The same mistake as before.
> + scr->total = new_capacity;
> + if (request)
> + update_fdomain_capacity_request(cpu);
> +}
> +
> +static ssize_t show_throttle_nsec(struct cpufreq_policy *policy, char *buf)
> +{
> + struct gov_data *gd = policy->governor_data;
> + return sprintf(buf, "%u\n", gd->throttle_nsec);
> +}
> +
> +static ssize_t store_throttle_nsec(struct cpufreq_policy *policy,
> + const char *buf, size_t count)
> +{
> + struct gov_data *gd = policy->governor_data;
> + unsigned int input;
> + int ret;
> +
> + ret = sscanf(buf, "%u", &input);
> +
> + if (ret != 1)
> + return -EINVAL;
> +
> + gd->throttle_nsec = input;
> + return count;
> +}
> +
> +static struct freq_attr sched_freq_throttle_nsec_attr =
> + __ATTR(throttle_nsec, 0644, show_throttle_nsec, store_throttle_nsec);
> +
> +static struct attribute *sched_freq_sysfs_attribs[] = {
> + &sched_freq_throttle_nsec_attr.attr,
> + NULL
> +};
> +
> +static struct attribute_group sched_freq_sysfs_group = {
> + .attrs = sched_freq_sysfs_attribs,
> + .name = "sched_freq",
> +};
> +
> +static int cpufreq_sched_policy_init(struct cpufreq_policy *policy)
> +{
> + struct gov_data *gd;
> + int ret;
> +
> + gd = kzalloc(sizeof(*gd), GFP_KERNEL);
> + if (!gd)
> + return -ENOMEM;
> + policy->governor_data = gd;
> + gd->policy = policy;
> + raw_spin_lock_init(&gd->fastpath_lock);
> + mutex_init(&gd->slowpath_lock);
> +
> + ret = sysfs_create_group(&policy->kobj, &sched_freq_sysfs_group);
> + if (ret)
> + goto err_mem;
> +
> + /*
> + * Set up schedfreq thread for slow path freq transitions if
> + * required by the driver.
> + */
> + if (cpufreq_driver_is_slow()) {
> + cpufreq_driver_slow = true;
> + gd->task = kthread_create(cpufreq_sched_thread, gd,
> + "kschedfreq:%d",
> + cpumask_first(policy->related_cpus));
> + if (IS_ERR_OR_NULL(gd->task)) {
> + pr_err("%s: failed to create kschedfreq thread\n",
> + __func__);
> + goto err_sysfs;
> + }
> + get_task_struct(gd->task);
> + kthread_bind_mask(gd->task, policy->related_cpus);
> + wake_up_process(gd->task);
> + init_irq_work(&gd->irq_work, cpufreq_sched_irq_work);
> + }
> + return 0;
> +
> +err_sysfs:
> + sysfs_remove_group(&policy->kobj, &sched_freq_sysfs_group);
> +err_mem:
> + policy->governor_data = NULL;
> + kfree(gd);
> + return -ENOMEM;
> +}
> +
> +static int cpufreq_sched_policy_exit(struct cpufreq_policy *policy)
> +{
> + struct gov_data *gd = policy->governor_data;
> +
> + /* Stop the schedfreq thread associated with this policy. */
> + if (cpufreq_driver_slow) {
> + kthread_stop(gd->task);
> + put_task_struct(gd->task);
> + }
> + sysfs_remove_group(&policy->kobj, &sched_freq_sysfs_group);
> + policy->governor_data = NULL;
> + kfree(gd);
> + return 0;
> +}
> +
> +static int cpufreq_sched_start(struct cpufreq_policy *policy)
> +{
> + struct gov_data *gd = policy->governor_data;
> + int cpu;
> +
> + /*
> + * The schedfreq static key is managed here so the global schedfreq
> + * lock must be taken - a per-policy lock such as policy->rwsem is
> + * not sufficient.
> + */
> + mutex_lock(&gov_enable_lock);
> +
> + gd->enabled = 1;
> +
> + /*
> + * Set up percpu information. Writing the percpu gd pointer will
> + * enable the fast path if the static key is already enabled.
> + */
> + for_each_cpu(cpu, policy->cpus) {
> + memset(&per_cpu(cpu_sched_capacity_reqs, cpu), 0,
> + sizeof(struct sched_capacity_reqs));
> + per_cpu(cpu_gov_data, cpu) = gd;
> + }
> +
> + if (enabled_policies == 0)
> + static_key_slow_inc(&__sched_freq);
> + enabled_policies++;
> + mutex_unlock(&gov_enable_lock);
> +
> + return 0;
> +}
> +
> +static void dummy(void *info) {}
> +
> +static int cpufreq_sched_stop(struct cpufreq_policy *policy)
> +{
> + struct gov_data *gd = policy->governor_data;
> + int cpu;
> +
> + /*
> + * The schedfreq static key is managed here so the global schedfreq
> + * lock must be taken - a per-policy lock such as policy->rwsem is
> + * not sufficient.
> + */
> + mutex_lock(&gov_enable_lock);
> +
> + /*
> + * The governor stop path may or may not hold policy->rwsem. There
> + * must be synchronization with the slow path however.
> + */
> + mutex_lock(&gd->slowpath_lock);
> +
> + /*
> + * Stop new entries into the hot path for all CPUs. This will
> + * potentially affect other policies which are still running but
> + * this is an infrequent operation.
> + */
> + static_key_slow_dec(&__sched_freq);
> + enabled_policies--;
> +
> + /*
> + * Ensure that all CPUs currently part of this policy are out
> + * of the hot path so that if this policy exits we can free gd.
> + */
> + preempt_disable();
> + smp_call_function_many(policy->cpus, dummy, NULL, true);
> + preempt_enable();
I'm not sure how this works, can you please tell me?
> +
> + /*
> + * Other CPUs in other policies may still have the schedfreq
> + * static key enabled. The percpu gd is used to signal which
> + * CPUs are enabled in the sched gov during the hot path.
> + */
> + for_each_cpu(cpu, policy->cpus)
> + per_cpu(cpu_gov_data, cpu) = NULL;
> +
> + /* Pause the slow path for this policy. */
> + gd->enabled = 0;
> +
> + if (enabled_policies)
> + static_key_slow_inc(&__sched_freq);
> + mutex_unlock(&gd->slowpath_lock);
> + mutex_unlock(&gov_enable_lock);
> +
> + return 0;
> +}
> +
> +static int cpufreq_sched_setup(struct cpufreq_policy *policy,
> + unsigned int event)
> +{
> + switch (event) {
> + case CPUFREQ_GOV_POLICY_INIT:
> + return cpufreq_sched_policy_init(policy);
> + case CPUFREQ_GOV_POLICY_EXIT:
> + return cpufreq_sched_policy_exit(policy);
> + case CPUFREQ_GOV_START:
> + return cpufreq_sched_start(policy);
> + case CPUFREQ_GOV_STOP:
> + return cpufreq_sched_stop(policy);
> + case CPUFREQ_GOV_LIMITS:
> + break;
> + }
> + return 0;
> +}
> +
> +#ifndef CONFIG_CPU_FREQ_DEFAULT_GOV_SCHED
> +static
> +#endif
> +struct cpufreq_governor cpufreq_gov_sched = {
> + .name = "sched",
> + .governor = cpufreq_sched_setup,
> + .owner = THIS_MODULE,
> +};
> +
> +static int __init cpufreq_sched_init(void)
> +{
> + return cpufreq_register_governor(&cpufreq_gov_sched);
> +}
> +
> +/* Try to make this the default governor */
> +fs_initcall(cpufreq_sched_init);
I have no comments to the rest of the patch.
Thanks,
Rafael
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-02-25 10:30 +0100 |
| Subject | Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection |
| Message-ID | <r60pc-1Hh-15@gated-at.bofh.it> |
| In reply to | #1342784 |
On Thu, Feb 25, 2016 at 04:55:57AM +0100, Rafael J. Wysocki wrote: > Well, I'm not familiar with static keys and how they work, so you'll need to > explain this part to me. See include/linux/jump_label.h, it has lots of text on them. There is also Documentation/static-keys.txt
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rjw@rjwysocki.net> |
|---|---|
| Date | 2016-02-25 22:10 +0100 |
| Subject | Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection |
| Message-ID | <r6bkC-1f7-9@gated-at.bofh.it> |
| In reply to | #1343018 |
On Thursday, February 25, 2016 10:21:50 AM Peter Zijlstra wrote: > On Thu, Feb 25, 2016 at 04:55:57AM +0100, Rafael J. Wysocki wrote: > > Well, I'm not familiar with static keys and how they work, so you'll need to > > explain this part to me. > > See include/linux/jump_label.h, it has lots of text on them. There is > also Documentation/static-keys.txt Thanks for the pointers! It looks like the author of the $subject patch hasn't looked at the latter document lately. In any case, IMO this might be used to hide the cpufreq_update_util() call sites from the scheduler code in case no one has set anything via cpufreq_set_update_util_data() for any CPUs. That essentially is when things like the performance governor are in use, so may be worth doing, but that's your judgement call mostly. :-) Thanks, Rafael
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-02-25 10:30 +0100 |
| Subject | Re: [RFCv7 PATCH 03/10] sched: scheduler-driven cpu frequency selection |
| Message-ID | <r60pc-1Hh-31@gated-at.bofh.it> |
| In reply to | #1342784 |
On Thu, Feb 25, 2016 at 04:55:57AM +0100, Rafael J. Wysocki wrote:
> > +static void dummy(void *info) {}
> > +
> > +static int cpufreq_sched_stop(struct cpufreq_policy *policy)
> > +{
> > + struct gov_data *gd = policy->governor_data;
> > + int cpu;
> > +
> > + /*
> > + * The schedfreq static key is managed here so the global schedfreq
> > + * lock must be taken - a per-policy lock such as policy->rwsem is
> > + * not sufficient.
> > + */
> > + mutex_lock(&gov_enable_lock);
> > +
> > + /*
> > + * The governor stop path may or may not hold policy->rwsem. There
> > + * must be synchronization with the slow path however.
> > + */
> > + mutex_lock(&gd->slowpath_lock);
> > +
> > + /*
> > + * Stop new entries into the hot path for all CPUs. This will
> > + * potentially affect other policies which are still running but
> > + * this is an infrequent operation.
> > + */
> > + static_key_slow_dec(&__sched_freq);
> > + enabled_policies--;
> > +
> > + /*
> > + * Ensure that all CPUs currently part of this policy are out
> > + * of the hot path so that if this policy exits we can free gd.
> > + */
> > + preempt_disable();
> > + smp_call_function_many(policy->cpus, dummy, NULL, true);
> > + preempt_enable();
>
> I'm not sure how this works, can you please tell me?
I think it relies on the fact that rq->lock disables IRQs, so if we've
managed to IPI all relevant CPUs, it means they cannot be inside a
rq->lock section.
Its vile though; one should not spray IPIs if one can avoid it. Such
things are much better done with RCU. Sure sync_sched() takes a little
longer, but this isn't a fast path by any measure.
> > +
> > + /*
> > + * Other CPUs in other policies may still have the schedfreq
> > + * static key enabled. The percpu gd is used to signal which
> > + * CPUs are enabled in the sched gov during the hot path.
> > + */
> > + for_each_cpu(cpu, policy->cpus)
> > + per_cpu(cpu_gov_data, cpu) = NULL;
> > +
> > + /* Pause the slow path for this policy. */
> > + gd->enabled = 0;
> > +
> > + if (enabled_policies)
> > + static_key_slow_inc(&__sched_freq);
> > + mutex_unlock(&gd->slowpath_lock);
> > + mutex_unlock(&gov_enable_lock);
> > +
> > + return 0;
> > +}
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web