Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1391619 > unrolled thread
| Started by | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| First post | 2016-04-30 14:50 +0200 |
| Last post | 2016-05-02 17:20 +0200 |
| Articles | 20 on this page of 46 — 7 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-04-30 14:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 09:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-01 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 11:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-07 11:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-08 10:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 04:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 05:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 09:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-10 09:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-10 09:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-10 17:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 05:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 06:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 11:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 12:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 06:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 10:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 17:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-02 17:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 16:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-03 17:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 12:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-06 00:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-06 21:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-09 10:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-04 19:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-05 11:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-05 16:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-06 19:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 19:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Ingo Molnar <mingo@kernel.org> - 2016-05-02 18:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 13:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 20:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:20 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-09 06:20 +0200 |
| Message-ID | <rwKPM-4Mr-7@gated-at.bofh.it> |
| In reply to | #1396570 |
On Mon, May 09, 2016 at 05:52:51AM +0200, Mike Galbraith wrote: > On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote: > > > In addition, I would argue maybe beefing up idle balancing is a more > > productive way to spread load, as work-stealing just does what needs > > to be done. And seems it has been (sub-unconsciously) neglected in this > > case, :) > > P.S. Nope, I'm dinging up multiple spots ;-) You bet, :)
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-02 10:50 +0200 |
| Message-ID | <ruhIe-dd-5@gated-at.bofh.it> |
| In reply to | #1391757 |
On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> Nah, tbench is just variance prone. It got dinged up at clients=cores
> on my desktop box, on 4 sockets the high end got seriously dinged up.
Ha!, check this:
root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo
NO_ORDER_IDLE > /debug/sched_features ; echo IDLE_CORE >
/debug/sched_features ; echo NO_FORCE_CORE > /debug/sched_features ;
tbench 20 -t 10
Throughput 5956.32 MB/sec 20 clients 20 procs max_latency=0.126 ms
root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo ORDER_IDLE >
/debug/sched_features ; echo IDLE_CORE > /debug/sched_features ; echo
NO_FORCE_CORE > /debug/sched_features ; tbench 20 -t 10
Throughput 5011.86 MB/sec 20 clients 20 procs max_latency=0.116 ms
That little ORDER_IDLE thing hurts silly. That's a little patch I had
lying about because some people complained that tasks hop around the
cache domain, instead of being stuck to a CPU.
I suspect what happens is that by all CPUs starting to look for idle at
the same place (the first cpu in the domain) they all find the same idle
cpu and things pile up.
The old behaviour, where they all start iterating from where they were
avoids some of that, at the cost of making tasks hop around.
Lets see if I can get the same behaviour out of the cpumask iteration
code..
---
kernel/sched/fair.c | 218 +++++++++++++++++++++++++++++++++++++----------
kernel/sched/features.h | 4 +
kernel/sched/idle_task.c | 4 +-
kernel/sched/sched.h | 1 +
kernel/time/tick-sched.c | 10 +--
5 files changed, 188 insertions(+), 49 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33ab..b7626a4 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ balance:
* One idle CPU per node is evaluated for a task numa move.
* Call select_idle_sibling to maybe find a better one.
*/
- if (!cur)
+ if (!cur) {
+ // XXX borken
env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+ }
assign:
assigned = true;
@@ -4491,6 +4493,17 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
}
#ifdef CONFIG_SMP
+
+/*
+ * Working cpumask for:
+ * load_balance,
+ * load_balance_newidle,
+ * select_idle_core.
+ *
+ * Assumes softirqs are disabled when in use.
+ */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+
#ifdef CONFIG_NO_HZ_COMMON
/*
* per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5175,187 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
}
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return false;
+
+ // XXX static key for !SMT topologies
+
+ return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+ int core = cpu_of(rq);
+ int cpu;
+
+ rcu_read_lock();
+ if (test_idle_cores(core))
+ goto unlock;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ if (cpu == core)
+ continue;
+
+ if (!idle_cpu(cpu))
+ goto unlock;
+ }
+
+ set_idle_cores(core);
+unlock:
+ rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, int target)
+{
+ struct cpumask *cpus = this_cpu_cpumask_var_ptr(load_balance_mask);
+ struct sched_domain *sd;
+ int core, cpu;
+
+ sd = rcu_dereference(per_cpu(sd_llc, target));
+ cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+ for_each_cpu(core, cpus) {
+ bool idle = true;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ cpumask_clear_cpu(cpu, cpus);
+ if (!idle_cpu(cpu))
+ idle = false;
+ }
+
+ if (idle)
+ break;
+ }
+
+ return core;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+static inline void clear_idle_cores(int cpu) { }
+static inline void set_idle_cores(int cpu) { }
+
+static inline bool test_idle_cores(int cpu)
+{
+ return false;
+}
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, int target)
+{
+ return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
/*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
*/
static int select_idle_sibling(struct task_struct *p, int target)
{
struct sched_domain *sd;
- struct sched_group *sg;
int i = task_cpu(p);
if (idle_cpu(target))
return target;
/*
- * If the prevous cpu is cache affine and idle, don't be stupid.
+ * If the previous cpu is cache affine and idle, don't be stupid.
*/
if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
return i;
- /*
- * Otherwise, iterate the domains and find an eligible idle cpu.
- *
- * A completely idle sched group at higher domains is more
- * desirable than an idle group at a lower level, because lower
- * domains have smaller groups and usually share hardware
- * resources which causes tasks to contend on them, e.g. x86
- * hyperthread siblings in the lowest domain (SMT) can contend
- * on the shared cpu pipeline.
- *
- * However, while we prefer idle groups at higher domains
- * finding an idle cpu at the lowest domain is still better than
- * returning 'target', which we've already established, isn't
- * idle.
- */
- sd = rcu_dereference(per_cpu(sd_llc, target));
- for_each_lower_domain(sd) {
- sg = sd->groups;
- do {
- if (!cpumask_intersects(sched_group_cpus(sg),
- tsk_cpus_allowed(p)))
- goto next;
+ i = target;
+ if (sched_feat(ORDER_IDLE))
+ i = per_cpu(sd_llc_id, target); /* first cpu in llc domain */
+ sd = rcu_dereference(per_cpu(sd_llc, i));
+ if (!sd)
+ return target;
+
+ if (sched_feat(OLD_IDLE)) {
+ struct sched_group *sg;
- /* Ensure the entire group is idle */
- for_each_cpu(i, sched_group_cpus(sg)) {
- if (i == target || !idle_cpu(i))
+ for_each_lower_domain(sd) {
+ sg = sd->groups;
+ do {
+ if (!cpumask_intersects(sched_group_cpus(sg),
+ tsk_cpus_allowed(p)))
goto next;
- }
- /*
- * It doesn't matter which cpu we pick, the
- * whole group is idle.
- */
- target = cpumask_first_and(sched_group_cpus(sg),
- tsk_cpus_allowed(p));
- goto done;
+ /* Ensure the entire group is idle */
+ for_each_cpu(i, sched_group_cpus(sg)) {
+ if (i == target || !idle_cpu(i))
+ goto next;
+ }
+
+ /*
+ * It doesn't matter which cpu we pick, the
+ * whole group is idle.
+ */
+ target = cpumask_first_and(sched_group_cpus(sg),
+ tsk_cpus_allowed(p));
+ goto done;
next:
- sg = sg->next;
- } while (sg != sd->groups);
- }
+ sg = sg->next;
+ } while (sg != sd->groups);
+ }
done:
+ return target;
+ }
+
+ /*
+ * If there are idle cores to be had, go find one.
+ */
+ if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
+ i = select_idle_core(p, target);
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
+
+ /*
+ * Failed to find an idle core; stop looking for one.
+ */
+ clear_idle_cores(target);
+ }
+
+ if (sched_feat(FORCE_CORE)) {
+ i = select_idle_core(p, target);
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
+ }
+
+ /*
+ * Otherwise, settle for anything idle in this cache domain.
+ */
+ for_each_cpu(i, sched_domain_span(sd)) {
+ if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+ continue;
+ if (idle_cpu(i))
+ return i;
+ }
+
return target;
}
@@ -7229,9 +7364,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
*/
#define MAX_PINNED_INTERVAL 512
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
static int need_active_balance(struct lb_env *env)
{
struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa..11ab7ab 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,7 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
SCHED_FEAT(LB_MIN, false)
SCHED_FEAT(ATTACH_AGE_LOAD, true)
+SCHED_FEAT(OLD_IDLE, false)
+SCHED_FEAT(ORDER_IDLE, false)
+SCHED_FEAT(IDLE_CORE, true)
+SCHED_FEAT(FORCE_CORE, false)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce949..cb394db 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
resched_curr(rq);
}
+extern void update_idle_core(struct rq *rq);
+
static struct task_struct *
pick_next_task_idle(struct rq *rq, struct task_struct *prev)
{
put_prev_task(rq, prev);
-
+ update_idle_core(rq);
schedstat_inc(rq, sched_goidle);
return rq->idle;
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fc..5994794 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
* Number of busy cpus in this group.
*/
atomic_t nr_busy_cpus;
+ int has_idle_cores;
unsigned long cpumask[0]; /* iteration mask */
};
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc..6e42cd2 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
WARN_ON_ONCE(irqs_disabled());
/*
- * Update the idle state in the scheduler domain hierarchy
- * when tick_nohz_stop_sched_tick() is called from the idle loop.
- * State will be updated to busy during the first busy tick after
- * exiting idle.
- */
+ * Update the idle state in the scheduler domain hierarchy
+ * when tick_nohz_stop_sched_tick() is called from the idle loop.
+ * State will be updated to busy during the first busy tick after
+ * exiting idle.
+ */
set_cpu_sd_state_idle();
local_irq_disable();
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-02 17:00 +0200 |
| Message-ID | <runuj-5gi-33@gated-at.bofh.it> |
| In reply to | #1392067 |
On Mon, 2016-05-02 at 10:46 +0200, Peter Zijlstra wrote:
> On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
>
> > Nah, tbench is just variance prone. It got dinged up at clients=cores
> > on my desktop box, on 4 sockets the high end got seriously dinged up.
>
>
> Ha!, check this:
>
> root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo
> NO_ORDER_IDLE > /debug/sched_features ; echo IDLE_CORE >
> /debug/sched_features ; echo NO_FORCE_CORE > /debug/sched_features ;
> tbench 20 -t 10
>
> Throughput 5956.32 MB/sec 20 clients 20 procs max_latency=0.126 ms
>
>
> root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo ORDER_IDLE >
> /debug/sched_features ; echo IDLE_CORE > /debug/sched_features ; echo
> NO_FORCE_CORE > /debug/sched_features ; tbench 20 -t 10
>
> Throughput 5011.86 MB/sec 20 clients 20 procs max_latency=0.116 ms
>
>
>
> That little ORDER_IDLE thing hurts silly. That's a little patch I had
> lying about because some people complained that tasks hop around the
> cache domain, instead of being stuck to a CPU.
>
> I suspect what happens is that by all CPUs starting to look for idle at
> the same place (the first cpu in the domain) they all find the same idle
> cpu and things pile up.
>
> The old behaviour, where they all start iterating from where they were
> avoids some of that, at the cost of making tasks hop around.
>
> Lets see if I can get the same behaviour out of the cpumask iteration
> code..
Order is one thing, but what the old behavior does first and foremost
is when the box starts getting really busy, only looking at target's
sibling shuts select_idle_sibling() down instead of letting it wreck
things. Once cores are moving, there are no large piles of anything
left to collect other than pain.
We really need a good way to know we're not gonna turn the box into a
shredder. The wake_wide() thing might help some, likely wants some
twiddling, in_interrupt() might be another time to try hard.
Anyway, the has_idle_cores business seems to shut select_idle_sibling()
down rather nicely when the the box gets busy. Forcing either core,
target's sibling or go fish turned in a top end win on 48 rq/socket.
Oh btw, did you know single socket boxen have no sd_busy? That doesn't
look right.
fromm:~/:[0]# for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done
Throughput 511.016 MB/sec 1 clients 1 procs max_latency=0.113 ms
Throughput 1042.03 MB/sec 2 clients 2 procs max_latency=0.098 ms
Throughput 1953.12 MB/sec 4 clients 4 procs max_latency=0.236 ms
Throughput 3694.99 MB/sec 8 clients 8 procs max_latency=0.308 ms
Throughput 7080.95 MB/sec 16 clients 16 procs max_latency=0.442 ms
Throughput 13444.7 MB/sec 32 clients 32 procs max_latency=1.417 ms
Throughput 20191.3 MB/sec 64 clients 64 procs max_latency=4.554 ms
Throughput 41115.4 MB/sec 128 clients 128 procs max_latency=13.414 ms
Throughput 66844.4 MB/sec 256 clients 256 procs max_latency=50.069 ms
5226 /*
5227 * If there are idle cores to be had, go find one.
5228 */
5229 if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
5230 i = select_idle_core(p, target);
5231 if ((unsigned)i < nr_cpumask_bits)
5232 return i;
5233
5234 /*
5235 * Failed to find an idle core; stop looking for one.
5236 */
5237 clear_idle_cores(target);
5238 }
5239 #if 1
5240 for_each_cpu(i, cpu_smt_mask(target)) {
5241 if (idle_cpu(i))
5242 return i;
5243 }
5244
5245 return target;
5246 #endif
5247
5248 if (sched_feat(FORCE_CORE)) {
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-02 17:00 +0200 |
| Message-ID | <runuk-5gi-47@gated-at.bofh.it> |
| In reply to | #1392265 |
On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote: > Oh btw, did you know single socket boxen have no sd_busy? That doesn't > look right. I suspected; didn't bother looking at yet. The 'problem' is that the LLC domain is the top-most, so it doesn't have a parent domain. I'm sure we can come up with something if we can get this all working right. And yes, I can get gains on various workloads with various options, I can even break all workloads, but I've so far completely failed on getting a win for everyone :/ In particular low count sysbench-psql (oltp test) vs tbench client==nr_cores is having me flummoxed for a bit.
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-05-02 17:50 +0200 |
| Message-ID | <ruogG-66V-17@gated-at.bofh.it> |
| In reply to | #1392268 |
On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote: > On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote: > > Oh btw, did you know single socket boxen have no sd_busy? That doesn't > > look right. > > I suspected; didn't bother looking at yet. The 'problem' is that the LLC > domain is the top-most, so it doesn't have a parent domain. I'm sure we > can come up with something if we can get this all working right. > > And yes, I can get gains on various workloads with various options, I > can even break all workloads, but I've so far completely failed on > getting a win for everyone :/ Adding in the task_hot() check to decide if scanning idle was a good idea ended up being really important. I'm happy to try a few variations here as well, do you have a more recent patch? -chris
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-03 16:40 +0200 |
| Message-ID | <ruJEu-1mn-9@gated-at.bofh.it> |
| In reply to | #1392320 |
On Mon, May 02, 2016 at 11:47:25AM -0400, Chris Mason wrote:
> On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote:
> > On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> > > Oh btw, did you know single socket boxen have no sd_busy? That doesn't
> > > look right.
> >
> > I suspected; didn't bother looking at yet. The 'problem' is that the LLC
> > domain is the top-most, so it doesn't have a parent domain. I'm sure we
> > can come up with something if we can get this all working right.
> >
> > And yes, I can get gains on various workloads with various options, I
> > can even break all workloads, but I've so far completely failed on
> > getting a win for everyone :/
>
> Adding in the task_hot() check to decide if scanning idle was a good
> idea ended up being really important
So I'm conflicted on this patch:
+static int bounce_to_target(struct task_struct *p, int cpu)
+{
+ s64 delta;
+
+ /*
+ * as the run queue gets bigger, its more and more likely that
+ * balance will have distributed things for us, and less likely
+ * that scanning all our CPUs for an idle one will find one.
+ * So, if nr_running > 1, just call this CPU good enough
+ */
+ if (cpu_rq(cpu)->cfs.nr_running > 1)
+ return 1;
+
+ /* taken from task_hot() */
+ delta = rq_clock_task(task_rq(p)) - p->se.exec_start;
+ return delta < (s64)sysctl_sched_migration_cost;
+}
This will work for you schbench workload because it sleep for 30ms while
the migration_cost thingy is 500us, therefore you'll trigger the full
LLC scan.
_However_, the migration_cost is supposed the model the cost of leaving
the LLC, so testing against that here seems wrong.
Let me go play with something that measures the cost of doing that LLC
scan and compares that against the sleepy time -- of course, now need to
go figure out how to do this clock thing without rq-lock pain.
+ if (package_sd && !bounce_to_target(p, target)) {
+ for_each_cpu_and(i, sched_domain_span(package_sd), tsk_cpus_allowed(p)) {
+ if (idle_cpu(i)) {
+ target = i;
+ break;
+ }
+
+ }
+ }
Also note your s/sd/package_sd/ rename is, strictly speaking, wrong.
Sure, on your current Intel system the LLC is the entire package, but
this is not true in general.
Take for instance the Intel Core2Quad and AMD Bulldozer thingies, they
had two dies in one package, and correspondingly two LLC domains in one
package.
(also, the Intel cluster-on-die thing can split the thing in two)
There were also the old P6 era SMP boards which had external LLC, where
you could have an LLC shared across multiple packages -- although I'm
thinking we'll never see that again, due to off package being far
toooooo slooooooow these days.
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-05-03 17:20 +0200 |
| Message-ID | <ruKhb-1Y8-1@gated-at.bofh.it> |
| In reply to | #1393459 |
On Tue, May 03, 2016 at 04:32:25PM +0200, Peter Zijlstra wrote:
> On Mon, May 02, 2016 at 11:47:25AM -0400, Chris Mason wrote:
> > On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote:
> > > On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> > > > Oh btw, did you know single socket boxen have no sd_busy? That doesn't
> > > > look right.
> > >
> > > I suspected; didn't bother looking at yet. The 'problem' is that the LLC
> > > domain is the top-most, so it doesn't have a parent domain. I'm sure we
> > > can come up with something if we can get this all working right.
> > >
> > > And yes, I can get gains on various workloads with various options, I
> > > can even break all workloads, but I've so far completely failed on
> > > getting a win for everyone :/
> >
> > Adding in the task_hot() check to decide if scanning idle was a good
> > idea ended up being really important
>
> So I'm conflicted on this patch:
>
> +static int bounce_to_target(struct task_struct *p, int cpu)
> +{
> + s64 delta;
> +
> + /*
> + * as the run queue gets bigger, its more and more likely that
> + * balance will have distributed things for us, and less likely
> + * that scanning all our CPUs for an idle one will find one.
> + * So, if nr_running > 1, just call this CPU good enough
> + */
> + if (cpu_rq(cpu)->cfs.nr_running > 1)
> + return 1;
The nr_running check is interesting. It is supposed to give the same
benefit as your "do we have anything idle?" variable, but without having
to constantly update a variable somewhere. I'll have to do a few runs
to verify (maybe a idle_scan_failed counter).
> +
> + /* taken from task_hot() */
> + delta = rq_clock_task(task_rq(p)) - p->se.exec_start;
> + return delta < (s64)sysctl_sched_migration_cost;
> +}
>
> This will work for you schbench workload because it sleep for 30ms while
> the migration_cost thingy is 500us, therefore you'll trigger the full
> LLC scan.
The task_hot checks don't do much for the sleeping schbench runs, but
they help a lot for this:
# pick a single core, in my case cpus 0,20 are the same core
# cpu_hog is any program that spins
#
taskset -c 20 cpu_hog &
# schbench -p 4 means message passing mode with 4 byte messages (like
# pipe test), no sleeps, just bouncing as fast as it can.
#
# make the scheduler choose between the sibling of the hog and cpu 1
#
taskset -c 0,1 schbench -p 4 -m 1 -t 1
Current mainline will stuff both schbench threads onto CPU 1, leaving
CPU 0 100% idle. My first patch with the minimal task_hot() checks
would sometimes pick CPU 0. My second patch that just directly calls
task_hot sticks to cpu1, which is ~3x faster than spreading it.
The full task_hot() checks also really help tbench.
>
> _However_, the migration_cost is supposed the model the cost of leaving
> the LLC, so testing against that here seems wrong.
>
> Let me go play with something that measures the cost of doing that LLC
> scan and compares that against the sleepy time -- of course, now need to
> go figure out how to do this clock thing without rq-lock pain.
>
>
>
> + if (package_sd && !bounce_to_target(p, target)) {
> + for_each_cpu_and(i, sched_domain_span(package_sd), tsk_cpus_allowed(p)) {
> + if (idle_cpu(i)) {
> + target = i;
> + break;
> + }
> +
> + }
> + }
>
> Also note your s/sd/package_sd/ rename is, strictly speaking, wrong.
> Sure, on your current Intel system the LLC is the entire package, but
> this is not true in general.
>
> Take for instance the Intel Core2Quad and AMD Bulldozer thingies, they
> had two dies in one package, and correspondingly two LLC domains in one
> package.
>
> (also, the Intel cluster-on-die thing can split the thing in two)
>
> There were also the old P6 era SMP boards which had external LLC, where
> you could have an LLC shared across multiple packages -- although I'm
> thinking we'll never see that again, due to off package being far
> toooooo slooooooow these days.
Gotcha, makes sense. I'll switch to llc_sd ;)
-chris
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-04 12:40 +0200 |
| Message-ID | <rv2nM-28B-15@gated-at.bofh.it> |
| In reply to | #1393494 |
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> > + if (cpu_rq(cpu)->cfs.nr_running > 1)
> > + return 1;
>
> The nr_running check is interesting. It is supposed to give the same
> benefit as your "do we have anything idle?" variable, but without having
> to constantly update a variable somewhere. I'll have to do a few runs
> to verify (maybe a idle_scan_failed counter).
Right; I got that.. I tried it and it doesn't seem to work as well. But
yeah, its better than nothing.
The reason I'm not too worried about the update_idle_core() thing is
that SMT threads share L1 so touching their sibling state isn't
typically expensive.
But yes, I'd need to find someone with SMT8 or something daft like that
to try this.
> The task_hot checks don't do much for the sleeping schbench runs, but
> they help a lot for this:
>
> # pick a single core, in my case cpus 0,20 are the same core
> # cpu_hog is any program that spins
> #
> taskset -c 20 cpu_hog &
>
> # schbench -p 4 means message passing mode with 4 byte messages (like
> # pipe test), no sleeps, just bouncing as fast as it can.
> #
> # make the scheduler choose between the sibling of the hog and cpu 1
> #
> taskset -c 0,1 schbench -p 4 -m 1 -t 1
>
> Current mainline will stuff both schbench threads onto CPU 1, leaving
> CPU 0 100% idle. My first patch with the minimal task_hot() checks
> would sometimes pick CPU 0. My second patch that just directly calls
> task_hot sticks to cpu1, which is ~3x faster than spreading it.
Urgh, another benchmark to play with ;-)
> The full task_hot() checks also really help tbench.
tbench wants select_idle_siblings() to just not exist; it goes happy
when you just return target.
tbench:
old (mainline like):
Throughput 875.822 MB/sec 2 clients 2 procs max_latency=0.117 ms
Throughput 2017.57 MB/sec 5 clients 5 procs max_latency=0.057 ms
Throughput 3954.66 MB/sec 10 clients 10 procs max_latency=0.094 ms
Throughput 5886.11 MB/sec 20 clients 20 procs max_latency=0.088 ms
Throughput 9095.57 MB/sec 40 clients 40 procs max_latency=0.864 ms
new:
Throughput 876.794 MB/sec 2 clients 2 procs max_latency=0.102 ms
Throughput 2048.73 MB/sec 5 clients 5 procs max_latency=0.095 ms
Throughput 3802.69 MB/sec 10 clients 10 procs max_latency=0.113 ms
Throughput 5521.81 MB/sec 20 clients 20 procs max_latency=0.091 ms
Throughput 10331.8 MB/sec 40 clients 40 procs max_latency=0.444 ms
nothing:
Throughput 759.532 MB/sec 2 clients 2 procs max_latency=0.210 ms
Throughput 1884.01 MB/sec 5 clients 5 procs max_latency=0.094 ms
Throughput 3931.31 MB/sec 10 clients 10 procs max_latency=0.091 ms
Throughput 6478.81 MB/sec 20 clients 20 procs max_latency=0.110 ms
Throughput 10001 MB/sec 40 clients 40 procs max_latency=0.148 ms
See the 20 client have a happy moment ;-) [ivb-ep: 2*10*2]
I've not quite figured out how to make the new bits switch off aggressive
enough to make tbench happy without hurting the others. More numbers:
sysbench-oltp-psql:
old (mainline like):
2: [30 secs] transactions: 53556 (1785.19 per sec.)
5: [30 secs] transactions: 118957 (3965.08 per sec.)
10: [30 secs] transactions: 241126 (8037.22 per sec.)
20: [30 secs] transactions: 383256 (12774.63 per sec.)
40: [30 secs] transactions: 539705 (17989.05 per sec.)
80: [30 secs] transactions: 541833 (18059.16 per sec.)
new:
2: [30 secs] transactions: 53012 (1767.03 per sec.)
5: [30 secs] transactions: 122057 (4068.49 per sec.)
10: [30 secs] transactions: 235781 (7859.09 per sec.)
20: [30 secs] transactions: 355967 (11864.99 per sec.)
40: [30 secs] transactions: 537327 (17909.80 per sec.)
80: [30 secs] transactions: 546017 (18198.82 per sec.)
schbench -m2 -t 20 -c 30000 -s 30000 -r 30:
old (mainline like):
Latency percentiles (usec)
50.0000th: 102
75.0000th: 109
90.0000th: 115
95.0000th: 118
*99.0000th: 5352
99.5000th: 12112
99.9000th: 13008
Over=0, min=0, max=27238
new:
Latency percentiles (usec)
50.0000th: 103
75.0000th: 109
90.0000th: 114
95.0000th: 116
*99.0000th: 120
99.5000th: 121
99.9000th: 124
Over=0, min=0, max=12939
---
include/linux/sched.h | 2 +
kernel/sched/core.c | 3 +
kernel/sched/fair.c | 267 +++++++++++++++++++++++++++++++++++++++--------
kernel/sched/features.h | 9 ++
kernel/sched/idle_task.c | 4 +-
kernel/sched/sched.h | 1 +
kernel/time/tick-sched.c | 10 +-
7 files changed, 246 insertions(+), 50 deletions(-)
diff --git a/include/linux/sched.h b/include/linux/sched.h
index ad9454d..e7ce1a0 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1068,6 +1068,8 @@ struct sched_domain {
u64 max_newidle_lb_cost;
unsigned long next_decay_max_lb_cost;
+ u64 avg_scan_cost; /* select_idle_sibling */
+
#ifdef CONFIG_SCHEDSTATS
/* load_balance() stats */
unsigned int lb_count[CPU_MAX_IDLE_TYPES];
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index c82ca6e..280e73e 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -7278,6 +7278,7 @@ static struct kmem_cache *task_group_cache __read_mostly;
#endif
DECLARE_PER_CPU(cpumask_var_t, load_balance_mask);
+DECLARE_PER_CPU(cpumask_var_t, select_idle_mask);
void __init sched_init(void)
{
@@ -7314,6 +7315,8 @@ void __init sched_init(void)
for_each_possible_cpu(i) {
per_cpu(load_balance_mask, i) = (cpumask_var_t)kzalloc_node(
cpumask_size(), GFP_KERNEL, cpu_to_node(i));
+ per_cpu(select_idle_mask, i) = (cpumask_var_t)kzalloc_node(
+ cpumask_size(), GFP_KERNEL, cpu_to_node(i));
}
#endif /* CONFIG_CPUMASK_OFFSTACK */
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33ab..9290fc8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ balance:
* One idle CPU per node is evaluated for a task numa move.
* Call select_idle_sibling to maybe find a better one.
*/
- if (!cur)
+ if (!cur) {
+ // XXX borken
env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+ }
assign:
assigned = true;
@@ -4491,6 +4493,11 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
}
#ifdef CONFIG_SMP
+
+/* Working cpumask for: load_balance, load_balance_newidle. */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+DEFINE_PER_CPU(cpumask_var_t, select_idle_mask);
+
#ifdef CONFIG_NO_HZ_COMMON
/*
* per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5169,240 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
}
+static int cpumask_next_wrap(int n, const struct cpumask *mask, int start, int *wrapped)
+{
+ int next;
+
+again:
+ next = find_next_bit(cpumask_bits(mask), nr_cpumask_bits, n+1);
+
+ if (*wrapped) {
+ if (next >= start)
+ return nr_cpumask_bits;
+ } else {
+ if (next >= nr_cpumask_bits) {
+ *wrapped = 1;
+ n = -1;
+ goto again;
+ }
+ }
+
+ return next;
+}
+
+#define for_each_cpu_wrap(cpu, mask, start, wrap) \
+ for ((wrap) = 0, (cpu) = (start)-1; \
+ (cpu) = cpumask_next_wrap((cpu), (mask), (start), &(wrap)), \
+ (cpu) < nr_cpumask_bits; )
+
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return false;
+
+ // XXX static key for !SMT topologies
+
+ return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+ int core = cpu_of(rq);
+ int cpu;
+
+ rcu_read_lock();
+ if (test_idle_cores(core))
+ goto unlock;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ if (cpu == core)
+ continue;
+
+ if (!idle_cpu(cpu))
+ goto unlock;
+ }
+
+ set_idle_cores(core);
+unlock:
+ rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, struct sched_domain *sd, int target)
+{
+ struct cpumask *cpus = this_cpu_cpumask_var_ptr(select_idle_mask);
+ int core, cpu, wrap;
+
+ if (!test_idle_cores(target))
+ return -1;
+
+ cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+
+ for_each_cpu_wrap(core, cpus, target, wrap) {
+ bool idle = true;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ cpumask_clear_cpu(cpu, cpus);
+ if (!idle_cpu(cpu))
+ idle = false;
+ }
+
+ if (idle)
+ return core;
+ }
+
+ /*
+ * Failed to find an idle core; stop looking for one.
+ */
+ clear_idle_cores(target);
+
+ return -1;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, struct sched_domain *sd, int target)
+{
+ return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
+static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, int target)
+{
+ struct sched_domain *this_sd = rcu_dereference(*this_cpu_ptr(&sd_llc));
+ u64 time, cost;
+ s64 delta;
+ int cpu, wrap;
+
+ if (sched_feat(AVG_CPU)) {
+ u64 avg_idle = this_rq()->avg_idle;
+ u64 avg_cost = this_sd->avg_scan_cost;
+
+ if (sched_feat(PRINT_AVG))
+ trace_printk("idle: %Ld cost: %Ld\n", avg_idle, avg_cost);
+
+ if (avg_idle / 32 < avg_cost)
+ return -1;
+ }
+
+ time = local_clock();
+
+ for_each_cpu_wrap(cpu, sched_domain_span(sd), target, wrap) {
+ if (!cpumask_test_cpu(cpu, tsk_cpus_allowed(p)))
+ continue;
+ if (idle_cpu(cpu))
+ break;
+ }
+
+ time = local_clock() - time;
+ cost = this_sd->avg_scan_cost;
+ delta = (s64)(time - cost) / 8;
+ /* trace_printk("time: %Ld cost: %Ld delta: %Ld\n", time, cost, delta); */
+ this_sd->avg_scan_cost += delta;
+
+ return cpu;
+}
+
/*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
*/
static int select_idle_sibling(struct task_struct *p, int target)
{
struct sched_domain *sd;
- struct sched_group *sg;
- int i = task_cpu(p);
+ int start, i = task_cpu(p);
if (idle_cpu(target))
return target;
/*
- * If the prevous cpu is cache affine and idle, don't be stupid.
+ * If the previous cpu is cache affine and idle, don't be stupid.
*/
if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
return i;
- /*
- * Otherwise, iterate the domains and find an eligible idle cpu.
- *
- * A completely idle sched group at higher domains is more
- * desirable than an idle group at a lower level, because lower
- * domains have smaller groups and usually share hardware
- * resources which causes tasks to contend on them, e.g. x86
- * hyperthread siblings in the lowest domain (SMT) can contend
- * on the shared cpu pipeline.
- *
- * However, while we prefer idle groups at higher domains
- * finding an idle cpu at the lowest domain is still better than
- * returning 'target', which we've already established, isn't
- * idle.
- */
- sd = rcu_dereference(per_cpu(sd_llc, target));
- for_each_lower_domain(sd) {
- sg = sd->groups;
- do {
- if (!cpumask_intersects(sched_group_cpus(sg),
- tsk_cpus_allowed(p)))
- goto next;
+ start = target;
+ if (sched_feat(ORDER_IDLE))
+ start = per_cpu(sd_llc_id, target); /* first cpu in llc domain */
- /* Ensure the entire group is idle */
- for_each_cpu(i, sched_group_cpus(sg)) {
- if (i == target || !idle_cpu(i))
+ sd = rcu_dereference(per_cpu(sd_llc, start));
+ if (!sd)
+ return target;
+
+ if (sched_feat(OLD_IDLE)) {
+ struct sched_group *sg;
+
+ for_each_lower_domain(sd) {
+ sg = sd->groups;
+ do {
+ if (!cpumask_intersects(sched_group_cpus(sg),
+ tsk_cpus_allowed(p)))
goto next;
- }
- /*
- * It doesn't matter which cpu we pick, the
- * whole group is idle.
- */
- target = cpumask_first_and(sched_group_cpus(sg),
- tsk_cpus_allowed(p));
- goto done;
+ /* Ensure the entire group is idle */
+ for_each_cpu(i, sched_group_cpus(sg)) {
+ if (i == target || !idle_cpu(i))
+ goto next;
+ }
+
+ /*
+ * It doesn't matter which cpu we pick, the
+ * whole group is idle.
+ */
+ target = cpumask_first_and(sched_group_cpus(sg),
+ tsk_cpus_allowed(p));
+ goto done;
next:
- sg = sg->next;
- } while (sg != sd->groups);
- }
+ sg = sg->next;
+ } while (sg != sd->groups);
+ }
done:
+ return target;
+ }
+
+ if (sched_feat(IDLE_CORE)) {
+ i = select_idle_core(p, sd, start);
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
+ }
+
+ if (sched_feat(IDLE_CPU)) {
+ i = select_idle_cpu(p, sd, start);
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
+ }
+
+ if (sched_feat(IDLE_SMT)) {
+ for_each_cpu(i, cpu_smt_mask(target)) {
+ if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+ continue;
+ if (idle_cpu(i))
+ return i;
+ }
+ }
+
return target;
}
@@ -7229,9 +7411,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
*/
#define MAX_PINNED_INTERVAL 512
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
static int need_active_balance(struct lb_env *env)
{
struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa..5dd10ec 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,12 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
SCHED_FEAT(LB_MIN, false)
SCHED_FEAT(ATTACH_AGE_LOAD, true)
+SCHED_FEAT(OLD_IDLE, false)
+SCHED_FEAT(ORDER_IDLE, false)
+
+SCHED_FEAT(IDLE_CORE, true)
+SCHED_FEAT(IDLE_CPU, true)
+SCHED_FEAT(AVG_CPU, true)
+SCHED_FEAT(PRINT_AVG, false)
+
+SCHED_FEAT(IDLE_SMT, false)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce949..cb394db 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
resched_curr(rq);
}
+extern void update_idle_core(struct rq *rq);
+
static struct task_struct *
pick_next_task_idle(struct rq *rq, struct task_struct *prev)
{
put_prev_task(rq, prev);
-
+ update_idle_core(rq);
schedstat_inc(rq, sched_goidle);
return rq->idle;
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fc..5994794 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
* Number of busy cpus in this group.
*/
atomic_t nr_busy_cpus;
+ int has_idle_cores;
unsigned long cpumask[0]; /* iteration mask */
};
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc..6e42cd2 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
WARN_ON_ONCE(irqs_disabled());
/*
- * Update the idle state in the scheduler domain hierarchy
- * when tick_nohz_stop_sched_tick() is called from the idle loop.
- * State will be updated to busy during the first busy tick after
- * exiting idle.
- */
+ * Update the idle state in the scheduler domain hierarchy
+ * when tick_nohz_stop_sched_tick() is called from the idle loop.
+ * State will be updated to busy during the first busy tick after
+ * exiting idle.
+ */
set_cpu_sd_state_idle();
local_irq_disable();
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-04 17:40 +0200 |
| Message-ID | <rv745-6DR-1@gated-at.bofh.it> |
| In reply to | #1394142 |
On Wed, May 04, 2016 at 12:37:01PM +0200, Peter Zijlstra wrote:
> +static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, int target)
> +{
> + struct sched_domain *this_sd = rcu_dereference(*this_cpu_ptr(&sd_llc));
> + u64 time, cost;
> + s64 delta;
> + int cpu, wrap;
> +
> + if (sched_feat(AVG_CPU)) {
> + u64 avg_idle = this_rq()->avg_idle;
> + u64 avg_cost = this_sd->avg_scan_cost;
> +
> + if (sched_feat(PRINT_AVG))
> + trace_printk("idle: %Ld cost: %Ld\n", avg_idle, avg_cost);
> +
> + if (avg_idle / 32 < avg_cost)
s/32/512/ + IDLE_SMT fixes a hackbench regression
hackbench, like tbench, doesn't like IDLE_CPU to trigger, but apparently
needs IDLE_SMT.
Bah, I could sort of explain 32 away, but 512 is firmly in the magic
value range :/
> + return -1;
> + }
> +
> + time = local_clock();
> +
> + for_each_cpu_wrap(cpu, sched_domain_span(sd), target, wrap) {
> + if (!cpumask_test_cpu(cpu, tsk_cpus_allowed(p)))
> + continue;
> + if (idle_cpu(cpu))
> + break;
> + }
> +
> + time = local_clock() - time;
> + cost = this_sd->avg_scan_cost;
> + delta = (s64)(time - cost) / 8;
> + /* trace_printk("time: %Ld cost: %Ld delta: %Ld\n", time, cost, delta); */
> + this_sd->avg_scan_cost += delta;
> +
> + return cpu;
> +}
[toc] | [prev] | [next] | [standalone]
| From | Matt Fleming <matt@codeblueprint.co.uk> |
|---|---|
| Date | 2016-05-06 00:10 +0200 |
| Message-ID | <rvzD4-9z-11@gated-at.bofh.it> |
| In reply to | #1394142 |
On Wed, 04 May, at 12:37:01PM, Peter Zijlstra wrote: > > tbench wants select_idle_siblings() to just not exist; it goes happy > when you just return target. I've been playing with this patch a little bit by hitting it with tbench on a Xeon, 12 cores with HT enabled, 2 sockets (48 cpus). I see a throughput improvement for 16, 32, 64, 128 and 256 clients when compared against mainline, so that's, OLD_IDLE, ORDER_IDLE, NO_IDLE_CORE, NO_IDLE_CPU, NO_IDLE_SMT vs. NO_OLD_IDLE, NO_ORDER_IDLE, IDLE_CORE, IDLE_CPU, IDLE_SMT See, [OLD] Throughput 5345.6 MB/sec 16 clients 16 procs max_latency=0.277 ms avg_latency=0.211853 ms [NEW] Throughput 5514.52 MB/sec 16 clients 16 procs max_latency=0.493 ms avg_latency=0.176441 ms [OLD] Throughput 7401.76 MB/sec 32 clients 32 procs max_latency=1.804 ms avg_latency=0.451147 ms [NEW] Throughput 10044.9 MB/sec 32 clients 32 procs max_latency=3.421 ms avg_latency=0.582529 ms [OLD] Throughput 13265.9 MB/sec 64 clients 64 procs max_latency=7.395 ms avg_latency=0.927147 ms [NEW] Throughput 13929.6 MB/sec 64 clients 64 procs max_latency=7.022 ms avg_latency=1.017059 ms [OLD] Throughput 12827.8 MB/sec 128 clients 128 procs max_latency=16.256 ms avg_latency=2.763706 ms [NEW] Throughput 13364.2 MB/sec 128 clients 128 procs max_latency=16.630 ms avg_latency=3.002971 ms [OLD] Throughput 12653.1 MB/sec 256 clients 256 procs max_latency=44.722 ms avg_latency=5.741647 ms [NEW] Throughput 12965.7 MB/sec 256 clients 256 procs max_latency=59.061 ms avg_latency=8.699118 ms For throughput changes to 1, 2, 4 and 8 clients it's more of a mixture with sometimes the old config winning and sometimes losing. [OLD] Throughput 488.819 MB/sec 1 clients 1 procs max_latency=0.191 ms avg_latency=0.058794 ms [NEW] Throughput 486.106 MB/sec 1 clients 1 procs max_latency=0.085 ms avg_latency=0.045794 ms [OLD] Throughput 925.987 MB/sec 2 clients 2 procs max_latency=0.201 ms avg_latency=0.090882 ms [NEW] Throughput 954.944 MB/sec 2 clients 2 procs max_latency=0.199 ms avg_latency=0.064294 ms [OLD] Throughput 1764.02 MB/sec 4 clients 4 procs max_latency=0.160 ms avg_latency=0.075206 ms [NEW] Throughput 1756.8 MB/sec 4 clients 4 procs max_latency=0.105 ms avg_latency=0.062382 ms [OLD] Throughput 3384.22 MB/sec 8 clients 8 procs max_latency=0.276 ms avg_latency=0.099441 ms [NEW] Throughput 3375.47 MB/sec 8 clients 8 procs max_latency=0.103 ms avg_latency=0.064176 ms Looking at latency, the new code consistently performs worse at the top end for 256 clients. Admittedly at that point the machine is pretty overloaded. Things are much better at the lower end. One thing I haven't yet done is twiddled the bits individually to see what the best combination is. Have you settled on the right settings yet?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-06 21:00 +0200 |
| Message-ID | <rvT8K-20P-17@gated-at.bofh.it> |
| In reply to | #1395398 |
On Thu, 2016-05-05 at 23:03 +0100, Matt Fleming wrote: > One thing I haven't yet done is twiddled the bits individually to see > what the best combination is. Have you settled on the right settings > yet? Lighter configs, revert sched/fair: Fix fairness issue on migration, twiddle knobs. Added an IDLE_SIBLING knob to ~virgin master.. only sorta virgin because I always throttle nohz. 1 x i4790 master for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done Throughput 871.785 MB/sec 1 clients 1 procs max_latency=0.324 ms Throughput 1514.5 MB/sec 2 clients 2 procs max_latency=0.411 ms Throughput 2722.43 MB/sec 4 clients 4 procs max_latency=2.400 ms Throughput 4334.46 MB/sec 8 clients 8 procs max_latency=3.561 ms echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features Throughput 1078.69 MB/sec 1 clients 1 procs max_latency=2.274 ms Throughput 2130.33 MB/sec 2 clients 2 procs max_latency=1.451 ms Throughput 3484.18 MB/sec 4 clients 4 procs max_latency=3.430 ms Throughput 4423.69 MB/sec 8 clients 8 procs max_latency=5.363 ms masterx for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done Throughput 707.673 MB/sec 1 clients 1 procs max_latency=2.279 ms Throughput 1503.55 MB/sec 2 clients 2 procs max_latency=0.695 ms Throughput 2527.73 MB/sec 4 clients 4 procs max_latency=2.321 ms Throughput 4291.26 MB/sec 8 clients 8 procs max_latency=3.815 ms echo NO_IDLE_CPU > /sys/kernel/debug/sched_features homer:~ # for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done Throughput 865.936 MB/sec 1 clients 1 procs max_latency=0.411 ms Throughput 1586.41 MB/sec 2 clients 2 procs max_latency=2.293 ms Throughput 2638.39 MB/sec 4 clients 4 procs max_latency=2.037 ms Throughput 4405.43 MB/sec 8 clients 8 procs max_latency=3.581 ms + echo NO_AVG_CPU > /sys/kernel/debug/sched_features + echo IDLE_SMT > /sys/kernel/debug/sched_features Throughput 697.126 MB/sec 1 clients 1 procs max_latency=2.220 ms Throughput 1562.82 MB/sec 2 clients 2 procs max_latency=0.526 ms Throughput 2620.62 MB/sec 4 clients 4 procs max_latency=6.460 ms Throughput 4345.13 MB/sec 8 clients 8 procs max_latency=27.921 ms 4 x E7-8890 master for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done Throughput 615.663 MB/sec 1 clients 1 procs max_latency=0.087 ms Throughput 1171.53 MB/sec 2 clients 2 procs max_latency=0.087 ms Throughput 2251.22 MB/sec 4 clients 4 procs max_latency=0.078 ms Throughput 4090.76 MB/sec 8 clients 8 procs max_latency=0.801 ms Throughput 7695.92 MB/sec 16 clients 16 procs max_latency=0.235 ms Throughput 15152 MB/sec 32 clients 32 procs max_latency=0.693 ms Throughput 21628.2 MB/sec 64 clients 64 procs max_latency=4.666 ms Throughput 43185.7 MB/sec 128 clients 128 procs max_latency=7.280 ms Throughput 72144.5 MB/sec 256 clients 256 procs max_latency=8.194 ms echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features Throughput 954.593 MB/sec 1 clients 1 procs max_latency=0.185 ms Throughput 1882.65 MB/sec 2 clients 2 procs max_latency=0.278 ms Throughput 3457.03 MB/sec 4 clients 4 procs max_latency=0.431 ms Throughput 6279.38 MB/sec 8 clients 8 procs max_latency=0.730 ms Throughput 11170.4 MB/sec 16 clients 16 procs max_latency=0.500 ms Throughput 21940.9 MB/sec 32 clients 32 procs max_latency=0.475 ms Throughput 41738.8 MB/sec 64 clients 64 procs max_latency=3.669 ms Throughput 67634.6 MB/sec 128 clients 128 procs max_latency=6.676 ms Throughput 76299.7 MB/sec 256 clients 256 procs max_latency=7.878 ms masterx for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done Throughput 587.956 MB/sec 1 clients 1 procs max_latency=0.124 ms Throughput 1140.16 MB/sec 2 clients 2 procs max_latency=0.476 ms Throughput 2296.03 MB/sec 4 clients 4 procs max_latency=0.142 ms Throughput 4116.65 MB/sec 8 clients 8 procs max_latency=0.464 ms Throughput 7820.27 MB/sec 16 clients 16 procs max_latency=0.238 ms Throughput 14899.2 MB/sec 32 clients 32 procs max_latency=0.321 ms Throughput 21909.8 MB/sec 64 clients 64 procs max_latency=0.905 ms Throughput 35495.2 MB/sec 128 clients 128 procs max_latency=6.158 ms Throughput 75863.2 MB/sec 256 clients 256 procs max_latency=7.650 ms echo NO_IDLE_CPU > /sys/kernel/debug/sched_features Throughput 555.15 MB/sec 1 clients 1 procs max_latency=0.096 ms Throughput 1195.12 MB/sec 2 clients 2 procs max_latency=0.131 ms Throughput 2276.97 MB/sec 4 clients 4 procs max_latency=0.105 ms Throughput 4248.14 MB/sec 8 clients 8 procs max_latency=0.131 ms Throughput 7860.86 MB/sec 16 clients 16 procs max_latency=0.210 ms Throughput 15178.6 MB/sec 32 clients 32 procs max_latency=0.229 ms Throughput 21523.9 MB/sec 64 clients 64 procs max_latency=0.842 ms Throughput 31082.1 MB/sec 128 clients 128 procs max_latency=7.311 ms Throughput 75887.9 MB/sec 256 clients 256 procs max_latency=7.764 ms + echo NO_AVG_CPU > /sys/kernel/debug/sched_features Throughput 598.063 MB/sec 1 clients 1 procs max_latency=0.131 ms Throughput 1140.2 MB/sec 2 clients 2 procs max_latency=0.092 ms Throughput 2268.68 MB/sec 4 clients 4 procs max_latency=0.170 ms Throughput 4259.7 MB/sec 8 clients 8 procs max_latency=0.212 ms Throughput 7904.15 MB/sec 16 clients 16 procs max_latency=0.191 ms Throughput 14840 MB/sec 32 clients 32 procs max_latency=0.279 ms Throughput 21701.5 MB/sec 64 clients 64 procs max_latency=0.856 ms Throughput 38945 MB/sec 128 clients 128 procs max_latency=7.501 ms Throughput 75669.4 MB/sec 256 clients 256 procs max_latency=14.984 ms + echo IDLE_SMT > /sys/kernel/debug/sched_features Throughput 592.799 MB/sec 1 clients 1 procs max_latency=0.120 ms Throughput 1208.28 MB/sec 2 clients 2 procs max_latency=0.078 ms Throughput 2319.22 MB/sec 4 clients 4 procs max_latency=0.141 ms Throughput 4196.64 MB/sec 8 clients 8 procs max_latency=0.253 ms Throughput 7816.47 MB/sec 16 clients 16 procs max_latency=0.117 ms Throughput 14990.8 MB/sec 32 clients 32 procs max_latency=0.189 ms Throughput 21809.4 MB/sec 64 clients 64 procs max_latency=0.832 ms Throughput 44813 MB/sec 128 clients 128 procs max_latency=7.930 ms Throughput 75978.1 MB/sec 256 clients 256 procs max_latency=7.337 ms
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-09 10:40 +0200 |
| Message-ID | <rwOTo-1x9-19@gated-at.bofh.it> |
| In reply to | #1396071 |
On Fri, May 06, 2016 at 08:54:38PM +0200, Mike Galbraith wrote: > master > Throughput 2722.43 MB/sec 4 clients 4 procs max_latency=2.400 ms > echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features > Throughput 3484.18 MB/sec 4 clients 4 procs max_latency=3.430 ms Yeah, I know about that bump, I just haven't managed to find a way to preserve that and keep all the other benchmarks ticking along :/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-09 11:00 +0200 |
| Message-ID | <rwPcK-1Kc-17@gated-at.bofh.it> |
| In reply to | #1396847 |
On Mon, 2016-05-09 at 10:33 +0200, Peter Zijlstra wrote: > On Fri, May 06, 2016 at 08:54:38PM +0200, Mike Galbraith wrote: > > master > > Throughput 2722.43 MB/sec 4 clients 4 procs max_latency=2.400 ms > > > echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features > > Throughput 3484.18 MB/sec 4 clients 4 procs max_latency=3.430 ms > > Yeah, I know about that bump, I just haven't managed to find a way to > preserve that and keep all the other benchmarks ticking along :/ Yup, L3 ain't L2. I haven't come up with a good metric either. Poo. Until that comes along, microbenchmarks can bugger off, real boxen don't just play high speed ping-pong with themselves for a living ;-) -Mike
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-04 17:50 +0200 |
| Message-ID | <rv7dM-6IW-9@gated-at.bofh.it> |
| In reply to | #1393494 |
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote: > # pick a single core, in my case cpus 0,20 are the same core > # cpu_hog is any program that spins > # > taskset -c 20 cpu_hog & > > # schbench -p 4 means message passing mode with 4 byte messages (like > # pipe test), no sleeps, just bouncing as fast as it can. > # > # make the scheduler choose between the sibling of the hog and cpu 1 > # > taskset -c 0,1 schbench -p 4 -m 1 -t 1 Will that schbench thingy print something? Mine doesn't seem to output anything, not actually exit, although it stops consuming CPU cycles at some point.
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-05-04 19:50 +0200 |
| Message-ID | <rv95U-8tM-15@gated-at.bofh.it> |
| In reply to | #1394487 |
On Wed, May 04, 2016 at 05:45:10PM +0200, Peter Zijlstra wrote:
> On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> > # pick a single core, in my case cpus 0,20 are the same core
> > # cpu_hog is any program that spins
> > #
> > taskset -c 20 cpu_hog &
> >
> > # schbench -p 4 means message passing mode with 4 byte messages (like
> > # pipe test), no sleeps, just bouncing as fast as it can.
> > #
> > # make the scheduler choose between the sibling of the hog and cpu 1
> > #
> > taskset -c 0,1 schbench -p 4 -m 1 -t 1
>
> Will that schbench thingy print something? Mine doesn't seem to output
> anything, not actually exit, although it stops consuming CPU cycles at
> some point.
>
>
It should, make sure you're at the top commit in git.
git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git
It's not recent so I'd be surprised if you weren't already there. The
default runtime is 30 seconds, but you can use -r to specify something
shorter.
It's possible I'm missing a wakeup to shut the whole thing down, but I
thought I fixed that.
./schbench -p 4 -m 1 -t 1
Latency percentiles (usec)
50.0000th: 5
75.0000th: 5
90.0000th: 5
95.0000th: 5
*99.0000th: 8
99.5000th: 15
99.9000th: 17
Over=0, min=0, max=652
avg worker transfer: 113768.27 ops/sec 444.41KB/s
-chris
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-05 11:40 +0200 |
| Message-ID | <rvnVg-5Rb-23@gated-at.bofh.it> |
| In reply to | #1394580 |
On Wed, May 04, 2016 at 01:46:16PM -0400, Chris Mason wrote: > It should, make sure you're at the top commit in git. > > git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git I did double check; I am on the top commit of that. I refetched and rebuild just to make tripple sure. > It's not recent so I'd be surprised if you weren't already there. The > default runtime is 30 seconds, but you can use -r to specify something > shorter. > > It's possible I'm missing a wakeup to shut the whole thing down, but I > thought I fixed that. Seems to still be missing, because: > ./schbench -p 4 -m 1 -t 1 > Latency percentiles (usec) > 50.0000th: 5 > 75.0000th: 5 > 90.0000th: 5 > 95.0000th: 5 > *99.0000th: 8 > 99.5000th: 15 > 99.9000th: 17 > Over=0, min=0, max=652 > avg worker transfer: 113768.27 ops/sec 444.41KB/s is not what mine does. I get ~25sec of cpu time and then it stalls forever. I'll try and have a prod at the program itself if you have no pending changes on your end.
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-05-05 16:00 +0200 |
| Message-ID | <rvrYS-113-15@gated-at.bofh.it> |
| In reply to | #1394973 |
On Thu, May 05, 2016 at 11:33:38AM +0200, Peter Zijlstra wrote: > On Wed, May 04, 2016 at 01:46:16PM -0400, Chris Mason wrote: > > It should, make sure you're at the top commit in git. > > > > git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git > > I did double check; I am on the top commit of that. I refetched and > rebuild just to make tripple sure. > > > It's not recent so I'd be surprised if you weren't already there. The > > default runtime is 30 seconds, but you can use -r to specify something > > shorter. > > > > It's possible I'm missing a wakeup to shut the whole thing down, but I > > thought I fixed that. > > Seems to still be missing, because: > > > ./schbench -p 4 -m 1 -t 1 > > Latency percentiles (usec) > > 50.0000th: 5 > > 75.0000th: 5 > > 90.0000th: 5 > > 95.0000th: 5 > > *99.0000th: 8 > > 99.5000th: 15 > > 99.9000th: 17 > > Over=0, min=0, max=652 > > avg worker transfer: 113768.27 ops/sec 444.41KB/s > > is not what mine does. I get ~25sec of cpu time and then it stalls > forever. > > I'll try and have a prod at the program itself if you have no pending > changes on your end. Sorry, I don't. Look at sleep_for_runtime() and how I test/set the global stopping variable in different places. I've almost certainly got someone waiting on a wakeup that'll never come. If all else fails, run_msg_thread() can pass a timeout to fwait() for a less error prone setup. I was just hoping to avoid the timers kernel side. -chris
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-06 09:20 +0200 |
| Message-ID | <rvIdj-JF-7@gated-at.bofh.it> |
| In reply to | #1395137 |
On Thu, May 05, 2016 at 09:58:44AM -0400, Chris Mason wrote: > > I'll try and have a prod at the program itself if you have no pending > > changes on your end. > > Sorry, I don't. Look at sleep_for_runtime() and how I test/set the > global stopping variable in different places. I've almost certainly got > someone waiting on a wakeup that'll never come. The below makes it go.. --- schbench.c | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/schbench.c b/schbench.c index a0e9f7e..f299959 100644 --- a/schbench.c +++ b/schbench.c @@ -49,7 +49,7 @@ static int pipe_test = 0; static unsigned int max_us = 50000; /* the message threads flip this to true when they decide runtime is up */ -static unsigned long stopping = 0; +static volatile unsigned long stopping = 0; /* @@ -746,8 +746,8 @@ static void sleep_for_runtime() else break; } - stopping = 1; __sync_synchronize(); + stopping = 1; } int main(int ac, char **av)
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-05-06 19:30 +0200 |
| Message-ID | <rvRJF-YG-15@gated-at.bofh.it> |
| In reply to | #1395653 |
On Fri, May 06, 2016 at 09:12:51AM +0200, Peter Zijlstra wrote: > ontent-Length: 973 > > On Thu, May 05, 2016 at 09:58:44AM -0400, Chris Mason wrote: > > > I'll try and have a prod at the program itself if you have no pending > > > changes on your end. > > > > Sorry, I don't. Look at sleep_for_runtime() and how I test/set the > > global stopping variable in different places. I've almost certainly got > > someone waiting on a wakeup that'll never come. > > The below makes it go.. Thanks Peter, pushed to git. -chris
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-06 09:30 +0200 |
| Message-ID | <rvIn0-P0-33@gated-at.bofh.it> |
| In reply to | #1393494 |
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> # pick a single core, in my case cpus 0,20 are the same core
> # cpu_hog is any program that spins
> #
> taskset -c 20 cpu_hog &
>
> # schbench -p 4 means message passing mode with 4 byte messages (like
> # pipe test), no sleeps, just bouncing as fast as it can.
> #
> # make the scheduler choose between the sibling of the hog and cpu 1
> #
> taskset -c 0,1 schbench -p 4 -m 1 -t 1
>
> Current mainline will stuff both schbench threads onto CPU 1, leaving
> CPU 0 100% idle. My first patch with the minimal task_hot() checks
> would sometimes pick CPU 0. My second patch that just directly calls
> task_hot sticks to cpu1, which is ~3x faster than spreading it.
Ok, with the thing fixed, my current patch seems to DTRT. If I trace
sched_migrate_task() I get:
$ grep schbench trace
doit-schbench-2-4042 [004] d..3 144541.309747: sched_migrate_task: comm=doit-schbench-2 pid=4042 prio=120 orig_cpu=4 dest_cpu=4
doit-schbench-2-4042 [004] d..2 144541.309772: sched_migrate_task: comm=doit-schbench-2 pid=4043 prio=120 orig_cpu=4 dest_cpu=11
doit-schbench-2-4042 [004] d..3 144541.309855: sched_migrate_task: comm=doit-schbench-2 pid=4042 prio=120 orig_cpu=4 dest_cpu=4
doit-schbench-2-4042 [004] d..2 144541.309882: sched_migrate_task: comm=doit-schbench-2 pid=4044 prio=120 orig_cpu=4 dest_cpu=5
migration/11-77 [011] d..4 144541.309974: sched_migrate_task: comm=doit-schbench-2 pid=4043 prio=120 orig_cpu=11 dest_cpu=12
migration/5-40 [005] d..4 144541.310013: sched_migrate_task: comm=doit-schbench-2 pid=4044 prio=120 orig_cpu=5 dest_cpu=6
schbench-4044 [001] d..3 144541.310995: sched_migrate_task: comm=schbench pid=4044 prio=120 orig_cpu=1 dest_cpu=1
schbench-4044 [001] d..2 144541.310999: sched_migrate_task: comm=schbench pid=4045 prio=120 orig_cpu=1 dest_cpu=1
schbench-4045 [001] d..3 144541.311232: sched_migrate_task: comm=schbench pid=4045 prio=120 orig_cpu=1 dest_cpu=1
schbench-4045 [001] d..2 144541.311234: sched_migrate_task: comm=schbench pid=4046 prio=120 orig_cpu=1 dest_cpu=1
So the thing gets put on cpu1 and never leaves.
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web