Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1391619 > unrolled thread

Re: sched: tweak select_idle_sibling to look for idle threads

Started byPeter Zijlstra <peterz@infradead.org>
First post2016-04-30 14:50 +0200
Last post2016-05-02 17:20 +0200
Articles 20 on this page of 46 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-04-30 14:50 +0200
    Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 09:20 +0200
      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-01 11:00 +0200
        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 11:30 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-07 11:10 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-08 10:10 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 04:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 05:50 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:10 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 09:50 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 11:00 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:40 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-10 09:10 +0200
                            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-10 09:50 +0200
                              Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-10 17:30 +0200
                                Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 05:00 +0200
                                  Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 06:20 +0200
                                    Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 11:10 +0200
                                      Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 12:00 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 06:00 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:20 +0200
      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 10:50 +0200
        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 17:00 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:00 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-02 17:50 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 16:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-03 17:20 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 12:40 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:40 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-06 00:10 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-06 21:00 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-09 10:40 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:00 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:50 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-04 19:50 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-05 11:40 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-05 16:00 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:20 +0200
                            Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-06 19:30 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:30 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 19:40 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:10 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Ingo Molnar <mingo@kernel.org> - 2016-05-02 18:10 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 13:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 20:30 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:20 +0200

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#1396578

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-09 06:20 +0200
Message-ID<rwKPM-4Mr-7@gated-at.bofh.it>
In reply to#1396570
On Mon, May 09, 2016 at 05:52:51AM +0200, Mike Galbraith wrote:
> On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote:
> 
> > In addition, I would argue maybe beefing up idle balancing is a more
> > productive way to spread load, as work-stealing just does what needs
> > to be done. And seems it has been (sub-unconsciously) neglected in this
> > case, :)
> 
> P.S. Nope, I'm dinging up multiple spots ;-)

You bet, :)

[toc] | [prev] | [next] | [standalone]


#1392067

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-02 10:50 +0200
Message-ID<ruhIe-dd-5@gated-at.bofh.it>
In reply to#1391757
On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:

> Nah, tbench is just variance prone.  It got dinged up at clients=cores
> on my desktop box, on 4 sockets the high end got seriously dinged up.


Ha!, check this:

root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo
NO_ORDER_IDLE > /debug/sched_features ; echo IDLE_CORE >
/debug/sched_features ; echo NO_FORCE_CORE > /debug/sched_features ;
tbench 20 -t 10

Throughput 5956.32 MB/sec  20 clients  20 procs  max_latency=0.126 ms


root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo ORDER_IDLE >
/debug/sched_features ; echo IDLE_CORE > /debug/sched_features ; echo
NO_FORCE_CORE > /debug/sched_features ; tbench 20 -t 10

Throughput 5011.86 MB/sec  20 clients  20 procs  max_latency=0.116 ms



That little ORDER_IDLE thing hurts silly. That's a little patch I had
lying about because some people complained that tasks hop around the
cache domain, instead of being stuck to a CPU.

I suspect what happens is that by all CPUs starting to look for idle at
the same place (the first cpu in the domain) they all find the same idle
cpu and things pile up.

The old behaviour, where they all start iterating from where they were
avoids some of that, at the cost of making tasks hop around.

Lets see if I can get the same behaviour out of the cpumask iteration
code..





---
 kernel/sched/fair.c      | 218 +++++++++++++++++++++++++++++++++++++----------
 kernel/sched/features.h  |   4 +
 kernel/sched/idle_task.c |   4 +-
 kernel/sched/sched.h     |   1 +
 kernel/time/tick-sched.c |  10 +--
 5 files changed, 188 insertions(+), 49 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33ab..b7626a4 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ balance:
 	 * One idle CPU per node is evaluated for a task numa move.
 	 * Call select_idle_sibling to maybe find a better one.
 	 */
-	if (!cur)
+	if (!cur) {
+		// XXX borken
 		env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+	}
 
 assign:
 	assigned = true;
@@ -4491,6 +4493,17 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 }
 
 #ifdef CONFIG_SMP
+
+/* 
+ * Working cpumask for:
+ *   load_balance, 
+ *   load_balance_newidle,
+ *   select_idle_core.
+ *
+ * Assumes softirqs are disabled when in use.
+ */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+
 #ifdef CONFIG_NO_HZ_COMMON
 /*
  * per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5175,187 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
 	return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
 }
 
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return false;
+
+	// XXX static key for !SMT topologies
+
+	return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+	int core = cpu_of(rq);
+	int cpu;
+
+	rcu_read_lock();
+	if (test_idle_cores(core))
+		goto unlock;
+
+	for_each_cpu(cpu, cpu_smt_mask(core)) {
+		if (cpu == core)
+			continue;
+
+		if (!idle_cpu(cpu))
+			goto unlock;
+	}
+
+	set_idle_cores(core);
+unlock:
+	rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, int target)
+{
+	struct cpumask *cpus = this_cpu_cpumask_var_ptr(load_balance_mask);
+	struct sched_domain *sd;
+	int core, cpu;
+
+	sd = rcu_dereference(per_cpu(sd_llc, target));
+	cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+	for_each_cpu(core, cpus) {
+		bool idle = true;
+
+		for_each_cpu(cpu, cpu_smt_mask(core)) {
+			cpumask_clear_cpu(cpu, cpus);
+			if (!idle_cpu(cpu))
+				idle = false;
+		}
+
+		if (idle)
+			break;
+	}
+
+	return core;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+static inline void clear_idle_cores(int cpu) { }
+static inline void set_idle_cores(int cpu) { }
+
+static inline bool test_idle_cores(int cpu)
+{
+	return false;
+}
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, int target)
+{
+	return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
 /*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
  */
 static int select_idle_sibling(struct task_struct *p, int target)
 {
 	struct sched_domain *sd;
-	struct sched_group *sg;
 	int i = task_cpu(p);
 
 	if (idle_cpu(target))
 		return target;
 
 	/*
-	 * If the prevous cpu is cache affine and idle, don't be stupid.
+	 * If the previous cpu is cache affine and idle, don't be stupid.
 	 */
 	if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
 		return i;
 
-	/*
-	 * Otherwise, iterate the domains and find an eligible idle cpu.
-	 *
-	 * A completely idle sched group at higher domains is more
-	 * desirable than an idle group at a lower level, because lower
-	 * domains have smaller groups and usually share hardware
-	 * resources which causes tasks to contend on them, e.g. x86
-	 * hyperthread siblings in the lowest domain (SMT) can contend
-	 * on the shared cpu pipeline.
-	 *
-	 * However, while we prefer idle groups at higher domains
-	 * finding an idle cpu at the lowest domain is still better than
-	 * returning 'target', which we've already established, isn't
-	 * idle.
-	 */
-	sd = rcu_dereference(per_cpu(sd_llc, target));
-	for_each_lower_domain(sd) {
-		sg = sd->groups;
-		do {
-			if (!cpumask_intersects(sched_group_cpus(sg),
-						tsk_cpus_allowed(p)))
-				goto next;
+	i = target;
+	if (sched_feat(ORDER_IDLE))
+		i = per_cpu(sd_llc_id, target); /* first cpu in llc domain */
+	sd = rcu_dereference(per_cpu(sd_llc, i));
+	if (!sd)
+		return target;
+
+	if (sched_feat(OLD_IDLE)) {
+		struct sched_group *sg;
 
-			/* Ensure the entire group is idle */
-			for_each_cpu(i, sched_group_cpus(sg)) {
-				if (i == target || !idle_cpu(i))
+		for_each_lower_domain(sd) {
+			sg = sd->groups;
+			do {
+				if (!cpumask_intersects(sched_group_cpus(sg),
+							tsk_cpus_allowed(p)))
 					goto next;
-			}
 
-			/*
-			 * It doesn't matter which cpu we pick, the
-			 * whole group is idle.
-			 */
-			target = cpumask_first_and(sched_group_cpus(sg),
-					tsk_cpus_allowed(p));
-			goto done;
+				/* Ensure the entire group is idle */
+				for_each_cpu(i, sched_group_cpus(sg)) {
+					if (i == target || !idle_cpu(i))
+						goto next;
+				}
+
+				/*
+				 * It doesn't matter which cpu we pick, the
+				 * whole group is idle.
+				 */
+				target = cpumask_first_and(sched_group_cpus(sg),
+						tsk_cpus_allowed(p));
+				goto done;
 next:
-			sg = sg->next;
-		} while (sg != sd->groups);
-	}
+				sg = sg->next;
+			} while (sg != sd->groups);
+		}
 done:
+		return target;
+	}
+
+	/*
+	 * If there are idle cores to be had, go find one.
+	 */
+	if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
+		i = select_idle_core(p, target);
+		if ((unsigned)i < nr_cpumask_bits)
+			return i;
+
+		/*
+		 * Failed to find an idle core; stop looking for one.
+		 */
+		clear_idle_cores(target);
+	}
+
+	if (sched_feat(FORCE_CORE)) {
+		i = select_idle_core(p, target);
+		if ((unsigned)i < nr_cpumask_bits)
+			return i;
+	}
+
+	/*
+	 * Otherwise, settle for anything idle in this cache domain.
+	 */
+	for_each_cpu(i, sched_domain_span(sd)) {
+		if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+			continue;
+		if (idle_cpu(i))
+			return i;
+	}
+
 	return target;
 }
 
@@ -7229,9 +7364,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
  */
 #define MAX_PINNED_INTERVAL	512
 
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
 static int need_active_balance(struct lb_env *env)
 {
 	struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa..11ab7ab 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,7 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
 SCHED_FEAT(LB_MIN, false)
 SCHED_FEAT(ATTACH_AGE_LOAD, true)
 
+SCHED_FEAT(OLD_IDLE, false)
+SCHED_FEAT(ORDER_IDLE, false)
+SCHED_FEAT(IDLE_CORE, true)
+SCHED_FEAT(FORCE_CORE, false)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce949..cb394db 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
 	resched_curr(rq);
 }
 
+extern void update_idle_core(struct rq *rq);
+
 static struct task_struct *
 pick_next_task_idle(struct rq *rq, struct task_struct *prev)
 {
 	put_prev_task(rq, prev);
-
+	update_idle_core(rq);
 	schedstat_inc(rq, sched_goidle);
 	return rq->idle;
 }
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fc..5994794 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
 	 * Number of busy cpus in this group.
 	 */
 	atomic_t nr_busy_cpus;
+	int	has_idle_cores;
 
 	unsigned long cpumask[0]; /* iteration mask */
 };
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc..6e42cd2 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
 	WARN_ON_ONCE(irqs_disabled());
 
 	/*
- 	 * Update the idle state in the scheduler domain hierarchy
- 	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
- 	 * State will be updated to busy during the first busy tick after
- 	 * exiting idle.
- 	 */
+	 * Update the idle state in the scheduler domain hierarchy
+	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
+	 * State will be updated to busy during the first busy tick after
+	 * exiting idle.
+	 */
 	set_cpu_sd_state_idle();
 
 	local_irq_disable();

[toc] | [prev] | [next] | [standalone]


#1392265

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-02 17:00 +0200
Message-ID<runuj-5gi-33@gated-at.bofh.it>
In reply to#1392067
On Mon, 2016-05-02 at 10:46 +0200, Peter Zijlstra wrote:
> On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> 
> > Nah, tbench is just variance prone.  It got dinged up at clients=cores
> > on my desktop box, on 4 sockets the high end got seriously dinged up.
> 
> 
> Ha!, check this:
> 
> root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo
> NO_ORDER_IDLE > /debug/sched_features ; echo IDLE_CORE >
> /debug/sched_features ; echo NO_FORCE_CORE > /debug/sched_features ;
> tbench 20 -t 10
> 
> Throughput 5956.32 MB/sec  20 clients  20 procs  max_latency=0.126 ms
> 
> 
> root@ivb-ep:~# echo OLD_IDLE > /debug/sched_features ; echo ORDER_IDLE >
> /debug/sched_features ; echo IDLE_CORE > /debug/sched_features ; echo
> NO_FORCE_CORE > /debug/sched_features ; tbench 20 -t 10
> 
> Throughput 5011.86 MB/sec  20 clients  20 procs  max_latency=0.116 ms
> 
> 
> 
> That little ORDER_IDLE thing hurts silly. That's a little patch I had
> lying about because some people complained that tasks hop around the
> cache domain, instead of being stuck to a CPU.
> 
> I suspect what happens is that by all CPUs starting to look for idle at
> the same place (the first cpu in the domain) they all find the same idle
> cpu and things pile up.
> 
> The old behaviour, where they all start iterating from where they were
> avoids some of that, at the cost of making tasks hop around.
> 
> Lets see if I can get the same behaviour out of the cpumask iteration
> code..

Order is one thing, but what the old behavior does first and foremost
is when the box starts getting really busy, only looking at target's
sibling shuts select_idle_sibling() down instead of letting it wreck
things.  Once cores are moving, there are no large piles of anything
left to collect other than pain.

We really need a good way to know we're not gonna turn the box into a
shredder.  The wake_wide() thing might help some, likely wants some
twiddling, in_interrupt() might be another time to try hard.

Anyway, the has_idle_cores business seems to shut select_idle_sibling()
down rather nicely when the the box gets busy.  Forcing either core,
target's sibling or go fish turned in a top end win on 48 rq/socket.

Oh btw, did you know single socket boxen have no sd_busy?  That doesn't
look right.

fromm:~/:[0]# for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done
Throughput 511.016 MB/sec  1 clients  1 procs  max_latency=0.113 ms
Throughput 1042.03 MB/sec  2 clients  2 procs  max_latency=0.098 ms
Throughput 1953.12 MB/sec  4 clients  4 procs  max_latency=0.236 ms
Throughput 3694.99 MB/sec  8 clients  8 procs  max_latency=0.308 ms
Throughput 7080.95 MB/sec  16 clients  16 procs  max_latency=0.442 ms
Throughput 13444.7 MB/sec  32 clients  32 procs  max_latency=1.417 ms
Throughput 20191.3 MB/sec  64 clients  64 procs  max_latency=4.554 ms
Throughput 41115.4 MB/sec  128 clients  128 procs  max_latency=13.414 ms
Throughput 66844.4 MB/sec  256 clients  256 procs  max_latency=50.069 ms

5226         /*
5227          * If there are idle cores to be had, go find one.
5228          */
5229         if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
5230                 i = select_idle_core(p, target);
5231                 if ((unsigned)i < nr_cpumask_bits)
5232                         return i;
5233  
5234                 /*
5235                  * Failed to find an idle core; stop looking for one.
5236                  */
5237                 clear_idle_cores(target);
5238         }
5239 #if 1
5240         for_each_cpu(i, cpu_smt_mask(target)) {
5241                 if (idle_cpu(i))
5242                         return i;
5243         }
5244  
5245         return target;
5246 #endif
5247  
5248         if (sched_feat(FORCE_CORE)) {

[toc] | [prev] | [next] | [standalone]


#1392268

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-02 17:00 +0200
Message-ID<runuk-5gi-47@gated-at.bofh.it>
In reply to#1392265
On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> Oh btw, did you know single socket boxen have no sd_busy?  That doesn't
> look right.

I suspected; didn't bother looking at yet. The 'problem' is that the LLC
domain is the top-most, so it doesn't have a parent domain. I'm sure we
can come up with something if we can get this all working right.

And yes, I can get gains on various workloads with various options, I
can even break all workloads, but I've so far completely failed on
getting a win for everyone :/

In particular low count sysbench-psql (oltp test) vs tbench
client==nr_cores is having me flummoxed for a bit.

[toc] | [prev] | [next] | [standalone]


#1392320

FromChris Mason <clm@fb.com>
Date2016-05-02 17:50 +0200
Message-ID<ruogG-66V-17@gated-at.bofh.it>
In reply to#1392268
On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote:
> On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> > Oh btw, did you know single socket boxen have no sd_busy?  That doesn't
> > look right.
> 
> I suspected; didn't bother looking at yet. The 'problem' is that the LLC
> domain is the top-most, so it doesn't have a parent domain. I'm sure we
> can come up with something if we can get this all working right.
> 
> And yes, I can get gains on various workloads with various options, I
> can even break all workloads, but I've so far completely failed on
> getting a win for everyone :/

Adding in the task_hot() check to decide if scanning idle was a good
idea ended up being really important.  I'm happy to try a few variations
here as well, do you have a more recent patch?

-chris

[toc] | [prev] | [next] | [standalone]


#1393459

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-03 16:40 +0200
Message-ID<ruJEu-1mn-9@gated-at.bofh.it>
In reply to#1392320
On Mon, May 02, 2016 at 11:47:25AM -0400, Chris Mason wrote:
> On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote:
> > On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> > > Oh btw, did you know single socket boxen have no sd_busy?  That doesn't
> > > look right.
> > 
> > I suspected; didn't bother looking at yet. The 'problem' is that the LLC
> > domain is the top-most, so it doesn't have a parent domain. I'm sure we
> > can come up with something if we can get this all working right.
> > 
> > And yes, I can get gains on various workloads with various options, I
> > can even break all workloads, but I've so far completely failed on
> > getting a win for everyone :/
> 
> Adding in the task_hot() check to decide if scanning idle was a good
> idea ended up being really important

So I'm conflicted on this patch:

+static int bounce_to_target(struct task_struct *p, int cpu)
+{
+       s64 delta;
+
+       /*
+        * as the run queue gets bigger, its more and more likely that
+        * balance will have distributed things for us, and less likely
+        * that scanning all our CPUs for an idle one will find one.
+        * So, if nr_running > 1, just call this CPU good enough
+        */
+       if (cpu_rq(cpu)->cfs.nr_running > 1)
+               return 1;
+
+       /* taken from task_hot() */
+       delta = rq_clock_task(task_rq(p)) - p->se.exec_start;
+       return delta < (s64)sysctl_sched_migration_cost;
+}

This will work for you schbench workload because it sleep for 30ms while
the migration_cost thingy is 500us, therefore you'll trigger the full
LLC scan.

_However_, the migration_cost is supposed the model the cost of leaving
the LLC, so testing against that here seems wrong.

Let me go play with something that measures the cost of doing that LLC
scan and compares that against the sleepy time -- of course, now need to
go figure out how to do this clock thing without rq-lock pain.



+       if (package_sd && !bounce_to_target(p, target)) {
+               for_each_cpu_and(i, sched_domain_span(package_sd), tsk_cpus_allowed(p)) {
+                       if (idle_cpu(i)) {
+                               target = i;
+                               break;
+                       }
+
+               }
+       }

Also note your s/sd/package_sd/ rename is, strictly speaking, wrong.
Sure, on your current Intel system the LLC is the entire package, but
this is not true in general.

Take for instance the Intel Core2Quad and AMD Bulldozer thingies, they
had two dies in one package, and correspondingly two LLC domains in one
package.

(also, the Intel cluster-on-die thing can split the thing in two)

There were also the old P6 era SMP boards which had external LLC, where
you could have an LLC shared across multiple packages -- although I'm
thinking we'll never see that again, due to off package being far
toooooo slooooooow these days.

[toc] | [prev] | [next] | [standalone]


#1393494

FromChris Mason <clm@fb.com>
Date2016-05-03 17:20 +0200
Message-ID<ruKhb-1Y8-1@gated-at.bofh.it>
In reply to#1393459
On Tue, May 03, 2016 at 04:32:25PM +0200, Peter Zijlstra wrote:
> On Mon, May 02, 2016 at 11:47:25AM -0400, Chris Mason wrote:
> > On Mon, May 02, 2016 at 04:58:17PM +0200, Peter Zijlstra wrote:
> > > On Mon, May 02, 2016 at 04:50:04PM +0200, Mike Galbraith wrote:
> > > > Oh btw, did you know single socket boxen have no sd_busy?  That doesn't
> > > > look right.
> > > 
> > > I suspected; didn't bother looking at yet. The 'problem' is that the LLC
> > > domain is the top-most, so it doesn't have a parent domain. I'm sure we
> > > can come up with something if we can get this all working right.
> > > 
> > > And yes, I can get gains on various workloads with various options, I
> > > can even break all workloads, but I've so far completely failed on
> > > getting a win for everyone :/
> > 
> > Adding in the task_hot() check to decide if scanning idle was a good
> > idea ended up being really important
> 
> So I'm conflicted on this patch:
> 
> +static int bounce_to_target(struct task_struct *p, int cpu)
> +{
> +       s64 delta;
> +
> +       /*
> +        * as the run queue gets bigger, its more and more likely that
> +        * balance will have distributed things for us, and less likely
> +        * that scanning all our CPUs for an idle one will find one.
> +        * So, if nr_running > 1, just call this CPU good enough
> +        */
> +       if (cpu_rq(cpu)->cfs.nr_running > 1)
> +               return 1;

The nr_running check is interesting.  It is supposed to give the same
benefit as your "do we have anything idle?" variable, but without having
to constantly update a variable somewhere.  I'll have to do a few runs
to verify (maybe a idle_scan_failed counter).

> +
> +       /* taken from task_hot() */
> +       delta = rq_clock_task(task_rq(p)) - p->se.exec_start;
> +       return delta < (s64)sysctl_sched_migration_cost;
> +}
> 
> This will work for you schbench workload because it sleep for 30ms while
> the migration_cost thingy is 500us, therefore you'll trigger the full
> LLC scan.

The task_hot checks don't do much for the sleeping schbench runs, but
they help a lot for this:

# pick a single core, in my case cpus 0,20 are the same core
# cpu_hog is any program that spins
#
taskset -c 20 cpu_hog &

# schbench -p 4 means message passing mode with 4 byte messages (like
# pipe test), no sleeps, just bouncing as fast as it can.
#
# make the scheduler choose between the sibling of the hog and cpu 1
#
taskset -c 0,1 schbench -p 4 -m 1 -t 1

Current mainline will stuff both schbench threads onto CPU 1, leaving
CPU 0 100% idle.  My first patch with the minimal task_hot() checks
would sometimes pick CPU 0.  My second patch that just directly calls
task_hot sticks to cpu1, which is ~3x faster than spreading it.

The full task_hot() checks also really help tbench.

> 
> _However_, the migration_cost is supposed the model the cost of leaving
> the LLC, so testing against that here seems wrong.
> 
> Let me go play with something that measures the cost of doing that LLC
> scan and compares that against the sleepy time -- of course, now need to
> go figure out how to do this clock thing without rq-lock pain.
> 
> 
> 
> +       if (package_sd && !bounce_to_target(p, target)) {
> +               for_each_cpu_and(i, sched_domain_span(package_sd), tsk_cpus_allowed(p)) {
> +                       if (idle_cpu(i)) {
> +                               target = i;
> +                               break;
> +                       }
> +
> +               }
> +       }
> 
> Also note your s/sd/package_sd/ rename is, strictly speaking, wrong.
> Sure, on your current Intel system the LLC is the entire package, but
> this is not true in general.
> 
> Take for instance the Intel Core2Quad and AMD Bulldozer thingies, they
> had two dies in one package, and correspondingly two LLC domains in one
> package.
> 
> (also, the Intel cluster-on-die thing can split the thing in two)
> 
> There were also the old P6 era SMP boards which had external LLC, where
> you could have an LLC shared across multiple packages -- although I'm
> thinking we'll never see that again, due to off package being far
> toooooo slooooooow these days.

Gotcha, makes sense.  I'll switch to llc_sd ;)

-chris

[toc] | [prev] | [next] | [standalone]


#1394142

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-04 12:40 +0200
Message-ID<rv2nM-28B-15@gated-at.bofh.it>
In reply to#1393494
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:

> > +       if (cpu_rq(cpu)->cfs.nr_running > 1)
> > +               return 1;
> 
> The nr_running check is interesting.  It is supposed to give the same
> benefit as your "do we have anything idle?" variable, but without having
> to constantly update a variable somewhere.  I'll have to do a few runs
> to verify (maybe a idle_scan_failed counter).

Right; I got that.. I tried it and it doesn't seem to work as well. But
yeah, its better than nothing.

The reason I'm not too worried about the update_idle_core() thing is
that SMT threads share L1 so touching their sibling state isn't
typically expensive.

But yes, I'd need to find someone with SMT8 or something daft like that
to try this.

> The task_hot checks don't do much for the sleeping schbench runs, but
> they help a lot for this:
> 
> # pick a single core, in my case cpus 0,20 are the same core
> # cpu_hog is any program that spins
> #
> taskset -c 20 cpu_hog &
> 
> # schbench -p 4 means message passing mode with 4 byte messages (like
> # pipe test), no sleeps, just bouncing as fast as it can.
> #
> # make the scheduler choose between the sibling of the hog and cpu 1
> #
> taskset -c 0,1 schbench -p 4 -m 1 -t 1
> 
> Current mainline will stuff both schbench threads onto CPU 1, leaving
> CPU 0 100% idle.  My first patch with the minimal task_hot() checks
> would sometimes pick CPU 0.  My second patch that just directly calls
> task_hot sticks to cpu1, which is ~3x faster than spreading it.

Urgh, another benchmark to play with ;-)

> The full task_hot() checks also really help tbench.

tbench wants select_idle_siblings() to just not exist; it goes happy
when you just return target.

tbench:

old (mainline like):

Throughput 875.822 MB/sec  2 clients  2 procs  max_latency=0.117 ms
Throughput 2017.57 MB/sec  5 clients  5 procs  max_latency=0.057 ms
Throughput 3954.66 MB/sec  10 clients  10 procs  max_latency=0.094 ms
Throughput 5886.11 MB/sec  20 clients  20 procs  max_latency=0.088 ms
Throughput 9095.57 MB/sec  40 clients  40 procs  max_latency=0.864 ms

new:

Throughput 876.794 MB/sec  2 clients  2 procs  max_latency=0.102 ms
Throughput 2048.73 MB/sec  5 clients  5 procs  max_latency=0.095 ms
Throughput 3802.69 MB/sec  10 clients  10 procs  max_latency=0.113 ms
Throughput 5521.81 MB/sec  20 clients  20 procs  max_latency=0.091 ms
Throughput 10331.8 MB/sec  40 clients  40 procs  max_latency=0.444 ms

nothing:

Throughput 759.532 MB/sec  2 clients  2 procs  max_latency=0.210 ms
Throughput 1884.01 MB/sec  5 clients  5 procs  max_latency=0.094 ms
Throughput 3931.31 MB/sec  10 clients  10 procs  max_latency=0.091 ms
Throughput 6478.81 MB/sec  20 clients  20 procs  max_latency=0.110 ms
Throughput 10001 MB/sec  40 clients  40 procs  max_latency=0.148 ms


See the 20 client have a happy moment ;-) [ivb-ep: 2*10*2]


I've not quite figured out how to make the new bits switch off aggressive
enough to make tbench happy without hurting the others. More numbers:


sysbench-oltp-psql:

old (mainline like):

  2: [30 secs]     transactions:                        53556  (1785.19 per sec.)
  5: [30 secs]     transactions:                        118957 (3965.08 per sec.)
 10: [30 secs]     transactions:                        241126 (8037.22 per sec.)
 20: [30 secs]     transactions:                        383256 (12774.63 per sec.)
 40: [30 secs]     transactions:                        539705 (17989.05 per sec.)
 80: [30 secs]     transactions:                        541833 (18059.16 per sec.)

new:

  2: [30 secs]     transactions:                        53012  (1767.03 per sec.)
  5: [30 secs]     transactions:                        122057 (4068.49 per sec.)
 10: [30 secs]     transactions:                        235781 (7859.09 per sec.)
 20: [30 secs]     transactions:                        355967 (11864.99 per sec.)
 40: [30 secs]     transactions:                        537327 (17909.80 per sec.)
 80: [30 secs]     transactions:                        546017 (18198.82 per sec.)



schbench -m2 -t 20 -c 30000 -s 30000 -r 30:

old (mainline like):

Latency percentiles (usec)
        50.0000th: 102
        75.0000th: 109
        90.0000th: 115
        95.0000th: 118
        *99.0000th: 5352
        99.5000th: 12112
        99.9000th: 13008
        Over=0, min=0, max=27238

new:

Latency percentiles (usec)
        50.0000th: 103
        75.0000th: 109
        90.0000th: 114
        95.0000th: 116
        *99.0000th: 120
        99.5000th: 121
        99.9000th: 124
        Over=0, min=0, max=12939


---
 include/linux/sched.h    |   2 +
 kernel/sched/core.c      |   3 +
 kernel/sched/fair.c      | 267 +++++++++++++++++++++++++++++++++++++++--------
 kernel/sched/features.h  |   9 ++
 kernel/sched/idle_task.c |   4 +-
 kernel/sched/sched.h     |   1 +
 kernel/time/tick-sched.c |  10 +-
 7 files changed, 246 insertions(+), 50 deletions(-)

diff --git a/include/linux/sched.h b/include/linux/sched.h
index ad9454d..e7ce1a0 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1068,6 +1068,8 @@ struct sched_domain {
 	u64 max_newidle_lb_cost;
 	unsigned long next_decay_max_lb_cost;
 
+	u64 avg_scan_cost;		/* select_idle_sibling */
+
 #ifdef CONFIG_SCHEDSTATS
 	/* load_balance() stats */
 	unsigned int lb_count[CPU_MAX_IDLE_TYPES];
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index c82ca6e..280e73e 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -7278,6 +7278,7 @@ static struct kmem_cache *task_group_cache __read_mostly;
 #endif
 
 DECLARE_PER_CPU(cpumask_var_t, load_balance_mask);
+DECLARE_PER_CPU(cpumask_var_t, select_idle_mask);
 
 void __init sched_init(void)
 {
@@ -7314,6 +7315,8 @@ void __init sched_init(void)
 	for_each_possible_cpu(i) {
 		per_cpu(load_balance_mask, i) = (cpumask_var_t)kzalloc_node(
 			cpumask_size(), GFP_KERNEL, cpu_to_node(i));
+		per_cpu(select_idle_mask, i) = (cpumask_var_t)kzalloc_node(
+			cpumask_size(), GFP_KERNEL, cpu_to_node(i));
 	}
 #endif /* CONFIG_CPUMASK_OFFSTACK */
 
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33ab..9290fc8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ balance:
 	 * One idle CPU per node is evaluated for a task numa move.
 	 * Call select_idle_sibling to maybe find a better one.
 	 */
-	if (!cur)
+	if (!cur) {
+		// XXX borken
 		env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+	}
 
 assign:
 	assigned = true;
@@ -4491,6 +4493,11 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 }
 
 #ifdef CONFIG_SMP
+
+/* Working cpumask for: load_balance, load_balance_newidle. */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+DEFINE_PER_CPU(cpumask_var_t, select_idle_mask);
+
 #ifdef CONFIG_NO_HZ_COMMON
 /*
  * per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5169,240 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
 	return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
 }
 
+static int cpumask_next_wrap(int n, const struct cpumask *mask, int start, int *wrapped)
+{
+	int next;
+
+again:
+	next = find_next_bit(cpumask_bits(mask), nr_cpumask_bits, n+1);
+
+	if (*wrapped) {
+		if (next >= start)
+			return nr_cpumask_bits;
+	} else {
+		if (next >= nr_cpumask_bits) {
+			*wrapped = 1;
+			n = -1;
+			goto again;
+		}
+	}
+
+	return next;
+}
+
+#define for_each_cpu_wrap(cpu, mask, start, wrap)				\
+	for ((wrap) = 0, (cpu) = (start)-1;					\
+		(cpu) = cpumask_next_wrap((cpu), (mask), (start), &(wrap)),	\
+		(cpu) < nr_cpumask_bits; )
+
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return false;
+
+	// XXX static key for !SMT topologies
+
+	return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+	int core = cpu_of(rq);
+	int cpu;
+
+	rcu_read_lock();
+	if (test_idle_cores(core))
+		goto unlock;
+
+	for_each_cpu(cpu, cpu_smt_mask(core)) {
+		if (cpu == core)
+			continue;
+
+		if (!idle_cpu(cpu))
+			goto unlock;
+	}
+
+	set_idle_cores(core);
+unlock:
+	rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, struct sched_domain *sd, int target)
+{
+	struct cpumask *cpus = this_cpu_cpumask_var_ptr(select_idle_mask);
+	int core, cpu, wrap;
+
+	if (!test_idle_cores(target))
+		return -1;
+
+	cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+
+	for_each_cpu_wrap(core, cpus, target, wrap) {
+		bool idle = true;
+
+		for_each_cpu(cpu, cpu_smt_mask(core)) {
+			cpumask_clear_cpu(cpu, cpus);
+			if (!idle_cpu(cpu))
+				idle = false;
+		}
+
+		if (idle)
+			return core;
+	}
+
+	/*
+	 * Failed to find an idle core; stop looking for one.
+	 */
+	clear_idle_cores(target);
+
+	return -1;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, struct sched_domain *sd, int target)
+{
+	return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
+static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, int target)
+{
+	struct sched_domain *this_sd = rcu_dereference(*this_cpu_ptr(&sd_llc));
+	u64 time, cost;
+	s64 delta;
+	int cpu, wrap;
+
+	if (sched_feat(AVG_CPU)) {
+		u64 avg_idle = this_rq()->avg_idle;
+		u64 avg_cost = this_sd->avg_scan_cost;
+
+		if (sched_feat(PRINT_AVG))
+			trace_printk("idle: %Ld cost: %Ld\n", avg_idle, avg_cost);
+
+		if (avg_idle / 32 < avg_cost)
+			return -1;
+	}
+
+	time = local_clock();
+
+	for_each_cpu_wrap(cpu, sched_domain_span(sd), target, wrap) {
+		if (!cpumask_test_cpu(cpu, tsk_cpus_allowed(p)))
+			continue;
+		if (idle_cpu(cpu))
+			break;
+	}
+
+	time = local_clock() - time;
+	cost = this_sd->avg_scan_cost;
+	delta = (s64)(time - cost) / 8;
+	/* trace_printk("time: %Ld cost: %Ld delta: %Ld\n", time, cost, delta); */
+	this_sd->avg_scan_cost += delta;
+
+	return cpu;
+}
+
 /*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
  */
 static int select_idle_sibling(struct task_struct *p, int target)
 {
 	struct sched_domain *sd;
-	struct sched_group *sg;
-	int i = task_cpu(p);
+	int start, i = task_cpu(p);
 
 	if (idle_cpu(target))
 		return target;
 
 	/*
-	 * If the prevous cpu is cache affine and idle, don't be stupid.
+	 * If the previous cpu is cache affine and idle, don't be stupid.
 	 */
 	if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
 		return i;
 
-	/*
-	 * Otherwise, iterate the domains and find an eligible idle cpu.
-	 *
-	 * A completely idle sched group at higher domains is more
-	 * desirable than an idle group at a lower level, because lower
-	 * domains have smaller groups and usually share hardware
-	 * resources which causes tasks to contend on them, e.g. x86
-	 * hyperthread siblings in the lowest domain (SMT) can contend
-	 * on the shared cpu pipeline.
-	 *
-	 * However, while we prefer idle groups at higher domains
-	 * finding an idle cpu at the lowest domain is still better than
-	 * returning 'target', which we've already established, isn't
-	 * idle.
-	 */
-	sd = rcu_dereference(per_cpu(sd_llc, target));
-	for_each_lower_domain(sd) {
-		sg = sd->groups;
-		do {
-			if (!cpumask_intersects(sched_group_cpus(sg),
-						tsk_cpus_allowed(p)))
-				goto next;
+	start = target;
+	if (sched_feat(ORDER_IDLE))
+		start = per_cpu(sd_llc_id, target); /* first cpu in llc domain */
 
-			/* Ensure the entire group is idle */
-			for_each_cpu(i, sched_group_cpus(sg)) {
-				if (i == target || !idle_cpu(i))
+	sd = rcu_dereference(per_cpu(sd_llc, start));
+	if (!sd)
+		return target;
+
+	if (sched_feat(OLD_IDLE)) {
+		struct sched_group *sg;
+
+		for_each_lower_domain(sd) {
+			sg = sd->groups;
+			do {
+				if (!cpumask_intersects(sched_group_cpus(sg),
+							tsk_cpus_allowed(p)))
 					goto next;
-			}
 
-			/*
-			 * It doesn't matter which cpu we pick, the
-			 * whole group is idle.
-			 */
-			target = cpumask_first_and(sched_group_cpus(sg),
-					tsk_cpus_allowed(p));
-			goto done;
+				/* Ensure the entire group is idle */
+				for_each_cpu(i, sched_group_cpus(sg)) {
+					if (i == target || !idle_cpu(i))
+						goto next;
+				}
+
+				/*
+				 * It doesn't matter which cpu we pick, the
+				 * whole group is idle.
+				 */
+				target = cpumask_first_and(sched_group_cpus(sg),
+						tsk_cpus_allowed(p));
+				goto done;
 next:
-			sg = sg->next;
-		} while (sg != sd->groups);
-	}
+				sg = sg->next;
+			} while (sg != sd->groups);
+		}
 done:
+		return target;
+	}
+
+	if (sched_feat(IDLE_CORE)) {
+		i = select_idle_core(p, sd, start);
+		if ((unsigned)i < nr_cpumask_bits)
+			return i;
+	}
+
+	if (sched_feat(IDLE_CPU)) {
+		i = select_idle_cpu(p, sd, start);
+		if ((unsigned)i < nr_cpumask_bits)
+			return i;
+	}
+
+	if (sched_feat(IDLE_SMT)) {
+		for_each_cpu(i, cpu_smt_mask(target)) {
+			if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+				continue;
+			if (idle_cpu(i))
+				return i;
+		}
+	}
+
 	return target;
 }
 
@@ -7229,9 +7411,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
  */
 #define MAX_PINNED_INTERVAL	512
 
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
 static int need_active_balance(struct lb_env *env)
 {
 	struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa..5dd10ec 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,12 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
 SCHED_FEAT(LB_MIN, false)
 SCHED_FEAT(ATTACH_AGE_LOAD, true)
 
+SCHED_FEAT(OLD_IDLE, false)
+SCHED_FEAT(ORDER_IDLE, false)
+
+SCHED_FEAT(IDLE_CORE, true)
+SCHED_FEAT(IDLE_CPU, true)
+SCHED_FEAT(AVG_CPU, true)
+SCHED_FEAT(PRINT_AVG, false)
+
+SCHED_FEAT(IDLE_SMT, false)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce949..cb394db 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
 	resched_curr(rq);
 }
 
+extern void update_idle_core(struct rq *rq);
+
 static struct task_struct *
 pick_next_task_idle(struct rq *rq, struct task_struct *prev)
 {
 	put_prev_task(rq, prev);
-
+	update_idle_core(rq);
 	schedstat_inc(rq, sched_goidle);
 	return rq->idle;
 }
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fc..5994794 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
 	 * Number of busy cpus in this group.
 	 */
 	atomic_t nr_busy_cpus;
+	int	has_idle_cores;
 
 	unsigned long cpumask[0]; /* iteration mask */
 };
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc..6e42cd2 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
 	WARN_ON_ONCE(irqs_disabled());
 
 	/*
- 	 * Update the idle state in the scheduler domain hierarchy
- 	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
- 	 * State will be updated to busy during the first busy tick after
- 	 * exiting idle.
- 	 */
+	 * Update the idle state in the scheduler domain hierarchy
+	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
+	 * State will be updated to busy during the first busy tick after
+	 * exiting idle.
+	 */
 	set_cpu_sd_state_idle();
 
 	local_irq_disable();

[toc] | [prev] | [next] | [standalone]


#1394470

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-04 17:40 +0200
Message-ID<rv745-6DR-1@gated-at.bofh.it>
In reply to#1394142
On Wed, May 04, 2016 at 12:37:01PM +0200, Peter Zijlstra wrote:

> +static int select_idle_cpu(struct task_struct *p, struct sched_domain *sd, int target)
> +{
> +	struct sched_domain *this_sd = rcu_dereference(*this_cpu_ptr(&sd_llc));
> +	u64 time, cost;
> +	s64 delta;
> +	int cpu, wrap;
> +
> +	if (sched_feat(AVG_CPU)) {
> +		u64 avg_idle = this_rq()->avg_idle;
> +		u64 avg_cost = this_sd->avg_scan_cost;
> +
> +		if (sched_feat(PRINT_AVG))
> +			trace_printk("idle: %Ld cost: %Ld\n", avg_idle, avg_cost);
> +
> +		if (avg_idle / 32 < avg_cost)

s/32/512/ + IDLE_SMT fixes a hackbench regression

hackbench, like tbench, doesn't like IDLE_CPU to trigger, but apparently
needs IDLE_SMT.

Bah, I could sort of explain 32 away, but 512 is firmly in the magic
value range :/

> +			return -1;
> +	}
> +
> +	time = local_clock();
> +
> +	for_each_cpu_wrap(cpu, sched_domain_span(sd), target, wrap) {
> +		if (!cpumask_test_cpu(cpu, tsk_cpus_allowed(p)))
> +			continue;
> +		if (idle_cpu(cpu))
> +			break;
> +	}
> +
> +	time = local_clock() - time;
> +	cost = this_sd->avg_scan_cost;
> +	delta = (s64)(time - cost) / 8;
> +	/* trace_printk("time: %Ld cost: %Ld delta: %Ld\n", time, cost, delta); */
> +	this_sd->avg_scan_cost += delta;
> +
> +	return cpu;
> +}

[toc] | [prev] | [next] | [standalone]


#1395398

FromMatt Fleming <matt@codeblueprint.co.uk>
Date2016-05-06 00:10 +0200
Message-ID<rvzD4-9z-11@gated-at.bofh.it>
In reply to#1394142
On Wed, 04 May, at 12:37:01PM, Peter Zijlstra wrote:
> 
> tbench wants select_idle_siblings() to just not exist; it goes happy
> when you just return target.

I've been playing with this patch a little bit by hitting it with
tbench on a Xeon, 12 cores with HT enabled, 2 sockets (48 cpus).

I see a throughput improvement for 16, 32, 64, 128 and 256 clients
when compared against mainline, so that's,

  OLD_IDLE, ORDER_IDLE, NO_IDLE_CORE, NO_IDLE_CPU, NO_IDLE_SMT

  vs.

  NO_OLD_IDLE, NO_ORDER_IDLE, IDLE_CORE, IDLE_CPU, IDLE_SMT

See,


 [OLD] Throughput 5345.6 MB/sec   16 clients  16 procs  max_latency=0.277 ms avg_latency=0.211853 ms
 [NEW] Throughput 5514.52 MB/sec  16 clients  16 procs  max_latency=0.493 ms avg_latency=0.176441 ms
 
 [OLD] Throughput 7401.76 MB/sec  32 clients  32 procs  max_latency=1.804 ms avg_latency=0.451147 ms
 [NEW] Throughput 10044.9 MB/sec  32 clients  32 procs  max_latency=3.421 ms avg_latency=0.582529 ms
 
 [OLD] Throughput 13265.9 MB/sec  64 clients  64 procs  max_latency=7.395 ms avg_latency=0.927147 ms
 [NEW] Throughput 13929.6 MB/sec  64 clients  64 procs  max_latency=7.022 ms avg_latency=1.017059 ms
 
 [OLD] Throughput 12827.8 MB/sec  128 clients  128 procs  max_latency=16.256 ms avg_latency=2.763706 ms
 [NEW] Throughput 13364.2 MB/sec  128 clients  128 procs  max_latency=16.630 ms avg_latency=3.002971 ms
 
 [OLD] Throughput 12653.1 MB/sec  256 clients  256 procs  max_latency=44.722 ms avg_latency=5.741647 ms
 [NEW] Throughput 12965.7 MB/sec  256 clients  256 procs  max_latency=59.061 ms avg_latency=8.699118 ms


For throughput changes to 1, 2, 4 and 8 clients it's more of a mixture
with sometimes the old config winning and sometimes losing.


 [OLD] Throughput 488.819 MB/sec  1 clients  1 procs  max_latency=0.191 ms avg_latency=0.058794 ms
 [NEW] Throughput 486.106 MB/sec  1 clients  1 procs  max_latency=0.085 ms avg_latency=0.045794 ms
 
 [OLD] Throughput 925.987 MB/sec  2 clients  2 procs  max_latency=0.201 ms avg_latency=0.090882 ms
 [NEW] Throughput 954.944 MB/sec  2 clients  2 procs  max_latency=0.199 ms avg_latency=0.064294 ms
 
 [OLD] Throughput 1764.02 MB/sec  4 clients  4 procs  max_latency=0.160 ms avg_latency=0.075206 ms
 [NEW] Throughput 1756.8 MB/sec   4 clients  4 procs  max_latency=0.105 ms avg_latency=0.062382 ms
 
 [OLD] Throughput 3384.22 MB/sec  8 clients  8 procs  max_latency=0.276 ms avg_latency=0.099441 ms
 [NEW] Throughput 3375.47 MB/sec  8 clients  8 procs  max_latency=0.103 ms avg_latency=0.064176 ms


Looking at latency, the new code consistently performs worse at the
top end for 256 clients. Admittedly at that point the machine is
pretty overloaded. Things are much better at the lower end.

One thing I haven't yet done is twiddled the bits individually to see
what the best combination is. Have you settled on the right settings
yet?

[toc] | [prev] | [next] | [standalone]


#1396071

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-06 21:00 +0200
Message-ID<rvT8K-20P-17@gated-at.bofh.it>
In reply to#1395398
On Thu, 2016-05-05 at 23:03 +0100, Matt Fleming wrote:

> One thing I haven't yet done is twiddled the bits individually to see
> what the best combination is. Have you settled on the right settings
> yet?

Lighter configs, revert sched/fair: Fix fairness issue on migration,
twiddle knobs.  Added an IDLE_SIBLING knob to ~virgin master.. only
sorta virgin because I always throttle nohz.

1 x i4790
master
for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done
Throughput 871.785 MB/sec  1 clients  1 procs  max_latency=0.324 ms
Throughput 1514.5 MB/sec  2 clients  2 procs  max_latency=0.411 ms
Throughput 2722.43 MB/sec  4 clients  4 procs  max_latency=2.400 ms
Throughput 4334.46 MB/sec  8 clients  8 procs  max_latency=3.561 ms

echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features
Throughput 1078.69 MB/sec  1 clients  1 procs  max_latency=2.274 ms
Throughput 2130.33 MB/sec  2 clients  2 procs  max_latency=1.451 ms
Throughput 3484.18 MB/sec  4 clients  4 procs  max_latency=3.430 ms
Throughput 4423.69 MB/sec  8 clients  8 procs  max_latency=5.363 ms


masterx
for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done
Throughput 707.673 MB/sec  1 clients  1 procs  max_latency=2.279 ms
Throughput 1503.55 MB/sec  2 clients  2 procs  max_latency=0.695 ms
Throughput 2527.73 MB/sec  4 clients  4 procs  max_latency=2.321 ms
Throughput 4291.26 MB/sec  8 clients  8 procs  max_latency=3.815 ms

echo NO_IDLE_CPU > /sys/kernel/debug/sched_features
homer:~ # for i in 1 2 4 8; do tbench.sh $i 30 2>&1|grep Throughput; done
Throughput 865.936 MB/sec  1 clients  1 procs  max_latency=0.411 ms
Throughput 1586.41 MB/sec  2 clients  2 procs  max_latency=2.293 ms
Throughput 2638.39 MB/sec  4 clients  4 procs  max_latency=2.037 ms
Throughput 4405.43 MB/sec  8 clients  8 procs  max_latency=3.581 ms

+ echo NO_AVG_CPU > /sys/kernel/debug/sched_features
+ echo IDLE_SMT > /sys/kernel/debug/sched_features
Throughput 697.126 MB/sec  1 clients  1 procs  max_latency=2.220 ms
Throughput 1562.82 MB/sec  2 clients  2 procs  max_latency=0.526 ms
Throughput 2620.62 MB/sec  4 clients  4 procs  max_latency=6.460 ms
Throughput 4345.13 MB/sec  8 clients  8 procs  max_latency=27.921 ms


4 x E7-8890
master
for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done
Throughput 615.663 MB/sec  1 clients  1 procs  max_latency=0.087 ms
Throughput 1171.53 MB/sec  2 clients  2 procs  max_latency=0.087 ms
Throughput 2251.22 MB/sec  4 clients  4 procs  max_latency=0.078 ms
Throughput 4090.76 MB/sec  8 clients  8 procs  max_latency=0.801 ms
Throughput 7695.92 MB/sec  16 clients  16 procs  max_latency=0.235 ms
Throughput 15152 MB/sec  32 clients  32 procs  max_latency=0.693 ms
Throughput 21628.2 MB/sec  64 clients  64 procs  max_latency=4.666 ms
Throughput 43185.7 MB/sec  128 clients  128 procs  max_latency=7.280 ms
Throughput 72144.5 MB/sec  256 clients  256 procs  max_latency=8.194 ms

echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features
Throughput 954.593 MB/sec  1 clients  1 procs  max_latency=0.185 ms
Throughput 1882.65 MB/sec  2 clients  2 procs  max_latency=0.278 ms
Throughput 3457.03 MB/sec  4 clients  4 procs  max_latency=0.431 ms
Throughput 6279.38 MB/sec  8 clients  8 procs  max_latency=0.730 ms
Throughput 11170.4 MB/sec  16 clients  16 procs  max_latency=0.500 ms
Throughput 21940.9 MB/sec  32 clients  32 procs  max_latency=0.475 ms
Throughput 41738.8 MB/sec  64 clients  64 procs  max_latency=3.669 ms
Throughput 67634.6 MB/sec  128 clients  128 procs  max_latency=6.676 ms
Throughput 76299.7 MB/sec  256 clients  256 procs  max_latency=7.878 ms

masterx
for i in 1 2 4 8 16 32 64 128 256; do tbench.sh $i 30 2>&1| grep Throughput; done
Throughput 587.956 MB/sec  1 clients  1 procs  max_latency=0.124 ms
Throughput 1140.16 MB/sec  2 clients  2 procs  max_latency=0.476 ms
Throughput 2296.03 MB/sec  4 clients  4 procs  max_latency=0.142 ms
Throughput 4116.65 MB/sec  8 clients  8 procs  max_latency=0.464 ms
Throughput 7820.27 MB/sec  16 clients  16 procs  max_latency=0.238 ms
Throughput 14899.2 MB/sec  32 clients  32 procs  max_latency=0.321 ms
Throughput 21909.8 MB/sec  64 clients  64 procs  max_latency=0.905 ms
Throughput 35495.2 MB/sec  128 clients  128 procs  max_latency=6.158 ms
Throughput 75863.2 MB/sec  256 clients  256 procs  max_latency=7.650 ms

echo NO_IDLE_CPU > /sys/kernel/debug/sched_features
Throughput 555.15 MB/sec  1 clients  1 procs  max_latency=0.096 ms
Throughput 1195.12 MB/sec  2 clients  2 procs  max_latency=0.131 ms
Throughput 2276.97 MB/sec  4 clients  4 procs  max_latency=0.105 ms
Throughput 4248.14 MB/sec  8 clients  8 procs  max_latency=0.131 ms
Throughput 7860.86 MB/sec  16 clients  16 procs  max_latency=0.210 ms
Throughput 15178.6 MB/sec  32 clients  32 procs  max_latency=0.229 ms
Throughput 21523.9 MB/sec  64 clients  64 procs  max_latency=0.842 ms
Throughput 31082.1 MB/sec  128 clients  128 procs  max_latency=7.311 ms
Throughput 75887.9 MB/sec  256 clients  256 procs  max_latency=7.764 ms

+ echo NO_AVG_CPU > /sys/kernel/debug/sched_features
Throughput 598.063 MB/sec  1 clients  1 procs  max_latency=0.131 ms
Throughput 1140.2 MB/sec  2 clients  2 procs  max_latency=0.092 ms
Throughput 2268.68 MB/sec  4 clients  4 procs  max_latency=0.170 ms
Throughput 4259.7 MB/sec  8 clients  8 procs  max_latency=0.212 ms
Throughput 7904.15 MB/sec  16 clients  16 procs  max_latency=0.191 ms
Throughput 14840 MB/sec  32 clients  32 procs  max_latency=0.279 ms
Throughput 21701.5 MB/sec  64 clients  64 procs  max_latency=0.856 ms
Throughput 38945 MB/sec  128 clients  128 procs  max_latency=7.501 ms
Throughput 75669.4 MB/sec  256 clients  256 procs  max_latency=14.984 ms

+ echo IDLE_SMT > /sys/kernel/debug/sched_features
Throughput 592.799 MB/sec  1 clients  1 procs  max_latency=0.120 ms
Throughput 1208.28 MB/sec  2 clients  2 procs  max_latency=0.078 ms
Throughput 2319.22 MB/sec  4 clients  4 procs  max_latency=0.141 ms
Throughput 4196.64 MB/sec  8 clients  8 procs  max_latency=0.253 ms
Throughput 7816.47 MB/sec  16 clients  16 procs  max_latency=0.117 ms
Throughput 14990.8 MB/sec  32 clients  32 procs  max_latency=0.189 ms
Throughput 21809.4 MB/sec  64 clients  64 procs  max_latency=0.832 ms
Throughput 44813 MB/sec  128 clients  128 procs  max_latency=7.930 ms
Throughput 75978.1 MB/sec  256 clients  256 procs  max_latency=7.337 ms

[toc] | [prev] | [next] | [standalone]


#1396847

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-09 10:40 +0200
Message-ID<rwOTo-1x9-19@gated-at.bofh.it>
In reply to#1396071
On Fri, May 06, 2016 at 08:54:38PM +0200, Mike Galbraith wrote:
> master
> Throughput 2722.43 MB/sec  4 clients  4 procs  max_latency=2.400 ms

> echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features
> Throughput 3484.18 MB/sec  4 clients  4 procs  max_latency=3.430 ms

Yeah, I know about that bump, I just haven't managed to find a way to
preserve that and keep all the other benchmarks ticking along :/

[toc] | [prev] | [next] | [standalone]


#1396862

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-09 11:00 +0200
Message-ID<rwPcK-1Kc-17@gated-at.bofh.it>
In reply to#1396847
On Mon, 2016-05-09 at 10:33 +0200, Peter Zijlstra wrote:
> On Fri, May 06, 2016 at 08:54:38PM +0200, Mike Galbraith wrote:
> > master
> > Throughput 2722.43 MB/sec  4 clients  4 procs  max_latency=2.400 ms
> 
> > echo NO_IDLE_SIBLING > /sys/kernel/debug/sched_features
> > Throughput 3484.18 MB/sec  4 clients  4 procs  max_latency=3.430 ms
> 
> Yeah, I know about that bump, I just haven't managed to find a way to
> preserve that and keep all the other benchmarks ticking along :/

Yup, L3 ain't L2.  I haven't come up with a good metric either. Poo. 
 Until that comes along, microbenchmarks can bugger off, real boxen
don't just play high speed ping-pong with themselves for a living ;-)

	-Mike

[toc] | [prev] | [next] | [standalone]


#1394487

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-04 17:50 +0200
Message-ID<rv7dM-6IW-9@gated-at.bofh.it>
In reply to#1393494
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> # pick a single core, in my case cpus 0,20 are the same core
> # cpu_hog is any program that spins
> #
> taskset -c 20 cpu_hog &
> 
> # schbench -p 4 means message passing mode with 4 byte messages (like
> # pipe test), no sleeps, just bouncing as fast as it can.
> #
> # make the scheduler choose between the sibling of the hog and cpu 1
> #
> taskset -c 0,1 schbench -p 4 -m 1 -t 1

Will that schbench thingy print something? Mine doesn't seem to output
anything, not actually exit, although it stops consuming CPU cycles at
some point.

[toc] | [prev] | [next] | [standalone]


#1394580

FromChris Mason <clm@fb.com>
Date2016-05-04 19:50 +0200
Message-ID<rv95U-8tM-15@gated-at.bofh.it>
In reply to#1394487
On Wed, May 04, 2016 at 05:45:10PM +0200, Peter Zijlstra wrote:
> On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> > # pick a single core, in my case cpus 0,20 are the same core
> > # cpu_hog is any program that spins
> > #
> > taskset -c 20 cpu_hog &
> > 
> > # schbench -p 4 means message passing mode with 4 byte messages (like
> > # pipe test), no sleeps, just bouncing as fast as it can.
> > #
> > # make the scheduler choose between the sibling of the hog and cpu 1
> > #
> > taskset -c 0,1 schbench -p 4 -m 1 -t 1
> 
> Will that schbench thingy print something? Mine doesn't seem to output
> anything, not actually exit, although it stops consuming CPU cycles at
> some point.
> 
> 

It should, make sure you're at the top commit in git.

git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git

It's not recent so I'd be surprised if you weren't already there.  The
default runtime is 30 seconds, but you can use -r to specify something
shorter.

It's possible I'm missing a wakeup to shut the whole thing down, but I
thought I fixed that.

 ./schbench -p 4 -m 1 -t 1
Latency percentiles (usec)
        50.0000th: 5
        75.0000th: 5
        90.0000th: 5
        95.0000th: 5
        *99.0000th: 8
        99.5000th: 15
        99.9000th: 17
        Over=0, min=0, max=652
avg worker transfer: 113768.27 ops/sec 444.41KB/s

-chris

[toc] | [prev] | [next] | [standalone]


#1394973

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-05 11:40 +0200
Message-ID<rvnVg-5Rb-23@gated-at.bofh.it>
In reply to#1394580
On Wed, May 04, 2016 at 01:46:16PM -0400, Chris Mason wrote:
> It should, make sure you're at the top commit in git.
> 
> git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git

I did double check; I am on the top commit of that. I refetched and
rebuild just to make tripple sure.

> It's not recent so I'd be surprised if you weren't already there.  The
> default runtime is 30 seconds, but you can use -r to specify something
> shorter.
> 
> It's possible I'm missing a wakeup to shut the whole thing down, but I
> thought I fixed that.

Seems to still be missing, because:

>  ./schbench -p 4 -m 1 -t 1
> Latency percentiles (usec)
>         50.0000th: 5
>         75.0000th: 5
>         90.0000th: 5
>         95.0000th: 5
>         *99.0000th: 8
>         99.5000th: 15
>         99.9000th: 17
>         Over=0, min=0, max=652
> avg worker transfer: 113768.27 ops/sec 444.41KB/s

is not what mine does. I get ~25sec of cpu time and then it stalls
forever.

I'll try and have a prod at the program itself if you have no pending
changes on your end.

[toc] | [prev] | [next] | [standalone]


#1395137

FromChris Mason <clm@fb.com>
Date2016-05-05 16:00 +0200
Message-ID<rvrYS-113-15@gated-at.bofh.it>
In reply to#1394973
On Thu, May 05, 2016 at 11:33:38AM +0200, Peter Zijlstra wrote:
> On Wed, May 04, 2016 at 01:46:16PM -0400, Chris Mason wrote:
> > It should, make sure you're at the top commit in git.
> > 
> > git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git
> 
> I did double check; I am on the top commit of that. I refetched and
> rebuild just to make tripple sure.
> 
> > It's not recent so I'd be surprised if you weren't already there.  The
> > default runtime is 30 seconds, but you can use -r to specify something
> > shorter.
> > 
> > It's possible I'm missing a wakeup to shut the whole thing down, but I
> > thought I fixed that.
> 
> Seems to still be missing, because:
> 
> >  ./schbench -p 4 -m 1 -t 1
> > Latency percentiles (usec)
> >         50.0000th: 5
> >         75.0000th: 5
> >         90.0000th: 5
> >         95.0000th: 5
> >         *99.0000th: 8
> >         99.5000th: 15
> >         99.9000th: 17
> >         Over=0, min=0, max=652
> > avg worker transfer: 113768.27 ops/sec 444.41KB/s
> 
> is not what mine does. I get ~25sec of cpu time and then it stalls
> forever.
> 
> I'll try and have a prod at the program itself if you have no pending
> changes on your end.

Sorry, I don't.  Look at sleep_for_runtime() and how I test/set the
global stopping variable in different places.  I've almost certainly got
someone waiting on a wakeup that'll never come.

If all else fails, run_msg_thread() can pass a timeout to fwait() for a
less error prone setup. I was just hoping to avoid the timers kernel side.

-chris

[toc] | [prev] | [next] | [standalone]


#1395653

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-06 09:20 +0200
Message-ID<rvIdj-JF-7@gated-at.bofh.it>
In reply to#1395137
On Thu, May 05, 2016 at 09:58:44AM -0400, Chris Mason wrote:
> > I'll try and have a prod at the program itself if you have no pending
> > changes on your end.
> 
> Sorry, I don't.  Look at sleep_for_runtime() and how I test/set the
> global stopping variable in different places.  I've almost certainly got
> someone waiting on a wakeup that'll never come.

The below makes it go..

---
 schbench.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/schbench.c b/schbench.c
index a0e9f7e..f299959 100644
--- a/schbench.c
+++ b/schbench.c
@@ -49,7 +49,7 @@ static int pipe_test = 0;
 static unsigned int max_us = 50000;
 
 /* the message threads flip this to true when they decide runtime is up */
-static unsigned long stopping = 0;
+static volatile unsigned long stopping = 0;
 
 
 /*
@@ -746,8 +746,8 @@ static void sleep_for_runtime()
 		else
 			break;
 	}
-	stopping = 1;
 	__sync_synchronize();
+	stopping = 1;
 }
 
 int main(int ac, char **av)

[toc] | [prev] | [next] | [standalone]


#1395998

FromChris Mason <clm@fb.com>
Date2016-05-06 19:30 +0200
Message-ID<rvRJF-YG-15@gated-at.bofh.it>
In reply to#1395653
On Fri, May 06, 2016 at 09:12:51AM +0200, Peter Zijlstra wrote:
> ontent-Length: 973
> 
> On Thu, May 05, 2016 at 09:58:44AM -0400, Chris Mason wrote:
> > > I'll try and have a prod at the program itself if you have no pending
> > > changes on your end.
> > 
> > Sorry, I don't.  Look at sleep_for_runtime() and how I test/set the
> > global stopping variable in different places.  I've almost certainly got
> > someone waiting on a wakeup that'll never come.
> 
> The below makes it go..

Thanks Peter, pushed to git.

-chris

[toc] | [prev] | [next] | [standalone]


#1395667

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-06 09:30 +0200
Message-ID<rvIn0-P0-33@gated-at.bofh.it>
In reply to#1393494
On Tue, May 03, 2016 at 11:11:53AM -0400, Chris Mason wrote:
> # pick a single core, in my case cpus 0,20 are the same core
> # cpu_hog is any program that spins
> #
> taskset -c 20 cpu_hog &
> 
> # schbench -p 4 means message passing mode with 4 byte messages (like
> # pipe test), no sleeps, just bouncing as fast as it can.
> #
> # make the scheduler choose between the sibling of the hog and cpu 1
> #
> taskset -c 0,1 schbench -p 4 -m 1 -t 1
> 
> Current mainline will stuff both schbench threads onto CPU 1, leaving
> CPU 0 100% idle.  My first patch with the minimal task_hot() checks
> would sometimes pick CPU 0.  My second patch that just directly calls
> task_hot sticks to cpu1, which is ~3x faster than spreading it.

Ok, with the thing fixed, my current patch seems to DTRT. If I trace
sched_migrate_task() I get:

$ grep schbench trace

 doit-schbench-2-4042  [004] d..3 144541.309747: sched_migrate_task: comm=doit-schbench-2 pid=4042 prio=120 orig_cpu=4 dest_cpu=4
 doit-schbench-2-4042  [004] d..2 144541.309772: sched_migrate_task: comm=doit-schbench-2 pid=4043 prio=120 orig_cpu=4 dest_cpu=11
 doit-schbench-2-4042  [004] d..3 144541.309855: sched_migrate_task: comm=doit-schbench-2 pid=4042 prio=120 orig_cpu=4 dest_cpu=4
 doit-schbench-2-4042  [004] d..2 144541.309882: sched_migrate_task: comm=doit-schbench-2 pid=4044 prio=120 orig_cpu=4 dest_cpu=5
    migration/11-77    [011] d..4 144541.309974: sched_migrate_task: comm=doit-schbench-2 pid=4043 prio=120 orig_cpu=11 dest_cpu=12
     migration/5-40    [005] d..4 144541.310013: sched_migrate_task: comm=doit-schbench-2 pid=4044 prio=120 orig_cpu=5 dest_cpu=6
        schbench-4044  [001] d..3 144541.310995: sched_migrate_task: comm=schbench pid=4044 prio=120 orig_cpu=1 dest_cpu=1
        schbench-4044  [001] d..2 144541.310999: sched_migrate_task: comm=schbench pid=4045 prio=120 orig_cpu=1 dest_cpu=1
        schbench-4045  [001] d..3 144541.311232: sched_migrate_task: comm=schbench pid=4045 prio=120 orig_cpu=1 dest_cpu=1
        schbench-4045  [001] d..2 144541.311234: sched_migrate_task: comm=schbench pid=4046 prio=120 orig_cpu=1 dest_cpu=1

So the thing gets put on cpu1 and never leaves.

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | linux.kernel


csiph-web