Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1391619 > unrolled thread

Re: sched: tweak select_idle_sibling to look for idle threads

Started byPeter Zijlstra <peterz@infradead.org>
First post2016-04-30 14:50 +0200
Last post2016-05-02 17:20 +0200
Articles 20 on this page of 46 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-04-30 14:50 +0200
    Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 09:20 +0200
      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-01 11:00 +0200
        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 11:30 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-07 11:10 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-08 10:10 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 04:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 05:50 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:10 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 09:50 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 11:00 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:40 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-10 09:10 +0200
                            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-10 09:50 +0200
                              Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-10 17:30 +0200
                                Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 05:00 +0200
                                  Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 06:20 +0200
                                    Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 11:10 +0200
                                      Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 12:00 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 06:00 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:20 +0200
      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 10:50 +0200
        Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 17:00 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:00 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-02 17:50 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 16:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-03 17:20 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 12:40 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:40 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-06 00:10 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-06 21:00 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-09 10:40 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:00 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:50 +0200
                    Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-04 19:50 +0200
                      Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-05 11:40 +0200
                        Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-05 16:00 +0200
                          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:20 +0200
                            Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-06 19:30 +0200
                  Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:30 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 19:40 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:10 +0200
            Re: sched: tweak select_idle_sibling to look for idle threads Ingo Molnar <mingo@kernel.org> - 2016-05-02 18:10 +0200
              Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 13:40 +0200
                Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 20:30 +0200
          Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:20 +0200

Page 1 of 3  [1] 2 3  Next page →


#1391619 — Re: sched: tweak select_idle_sibling to look for idle threads

FromPeter Zijlstra <peterz@infradead.org>
Date2016-04-30 14:50 +0200
SubjectRe: sched: tweak select_idle_sibling to look for idle threads
Message-ID<rtCvn-7R6-1@gated-at.bofh.it>
On Sat, Apr 09, 2016 at 03:05:54PM -0400, Chris Mason wrote:
> select_task_rq_fair() can leave cpu utilization a little lumpy,
> especially as the workload ramps up to the maximum capacity of the
> machine.  The end result can be high p99 response times as apps
> wait to get scheduled, even when boxes are mostly idle.
> 
> I wrote schbench to try and measure this:
> 
> git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git

Can you guys have a play with this; I think one and two node tbench are
good, but I seem to be getting significant run to run variance on that,
so maybe I'm not doing it right.

schbench numbers with: ./schbench -m2 -t 20 -c 30000 -s 30000 -r 30
on my ivb-ep (2 sockets, 10 cores/socket, 2 threads/core) appear to be
decent.

I've also not ran anything other than schbench/tbench so maybe I
completely wrecked something else (as per usual..).

I've not thought about that bounce_to_target() thing much.. I'll go give
that a ponder.

---
 kernel/sched/fair.c      | 180 +++++++++++++++++++++++++++++++++++------------
 kernel/sched/features.h  |   1 +
 kernel/sched/idle_task.c |   4 +-
 kernel/sched/sched.h     |   1 +
 kernel/time/tick-sched.c |  10 +--
 5 files changed, 146 insertions(+), 50 deletions(-)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33abce650..b9d8d1dc5183 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ static void task_numa_compare(struct task_numa_env *env,
 	 * One idle CPU per node is evaluated for a task numa move.
 	 * Call select_idle_sibling to maybe find a better one.
 	 */
-	if (!cur)
+	if (!cur) {
+		// XXX borken
 		env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+	}
 
 assign:
 	assigned = true;
@@ -4491,6 +4493,17 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
 }
 
 #ifdef CONFIG_SMP
+
+/*
+ * Working cpumask for:
+ *   load_balance,
+ *   load_balance_newidle,
+ *   select_idle_core.
+ *
+ * Assumes softirqs are disabled when in use.
+ */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+
 #ifdef CONFIG_NO_HZ_COMMON
 /*
  * per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5175,147 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
 	return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
 }
 
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return;
+
+	WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+	struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+	if (!sd)
+		return false;
+
+	// XXX static key for !SMT topologies
+
+	return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+	int core = cpu_of(rq);
+	int cpu;
+
+	rcu_read_lock();
+	if (test_idle_cores(core))
+		goto unlock;
+
+	for_each_cpu(cpu, cpu_smt_mask(core)) {
+		if (cpu == core)
+			continue;
+
+		if (!idle_cpu(cpu))
+			goto unlock;
+	}
+
+	set_idle_cores(core);
+unlock:
+	rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, int target)
+{
+	struct cpumask *cpus = this_cpu_cpumask_var_ptr(load_balance_mask);
+	struct sched_domain *sd;
+	int core, cpu;
+
+	sd = rcu_dereference(per_cpu(sd_llc, target));
+	cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+	for_each_cpu(core, cpus) {
+		bool idle = true;
+
+		for_each_cpu(cpu, cpu_smt_mask(core)) {
+			cpumask_clear_cpu(cpu, cpus);
+			if (!idle_cpu(cpu))
+				idle = false;
+		}
+
+		if (idle)
+			break;
+	}
+
+	return core;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+static inline void clear_idle_cores(int cpu) { }
+static inline void set_idle_cores(int cpu) { }
+
+static inline bool test_idle_cores(int cpu)
+{
+	return false;
+}
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, int target)
+{
+	return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
 /*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
  */
 static int select_idle_sibling(struct task_struct *p, int target)
 {
 	struct sched_domain *sd;
-	struct sched_group *sg;
 	int i = task_cpu(p);
 
 	if (idle_cpu(target))
 		return target;
 
 	/*
-	 * If the prevous cpu is cache affine and idle, don't be stupid.
+	 * If the previous cpu is cache affine and idle, don't be stupid.
 	 */
 	if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
 		return i;
 
+	sd = rcu_dereference(per_cpu(sd_llc, target));
+	if (!sd)
+		return target;
+
 	/*
-	 * Otherwise, iterate the domains and find an eligible idle cpu.
-	 *
-	 * A completely idle sched group at higher domains is more
-	 * desirable than an idle group at a lower level, because lower
-	 * domains have smaller groups and usually share hardware
-	 * resources which causes tasks to contend on them, e.g. x86
-	 * hyperthread siblings in the lowest domain (SMT) can contend
-	 * on the shared cpu pipeline.
-	 *
-	 * However, while we prefer idle groups at higher domains
-	 * finding an idle cpu at the lowest domain is still better than
-	 * returning 'target', which we've already established, isn't
-	 * idle.
+	 * If there are idle cores to be had, go find one.
 	 */
-	sd = rcu_dereference(per_cpu(sd_llc, target));
-	for_each_lower_domain(sd) {
-		sg = sd->groups;
-		do {
-			if (!cpumask_intersects(sched_group_cpus(sg),
-						tsk_cpus_allowed(p)))
-				goto next;
-
-			/* Ensure the entire group is idle */
-			for_each_cpu(i, sched_group_cpus(sg)) {
-				if (i == target || !idle_cpu(i))
-					goto next;
-			}
+	if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
+		i = select_idle_core(p, target);
+		if ((unsigned)i < nr_cpumask_bits)
+			return i;
 
-			/*
-			 * It doesn't matter which cpu we pick, the
-			 * whole group is idle.
-			 */
-			target = cpumask_first_and(sched_group_cpus(sg),
-					tsk_cpus_allowed(p));
-			goto done;
-next:
-			sg = sg->next;
-		} while (sg != sd->groups);
+		/*
+		 * Failed to find an idle core; stop looking for one.
+		 */
+		clear_idle_cores(target);
 	}
-done:
+
+	/*
+	 * Otherwise, settle for anything idle in this cache domain.
+	 */
+	for_each_cpu(i, sched_domain_span(sd)) {
+		if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+			continue;
+		if (idle_cpu(i))
+			return i;
+	}
+
 	return target;
 }
 
@@ -7229,9 +7324,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
  */
 #define MAX_PINNED_INTERVAL	512
 
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
 static int need_active_balance(struct lb_env *env)
 {
 	struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa46c2f..76bb8814649a 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,4 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
 SCHED_FEAT(LB_MIN, false)
 SCHED_FEAT(ATTACH_AGE_LOAD, true)
 
+SCHED_FEAT(IDLE_CORE, true)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce94931f1b..cb394db407e4 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
 	resched_curr(rq);
 }
 
+extern void update_idle_core(struct rq *rq);
+
 static struct task_struct *
 pick_next_task_idle(struct rq *rq, struct task_struct *prev)
 {
 	put_prev_task(rq, prev);
-
+	update_idle_core(rq);
 	schedstat_inc(rq, sched_goidle);
 	return rq->idle;
 }
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fcaa0e8..5994794bfc85 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
 	 * Number of busy cpus in this group.
 	 */
 	atomic_t nr_busy_cpus;
+	int	has_idle_cores;
 
 	unsigned long cpumask[0]; /* iteration mask */
 };
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc53bc4..6e42cd218ba5 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
 	WARN_ON_ONCE(irqs_disabled());
 
 	/*
- 	 * Update the idle state in the scheduler domain hierarchy
- 	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
- 	 * State will be updated to busy during the first busy tick after
- 	 * exiting idle.
- 	 */
+	 * Update the idle state in the scheduler domain hierarchy
+	 * when tick_nohz_stop_sched_tick() is called from the idle loop.
+	 * State will be updated to busy during the first busy tick after
+	 * exiting idle.
+	 */
 	set_cpu_sd_state_idle();
 
 	local_irq_disable();

[toc] | [next] | [standalone]


#1391757

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-01 09:20 +0200
Message-ID<rtTPz-5sM-1@gated-at.bofh.it>
In reply to#1391619
On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote:
> On Sat, Apr 09, 2016 at 03:05:54PM -0400, Chris Mason wrote:
> > select_task_rq_fair() can leave cpu utilization a little lumpy,
> > especially as the workload ramps up to the maximum capacity of the
> > machine.  The end result can be high p99 response times as apps
> > wait to get scheduled, even when boxes are mostly idle.
> > 
> > I wrote schbench to try and measure this:
> > 
> > git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git
> 
> Can you guys have a play with this; I think one and two node tbench are
> good, but I seem to be getting significant run to run variance on that,
> so maybe I'm not doing it right.

Nah, tbench is just variance prone.  It got dinged up at clients=cores
on my desktop box, on 4 sockets the high end got seriously dinged up.

tbench

1 x i4790
master                          avg
1     714       684     688     695   1.000
2    1260      1234    1284    1259   1.000
4    2238      2301    2286    2275   1.000
8    3388      3418    3396    3400   1.000

masterx
1     690       701     701     697   1.002
2    1287      1332    1235    1284   1.019
4    2014      2006    1999    2006   0.881
8    3388      3385    3404    3392   0.997

4 x E7-8890
master                          avg
1     524      524      523     523   1.000
2    1049     1053     1045    1049   1.000
4    2064     2081     2091    2078   1.000
8    3737     3813     3746    3765   1.000
16   7129     7028     7082    7079   1.000
32  13718    13730    13578   13675   1.000
64  21397    21435    21519   21450   1.000
128 39846    38397    39026   39089   1.000
256 59509    59797    59344   59550   1.000

masterx                         avg
1     505      507      501     504   0.963   1.000
2    1036     1027     1039    1034   0.985   1.000
4    1977     2001     1992    1990   0.957   1.000
8    3734     3802     3778    3771   1.001   1.000
16   7124     7079     7071    7091   1.001   1.000
32  13549    13758    13364   13557   0.991   1.000
64  21975    22161    22100   22078   1.029   1.000
128 23066    23044    23028   23046   0.589   1.000
256 29905    29630    30753   30096   0.505   1.000

masterx NO_IDLE_CORE            avg
1     500      521      502     507   0.969   1.005
2    1012      996     1043    1017   0.969   0.983
4    1988     1992     1995    1991   0.958   1.000
8    3834     3758     3671    3754   0.997   0.995
16   7160     7206     7168    7178   1.013   1.012
32  13788    13773    13672   13744   1.005   1.013
64  21771    21845    21826   21814   1.016   0.988
128 23248    23136    23133   23172   0.592   1.005
256 28683    30013    31850   30182   0.506   1.002

[toc] | [prev] | [next] | [standalone]


#1391776

FromPeter Zijlstra <peterz@infradead.org>
Date2016-05-01 11:00 +0200
Message-ID<rtVom-6wL-9@gated-at.bofh.it>
In reply to#1391757
On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote:

> > Can you guys have a play with this; I think one and two node tbench are
> > good, but I seem to be getting significant run to run variance on that,
> > so maybe I'm not doing it right.
> 
> Nah, tbench is just variance prone.  It got dinged up at clients=cores
> on my desktop box, on 4 sockets the high end got seriously dinged up.

Ouch, yeah, big hurt. Lets try that again... :-)

[toc] | [prev] | [next] | [standalone]


#1391778

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-01 11:30 +0200
Message-ID<rtVRo-7cO-1@gated-at.bofh.it>
In reply to#1391776
On Sun, 2016-05-01 at 10:53 +0200, Peter Zijlstra wrote:
> On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> > On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote:
> 
> > > Can you guys have a play with this; I think one and two node tbench are
> > > good, but I seem to be getting significant run to run variance on that,
> > > so maybe I'm not doing it right.
> > 
> > Nah, tbench is just variance prone.  It got dinged up at clients=cores
> > on my desktop box, on 4 sockets the high end got seriously dinged up.
> 
> Ouch, yeah, big hurt. Lets try that again... :-)

Yeah, box could use a little bandaid and a hug :)

Playing with Chris' benchmark, seems the biggest problem is that we
don't buddy up waker of many and it's wakees in a node.. ie the wake
wide thing isn't necessarily our friend when there are multiple wakers
of many.  If I run an instance per node with one mother of all work in
autobench mode, it works exactly as you'd expect, game over is when
wakees = socket size. It never get's near that point if I let things
wander, it beats itself up well before we get there.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1396286

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-07 11:10 +0200
Message-ID<rw6pj-7ak-3@gated-at.bofh.it>
In reply to#1391778
On Sun, May 01, 2016 at 11:20:25AM +0200, Mike Galbraith wrote:
> On Sun, 2016-05-01 at 10:53 +0200, Peter Zijlstra wrote:
> > On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> > > On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote:
> > 
> > > > Can you guys have a play with this; I think one and two node tbench are
> > > > good, but I seem to be getting significant run to run variance on that,
> > > > so maybe I'm not doing it right.
> > > 
> > > Nah, tbench is just variance prone.  It got dinged up at clients=cores
> > > on my desktop box, on 4 sockets the high end got seriously dinged up.
> > 
> > Ouch, yeah, big hurt. Lets try that again... :-)
> 
> Yeah, box could use a little bandaid and a hug :)
> 
> Playing with Chris' benchmark, seems the biggest problem is that we
> don't buddy up waker of many and it's wakees in a node.. ie the wake
> wide thing isn't necessarily our friend when there are multiple wakers
> of many.  If I run an instance per node with one mother of all work in
> autobench mode, it works exactly as you'd expect, game over is when
> wakees = socket size. It never get's near that point if I let things
> wander, it beats itself up well before we get there.

Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
but the numbers are so wild to easily break the fragile condition, like:

if (master * 100 < slave * factor * 110)
        return 0;

And since you accumulate wakee number (and decay at HZ), this check tends to
not satisfy ever?

if (slave < factor)
	return 0;

[toc] | [prev] | [next] | [standalone]


#1396389

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-08 10:10 +0200
Message-ID<rwrWO-37i-3@gated-at.bofh.it>
In reply to#1396286
On Sat, 2016-05-07 at 09:24 +0800, Yuyang Du wrote:
> On Sun, May 01, 2016 at 11:20:25AM +0200, Mike Galbraith wrote:

> > Playing with Chris' benchmark, seems the biggest problem is that we
> > don't buddy up waker of many and it's wakees in a node.. ie the wake
> > wide thing isn't necessarily our friend when there are multiple wakers
> > of many.  If I run an instance per node with one mother of all work in
> > autobench mode, it works exactly as you'd expect, game over is when
> > wakees = socket size. It never get's near that point if I let things
> > wander, it beats itself up well before we get there.
> 
> Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
> but the numbers are so wild to easily break the fragile condition, like:

Seems lockless traversal and averages just lets multiple CPUs select
the same spot.  An atomic reservation (feature) when looking for an
idle spot (also for fork) might fix it up.  Run the thing as RT,
push/pull ensures that it reaches box saturation regardless of the
number of messaging threads, whereas with fair class, any number > 1
will certainly stack tasks before the box is saturated.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1396556

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-09 04:40 +0200
Message-ID<rwJgZ-2Lv-5@gated-at.bofh.it>
In reply to#1396389
On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote:
> > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
> > but the numbers are so wild to easily break the fragile condition, like:
> 
> Seems lockless traversal and averages just lets multiple CPUs select
> the same spot.  An atomic reservation (feature) when looking for an
> idle spot (also for fork) might fix it up.  Run the thing as RT,
> push/pull ensures that it reaches box saturation regardless of the
> number of messaging threads, whereas with fair class, any number > 1
> will certainly stack tasks before the box is saturated.

Yes, good idea, bringing order to the race to grab idle CPU is absolutely
helpful.

In addition, I would argue maybe beefing up idle balancing is a more
productive way to spread load, as work-stealing just does what needs
to be done. And seems it has been (sub-unconsciously) neglected in this
case, :)

Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so,
the slave will be 0 forever (as last_wakee is never flipped).

Basically whenever a waker has more than 1 wakee, the wakee_flips
will comfortably grow very large (with last_wakee alternating),
whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0.

So recording only the last_wakee seems not right unless you have other
good reason. If not the latter, counting waking wakee times should be
better, and then allow the statistics to happily play.

[toc] | [prev] | [next] | [standalone]


#1396569

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-09 05:50 +0200
Message-ID<rwKmK-3VJ-7@gated-at.bofh.it>
In reply to#1396556
On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote:
> On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote:
> > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
> > > but the numbers are so wild to easily break the fragile condition, like:
> > 
> > Seems lockless traversal and averages just lets multiple CPUs select
> > the same spot.  An atomic reservation (feature) when looking for an
> > idle spot (also for fork) might fix it up.  Run the thing as RT,
> > push/pull ensures that it reaches box saturation regardless of the
> > number of messaging threads, whereas with fair class, any number > 1
> > will certainly stack tasks before the box is saturated.
> 
> Yes, good idea, bringing order to the race to grab idle CPU is absolutely
> helpful.

Well, good ideas work, as yet this one helps jack diddly spit.

> In addition, I would argue maybe beefing up idle balancing is a more
> productive way to spread load, as work-stealing just does what needs
> to be done. And seems it has been (sub-unconsciously) neglected in this
> case, :)
> 
> Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so,
> the slave will be 0 forever (as last_wakee is never flipped).

Yeah, it's irrelevant here, this load is all about instantaneous state.
 I could use a bit more of that, reserving on the wakeup side won't
help this benchmark until everything else cares.  One stack, and it's
game over.  It could help generic utilization and latency some.. but it
seems kinda unlikely it'll be worth the cycle expenditure.

> Basically whenever a waker has more than 1 wakee, the wakee_flips
> will comfortably grow very large (with last_wakee alternating),
> whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0.

Yup, it is a heuristic, and like all of those, imperfect.  I've watched
it improving utilization in the wild though, so won't mind that until I
catch it doing really bad things.

> So recording only the last_wakee seems not right unless you have other
> good reason. If not the latter, counting waking wakee times should be
> better, and then allow the statistics to happily play.

[toc] | [prev] | [next] | [standalone]


#1396572

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-09 06:10 +0200
Message-ID<rwKG6-4Ch-1@gated-at.bofh.it>
In reply to#1396569
On Mon, May 09, 2016 at 05:45:40AM +0200, Mike Galbraith wrote:
> On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote:
> > On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote:
> > > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
> > > > but the numbers are so wild to easily break the fragile condition, like:
> > > 
> > > Seems lockless traversal and averages just lets multiple CPUs select
> > > the same spot.  An atomic reservation (feature) when looking for an
> > > idle spot (also for fork) might fix it up.  Run the thing as RT,
> > > push/pull ensures that it reaches box saturation regardless of the
> > > number of messaging threads, whereas with fair class, any number > 1
> > > will certainly stack tasks before the box is saturated.
> > 
> > Yes, good idea, bringing order to the race to grab idle CPU is absolutely
> > helpful.
> 
> Well, good ideas work, as yet this one helps jack diddly spit.

Then a valid question is whether it is this selection screwed up in case
like this, as it should necessarily always be asked.
 
> > In addition, I would argue maybe beefing up idle balancing is a more
> > productive way to spread load, as work-stealing just does what needs
> > to be done. And seems it has been (sub-unconsciously) neglected in this
> > case, :)
> > 
> > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so,
> > the slave will be 0 forever (as last_wakee is never flipped).
> 
> Yeah, it's irrelevant here, this load is all about instantaneous state.
>  I could use a bit more of that, reserving on the wakeup side won't
> help this benchmark until everything else cares.  One stack, and it's
> game over.  It could help generic utilization and latency some.. but it
> seems kinda unlikely it'll be worth the cycle expenditure.

Yes and no, it depends on how efficient work-stealing is, compared to
selection, but remember, at the end of the day, the wakee CPU measures the
latency, that CPU does not care it is selected or it steals.
 
> > Basically whenever a waker has more than 1 wakee, the wakee_flips
> > will comfortably grow very large (with last_wakee alternating),
> > whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0.
> 
> Yup, it is a heuristic, and like all of those, imperfect.  I've watched
> it improving utilization in the wild though, so won't mind that until I
> catch it doing really bad things.
 
> > So recording only the last_wakee seems not right unless you have other
> > good reason. If not the latter, counting waking wakee times should be
> > better, and then allow the statistics to happily play.

En... should we try remove recording last_wakee?

[toc] | [prev] | [next] | [standalone]


#1396787

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-09 09:50 +0200
Message-ID<rwO71-yr-43@gated-at.bofh.it>
In reply to#1396572
On Mon, 2016-05-09 at 04:22 +0800, Yuyang Du wrote:
> On Mon, May 09, 2016 at 05:45:40AM +0200, Mike Galbraith wrote:
> > On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote:
> > > On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote:
> > > > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
> > > > > but the numbers are so wild to easily break the fragile condition, like:
> > > > 
> > > > Seems lockless traversal and averages just lets multiple CPUs select
> > > > the same spot.  An atomic reservation (feature) when looking for an
> > > > idle spot (also for fork) might fix it up.  Run the thing as RT,
> > > > push/pull ensures that it reaches box saturation regardless of the
> > > > number of messaging threads, whereas with fair class, any number > 1
> > > > will certainly stack tasks before the box is saturated.
> > > 
> > > Yes, good idea, bringing order to the race to grab idle CPU is absolutely
> > > helpful.
> > 
> > Well, good ideas work, as yet this one helps jack diddly spit.
> 
> Then a valid question is whether it is this selection screwed up in case
> like this, as it should necessarily always be asked.

That's a given, it's just a question of how to do a bit better cheaply.
 
> > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so,
> > > the slave will be 0 forever (as last_wakee is never flipped).
> > 
> > Yeah, it's irrelevant here, this load is all about instantaneous state.
> >  I could use a bit more of that, reserving on the wakeup side won't
> > help this benchmark until everything else cares.  One stack, and it's
> > game over.  It could help generic utilization and latency some.. but it
> > seems kinda unlikely it'll be worth the cycle expenditure.
> 
> Yes and no, it depends on how efficient work-stealing is, compared to
> selection, but remember, at the end of the day, the wakee CPU measures the
> latency, that CPU does not care it is selected or it steals.

In a perfect world, running only Chris' benchmark on an otherwise idle
box, there would never _be_ any work to steal.  In the real world, we
smooth utilization, optimistically peek at this/that, and intentionally
throttle idle balancing (etc etc), which adds up to an imperfect world
for this (based on real world load) benchmark.

> En... should we try remove recording last_wakee?

The more the merrier, go for it! :)

	-Mike

[toc] | [prev] | [next] | [standalone]


#1396861

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-09 11:00 +0200
Message-ID<rwPcK-1Kc-19@gated-at.bofh.it>
In reply to#1396787
On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote:
> > Then a valid question is whether it is this selection screwed up in case
> > like this, as it should necessarily always be asked.
> 
> That's a given, it's just a question of how to do a bit better cheaply.
>  
> > > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so,
> > > > the slave will be 0 forever (as last_wakee is never flipped).
> > > 
> > > Yeah, it's irrelevant here, this load is all about instantaneous state.
> > >  I could use a bit more of that, reserving on the wakeup side won't
> > > help this benchmark until everything else cares.  One stack, and it's
> > > game over.  It could help generic utilization and latency some.. but it
> > > seems kinda unlikely it'll be worth the cycle expenditure.
> > 
> > Yes and no, it depends on how efficient work-stealing is, compared to
> > selection, but remember, at the end of the day, the wakee CPU measures the
> > latency, that CPU does not care it is selected or it steals.
> 
> In a perfect world, running only Chris' benchmark on an otherwise idle
> box, there would never _be_ any work to steal. 

What is the perfect world like? I don't get what you mean.

> In the real world, we
> smooth utilization, optimistically peek at this/that, and intentionally
> throttle idle balancing (etc etc), which adds up to an imperfect world
> for this (based on real world load) benchmark.
 
So, is this a shout-out: these parts should be coordinated better?

> > En... should we try remove recording last_wakee?
> 
> The more the merrier, go for it! :)
 
Nuh, really, this heuristic is too heuristic, :) 
The totality of all possible cases is scary.

Just for a general M:N two-way waker-wakee relationship, not recording
last_wakee may work well generally. E.g., currently, on a 2-socket (24-thread
per socket) 1:24 and 1:48 can't really be differentiated, whereas 1:24 and
2:48 are completely different.

Am I understanding correctly?

[toc] | [prev] | [next] | [standalone]


#1396885

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-09 11:40 +0200
Message-ID<rwPPs-2FZ-27@gated-at.bofh.it>
In reply to#1396861
On Mon, 2016-05-09 at 09:13 +0800, Yuyang Du wrote:
> On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote:

> > In a perfect world, running only Chris' benchmark on an otherwise idle
> > box, there would never _be_ any work to steal. 
> 
> What is the perfect world like? I don't get what you mean.

In a perfect world from this benchmark's perspective, when you fork or
wake while box is underutilized, wakee/child lands on an idle CPU.  To
this benchmark, anything else is broken.
 
> > In the real world, we
> > smooth utilization, optimistically peek at this/that, and intentionally
> > throttle idle balancing (etc etc), which adds up to an imperfect world
> > for this (based on real world load) benchmark.
>  
> So, is this a shout-out: these parts should be coordinated better?

Switching to instantaneous load along with the cpu reservation hackery
made Chris's benchmark a happy camper.  Is that the answer?  Nope, just
verification of the where the problem lives.

> > > En... should we try remove recording last_wakee?
> > 
> > The more the merrier, go for it! :)
>  
> Nuh, really, this heuristic is too heuristic, :) 
> The totality of all possible cases is scary.

Well, make it better.  The author provided evidence when it was born.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1397747

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-10 09:10 +0200
Message-ID<rx9XQ-5Zf-21@gated-at.bofh.it>
In reply to#1396885
On Mon, May 09, 2016 at 11:39:05AM +0200, Mike Galbraith wrote:
> On Mon, 2016-05-09 at 09:13 +0800, Yuyang Du wrote:
> > On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote:
> 
> > > In a perfect world, running only Chris' benchmark on an otherwise idle
> > > box, there would never _be_ any work to steal. 
> > 
> > What is the perfect world like? I don't get what you mean.
> 
> In a perfect world from this benchmark's perspective, when you fork or
> wake while box is underutilized, wakee/child lands on an idle CPU.  To
> this benchmark, anything else is broken.
>  
> > > In the real world, we
> > > smooth utilization, optimistically peek at this/that, and intentionally
> > > throttle idle balancing (etc etc), which adds up to an imperfect world
> > > for this (based on real world load) benchmark.
> >  
> > So, is this a shout-out: these parts should be coordinated better?
> 
> Switching to instantaneous load along with the cpu reservation hackery
> made Chris's benchmark a happy camper.  Is that the answer?  Nope, just
> verification of the where the problem lives.

By cpu reservation, you mean the various averages in select_task_rq_fair?
It does seem a lot of cleanup should be done.
 
> > > > En... should we try remove recording last_wakee?
> > > 
> > > The more the merrier, go for it! :)
> >  
> > Nuh, really, this heuristic is too heuristic, :) 
> > The totality of all possible cases is scary.
> 
> Well, make it better.  The author provided evidence when it was born.

I have to think this through, hot-potato. Maybe even droping it does not
sound outrageous.

[toc] | [prev] | [next] | [standalone]


#1397795

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-10 09:50 +0200
Message-ID<rxaAz-6l0-51@gated-at.bofh.it>
In reply to#1397747
On Tue, 2016-05-10 at 07:26 +0800, Yuyang Du wrote:

> By cpu reservation, you mean the various averages in select_task_rq_fair?
> It does seem a lot of cleanup should be done.

Nah, I meant claiming an idle cpu with cmpxchg().  It's mostly the
average load business that leads to premature stacking though, the
reservation thingy more or less just wastes cycles.  Only whacking
cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t
<near socket size> -a work well.  'Course a rock in its gearbox also
rendered load balancing fairly busted for the general case :)

	-Mike

[toc] | [prev] | [next] | [standalone]


#1398226

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-05-10 17:30 +0200
Message-ID<rxhLI-517-13@gated-at.bofh.it>
In reply to#1397795
On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:

>  Only whacking
> cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t
> <near socket size> -a work well.  'Course a rock in its gearbox also
> rendered load balancing fairly busted for the general case :)

Smaller rock doesn't injure heavy tbench, but more importantly, still
demonstrates the issue when you want full spread.

schbench -m4 -t38 -a

cputime 30000 threads 38 p99 177
cputime 30000 threads 39 p99 10160

LB_TIP_AVG_HIGH
cputime 30000 threads 38 p99 193
cputime 30000 threads 39 p99 184
cputime 30000 threads 40 p99 203
cputime 30000 threads 41 p99 202
cputime 30000 threads 42 p99 205
cputime 30000 threads 43 p99 218
cputime 30000 threads 44 p99 237
cputime 30000 threads 45 p99 245
cputime 30000 threads 46 p99 262
cputime 30000 threads 47 p99 296
cputime 30000 threads 48 p99 3308

47*4+4=nr_cpus yay

---
 kernel/sched/fair.c     |    3 +++
 kernel/sched/features.h |    1 +
 2 files changed, 4 insertions(+)

--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
 
 static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
 {
+	if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
+		return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
+							 cfs_rq->load.weight/2);
 	return cfs_rq->runnable_load_avg;
 }
 
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -67,6 +67,7 @@ SCHED_FEAT(RT_PUSH_IPI, true)
 SCHED_FEAT(FORCE_SD_OVERLAP, false)
 SCHED_FEAT(RT_RUNTIME_SHARE, true)
 SCHED_FEAT(LB_MIN, false)
+SCHED_FEAT(LB_TIP_AVG_HIGH, false)
 SCHED_FEAT(ATTACH_AGE_LOAD, true)
 
 SCHED_FEAT(OLD_IDLE, false)

[toc] | [prev] | [next] | [standalone]


#1398650

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-11 05:00 +0200
Message-ID<rxsxs-7Dd-9@gated-at.bofh.it>
In reply to#1398226
On Tue, May 10, 2016 at 05:26:05PM +0200, Mike Galbraith wrote:
> On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:
> 
> >  Only whacking
> > cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t
> > <near socket size> -a work well.  'Course a rock in its gearbox also
> > rendered load balancing fairly busted for the general case :)
> 
> Smaller rock doesn't injure heavy tbench, but more importantly, still
> demonstrates the issue when you want full spread.
> 
> schbench -m4 -t38 -a
> 
> cputime 30000 threads 38 p99 177
> cputime 30000 threads 39 p99 10160
> 
> LB_TIP_AVG_HIGH
> cputime 30000 threads 38 p99 193
> cputime 30000 threads 39 p99 184
> cputime 30000 threads 40 p99 203
> cputime 30000 threads 41 p99 202
> cputime 30000 threads 42 p99 205
> cputime 30000 threads 43 p99 218
> cputime 30000 threads 44 p99 237
> cputime 30000 threads 45 p99 245
> cputime 30000 threads 46 p99 262
> cputime 30000 threads 47 p99 296
> cputime 30000 threads 48 p99 3308
> 
> 47*4+4=nr_cpus yay
 
yay... and haha, "a perfect world"...

> ---
>  kernel/sched/fair.c     |    3 +++
>  kernel/sched/features.h |    1 +
>  2 files changed, 4 insertions(+)
> 
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
>  
>  static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
>  {
> +	if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> +		return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> +							 cfs_rq->load.weight/2);
>  	return cfs_rq->runnable_load_avg;
>  }
  
cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
in gearbox that only impedes speed up, but also speed down.

But I really don't know the load references in select_task_rq() should be
what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
and just wanting an idle cpu. ?

[toc] | [prev] | [next] | [standalone]


#1398680

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-05-11 06:20 +0200
Message-ID<rxtMS-Fp-3@gated-at.bofh.it>
In reply to#1398650
On Wed, 2016-05-11 at 03:16 +0800, Yuyang Du wrote:
> On Tue, May 10, 2016 at 05:26:05PM +0200, Mike Galbraith wrote:
> > On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:
> > 
> > >  Only whacking
> > > cfs_rq_runnable_load_avg() with a rock makes schbench -m  -t
> > >  -a work well.  'Course a rock in its gearbox also
> > > rendered load balancing fairly busted for the general case :)
> > 
> > Smaller rock doesn't injure heavy tbench, but more importantly, still
> > demonstrates the issue when you want full spread.
> > 
> > schbench -m4 -t38 -a
> > 
> > cputime 30000 threads 38 p99 177
> > cputime 30000 threads 39 p99 10160
> > 
> > LB_TIP_AVG_HIGH
> > cputime 30000 threads 38 p99 193
> > cputime 30000 threads 39 p99 184
> > cputime 30000 threads 40 p99 203
> > cputime 30000 threads 41 p99 202
> > cputime 30000 threads 42 p99 205
> > cputime 30000 threads 43 p99 218
> > cputime 30000 threads 44 p99 237
> > cputime 30000 threads 45 p99 245
> > cputime 30000 threads 46 p99 262
> > cputime 30000 threads 47 p99 296
> > cputime 30000 threads 48 p99 3308
> > 
> > 47*4+4=nr_cpus yay
>  
> yay... and haha, "a perfect world"...

Yup.. for this load.

> > ---
> >  kernel/sched/fair.c     |    3 +++
> >  kernel/sched/features.h |    1 +
> >  2 files changed, 4 insertions(+)
> > 
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
> >  
> >  static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
> >  {
> > +> > 	> > if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> > +> > 	> > 	> > return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> > +> > 	> > 	> > 	> > 	> > 	> > 	> > 	> >  cfs_rq->load.weight/2);
> >  > > 	> > return cfs_rq->runnable_load_avg;
> >  }
>   
> cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
> than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
> in gearbox that only impedes speed up, but also speed down.

Yeah, just like everything else, it'll cuts both ways (why you can't
win the sched game).  If I can believe tbench, at tasks=cpus, reducing
lag increased utilization and reduced latency a wee bit, as did the
reserve thing once a booboo got fixed up.  Makes sense, robbing Peter
to pay Paul should work out better for Paul.

NO_LB_TIP_AVG_HIGH
Throughput 27132.9 MB/sec  96 clients  96 procs  max_latency=7.656 ms
Throughput 28464.1 MB/sec  96 clients  96 procs  max_latency=9.905 ms
Throughput 25369.8 MB/sec  96 clients  96 procs  max_latency=7.192 ms
Throughput 25670.3 MB/sec  96 clients  96 procs  max_latency=5.874 ms
Throughput 29309.3 MB/sec  96 clients  96 procs  max_latency=1.331 ms
avg        27189   1.000                                     6.391   1.000

NO_LB_TIP_AVG_HIGH IDLE_RESERVE
Throughput 24437.5 MB/sec  96 clients  96 procs  max_latency=1.837 ms
Throughput 29464.7 MB/sec  96 clients  96 procs  max_latency=1.594 ms
Throughput 28023.6 MB/sec  96 clients  96 procs  max_latency=1.494 ms
Throughput 28299.0 MB/sec  96 clients  96 procs  max_latency=10.404 ms
Throughput 29072.1 MB/sec  96 clients  96 procs  max_latency=5.575 ms
avg        27859   1.024                                     4.180   0.654

LB_TIP_AVG_HIGH NO_IDLE_RESERVE
Throughput 29068.1 MB/sec  96 clients  96 procs  max_latency=5.599 ms
Throughput 26435.6 MB/sec  96 clients  96 procs  max_latency=3.703 ms
Throughput 23930.0 MB/sec  96 clients  96 procs  max_latency=7.742 ms
Throughput 29464.2 MB/sec  96 clients  96 procs  max_latency=1.549 ms
Throughput 24250.9 MB/sec  96 clients  96 procs  max_latency=1.518 ms
avg        26629   0.979                                     4.022   0.629

LB_TIP_AVG_HIGH IDLE_RESERVE
Throughput 30340.1 MB/sec  96 clients  96 procs  max_latency=1.465 ms
Throughput 29042.9 MB/sec  96 clients  96 procs  max_latency=4.515 ms
Throughput 26718.7 MB/sec  96 clients  96 procs  max_latency=1.822 ms
Throughput 28694.4 MB/sec  96 clients  96 procs  max_latency=1.503 ms
Throughput 28918.2 MB/sec  96 clients  96 procs  max_latency=7.599 ms
avg        28742   1.057                                     3.380   0.528

> But I really don't know the load references in select_task_rq() should be
> what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
> and just wanting an idle cpu. ?

Depends on the goal.  For both, load lagging reality means the high
frequency component is squelched, meaning less migration cost, but also
higher latency due to stacking.  It's a tradeoff where Chris' latency
is everything" benchmark, and _maybe_ the real world load it's based
upon is on Peter's end of the rob Peter to pay Paul transaction.  The
benchmark says it definitely is, the real world load may have already
been fixed up by the select_idle_sibling() rewrite.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1398811

FromYuyang Du <yuyang.du@intel.com>
Date2016-05-11 11:10 +0200
Message-ID<rxyjx-58a-23@gated-at.bofh.it>
In reply to#1398680
On Wed, May 11, 2016 at 06:17:51AM +0200, Mike Galbraith wrote:
> > >  static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
> > >  {
> > > +> > 	> > if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> > > +> > 	> > 	> > return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> > > +> > 	> > 	> > 	> > 	> > 	> > 	> > 	> >  cfs_rq->load.weight/2);
> > >  > > 	> > return cfs_rq->runnable_load_avg;
> > >  }
> >   
> > cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
> > than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
> > in gearbox that only impedes speed up, but also speed down.
> 
> Yeah, just like everything else, it'll cuts both ways (why you can't
> win the sched game).  If I can believe tbench, at tasks=cpus, reducing
> lag increased utilization and reduced latency a wee bit, as did the
> reserve thing once a booboo got fixed up.

Ok, so you have a secret IDLE_RESERVE? Good luck and show it, ;)

> Makes sense, robbing Peter
> to pay Paul should work out better for Paul.
> 
> NO_LB_TIP_AVG_HIGH
> Throughput 27132.9 MB/sec  96 clients  96 procs  max_latency=7.656 ms
> Throughput 28464.1 MB/sec  96 clients  96 procs  max_latency=9.905 ms
> Throughput 25369.8 MB/sec  96 clients  96 procs  max_latency=7.192 ms
> Throughput 25670.3 MB/sec  96 clients  96 procs  max_latency=5.874 ms
> Throughput 29309.3 MB/sec  96 clients  96 procs  max_latency=1.331 ms
> avg        27189   1.000                                     6.391   1.000
> 
> NO_LB_TIP_AVG_HIGH IDLE_RESERVE
> Throughput 24437.5 MB/sec  96 clients  96 procs  max_latency=1.837 ms
> Throughput 29464.7 MB/sec  96 clients  96 procs  max_latency=1.594 ms
> Throughput 28023.6 MB/sec  96 clients  96 procs  max_latency=1.494 ms
> Throughput 28299.0 MB/sec  96 clients  96 procs  max_latency=10.404 ms
> Throughput 29072.1 MB/sec  96 clients  96 procs  max_latency=5.575 ms
> avg        27859   1.024                                     4.180   0.654
> 
> LB_TIP_AVG_HIGH NO_IDLE_RESERVE
> Throughput 29068.1 MB/sec  96 clients  96 procs  max_latency=5.599 ms
> Throughput 26435.6 MB/sec  96 clients  96 procs  max_latency=3.703 ms
> Throughput 23930.0 MB/sec  96 clients  96 procs  max_latency=7.742 ms
> Throughput 29464.2 MB/sec  96 clients  96 procs  max_latency=1.549 ms
> Throughput 24250.9 MB/sec  96 clients  96 procs  max_latency=1.518 ms
> avg        26629   0.979                                     4.022   0.629
> 
> LB_TIP_AVG_HIGH IDLE_RESERVE
> Throughput 30340.1 MB/sec  96 clients  96 procs  max_latency=1.465 ms
> Throughput 29042.9 MB/sec  96 clients  96 procs  max_latency=4.515 ms
> Throughput 26718.7 MB/sec  96 clients  96 procs  max_latency=1.822 ms
> Throughput 28694.4 MB/sec  96 clients  96 procs  max_latency=1.503 ms
> Throughput 28918.2 MB/sec  96 clients  96 procs  max_latency=7.599 ms
> avg        28742   1.057                                     3.380   0.528
> 
> > But I really don't know the load references in select_task_rq() should be
> > what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
> > and just wanting an idle cpu. ?
> 
> Depends on the goal.  For both, load lagging reality means the high
> frequency component is squelched, meaning less migration cost, but also
> higher latency due to stacking.  It's a tradeoff where Chris' latency
> is everything" benchmark, and _maybe_ the real world load it's based
> upon is on Peter's end of the rob Peter to pay Paul transaction.  The
> benchmark says it definitely is, the real world load may have already
> been fixed up by the select_idle_sibling() rewrite.
 
Obviously, load avgs are good at balancing in a larger scale in a timeframe,
so they should be used in comparing/balancing sd's not cpus. However, this
is not the case currently: avgs are mixed with idle cpu/core selection, so
I think better job can be done before and after select_idle_sibling().

For example, I don't know what the complex wake_affine() is really doing for
what. Am i missing something, you think?

Kudos to select_idle_sibling() rewrite, like Peter said, a second step and
an even third step scans are really helping, in addition to many cleanups
and refactors.

[toc] | [prev] | [next] | [standalone]


#1398886

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2016-05-11 12:00 +0200
Message-ID<rxz5V-5GG-13@gated-at.bofh.it>
In reply to#1398811
On Wed, 2016-05-11 at 09:23 +0800, Yuyang Du wrote:

> > Yeah, just like everything else, it'll cuts both ways (why you can't
> > win the sched game).  If I can believe tbench, at tasks=cpus, reducing
> > lag increased utilization and reduced latency a wee bit, as did the
> > reserve thing once a booboo got fixed up.
> 
> Ok, so you have a secret IDLE_RESERVE? Good luck and show it, ;)

Nothing sexy, just cpmxchg(), with the obvious test/set/clear spots.

cmpxchg(&cpu_rq(cpu)->idle_latch, cpu, nr_cpu_ids)

> Depends on the goal.  For both, load lagging reality means the high
> > frequency component is squelched, meaning less migration cost, but also
> > higher latency due to stacking.  It's a tradeoff where Chris' latency
> > is everything" benchmark, and _maybe_ the real world load it's based
> > upon is on Peter's end of the rob Peter to pay Paul transaction.  The
> > benchmark says it definitely is, the real world load may have already
> > been fixed up by the select_idle_sibling() rewrite.
>  
> Obviously, load avgs are good at balancing in a larger scale in a timeframe,
> so they should be used in comparing/balancing sd's not cpus. However, this
> is not the case currently: avgs are mixed with idle cpu/core selection, so
> I think better job can be done before and after select_idle_sibling().
> 
> For example, I don't know what the complex wake_affine() is really doing for
> what. Am i missing something, you think?

wake_affine() just says no to keep us from pulling the whole load to
one cache, starting massive tug-o-war with LB and nuking throughput. 
 Everybody wants hot data, but they can't all have it and scale.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1396570

FromMike Galbraith <mgalbraith@suse.de>
Date2016-05-09 06:00 +0200
Message-ID<rwKwp-44c-3@gated-at.bofh.it>
In reply to#1396556
On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote:

> In addition, I would argue maybe beefing up idle balancing is a more
> productive way to spread load, as work-stealing just does what needs
> to be done. And seems it has been (sub-unconsciously) neglected in this
> case, :)

P.S. Nope, I'm dinging up multiple spots ;-)

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web