Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1391619 > unrolled thread
| Started by | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| First post | 2016-04-30 14:50 +0200 |
| Last post | 2016-05-02 17:20 +0200 |
| Articles | 20 on this page of 46 — 7 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-04-30 14:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 09:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-01 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-01 11:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-07 11:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-08 10:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 04:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 05:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 09:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-10 09:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-10 09:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-10 17:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 05:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 06:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-11 11:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <umgwanakikbuti@gmail.com> - 2016-05-11 12:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 06:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Yuyang Du <yuyang.du@intel.com> - 2016-05-09 06:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 10:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 17:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-02 17:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 16:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-03 17:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 12:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-06 00:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-06 21:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-09 10:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-09 11:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-04 17:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-04 19:50 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-05 11:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-05 16:00 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:20 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Chris Mason <clm@fb.com> - 2016-05-06 19:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-06 09:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Mike Galbraith <mgalbraith@suse.de> - 2016-05-02 19:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Ingo Molnar <mingo@kernel.org> - 2016-05-02 18:10 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 13:40 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-03 20:30 +0200
Re: sched: tweak select_idle_sibling to look for idle threads Peter Zijlstra <peterz@infradead.org> - 2016-05-02 17:20 +0200
Page 1 of 3 [1] 2 3 Next page →
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-30 14:50 +0200 |
| Subject | Re: sched: tweak select_idle_sibling to look for idle threads |
| Message-ID | <rtCvn-7R6-1@gated-at.bofh.it> |
On Sat, Apr 09, 2016 at 03:05:54PM -0400, Chris Mason wrote:
> select_task_rq_fair() can leave cpu utilization a little lumpy,
> especially as the workload ramps up to the maximum capacity of the
> machine. The end result can be high p99 response times as apps
> wait to get scheduled, even when boxes are mostly idle.
>
> I wrote schbench to try and measure this:
>
> git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git
Can you guys have a play with this; I think one and two node tbench are
good, but I seem to be getting significant run to run variance on that,
so maybe I'm not doing it right.
schbench numbers with: ./schbench -m2 -t 20 -c 30000 -s 30000 -r 30
on my ivb-ep (2 sockets, 10 cores/socket, 2 threads/core) appear to be
decent.
I've also not ran anything other than schbench/tbench so maybe I
completely wrecked something else (as per usual..).
I've not thought about that bounce_to_target() thing much.. I'll go give
that a ponder.
---
kernel/sched/fair.c | 180 +++++++++++++++++++++++++++++++++++------------
kernel/sched/features.h | 1 +
kernel/sched/idle_task.c | 4 +-
kernel/sched/sched.h | 1 +
kernel/time/tick-sched.c | 10 +--
5 files changed, 146 insertions(+), 50 deletions(-)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index b8a33abce650..b9d8d1dc5183 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -1501,8 +1501,10 @@ static void task_numa_compare(struct task_numa_env *env,
* One idle CPU per node is evaluated for a task numa move.
* Call select_idle_sibling to maybe find a better one.
*/
- if (!cur)
+ if (!cur) {
+ // XXX borken
env->dst_cpu = select_idle_sibling(env->p, env->dst_cpu);
+ }
assign:
assigned = true;
@@ -4491,6 +4493,17 @@ static void dequeue_task_fair(struct rq *rq, struct task_struct *p, int flags)
}
#ifdef CONFIG_SMP
+
+/*
+ * Working cpumask for:
+ * load_balance,
+ * load_balance_newidle,
+ * select_idle_core.
+ *
+ * Assumes softirqs are disabled when in use.
+ */
+DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
+
#ifdef CONFIG_NO_HZ_COMMON
/*
* per rq 'load' arrray crap; XXX kill this.
@@ -5162,65 +5175,147 @@ find_idlest_cpu(struct sched_group *group, struct task_struct *p, int this_cpu)
return shallowest_idle_cpu != -1 ? shallowest_idle_cpu : least_loaded_cpu;
}
+#ifdef CONFIG_SCHED_SMT
+
+static inline void clear_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 0);
+}
+
+static inline void set_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return;
+
+ WRITE_ONCE(sd->groups->sgc->has_idle_cores, 1);
+}
+
+static inline bool test_idle_cores(int cpu)
+{
+ struct sched_domain *sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ if (!sd)
+ return false;
+
+ // XXX static key for !SMT topologies
+
+ return READ_ONCE(sd->groups->sgc->has_idle_cores);
+}
+
+void update_idle_core(struct rq *rq)
+{
+ int core = cpu_of(rq);
+ int cpu;
+
+ rcu_read_lock();
+ if (test_idle_cores(core))
+ goto unlock;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ if (cpu == core)
+ continue;
+
+ if (!idle_cpu(cpu))
+ goto unlock;
+ }
+
+ set_idle_cores(core);
+unlock:
+ rcu_read_unlock();
+}
+
+static int select_idle_core(struct task_struct *p, int target)
+{
+ struct cpumask *cpus = this_cpu_cpumask_var_ptr(load_balance_mask);
+ struct sched_domain *sd;
+ int core, cpu;
+
+ sd = rcu_dereference(per_cpu(sd_llc, target));
+ cpumask_and(cpus, sched_domain_span(sd), tsk_cpus_allowed(p));
+ for_each_cpu(core, cpus) {
+ bool idle = true;
+
+ for_each_cpu(cpu, cpu_smt_mask(core)) {
+ cpumask_clear_cpu(cpu, cpus);
+ if (!idle_cpu(cpu))
+ idle = false;
+ }
+
+ if (idle)
+ break;
+ }
+
+ return core;
+}
+
+#else /* CONFIG_SCHED_SMT */
+
+static inline void clear_idle_cores(int cpu) { }
+static inline void set_idle_cores(int cpu) { }
+
+static inline bool test_idle_cores(int cpu)
+{
+ return false;
+}
+
+void update_idle_core(struct rq *rq) { }
+
+static inline int select_idle_core(struct task_struct *p, int target)
+{
+ return -1;
+}
+
+#endif /* CONFIG_SCHED_SMT */
+
/*
- * Try and locate an idle CPU in the sched_domain.
+ * Try and locate an idle core/thread in the LLC cache domain.
*/
static int select_idle_sibling(struct task_struct *p, int target)
{
struct sched_domain *sd;
- struct sched_group *sg;
int i = task_cpu(p);
if (idle_cpu(target))
return target;
/*
- * If the prevous cpu is cache affine and idle, don't be stupid.
+ * If the previous cpu is cache affine and idle, don't be stupid.
*/
if (i != target && cpus_share_cache(i, target) && idle_cpu(i))
return i;
+ sd = rcu_dereference(per_cpu(sd_llc, target));
+ if (!sd)
+ return target;
+
/*
- * Otherwise, iterate the domains and find an eligible idle cpu.
- *
- * A completely idle sched group at higher domains is more
- * desirable than an idle group at a lower level, because lower
- * domains have smaller groups and usually share hardware
- * resources which causes tasks to contend on them, e.g. x86
- * hyperthread siblings in the lowest domain (SMT) can contend
- * on the shared cpu pipeline.
- *
- * However, while we prefer idle groups at higher domains
- * finding an idle cpu at the lowest domain is still better than
- * returning 'target', which we've already established, isn't
- * idle.
+ * If there are idle cores to be had, go find one.
*/
- sd = rcu_dereference(per_cpu(sd_llc, target));
- for_each_lower_domain(sd) {
- sg = sd->groups;
- do {
- if (!cpumask_intersects(sched_group_cpus(sg),
- tsk_cpus_allowed(p)))
- goto next;
-
- /* Ensure the entire group is idle */
- for_each_cpu(i, sched_group_cpus(sg)) {
- if (i == target || !idle_cpu(i))
- goto next;
- }
+ if (sched_feat(IDLE_CORE) && test_idle_cores(target)) {
+ i = select_idle_core(p, target);
+ if ((unsigned)i < nr_cpumask_bits)
+ return i;
- /*
- * It doesn't matter which cpu we pick, the
- * whole group is idle.
- */
- target = cpumask_first_and(sched_group_cpus(sg),
- tsk_cpus_allowed(p));
- goto done;
-next:
- sg = sg->next;
- } while (sg != sd->groups);
+ /*
+ * Failed to find an idle core; stop looking for one.
+ */
+ clear_idle_cores(target);
}
-done:
+
+ /*
+ * Otherwise, settle for anything idle in this cache domain.
+ */
+ for_each_cpu(i, sched_domain_span(sd)) {
+ if (!cpumask_test_cpu(i, tsk_cpus_allowed(p)))
+ continue;
+ if (idle_cpu(i))
+ return i;
+ }
+
return target;
}
@@ -7229,9 +7324,6 @@ static struct rq *find_busiest_queue(struct lb_env *env,
*/
#define MAX_PINNED_INTERVAL 512
-/* Working cpumask for load_balance and load_balance_newidle. */
-DEFINE_PER_CPU(cpumask_var_t, load_balance_mask);
-
static int need_active_balance(struct lb_env *env)
{
struct sched_domain *sd = env->sd;
diff --git a/kernel/sched/features.h b/kernel/sched/features.h
index 69631fa46c2f..76bb8814649a 100644
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -69,3 +69,4 @@ SCHED_FEAT(RT_RUNTIME_SHARE, true)
SCHED_FEAT(LB_MIN, false)
SCHED_FEAT(ATTACH_AGE_LOAD, true)
+SCHED_FEAT(IDLE_CORE, true)
diff --git a/kernel/sched/idle_task.c b/kernel/sched/idle_task.c
index 47ce94931f1b..cb394db407e4 100644
--- a/kernel/sched/idle_task.c
+++ b/kernel/sched/idle_task.c
@@ -23,11 +23,13 @@ static void check_preempt_curr_idle(struct rq *rq, struct task_struct *p, int fl
resched_curr(rq);
}
+extern void update_idle_core(struct rq *rq);
+
static struct task_struct *
pick_next_task_idle(struct rq *rq, struct task_struct *prev)
{
put_prev_task(rq, prev);
-
+ update_idle_core(rq);
schedstat_inc(rq, sched_goidle);
return rq->idle;
}
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 69da6fcaa0e8..5994794bfc85 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -866,6 +866,7 @@ struct sched_group_capacity {
* Number of busy cpus in this group.
*/
atomic_t nr_busy_cpus;
+ int has_idle_cores;
unsigned long cpumask[0]; /* iteration mask */
};
diff --git a/kernel/time/tick-sched.c b/kernel/time/tick-sched.c
index 31872bc53bc4..6e42cd218ba5 100644
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
WARN_ON_ONCE(irqs_disabled());
/*
- * Update the idle state in the scheduler domain hierarchy
- * when tick_nohz_stop_sched_tick() is called from the idle loop.
- * State will be updated to busy during the first busy tick after
- * exiting idle.
- */
+ * Update the idle state in the scheduler domain hierarchy
+ * when tick_nohz_stop_sched_tick() is called from the idle loop.
+ * State will be updated to busy during the first busy tick after
+ * exiting idle.
+ */
set_cpu_sd_state_idle();
local_irq_disable();
[toc] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-01 09:20 +0200 |
| Message-ID | <rtTPz-5sM-1@gated-at.bofh.it> |
| In reply to | #1391619 |
On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote: > On Sat, Apr 09, 2016 at 03:05:54PM -0400, Chris Mason wrote: > > select_task_rq_fair() can leave cpu utilization a little lumpy, > > especially as the workload ramps up to the maximum capacity of the > > machine. The end result can be high p99 response times as apps > > wait to get scheduled, even when boxes are mostly idle. > > > > I wrote schbench to try and measure this: > > > > git://git.kernel.org/pub/scm/linux/kernel/git/mason/schbench.git > > Can you guys have a play with this; I think one and two node tbench are > good, but I seem to be getting significant run to run variance on that, > so maybe I'm not doing it right. Nah, tbench is just variance prone. It got dinged up at clients=cores on my desktop box, on 4 sockets the high end got seriously dinged up. tbench 1 x i4790 master avg 1 714 684 688 695 1.000 2 1260 1234 1284 1259 1.000 4 2238 2301 2286 2275 1.000 8 3388 3418 3396 3400 1.000 masterx 1 690 701 701 697 1.002 2 1287 1332 1235 1284 1.019 4 2014 2006 1999 2006 0.881 8 3388 3385 3404 3392 0.997 4 x E7-8890 master avg 1 524 524 523 523 1.000 2 1049 1053 1045 1049 1.000 4 2064 2081 2091 2078 1.000 8 3737 3813 3746 3765 1.000 16 7129 7028 7082 7079 1.000 32 13718 13730 13578 13675 1.000 64 21397 21435 21519 21450 1.000 128 39846 38397 39026 39089 1.000 256 59509 59797 59344 59550 1.000 masterx avg 1 505 507 501 504 0.963 1.000 2 1036 1027 1039 1034 0.985 1.000 4 1977 2001 1992 1990 0.957 1.000 8 3734 3802 3778 3771 1.001 1.000 16 7124 7079 7071 7091 1.001 1.000 32 13549 13758 13364 13557 0.991 1.000 64 21975 22161 22100 22078 1.029 1.000 128 23066 23044 23028 23046 0.589 1.000 256 29905 29630 30753 30096 0.505 1.000 masterx NO_IDLE_CORE avg 1 500 521 502 507 0.969 1.005 2 1012 996 1043 1017 0.969 0.983 4 1988 1992 1995 1991 0.958 1.000 8 3834 3758 3671 3754 0.997 0.995 16 7160 7206 7168 7178 1.013 1.012 32 13788 13773 13672 13744 1.005 1.013 64 21771 21845 21826 21814 1.016 0.988 128 23248 23136 23133 23172 0.592 1.005 256 28683 30013 31850 30182 0.506 1.002
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-01 11:00 +0200 |
| Message-ID | <rtVom-6wL-9@gated-at.bofh.it> |
| In reply to | #1391757 |
On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote: > On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote: > > Can you guys have a play with this; I think one and two node tbench are > > good, but I seem to be getting significant run to run variance on that, > > so maybe I'm not doing it right. > > Nah, tbench is just variance prone. It got dinged up at clients=cores > on my desktop box, on 4 sockets the high end got seriously dinged up. Ouch, yeah, big hurt. Lets try that again... :-)
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-01 11:30 +0200 |
| Message-ID | <rtVRo-7cO-1@gated-at.bofh.it> |
| In reply to | #1391776 |
On Sun, 2016-05-01 at 10:53 +0200, Peter Zijlstra wrote: > On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote: > > On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote: > > > > Can you guys have a play with this; I think one and two node tbench are > > > good, but I seem to be getting significant run to run variance on that, > > > so maybe I'm not doing it right. > > > > Nah, tbench is just variance prone. It got dinged up at clients=cores > > on my desktop box, on 4 sockets the high end got seriously dinged up. > > Ouch, yeah, big hurt. Lets try that again... :-) Yeah, box could use a little bandaid and a hug :) Playing with Chris' benchmark, seems the biggest problem is that we don't buddy up waker of many and it's wakees in a node.. ie the wake wide thing isn't necessarily our friend when there are multiple wakers of many. If I run an instance per node with one mother of all work in autobench mode, it works exactly as you'd expect, game over is when wakees = socket size. It never get's near that point if I let things wander, it beats itself up well before we get there. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-07 11:10 +0200 |
| Message-ID | <rw6pj-7ak-3@gated-at.bofh.it> |
| In reply to | #1391778 |
On Sun, May 01, 2016 at 11:20:25AM +0200, Mike Galbraith wrote:
> On Sun, 2016-05-01 at 10:53 +0200, Peter Zijlstra wrote:
> > On Sun, May 01, 2016 at 09:12:33AM +0200, Mike Galbraith wrote:
> > > On Sat, 2016-04-30 at 14:47 +0200, Peter Zijlstra wrote:
> >
> > > > Can you guys have a play with this; I think one and two node tbench are
> > > > good, but I seem to be getting significant run to run variance on that,
> > > > so maybe I'm not doing it right.
> > >
> > > Nah, tbench is just variance prone. It got dinged up at clients=cores
> > > on my desktop box, on 4 sockets the high end got seriously dinged up.
> >
> > Ouch, yeah, big hurt. Lets try that again... :-)
>
> Yeah, box could use a little bandaid and a hug :)
>
> Playing with Chris' benchmark, seems the biggest problem is that we
> don't buddy up waker of many and it's wakees in a node.. ie the wake
> wide thing isn't necessarily our friend when there are multiple wakers
> of many. If I run an instance per node with one mother of all work in
> autobench mode, it works exactly as you'd expect, game over is when
> wakees = socket size. It never get's near that point if I let things
> wander, it beats itself up well before we get there.
Maybe give the criteria a bit margin, not just wakees tend to equal llc_size,
but the numbers are so wild to easily break the fragile condition, like:
if (master * 100 < slave * factor * 110)
return 0;
And since you accumulate wakee number (and decay at HZ), this check tends to
not satisfy ever?
if (slave < factor)
return 0;
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-08 10:10 +0200 |
| Message-ID | <rwrWO-37i-3@gated-at.bofh.it> |
| In reply to | #1396286 |
On Sat, 2016-05-07 at 09:24 +0800, Yuyang Du wrote: > On Sun, May 01, 2016 at 11:20:25AM +0200, Mike Galbraith wrote: > > Playing with Chris' benchmark, seems the biggest problem is that we > > don't buddy up waker of many and it's wakees in a node.. ie the wake > > wide thing isn't necessarily our friend when there are multiple wakers > > of many. If I run an instance per node with one mother of all work in > > autobench mode, it works exactly as you'd expect, game over is when > > wakees = socket size. It never get's near that point if I let things > > wander, it beats itself up well before we get there. > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size, > but the numbers are so wild to easily break the fragile condition, like: Seems lockless traversal and averages just lets multiple CPUs select the same spot. An atomic reservation (feature) when looking for an idle spot (also for fork) might fix it up. Run the thing as RT, push/pull ensures that it reaches box saturation regardless of the number of messaging threads, whereas with fair class, any number > 1 will certainly stack tasks before the box is saturated. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-09 04:40 +0200 |
| Message-ID | <rwJgZ-2Lv-5@gated-at.bofh.it> |
| In reply to | #1396389 |
On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote: > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size, > > but the numbers are so wild to easily break the fragile condition, like: > > Seems lockless traversal and averages just lets multiple CPUs select > the same spot. An atomic reservation (feature) when looking for an > idle spot (also for fork) might fix it up. Run the thing as RT, > push/pull ensures that it reaches box saturation regardless of the > number of messaging threads, whereas with fair class, any number > 1 > will certainly stack tasks before the box is saturated. Yes, good idea, bringing order to the race to grab idle CPU is absolutely helpful. In addition, I would argue maybe beefing up idle balancing is a more productive way to spread load, as work-stealing just does what needs to be done. And seems it has been (sub-unconsciously) neglected in this case, :) Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so, the slave will be 0 forever (as last_wakee is never flipped). Basically whenever a waker has more than 1 wakee, the wakee_flips will comfortably grow very large (with last_wakee alternating), whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0. So recording only the last_wakee seems not right unless you have other good reason. If not the latter, counting waking wakee times should be better, and then allow the statistics to happily play.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-09 05:50 +0200 |
| Message-ID | <rwKmK-3VJ-7@gated-at.bofh.it> |
| In reply to | #1396556 |
On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote: > On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote: > > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size, > > > but the numbers are so wild to easily break the fragile condition, like: > > > > Seems lockless traversal and averages just lets multiple CPUs select > > the same spot. An atomic reservation (feature) when looking for an > > idle spot (also for fork) might fix it up. Run the thing as RT, > > push/pull ensures that it reaches box saturation regardless of the > > number of messaging threads, whereas with fair class, any number > 1 > > will certainly stack tasks before the box is saturated. > > Yes, good idea, bringing order to the race to grab idle CPU is absolutely > helpful. Well, good ideas work, as yet this one helps jack diddly spit. > In addition, I would argue maybe beefing up idle balancing is a more > productive way to spread load, as work-stealing just does what needs > to be done. And seems it has been (sub-unconsciously) neglected in this > case, :) > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so, > the slave will be 0 forever (as last_wakee is never flipped). Yeah, it's irrelevant here, this load is all about instantaneous state. I could use a bit more of that, reserving on the wakeup side won't help this benchmark until everything else cares. One stack, and it's game over. It could help generic utilization and latency some.. but it seems kinda unlikely it'll be worth the cycle expenditure. > Basically whenever a waker has more than 1 wakee, the wakee_flips > will comfortably grow very large (with last_wakee alternating), > whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0. Yup, it is a heuristic, and like all of those, imperfect. I've watched it improving utilization in the wild though, so won't mind that until I catch it doing really bad things. > So recording only the last_wakee seems not right unless you have other > good reason. If not the latter, counting waking wakee times should be > better, and then allow the statistics to happily play.
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-09 06:10 +0200 |
| Message-ID | <rwKG6-4Ch-1@gated-at.bofh.it> |
| In reply to | #1396569 |
On Mon, May 09, 2016 at 05:45:40AM +0200, Mike Galbraith wrote: > On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote: > > On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote: > > > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size, > > > > but the numbers are so wild to easily break the fragile condition, like: > > > > > > Seems lockless traversal and averages just lets multiple CPUs select > > > the same spot. An atomic reservation (feature) when looking for an > > > idle spot (also for fork) might fix it up. Run the thing as RT, > > > push/pull ensures that it reaches box saturation regardless of the > > > number of messaging threads, whereas with fair class, any number > 1 > > > will certainly stack tasks before the box is saturated. > > > > Yes, good idea, bringing order to the race to grab idle CPU is absolutely > > helpful. > > Well, good ideas work, as yet this one helps jack diddly spit. Then a valid question is whether it is this selection screwed up in case like this, as it should necessarily always be asked. > > In addition, I would argue maybe beefing up idle balancing is a more > > productive way to spread load, as work-stealing just does what needs > > to be done. And seems it has been (sub-unconsciously) neglected in this > > case, :) > > > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so, > > the slave will be 0 forever (as last_wakee is never flipped). > > Yeah, it's irrelevant here, this load is all about instantaneous state. > I could use a bit more of that, reserving on the wakeup side won't > help this benchmark until everything else cares. One stack, and it's > game over. It could help generic utilization and latency some.. but it > seems kinda unlikely it'll be worth the cycle expenditure. Yes and no, it depends on how efficient work-stealing is, compared to selection, but remember, at the end of the day, the wakee CPU measures the latency, that CPU does not care it is selected or it steals. > > Basically whenever a waker has more than 1 wakee, the wakee_flips > > will comfortably grow very large (with last_wakee alternating), > > whereas when a waker has 0 or 1 wakee, the wakee_flips will just be 0. > > Yup, it is a heuristic, and like all of those, imperfect. I've watched > it improving utilization in the wild though, so won't mind that until I > catch it doing really bad things. > > So recording only the last_wakee seems not right unless you have other > > good reason. If not the latter, counting waking wakee times should be > > better, and then allow the statistics to happily play. En... should we try remove recording last_wakee?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-09 09:50 +0200 |
| Message-ID | <rwO71-yr-43@gated-at.bofh.it> |
| In reply to | #1396572 |
On Mon, 2016-05-09 at 04:22 +0800, Yuyang Du wrote: > On Mon, May 09, 2016 at 05:45:40AM +0200, Mike Galbraith wrote: > > On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote: > > > On Sun, May 08, 2016 at 10:08:55AM +0200, Mike Galbraith wrote: > > > > > Maybe give the criteria a bit margin, not just wakees tend to equal llc_size, > > > > > but the numbers are so wild to easily break the fragile condition, like: > > > > > > > > Seems lockless traversal and averages just lets multiple CPUs select > > > > the same spot. An atomic reservation (feature) when looking for an > > > > idle spot (also for fork) might fix it up. Run the thing as RT, > > > > push/pull ensures that it reaches box saturation regardless of the > > > > number of messaging threads, whereas with fair class, any number > 1 > > > > will certainly stack tasks before the box is saturated. > > > > > > Yes, good idea, bringing order to the race to grab idle CPU is absolutely > > > helpful. > > > > Well, good ideas work, as yet this one helps jack diddly spit. > > Then a valid question is whether it is this selection screwed up in case > like this, as it should necessarily always be asked. That's a given, it's just a question of how to do a bit better cheaply. > > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so, > > > the slave will be 0 forever (as last_wakee is never flipped). > > > > Yeah, it's irrelevant here, this load is all about instantaneous state. > > I could use a bit more of that, reserving on the wakeup side won't > > help this benchmark until everything else cares. One stack, and it's > > game over. It could help generic utilization and latency some.. but it > > seems kinda unlikely it'll be worth the cycle expenditure. > > Yes and no, it depends on how efficient work-stealing is, compared to > selection, but remember, at the end of the day, the wakee CPU measures the > latency, that CPU does not care it is selected or it steals. In a perfect world, running only Chris' benchmark on an otherwise idle box, there would never _be_ any work to steal. In the real world, we smooth utilization, optimistically peek at this/that, and intentionally throttle idle balancing (etc etc), which adds up to an imperfect world for this (based on real world load) benchmark. > En... should we try remove recording last_wakee? The more the merrier, go for it! :) -Mike
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-09 11:00 +0200 |
| Message-ID | <rwPcK-1Kc-19@gated-at.bofh.it> |
| In reply to | #1396787 |
On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote: > > Then a valid question is whether it is this selection screwed up in case > > like this, as it should necessarily always be asked. > > That's a given, it's just a question of how to do a bit better cheaply. > > > > > Regarding wake_wide(), it seems the M:N is 1:24, not 6:6*24, if so, > > > > the slave will be 0 forever (as last_wakee is never flipped). > > > > > > Yeah, it's irrelevant here, this load is all about instantaneous state. > > > I could use a bit more of that, reserving on the wakeup side won't > > > help this benchmark until everything else cares. One stack, and it's > > > game over. It could help generic utilization and latency some.. but it > > > seems kinda unlikely it'll be worth the cycle expenditure. > > > > Yes and no, it depends on how efficient work-stealing is, compared to > > selection, but remember, at the end of the day, the wakee CPU measures the > > latency, that CPU does not care it is selected or it steals. > > In a perfect world, running only Chris' benchmark on an otherwise idle > box, there would never _be_ any work to steal. What is the perfect world like? I don't get what you mean. > In the real world, we > smooth utilization, optimistically peek at this/that, and intentionally > throttle idle balancing (etc etc), which adds up to an imperfect world > for this (based on real world load) benchmark. So, is this a shout-out: these parts should be coordinated better? > > En... should we try remove recording last_wakee? > > The more the merrier, go for it! :) Nuh, really, this heuristic is too heuristic, :) The totality of all possible cases is scary. Just for a general M:N two-way waker-wakee relationship, not recording last_wakee may work well generally. E.g., currently, on a 2-socket (24-thread per socket) 1:24 and 1:48 can't really be differentiated, whereas 1:24 and 2:48 are completely different. Am I understanding correctly?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-09 11:40 +0200 |
| Message-ID | <rwPPs-2FZ-27@gated-at.bofh.it> |
| In reply to | #1396861 |
On Mon, 2016-05-09 at 09:13 +0800, Yuyang Du wrote: > On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote: > > In a perfect world, running only Chris' benchmark on an otherwise idle > > box, there would never _be_ any work to steal. > > What is the perfect world like? I don't get what you mean. In a perfect world from this benchmark's perspective, when you fork or wake while box is underutilized, wakee/child lands on an idle CPU. To this benchmark, anything else is broken. > > In the real world, we > > smooth utilization, optimistically peek at this/that, and intentionally > > throttle idle balancing (etc etc), which adds up to an imperfect world > > for this (based on real world load) benchmark. > > So, is this a shout-out: these parts should be coordinated better? Switching to instantaneous load along with the cpu reservation hackery made Chris's benchmark a happy camper. Is that the answer? Nope, just verification of the where the problem lives. > > > En... should we try remove recording last_wakee? > > > > The more the merrier, go for it! :) > > Nuh, really, this heuristic is too heuristic, :) > The totality of all possible cases is scary. Well, make it better. The author provided evidence when it was born. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-10 09:10 +0200 |
| Message-ID | <rx9XQ-5Zf-21@gated-at.bofh.it> |
| In reply to | #1396885 |
On Mon, May 09, 2016 at 11:39:05AM +0200, Mike Galbraith wrote: > On Mon, 2016-05-09 at 09:13 +0800, Yuyang Du wrote: > > On Mon, May 09, 2016 at 09:44:13AM +0200, Mike Galbraith wrote: > > > > In a perfect world, running only Chris' benchmark on an otherwise idle > > > box, there would never _be_ any work to steal. > > > > What is the perfect world like? I don't get what you mean. > > In a perfect world from this benchmark's perspective, when you fork or > wake while box is underutilized, wakee/child lands on an idle CPU. To > this benchmark, anything else is broken. > > > > In the real world, we > > > smooth utilization, optimistically peek at this/that, and intentionally > > > throttle idle balancing (etc etc), which adds up to an imperfect world > > > for this (based on real world load) benchmark. > > > > So, is this a shout-out: these parts should be coordinated better? > > Switching to instantaneous load along with the cpu reservation hackery > made Chris's benchmark a happy camper. Is that the answer? Nope, just > verification of the where the problem lives. By cpu reservation, you mean the various averages in select_task_rq_fair? It does seem a lot of cleanup should be done. > > > > En... should we try remove recording last_wakee? > > > > > > The more the merrier, go for it! :) > > > > Nuh, really, this heuristic is too heuristic, :) > > The totality of all possible cases is scary. > > Well, make it better. The author provided evidence when it was born. I have to think this through, hot-potato. Maybe even droping it does not sound outrageous.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-10 09:50 +0200 |
| Message-ID | <rxaAz-6l0-51@gated-at.bofh.it> |
| In reply to | #1397747 |
On Tue, 2016-05-10 at 07:26 +0800, Yuyang Du wrote: > By cpu reservation, you mean the various averages in select_task_rq_fair? > It does seem a lot of cleanup should be done. Nah, I meant claiming an idle cpu with cmpxchg(). It's mostly the average load business that leads to premature stacking though, the reservation thingy more or less just wastes cycles. Only whacking cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t <near socket size> -a work well. 'Course a rock in its gearbox also rendered load balancing fairly busted for the general case :) -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-05-10 17:30 +0200 |
| Message-ID | <rxhLI-517-13@gated-at.bofh.it> |
| In reply to | #1397795 |
On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:
> Only whacking
> cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t
> <near socket size> -a work well. 'Course a rock in its gearbox also
> rendered load balancing fairly busted for the general case :)
Smaller rock doesn't injure heavy tbench, but more importantly, still
demonstrates the issue when you want full spread.
schbench -m4 -t38 -a
cputime 30000 threads 38 p99 177
cputime 30000 threads 39 p99 10160
LB_TIP_AVG_HIGH
cputime 30000 threads 38 p99 193
cputime 30000 threads 39 p99 184
cputime 30000 threads 40 p99 203
cputime 30000 threads 41 p99 202
cputime 30000 threads 42 p99 205
cputime 30000 threads 43 p99 218
cputime 30000 threads 44 p99 237
cputime 30000 threads 45 p99 245
cputime 30000 threads 46 p99 262
cputime 30000 threads 47 p99 296
cputime 30000 threads 48 p99 3308
47*4+4=nr_cpus yay
---
kernel/sched/fair.c | 3 +++
kernel/sched/features.h | 1 +
2 files changed, 4 insertions(+)
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
{
+ if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
+ return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
+ cfs_rq->load.weight/2);
return cfs_rq->runnable_load_avg;
}
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -67,6 +67,7 @@ SCHED_FEAT(RT_PUSH_IPI, true)
SCHED_FEAT(FORCE_SD_OVERLAP, false)
SCHED_FEAT(RT_RUNTIME_SHARE, true)
SCHED_FEAT(LB_MIN, false)
+SCHED_FEAT(LB_TIP_AVG_HIGH, false)
SCHED_FEAT(ATTACH_AGE_LOAD, true)
SCHED_FEAT(OLD_IDLE, false)
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-11 05:00 +0200 |
| Message-ID | <rxsxs-7Dd-9@gated-at.bofh.it> |
| In reply to | #1398226 |
On Tue, May 10, 2016 at 05:26:05PM +0200, Mike Galbraith wrote:
> On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:
>
> > Only whacking
> > cfs_rq_runnable_load_avg() with a rock makes schbench -m <sockets> -t
> > <near socket size> -a work well. 'Course a rock in its gearbox also
> > rendered load balancing fairly busted for the general case :)
>
> Smaller rock doesn't injure heavy tbench, but more importantly, still
> demonstrates the issue when you want full spread.
>
> schbench -m4 -t38 -a
>
> cputime 30000 threads 38 p99 177
> cputime 30000 threads 39 p99 10160
>
> LB_TIP_AVG_HIGH
> cputime 30000 threads 38 p99 193
> cputime 30000 threads 39 p99 184
> cputime 30000 threads 40 p99 203
> cputime 30000 threads 41 p99 202
> cputime 30000 threads 42 p99 205
> cputime 30000 threads 43 p99 218
> cputime 30000 threads 44 p99 237
> cputime 30000 threads 45 p99 245
> cputime 30000 threads 46 p99 262
> cputime 30000 threads 47 p99 296
> cputime 30000 threads 48 p99 3308
>
> 47*4+4=nr_cpus yay
yay... and haha, "a perfect world"...
> ---
> kernel/sched/fair.c | 3 +++
> kernel/sched/features.h | 1 +
> 2 files changed, 4 insertions(+)
>
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
>
> static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
> {
> + if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> + return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> + cfs_rq->load.weight/2);
> return cfs_rq->runnable_load_avg;
> }
cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
in gearbox that only impedes speed up, but also speed down.
But I really don't know the load references in select_task_rq() should be
what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
and just wanting an idle cpu. ?
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-05-11 06:20 +0200 |
| Message-ID | <rxtMS-Fp-3@gated-at.bofh.it> |
| In reply to | #1398650 |
On Wed, 2016-05-11 at 03:16 +0800, Yuyang Du wrote:
> On Tue, May 10, 2016 at 05:26:05PM +0200, Mike Galbraith wrote:
> > On Tue, 2016-05-10 at 09:49 +0200, Mike Galbraith wrote:
> >
> > > Only whacking
> > > cfs_rq_runnable_load_avg() with a rock makes schbench -m -t
> > > -a work well. 'Course a rock in its gearbox also
> > > rendered load balancing fairly busted for the general case :)
> >
> > Smaller rock doesn't injure heavy tbench, but more importantly, still
> > demonstrates the issue when you want full spread.
> >
> > schbench -m4 -t38 -a
> >
> > cputime 30000 threads 38 p99 177
> > cputime 30000 threads 39 p99 10160
> >
> > LB_TIP_AVG_HIGH
> > cputime 30000 threads 38 p99 193
> > cputime 30000 threads 39 p99 184
> > cputime 30000 threads 40 p99 203
> > cputime 30000 threads 41 p99 202
> > cputime 30000 threads 42 p99 205
> > cputime 30000 threads 43 p99 218
> > cputime 30000 threads 44 p99 237
> > cputime 30000 threads 45 p99 245
> > cputime 30000 threads 46 p99 262
> > cputime 30000 threads 47 p99 296
> > cputime 30000 threads 48 p99 3308
> >
> > 47*4+4=nr_cpus yay
>
> yay... and haha, "a perfect world"...
Yup.. for this load.
> > ---
> > kernel/sched/fair.c | 3 +++
> > kernel/sched/features.h | 1 +
> > 2 files changed, 4 insertions(+)
> >
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -3027,6 +3027,9 @@ void remove_entity_load_avg(struct sched
> >
> > static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
> > {
> > +> > > > if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> > +> > > > > > return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> > +> > > > > > > > > > > > > > > > cfs_rq->load.weight/2);
> > > > > > return cfs_rq->runnable_load_avg;
> > }
>
> cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
> than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
> in gearbox that only impedes speed up, but also speed down.
Yeah, just like everything else, it'll cuts both ways (why you can't
win the sched game). If I can believe tbench, at tasks=cpus, reducing
lag increased utilization and reduced latency a wee bit, as did the
reserve thing once a booboo got fixed up. Makes sense, robbing Peter
to pay Paul should work out better for Paul.
NO_LB_TIP_AVG_HIGH
Throughput 27132.9 MB/sec 96 clients 96 procs max_latency=7.656 ms
Throughput 28464.1 MB/sec 96 clients 96 procs max_latency=9.905 ms
Throughput 25369.8 MB/sec 96 clients 96 procs max_latency=7.192 ms
Throughput 25670.3 MB/sec 96 clients 96 procs max_latency=5.874 ms
Throughput 29309.3 MB/sec 96 clients 96 procs max_latency=1.331 ms
avg 27189 1.000 6.391 1.000
NO_LB_TIP_AVG_HIGH IDLE_RESERVE
Throughput 24437.5 MB/sec 96 clients 96 procs max_latency=1.837 ms
Throughput 29464.7 MB/sec 96 clients 96 procs max_latency=1.594 ms
Throughput 28023.6 MB/sec 96 clients 96 procs max_latency=1.494 ms
Throughput 28299.0 MB/sec 96 clients 96 procs max_latency=10.404 ms
Throughput 29072.1 MB/sec 96 clients 96 procs max_latency=5.575 ms
avg 27859 1.024 4.180 0.654
LB_TIP_AVG_HIGH NO_IDLE_RESERVE
Throughput 29068.1 MB/sec 96 clients 96 procs max_latency=5.599 ms
Throughput 26435.6 MB/sec 96 clients 96 procs max_latency=3.703 ms
Throughput 23930.0 MB/sec 96 clients 96 procs max_latency=7.742 ms
Throughput 29464.2 MB/sec 96 clients 96 procs max_latency=1.549 ms
Throughput 24250.9 MB/sec 96 clients 96 procs max_latency=1.518 ms
avg 26629 0.979 4.022 0.629
LB_TIP_AVG_HIGH IDLE_RESERVE
Throughput 30340.1 MB/sec 96 clients 96 procs max_latency=1.465 ms
Throughput 29042.9 MB/sec 96 clients 96 procs max_latency=4.515 ms
Throughput 26718.7 MB/sec 96 clients 96 procs max_latency=1.822 ms
Throughput 28694.4 MB/sec 96 clients 96 procs max_latency=1.503 ms
Throughput 28918.2 MB/sec 96 clients 96 procs max_latency=7.599 ms
avg 28742 1.057 3.380 0.528
> But I really don't know the load references in select_task_rq() should be
> what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
> and just wanting an idle cpu. ?
Depends on the goal. For both, load lagging reality means the high
frequency component is squelched, meaning less migration cost, but also
higher latency due to stacking. It's a tradeoff where Chris' latency
is everything" benchmark, and _maybe_ the real world load it's based
upon is on Peter's end of the rob Peter to pay Paul transaction. The
benchmark says it definitely is, the real world load may have already
been fixed up by the select_idle_sibling() rewrite.
-Mike
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2016-05-11 11:10 +0200 |
| Message-ID | <rxyjx-58a-23@gated-at.bofh.it> |
| In reply to | #1398680 |
On Wed, May 11, 2016 at 06:17:51AM +0200, Mike Galbraith wrote:
> > > static inline unsigned long cfs_rq_runnable_load_avg(struct cfs_rq *cfs_rq)
> > > {
> > > +> > > > if (sched_feat(LB_TIP_AVG_HIGH) && cfs_rq->load.weight > cfs_rq->runnable_load_avg*2)
> > > +> > > > > > return cfs_rq->runnable_load_avg + min_t(unsigned long, NICE_0_LOAD,
> > > +> > > > > > > > > > > > > > > > cfs_rq->load.weight/2);
> > > > > > > return cfs_rq->runnable_load_avg;
> > > }
> >
> > cfs_rq->runnable_load_avg is for sure no greater than (in this case much less
> > than, maybe 1/2 of) load.weight, whereas load_avg is not necessarily a rock
> > in gearbox that only impedes speed up, but also speed down.
>
> Yeah, just like everything else, it'll cuts both ways (why you can't
> win the sched game). If I can believe tbench, at tasks=cpus, reducing
> lag increased utilization and reduced latency a wee bit, as did the
> reserve thing once a booboo got fixed up.
Ok, so you have a secret IDLE_RESERVE? Good luck and show it, ;)
> Makes sense, robbing Peter
> to pay Paul should work out better for Paul.
>
> NO_LB_TIP_AVG_HIGH
> Throughput 27132.9 MB/sec 96 clients 96 procs max_latency=7.656 ms
> Throughput 28464.1 MB/sec 96 clients 96 procs max_latency=9.905 ms
> Throughput 25369.8 MB/sec 96 clients 96 procs max_latency=7.192 ms
> Throughput 25670.3 MB/sec 96 clients 96 procs max_latency=5.874 ms
> Throughput 29309.3 MB/sec 96 clients 96 procs max_latency=1.331 ms
> avg 27189 1.000 6.391 1.000
>
> NO_LB_TIP_AVG_HIGH IDLE_RESERVE
> Throughput 24437.5 MB/sec 96 clients 96 procs max_latency=1.837 ms
> Throughput 29464.7 MB/sec 96 clients 96 procs max_latency=1.594 ms
> Throughput 28023.6 MB/sec 96 clients 96 procs max_latency=1.494 ms
> Throughput 28299.0 MB/sec 96 clients 96 procs max_latency=10.404 ms
> Throughput 29072.1 MB/sec 96 clients 96 procs max_latency=5.575 ms
> avg 27859 1.024 4.180 0.654
>
> LB_TIP_AVG_HIGH NO_IDLE_RESERVE
> Throughput 29068.1 MB/sec 96 clients 96 procs max_latency=5.599 ms
> Throughput 26435.6 MB/sec 96 clients 96 procs max_latency=3.703 ms
> Throughput 23930.0 MB/sec 96 clients 96 procs max_latency=7.742 ms
> Throughput 29464.2 MB/sec 96 clients 96 procs max_latency=1.549 ms
> Throughput 24250.9 MB/sec 96 clients 96 procs max_latency=1.518 ms
> avg 26629 0.979 4.022 0.629
>
> LB_TIP_AVG_HIGH IDLE_RESERVE
> Throughput 30340.1 MB/sec 96 clients 96 procs max_latency=1.465 ms
> Throughput 29042.9 MB/sec 96 clients 96 procs max_latency=4.515 ms
> Throughput 26718.7 MB/sec 96 clients 96 procs max_latency=1.822 ms
> Throughput 28694.4 MB/sec 96 clients 96 procs max_latency=1.503 ms
> Throughput 28918.2 MB/sec 96 clients 96 procs max_latency=7.599 ms
> avg 28742 1.057 3.380 0.528
>
> > But I really don't know the load references in select_task_rq() should be
> > what kind. So maybe the real issue is a mix of them, i.e., conflated balancing
> > and just wanting an idle cpu. ?
>
> Depends on the goal. For both, load lagging reality means the high
> frequency component is squelched, meaning less migration cost, but also
> higher latency due to stacking. It's a tradeoff where Chris' latency
> is everything" benchmark, and _maybe_ the real world load it's based
> upon is on Peter's end of the rob Peter to pay Paul transaction. The
> benchmark says it definitely is, the real world load may have already
> been fixed up by the select_idle_sibling() rewrite.
Obviously, load avgs are good at balancing in a larger scale in a timeframe,
so they should be used in comparing/balancing sd's not cpus. However, this
is not the case currently: avgs are mixed with idle cpu/core selection, so
I think better job can be done before and after select_idle_sibling().
For example, I don't know what the complex wake_affine() is really doing for
what. Am i missing something, you think?
Kudos to select_idle_sibling() rewrite, like Peter said, a second step and
an even third step scans are really helping, in addition to many cleanups
and refactors.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2016-05-11 12:00 +0200 |
| Message-ID | <rxz5V-5GG-13@gated-at.bofh.it> |
| In reply to | #1398811 |
On Wed, 2016-05-11 at 09:23 +0800, Yuyang Du wrote: > > Yeah, just like everything else, it'll cuts both ways (why you can't > > win the sched game). If I can believe tbench, at tasks=cpus, reducing > > lag increased utilization and reduced latency a wee bit, as did the > > reserve thing once a booboo got fixed up. > > Ok, so you have a secret IDLE_RESERVE? Good luck and show it, ;) Nothing sexy, just cpmxchg(), with the obvious test/set/clear spots. cmpxchg(&cpu_rq(cpu)->idle_latch, cpu, nr_cpu_ids) > Depends on the goal. For both, load lagging reality means the high > > frequency component is squelched, meaning less migration cost, but also > > higher latency due to stacking. It's a tradeoff where Chris' latency > > is everything" benchmark, and _maybe_ the real world load it's based > > upon is on Peter's end of the rob Peter to pay Paul transaction. The > > benchmark says it definitely is, the real world load may have already > > been fixed up by the select_idle_sibling() rewrite. > > Obviously, load avgs are good at balancing in a larger scale in a timeframe, > so they should be used in comparing/balancing sd's not cpus. However, this > is not the case currently: avgs are mixed with idle cpu/core selection, so > I think better job can be done before and after select_idle_sibling(). > > For example, I don't know what the complex wake_affine() is really doing for > what. Am i missing something, you think? wake_affine() just says no to keep us from pulling the whole load to one cache, starting massive tug-o-war with LB and nuking throughput. Everybody wants hot data, but they can't all have it and scale. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <mgalbraith@suse.de> |
|---|---|
| Date | 2016-05-09 06:00 +0200 |
| Message-ID | <rwKwp-44c-3@gated-at.bofh.it> |
| In reply to | #1396556 |
On Mon, 2016-05-09 at 02:57 +0800, Yuyang Du wrote: > In addition, I would argue maybe beefing up idle balancing is a more > productive way to spread load, as work-stealing just does what needs > to be done. And seems it has been (sub-unconsciously) neglected in this > case, :) P.S. Nope, I'm dinging up multiple spots ;-)
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web