Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1426544 > unrolled thread

[RFC 0/4] sched/fair: Rebalance tasks based on affinity

Started byJiri Olsa <jolsa@kernel.org>
First post2016-06-20 14:30 +0200
Last post2016-06-20 20:50 +0200
Articles 6 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [RFC 0/4] sched/fair: Rebalance tasks based on affinity Jiri Olsa <jolsa@kernel.org> - 2016-06-20 14:30 +0200
    [PATCH 4/4] sched/fair: Add schedstat debug values for REBALANCE_AFFINITY Jiri Olsa <jolsa@kernel.org> - 2016-06-20 14:30 +0200
    [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks Jiri Olsa <jolsa@kernel.org> - 2016-06-20 14:30 +0200
      Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance  callbacks Jiri Olsa <jolsa@redhat.com> - 2016-06-20 18:40 +0200
        Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-06-21 11:10 +0200
      Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance  callbacks Peter Zijlstra <peterz@infradead.org> - 2016-06-20 20:50 +0200

#1426544 — [RFC 0/4] sched/fair: Rebalance tasks based on affinity

FromJiri Olsa <jolsa@kernel.org>
Date2016-06-20 14:30 +0200
Subject[RFC 0/4] sched/fair: Rebalance tasks based on affinity
Message-ID<rM6uZ-8aw-7@gated-at.bofh.it>
hi,
please consider this to be more of a question
for opinions on how to solve this issue ;-)

I'm following up on my previous post:
  http://marc.info/?t=145975837400001&r=1&w=2

The patchset is working for my testcase and our
other tests looks good so far, but I'm not sure
I haven't broken something else or used the right
wheel to implement this.

The issue itself and the fix are described in
changelog of patch 3/4.

thanks for any feedback,
jirka


---
Jiri Olsa (4):
      sched/fair: Introduce sched_entity::dont_balance
      sched/fair: Introduce idle enter/exit balance callbacks
      sched/fair: Add REBALANCE_AFFINITY rebalancing code
      sched/fair: Add schedstat debug values for REBALANCE_AFFINITY

 include/linux/sched.h   |   4 +++
 kernel/sched/debug.c    |   4 +++
 kernel/sched/fair.c     | 170 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++------
 kernel/sched/features.h |   1 +
 kernel/sched/idle.c     |   2 ++
 kernel/sched/sched.h    |   8 ++++++
 6 files changed, 181 insertions(+), 8 deletions(-)

[toc] | [next] | [standalone]


#1426551 — [PATCH 4/4] sched/fair: Add schedstat debug values for REBALANCE_AFFINITY

FromJiri Olsa <jolsa@kernel.org>
Date2016-06-20 14:30 +0200
Subject[PATCH 4/4] sched/fair: Add schedstat debug values for REBALANCE_AFFINITY
Message-ID<rM6v0-8aw-67@gated-at.bofh.it>
In reply to#1426544
It was helpful to watch several stats when testing
this feature. Adding the most useful.

Signed-off-by: Jiri Olsa <jolsa@kernel.org>
---
 include/linux/sched.h |  2 ++
 kernel/sched/debug.c  |  4 ++++
 kernel/sched/fair.c   | 15 ++++++++++++++-
 kernel/sched/sched.h  |  5 +++++
 4 files changed, 25 insertions(+), 1 deletion(-)

diff --git a/include/linux/sched.h b/include/linux/sched.h
index 0e6ac882283b..4b772820436b 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1305,6 +1305,8 @@ struct sched_statistics {
 	u64			nr_failed_migrations_running;
 	u64			nr_failed_migrations_hot;
 	u64			nr_forced_migrations;
+	u64			nr_dont_balance;
+	u64			nr_balanced_affinity;
 
 	u64			nr_wakeups;
 	u64			nr_wakeups_sync;
diff --git a/kernel/sched/debug.c b/kernel/sched/debug.c
index 2a0a9995256d..5947558dab65 100644
--- a/kernel/sched/debug.c
+++ b/kernel/sched/debug.c
@@ -635,6 +635,9 @@ do {									\
 		P(sched_goidle);
 		P(ttwu_count);
 		P(ttwu_local);
+		P(nr_dont_balance);
+		P(nr_affinity_out);
+		P(nr_affinity_in);
 	}
 
 #undef P
@@ -912,6 +915,7 @@ void proc_sched_show_task(struct task_struct *p, struct seq_file *m)
 		P(se.statistics.nr_wakeups_affine_attempts);
 		P(se.statistics.nr_wakeups_passive);
 		P(se.statistics.nr_wakeups_idle);
+		P(se.statistics.nr_dont_balance);
 
 		avg_atom = p->se.sum_exec_runtime;
 		if (nr_switches)
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 736e525e189c..a4f1ed403f1e 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -902,7 +902,15 @@ update_stats_curr_start(struct cfs_rq *cfs_rq, struct sched_entity *se)
 
 static bool dont_balance(struct task_struct *p)
 {
-	return sched_feat(REBALANCE_AFFINITY) && p->se.dont_balance;
+	bool dont_balance = false;
+
+	if (sched_feat(REBALANCE_AFFINITY) && p->se.dont_balance) {
+		dont_balance = true;
+		schedstat_inc(task_rq(p), nr_dont_balance);
+		schedstat_inc(p, se.statistics.nr_dont_balance);
+	}
+
+	return dont_balance;
 }
 
 /**************************************************
@@ -7903,6 +7911,7 @@ static void rebalance_affinity(struct rq *rq)
 		if (cpu >= nr_cpu_ids)
 			continue;
 
+		schedstat_inc(rq, nr_affinity_out);
 		__detach_task(p, rq, cpu);
 		raw_spin_unlock(&rq->lock);
 
@@ -7911,6 +7920,10 @@ static void rebalance_affinity(struct rq *rq)
 		raw_spin_lock(&dst_rq->lock);
 		attach_task(dst_rq, p);
 		p->se.dont_balance = true;
+
+		schedstat_inc(p, se.statistics.nr_balanced_affinity);
+		schedstat_inc(dst_rq, nr_affinity_in);
+
 		raw_spin_unlock(&dst_rq->lock);
 
 		local_irq_restore(flags);
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index d1a6224cd140..f086e233f7e6 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -701,6 +701,11 @@ struct rq {
 	/* try_to_wake_up() stats */
 	unsigned int ttwu_count;
 	unsigned int ttwu_local;
+
+	/* rebalance_affinity() stats */
+	unsigned int nr_dont_balance;
+	unsigned int nr_affinity_out;
+	unsigned int nr_affinity_in;
 #endif
 
 #ifdef CONFIG_SMP
-- 
2.4.11

[toc] | [prev] | [next] | [standalone]


#1426552 — [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks

FromJiri Olsa <jolsa@kernel.org>
Date2016-06-20 14:30 +0200
Subject[PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks
Message-ID<rM6v0-8aw-69@gated-at.bofh.it>
In reply to#1426544
Introducing idle enter/exit balance callbacks to keep
balance.idle_cpus_mask cpumask of current idle cpus
in system.

It's used only when REBALANCE_AFFINITY feature is
switched on. The code functionality of this feature
is introduced in following patch.

Signed-off-by: Jiri Olsa <jolsa@kernel.org>
---
 kernel/sched/fair.c  | 32 ++++++++++++++++++++++++++++++++
 kernel/sched/idle.c  |  2 ++
 kernel/sched/sched.h |  3 +++
 3 files changed, 37 insertions(+)

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index f19c9435c64d..78c4127f2f3a 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -7802,6 +7802,37 @@ static inline int on_null_domain(struct rq *rq)
 	return unlikely(!rcu_dereference_sched(rq->sd));
 }
 
+static struct {
+	cpumask_var_t	idle_cpus_mask;
+	atomic_t	nr_cpus;
+} balance ____cacheline_aligned;
+
+void sched_idle_enter(int cpu)
+{
+	if (!sched_feat(REBALANCE_AFFINITY))
+		return;
+
+	if (!cpu_active(cpu))
+		return;
+
+	if (on_null_domain(cpu_rq(cpu)))
+		return;
+
+	cpumask_set_cpu(cpu, balance.idle_cpus_mask);
+	atomic_inc(&balance.nr_cpus);
+}
+
+void sched_idle_exit(int cpu)
+{
+	if (!sched_feat(REBALANCE_AFFINITY))
+		return;
+
+	if (likely(cpumask_test_cpu(cpu, balance.idle_cpus_mask))) {
+		cpumask_clear_cpu(cpu, balance.idle_cpus_mask);
+		atomic_dec(&balance.nr_cpus);
+	}
+}
+
 #ifdef CONFIG_NO_HZ_COMMON
 /*
  * idle load balancing details
@@ -8731,6 +8762,7 @@ __init void init_sched_fair_class(void)
 	nohz.next_balance = jiffies;
 	zalloc_cpumask_var(&nohz.idle_cpus_mask, GFP_NOWAIT);
 #endif
+	zalloc_cpumask_var(&balance.idle_cpus_mask, GFP_NOWAIT);
 #endif /* SMP */
 
 }
diff --git a/kernel/sched/idle.c b/kernel/sched/idle.c
index db4ff7c100b9..0cc0109eacd3 100644
--- a/kernel/sched/idle.c
+++ b/kernel/sched/idle.c
@@ -216,6 +216,7 @@ static void cpu_idle_loop(void)
 		__current_set_polling();
 		quiet_vmstat();
 		tick_nohz_idle_enter();
+		sched_idle_enter(cpu);
 
 		while (!need_resched()) {
 			check_pgt_cache();
@@ -256,6 +257,7 @@ static void cpu_idle_loop(void)
 		 */
 		preempt_set_need_resched();
 		tick_nohz_idle_exit();
+		sched_idle_exit(cpu);
 		__current_clr_polling();
 
 		/*
diff --git a/kernel/sched/sched.h b/kernel/sched/sched.h
index 72f1f3087b04..d1a6224cd140 100644
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -1692,6 +1692,9 @@ extern void init_dl_rq(struct dl_rq *dl_rq);
 extern void cfs_bandwidth_usage_inc(void);
 extern void cfs_bandwidth_usage_dec(void);
 
+void sched_idle_enter(int cpu);
+void sched_idle_exit(int cpu);
+
 #ifdef CONFIG_NO_HZ_COMMON
 enum rq_nohz_flag_bits {
 	NOHZ_TICK_STOPPED,
-- 
2.4.11

[toc] | [prev] | [next] | [standalone]


#1426790 — Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks

FromJiri Olsa <jolsa@redhat.com>
Date2016-06-20 18:40 +0200
SubjectRe: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks
Message-ID<rMaoV-2aV-13@gated-at.bofh.it>
In reply to#1426552
On Mon, Jun 20, 2016 at 04:30:03PM +0200, Peter Zijlstra wrote:
> On Mon, Jun 20, 2016 at 02:15:12PM +0200, Jiri Olsa wrote:
> > Introducing idle enter/exit balance callbacks to keep
> > balance.idle_cpus_mask cpumask of current idle cpus
> > in system.
> > 
> > It's used only when REBALANCE_AFFINITY feature is
> > switched on. The code functionality of this feature
> > is introduced in following patch.
> > 
> > Signed-off-by: Jiri Olsa <jolsa@kernel.org>
> > ---
> >  kernel/sched/fair.c  | 32 ++++++++++++++++++++++++++++++++
> >  kernel/sched/idle.c  |  2 ++
> >  kernel/sched/sched.h |  3 +++
> >  3 files changed, 37 insertions(+)
> > 
> > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> > index f19c9435c64d..78c4127f2f3a 100644
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -7802,6 +7802,37 @@ static inline int on_null_domain(struct rq *rq)
> >  	return unlikely(!rcu_dereference_sched(rq->sd));
> >  }
> >  
> > +static struct {
> > +	cpumask_var_t	idle_cpus_mask;
> > +	atomic_t	nr_cpus;
> > +} balance ____cacheline_aligned;
> 
> How is this different from the nohz idle cpu mask?

well, the nohz idle cpu mask is deep in the nohz code:

tick_irq_exit
  if ((idle_cpu(cpu) && !need_resched()) || tick_nohz_full_cpu(cpu)) {
    if (!in_interrupt())
      tick_nohz_irq_exit
        if (ts->inidle) {
          __tick_nohz_idle_enter(ts)
            tick_nohz_idle_enter
             __tick_nohz_idle_enter
               if (can_stop_idle_tick(cpu, ts)) {
                 tick_nohz_stop_sched_tick
                   if (ts->tick_stopped)
                     nohz_balance_enter_idle(cpu) {
                       set idle mask for cpu
                     }
        } else {
                tick_nohz_full_update_tick(ts);
                 if (can_stop_full_tick(ts))
                    tick_nohz_stop_sched_tick
                      if (ts->tick_stopped)
                        nohz_balance_enter_idle(cpu) {
                          set idle mask for cpu
                        }
	}

cpu_idle_loop
  tick_nohz_idle_enter
     __tick_nohz_idle_enter(ts)
       tick_nohz_idle_enter
        __tick_nohz_idle_enter
          if (can_stop_idle_tick(cpu, ts)) {
            tick_nohz_stop_sched_tick
              if (ts->tick_stopped)
                nohz_balance_enter_idle(cpu) {
                  set idle mask for cpu
                }


sched_cpu_dying
  nohz_balance_exit_idle(cpu) {
    unset idle mask for cpu
  }

trigger_load_balance
  nohz_kick_needed
    nohz_balance_exit_idle(cpu) {
      unset idle mask for cpu
    }

... I might have missed some of the paths


so it might not be that easy to switch it to use the easy
change I added as part of this RFC, but AFAIU it should be
the same idle mask, but this approach might be too naive
and miss some idle enter/exit paths.. CC-ing Frederic

thanks,
jirka

[toc] | [prev] | [next] | [standalone]


#1427486 — Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks

FromPeter Zijlstra <peterz@infradead.org>
Date2016-06-21 11:10 +0200
SubjectRe: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks
Message-ID<rMpQZ-3Mw-1@gated-at.bofh.it>
In reply to#1426790
On Mon, Jun 20, 2016 at 06:27:41PM +0200, Jiri Olsa wrote:
> so it might not be that easy to switch it to use the easy
> change I added as part of this RFC, but AFAIU it should be
> the same idle mask, but this approach might be too naive
> and miss some idle enter/exit paths.. CC-ing Frederic

Right, for RFC this will do; but I really want to avoid adding another
global atomic bitmask on the idle path.

In any case, I'll go look over the actual patches sometime soon, I still
got a few other patches series to stare at first.

[toc] | [prev] | [next] | [standalone]


#1426928 — Re: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks

FromPeter Zijlstra <peterz@infradead.org>
Date2016-06-20 20:50 +0200
SubjectRe: [PATCH 2/4] sched/fair: Introduce idle enter/exit balance callbacks
Message-ID<rMaoV-2aV-15@gated-at.bofh.it>
In reply to#1426552
On Mon, Jun 20, 2016 at 02:15:12PM +0200, Jiri Olsa wrote:
> Introducing idle enter/exit balance callbacks to keep
> balance.idle_cpus_mask cpumask of current idle cpus
> in system.
> 
> It's used only when REBALANCE_AFFINITY feature is
> switched on. The code functionality of this feature
> is introduced in following patch.
> 
> Signed-off-by: Jiri Olsa <jolsa@kernel.org>
> ---
>  kernel/sched/fair.c  | 32 ++++++++++++++++++++++++++++++++
>  kernel/sched/idle.c  |  2 ++
>  kernel/sched/sched.h |  3 +++
>  3 files changed, 37 insertions(+)
> 
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index f19c9435c64d..78c4127f2f3a 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -7802,6 +7802,37 @@ static inline int on_null_domain(struct rq *rq)
>  	return unlikely(!rcu_dereference_sched(rq->sd));
>  }
>  
> +static struct {
> +	cpumask_var_t	idle_cpus_mask;
> +	atomic_t	nr_cpus;
> +} balance ____cacheline_aligned;

How is this different from the nohz idle cpu mask?

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web