Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1396938 > unrolled thread
| Started by | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| First post | 2016-05-09 13:00 +0200 |
| Last post | 2016-05-25 18:30 +0200 |
| Articles | 20 on this page of 30 — 8 participants |
Back to article view | Back to linux.kernel
[RFC][PATCH 0/7] sched: select_idle_siblings rewrite Peter Zijlstra <peterz@infradead.org> - 2016-05-09 13:00 +0200
[RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-09 13:00 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-11 14:00 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-11 14:40 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-11 20:20 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-11 20:30 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Michael Neuling <mikey@neuling.org> - 2016-05-12 04:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-12 07:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Michael Neuling <mikey@neuling.org> - 2016-05-12 13:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-12 13:40 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Michael Neuling <mikey@neuling.org> - 2016-05-13 02:20 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-16 16:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-17 12:30 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Srikar Dronamraju <srikar@linux.vnet.ibm.com> - 2016-05-17 13:00 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-17 13:20 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-11 19:40 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Matt Fleming <matt@codeblueprint.co.uk> - 2016-05-11 20:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Dietmar Eggemann <dietmar.eggemann@arm.com> - 2016-05-16 17:40 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-16 19:10 +0200
Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared Dietmar Eggemann <dietmar.eggemann@arm.com> - 2016-05-16 19:30 +0200
[RFC][PATCH 1/7] sched: Remove unused @cpu argument from destroy_sched_domain*() Peter Zijlstra <peterz@infradead.org> - 2016-05-09 13:00 +0200
[RFC][PATCH 3/7] sched: Introduce struct sched_domain_shared Peter Zijlstra <peterz@infradead.org> - 2016-05-09 13:00 +0200
[RFC][PATCH 6/7] sched: Optimize SCHED_SMT Peter Zijlstra <peterz@infradead.org> - 2016-05-09 13:00 +0200
Re: [RFC][PATCH 0/7] sched: select_idle_siblings rewrite Chris Mason <clm@fb.com> - 2016-05-10 03:00 +0200
Re: [RFC][PATCH 0/7] sched: select_idle_siblings rewrite Chris Mason <clm@fb.com> - 2016-05-11 16:30 +0200
[RFC][PATCH 8/7] sched/fair: Use utilization distance to filter affine sync wakeups Mike Galbraith <mgalbraith@suse.de> - 2016-05-18 08:00 +0200
Re: [RFC][PATCH 8/7] sched/fair: Use utilization distance to filter affine sync wakeups Rik van Riel <riel@redhat.com> - 2016-05-19 23:50 +0200
Re: [RFC][PATCH 8/7] sched/fair: Use utilization distance to filter affine sync wakeups Mike Galbraith <mgalbraith@suse.de> - 2016-05-20 05:00 +0200
Re: [RFC][PATCH 0/7] sched: select_idle_siblings rewrite Chris Mason <clm@fb.com> - 2016-05-25 17:00 +0200
Re: [RFC][PATCH 0/7] sched: select_idle_siblings rewrite Peter Zijlstra <peterz@infradead.org> - 2016-05-25 18:30 +0200
Page 1 of 2 [1] 2 Next page →
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-09 13:00 +0200 |
| Subject | [RFC][PATCH 0/7] sched: select_idle_siblings rewrite |
| Message-ID | <rwR4S-3QS-9@gated-at.bofh.it> |
Hai, here be a semi coherent patch series for the recent select_idle_siblings() tinkering. Happy benchmarking.. --- include/linux/sched.h | 10 ++ kernel/sched/core.c | 94 +++++++++++---- kernel/sched/fair.c | 298 ++++++++++++++++++++++++++++++++++++++--------- kernel/sched/idle_task.c | 2 +- kernel/sched/sched.h | 23 +++- kernel/time/tick-sched.c | 10 +- 6 files changed, 348 insertions(+), 89 deletions(-)
[toc] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-09 13:00 +0200 |
| Subject | [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rwR4T-3QS-23@gated-at.bofh.it> |
| In reply to | #1396938 |
Move the nr_busy_cpus thing from its hacky sd->parent->groups->sgc
location into the much more natural sched_domain_shared location.
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
---
include/linux/sched.h | 1 +
kernel/sched/core.c | 10 +++++-----
kernel/sched/fair.c | 22 ++++++++++++----------
kernel/sched/sched.h | 6 +-----
kernel/time/tick-sched.c | 10 +++++-----
5 files changed, 24 insertions(+), 25 deletions(-)
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1059,6 +1059,7 @@ struct sched_group;
struct sched_domain_shared {
atomic_t ref;
+ atomic_t nr_busy_cpus;
};
struct sched_domain {
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -5866,14 +5866,14 @@ static void destroy_sched_domains(struct
DEFINE_PER_CPU(struct sched_domain *, sd_llc);
DEFINE_PER_CPU(int, sd_llc_size);
DEFINE_PER_CPU(int, sd_llc_id);
+DEFINE_PER_CPU(struct sched_domain_shared *, sd_llc_shared);
DEFINE_PER_CPU(struct sched_domain *, sd_numa);
-DEFINE_PER_CPU(struct sched_domain *, sd_busy);
DEFINE_PER_CPU(struct sched_domain *, sd_asym);
static void update_top_cache_domain(int cpu)
{
+ struct sched_domain_shared *sds = NULL;
struct sched_domain *sd;
- struct sched_domain *busy_sd = NULL;
int id = cpu;
int size = 1;
@@ -5881,13 +5881,13 @@ static void update_top_cache_domain(int
if (sd) {
id = cpumask_first(sched_domain_span(sd));
size = cpumask_weight(sched_domain_span(sd));
- busy_sd = sd->parent; /* sd_busy */
+ sds = sd->shared;
}
- rcu_assign_pointer(per_cpu(sd_busy, cpu), busy_sd);
rcu_assign_pointer(per_cpu(sd_llc, cpu), sd);
per_cpu(sd_llc_size, cpu) = size;
per_cpu(sd_llc_id, cpu) = id;
+ rcu_assign_pointer(per_cpu(sd_llc_shared, cpu), sds);
sd = lowest_flag_domain(cpu, SD_NUMA);
rcu_assign_pointer(per_cpu(sd_numa, cpu), sd);
@@ -6184,7 +6184,6 @@ static void init_sched_groups_capacity(i
return;
update_group_capacity(sd, cpu);
- atomic_set(&sg->sgc->nr_busy_cpus, sg->group_weight);
}
/*
@@ -6388,6 +6387,7 @@ sd_init(struct sched_domain_topology_lev
sd->shared = *per_cpu_ptr(sdd->sds, sd_id);
atomic_inc(&sd->shared->ref);
+ atomic_set(&sd->shared->nr_busy_cpus, sd_weight);
#ifdef CONFIG_NUMA
} else if (sd->flags & SD_NUMA) {
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -7842,13 +7842,13 @@ static inline void set_cpu_sd_state_busy
int cpu = smp_processor_id();
rcu_read_lock();
- sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ sd = rcu_dereference(per_cpu(sd_llc, cpu));
if (!sd || !sd->nohz_idle)
goto unlock;
sd->nohz_idle = 0;
- atomic_inc(&sd->groups->sgc->nr_busy_cpus);
+ atomic_inc(&sd->shared->nr_busy_cpus);
unlock:
rcu_read_unlock();
}
@@ -7859,13 +7859,13 @@ void set_cpu_sd_state_idle(void)
int cpu = smp_processor_id();
rcu_read_lock();
- sd = rcu_dereference(per_cpu(sd_busy, cpu));
+ sd = rcu_dereference(per_cpu(sd_llc, cpu));
if (!sd || sd->nohz_idle)
goto unlock;
sd->nohz_idle = 1;
- atomic_dec(&sd->groups->sgc->nr_busy_cpus);
+ atomic_dec(&sd->shared->nr_busy_cpus);
unlock:
rcu_read_unlock();
}
@@ -8092,8 +8092,8 @@ static void nohz_idle_balance(struct rq
static inline bool nohz_kick_needed(struct rq *rq)
{
unsigned long now = jiffies;
+ struct sched_domain_shared *sds;
struct sched_domain *sd;
- struct sched_group_capacity *sgc;
int nr_busy, cpu = rq->cpu;
bool kick = false;
@@ -8121,11 +8121,13 @@ static inline bool nohz_kick_needed(stru
return true;
rcu_read_lock();
- sd = rcu_dereference(per_cpu(sd_busy, cpu));
- if (sd) {
- sgc = sd->groups->sgc;
- nr_busy = atomic_read(&sgc->nr_busy_cpus);
-
+ sds = rcu_dereference(per_cpu(sd_llc_shared, cpu));
+ if (sds) {
+ /*
+ * XXX: write a coherent comment on why we do this.
+ * See also: http:lkml.kernel.org/r/20111202010832.602203411@sbsiddha-desk.sc.intel.com
+ */
+ nr_busy = atomic_read(&sds->nr_busy_cpus);
if (nr_busy > 1) {
kick = true;
goto unlock;
--- a/kernel/sched/sched.h
+++ b/kernel/sched/sched.h
@@ -856,8 +856,8 @@ static inline struct sched_domain *lowes
DECLARE_PER_CPU(struct sched_domain *, sd_llc);
DECLARE_PER_CPU(int, sd_llc_size);
DECLARE_PER_CPU(int, sd_llc_id);
+DECLARE_PER_CPU(struct sched_domain_shared *, sd_llc_shared);
DECLARE_PER_CPU(struct sched_domain *, sd_numa);
-DECLARE_PER_CPU(struct sched_domain *, sd_busy);
DECLARE_PER_CPU(struct sched_domain *, sd_asym);
struct sched_group_capacity {
@@ -869,10 +869,6 @@ struct sched_group_capacity {
unsigned int capacity;
unsigned long next_update;
int imbalance; /* XXX unrelated to capacity but shared group state */
- /*
- * Number of busy cpus in this group.
- */
- atomic_t nr_busy_cpus;
unsigned long cpumask[0]; /* iteration mask */
};
--- a/kernel/time/tick-sched.c
+++ b/kernel/time/tick-sched.c
@@ -933,11 +933,11 @@ void tick_nohz_idle_enter(void)
WARN_ON_ONCE(irqs_disabled());
/*
- * Update the idle state in the scheduler domain hierarchy
- * when tick_nohz_stop_sched_tick() is called from the idle loop.
- * State will be updated to busy during the first busy tick after
- * exiting idle.
- */
+ * Update the idle state in the scheduler domain hierarchy
+ * when tick_nohz_stop_sched_tick() is called from the idle loop.
+ * State will be updated to busy during the first busy tick after
+ * exiting idle.
+ */
set_cpu_sd_state_idle();
local_irq_disable();
[toc] | [prev] | [next] | [standalone]
| From | Matt Fleming <matt@codeblueprint.co.uk> |
|---|---|
| Date | 2016-05-11 14:00 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxAY1-7z7-1@gated-at.bofh.it> |
| In reply to | #1396939 |
On Mon, 09 May, at 12:48:11PM, Peter Zijlstra wrote:
> Move the nr_busy_cpus thing from its hacky sd->parent->groups->sgc
> location into the much more natural sched_domain_shared location.
>
> Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
> ---
> include/linux/sched.h | 1 +
> kernel/sched/core.c | 10 +++++-----
> kernel/sched/fair.c | 22 ++++++++++++----------
> kernel/sched/sched.h | 6 +-----
> kernel/time/tick-sched.c | 10 +++++-----
> 5 files changed, 24 insertions(+), 25 deletions(-)
>
> --- a/include/linux/sched.h
> +++ b/include/linux/sched.h
> @@ -1059,6 +1059,7 @@ struct sched_group;
>
> struct sched_domain_shared {
> atomic_t ref;
> + atomic_t nr_busy_cpus;
> };
>
> struct sched_domain {
[...]
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -7842,13 +7842,13 @@ static inline void set_cpu_sd_state_busy
> int cpu = smp_processor_id();
>
> rcu_read_lock();
> - sd = rcu_dereference(per_cpu(sd_busy, cpu));
> + sd = rcu_dereference(per_cpu(sd_llc, cpu));
>
> if (!sd || !sd->nohz_idle)
> goto unlock;
> sd->nohz_idle = 0;
>
> - atomic_inc(&sd->groups->sgc->nr_busy_cpus);
> + atomic_inc(&sd->shared->nr_busy_cpus);
> unlock:
> rcu_read_unlock();
> }
This breaks my POWER7 box which presumably doesn't have SD_SHARE_PKG_RESOURCES,
NIP [c00000000012de68] .set_cpu_sd_state_idle+0x58/0x80
LR [c00000000017ded4] .tick_nohz_idle_enter+0x24/0x90
Call Trace:
[c0000007774b7cf0] [c0000007774b4080] 0xc0000007774b4080 (unreliable)
[c0000007774b7d60] [c00000000017ded4] .tick_nohz_idle_enter+0x24/0x90
[c0000007774b7dd0] [c000000000137200] .cpu_startup_entry+0xe0/0x440
[c0000007774b7ee0] [c00000000004739c] .start_secondary+0x35c/0x3a0
[c0000007774b7f90] [c000000000008bfc] start_secondary_prolog+0x10/0x14
The following fixes it,
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 978b3ef2d87e..d27153adee4d 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -7920,7 +7920,8 @@ static inline void set_cpu_sd_state_busy(void)
goto unlock;
sd->nohz_idle = 0;
- atomic_inc(&sd->shared->nr_busy_cpus);
+ if (sd->shared)
+ atomic_inc(&sd->shared->nr_busy_cpus);
unlock:
rcu_read_unlock();
}
@@ -7937,7 +7938,8 @@ void set_cpu_sd_state_idle(void)
goto unlock;
sd->nohz_idle = 1;
- atomic_dec(&sd->shared->nr_busy_cpus);
+ if (sd->shared)
+ atomic_dec(&sd->shared->nr_busy_cpus);
unlock:
rcu_read_unlock();
}
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-11 14:40 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxBAM-8ht-43@gated-at.bofh.it> |
| In reply to | #1398979 |
On Wed, May 11, 2016 at 12:55:56PM +0100, Matt Fleming wrote: > > --- a/kernel/sched/fair.c > > +++ b/kernel/sched/fair.c > > @@ -7842,13 +7842,13 @@ static inline void set_cpu_sd_state_busy > > int cpu = smp_processor_id(); > > > > rcu_read_lock(); > > - sd = rcu_dereference(per_cpu(sd_busy, cpu)); > > + sd = rcu_dereference(per_cpu(sd_llc, cpu)); > > > > if (!sd || !sd->nohz_idle) > > goto unlock; > > sd->nohz_idle = 0; > > > > - atomic_inc(&sd->groups->sgc->nr_busy_cpus); > > + atomic_inc(&sd->shared->nr_busy_cpus); > > unlock: > > rcu_read_unlock(); > > } > > This breaks my POWER7 box which presumably doesn't have SD_SHARE_PKG_RESOURCES, > Hmm, PPC folks; what does your topology look like? Currently your sched_domain_topology, as per arch/powerpc/kernel/smp.c seems to suggest your cores do not share cache at all. https://en.wikipedia.org/wiki/POWER7 seems to agree and states "4 MB L3 cache per C1 core" And http://www-03.ibm.com/systems/resources/systems_power_software_i_perfmgmt_underthehood.pdf also explicitly draws pictures with the L3 per core. _however_, that same document describes L3 inter-core fill and lateral cast-out, which sounds like the L3s work together to form a node wide caching system. Do we want to model this co-operative L3 slices thing as a sort of node-wide LLC for the purpose of the scheduler ? While we should definitely fix the assumption that an LLC exists (and I need to look at why it isn't set to the core domain instead as well), the scheduler does try and scale things by 'assuming' LLC := node. It does this for NOHZ, and these here patches under discussion would be doing the same for idle-core state. Would this make sense for power, or should we somehow think of something else?
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-11 20:20 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxGTN-5ei-19@gated-at.bofh.it> |
| In reply to | #1399024 |
On Wed, May 11, 2016 at 02:33:45PM +0200, Peter Zijlstra wrote: > Do we want to model this co-operative L3 slices thing as a sort of > node-wide LLC for the purpose of the scheduler ? > > the scheduler does try and scale things by 'assuming' LLC := node. So this whole series is about selecting idle CPUs to run stuff on. With the current PPC setup we would never consider placing a task outside of its core. If we were to add a node wide cache domain, the scheduler would look to place a task across cores if there is an idle core to be had. Would something like that be beneficial?
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-11 20:30 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxH3s-5o7-3@gated-at.bofh.it> |
| In reply to | #1399024 |
On Wed, May 11, 2016 at 02:33:45PM +0200, Peter Zijlstra wrote: > Hmm, PPC folks; what does your topology look like? > > Currently your sched_domain_topology, as per arch/powerpc/kernel/smp.c > seems to suggest your cores do not share cache at all. > > https://en.wikipedia.org/wiki/POWER7 seems to agree and states > > "4 MB L3 cache per C1 core" > > And http://www-03.ibm.com/systems/resources/systems_power_software_i_perfmgmt_underthehood.pdf > also explicitly draws pictures with the L3 per core. > > _however_, that same document describes L3 inter-core fill and lateral > cast-out, which sounds like the L3s work together to form a node wide > caching system. > > Do we want to model this co-operative L3 slices thing as a sort of > node-wide LLC for the purpose of the scheduler ? Going back a generation; Power6 seems to have a shared L3 (off package) between the two cores on the package. The current topology does not reflect that at all. And going forward a generation; Power8 seems to share the per-core (chiplet) L3 amonst all cores (chiplets) + is has the centaur (memory controller) 16M L4. So it seems the current topology setup is not describing these chips very well. Also note that the arch topology code can runtime select a topology, so you could make that topo setup micro-arch specific.
[toc] | [prev] | [next] | [standalone]
| From | Michael Neuling <mikey@neuling.org> |
|---|---|
| Date | 2016-05-12 04:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxOeC-3YK-5@gated-at.bofh.it> |
| In reply to | #1399403 |
On Wed, 2016-05-11 at 20:24 +0200, Peter Zijlstra wrote: > On Wed, May 11, 2016 at 02:33:45PM +0200, Peter Zijlstra wrote: > > > > Hmm, PPC folks; what does your topology look like? > > > > Currently your sched_domain_topology, as per arch/powerpc/kernel/smp.c > > seems to suggest your cores do not share cache at all. > > > > https://en.wikipedia.org/wiki/POWER7 seems to agree and states > > > > "4 MB L3 cache per C1 core" > > > > And http://www-03.ibm.com/systems/resources/systems_power_software_i_pe > > rfmgmt_underthehood.pdf > > also explicitly draws pictures with the L3 per core. > > > > _however_, that same document describes L3 inter-core fill and lateral > > cast-out, which sounds like the L3s work together to form a node wide > > caching system. > > > > Do we want to model this co-operative L3 slices thing as a sort of > > node-wide LLC for the purpose of the scheduler ? > Going back a generation; Power6 seems to have a shared L3 (off package) > between the two cores on the package. The current topology does not > reflect that at all. > > And going forward a generation; Power8 seems to share the per-core > (chiplet) L3 amonst all cores (chiplets) + is has the centaur (memory > controller) 16M L4. Yep, L1/L2/L3 is per core on POWER8 and POWER7. POWER6 and POWER5 (both dual core chips) had a shared off chip cache The POWER8 L4 is really a bit different as it's out in the memory controller. It's more of a memory DIMM buffer as it can only cache data associated with the physical addresses on those DIMMS. > So it seems the current topology setup is not describing these chips > very well. Also note that the arch topology code can runtime select a > topology, so you could make that topo setup micro-arch specific. We are planning on making some topology changes for the upcoming P9 which will share L2/L3 amongst pairs of cores (24 cores per chip). FWIW our P9 upstreaming is still in it's infancy since P9 is not released yet. Mike
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-12 07:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxR2O-7ar-1@gated-at.bofh.it> |
| In reply to | #1399641 |
On Thu, May 12, 2016 at 12:05:37PM +1000, Michael Neuling wrote: > On Wed, 2016-05-11 at 20:24 +0200, Peter Zijlstra wrote: > > On Wed, May 11, 2016 at 02:33:45PM +0200, Peter Zijlstra wrote: > > > > > > Hmm, PPC folks; what does your topology look like? > > > > > > Currently your sched_domain_topology, as per arch/powerpc/kernel/smp.c > > > seems to suggest your cores do not share cache at all. > > > > > > https://en.wikipedia.org/wiki/POWER7 seems to agree and states > > > > > > "4 MB L3 cache per C1 core" > > > > > > And http://www-03.ibm.com/systems/resources/systems_power_software_i_pe > > > rfmgmt_underthehood.pdf > > > also explicitly draws pictures with the L3 per core. > > > > > > _however_, that same document describes L3 inter-core fill and lateral > > > cast-out, which sounds like the L3s work together to form a node wide > > > caching system. > > > > > > Do we want to model this co-operative L3 slices thing as a sort of > > > node-wide LLC for the purpose of the scheduler ? > > Going back a generation; Power6 seems to have a shared L3 (off package) > > between the two cores on the package. The current topology does not > > reflect that at all. > > > > And going forward a generation; Power8 seems to share the per-core > > (chiplet) L3 amonst all cores (chiplets) + is has the centaur (memory > > controller) 16M L4. > > Yep, L1/L2/L3 is per core on POWER8 and POWER7. POWER6 and POWER5 (both > dual core chips) had a shared off chip cache But as per the above, Power7 and Power8 have explicit logic to share the per-core L3 with the other cores. How effective is that? From some of the slides/documents i've looked at the L3s are connected with a high-speed fabric. Suggesting that the cross-core sharing should be fairly efficient. In which case it would make sense to treat/model the combined L3 as a single large LLC covering all cores.
[toc] | [prev] | [next] | [standalone]
| From | Michael Neuling <mikey@neuling.org> |
|---|---|
| Date | 2016-05-12 13:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxWFc-4nU-17@gated-at.bofh.it> |
| In reply to | #1399668 |
On Thu, 2016-05-12 at 07:07 +0200, Peter Zijlstra wrote: > On Thu, May 12, 2016 at 12:05:37PM +1000, Michael Neuling wrote: > > > > On Wed, 2016-05-11 at 20:24 +0200, Peter Zijlstra wrote: > > > > > > On Wed, May 11, 2016 at 02:33:45PM +0200, Peter Zijlstra wrote: > > > > > > > > > > > > Hmm, PPC folks; what does your topology look like? > > > > > > > > Currently your sched_domain_topology, as per > > > > arch/powerpc/kernel/smp.c > > > > seems to suggest your cores do not share cache at all. > > > > > > > > https://en.wikipedia.org/wiki/POWER7 seems to agree and states > > > > > > > > "4 MB L3 cache per C1 core" > > > > > > > > And http://www-03.ibm.com/systems/resources/systems_power_software_ > > > > i_pe > > > > rfmgmt_underthehood.pdf > > > > also explicitly draws pictures with the L3 per core. > > > > > > > > _however_, that same document describes L3 inter-core fill and > > > > lateral > > > > cast-out, which sounds like the L3s work together to form a node > > > > wide > > > > caching system. > > > > > > > > Do we want to model this co-operative L3 slices thing as a sort of > > > > node-wide LLC for the purpose of the scheduler ? > > > Going back a generation; Power6 seems to have a shared L3 (off > > > package) > > > between the two cores on the package. The current topology does not > > > reflect that at all. > > > > > > And going forward a generation; Power8 seems to share the per-core > > > (chiplet) L3 amonst all cores (chiplets) + is has the centaur (memory > > > controller) 16M L4. > > Yep, L1/L2/L3 is per core on POWER8 and POWER7. POWER6 and POWER5 > > (both > > dual core chips) had a shared off chip cache > But as per the above, Power7 and Power8 have explicit logic to share the > per-core L3 with the other cores. > > How effective is that? From some of the slides/documents i've looked at > the L3s are connected with a high-speed fabric. Suggesting that the > cross-core sharing should be fairly efficient. I'm not sure. I thought it was mostly private but if another core was sleeping or not experiencing much cache pressure, another core could use it for some things. But I'm fuzzy on the the exact properties, sorry. > In which case it would make sense to treat/model the combined L3 as a > single large LLC covering all cores. Are you thinking it would be much cheaper to migrate a task to another core inside this chip, than to off chip? Mikey
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-12 13:40 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxX8e-4Hz-21@gated-at.bofh.it> |
| In reply to | #1399922 |
On Thu, May 12, 2016 at 09:07:52PM +1000, Michael Neuling wrote: > On Thu, 2016-05-12 at 07:07 +0200, Peter Zijlstra wrote: > > But as per the above, Power7 and Power8 have explicit logic to share the > > per-core L3 with the other cores. > > > > How effective is that? From some of the slides/documents i've looked at > > the L3s are connected with a high-speed fabric. Suggesting that the > > cross-core sharing should be fairly efficient. > > I'm not sure. I thought it was mostly private but if another core was > sleeping or not experiencing much cache pressure, another core could use it > for some things. But I'm fuzzy on the the exact properties, sorry. Right; I'm going by bits and pieces found on the tubes, so I'm just guessing ;-) But it sounds like these L3s are nowhere close to what Intel does with their L3, where each core has an L3 slice, and slices are connected on a ring to form a unified/shared cache across all cores. http://www.realworldtech.com/sandy-bridge/8/ > > In which case it would make sense to treat/model the combined L3 as a > > single large LLC covering all cores. > > Are you thinking it would be much cheaper to migrate a task to another core > inside this chip, than to off chip? Basically; and if so, if its cheap enough to shoot a task to an idle core to avoid queueing. Assuming there still is some cache residency on the old core, the inter-core fill should be much cheaper than fetching it off package (either remote cache or dram). Or at least; so goes my reasoning based on my google results.
[toc] | [prev] | [next] | [standalone]
| From | Michael Neuling <mikey@neuling.org> |
|---|---|
| Date | 2016-05-13 02:20 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <ry8ZI-dE-5@gated-at.bofh.it> |
| In reply to | #1399951 |
On Thu, 2016-05-12 at 13:33 +0200, Peter Zijlstra wrote: > On Thu, May 12, 2016 at 09:07:52PM +1000, Michael Neuling wrote: > > > > On Thu, 2016-05-12 at 07:07 +0200, Peter Zijlstra wrote: > > > > > > > > But as per the above, Power7 and Power8 have explicit logic to share > > > the > > > per-core L3 with the other cores. > > > > > > How effective is that? From some of the slides/documents i've looked > > > at > > > the L3s are connected with a high-speed fabric. Suggesting that the > > > cross-core sharing should be fairly efficient. > > I'm not sure. I thought it was mostly private but if another core was > > sleeping or not experiencing much cache pressure, another core could > > use it > > for some things. But I'm fuzzy on the the exact properties, sorry. > Right; I'm going by bits and pieces found on the tubes, so I'm just > guessing ;-) > > But it sounds like these L3s are nowhere close to what Intel does with > their L3, where each core has an L3 slice, and slices are connected on a > ring to form a unified/shared cache across all cores. > > http://www.realworldtech.com/sandy-bridge/8/ The POWER8 user manual is what you want to look at: https://www.setphaserstostun.org/power8/POWER8_UM_v1.3_16MAR2016_pub.pdf There is a section 10. "L3 Cache Overview" starting on page 128. In there it talks about L3.0 which is using the local cores L3. L3.1 which is using some other cores L3. Once the L3.0 is full, we can cast out to an L3.1 (ie. the cache on another core). L3.1 can also provide data for reads. ECO mode (section 10.4) is what I was talking about for sleeping/unused cores. That's more of a boot time (firmware option) than something we can dynamically play with at runtime (I believe), so it's not something I think is relevant here. > > > > > > > > In which case it would make sense to treat/model the combined L3 as a > > > single large LLC covering all cores. > > Are you thinking it would be much cheaper to migrate a task to another > > core > > inside this chip, than to off chip? > Basically; and if so, if its cheap enough to shoot a task to an idle > core to avoid queueing. Assuming there still is some cache residency on > the old core, the inter-core fill should be much cheaper than fetching > it off package (either remote cache or dram). So I think that will apply on POWER8. In 10.4.2 it says "The L3.1 ECO Caches will be snooped and provide intervention data similar to the L2 and L3.0 caches on the chip" That should be much faster than going to another chip or DIMM. So migrating to another core on the same chip should be faster than off chip. Mikey > Or at least; so goes my reasoning based on my google results. >
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-16 16:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzrnz-82y-1@gated-at.bofh.it> |
| In reply to | #1400430 |
On Fri, May 13, 2016 at 10:12:26AM +1000, Michael Neuling wrote:
> > Basically; and if so, if its cheap enough to shoot a task to an idle
> > core to avoid queueing. Assuming there still is some cache residency on
> > the old core, the inter-core fill should be much cheaper than fetching
> > it off package (either remote cache or dram).
>
> So I think that will apply on POWER8.
>
> In 10.4.2 it says "The L3.1 ECO Caches will be snooped and provide
> intervention data similar to the L2 and L3.0 caches on the
> chip" That should be much faster than going to another chip or DIMM.
>
> So migrating to another core on the same chip should be faster than off
> chip.
OK; so something like the below might be what you want to play with.
---
arch/powerpc/kernel/smp.c | 22 +++++++++++++++++++++-
1 file changed, 21 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/kernel/smp.c b/arch/powerpc/kernel/smp.c
index 55c924b65f71..1a54fa8a3323 100644
--- a/arch/powerpc/kernel/smp.c
+++ b/arch/powerpc/kernel/smp.c
@@ -782,6 +782,23 @@ static struct sched_domain_topology_level powerpc_topology[] = {
{ NULL, },
};
+static struct sched_domain_topology_level powerpc8_topology[] = {
+#ifdef CONFIG_SCHED_SMT
+ { cpu_smt_mask, powerpc_smt_flags, SD_INIT_NAME(SMT) },
+#endif
+#ifdef CONFIG_SCHED_MC
+ /*
+ * Model the L3.1 cache and sets the LLC as the whole package.
+ *
+ * This also ensures we try and move woken tasks to idle cores inside
+ * the package to avoid queueing.
+ */
+ { cpu_coregroup_mask, cpu_core_flags, SD_INIT_NAME(MC) },
+#endif
+ { cpu_cpu_mask, SD_INIT_NAME(DIE) },
+ { NULL, },
+};
+
void __init smp_cpus_done(unsigned int max_cpus)
{
cpumask_var_t old_mask;
@@ -806,7 +823,10 @@ void __init smp_cpus_done(unsigned int max_cpus)
dump_numa_cpu_topology();
- set_sched_topology(powerpc_topology);
+ if (cpu_has_feature(CPU_FTRS_POWER8))
+ set_sched_topology(powerpc8_topology);
+ else
+ set_sched_topology(powerpc_topology);
}
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-17 12:30 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzKqe-3kb-7@gated-at.bofh.it> |
| In reply to | #1401514 |
On Mon, May 16, 2016 at 04:00:32PM +0200, Peter Zijlstra wrote:
> On Fri, May 13, 2016 at 10:12:26AM +1000, Michael Neuling wrote:
>
> > > Basically; and if so, if its cheap enough to shoot a task to an idle
> > > core to avoid queueing. Assuming there still is some cache residency on
> > > the old core, the inter-core fill should be much cheaper than fetching
> > > it off package (either remote cache or dram).
> >
> > So I think that will apply on POWER8.
> >
> > In 10.4.2 it says "The L3.1 ECO Caches will be snooped and provide
> > intervention data similar to the L2 and L3.0 caches on the
> > chip" That should be much faster than going to another chip or DIMM.
> >
> > So migrating to another core on the same chip should be faster than off
> > chip.
>
> OK; so something like the below might be what you want to play with.
>
> ---
Maybe even something like so; which would make Power <= 6 use the
default topology and result in a shared LLC between the on package cores
etc..
Power7 is then special for not having a shared L3 but having the
asymmetric SMT thing and Power8 again gains the shared L3 while
retaining the asymmetric SMT stuff.
---
arch/powerpc/kernel/smp.c | 34 +++++++++++++++++++++++++++++++++-
1 file changed, 33 insertions(+), 1 deletion(-)
diff --git a/arch/powerpc/kernel/smp.c b/arch/powerpc/kernel/smp.c
index 55c924b65f71..09ce422fc29c 100644
--- a/arch/powerpc/kernel/smp.c
+++ b/arch/powerpc/kernel/smp.c
@@ -782,6 +782,23 @@ static struct sched_domain_topology_level powerpc_topology[] = {
{ NULL, },
};
+static struct sched_domain_topology_level powerpc8_topology[] = {
+#ifdef CONFIG_SCHED_SMT
+ { cpu_smt_mask, powerpc_smt_flags, SD_INIT_NAME(SMT) },
+#endif
+#ifdef CONFIG_SCHED_MC
+ /*
+ * Model the L3.1 cache and sets the LLC as the whole package.
+ *
+ * This also ensures we try and move woken tasks to idle cores inside
+ * the package to avoid queueing.
+ */
+ { cpu_coregroup_mask, cpu_core_flags, SD_INIT_NAME(MC) },
+#endif
+ { cpu_cpu_mask, SD_INIT_NAME(DIE) },
+ { NULL, },
+};
+
void __init smp_cpus_done(unsigned int max_cpus)
{
cpumask_var_t old_mask;
@@ -806,7 +823,22 @@ void __init smp_cpus_done(unsigned int max_cpus)
dump_numa_cpu_topology();
- set_sched_topology(powerpc_topology);
+ if (cpu_has_feature(CPU_FTRS_POWER7)) {
+ /*
+ * Power7 topology is special because it doesn't have
+ * a shared L3 between cores and has the ASYM(metric)
+ * SMT thing.
+ */
+ set_sched_topology(powerpc_topology);
+ }
+
+ if (cpu_has_feature(CPU_FTRS_POWER8)) {
+ /*
+ * Power8 differs from Power7 in that it does sort-of
+ * have a shared L3 again.
+ */
+ set_sched_topology(powerpc8_topology);
+ }
}
[toc] | [prev] | [next] | [standalone]
| From | Srikar Dronamraju <srikar@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-05-17 13:00 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzKTg-3vk-19@gated-at.bofh.it> |
| In reply to | #1402271 |
> > Maybe even something like so; which would make Power <= 6 use the > default topology and result in a shared LLC between the on package cores > etc.. > > Power7 is then special for not having a shared L3 but having the > asymmetric SMT thing and Power8 again gains the shared L3 while > retaining the asymmetric SMT stuff. Asymmetric SMT is only for Power 7. Its not enabled for Power 8. On Power8, if there are only 2 active threads assume thread 0 and thread 4 running in smt 8 mode, it would internally run as if running in smt 2 mode without any scheduler tweaks. (This is unlike power 7). -- Thanks and Regards Srikar Dronamraju
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-17 13:20 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzLcC-3QQ-21@gated-at.bofh.it> |
| In reply to | #1402287 |
On Tue, May 17, 2016 at 04:22:04PM +0530, Srikar Dronamraju wrote: > > > > Maybe even something like so; which would make Power <= 6 use the > > default topology and result in a shared LLC between the on package cores > > etc.. > > > > Power7 is then special for not having a shared L3 but having the > > asymmetric SMT thing and Power8 again gains the shared L3 while > > retaining the asymmetric SMT stuff. > > > Asymmetric SMT is only for Power 7. Its not enabled for Power 8. > On Power8, if there are only 2 active threads assume thread 0 and thread > 4 running in smt 8 mode, it would internally run as if running in smt 2 > mode without any scheduler tweaks. (This is unlike power 7). Ah, good to know. In that case, you can use the generic topology for everything except power7. IOW, just make that existing set_sched_topology() call depend on CPU_FTRS_POWER7 or so (if that is the correct magic incantation).
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-11 19:40 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxGh4-4oy-17@gated-at.bofh.it> |
| In reply to | #1398979 |
On Wed, May 11, 2016 at 12:55:56PM +0100, Matt Fleming wrote:
> This breaks my POWER7 box which presumably doesn't have SD_SHARE_PKG_RESOURCES,
> index 978b3ef2d87e..d27153adee4d 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -7920,7 +7920,8 @@ static inline void set_cpu_sd_state_busy(void)
> goto unlock;
> sd->nohz_idle = 0;
>
> - atomic_inc(&sd->shared->nr_busy_cpus);
> + if (sd->shared)
> + atomic_inc(&sd->shared->nr_busy_cpus);
> unlock:
> rcu_read_unlock();
> }
Ah, no, the problem is that while it does have SHARE_PKG_RESOURCES (in
its SMT domain -- SMT threads share all cache after all), I failed to
connect the sched_domain_shared structure for it.
Does something like this also work?
---
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -6389,10 +6389,6 @@ sd_init(struct sched_domain_topology_lev
sd->cache_nice_tries = 1;
sd->busy_idx = 2;
- sd->shared = *per_cpu_ptr(sdd->sds, sd_id);
- atomic_inc(&sd->shared->ref);
- atomic_set(&sd->shared->nr_busy_cpus, sd_weight);
-
#ifdef CONFIG_NUMA
} else if (sd->flags & SD_NUMA) {
sd->cache_nice_tries = 2;
@@ -6414,6 +6410,16 @@ sd_init(struct sched_domain_topology_lev
sd->idle_idx = 1;
}
+ /*
+ * For all levels sharing cache; connect a sched_domain_shared
+ * instance.
+ */
+ if (sd->flags & SH_SHARED_PKG_RESOURCES) {
+ sd->shared = *per_cpu_ptr(sdd->sds, sd_id);
+ atomic_inc(&sd->shared->ref);
+ atomic_set(&sd->shared->nr_busy_cpus, sd_weight);
+ }
+
sd->private = sdd;
return sd;
[toc] | [prev] | [next] | [standalone]
| From | Matt Fleming <matt@codeblueprint.co.uk> |
|---|---|
| Date | 2016-05-11 20:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rxGK6-57i-3@gated-at.bofh.it> |
| In reply to | #1399387 |
On Wed, 11 May, at 07:37:52PM, Peter Zijlstra wrote: > On Wed, May 11, 2016 at 12:55:56PM +0100, Matt Fleming wrote: > > > This breaks my POWER7 box which presumably doesn't have SD_SHARE_PKG_RESOURCES, > > > index 978b3ef2d87e..d27153adee4d 100644 > > --- a/kernel/sched/fair.c > > +++ b/kernel/sched/fair.c > > @@ -7920,7 +7920,8 @@ static inline void set_cpu_sd_state_busy(void) > > goto unlock; > > sd->nohz_idle = 0; > > > > - atomic_inc(&sd->shared->nr_busy_cpus); > > + if (sd->shared) > > + atomic_inc(&sd->shared->nr_busy_cpus); > > unlock: > > rcu_read_unlock(); > > } > > > Ah, no, the problem is that while it does have SHARE_PKG_RESOURCES (in > its SMT domain -- SMT threads share all cache after all), I failed to > connect the sched_domain_shared structure for it. > > Does something like this also work? Yep, it does.
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2016-05-16 17:40 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzsMG-lp-5@gated-at.bofh.it> |
| In reply to | #1396939 |
On 09/05/16 11:48, Peter Zijlstra wrote:
Couldn't you just always access sd->shared via
sd = rcu_dereference(per_cpu(sd_llc, cpu)) for
updating nr_busy_cpus?
The call_rcu() thing is on the sd any way.
@@ -5879,7 +5879,6 @@ static void destroy_sched_domains(struct sched_domain *sd)
DEFINE_PER_CPU(struct sched_domain *, sd_llc);
DEFINE_PER_CPU(int, sd_llc_size);
DEFINE_PER_CPU(int, sd_llc_id);
-DEFINE_PER_CPU(struct sched_domain_shared *, sd_llc_shared);
DEFINE_PER_CPU(struct sched_domain *, sd_numa);
DEFINE_PER_CPU(struct sched_domain *, sd_asym);
@@ -5900,7 +5899,6 @@ static void update_top_cache_domain(int cpu)
rcu_assign_pointer(per_cpu(sd_llc, cpu), sd);
per_cpu(sd_llc_size, cpu) = size;
per_cpu(sd_llc_id, cpu) = id;
- rcu_assign_pointer(per_cpu(sd_llc_shared, cpu), sds);
sd = lowest_flag_domain(cpu, SD_NUMA);
rcu_assign_pointer(per_cpu(sd_numa, cpu), sd);
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 0d8dad2972b6..5aed6089dae8 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -8136,7 +8136,6 @@ static void nohz_idle_balance(struct rq *this_rq, enum cpu_idle_type idle)
static inline bool nohz_kick_needed(struct rq *rq)
{
unsigned long now = jiffies;
- struct sched_domain_shared *sds;
struct sched_domain *sd;
int nr_busy, cpu = rq->cpu;
bool kick = false;
@@ -8165,13 +8164,13 @@ static inline bool nohz_kick_needed(struct rq *rq)
return true;
rcu_read_lock();
- sds = rcu_dereference(per_cpu(sd_llc_shared, cpu));
- if (sds) {
+ sd = rcu_dereference(per_cpu(sd_llc, cpu));
+ if (sd) {
/*
* XXX: write a coherent comment on why we do this.
* See also: http:lkml.kernel.org/r/20111202010832.602203411@sbsiddha-desk.sc.intel.com
*/
- nr_busy = atomic_read(&sds->nr_busy_cpus);
+ nr_busy = atomic_read(&sd->shared->nr_busy_cpus);
if (nr_busy > 1) {
kick = true;
goto unlock;
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-16 19:10 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzubM-1me-13@gated-at.bofh.it> |
| In reply to | #1401576 |
On Mon, May 16, 2016 at 04:31:08PM +0100, Dietmar Eggemann wrote: > On 09/05/16 11:48, Peter Zijlstra wrote: > > Couldn't you just always access sd->shared via > sd = rcu_dereference(per_cpu(sd_llc, cpu)) for > updating nr_busy_cpus? Sure; but why would I want to add that extra dereference? Note that in the next patch I add more users of sd_llc_shared.
[toc] | [prev] | [next] | [standalone]
| From | Dietmar Eggemann <dietmar.eggemann@arm.com> |
|---|---|
| Date | 2016-05-16 19:30 +0200 |
| Subject | Re: [RFC][PATCH 4/7] sched: Replace sd_busy/nr_busy_cpus with sched_domain_shared |
| Message-ID | <rzuve-1tB-35@gated-at.bofh.it> |
| In reply to | #1401643 |
On 16/05/16 18:02, Peter Zijlstra wrote: > On Mon, May 16, 2016 at 04:31:08PM +0100, Dietmar Eggemann wrote: >> On 09/05/16 11:48, Peter Zijlstra wrote: >> >> Couldn't you just always access sd->shared via >> sd = rcu_dereference(per_cpu(sd_llc, cpu)) for >> updating nr_busy_cpus? > > Sure; but why would I want to add that extra dereference? Note that in > the next patch I add more users of sd_llc_shared. > I see ... I thought because you do this already in set_cpu_sd_state_[busy|idle]. But there you need the sd reference to set sd->nohz_idle already.
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web