Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1719107 > unrolled thread
| Started by | Kemi Wang <kemi.wang@intel.com> |
|---|---|
| First post | 2017-08-24 12:10 +0200 |
| Last post | 2017-08-25 10:10 +0200 |
| Articles | 4 — 2 participants |
Back to article view | Back to linux.kernel
[PATCH v2 0/3] Separate NUMA statistics from zone statistics Kemi Wang <kemi.wang@intel.com> - 2017-08-24 12:10 +0200
[PATCH v2 3/3] mm: Consider the number in local CPUs when *reads* NUMA stats Kemi Wang <kemi.wang@intel.com> - 2017-08-24 12:10 +0200
[PATCH v2 2/3] mm: Update NUMA counter threshold size Kemi Wang <kemi.wang@intel.com> - 2017-08-24 12:10 +0200
Re: [PATCH v2 0/3] Separate NUMA statistics from zone statistics Mel Gorman <mgorman@techsingularity.net> - 2017-08-25 10:10 +0200
| From | Kemi Wang <kemi.wang@intel.com> |
|---|---|
| Date | 2017-08-24 12:10 +0200 |
| Subject | [PATCH v2 0/3] Separate NUMA statistics from zone statistics |
| Message-ID | <uhXfk-52T-13@gated-at.bofh.it> |
Each page allocation updates a set of per-zone statistics with a call to
zone_statistics(). As discussed in 2017 MM summit, these are a substantial
source of overhead in the page allocator and are very rarely consumed. This
significant overhead in cache bouncing caused by zone counters (NUMA
associated counters) update in parallel in multi-threaded page allocation
(pointed out by Dave Hansen).
A link to the MM summit slides:
http://people.netfilter.org/hawk/presentations/MM-summit2017/MM-summit2017
-JesperBrouer.pdf
To mitigate this overhead, this patchset separates NUMA statistics from
zone statistics framework, and update NUMA counter threshold to a fixed
size of MAX_U16 - 2, as a small threshold greatly increases the update
frequency of the global counter from local per cpu counter (suggested by
Ying Huang). The rationality is that these statistics counters don't need
to be read often, unlike other VM counters, so it's not a problem to use a
large threshold and make readers more expensive.
With this patchset, we see 31.3% drop of CPU cycles(537-->369, see below)
for per single page allocation and reclaim on Jesper's page_bench03
benchmark. Meanwhile, this patchset keeps the same style of virtual memory
statistics with little end-user-visible effects (only move the numa stats
to show behind zone page stats, see the first patch for details).
I did an experiment of single page allocation and reclaim concurrently
using Jesper's page_bench03 benchmark on a 2-Socket Broadwell-based server
(88 processors with 126G memory) with different size of threshold of pcp
counter.
Benchmark provided by Jesper D Brouer(increase loop times to 10000000):
https://github.com/netoptimizer/prototype-kernel/tree/master/kernel/mm/
bench
Threshold CPU cycles Throughput(88 threads)
32 799 241760478
64 640 301628829
125 537 358906028 <==> system by default
256 468 412397590
512 428 450550704
4096 399 482520943
20000 394 489009617
30000 395 488017817
65533 369(-31.3%) 521661345(+45.3%) <==> with this patchset
N/A 342(-36.3%) 562900157(+56.8%) <==> disable zone_statistics
Kemi Wang (3):
mm: Change the call sites of numa statistics items
mm: Update NUMA counter threshold size
mm: Consider the number in local CPUs when *reads* NUMA stats
drivers/base/node.c | 22 ++++---
include/linux/mmzone.h | 24 +++++---
include/linux/vmstat.h | 33 +++++++++++
mm/page_alloc.c | 10 ++--
mm/vmstat.c | 152 +++++++++++++++++++++++++++++++++++++++++++++++--
5 files changed, 217 insertions(+), 24 deletions(-)
--
2.7.4
[toc] | [next] | [standalone]
| From | Kemi Wang <kemi.wang@intel.com> |
|---|---|
| Date | 2017-08-24 12:10 +0200 |
| Subject | [PATCH v2 3/3] mm: Consider the number in local CPUs when *reads* NUMA stats |
| Message-ID | <uhXfk-52T-27@gated-at.bofh.it> |
| In reply to | #1719107 |
To avoid deviation, the per cpu number of NUMA stats in vm_numa_stat_diff[]
is included when a user *reads* the NUMA stats.
Since NUMA stats does not be read by users frequently, and kernel does not
need it to make a decision, it will not be a problem to make the readers
more expensive.
Changelog:
v2:
a) new creation.
Signed-off-by: Kemi Wang <kemi.wang@intel.com>
---
include/linux/vmstat.h | 6 +++++-
mm/vmstat.c | 9 +++++++--
2 files changed, 12 insertions(+), 3 deletions(-)
diff --git a/include/linux/vmstat.h b/include/linux/vmstat.h
index a29bd98..72e9ca6 100644
--- a/include/linux/vmstat.h
+++ b/include/linux/vmstat.h
@@ -125,10 +125,14 @@ static inline unsigned long global_numa_state(enum numa_stat_item item)
return x;
}
-static inline unsigned long zone_numa_state(struct zone *zone,
+static inline unsigned long zone_numa_state_snapshot(struct zone *zone,
enum numa_stat_item item)
{
long x = atomic_long_read(&zone->vm_numa_stat[item]);
+ int cpu;
+
+ for_each_online_cpu(cpu)
+ x += per_cpu_ptr(zone->pageset, cpu)->vm_numa_stat_diff[item];
return x;
}
diff --git a/mm/vmstat.c b/mm/vmstat.c
index b015f39..abeab81 100644
--- a/mm/vmstat.c
+++ b/mm/vmstat.c
@@ -895,6 +895,10 @@ unsigned long sum_zone_node_page_state(int node,
return count;
}
+/*
+ * Determine the per node value of a numa stat item. To avoid deviation,
+ * the per cpu stat number in vm_numa_stat_diff[] is also included.
+ */
unsigned long sum_zone_numa_state(int node,
enum numa_stat_item item)
{
@@ -903,7 +907,7 @@ unsigned long sum_zone_numa_state(int node,
unsigned long count = 0;
for (i = 0; i < MAX_NR_ZONES; i++)
- count += zone_numa_state(zones + i, item);
+ count += zone_numa_state_snapshot(zones + i, item);
return count;
}
@@ -1534,7 +1538,7 @@ static void zoneinfo_show_print(struct seq_file *m, pg_data_t *pgdat,
for (i = 0; i < NR_VM_NUMA_STAT_ITEMS; i++)
seq_printf(m, "\n %-12s %lu",
vmstat_text[i + NR_VM_ZONE_STAT_ITEMS],
- zone_numa_state(zone, i));
+ zone_numa_state_snapshot(zone, i));
#endif
seq_printf(m, "\n pagesets");
@@ -1790,6 +1794,7 @@ static bool need_update(int cpu)
#ifdef CONFIG_NUMA
BUILD_BUG_ON(sizeof(p->vm_numa_stat_diff[0]) != 2);
#endif
+
/*
* The fast way of checking if there are any vmstat diffs.
* This works because the diffs are byte sized items.
--
2.7.4
[toc] | [prev] | [next] | [standalone]
| From | Kemi Wang <kemi.wang@intel.com> |
|---|---|
| Date | 2017-08-24 12:10 +0200 |
| Subject | [PATCH v2 2/3] mm: Update NUMA counter threshold size |
| Message-ID | <uhXfk-52T-21@gated-at.bofh.it> |
| In reply to | #1719107 |
There is significant overhead in cache bouncing caused by zone counters
(NUMA associated counters) update in parallel in multi-threaded page
allocation (suggested by Dave Hansen).
This patch updates NUMA counter threshold to a fixed size of MAX_U16 - 2,
as a small threshold greatly increases the update frequency of the global
counter from local per cpu counter(suggested by Ying Huang).
The rationality is that these statistics counters don't affect the kernel's
decision, unlike other VM counters, so it's not a problem to use a large
threshold.
With this patchset, we see 31.3% drop of CPU cycles(537-->369) for per
single page allocation and reclaim on Jesper's page_bench03 benchmark.
Benchmark provided by Jesper D Brouer(increase loop times to 10000000):
https://github.com/netoptimizer/prototype-kernel/tree/master/kernel/mm/
bench
Threshold CPU cycles Throughput(88 threads)
32 799 241760478
64 640 301628829
125 537 358906028 <==> system by default (base)
256 468 412397590
512 428 450550704
4096 399 482520943
20000 394 489009617
30000 395 488017817
65533 369(-31.3%) 521661345(+45.3%) <==> with this patchset
N/A 342(-36.3%) 562900157(+56.8%) <==> disable zone_statistics
Changelog:
v2:
a) Change the type of vm_numa_stat_diff[] from s16 to u16, since numa
stats counter is always a incremental field.
b) Remove numa_stat_threshold field in struct per_cpu_pageset, since it
is a constant value and rarely be changed.
c) Cut down instructions in __inc_numa_state() due to the incremental
numa counter and the consistant numa threshold.
d) Move zone_numa_state_snapshot() to an individual patch, since it
does not appear to be related to this patch.
Signed-off-by: Kemi Wang <kemi.wang@intel.com>
Suggested-by: Dave Hansen <dave.hansen@intel.com>
Suggested-by: Ying Huang <ying.huang@intel.com>
---
include/linux/mmzone.h | 3 +--
mm/vmstat.c | 28 ++++++++++------------------
2 files changed, 11 insertions(+), 20 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 582f6d9..c386ec4 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -282,8 +282,7 @@ struct per_cpu_pageset {
struct per_cpu_pages pcp;
#ifdef CONFIG_NUMA
s8 expire;
- s8 numa_stat_threshold;
- s8 vm_numa_stat_diff[NR_VM_NUMA_STAT_ITEMS];
+ u16 vm_numa_stat_diff[NR_VM_NUMA_STAT_ITEMS];
#endif
#ifdef CONFIG_SMP
s8 stat_threshold;
diff --git a/mm/vmstat.c b/mm/vmstat.c
index 0c3b54b..b015f39 100644
--- a/mm/vmstat.c
+++ b/mm/vmstat.c
@@ -30,6 +30,8 @@
#include "internal.h"
+#define NUMA_STATS_THRESHOLD (U16_MAX - 2)
+
#ifdef CONFIG_VM_EVENT_COUNTERS
DEFINE_PER_CPU(struct vm_event_state, vm_event_states) = {{0}};
EXPORT_PER_CPU_SYMBOL(vm_event_states);
@@ -194,10 +196,7 @@ void refresh_zone_stat_thresholds(void)
per_cpu_ptr(zone->pageset, cpu)->stat_threshold
= threshold;
-#ifdef CONFIG_NUMA
- per_cpu_ptr(zone->pageset, cpu)->numa_stat_threshold
- = threshold;
-#endif
+
/* Base nodestat threshold on the largest populated zone. */
pgdat_threshold = per_cpu_ptr(pgdat->per_cpu_nodestats, cpu)->stat_threshold;
per_cpu_ptr(pgdat->per_cpu_nodestats, cpu)->stat_threshold
@@ -231,14 +230,9 @@ void set_pgdat_percpu_threshold(pg_data_t *pgdat,
continue;
threshold = (*calculate_pressure)(zone);
- for_each_online_cpu(cpu) {
+ for_each_online_cpu(cpu)
per_cpu_ptr(zone->pageset, cpu)->stat_threshold
= threshold;
-#ifdef CONFIG_NUMA
- per_cpu_ptr(zone->pageset, cpu)->numa_stat_threshold
- = threshold;
-#endif
- }
}
}
@@ -872,16 +866,14 @@ void __inc_numa_state(struct zone *zone,
enum numa_stat_item item)
{
struct per_cpu_pageset __percpu *pcp = zone->pageset;
- s8 __percpu *p = pcp->vm_numa_stat_diff + item;
- s8 v, t;
+ u16 __percpu *p = pcp->vm_numa_stat_diff + item;
+ u16 v;
v = __this_cpu_inc_return(*p);
- t = __this_cpu_read(pcp->numa_stat_threshold);
- if (unlikely(v > t)) {
- s8 overstep = t >> 1;
- zone_numa_state_add(v + overstep, zone, item);
- __this_cpu_write(*p, -overstep);
+ if (unlikely(v > NUMA_STATS_THRESHOLD)) {
+ zone_numa_state_add(v, zone, item);
+ __this_cpu_write(*p, 0);
}
}
@@ -1796,7 +1788,7 @@ static bool need_update(int cpu)
BUILD_BUG_ON(sizeof(p->vm_stat_diff[0]) != 1);
#ifdef CONFIG_NUMA
- BUILD_BUG_ON(sizeof(p->vm_numa_stat_diff[0]) != 1);
+ BUILD_BUG_ON(sizeof(p->vm_numa_stat_diff[0]) != 2);
#endif
/*
* The fast way of checking if there are any vmstat diffs.
--
2.7.4
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-08-25 10:10 +0200 |
| Message-ID | <uihQK-1wL-19@gated-at.bofh.it> |
| In reply to | #1719107 |
On Thu, Aug 24, 2017 at 05:59:58PM +0800, Kemi Wang wrote: > Each page allocation updates a set of per-zone statistics with a call to > zone_statistics(). As discussed in 2017 MM summit, these are a substantial > source of overhead in the page allocator and are very rarely consumed. This > significant overhead in cache bouncing caused by zone counters (NUMA > associated counters) update in parallel in multi-threaded page allocation > (pointed out by Dave Hansen). > For the series; Acked-by: Mel Gorman <mgorman@techsingularity.net> -- Mel Gorman SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web