Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1728354 > unrolled thread
| Started by | kan.liang@intel.com |
|---|---|
| First post | 2017-09-07 20:00 +0200 |
| Last post | 2017-09-07 20:00 +0200 |
| Articles | 2 — 1 participant |
Back to article view | Back to linux.kernel
[PATCH RFC 00/10] perf top optimization kan.liang@intel.com - 2017-09-07 20:00 +0200
[PATCH RFC 09/10] perf top: add option to set the number of thread for event synthesize kan.liang@intel.com - 2017-09-07 20:00 +0200
| From | kan.liang@intel.com |
|---|---|
| Date | 2017-09-07 20:00 +0200 |
| Subject | [PATCH RFC 00/10] perf top optimization |
| Message-ID | <un9fQ-3ll-15@gated-at.bofh.it> |
From: Kan Liang <kan.liang@intel.com>
The patch series intends to fix the severe performance issue in
Knights Landing/Mill, when monitoring in heavy load system.
perf top costs a few minutes to show the result, which is
unacceptable.
With the patch series applied, the latency will reduces to
several seconds.
machine__synthesize_threads and perf_top__mmap_read costs most of
the perf top time (> 99%).
Patch 1-9 do the optimization for machine__synthesize_threads.
Patch 10 does the optimization for perf_top__mmap_read.
Optimization for machine__synthesize_threads
- Multithreading the whole process.
- The threads number is set to the max online CPU# by default.
User can change the threads number through the new option.
- Introduces hashtable for machine threads to reduce the lock
contention.
- The optimization can also benefit other platforms and other
perf tools, like perf record. But this patch series doesn't
do the optimization for other tools. It can be done later
separately.
- With this optimization applied, there is a 1.56x speedup in
Knights Mill with heavy workload.
Optimization for perf_top__mmap_read
- switch back to overwrite mode
For non overwrite mode, it tries to read everything in the ring buffer
and does not check the messup. Once there are lots of samples delivered
shortly, the processing time could be very long.
Considering the real time requirement for perf top, it should switch
back to overwrite mode.
- With this optimization applied, there is a huge 53.8x speedup in
Knights Mill with heavy workload.
- With this optimization applied, the latency of perf_top__mmap_read is
less than the default perf top fresh time (2s) in Knights Mill with
heavy workload.
Here are the perf top latency test result on Knights Mill and Skylake server
The heavy workload is to compile Linux kernel as below
"sudo nice make -j$(grep -c '^processor' /proc/cpuinfo)"
Then, "sudo perf top"
The latency period is the time between perf top launched and the first
profiling result shown.
- Latency on Knights Mill (272 CPUs)
Original(s) With patch(s) Speedup
272.68 16.48 16.54x
- Latency on Skylake server (192 CPUs)
Original(s) With patch(s) Speedup
12.28 2.96 4.15x
Kan Liang (10):
perf tools: hashtable for machine threads
perf tools: using scandir to replace readdir
petf tools: using comm_str to replace comm in hist_entry
petf tools: introduce a new function to set namespaces id
perf tools: lock to protect thread list
perf tools: lock to protect comm_str rb tree
perf tools: change machine comm_exec type to atomic
perf top: implement multithreading for perf_event__synthesize_threads
perf top: add option to set the number of thread for event synthesize
perf top: switch back to overwrite mode
tools/perf/builtin-kvm.c | 3 +-
tools/perf/builtin-record.c | 2 +-
tools/perf/builtin-top.c | 9 +-
tools/perf/builtin-trace.c | 21 +++--
tools/perf/tests/mmap-thread-lookup.c | 2 +-
tools/perf/ui/browsers/hists.c | 2 +-
tools/perf/util/comm.c | 18 +++-
tools/perf/util/event.c | 142 ++++++++++++++++++++++++------
tools/perf/util/event.h | 14 ++-
tools/perf/util/evlist.c | 5 +-
tools/perf/util/hist.c | 10 +--
tools/perf/util/machine.c | 158 +++++++++++++++++++++-------------
tools/perf/util/machine.h | 34 ++++++--
tools/perf/util/rb_resort.h | 5 +-
tools/perf/util/sort.c | 8 +-
tools/perf/util/sort.h | 2 +-
tools/perf/util/thread.c | 68 ++++++++++++---
tools/perf/util/thread.h | 6 +-
tools/perf/util/top.h | 1 +
19 files changed, 368 insertions(+), 142 deletions(-)
--
2.5.5
[toc] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2017-09-07 20:00 +0200 |
| Subject | [PATCH RFC 09/10] perf top: add option to set the number of thread for event synthesize |
| Message-ID | <un9fS-3ll-49@gated-at.bofh.it> |
| In reply to | #1728354 |
From: Kan Liang <kan.liang@intel.com>
Using UINT_MAX to indicate the default thread#, which is the max number
of online CPU.
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/builtin-top.c | 5 ++++-
tools/perf/util/event.c | 5 ++++-
tools/perf/util/top.h | 1 +
3 files changed, 9 insertions(+), 2 deletions(-)
diff --git a/tools/perf/builtin-top.c b/tools/perf/builtin-top.c
index 4b8fdc1..5f59aa7 100644
--- a/tools/perf/builtin-top.c
+++ b/tools/perf/builtin-top.c
@@ -961,7 +961,7 @@ static int __cmd_top(struct perf_top *top)
machine__synthesize_threads(&top->session->machines.host, &opts->target,
top->evlist->threads, false,
opts->proc_map_timeout,
- (unsigned int)sysconf(_SC_NPROCESSORS_ONLN));
+ top->nr_threads_synthesize);
if (perf_hpp_list.socket) {
ret = perf_env__read_cpu_topology_map(&perf_env);
@@ -1114,6 +1114,7 @@ int cmd_top(int argc, const char **argv)
},
.max_stack = sysctl_perf_event_max_stack,
.sym_pcnt_filter = 5,
+ .nr_threads_synthesize = UINT_MAX,
};
struct record_opts *opts = &top.record_opts;
struct target *target = &opts->target;
@@ -1223,6 +1224,8 @@ int cmd_top(int argc, const char **argv)
OPT_BOOLEAN(0, "hierarchy", &symbol_conf.report_hierarchy,
"Show entries in a hierarchy"),
OPT_BOOLEAN(0, "force", &symbol_conf.force, "don't complain, do it"),
+ OPT_UINTEGER(0, "num-thread-synthesize", &top.nr_threads_synthesize,
+ "number of thread to run event synthesize"),
OPT_END()
};
const char * const top_usage[] = {
diff --git a/tools/perf/util/event.c b/tools/perf/util/event.c
index 4f565ff..2104603 100644
--- a/tools/perf/util/event.c
+++ b/tools/perf/util/event.c
@@ -777,7 +777,10 @@ int perf_event__synthesize_threads(struct perf_tool *tool,
if (n < 0)
return err;
- thread_nr = nr_threads_synthesize;
+ if (nr_threads_synthesize == UINT_MAX)
+ thread_nr = sysconf(_SC_NPROCESSORS_ONLN);
+ else
+ thread_nr = nr_threads_synthesize;
if (thread_nr <= 0)
thread_nr = 1;
if (thread_nr > n)
diff --git a/tools/perf/util/top.h b/tools/perf/util/top.h
index 9bdfb78..f4296e1 100644
--- a/tools/perf/util/top.h
+++ b/tools/perf/util/top.h
@@ -37,6 +37,7 @@ struct perf_top {
int sym_pcnt_filter;
const char *sym_filter;
float min_percent;
+ unsigned int nr_threads_synthesize;
};
#define CONSOLE_CLEAR "[H[2J"
--
2.5.5
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web