Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1260609 > unrolled thread
| Started by | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| First post | 2015-11-02 14:00 +0100 |
| Last post | 2015-11-02 23:50 +0100 |
| Articles | 15 — 3 participants |
Back to article view | Back to linux.kernel
[RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-02 14:00 +0100
[RFC/PATCH v2 4/4] perf report: Add callchain value option Namhyung Kim <namhyung@kernel.org> - 2015-11-02 14:00 +0100
[RFC/PATCH v2 3/4] perf callchain: Add count fields to struct callchain_node Namhyung Kim <namhyung@kernel.org> - 2015-11-02 14:00 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Brendan Gregg <brendan.d.gregg@gmail.com> - 2015-11-02 21:40 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Arnaldo Carvalho de Melo <acme@kernel.org> - 2015-11-02 22:40 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-02 23:20 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Arnaldo Carvalho de Melo <acme@kernel.org> - 2015-11-02 23:30 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-02 23:50 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Arnaldo Carvalho de Melo <acme@kernel.org> - 2015-11-03 00:10 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-03 00:50 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Arnaldo Carvalho de Melo <acme@kernel.org> - 2015-11-03 01:50 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-03 02:40 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Arnaldo Carvalho de Melo <acme@kernel.org> - 2015-11-03 02:50 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Namhyung Kim <namhyung@kernel.org> - 2015-11-03 04:20 +0100
Re: [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) Brendan Gregg <brendan.d.gregg@gmail.com> - 2015-11-02 23:50 +0100
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-02 14:00 +0100 |
| Subject | [RFC/PATCH 0/4] perf report: Support folded callchain output (v2) |
| Message-ID | <qqmSl-5qi-5@gated-at.bofh.it> |
Hello,
This is what Brendan requested on the perf-users mailing list [1] to
support FlameGraphs [2] more efficiently. This patchset adds a few
more callchain options to adjust the output for it.
At first, 'folded' output mode was added. The folded output puts all
calchain nodes in a line separated by semicolons, a space and the
value. Now it only supports --stdio as other UI provides some way of
folding/expanding callchains dynamically.
The value is now can be one of 'percent', 'period', or 'count'. The
percent is current default output and the period is the raw number of
sample periods. The count is the number of samples for each callchain.
Here's an example:
$ perf report --no-children --show-nr-samples --stdio -g folded,count
...
39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
$ perf report --no-children --stdio -g percent
...
39.93% swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--28.63%-- start_secondary
|
--11.30%-- rest_init
$ perf report --no-children --stdio --show-total-period -g period
...
39.93% 13018705 swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--9334403-- start_secondary
|
--3684302-- rest_init
$ perf report --no-children --stdio --show-nr-samples -g count
...
39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--57-- start_secondary
|
--23-- rest_init
You can get it from 'perf/callchain-fold-v2' branch on my tree:
git://git.kernel.org/pub/scm/linux/kernel/git/namhyung/linux-perf.git
Any comments are welcome, thanks
Namhyung
[1] http://www.spinics.net/lists/linux-perf-users/msg02498.html
[2] http://www.brendangregg.com/FlameGraphs/cpuflamegraphs.html
Namhyung Kim (4):
perf report: Support folded callchain mode on --stdio
perf callchain: Abstract callchain print function
perf callchain: Add count fields to struct callchain_node
perf report: Add callchain value option
tools/perf/Documentation/perf-report.txt | 13 +++--
tools/perf/builtin-report.c | 4 +-
tools/perf/ui/browsers/hists.c | 8 +--
tools/perf/ui/gtk/hists.c | 8 +--
tools/perf/ui/stdio/hist.c | 91 ++++++++++++++++++++++++++------
tools/perf/util/callchain.c | 87 +++++++++++++++++++++++++++++-
tools/perf/util/callchain.h | 24 ++++++++-
tools/perf/util/util.c | 3 +-
8 files changed, 204 insertions(+), 34 deletions(-)
--
2.6.2
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-02 14:00 +0100 |
| Subject | [RFC/PATCH v2 4/4] perf report: Add callchain value option |
| Message-ID | <qqmSm-5qi-27@gated-at.bofh.it> |
| In reply to | #1260609 |
Now -g/--call-graph option supports how to display callchain values.
Possible values are 'percent', 'period' and 'count'. The percent is
same as before and it's the default behavior. The period displays the
raw period value rather than the percentage. The count displays the
number of occurrences.
$ perf report --no-children --stdio -g percent
...
39.93% swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--28.63%-- start_secondary
|
--11.30%-- rest_init
$ perf report --no-children --show-total-period --stdio -g period
...
39.93% 13018705 swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--9334403-- start_secondary
|
--3684302-- rest_init
$ perf report --no-children --show-nr-samples --stdio -g count
...
39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
|
---intel_idle
cpuidle_enter_state
cpuidle_enter
call_cpuidle
cpu_startup_entry
|
|--57-- start_secondary
|
--23-- rest_init
Cc: Brendan Gregg <brendan.d.gregg@gmail.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
---
tools/perf/Documentation/perf-report.txt | 13 ++++---
tools/perf/builtin-report.c | 4 +--
tools/perf/ui/stdio/hist.c | 8 +++++
tools/perf/util/callchain.c | 60 +++++++++++++++++++++++++++-----
tools/perf/util/callchain.h | 10 +++++-
tools/perf/util/util.c | 3 +-
6 files changed, 82 insertions(+), 16 deletions(-)
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 5ce8da1e1256..bb9fd23a105e 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -170,11 +170,11 @@ OPTIONS
Dump raw trace in ASCII.
-g::
---call-graph=<print_type,threshold[,print_limit],order,sort_key,branch>::
+--call-graph=<print_type,threshold[,print_limit],order,sort_key[,branch],value>::
Display call chains using type, min percent threshold, print limit,
- call order, sort key and branch. Note that ordering of parameters is not
- fixed so any parement can be given in an arbitraty order. One exception
- is the print_limit which should be preceded by threshold.
+ call order, sort key, optional branch and value. Note that ordering of
+ parameters is not fixed so any parement can be given in an arbitraty order.
+ One exception is the print_limit which should be preceded by threshold.
print_type can be either:
- flat: single column, linear exposure of call chains.
@@ -204,6 +204,11 @@ OPTIONS
- branch: include last branch information in callgraph when available.
Usually more convenient to use --branch-history for this.
+ value can be:
+ - percent: diplay overhead percent (default)
+ - period: display event period
+ - count: display evnt count
+
--children::
Accumulate callchain of children to parent entry so that then can
show up in the output. The output will have a new "Children" column
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 2853ad2bd435..3dd4bb4ded1a 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -625,7 +625,7 @@ parse_percent_limit(const struct option *opt, const char *str,
return 0;
}
-#define CALLCHAIN_DEFAULT_OPT "graph,0.5,caller,function"
+#define CALLCHAIN_DEFAULT_OPT "graph,0.5,caller,function,percent"
const char report_callchain_help[] = "Display call graph (stack chain/backtrace):\n\n"
CALLCHAIN_REPORT_HELP
@@ -708,7 +708,7 @@ int cmd_report(int argc, const char **argv, const char *prefix __maybe_unused)
OPT_BOOLEAN('x', "exclude-other", &symbol_conf.exclude_other,
"Only display entries with parent-match"),
OPT_CALLBACK_DEFAULT('g', "call-graph", &report,
- "print_type,threshold[,print_limit],order,sort_key[,branch]",
+ "print_type,threshold[,print_limit],order,sort_key[,branch],value",
report_callchain_help, &report_parse_callchain_opt,
callchain_default_opt),
OPT_BOOLEAN(0, "children", &symbol_conf.cumulate_callchain,
diff --git a/tools/perf/ui/stdio/hist.c b/tools/perf/ui/stdio/hist.c
index e84ca21252d3..2104b09d41a8 100644
--- a/tools/perf/ui/stdio/hist.c
+++ b/tools/perf/ui/stdio/hist.c
@@ -88,6 +88,7 @@ static size_t __callchain__fprintf_graph(FILE *fp, struct rb_root *root,
size_t ret = 0;
int i;
uint entries_printed = 0;
+ int cumul_count = 0;
remaining = total_samples;
@@ -99,6 +100,7 @@ static size_t __callchain__fprintf_graph(FILE *fp, struct rb_root *root,
child = rb_entry(node, struct callchain_node, rb_node);
cumul = callchain_cumul_hits(child);
remaining -= cumul;
+ cumul_count += callchain_cumul_counts(child);
/*
* The depth mask manages the output of pipes that show
@@ -148,6 +150,12 @@ static size_t __callchain__fprintf_graph(FILE *fp, struct rb_root *root,
if (!rem_sq_bracket)
return ret;
+ if (callchain_param.value == CCVAL_COUNT) {
+ rem_node.count = child->parent->children_count - cumul_count;
+ if (rem_node.count <= 0)
+ return ret;
+ }
+
new_depth_mask &= ~(1 << (depth - 1));
ret += ipchain__fprintf_graph(fp, &rem_node, &rem_hits, depth,
new_depth_mask, 0, total_samples,
diff --git a/tools/perf/util/callchain.c b/tools/perf/util/callchain.c
index 0a97d77509bd..7f0a89584f1b 100644
--- a/tools/perf/util/callchain.c
+++ b/tools/perf/util/callchain.c
@@ -83,6 +83,23 @@ static int parse_callchain_sort_key(const char *value)
return -1;
}
+static int parse_callchain_value(const char *value)
+{
+ if (!strncmp(value, "percent", strlen(value))) {
+ callchain_param.value = CCVAL_PERCENT;
+ return 0;
+ }
+ if (!strncmp(value, "period", strlen(value))) {
+ callchain_param.value = CCVAL_PERIOD;
+ return 0;
+ }
+ if (!strncmp(value, "count", strlen(value))) {
+ callchain_param.value = CCVAL_COUNT;
+ return 0;
+ }
+ return -1;
+}
+
static int
__parse_callchain_report_opt(const char *arg, bool allow_record_opt)
{
@@ -106,7 +123,8 @@ __parse_callchain_report_opt(const char *arg, bool allow_record_opt)
if (!parse_callchain_mode(tok) ||
!parse_callchain_order(tok) ||
- !parse_callchain_sort_key(tok)) {
+ !parse_callchain_sort_key(tok) ||
+ !parse_callchain_value(tok)) {
/* parsing ok - move on to the next */
try_stack_size = false;
goto next;
@@ -819,12 +837,26 @@ char *callchain_node__sprintf_value(struct callchain_node *node,
char *bf, size_t bfsize, u64 total)
{
double percent = 0.0;
- u64 cumul = callchain_cumul_hits(node);
+ u64 period = callchain_cumul_hits(node);
+ int count = callchain_cumul_counts(node);
if (total)
- percent = cumul * 100.0 / total;
+ percent = period * 100.0 / total;
+ if (callchain_param.mode == CHAIN_FOLDED)
+ count = node->count;
- scnprintf(bf, bfsize, "%6.2f%%", percent);
+ switch (callchain_param.value) {
+ case CCVAL_PERIOD:
+ scnprintf(bf, bfsize, "%"PRIu64, period);
+ break;
+ case CCVAL_COUNT:
+ scnprintf(bf, bfsize, "%u", count);
+ break;
+ case CCVAL_PERCENT:
+ default:
+ scnprintf(bf, bfsize, "%.2f%%", percent);
+ break;
+ }
return bf;
}
@@ -832,12 +864,24 @@ int callchain_node__fprintf_value(struct callchain_node *node,
FILE *fp, u64 total)
{
double percent = 0.0;
- u64 cumul = callchain_cumul_hits(node);
+ u64 period = callchain_cumul_hits(node);
+ int count = callchain_cumul_counts(node);
if (total)
- percent = cumul * 100.0 / total;
-
- return percent_color_fprintf(fp, "%.2f%%", percent);
+ percent = period * 100.0 / total;
+ if (callchain_param.mode == CHAIN_FOLDED)
+ count = node->count;
+
+ switch (callchain_param.value) {
+ case CCVAL_PERIOD:
+ return fprintf(fp, "%"PRIu64, period);
+ case CCVAL_COUNT:
+ return fprintf(fp, "%u", count);
+ case CCVAL_PERCENT:
+ default:
+ return percent_color_fprintf(fp, "%.2f%%", percent);
+ }
+ return 0;
}
static void free_callchain_node(struct callchain_node *node)
diff --git a/tools/perf/util/callchain.h b/tools/perf/util/callchain.h
index 2f948f0ff034..e8533e328a47 100644
--- a/tools/perf/util/callchain.h
+++ b/tools/perf/util/callchain.h
@@ -29,7 +29,8 @@
HELP_PAD "print_limit:\tmaximum number of call graph entry (<number>)\n" \
HELP_PAD "order:\t\tcall graph order (caller|callee)\n" \
HELP_PAD "sort_key:\tcall graph sort key (function|address)\n" \
- HELP_PAD "branch:\t\tinclude last branch info to call graph (branch)\n"
+ HELP_PAD "branch:\t\tinclude last branch info to call graph (branch)\n" \
+ HELP_PAD "value:\t\tcall graph value (percent|period|count)\n"
enum perf_call_graph_mode {
CALLCHAIN_NONE,
@@ -81,6 +82,12 @@ enum chain_key {
CCKEY_ADDRESS
};
+enum chain_value {
+ CCVAL_PERCENT,
+ CCVAL_PERIOD,
+ CCVAL_COUNT,
+};
+
struct callchain_param {
bool enabled;
enum perf_call_graph_mode record_mode;
@@ -93,6 +100,7 @@ struct callchain_param {
bool order_set;
enum chain_key key;
bool branch_callstack;
+ enum chain_value value;
};
extern struct callchain_param callchain_param;
diff --git a/tools/perf/util/util.c b/tools/perf/util/util.c
index cd12c25e4ea4..174912f87913 100644
--- a/tools/perf/util/util.c
+++ b/tools/perf/util/util.c
@@ -20,7 +20,8 @@ struct callchain_param callchain_param = {
.mode = CHAIN_GRAPH_ABS,
.min_percent = 0.5,
.order = ORDER_CALLEE,
- .key = CCKEY_FUNCTION
+ .key = CCKEY_FUNCTION,
+ .value = CCVAL_PERCENT,
};
/*
--
2.6.2
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-02 14:00 +0100 |
| Subject | [RFC/PATCH v2 3/4] perf callchain: Add count fields to struct callchain_node |
| Message-ID | <qqmSm-5qi-35@gated-at.bofh.it> |
| In reply to | #1260609 |
It's to track the count of occurrences of the callchains.
Cc: Brendan Gregg <brendan.d.gregg@gmail.com>
Signed-off-by: Namhyung Kim <namhyung@kernel.org>
---
tools/perf/util/callchain.c | 10 ++++++++++
tools/perf/util/callchain.h | 7 +++++++
2 files changed, 17 insertions(+)
diff --git a/tools/perf/util/callchain.c b/tools/perf/util/callchain.c
index 44184d198855..0a97d77509bd 100644
--- a/tools/perf/util/callchain.c
+++ b/tools/perf/util/callchain.c
@@ -437,6 +437,8 @@ add_child(struct callchain_node *parent,
new->children_hit = 0;
new->hit = period;
+ new->children_count = 0;
+ new->count = 1;
return new;
}
@@ -484,6 +486,9 @@ split_add_child(struct callchain_node *parent,
parent->children_hit = callchain_cumul_hits(new);
new->val_nr = parent->val_nr - idx_local;
parent->val_nr = idx_local;
+ new->count = parent->count;
+ new->children_count = parent->children_count;
+ parent->children_count = callchain_cumul_counts(new);
/* create a new child for the new branch if any */
if (idx_total < cursor->nr) {
@@ -494,6 +499,8 @@ split_add_child(struct callchain_node *parent,
parent->hit = 0;
parent->children_hit += period;
+ parent->count = 0;
+ parent->children_count += 1;
node = callchain_cursor_current(cursor);
new = add_child(parent, cursor, period);
@@ -516,6 +523,7 @@ split_add_child(struct callchain_node *parent,
rb_insert_color(&new->rb_node_in, &parent->rb_root_in);
} else {
parent->hit = period;
+ parent->count = 1;
}
}
@@ -562,6 +570,7 @@ append_chain_children(struct callchain_node *root,
inc_children_hit:
root->children_hit += period;
+ root->children_count++;
}
static int
@@ -614,6 +623,7 @@ append_chain(struct callchain_node *root,
/* we match 100% of the path, increment the hit */
if (matches == root->val_nr && cursor->pos == cursor->nr) {
root->hit += period;
+ root->count++;
return 0;
}
diff --git a/tools/perf/util/callchain.h b/tools/perf/util/callchain.h
index 3a90a57f6213..2f948f0ff034 100644
--- a/tools/perf/util/callchain.h
+++ b/tools/perf/util/callchain.h
@@ -60,6 +60,8 @@ struct callchain_node {
struct rb_root rb_root_in; /* input tree of children */
struct rb_root rb_root; /* sorted output tree of children */
unsigned int val_nr;
+ unsigned int count;
+ unsigned int children_count;
u64 hit;
u64 children_hit;
};
@@ -145,6 +147,11 @@ static inline u64 callchain_cumul_hits(struct callchain_node *node)
return node->hit + node->children_hit;
}
+static inline int callchain_cumul_counts(struct callchain_node *node)
+{
+ return node->count + node->children_count;
+}
+
int callchain_register_param(struct callchain_param *param);
int callchain_append(struct callchain_root *root,
struct callchain_cursor *cursor,
--
2.6.2
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Brendan Gregg <brendan.d.gregg@gmail.com> |
|---|---|
| Date | 2015-11-02 21:40 +0100 |
| Message-ID | <qqu3w-1uH-13@gated-at.bofh.it> |
| In reply to | #1260609 |
G'Day Namhyung, On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: > Hello, > > This is what Brendan requested on the perf-users mailing list [1] to > support FlameGraphs [2] more efficiently. This patchset adds a few > more callchain options to adjust the output for it. > > At first, 'folded' output mode was added. The folded output puts all > calchain nodes in a line separated by semicolons, a space and the > value. Now it only supports --stdio as other UI provides some way of > folding/expanding callchains dynamically. > > The value is now can be one of 'percent', 'period', or 'count'. The > percent is current default output and the period is the raw number of > sample periods. The count is the number of samples for each callchain. > > Here's an example: > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > ... > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 Thanks! So for the folded output I don't need the summary line (the row of columns printed by hist_entry__snprintf()), and don't need anything except folded stacks and the counts. If working with the existing stdio interface is making it harder than it needs to be, might it be easier to make it a separate interface (ui/folded), that just emitted the folded output? Just an idea. This existing patchset is working for me, I'd just be filtering the output. Having the option for percentages and periods is nice. I can envisage using periods (for latency flame graphs). Brendan -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2015-11-02 22:40 +0100 |
| Message-ID | <qquZB-250-27@gated-at.bofh.it> |
| In reply to | #1260950 |
Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu: > G'Day Namhyung, > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: > > Hello, > > > > This is what Brendan requested on the perf-users mailing list [1] to > > support FlameGraphs [2] more efficiently. This patchset adds a few > > more callchain options to adjust the output for it. > > > > At first, 'folded' output mode was added. The folded output puts all > > calchain nodes in a line separated by semicolons, a space and the > > value. Now it only supports --stdio as other UI provides some way of > > folding/expanding callchains dynamically. > > > > The value is now can be one of 'percent', 'period', or 'count'. The > > percent is current default output and the period is the raw number of > > sample periods. The count is the number of samples for each callchain. > > > > Here's an example: > > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > > ... > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > Thanks! > > So for the folded output I don't need the summary line (the row of > columns printed by hist_entry__snprintf()), and don't need anything > except folded stacks and the counts. If working with the existing > stdio interface is making it harder than it needs to be, might it be I don't think it so, just add some flag asking for that hist_entry__snprintf() to be supressed, ideas for a long option name? Having it as Namhyung did may have value for some people as a more compact way to show the callchains together with the hist_entry line. With this in mind, do you have any other issues with Namhyung's patchkit? An acked-by/tested-by you would be nice to have, and then we could work out the new option to suppress that hist_entry__snprintf() in a follow up patch. > easier to make it a separate interface (ui/folded), that just emitted > the folded output? Just an idea. This existing patchset is working for > me, I'd just be filtering the output. > > Having the option for percentages and periods is nice. I can envisage > using periods (for latency flame graphs). You mean in the callchain lines? - Arnaldo -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-02 23:20 +0100 |
| Message-ID | <qqvCi-2xp-19@gated-at.bofh.it> |
| In reply to | #1260997 |
Hi Arnaldo,
On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote:
> Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu:
> > G'Day Namhyung,
> >
> > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote:
> > > Hello,
> > >
> > > This is what Brendan requested on the perf-users mailing list [1] to
> > > support FlameGraphs [2] more efficiently. This patchset adds a few
> > > more callchain options to adjust the output for it.
> > >
> > > At first, 'folded' output mode was added. The folded output puts all
> > > calchain nodes in a line separated by semicolons, a space and the
> > > value. Now it only supports --stdio as other UI provides some way of
> > > folding/expanding callchains dynamically.
> > >
> > > The value is now can be one of 'percent', 'period', or 'count'. The
> > > percent is current default output and the period is the raw number of
> > > sample periods. The count is the number of samples for each callchain.
> > >
> > > Here's an example:
> > >
> > > $ perf report --no-children --show-nr-samples --stdio -g folded,count
> > > ...
> > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
> > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
> > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
> >
> > Thanks!
> >
> > So for the folded output I don't need the summary line (the row of
> > columns printed by hist_entry__snprintf()), and don't need anything
> > except folded stacks and the counts. If working with the existing
> > stdio interface is making it harder than it needs to be, might it be
>
> I don't think it so, just add some flag asking for that
> hist_entry__snprintf() to be supressed, ideas for a long option name?
>
> Having it as Namhyung did may have value for some people as a more
> compact way to show the callchains together with the hist_entry line.
Yeah, I'd keep the hist entry line unless it's too hard to
parse/filter. IMHO it's just a way to show callchains, so no need to
have separate output mode..
Brendan, I guess you still need to know other info like cpu or pid, no?
And I feel like it'd be better to put the count before the callchains
for consistency like below. Is it OK to you?
$ perf report --no-children --show-nr-samples --stdio -g folded,count
...
39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
57 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary
23 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;...
>
> With this in mind, do you have any other issues with Namhyung's
> patchkit? An acked-by/tested-by you would be nice to have, and then we
> could work out the new option to suppress that hist_entry__snprintf()
> in a follow up patch.
>
> > easier to make it a separate interface (ui/folded), that just emitted
> > the folded output? Just an idea. This existing patchset is working for
> > me, I'd just be filtering the output.
> >
> > Having the option for percentages and periods is nice. I can envisage
> > using periods (for latency flame graphs).
Glad to see you like it. :)
Thanks,
Namhyung
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2015-11-02 23:30 +0100 |
| Message-ID | <qqvLX-2Au-7@gated-at.bofh.it> |
| In reply to | #1261022 |
Hi Namhyung,
Em Tue, Nov 03, 2015 at 07:12:04AM +0900, Namhyung Kim escreveu:
> On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote:
> > Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu:
> > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote:
> > > > This is what Brendan requested on the perf-users mailing list [1] to
> > > > support FlameGraphs [2] more efficiently. This patchset adds a few
> > > > more callchain options to adjust the output for it.
> > > > At first, 'folded' output mode was added. The folded output puts all
> > > > calchain nodes in a line separated by semicolons, a space and the
> > > > value. Now it only supports --stdio as other UI provides some way of
> > > > folding/expanding callchains dynamically.
> > > > The value is now can be one of 'percent', 'period', or 'count'. The
> > > > percent is current default output and the period is the raw number of
> > > > sample periods. The count is the number of samples for each callchain.
> > > > Here's an example:
> > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count
> > > > ...
> > > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
> > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
> > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
> > > So for the folded output I don't need the summary line (the row of
> > > columns printed by hist_entry__snprintf()), and don't need anything
> > > except folded stacks and the counts. If working with the existing
> > > stdio interface is making it harder than it needs to be, might it be
> > I don't think it so, just add some flag asking for that
> > hist_entry__snprintf() to be supressed, ideas for a long option name?
> > Having it as Namhyung did may have value for some people as a more
> > compact way to show the callchains together with the hist_entry line.
> Yeah, I'd keep the hist entry line unless it's too hard to
> parse/filter. IMHO it's just a way to show callchains, so no need to
What I suggested was to have something like:
$ perf report --no-children --no-hists --stdio -g folded,count
^^^^^^^^^^
^^^^^^^^^^
...
intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
I.e. the first entry in the callchain is 'intel_idle', just like in what
Brendan called the 'summary line', i.e. reduntant when what he wants its
just all the callchains and how many times they were sampled.
> have separate output mode..
> Brendan, I guess you still need to know other info like cpu or pid, no?
Possibly, but just with the callchains he has enough info for the basic
flame graph, no?
> And I feel like it'd be better to put the count before the callchains
> for consistency like below. Is it OK to you?
Consistency with what?
The main thing here is the callchain, all the other stuff are things
related to it, so showing it first makes sense to me.
Having some way to list the desired info to have for each callchain may
be interesting, and if he could do it like:
-g folded,count,cpu,other,fields
then he would know how to parse the per-callchain info at the end of
each line, right?
- Arnaldo
>
> $ perf report --no-children --show-nr-samples --stdio -g folded,count
> ...
> 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
> 57 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary
> 23 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;...
>
>
> >
> > With this in mind, do you have any other issues with Namhyung's
> > patchkit? An acked-by/tested-by you would be nice to have, and then we
> > could work out the new option to suppress that hist_entry__snprintf()
> > in a follow up patch.
> >
> > > easier to make it a separate interface (ui/folded), that just emitted
> > > the folded output? Just an idea. This existing patchset is working for
> > > me, I'd just be filtering the output.
> > >
> > > Having the option for percentages and periods is nice. I can envisage
> > > using periods (for latency flame graphs).
>
> Glad to see you like it. :)
>
> Thanks,
> Namhyung
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-02 23:50 +0100 |
| Message-ID | <qqw5k-2HW-1@gated-at.bofh.it> |
| In reply to | #1261027 |
On Mon, Nov 02, 2015 at 07:28:42PM -0300, Arnaldo Carvalho de Melo wrote: > Hi Namhyung, > > Em Tue, Nov 03, 2015 at 07:12:04AM +0900, Namhyung Kim escreveu: > > On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote: > > > Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu: > > > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: > > > > > This is what Brendan requested on the perf-users mailing list [1] to > > > > > support FlameGraphs [2] more efficiently. This patchset adds a few > > > > > more callchain options to adjust the output for it. > > > > > > At first, 'folded' output mode was added. The folded output puts all > > > > > calchain nodes in a line separated by semicolons, a space and the > > > > > value. Now it only supports --stdio as other UI provides some way of > > > > > folding/expanding callchains dynamically. > > > > > > The value is now can be one of 'percent', 'period', or 'count'. The > > > > > percent is current default output and the period is the raw number of > > > > > sample periods. The count is the number of samples for each callchain. > > > > > > Here's an example: > > > > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > > > > > ... > > > > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > > > > So for the folded output I don't need the summary line (the row of > > > > columns printed by hist_entry__snprintf()), and don't need anything > > > > except folded stacks and the counts. If working with the existing > > > > stdio interface is making it harder than it needs to be, might it be > > > > I don't think it so, just add some flag asking for that > > > hist_entry__snprintf() to be supressed, ideas for a long option name? > > > > Having it as Namhyung did may have value for some people as a more > > > compact way to show the callchains together with the hist_entry line. > > > Yeah, I'd keep the hist entry line unless it's too hard to > > parse/filter. IMHO it's just a way to show callchains, so no need to > > What I suggested was to have something like: > > $ perf report --no-children --no-hists --stdio -g folded,count > ^^^^^^^^^^ > ^^^^^^^^^^ > ... > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > I.e. the first entry in the callchain is 'intel_idle', just like in what > Brendan called the 'summary line', i.e. reduntant when what he wants its > just all the callchains and how many times they were sampled. Yep, I know. But isn't 'perf report' all for seeing hist lines? :) I'm not insisting it strongly, but it's a bit strange for me if perf report doesn't show any hist lines.. > > > have separate output mode.. > > > Brendan, I guess you still need to know other info like cpu or pid, no? > > Possibly, but just with the callchains he has enough info for the basic > flame graph, no? > > > And I feel like it'd be better to put the count before the callchains > > for consistency like below. Is it OK to you? > > Consistency with what? Oh, I meant consistency with other callchain output style like graph, fractal or flat - They all show the numbers before callchains. And I think it's easier to read for human. :) > > The main thing here is the callchain, all the other stuff are things > related to it, so showing it first makes sense to me. > > Having some way to list the desired info to have for each callchain may > be interesting, and if he could do it like: > > -g folded,count,cpu,other,fields > > then he would know how to parse the per-callchain info at the end of > each line, right? Hmm.. looks like that it ends up having redundant info. I don't think it's generally useful to other 'perf report' stuffs. Wouldn't it be better just adding minimal support and let the external tool parse the output? Thanks, Namhyung -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2015-11-03 00:10 +0100 |
| Message-ID | <qqwoF-34I-9@gated-at.bofh.it> |
| In reply to | #1261029 |
Em Tue, Nov 03, 2015 at 07:49:27AM +0900, Namhyung Kim escreveu:
> On Mon, Nov 02, 2015 at 07:28:42PM -0300, Arnaldo Carvalho de Melo wrote:
> > Hi Namhyung,
> >
> > Em Tue, Nov 03, 2015 at 07:12:04AM +0900, Namhyung Kim escreveu:
> > > On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote:
> > > > Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu:
> > > > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote:
> > > > > > This is what Brendan requested on the perf-users mailing list [1] to
> > > > > > support FlameGraphs [2] more efficiently. This patchset adds a few
> > > > > > more callchain options to adjust the output for it.
> >
> > > > > > At first, 'folded' output mode was added. The folded output puts all
> > > > > > calchain nodes in a line separated by semicolons, a space and the
> > > > > > value. Now it only supports --stdio as other UI provides some way of
> > > > > > folding/expanding callchains dynamically.
> >
> > > > > > The value is now can be one of 'percent', 'period', or 'count'. The
> > > > > > percent is current default output and the period is the raw number of
> > > > > > sample periods. The count is the number of samples for each callchain.
> >
> > > > > > Here's an example:
> >
> > > > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count
> > > > > > ...
> > > > > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel
> > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
> > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
> >
> > > > > So for the folded output I don't need the summary line (the row of
> > > > > columns printed by hist_entry__snprintf()), and don't need anything
> > > > > except folded stacks and the counts. If working with the existing
> > > > > stdio interface is making it harder than it needs to be, might it be
> >
> > > > I don't think it so, just add some flag asking for that
> > > > hist_entry__snprintf() to be supressed, ideas for a long option name?
> >
> > > > Having it as Namhyung did may have value for some people as a more
> > > > compact way to show the callchains together with the hist_entry line.
> >
> > > Yeah, I'd keep the hist entry line unless it's too hard to
> > > parse/filter. IMHO it's just a way to show callchains, so no need to
> >
> > What I suggested was to have something like:
> >
> > $ perf report --no-children --no-hists --stdio -g folded,count
> > ^^^^^^^^^^
> > ^^^^^^^^^^
> > ...
> > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57
> > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23
> >
> > I.e. the first entry in the callchain is 'intel_idle', just like in what
> > Brendan called the 'summary line', i.e. reduntant when what he wants its
> > just all the callchains and how many times they were sampled.
>
> Yep, I know. But isn't 'perf report' all for seeing hist lines? :)
Well, so far, yes, but he is presenting a usecase where what we want to
see is just callchains, and we can achieve that rather easily, no?
> I'm not insisting it strongly, but it's a bit strange for me if perf
> report doesn't show any hist lines..
If that is of no use in this use case, why not?
> > > have separate output mode..
> >
> > > Brendan, I guess you still need to know other info like cpu or pid, no?
> >
> > Possibly, but just with the callchains he has enough info for the basic
> > flame graph, no?
> >
> > > And I feel like it'd be better to put the count before the callchains
> > > for consistency like below. Is it OK to you?
> >
> > Consistency with what?
>
> Oh, I meant consistency with other callchain output style like graph,
> fractal or flat - They all show the numbers before callchains. And I
> think it's easier to read for human. :)
Well, As I said, isn't the main object here the callchain? :-)
And Brendan's request is for a something to be consumed by scripts, i.e.
something like we have for perf stat:
For humans:
[root@felicio ~]# perf stat -e cycles -I 1000 -a
# time counts unit events
1.000304391 1,820,038 cycles
2.000490191 1,005,477,007 cycles
3.000657813 1,717,007 cycles
^C 3.917890293 2,804,034 cycles
For machines/scripts:
[root@felicio ~]# perf stat -x, -e cycles -I 1000 -a
1.000291954,1923360,,cycles,3998167210,100.00
2.000477154,1005608105,,cycles,3998475482,100.00
3.000612612,1345483,,cycles,3998332391,100.00
4.000744469,1005046913,,cycles,3998258199,100.00
^C 4.331684347,1551327,,cycles,3463190970,100.00
[root@felicio ~]#
> > The main thing here is the callchain, all the other stuff are things
> > related to it, so showing it first makes sense to me.
> >
> > Having some way to list the desired info to have for each callchain may
> > be interesting, and if he could do it like:
> >
> > -g folded,count,cpu,other,fields
> >
> > then he would know how to parse the per-callchain info at the end of
> > each line, right?
>
> Hmm.. looks like that it ends up having redundant info. I don't think
What is redundant, and with with what?
> it's generally useful to other 'perf report' stuffs. Wouldn't it be
> better just adding minimal support and let the external tool parse the
> output?
Oh well, perhaps we could have a 'perf callchain' tool that would be
centered on callchains and would provided one line per callchain, which
would have:
callchain;seprarated;colons series,of,desired,fields,for,this,callchain
Which would reuse heavily the 'perf report' / 'perf top' code for
histograms, no?
I still think that this is a 'perf report' thing, but one that is
centered in callchains, and that is to be consumed by scripts, not
humans.
- Arnaldo
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-03 00:50 +0100 |
| Message-ID | <qqx1n-3iZ-1@gated-at.bofh.it> |
| In reply to | #1261048 |
On Mon, Nov 02, 2015 at 08:04:36PM -0300, Arnaldo Carvalho de Melo wrote: > Em Tue, Nov 03, 2015 at 07:49:27AM +0900, Namhyung Kim escreveu: > > On Mon, Nov 02, 2015 at 07:28:42PM -0300, Arnaldo Carvalho de Melo wrote: > > > Hi Namhyung, > > > > > > Em Tue, Nov 03, 2015 at 07:12:04AM +0900, Namhyung Kim escreveu: > > > > On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote: > > > > > Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu: > > > > > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: > > > > > > > This is what Brendan requested on the perf-users mailing list [1] to > > > > > > > support FlameGraphs [2] more efficiently. This patchset adds a few > > > > > > > more callchain options to adjust the output for it. > > > > > > > > > > At first, 'folded' output mode was added. The folded output puts all > > > > > > > calchain nodes in a line separated by semicolons, a space and the > > > > > > > value. Now it only supports --stdio as other UI provides some way of > > > > > > > folding/expanding callchains dynamically. > > > > > > > > > > The value is now can be one of 'percent', 'period', or 'count'. The > > > > > > > percent is current default output and the period is the raw number of > > > > > > > sample periods. The count is the number of samples for each callchain. > > > > > > > > > > Here's an example: > > > > > > > > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > > > > > > > ... > > > > > > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > > > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > > > > > > > > So for the folded output I don't need the summary line (the row of > > > > > > columns printed by hist_entry__snprintf()), and don't need anything > > > > > > except folded stacks and the counts. If working with the existing > > > > > > stdio interface is making it harder than it needs to be, might it be > > > > > > > > I don't think it so, just add some flag asking for that > > > > > hist_entry__snprintf() to be supressed, ideas for a long option name? > > > > > > > > Having it as Namhyung did may have value for some people as a more > > > > > compact way to show the callchains together with the hist_entry line. > > > > > > > Yeah, I'd keep the hist entry line unless it's too hard to > > > > parse/filter. IMHO it's just a way to show callchains, so no need to > > > > > > What I suggested was to have something like: > > > > > > $ perf report --no-children --no-hists --stdio -g folded,count > > > ^^^^^^^^^^ > > > ^^^^^^^^^^ > > > ... > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > > > > > I.e. the first entry in the callchain is 'intel_idle', just like in what > > > Brendan called the 'summary line', i.e. reduntant when what he wants its > > > just all the callchains and how many times they were sampled. > > > > Yep, I know. But isn't 'perf report' all for seeing hist lines? :) > > Well, so far, yes, but he is presenting a usecase where what we want to > see is just callchains, and we can achieve that rather easily, no? But it's also easy to filter from the script side. > > > I'm not insisting it strongly, but it's a bit strange for me if perf > > report doesn't show any hist lines.. > > If that is of no use in this use case, why not? Well, I think FlameGraphs is a rather unusual case and folded output seems useful to other use cases too. > > > > > have separate output mode.. > > > > > > > Brendan, I guess you still need to know other info like cpu or pid, no? > > > > > > Possibly, but just with the callchains he has enough info for the basic > > > flame graph, no? > > > > > > > And I feel like it'd be better to put the count before the callchains > > > > for consistency like below. Is it OK to you? > > > > > > Consistency with what? > > > > Oh, I meant consistency with other callchain output style like graph, > > fractal or flat - They all show the numbers before callchains. And I > > think it's easier to read for human. :) > > Well, As I said, isn't the main object here the callchain? :-) > > And Brendan's request is for a something to be consumed by scripts, i.e. > something like we have for perf stat: > > For humans: > > [root@felicio ~]# perf stat -e cycles -I 1000 -a > # time counts unit events > 1.000304391 1,820,038 cycles > 2.000490191 1,005,477,007 cycles > 3.000657813 1,717,007 cycles > ^C 3.917890293 2,804,034 cycles > > For machines/scripts: > > [root@felicio ~]# perf stat -x, -e cycles -I 1000 -a > 1.000291954,1923360,,cycles,3998167210,100.00 > 2.000477154,1005608105,,cycles,3998475482,100.00 > 3.000612612,1345483,,cycles,3998332391,100.00 > 4.000744469,1005046913,,cycles,3998258199,100.00 > ^C 4.331684347,1551327,,cycles,3463190970,100.00 > > [root@felicio ~]# Yes, I thought about it too. Maybe -t/--field-separator option can be used to separate folded callchains too. > > > > > The main thing here is the callchain, all the other stuff are things > > > related to it, so showing it first makes sense to me. > > > > > > Having some way to list the desired info to have for each callchain may > > > be interesting, and if he could do it like: > > > > > > -g folded,count,cpu,other,fields > > > > > > then he would know how to parse the per-callchain info at the end of > > > each line, right? > > > > Hmm.. looks like that it ends up having redundant info. I don't think > > What is redundant, and with with what? When it's used with normal perf report cases, those other info in callchain lines are redundant to hist lines. Also if a hist entry has many callchains, each callchain lines will have same info in other fields. > > > it's generally useful to other 'perf report' stuffs. Wouldn't it be > > better just adding minimal support and let the external tool parse the > > output? > > Oh well, perhaps we could have a 'perf callchain' tool that would be > centered on callchains and would provided one line per callchain, which > would have: > > callchain;seprarated;colons series,of,desired,fields,for,this,callchain > > Which would reuse heavily the 'perf report' / 'perf top' code for > histograms, no? I guess the callchain code is pretty isolated or can be isolated easily though. > > I still think that this is a 'perf report' thing, but one that is > centered in callchains, and that is to be consumed by scripts, not > humans. Agreed. I'm just looking for a way to support it with minimal change. :) Thanks, Namhyung -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2015-11-03 01:50 +0100 |
| Message-ID | <qqxXr-3T2-1@gated-at.bofh.it> |
| In reply to | #1261067 |
Em Tue, Nov 03, 2015 at 08:46:06AM +0900, Namhyung Kim escreveu: > On Mon, Nov 02, 2015 at 08:04:36PM -0300, Arnaldo Carvalho de Melo wrote: > > Em Tue, Nov 03, 2015 at 07:49:27AM +0900, Namhyung Kim escreveu: > > > On Mon, Nov 02, 2015 at 07:28:42PM -0300, Arnaldo Carvalho de Melo wrote: > > > > Em Tue, Nov 03, 2015 at 07:12:04AM +0900, Namhyung Kim escreveu: > > > > > On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote: > > > > > > Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu: > > > > > > > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: > > > > > > > > This is what Brendan requested on the perf-users mailing list [1] to > > > > > > > > support FlameGraphs [2] more efficiently. This patchset adds a few > > > > > > > > more callchain options to adjust the output for it. > > > > > > > > > > > > At first, 'folded' output mode was added. The folded output puts all > > > > > > > > calchain nodes in a line separated by semicolons, a space and the > > > > > > > > value. Now it only supports --stdio as other UI provides some way of > > > > > > > > folding/expanding callchains dynamically. > > > > > > > > > > > > The value is now can be one of 'percent', 'period', or 'count'. The > > > > > > > > percent is current default output and the period is the raw number of > > > > > > > > sample periods. The count is the number of samples for each callchain. > > > > > > > > > > > > Here's an example: > > > > > > > > > > > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > > > > > > > > ... > > > > > > > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > > > > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > > > > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > > > > > > So for the folded output I don't need the summary line (the row of > > > > > > > columns printed by hist_entry__snprintf()), and don't need anything > > > > > > > except folded stacks and the counts. If working with the existing > > > > > > > stdio interface is making it harder than it needs to be, might it be > > > > > > I don't think it so, just add some flag asking for that > > > > > > hist_entry__snprintf() to be supressed, ideas for a long option name? > > > > > > Having it as Namhyung did may have value for some people as a more > > > > > > compact way to show the callchains together with the hist_entry line. > > > > > Yeah, I'd keep the hist entry line unless it's too hard to > > > > > parse/filter. IMHO it's just a way to show callchains, so no need to > > > > What I suggested was to have something like: > > > > $ perf report --no-children --no-hists --stdio -g folded,count > > > > ^^^^^^^^^^ > > > > ^^^^^^^^^^ > > > > ... > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 > > > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 > > > > > > > > I.e. the first entry in the callchain is 'intel_idle', just like in what > > > > Brendan called the 'summary line', i.e. reduntant when what he wants its > > > > just all the callchains and how many times they were sampled. > > > Yep, I know. But isn't 'perf report' all for seeing hist lines? :) > > Well, so far, yes, but he is presenting a usecase where what we want to > > see is just callchains, and we can achieve that rather easily, no? > But it's also easy to filter from the script side. Why not go all the way and provide just what the script wants? > > > I'm not insisting it strongly, but it's a bit strange for me if perf > > > report doesn't show any hist lines.. > > > > If that is of no use in this use case, why not? > > Well, I think FlameGraphs is a rather unusual case and folded output > seems useful to other use cases too. Sure thing, I agreed with that, its just one flag to tell if the hist_entry__snprintf should be used or not. > > > > > have separate output mode.. > > > > > > > > > Brendan, I guess you still need to know other info like cpu or pid, no? > > > > > > > > Possibly, but just with the callchains he has enough info for the basic > > > > flame graph, no? > > > > > > > > > And I feel like it'd be better to put the count before the callchains > > > > > for consistency like below. Is it OK to you? > > > > > > > > Consistency with what? > > > > > > Oh, I meant consistency with other callchain output style like graph, > > > fractal or flat - They all show the numbers before callchains. And I > > > think it's easier to read for human. :) > > > > Well, As I said, isn't the main object here the callchain? :-) > > > > And Brendan's request is for a something to be consumed by scripts, i.e. > > something like we have for perf stat: > > > > For humans: > > > > [root@felicio ~]# perf stat -e cycles -I 1000 -a > > # time counts unit events > > 1.000304391 1,820,038 cycles > > 2.000490191 1,005,477,007 cycles > > 3.000657813 1,717,007 cycles > > ^C 3.917890293 2,804,034 cycles > > > > For machines/scripts: > > > > [root@felicio ~]# perf stat -x, -e cycles -I 1000 -a > > 1.000291954,1923360,,cycles,3998167210,100.00 > > 2.000477154,1005608105,,cycles,3998475482,100.00 > > 3.000612612,1345483,,cycles,3998332391,100.00 > > 4.000744469,1005046913,,cycles,3998258199,100.00 > > ^C 4.331684347,1551327,,cycles,3463190970,100.00 > > > > [root@felicio ~]# > Yes, I thought about it too. Maybe -t/--field-separator option can be > used to separate folded callchains too. What I meant here was: for humans, we don't want a field separator, and we want headers, we want alignment, etc, while for scripts, its better something easily parseable and with a record per line, no alignment is needed, etc. > > > > The main thing here is the callchain, all the other stuff are things > > > > related to it, so showing it first makes sense to me. > > > > > > > > Having some way to list the desired info to have for each callchain may > > > > be interesting, and if he could do it like: > > > > > > > > -g folded,count,cpu,other,fields > > > > > > > > then he would know how to parse the per-callchain info at the end of > > > > each line, right? > > > > > > Hmm.. looks like that it ends up having redundant info. I don't think > > > > What is redundant, and with with what? > > When it's used with normal perf report cases, those other info in > callchain lines are redundant to hist lines. Also if a hist entry has Sure, but if the user doesn't want to see the output of hist_entry__snprintf()... :-) > many callchains, each callchain lines will have same info in other fields. Sure, but that would be what the script expects to consume, i.e. one line per callchain. > > > it's generally useful to other 'perf report' stuffs. Wouldn't it be > > > better just adding minimal support and let the external tool parse the > > > output? > > > > Oh well, perhaps we could have a 'perf callchain' tool that would be > > centered on callchains and would provided one line per callchain, which > > would have: > > > > callchain;seprarated;colons series,of,desired,fields,for,this,callchain > > > > Which would reuse heavily the 'perf report' / 'perf top' code for > > histograms, no? > I guess the callchain code is pretty isolated or can be isolated > easily though. > > I still think that this is a 'perf report' thing, but one that is > > centered in callchains, and that is to be consumed by scripts, not > > humans. > Agreed. > I'm just looking for a way to support it with minimal change. :) Hey, me too. A --no-hists flag looks like a quickie, no need to isolate callchain code, or anything like that, just one long option switch and we get what we need. - Arnaldo -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-03 02:40 +0100 |
| Message-ID | <qqyJQ-4pO-13@gated-at.bofh.it> |
| In reply to | #1261089 |
On Mon, Nov 02, 2015 at 09:46:47PM -0300, Arnaldo Carvalho de Melo wrote: > Em Tue, Nov 03, 2015 at 08:46:06AM +0900, Namhyung Kim escreveu: > > On Mon, Nov 02, 2015 at 08:04:36PM -0300, Arnaldo Carvalho de Melo wrote: > > > I still think that this is a 'perf report' thing, but one that is > > > centered in callchains, and that is to be consumed by scripts, not > > > humans. > > > Agreed. > > > I'm just looking for a way to support it with minimal change. :) > > Hey, me too. A --no-hists flag looks like a quickie, no need to isolate > callchain code, or anything like that, just one long option switch and > we get what we need. Hmm.. okay. Let me think about the --no-hists flags then. What do you want to do if the --no-hists flags is used without folded callchain mode or other than --stdio? And if you want to print other info in the callchains, what would be the output of non-folded mode? I think the simplest solution would be supporting the folded mode only and error out other cases. Is it ok to you? Thanks, Namhyung -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2015-11-03 02:50 +0100 |
| Message-ID | <qqyTw-4t9-3@gated-at.bofh.it> |
| In reply to | #1261115 |
Em Tue, Nov 03, 2015 at 10:35:35AM +0900, Namhyung Kim escreveu: > On Mon, Nov 02, 2015 at 09:46:47PM -0300, Arnaldo Carvalho de Melo wrote: > > Em Tue, Nov 03, 2015 at 08:46:06AM +0900, Namhyung Kim escreveu: > > > On Mon, Nov 02, 2015 at 08:04:36PM -0300, Arnaldo Carvalho de Melo wrote: > > > > I still think that this is a 'perf report' thing, but one that is > > > > centered in callchains, and that is to be consumed by scripts, not > > > > humans. > > > > > Agreed. > > > > > I'm just looking for a way to support it with minimal change. :) > > > > Hey, me too. A --no-hists flag looks like a quickie, no need to isolate > > callchain code, or anything like that, just one long option switch and > > we get what we need. > > Hmm.. okay. Let me think about the --no-hists flags then. > > What do you want to do if the --no-hists flags is used without folded > callchain mode or other than --stdio? What the user asked it to, to not show what hist_entry__snprintf() produces, i.e. just the callchains. Its left to the user to decide if that output is good for whatever purpose it has in mind. We, from this discussion, know that suppressing it when using with folded callchains, is useful at least for Brendan's scripts :-) > And if you want to print other info in the callchains, what would be > the output of non-folded mode? > I think the simplest solution would be supporting the folded mode only > and error out other cases. Is it ok to you? Well, the other info, if it comes at the end, may even be useful in non folded mode, no? If it is not, then the user will not use it, i.e. some combinations may not produce useful results, but if we want to have more flexibility to support usecases like Brendan's, and I think we want, without making the existing code overly complex, then why not? - Arnaldo -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Namhyung Kim <namhyung@kernel.org> |
|---|---|
| Date | 2015-11-03 04:20 +0100 |
| Message-ID | <qqAiC-5zj-11@gated-at.bofh.it> |
| In reply to | #1261124 |
On Mon, Nov 02, 2015 at 10:46:00PM -0300, Arnaldo Carvalho de Melo wrote:
> Em Tue, Nov 03, 2015 at 10:35:35AM +0900, Namhyung Kim escreveu:
> > On Mon, Nov 02, 2015 at 09:46:47PM -0300, Arnaldo Carvalho de Melo wrote:
> > > Em Tue, Nov 03, 2015 at 08:46:06AM +0900, Namhyung Kim escreveu:
> > > > On Mon, Nov 02, 2015 at 08:04:36PM -0300, Arnaldo Carvalho de Melo wrote:
> > > > > I still think that this is a 'perf report' thing, but one that is
> > > > > centered in callchains, and that is to be consumed by scripts, not
> > > > > humans.
> > >
> > > > Agreed.
> > >
> > > > I'm just looking for a way to support it with minimal change. :)
> > >
> > > Hey, me too. A --no-hists flag looks like a quickie, no need to isolate
> > > callchain code, or anything like that, just one long option switch and
> > > we get what we need.
> >
> > Hmm.. okay. Let me think about the --no-hists flags then.
> >
> > What do you want to do if the --no-hists flags is used without folded
> > callchain mode or other than --stdio?
>
> What the user asked it to, to not show what hist_entry__snprintf()
> produces, i.e. just the callchains.
>
> Its left to the user to decide if that output is good for whatever
> purpose it has in mind.
OK, will add it in a follow-up patch after checking TUI and GTK.
>
> We, from this discussion, know that suppressing it when using with
> folded callchains, is useful at least for Brendan's scripts :-)
OK
>
> > And if you want to print other info in the callchains, what would be
> > the output of non-folded mode?
>
> > I think the simplest solution would be supporting the folded mode only
> > and error out other cases. Is it ok to you?
>
> Well, the other info, if it comes at the end, may even be useful in non
> folded mode, no?
At the end? Brendan wanted to have it first and I think it'd be
better to show first.
Anyway, this other info depends on the sort keys - IOW it cannot show
task comm name if user gave sort keys without comm like '-s cpu'. So
how about adding 'info' or 'context' (or whatever name it) option to
-g/--call-graph to show info selected by sort keys.
For example,
$ perf report --no-children --stdio -s comm,dso -g folded,info --no-hists
28.63% swapper,[kernel.vmlinux] intel_idle;cpuidle_enter_state;...
11.30% swapper,[kernel.vmlinux] intel_idle;cpuidle_enter_state;...
$ perf report --no-children --stdio -s pid,sym -g info
...
39.93% swapper [k] intel_idle
<0:swapper,intel_idle>
|
|---intel_idel
cpuidle_enter_state
...
What do you think?
Thanks,
Namhyung
>
> If it is not, then the user will not use it, i.e. some combinations may
> not produce useful results, but if we want to have more flexibility to
> support usecases like Brendan's, and I think we want, without making the
> existing code overly complex, then why not?
>
> - Arnaldo
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Brendan Gregg <brendan.d.gregg@gmail.com> |
|---|---|
| Date | 2015-11-02 23:50 +0100 |
| Message-ID | <qqw5k-2HW-7@gated-at.bofh.it> |
| In reply to | #1261022 |
On Mon, Nov 2, 2015 at 2:12 PM, Namhyung Kim <namhyung@kernel.org> wrote: > Hi Arnaldo, > > On Mon, Nov 02, 2015 at 06:30:21PM -0300, Arnaldo Carvalho de Melo wrote: >> Em Mon, Nov 02, 2015 at 12:37:28PM -0800, Brendan Gregg escreveu: >> > G'Day Namhyung, >> > >> > On Mon, Nov 2, 2015 at 4:57 AM, Namhyung Kim <namhyung@kernel.org> wrote: >> > > Hello, >> > > >> > > This is what Brendan requested on the perf-users mailing list [1] to >> > > support FlameGraphs [2] more efficiently. This patchset adds a few >> > > more callchain options to adjust the output for it. >> > > >> > > At first, 'folded' output mode was added. The folded output puts all >> > > calchain nodes in a line separated by semicolons, a space and the >> > > value. Now it only supports --stdio as other UI provides some way of >> > > folding/expanding callchains dynamically. >> > > >> > > The value is now can be one of 'percent', 'period', or 'count'. The >> > > percent is current default output and the period is the raw number of >> > > sample periods. The count is the number of samples for each callchain. >> > > >> > > Here's an example: >> > > >> > > $ perf report --no-children --show-nr-samples --stdio -g folded,count >> > > ... >> > > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel >> > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary 57 >> > > intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... 23 >> > >> > Thanks! >> > >> > So for the folded output I don't need the summary line (the row of >> > columns printed by hist_entry__snprintf()), and don't need anything >> > except folded stacks and the counts. If working with the existing >> > stdio interface is making it harder than it needs to be, might it be >> >> I don't think it so, just add some flag asking for that >> hist_entry__snprintf() to be supressed, ideas for a long option name? >> >> Having it as Namhyung did may have value for some people as a more >> compact way to show the callchains together with the hist_entry line. > > Yeah, I'd keep the hist entry line unless it's too hard to > parse/filter. IMHO it's just a way to show callchains, so no need to > have separate output mode.. Ok, good point, it can be thought of as a different stack representation format. > > Brendan, I guess you still need to know other info like cpu or pid, no? > Yes, I just realized that I either include the process name (Command column) or name-PID, as the first folded element. Eg, output can be: mkdir;getopt_long;page_fault;do_page_fault;__do_page_fault;filemap_map_pages 3 Or: mkdir-21918;getopt_long;page_fault;do_page_fault;__do_page_fault;filemap_map_pages 2 Usually the first, but sometimes it's helpful to split on PID as well. As for what to call such options (which may be a follow on patch anyway) ... maybe something like: "folded": fold stacks as single lines "nameonly,folded": suppress summary line and include process name in the folded stack "pidonly,folded": suppress summary line and include process_name-PID in the folded stack > And I feel like it'd be better to put the count before the callchains > for consistency like below. Is it OK to you? > > $ perf report --no-children --show-nr-samples --stdio -g folded,count > ... > 39.93% 80 swapper [kernel.vmlinux] [k] intel_idel > 57 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;start_secondary > 23 intel_idle;cpuidle_enter_state;cpuidle_enter;call_cpuidle;cpu_startup_entry;rest_init;... > If it was printing with the perf report summary, sure, but if we have a way to only emit folded output, then counts last would be perfect and maybe a bit more intuitive (key then value). > >> >> With this in mind, do you have any other issues with Namhyung's >> patchkit? An acked-by/tested-by you would be nice to have, and then we >> could work out the new option to suppress that hist_entry__snprintf() >> in a follow up patch. Acked and tested, yes. Looks like I'd be using caller ordering, eg, to get lines like this: __GI___libc_read;entry_SYSCALL_64_fastpath;sys_read;vfs_read;__vfs_read;urandom_read;extract_entropy_user;extract_buf;check_events;xen_hypercall_xen_version 91 Which I can do just by using "-g folded,count,caller". >> >> > easier to make it a separate interface (ui/folded), that just emitted >> > the folded output? Just an idea. This existing patchset is working for >> > me, I'd just be filtering the output. >> > >> > Having the option for percentages and periods is nice. I can envisage >> > using periods (for latency flame graphs). > > Glad to see you like it. :) > > Thanks, > Namhyung -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web