Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1500566 > unrolled thread
| Started by | Andi Kleen <andi@firstfloor.org> |
|---|---|
| First post | 2016-10-13 23:20 +0200 |
| Last post | 2016-10-17 13:10 +0200 |
| Articles | 18 on this page of 38 — 5 participants |
Back to article view | Back to linux.kernel
Support Intel uncore event lists Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
[PATCH 07/10] perf, tools: Collapse identically named events in perf stat Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:00 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:30 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Andi Kleen <andi@firstfloor.org> - 2016-10-17 18:40 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Jiri Olsa <jolsa@redhat.com> - 2016-10-17 19:30 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Andi Kleen <ak@linux.intel.com> - 2016-10-17 20:20 +0200
Re: [PATCH 07/10] perf, tools: Collapse identically named events in perf stat Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:30 +0200
[PATCH 08/10] perf, tools: Expand PMU events by prefix match Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 08/10] perf, tools: Expand PMU events by prefix match Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:40 +0200
Re: [PATCH 08/10] perf, tools: Expand PMU events by prefix match Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:50 +0200
Re: [PATCH 08/10] perf, tools: Expand PMU events by prefix match Andi Kleen <andi@firstfloor.org> - 2016-10-17 19:20 +0200
Re: [PATCH 08/10] perf, tools: Expand PMU events by prefix match Jiri Olsa <jolsa@redhat.com> - 2016-10-17 19:30 +0200
[PATCH 02/10] perf, tools: Only print Using CPUID message once Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 02/10] perf, tools: Only print Using CPUID message once Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-10-14 17:50 +0200
[tip:perf/core] perf pmu: Only print Using CPUID message once tip-bot for Andi Kleen <tipbot@zytor.com> - 2016-10-24 21:10 +0200
[PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:50 +0200
Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Andi Kleen <andi@firstfloor.org> - 2016-10-17 18:30 +0200
Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Jiri Olsa <jolsa@redhat.com> - 2016-10-17 19:50 +0200
Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Jiri Olsa <jolsa@redhat.com> - 2016-10-17 19:50 +0200
Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list Andi Kleen <andi@firstfloor.org> - 2016-10-17 20:30 +0200
[PATCH 04/10] perf, tools: Support per pmu json aliases Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 04/10] perf, tools: Support per pmu json aliases Jiri Olsa <jolsa@redhat.com> - 2016-10-14 14:40 +0200
[PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus Jiri Olsa <jolsa@redhat.com> - 2016-10-17 11:40 +0200
Re: [PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus Jiri Olsa <jolsa@redhat.com> - 2016-10-17 12:30 +0200
[PATCH 01/10] perf, tools: Factor out scale conversion code Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-10-14 17:40 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Andi Kleen <andi@firstfloor.org> - 2016-10-14 17:50 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-10-14 18:20 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Andi Kleen <ak@linux.intel.com> - 2016-10-14 18:20 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-10-14 18:30 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-10-14 18:40 +0200
Re: [PATCH 01/10] perf, tools: Factor out scale conversion code Andi Kleen <andi@firstfloor.org> - 2016-10-14 18:40 +0200
[PATCH 10/10] perf, tools, stat: Output generic dividedby metric Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
[PATCH 06/10] perf, tools: Add debug support for outputing alias string Andi Kleen <andi@firstfloor.org> - 2016-10-13 23:20 +0200
Re: Support Intel uncore event lists Jiri Olsa <jolsa@redhat.com> - 2016-10-17 13:10 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-10-17 19:50 +0200 |
| Subject | Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list |
| Message-ID | <stkcW-6Af-25@gated-at.bofh.it> |
| In reply to | #1502174 |
On Mon, Oct 17, 2016 at 09:27:54AM -0700, Andi Kleen wrote: > On Mon, Oct 17, 2016 at 01:44:43PM +0200, Jiri Olsa wrote: > > On Thu, Oct 13, 2016 at 02:15:31PM -0700, Andi Kleen wrote: > > > From: Andi Kleen <ak@linux.intel.com> > > > > > > Add support for parsing the DividedBy header in the JSON event lists and > > > storing them in the alias structure. > > > > I wish you'd add JSON tags always one by one as you did in here ;-) > > > > however Ithink we'll need more info here: > > - what's the value? > > - what's it going to be used for? > > That's all described in the next patch. But I can copy the description. > > > - looks like formula stuff, why post processing via python/perl can't be used in this case? > > It would be fairly complicated to interface that with event lists, and > also still wouldn't work with standard perf stat. > > DividedBy already covers the majority of interesting cases and fits > nicely with the existing frame work. If we wanted more complex > formulas something with python would be probably needed, but I don't see > the need yet. so.. - you put 'DividedBy' into JSON event's defition any further explanation how or why the format we use for event defs will be used now used to describe ratios - then you force perf stat to merge together all 'same' uncore events to get just one number.. - then you display that ratio (just the number) in perf stat metrics output without any explanation or description I dont see that as a nicely fit, more like hack please let's go first to discuss the DividedBy being included in JSON defs, which is fragile topic to begin with thanks, jirka
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-17 20:30 +0200 |
| Subject | Re: [PATCH 09/10] perf, tools: Support DividedBy header in JSON event list |
| Message-ID | <stkPD-7cX-15@gated-at.bofh.it> |
| In reply to | #1502316 |
>
> so..
>
> - you put 'DividedBy' into JSON event's defition any further
> explanation how or why the format we use for event defs will
> be used now used to describe ratios
>
> - then you force perf stat to merge together all 'same' uncore events
> to get just one number..
The ratios don't need that. They work fine without merging.
It's an independent feature.
The merging is just a convenience feature for some of the uncore pmus to make
the output much more readable. For example the cbox pmu is duplicated
for each core, and we have systems with 21 cores per socket now.
So without merging you end up with something like the output below.
The second variant is much more readable.
>
> - then you display that ratio (just the number) in perf stat metrics output
> without any explanation or description
The event name is already expressive enough. I find it fairly
straight forward that there is a ratio attached with a count.
>
> I dont see that as a nicely fit, more like hack
I would call it a simple solution that works well.
I don't see any easy path for full scripting, and also I think
it would be vastly overengineered here needing a lot of
infastructure that isn't really needed.
-Andi
% perf stat --no-merge -a -e unc_c_llc_lookup.any sleep 1
Performance counter stats for 'system wide':
694,976 Bytes unc_c_llc_lookup.any
706,304 Bytes unc_c_llc_lookup.any
956,608 Bytes unc_c_llc_lookup.any
782,720 Bytes unc_c_llc_lookup.any
605,696 Bytes unc_c_llc_lookup.any
442,816 Bytes unc_c_llc_lookup.any
659,328 Bytes unc_c_llc_lookup.any
509,312 Bytes unc_c_llc_lookup.any
263,936 Bytes unc_c_llc_lookup.any
592,448 Bytes unc_c_llc_lookup.any
672,448 Bytes unc_c_llc_lookup.any
608,640 Bytes unc_c_llc_lookup.any
641,024 Bytes unc_c_llc_lookup.any
856,896 Bytes unc_c_llc_lookup.any
808,832 Bytes unc_c_llc_lookup.any
684,864 Bytes unc_c_llc_lookup.any
710,464 Bytes unc_c_llc_lookup.any
538,304 Bytes unc_c_llc_lookup.any
1.002577660 seconds time elapsed
% perf stat -a -e unc_c_llc_lookup.any sleep 1
Performance counter stats for 'system wide':
2,685,120 Bytes unc_c_llc_lookup.any
1.002648032 seconds time elapsed
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-13 23:20 +0200 |
| Subject | [PATCH 04/10] perf, tools: Support per pmu json aliases |
| Message-ID | <srVzY-qJ-33@gated-at.bofh.it> |
| In reply to | #1500566 |
From: Andi Kleen <ak@linux.intel.com>
Add support for registering json aliases per PMU. Any alias
with an unit matching the prefix is registered to the PMU.
Uncore has multiple instances of most units, so all
these aliases get registered for each individual PMU
(this is important later to run the event on every instance
of the PMU).
To avoid printing the events multiple times in perf list
filter out duplicated events during printing.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/util/pmu.c | 17 +++++++++++------
1 file changed, 11 insertions(+), 6 deletions(-)
diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
index 363cb7b0ccc7..f8a052a793b1 100644
--- a/tools/perf/util/pmu.c
+++ b/tools/perf/util/pmu.c
@@ -509,7 +509,7 @@ char * __weak get_cpuid_str(void)
* to the current running CPU. Then, add all PMU events from that table
* as aliases.
*/
-static void pmu_add_cpu_aliases(struct list_head *head)
+static void pmu_add_cpu_aliases(struct list_head *head, const char *name)
{
int i;
struct pmu_events_map *map;
@@ -575,6 +575,7 @@ static struct perf_pmu *pmu_lookup(const char *name)
LIST_HEAD(format);
LIST_HEAD(aliases);
__u32 type;
+ int noff = 0;
/*
* The pmu data we store & need consists of the pmu
@@ -584,15 +585,16 @@ static struct perf_pmu *pmu_lookup(const char *name)
if (pmu_format(name, &format))
return NULL;
- if (pmu_aliases(name, &aliases))
+ if (pmu_type(name, &type))
return NULL;
- if (!strcmp(name, "cpu"))
- pmu_add_cpu_aliases(&aliases);
-
- if (pmu_type(name, &type))
+ if (pmu_aliases(name, &aliases))
return NULL;
+ if (!strncmp(name, "uncore_", 7))
+ noff = 7;
+
+ pmu_add_cpu_aliases(&aliases, name + noff);
pmu = zalloc(sizeof(*pmu));
if (!pmu)
return NULL;
@@ -1188,6 +1190,9 @@ void print_pmu_events(const char *event_glob, bool name_only, bool quiet_flag,
len = j;
qsort(aliases, len, sizeof(struct sevent), cmp_sevent);
for (j = 0; j < len; j++) {
+ /* Skip duplicates */
+ if (j > 0 && !strcmp(aliases[j].name, aliases[j - 1].name))
+ continue;
if (name_only) {
printf("%s ", aliases[j].name);
continue;
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-10-14 14:40 +0200 |
| Subject | Re: [PATCH 04/10] perf, tools: Support per pmu json aliases |
| Message-ID | <ss9Wi-1g1-23@gated-at.bofh.it> |
| In reply to | #1500572 |
On Thu, Oct 13, 2016 at 02:15:26PM -0700, Andi Kleen wrote:
> From: Andi Kleen <ak@linux.intel.com>
>
> Add support for registering json aliases per PMU. Any alias
> with an unit matching the prefix is registered to the PMU.
> Uncore has multiple instances of most units, so all
> these aliases get registered for each individual PMU
> (this is important later to run the event on every instance
> of the PMU).
>
> To avoid printing the events multiple times in perf list
> filter out duplicated events during printing.
>
> Signed-off-by: Andi Kleen <ak@linux.intel.com>
> ---
> tools/perf/util/pmu.c | 17 +++++++++++------
> 1 file changed, 11 insertions(+), 6 deletions(-)
>
> diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
> index 363cb7b0ccc7..f8a052a793b1 100644
> --- a/tools/perf/util/pmu.c
> +++ b/tools/perf/util/pmu.c
> @@ -509,7 +509,7 @@ char * __weak get_cpuid_str(void)
> * to the current running CPU. Then, add all PMU events from that table
> * as aliases.
> */
> -static void pmu_add_cpu_aliases(struct list_head *head)
> +static void pmu_add_cpu_aliases(struct list_head *head, const char *name)
> {
> int i;
> struct pmu_events_map *map;
> @@ -575,6 +575,7 @@ static struct perf_pmu *pmu_lookup(const char *name)
> LIST_HEAD(format);
> LIST_HEAD(aliases);
> __u32 type;
> + int noff = 0;
>
> /*
> * The pmu data we store & need consists of the pmu
> @@ -584,15 +585,16 @@ static struct perf_pmu *pmu_lookup(const char *name)
> if (pmu_format(name, &format))
> return NULL;
>
> - if (pmu_aliases(name, &aliases))
> + if (pmu_type(name, &type))
> return NULL;
>
> - if (!strcmp(name, "cpu"))
> - pmu_add_cpu_aliases(&aliases);
> -
> - if (pmu_type(name, &type))
> + if (pmu_aliases(name, &aliases))
> return NULL;
>
> + if (!strncmp(name, "uncore_", 7))
> + noff = 7;
please do this (best in a function) within pmu_add_cpu_aliases
in the check itself:
if (pe->pmu && strncmp(pe->pmu, name, strlen(pe->pmu)))
continue;
Also any chance the json Unit field could have a uncore prefix already?
would be nice to see more info about those fields in changelog
thanks,
jirka
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-13 23:20 +0200 |
| Subject | [PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus |
| Message-ID | <srVzY-qJ-19@gated-at.bofh.it> |
| In reply to | #1500566 |
From: Andi Kleen <ak@linux.intel.com>
The code for handling pmu aliases without specifying
the PMU hardcoded only supported the cpu PMU.
This patch extends it to work for all PMUs. We always
duplicate the event for all PMUs that have an matching alias.
This allows to automatically expand an alias for all instances
of a PMU (so for example you can monitor all cache boxes with
a single event)
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/util/parse-events.c | 46 ++++++++++++++++++++++++------------------
tools/perf/util/parse-events.y | 32 ++++++++++++++++++++++-------
2 files changed, 51 insertions(+), 27 deletions(-)
diff --git a/tools/perf/util/parse-events.c b/tools/perf/util/parse-events.c
index 4e778eae1510..a2bbd17a0dc3 100644
--- a/tools/perf/util/parse-events.c
+++ b/tools/perf/util/parse-events.c
@@ -1497,35 +1497,41 @@ static void perf_pmu__parse_init(void)
struct perf_pmu_alias *alias;
int len = 0;
- pmu = perf_pmu__find("cpu");
- if ((pmu == NULL) || list_empty(&pmu->aliases)) {
+ pmu = NULL;
+ while ((pmu = perf_pmu__scan(pmu)) != NULL) {
+ list_for_each_entry(alias, &pmu->aliases, list) {
+ if (strchr(alias->name, '-'))
+ len++;
+ len++;
+ }
+ }
+
+ if (len == 0) {
perf_pmu_events_list_num = -1;
return;
}
- list_for_each_entry(alias, &pmu->aliases, list) {
- if (strchr(alias->name, '-'))
- len++;
- len++;
- }
perf_pmu_events_list = malloc(sizeof(struct perf_pmu_event_symbol) * len);
if (!perf_pmu_events_list)
return;
perf_pmu_events_list_num = len;
len = 0;
- list_for_each_entry(alias, &pmu->aliases, list) {
- struct perf_pmu_event_symbol *p = perf_pmu_events_list + len;
- char *tmp = strchr(alias->name, '-');
-
- if (tmp != NULL) {
- SET_SYMBOL(strndup(alias->name, tmp - alias->name),
- PMU_EVENT_SYMBOL_PREFIX);
- p++;
- SET_SYMBOL(strdup(++tmp), PMU_EVENT_SYMBOL_SUFFIX);
- len += 2;
- } else {
- SET_SYMBOL(strdup(alias->name), PMU_EVENT_SYMBOL);
- len++;
+ pmu = NULL;
+ while ((pmu = perf_pmu__scan(pmu)) != NULL) {
+ list_for_each_entry(alias, &pmu->aliases, list) {
+ struct perf_pmu_event_symbol *p = perf_pmu_events_list + len;
+ char *tmp = strchr(alias->name, '-');
+
+ if (tmp != NULL) {
+ SET_SYMBOL(strndup(alias->name, tmp - alias->name),
+ PMU_EVENT_SYMBOL_PREFIX);
+ p++;
+ SET_SYMBOL(strdup(++tmp), PMU_EVENT_SYMBOL_SUFFIX);
+ len += 2;
+ } else {
+ SET_SYMBOL(strdup(alias->name), PMU_EVENT_SYMBOL);
+ len++;
+ }
}
}
qsort(perf_pmu_events_list, len,
diff --git a/tools/perf/util/parse-events.y b/tools/perf/util/parse-events.y
index 879115f93edc..f3b5ec901600 100644
--- a/tools/perf/util/parse-events.y
+++ b/tools/perf/util/parse-events.y
@@ -12,6 +12,7 @@
#include <linux/list.h>
#include <linux/types.h>
#include "util.h"
+#include "pmu.h"
#include "parse-events.h"
#include "parse-events-bison.h"
@@ -236,15 +237,32 @@ PE_KERNEL_PMU_EVENT sep_dc
struct list_head *head;
struct parse_events_term *term;
struct list_head *list;
+ struct perf_pmu *pmu = NULL;
+ int ok = 0;
- ALLOC_LIST(head);
- ABORT_ON(parse_events_term__num(&term, PARSE_EVENTS__TERM_TYPE_USER,
- $1, 1, &@1, NULL));
- list_add_tail(&term->list, head);
-
+ /* Add it for all PMUs that support the alias */
ALLOC_LIST(list);
- ABORT_ON(parse_events_add_pmu(data, list, "cpu", head));
- parse_events_terms__delete(head);
+ while ((pmu = perf_pmu__scan(pmu)) != NULL) {
+ struct perf_pmu_alias *alias;
+
+ list_for_each_entry(alias, &pmu->aliases, list) {
+ if (!strcasecmp(alias->name, $1)) {
+ ALLOC_LIST(head);
+ ABORT_ON(parse_events_term__num(&term, PARSE_EVENTS__TERM_TYPE_USER,
+ $1, 1, &@1, NULL));
+ list_add_tail(&term->list, head);
+
+ if (!parse_events_add_pmu(data, list,
+ pmu->name, head)) {
+ ok++;
+ }
+
+ parse_events_terms__delete(head);
+ }
+ }
+ }
+ if (!ok)
+ YYABORT;
$$ = list;
}
|
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-10-17 11:40 +0200 |
| Subject | Re: [PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus |
| Message-ID | <stcyJ-1Fs-15@gated-at.bofh.it> |
| In reply to | #1500573 |
On Thu, Oct 13, 2016 at 02:15:27PM -0700, Andi Kleen wrote:
SNIP
> @@ -236,15 +237,32 @@ PE_KERNEL_PMU_EVENT sep_dc
> struct list_head *head;
> struct parse_events_term *term;
> struct list_head *list;
> + struct perf_pmu *pmu = NULL;
> + int ok = 0;
>
> - ALLOC_LIST(head);
> - ABORT_ON(parse_events_term__num(&term, PARSE_EVENTS__TERM_TYPE_USER,
> - $1, 1, &@1, NULL));
> - list_add_tail(&term->list, head);
> -
> + /* Add it for all PMUs that support the alias */
> ALLOC_LIST(list);
> - ABORT_ON(parse_events_add_pmu(data, list, "cpu", head));
> - parse_events_terms__delete(head);
> + while ((pmu = perf_pmu__scan(pmu)) != NULL) {
> + struct perf_pmu_alias *alias;
> +
> + list_for_each_entry(alias, &pmu->aliases, list) {
> + if (!strcasecmp(alias->name, $1)) {
> + ALLOC_LIST(head);
> + ABORT_ON(parse_events_term__num(&term, PARSE_EVENTS__TERM_TYPE_USER,
> + $1, 1, &@1, NULL));
> + list_add_tail(&term->list, head);
> +
> + if (!parse_events_add_pmu(data, list,
> + pmu->name, head)) {
> + ok++;
> + }
> +
> + parse_events_terms__delete(head);
> + }
> + }
> + }
> + if (!ok)
> + YYABORT;
> $$ = list;
> }
> |
how about the next rule:
PE_PMU_EVENT_PRE '-' PE_PMU_EVENT_SUF sep_dc
I think that must be changed as well
jirka
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-10-17 12:30 +0200 |
| Subject | Re: [PATCH 05/10] perf, tools: Support event aliases for non cpu// pmus |
| Message-ID | <stdl7-2ek-27@gated-at.bofh.it> |
| In reply to | #1500573 |
On Thu, Oct 13, 2016 at 02:15:27PM -0700, Andi Kleen wrote: > From: Andi Kleen <ak@linux.intel.com> > > The code for handling pmu aliases without specifying > the PMU hardcoded only supported the cpu PMU. > > This patch extends it to work for all PMUs. We always > duplicate the event for all PMUs that have an matching alias. > This allows to automatically expand an alias for all instances > of a PMU (so for example you can monitor all cache boxes with > a single event) could you please put some examples of new usage into changelog.. it's be easier to sell it thanks, jirka
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-13 23:20 +0200 |
| Subject | [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <srVzY-qJ-35@gated-at.bofh.it> |
| In reply to | #1500566 |
From: Andi Kleen <ak@linux.intel.com>
Move the scale factor parsing code to an own function
to reuse it in an upcoming patch.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/util/pmu.c | 64 +++++++++++++++++++++++++++------------------------
1 file changed, 34 insertions(+), 30 deletions(-)
diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
index b1474dcadfa2..9adae7e7477c 100644
--- a/tools/perf/util/pmu.c
+++ b/tools/perf/util/pmu.c
@@ -94,32 +94,10 @@ static int pmu_format(const char *name, struct list_head *format)
return 0;
}
-static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *name)
+static double convert_scale(const char *scale, char **end)
{
- struct stat st;
- ssize_t sret;
- char scale[128];
- int fd, ret = -1;
- char path[PATH_MAX];
char *lc;
-
- snprintf(path, PATH_MAX, "%s/%s.scale", dir, name);
-
- fd = open(path, O_RDONLY);
- if (fd == -1)
- return -1;
-
- if (fstat(fd, &st) < 0)
- goto error;
-
- sret = read(fd, scale, sizeof(scale)-1);
- if (sret < 0)
- goto error;
-
- if (scale[sret - 1] == '\n')
- scale[sret - 1] = '\0';
- else
- scale[sret] = '\0';
+ double sval;
/*
* save current locale
@@ -132,10 +110,8 @@ static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *
* call below.
*/
lc = strdup(lc);
- if (!lc) {
- ret = -ENOMEM;
- goto error;
- }
+ if (!lc)
+ return 1.0;
/*
* force to C locale to ensure kernel
@@ -144,13 +120,41 @@ static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *
*/
setlocale(LC_NUMERIC, "C");
- alias->scale = strtod(scale, NULL);
+ sval = strtod(scale, end);
/* restore locale */
setlocale(LC_NUMERIC, lc);
-
free(lc);
+ return sval;
+}
+
+static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *name)
+{
+ struct stat st;
+ ssize_t sret;
+ char scale[128];
+ int fd, ret = -1;
+ char path[PATH_MAX];
+
+ snprintf(path, PATH_MAX, "%s/%s.scale", dir, name);
+
+ fd = open(path, O_RDONLY);
+ if (fd == -1)
+ return -1;
+
+ if (fstat(fd, &st) < 0)
+ goto error;
+
+ sret = read(fd, scale, sizeof(scale)-1);
+ if (sret < 0)
+ goto error;
+
+ if (scale[sret - 1] == '\n')
+ scale[sret - 1] = '\0';
+ else
+ scale[sret] = '\0';
+ alias->scale = convert_scale(scale, NULL);
ret = 0;
error:
close(fd);
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2016-10-14 17:40 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <sscKu-36Y-1@gated-at.bofh.it> |
| In reply to | #1500575 |
Em Thu, Oct 13, 2016 at 02:15:23PM -0700, Andi Kleen escreveu:
> From: Andi Kleen <ak@linux.intel.com>
>
> Move the scale factor parsing code to an own function
> to reuse it in an upcoming patch.
>
> Signed-off-by: Andi Kleen <ak@linux.intel.com>
> ---
> tools/perf/util/pmu.c | 64 +++++++++++++++++++++++++++------------------------
> 1 file changed, 34 insertions(+), 30 deletions(-)
>
> diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
> index b1474dcadfa2..9adae7e7477c 100644
> --- a/tools/perf/util/pmu.c
> +++ b/tools/perf/util/pmu.c
> @@ -94,32 +94,10 @@ static int pmu_format(const char *name, struct list_head *format)
> return 0;
> }
>
> -static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *name)
> +static double convert_scale(const char *scale, char **end)
> {
> - struct stat st;
> - ssize_t sret;
> - char scale[128];
> - int fd, ret = -1;
> - char path[PATH_MAX];
> char *lc;
> -
> - snprintf(path, PATH_MAX, "%s/%s.scale", dir, name);
> -
> - fd = open(path, O_RDONLY);
> - if (fd == -1)
> - return -1;
> -
> - if (fstat(fd, &st) < 0)
> - goto error;
> -
> - sret = read(fd, scale, sizeof(scale)-1);
> - if (sret < 0)
> - goto error;
> -
> - if (scale[sret - 1] == '\n')
> - scale[sret - 1] = '\0';
> - else
> - scale[sret] = '\0';
> + double sval;
>
> /*
> * save current locale
> @@ -132,10 +110,8 @@ static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *
> * call below.
> */
> lc = strdup(lc);
> - if (!lc) {
> - ret = -ENOMEM;
> - goto error;
> - }
> + if (!lc)
> + return 1.0;
So if we can't convert the scale because of an allocation failure
related to locale issues we silently trow it away and do no scale at
all?
- Arnaldo
>
> /*
> * force to C locale to ensure kernel
> @@ -144,13 +120,41 @@ static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *
> */
> setlocale(LC_NUMERIC, "C");
>
> - alias->scale = strtod(scale, NULL);
> + sval = strtod(scale, end);
>
> /* restore locale */
> setlocale(LC_NUMERIC, lc);
> -
> free(lc);
> + return sval;
> +}
> +
> +static int perf_pmu__parse_scale(struct perf_pmu_alias *alias, char *dir, char *name)
> +{
> + struct stat st;
> + ssize_t sret;
> + char scale[128];
> + int fd, ret = -1;
> + char path[PATH_MAX];
> +
> + snprintf(path, PATH_MAX, "%s/%s.scale", dir, name);
> +
> + fd = open(path, O_RDONLY);
> + if (fd == -1)
> + return -1;
> +
> + if (fstat(fd, &st) < 0)
> + goto error;
> +
> + sret = read(fd, scale, sizeof(scale)-1);
> + if (sret < 0)
> + goto error;
> +
> + if (scale[sret - 1] == '\n')
> + scale[sret - 1] = '\0';
> + else
> + scale[sret] = '\0';
>
> + alias->scale = convert_scale(scale, NULL);
> ret = 0;
> error:
> close(fd);
> --
> 2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-14 17:50 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <sscUa-3ay-13@gated-at.bofh.it> |
| In reply to | #1501062 |
> So if we can't convert the scale because of an allocation failure > related to locale issues we silently trow it away and do no scale at > all? That is right. If your machine is thrashing to death in a OOM this is your smallest problem. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2016-10-14 18:20 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <ssdnc-3zJ-3@gated-at.bofh.it> |
| In reply to | #1501065 |
Em Fri, Oct 14, 2016 at 08:45:15AM -0700, Andi Kleen escreveu: > > So if we can't convert the scale because of an allocation failure > > related to locale issues we silently trow it away and do no scale at > > all? > That is right. If your machine is thrashing to death in a OOM this > is your smallest problem. Please keep it was before, i.e. return an error value, and bail out. It was like that before, why introduce these kinds of silent "do something else in an unlikely case" handling? - Arnaldo
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2016-10-14 18:20 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <ssdnc-3zJ-5@gated-at.bofh.it> |
| In reply to | #1501079 |
On Fri, Oct 14, 2016 at 01:08:42PM -0300, Arnaldo Carvalho de Melo wrote: > Em Fri, Oct 14, 2016 at 08:45:15AM -0700, Andi Kleen escreveu: > > > So if we can't convert the scale because of an allocation failure > > > related to locale issues we silently trow it away and do no scale at > > > all? > > > That is right. If your machine is thrashing to death in a OOM this > > is your smallest problem. > > Please keep it was before, i.e. return an error value, and bail out. It > was like that before, why introduce these kinds of silent "do something > else in an unlikely case" handling? Ok. I will fix it. But just for the record I don't think this fine grained memory error handling makes any sense for perf. It is needed in the kernel, but it's not appropiate for user programs: - Usually when you're out of memory then every thing afterward that needs memory will fail too, so there's no sane way to continue, - When you run out of memory in user space you usually get killed at some point anyways because the OOM killer kicks in. - The only exception is that you run out of VA space, but then the point above applies. - These error paths are all untested and most likely a significant fraction of them is broken because untested code is often broken. What most user space does is to just have malloc wrappers that exit when you run out of memory with an error message. That's nearly always the right strategy for user programs. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2016-10-14 18:30 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <ssdwS-3Da-7@gated-at.bofh.it> |
| In reply to | #1501080 |
Em Fri, Oct 14, 2016 at 09:15:16AM -0700, Andi Kleen escreveu:
> On Fri, Oct 14, 2016 at 01:08:42PM -0300, Arnaldo Carvalho de Melo wrote:
> > Em Fri, Oct 14, 2016 at 08:45:15AM -0700, Andi Kleen escreveu:
> > > > So if we can't convert the scale because of an allocation failure
> > > > related to locale issues we silently trow it away and do no scale at
> > > > all?
> >
> > > That is right. If your machine is thrashing to death in a OOM this
> > > is your smallest problem.
> >
> > Please keep it was before, i.e. return an error value, and bail out. It
> > was like that before, why introduce these kinds of silent "do something
> > else in an unlikely case" handling?
>
> Ok. I will fix it.
>
> But just for the record I don't think this fine grained memory error
> handling makes any sense for perf. It is needed in the kernel, but
> it's not appropiate for user programs:
>
> - Usually when you're out of memory then every thing afterward
> that needs memory will fail too, so there's no sane way to continue,
> - When you run out of memory in user space you usually get killed
> at some point anyways because the OOM killer kicks in.
> - The only exception is that you run out of VA space, but then the point
> above applies.
> - These error paths are all untested and most likely a significant
> fraction of them is broken because untested code is often broken.
>
> What most user space does is to just have malloc wrappers that
> exit when you run out of memory with an error message. That's nearly
> always the right strategy for user programs.
I disagree, and I usually try not to differentiate that much if I'm
programming for the kernel or for userspace, in fact, I try as much as
possible to reduce the gap of writing for the kernel or writing for
userspace.
I tools/ specifically, I try to use the same constructs, list.h, rbtree,
WARN_, pr_, etc, not panic()'ing, etc.
In this specific case, doing a:
printf("life is hard, I give up"); exit(1);
would be preferrable to silently not doing what was asked for, namely do
the scaling, with that said, please keep the original way of handling it
and just return an error when what was asked for can't be done.
Thanks,
- Arnaldo
[toc] | [prev] | [next] | [standalone]
| From | Arnaldo Carvalho de Melo <acme@kernel.org> |
|---|---|
| Date | 2016-10-14 18:40 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <ssdGy-3GF-39@gated-at.bofh.it> |
| In reply to | #1501086 |
Em Fri, Oct 14, 2016 at 09:36:42AM -0700, Andi Kleen escreveu:
> > I tools/ specifically, I try to use the same constructs, list.h, rbtree,
> > WARN_, pr_, etc, not panic()'ing, etc.
> >
> > In this specific case, doing a:
> >
> > printf("life is hard, I give up"); exit(1);
>
> Ok that would be just what a wrapper does, only open coded.
>
> How about we just add helpers for these cases?
>
> xstrdup
> xmalloc
> xasprintf
Nope, we had those, I removed them already.
- Arnaldo
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-14 18:40 +0200 |
| Subject | Re: [PATCH 01/10] perf, tools: Factor out scale conversion code |
| Message-ID | <ssdGy-3GF-41@gated-at.bofh.it> |
| In reply to | #1501086 |
> I tools/ specifically, I try to use the same constructs, list.h, rbtree,
> WARN_, pr_, etc, not panic()'ing, etc.
>
> In this specific case, doing a:
>
> printf("life is hard, I give up"); exit(1);
Ok that would be just what a wrapper does, only open coded.
How about we just add helpers for these cases?
xstrdup
xmalloc
xasprintf
etc.
-Andi
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-13 23:20 +0200 |
| Subject | [PATCH 10/10] perf, tools, stat: Output generic dividedby metric |
| Message-ID | <srVzY-qJ-23@gated-at.bofh.it> |
| In reply to | #1500566 |
From: Andi Kleen <ak@linux.intel.com>
Add generic infrastructure to perf stat to output ratios for "DividedBy"
entries in the event lists. Many events are more useful as ratios
than in raw form, typically some count in relation to total ticks.
Transfer the dividedby information from the alias to the evsel.
We mark the events that need to be collected for DividedBy, and also
link the events using them with a pointer. The code is careful
to always prefer the right event in the same group to minimize
multiplexing errors.
Then add a rblist to the stat shadow code that remembers stats based
on the cpu and context.
Then finally update and retrieve and print these ratios similarly to the
existing hardcoded perf metrics.
Normally we just output the ratio as percent without further commentary,
but for --metric-only this would lead to empty columns. So for this
case use the original event as description.
So far there is no attempt to automatically add the DividedBy event,
if it is missing, however we suggest it to the user.
$ perf stat -a -I 1000 -e '{unc_p_clockticks,unc_p_freq_max_os_cycles}'
1.000228813 800,139,950 unc_p_clockticks
1.000228813 789,833,783 unc_p_freq_max_os_cycles # 98.7%
2.000654229 800,308,990 unc_p_clockticks
2.000654229 396,214,238 unc_p_freq_max_os_cycles # 49.5%
$ perf stat -a -I 1000 -e '{unc_p_clockticks,unc_p_freq_max_os_cycles}' --metric-only
1.000206740 48.0%
2.000451543 48.1%
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/builtin-stat.c | 3 +
tools/perf/util/evsel.c | 3 +
tools/perf/util/evsel.h | 3 +
tools/perf/util/parse-events.c | 1 +
tools/perf/util/pmu.c | 2 +
tools/perf/util/pmu.h | 1 +
tools/perf/util/stat-shadow.c | 139 +++++++++++++++++++++++++++++++++++++++++
tools/perf/util/stat.h | 2 +
8 files changed, 154 insertions(+)
diff --git a/tools/perf/builtin-stat.c b/tools/perf/builtin-stat.c
index 76304f27c090..b6702a8e0031 100644
--- a/tools/perf/builtin-stat.c
+++ b/tools/perf/builtin-stat.c
@@ -1141,6 +1141,7 @@ static void printout(int id, int nr, struct perf_evsel *counter, double uval,
out.print_metric = pm;
out.new_line = nl;
out.ctx = &os;
+ out.force_header = false;
if (csv_output && !metric_only) {
print_noise(counter, noise);
@@ -1458,6 +1459,7 @@ static void print_metric_headers(const char *prefix, bool no_indent)
out.ctx = &os;
out.print_metric = print_metric_header;
out.new_line = new_line_metric;
+ out.force_header = true;
os.evsel = counter;
perf_stat__print_shadow_stats(counter, 0,
0,
@@ -2440,6 +2442,7 @@ int cmd_stat(int argc, const char **argv, const char *prefix __maybe_unused)
argc = parse_options_subcommand(argc, argv, stat_options, stat_subcommands,
(const char **) stat_usage,
PARSE_OPT_STOP_AT_NON_OPTION);
+ perf_stat__collect_dividedby(evsel_list);
perf_stat__init_shadow_stats();
if (csv_sep) {
diff --git a/tools/perf/util/evsel.c b/tools/perf/util/evsel.c
index 8bc271141d9d..3484c0c67f8d 100644
--- a/tools/perf/util/evsel.c
+++ b/tools/perf/util/evsel.c
@@ -235,6 +235,9 @@ void perf_evsel__init(struct perf_evsel *evsel,
evsel->sample_size = __perf_evsel__sample_size(attr->sample_type);
perf_evsel__calc_id_pos(evsel);
evsel->cmdline_group_boundary = false;
+ evsel->dividedby = NULL;
+ evsel->div_event = NULL;
+ evsel->collect_stat = false;
}
struct perf_evsel *perf_evsel__new_idx(struct perf_event_attr *attr, int idx)
diff --git a/tools/perf/util/evsel.h b/tools/perf/util/evsel.h
index 4e3158fe79c2..15afaf6a393e 100644
--- a/tools/perf/util/evsel.h
+++ b/tools/perf/util/evsel.h
@@ -129,6 +129,9 @@ struct perf_evsel {
struct list_head config_terms;
int bpf_fd;
bool alias;
+ const char * dividedby;
+ struct perf_evsel *div_event;
+ bool collect_stat;
};
union u64_swap {
diff --git a/tools/perf/util/parse-events.c b/tools/perf/util/parse-events.c
index 8b2333278988..59234666b743 100644
--- a/tools/perf/util/parse-events.c
+++ b/tools/perf/util/parse-events.c
@@ -1245,6 +1245,7 @@ int parse_events_add_pmu(struct parse_events_evlist *data,
evsel->scale = info.scale;
evsel->per_pkg = info.per_pkg;
evsel->snapshot = info.snapshot;
+ evsel->dividedby = info.dividedby;
}
return evsel ? 0 : -ENOMEM;
diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
index d298e7413a80..5c30a6ceee0c 100644
--- a/tools/perf/util/pmu.c
+++ b/tools/perf/util/pmu.c
@@ -980,6 +980,7 @@ int perf_pmu__check_alias(struct perf_pmu *pmu, struct list_head *head_terms,
info->unit = NULL;
info->scale = 0.0;
info->snapshot = false;
+ info->dividedby = NULL;
list_for_each_entry_safe(term, h, head_terms, list) {
alias = pmu_find_alias(pmu, term);
@@ -995,6 +996,7 @@ int perf_pmu__check_alias(struct perf_pmu *pmu, struct list_head *head_terms,
if (alias->per_pkg)
info->per_pkg = true;
+ info->dividedby = alias->dividedby;
list_del(&term->list);
free(term);
diff --git a/tools/perf/util/pmu.h b/tools/perf/util/pmu.h
index faf8a7f97d03..5fbdb65064db 100644
--- a/tools/perf/util/pmu.h
+++ b/tools/perf/util/pmu.h
@@ -31,6 +31,7 @@ struct perf_pmu {
struct perf_pmu_info {
const char *unit;
+ const char *dividedby;
double scale;
bool per_pkg;
bool snapshot;
diff --git a/tools/perf/util/stat-shadow.c b/tools/perf/util/stat-shadow.c
index 8a2bbd2a4d82..e1f9d1bc1ef2 100644
--- a/tools/perf/util/stat-shadow.c
+++ b/tools/perf/util/stat-shadow.c
@@ -3,6 +3,8 @@
#include "stat.h"
#include "color.h"
#include "pmu.h"
+#include "rblist.h"
+#include "evlist.h"
enum {
CTX_BIT_USER = 1 << 0,
@@ -41,13 +43,73 @@ static struct stats runtime_topdown_slots_issued[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_topdown_slots_retired[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_topdown_fetch_bubbles[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_topdown_recovery_bubbles[NUM_CTX][MAX_NR_CPUS];
+static struct rblist runtime_saved_values;
static bool have_frontend_stalled;
struct stats walltime_nsecs_stats;
+struct saved_value {
+ struct rb_node rb_node;
+ struct perf_evsel *evsel;
+ int cpu;
+ int ctx;
+ struct stats stats;
+};
+
+static int saved_value_cmp(struct rb_node *rb_node, const void *entry)
+{
+ struct saved_value *a = container_of(rb_node,
+ struct saved_value,
+ rb_node);
+ const struct saved_value *b = entry;
+
+ if (a->ctx != b->ctx)
+ return a->ctx - b->ctx;
+ if (a->cpu != b->cpu)
+ return a->cpu - b->cpu;
+ return a->evsel - b->evsel;
+}
+
+static struct rb_node *saved_value_new(struct rblist *rblist __maybe_unused,
+ const void *entry)
+{
+ struct saved_value *nd = malloc(sizeof(struct saved_value));
+
+ if (!nd)
+ return NULL;
+ memcpy(nd, entry, sizeof(struct saved_value));
+ return &nd->rb_node;
+}
+
+static struct saved_value *saved_value_lookup(struct perf_evsel *evsel,
+ int cpu, int ctx,
+ bool create)
+{
+ struct rb_node *nd;
+ struct saved_value dm = {
+ .cpu = cpu,
+ .ctx = ctx,
+ .evsel = evsel,
+ };
+ nd = rblist__find(&runtime_saved_values, &dm);
+ if (nd)
+ return container_of(nd, struct saved_value, rb_node);
+ if (create) {
+ rblist__add_node(&runtime_saved_values, &dm);
+ nd = rblist__find(&runtime_saved_values, &dm);
+ if (nd)
+ return container_of(nd, struct saved_value, rb_node);
+ }
+ return NULL;
+}
+
void perf_stat__init_shadow_stats(void)
{
have_frontend_stalled = pmu_have_event("cpu", "stalled-cycles-frontend");
+ rblist__init(&runtime_saved_values);
+ runtime_saved_values.node_cmp = saved_value_cmp;
+ runtime_saved_values.node_new = saved_value_new;
+ /* No delete for now */
}
static int evsel_context(struct perf_evsel *evsel)
@@ -70,6 +132,8 @@ static int evsel_context(struct perf_evsel *evsel)
void perf_stat__reset_shadow_stats(void)
{
+ struct rb_node *pos, *next;
+
memset(runtime_nsecs_stats, 0, sizeof(runtime_nsecs_stats));
memset(runtime_cycles_stats, 0, sizeof(runtime_cycles_stats));
memset(runtime_stalled_cycles_front_stats, 0, sizeof(runtime_stalled_cycles_front_stats));
@@ -92,6 +156,15 @@ void perf_stat__reset_shadow_stats(void)
memset(runtime_topdown_slots_issued, 0, sizeof(runtime_topdown_slots_issued));
memset(runtime_topdown_fetch_bubbles, 0, sizeof(runtime_topdown_fetch_bubbles));
memset(runtime_topdown_recovery_bubbles, 0, sizeof(runtime_topdown_recovery_bubbles));
+
+ next = rb_first(&runtime_saved_values.entries);
+ while (next) {
+ pos = next;
+ next = rb_next(pos);
+ memset(&container_of(pos, struct saved_value, rb_node)->stats,
+ 0,
+ sizeof(struct stats));
+ }
}
/*
@@ -143,6 +216,12 @@ void perf_stat__update_shadow_stats(struct perf_evsel *counter, u64 *count,
update_stats(&runtime_dtlb_cache_stats[ctx][cpu], count[0]);
else if (perf_evsel__match(counter, HW_CACHE, HW_CACHE_ITLB))
update_stats(&runtime_itlb_cache_stats[ctx][cpu], count[0]);
+
+ if (counter->collect_stat) {
+ struct saved_value *v = saved_value_lookup(counter, cpu, ctx,
+ true);
+ update_stats(&v->stats, count[0]);
+ }
}
/* used for get_ratio_color() */
@@ -172,6 +251,55 @@ static const char *get_ratio_color(enum grc_type type, double ratio)
return color;
}
+static struct perf_evsel *perf_stat__find_event(struct perf_evlist *evsel_list,
+ const char *name)
+{
+ struct perf_evsel *c2;
+
+ evlist__for_each_entry (evsel_list, c2) {
+ if (!strcasecmp(c2->name, name))
+ return c2;
+ }
+ return NULL;
+}
+
+/* Mark DividedBy target events and link events using them to them. */
+void perf_stat__collect_dividedby(struct perf_evlist *evsel_list)
+{
+ struct perf_evsel *counter, *leader, *c2;
+ bool found;
+
+ evlist__for_each_entry(evsel_list, counter) {
+ leader = counter->leader;
+ if (!counter->dividedby)
+ continue;
+ found = false;
+ if (leader) {
+ /* Search in group */
+ for_each_group_member (c2, leader) {
+ if (!strcasecmp(c2->name, counter->dividedby)) {
+ found = true;
+ break;
+ }
+ }
+ }
+ if (!found) {
+ /* Search ignoring groups */
+ c2 = perf_stat__find_event(evsel_list, counter->dividedby);
+ }
+ if (!c2) {
+ /* Could try to automatically add the event here. */
+ fprintf(stderr, "Add %s to groups to get ratios for %s\n",
+ counter->dividedby,
+ counter->name);
+ counter->dividedby = NULL;
+ continue;
+ }
+ counter->div_event = c2;
+ c2->collect_stat = true;
+ }
+}
+
static void print_stalled_cycles_frontend(int cpu,
struct perf_evsel *evsel, double avg,
struct perf_stat_output_ctx *out)
@@ -614,6 +742,17 @@ void perf_stat__print_shadow_stats(struct perf_evsel *evsel,
be_bound * 100.);
else
print_metric(ctxp, NULL, NULL, name, 0);
+ } else if (evsel->dividedby) {
+ struct saved_value *v = saved_value_lookup(evsel->div_event, cpu, ctx,
+ false);
+ if (v) {
+ total = avg_stats(&v->stats);
+ if (total)
+ ratio = avg / total;
+ print_metric(ctxp, NULL, "%8.1f%%",
+ out->force_header ? evsel->name : "",
+ ratio * 100.);
+ }
} else if (runtime_nsecs_stats[cpu].n != 0) {
char unit = 'M';
char unit_buf[10];
diff --git a/tools/perf/util/stat.h b/tools/perf/util/stat.h
index c29bb94c48a4..d79e03ccd644 100644
--- a/tools/perf/util/stat.h
+++ b/tools/perf/util/stat.h
@@ -85,11 +85,13 @@ struct perf_stat_output_ctx {
void *ctx;
print_metric_t print_metric;
new_line_t new_line;
+ bool force_header;
};
void perf_stat__print_shadow_stats(struct perf_evsel *evsel,
double avg, int cpu,
struct perf_stat_output_ctx *out);
+void perf_stat__collect_dividedby(struct perf_evlist *);
int perf_evlist__alloc_stats(struct perf_evlist *evlist, bool alloc_raw);
void perf_evlist__free_stats(struct perf_evlist *evlist);
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-10-13 23:20 +0200 |
| Subject | [PATCH 06/10] perf, tools: Add debug support for outputing alias string |
| Message-ID | <srVzY-qJ-15@gated-at.bofh.it> |
| In reply to | #1500566 |
From: Andi Kleen <ak@linux.intel.com>
For debugging and testing it is useful to see the converted
alias string. Add support to perf stat/record and perf list to print
the alias conversion. The text string is saved in the alias structure.
For perf stat/record it is folded into the normal -v. For perf list
-v was taken, so we use --debug.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/builtin-list.c | 3 +++
tools/perf/util/parse-events.y | 3 +++
tools/perf/util/pmu.c | 8 ++++++++
tools/perf/util/pmu.h | 1 +
4 files changed, 15 insertions(+)
diff --git a/tools/perf/builtin-list.c b/tools/perf/builtin-list.c
index ba9322ff858b..3b9d98b5feef 100644
--- a/tools/perf/builtin-list.c
+++ b/tools/perf/builtin-list.c
@@ -14,6 +14,7 @@
#include "util/parse-events.h"
#include "util/cache.h"
#include "util/pmu.h"
+#include "util/debug.h"
#include <subcmd/parse-options.h>
static bool desc_flag = true;
@@ -29,6 +30,8 @@ int cmd_list(int argc, const char **argv, const char *prefix __maybe_unused)
"Print extra event descriptions. --no-desc to not print."),
OPT_BOOLEAN('v', "long-desc", &long_desc_flag,
"Print longer event descriptions."),
+ OPT_INCR(0, "debug", &verbose,
+ "Enable debugging output"),
OPT_END()
};
const char * const list_usage[] = {
diff --git a/tools/perf/util/parse-events.y b/tools/perf/util/parse-events.y
index f3b5ec901600..3a5196380609 100644
--- a/tools/perf/util/parse-events.y
+++ b/tools/perf/util/parse-events.y
@@ -13,6 +13,7 @@
#include <linux/types.h>
#include "util.h"
#include "pmu.h"
+#include "debug.h"
#include "parse-events.h"
#include "parse-events-bison.h"
@@ -254,6 +255,8 @@ PE_KERNEL_PMU_EVENT sep_dc
if (!parse_events_add_pmu(data, list,
pmu->name, head)) {
+ pr_debug("%s -> %s/%s/\n", $1,
+ pmu->name, alias->str);
ok++;
}
diff --git a/tools/perf/util/pmu.c b/tools/perf/util/pmu.c
index f8a052a793b1..dc93c7d4a799 100644
--- a/tools/perf/util/pmu.c
+++ b/tools/perf/util/pmu.c
@@ -272,6 +272,8 @@ static int __perf_pmu__new_alias(struct list_head *list, char *dir, char *name,
snprintf(alias->unit, sizeof(alias->unit), "%s", unit);
}
alias->per_pkg = perpkg && sscanf(perpkg, "%d", &num) == 1 && num == 1;
+ alias->str = strdup(val);
+
list_add_tail(&alias->list, list);
return 0;
@@ -1082,6 +1084,8 @@ struct sevent {
char *name;
char *desc;
char *topic;
+ char *str;
+ char *pmu;
};
static int cmp_sevent(const void *a, const void *b)
@@ -1176,6 +1180,8 @@ void print_pmu_events(const char *event_glob, bool name_only, bool quiet_flag,
aliases[j].desc = long_desc ? alias->long_desc :
alias->desc;
aliases[j].topic = alias->topic;
+ aliases[j].str = alias->str;
+ aliases[j].pmu = pmu->name;
j++;
}
if (pmu->selectable &&
@@ -1210,6 +1216,8 @@ void print_pmu_events(const char *event_glob, bool name_only, bool quiet_flag,
printf("%*s", 8, "[");
wordwrap(aliases[j].desc, 8, columns, 0);
printf("]\n");
+ if (verbose)
+ printf("%*s%s/%s/\n", 8, "", aliases[j].pmu, aliases[j].str);
} else
printf(" %-50s [Kernel PMU event]\n", aliases[j].name);
printed++;
diff --git a/tools/perf/util/pmu.h b/tools/perf/util/pmu.h
index 25712034c815..00852ddc7741 100644
--- a/tools/perf/util/pmu.h
+++ b/tools/perf/util/pmu.h
@@ -43,6 +43,7 @@ struct perf_pmu_alias {
char *desc;
char *long_desc;
char *topic;
+ char *str;
struct list_head terms; /* HEAD struct parse_events_term -> list */
struct list_head list; /* ELEM */
char unit[UNIT_MAX_LEN+1];
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-10-17 13:10 +0200 |
| Message-ID | <stdXP-2H4-17@gated-at.bofh.it> |
| In reply to | #1500566 |
On Thu, Oct 13, 2016 at 02:15:22PM -0700, Andi Kleen wrote: > This adds uncore support on top of the recently merged JSON event list > infrastructure for core events. Uncore is everything outside the core, > including memory controllers, PCI, interconnect etc. > > Uncore is more complicated to handle than core events because it uses > many duplicated PMUs, which leads to long event lists and verbose duplicated > outputs. > > In fact previously it was nearly unusable for many cases without special > tools to generate event list and aggregate data (such as > https://github.com/andikleen/pmu-tools/tree/master/ucevent) > > With this patchkit we add: > - Basic support for uncore events in JSON events > - Support aliases that get duplicated over many PMUs transparently > - Support summing up duplicated PMUs per socket > - Support extending the perf stat builtin metrics with simple ratios > specified in the event list. This covers the vast majority of useful > metrics. > > So far mainly servers are supported. Also this is not using full event lists > (which are full of very obscure events) but only for a smaller subset of > curated useful and understandable metrics. > > The actual event lists are not posted, but available at > git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc perf/intel-uncore-json-files-1 > > The code is available here > git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc perf/builtin-json-15 perf test 5 is failing [jolsa@krava perf]$ sudo ./perf test 5 -v ... mem-loads -> cpu/event=0xcd,umask=0x1,ldlat=3/ failed to parse event 'mem-snp-hit:u,cpu/event=mem-snp-hit/u', err 1 test child finished with 1 ---- end ---- parse events tests: FAILED! jirka
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web