Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1292670 > unrolled thread
| Started by | Andi Kleen <andi@firstfloor.org> |
|---|---|
| First post | 2015-12-16 02:00 +0100 |
| Last post | 2015-12-18 22:40 +0100 |
| Articles | 9 on this page of 29 — 4 participants |
Back to article view | Back to linux.kernel
Add top down metrics to perf stat v2 Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
[PATCH 07/10] perf, tools, stat: Add extra output of counter values with -v Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
[PATCH 10/10] x86, perf: Add Top Down events to Intel Atom Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
[PATCH 04/10] perf, tools, stat: Scale values by unit before metrics Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
[PATCH 09/10] x86, perf: Add Top Down events to Intel Core Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
[PATCH 03/10] perf, tools, stat: Avoid fractional digits for integer scales Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: [PATCH 03/10] perf, tools, stat: Avoid fractional digits for integer scales Stephane Eranian <eranian@google.com> - 2015-12-16 15:30 +0100
[PATCH 05/10] perf, tools, stat: Basic support for TopDown in perf stat Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: [PATCH 05/10] perf, tools, stat: Basic support for TopDown in perf stat Stephane Eranian <eranian@google.com> - 2015-12-16 15:40 +0100
Re: [PATCH 05/10] perf, tools, stat: Basic support for TopDown in perf stat Andi Kleen <andi@firstfloor.org> - 2015-12-16 22:30 +0100
Re: [PATCH 05/10] perf, tools, stat: Basic support for TopDown in perf stat Stephane Eranian <eranian@google.com> - 2015-12-17 10:30 +0100
Re: [PATCH 05/10] perf, tools, stat: Basic support for TopDown in perf stat Andi Kleen <andi@firstfloor.org> - 2015-12-17 15:10 +0100
[PATCH 08/10] x86, perf: Support sysfs files depending on SMT status Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: [PATCH 08/10] x86, perf: Support sysfs files depending on SMT status Peter Zijlstra <peterz@infradead.org> - 2015-12-16 10:00 +0100
Re: [PATCH 08/10] x86, perf: Support sysfs files depending on SMT status Stephane Eranian <eranian@google.com> - 2015-12-16 13:50 +0100
Re: [PATCH 08/10] x86, perf: Support sysfs files depending on SMT status Andi Kleen <andi@firstfloor.org> - 2015-12-16 17:30 +0100
[PATCH 01/10] perf, tools: Dont stop PMU parsing on alias parse error Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: [PATCH 01/10] perf, tools: Dont stop PMU parsing on alias parse error Jiri Olsa <jolsa@redhat.com> - 2015-12-21 17:10 +0100
[PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: [PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases Stephane Eranian <eranian@google.com> - 2015-12-16 15:20 +0100
Re: [PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases Jiri Olsa <jolsa@redhat.com> - 2015-12-21 17:20 +0100
Re: [PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases Jiri Olsa <jolsa@redhat.com> - 2015-12-21 17:20 +0100
[PATCH 06/10] perf, tools, stat: Add computation of TopDown formulas Andi Kleen <andi@firstfloor.org> - 2015-12-16 02:00 +0100
Re: Add top down metrics to perf stat v2 Stephane Eranian <eranian@google.com> - 2015-12-17 11:30 +0100
Re: Add top down metrics to perf stat v2 Andi Kleen <andi@firstfloor.org> - 2015-12-17 15:10 +0100
Re: Add top down metrics to perf stat v2 Stephane Eranian <eranian@google.com> - 2015-12-18 00:40 +0100
Re: Add top down metrics to perf stat v2 Andi Kleen <andi@firstfloor.org> - 2015-12-18 03:00 +0100
Re: Add top down metrics to perf stat v2 Stephane Eranian <eranian@google.com> - 2015-12-18 10:40 +0100
Re: Add top down metrics to perf stat v2 Andi Kleen <andi@firstfloor.org> - 2015-12-18 22:40 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2015-12-21 17:20 +0100 |
| Subject | Re: [PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases |
| Message-ID | <qIblM-jv-3@gated-at.bofh.it> |
| In reply to | #1292683 |
On Tue, Dec 15, 2015 at 04:54:18PM -0800, Andi Kleen wrote: > From: Andi Kleen <ak@linux.intel.com> > > When an event alias is used that the kernel marked as .agg-per-core, force > --per-core mode (and also require -a and forbid cgroups or per thread mode). > This in term means, --topdown forces --per-core mode. > > This is needed for TopDown in SMT mode, because it needs to measure > all threads in a core together and merge the values to compute the correct > percentages of how the pipeline is limited. > > We do this if any alias is agg-per-core. > > Add the code to parse the .agg-per-core attributes and propagate > the information to the evsel. Then the main stat code does > the necessary checks and forces per core mode. > > Open issue: in combination with -C ... we get wrong values. I think that's > a existing bug that needs to be debugged/fixed separately. could you please be more specific on what's failing for you? thanks, jirka -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2015-12-21 17:20 +0100 |
| Subject | Re: [PATCH 02/10] perf, tools, stat: Force --per-core mode for .agg-per-core aliases |
| Message-ID | <qIblN-jv-21@gated-at.bofh.it> |
| In reply to | #1292683 |
On Tue, Dec 15, 2015 at 04:54:18PM -0800, Andi Kleen wrote: > From: Andi Kleen <ak@linux.intel.com> > > When an event alias is used that the kernel marked as .agg-per-core, force > --per-core mode (and also require -a and forbid cgroups or per thread mode). > This in term means, --topdown forces --per-core mode. > > This is needed for TopDown in SMT mode, because it needs to measure > all threads in a core together and merge the values to compute the correct > percentages of how the pipeline is limited. > > We do this if any alias is agg-per-core. > > Add the code to parse the .agg-per-core attributes and propagate > the information to the evsel. Then the main stat code does > the necessary checks and forces per core mode. please split it into 2 patches: - support .aggr-per-core alias parsing - force --per-core mode for .agg-per-core aliases thanks, jirka -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2015-12-16 02:00 +0100 |
| Subject | [PATCH 06/10] perf, tools, stat: Add computation of TopDown formulas |
| Message-ID | <qG8BK-3aD-35@gated-at.bofh.it> |
| In reply to | #1292670 |
From: Andi Kleen <ak@linux.intel.com>
Implement the TopDown formulas in perf stat. The topdown basic metrics
reported by the kernel are collected, and the formulas are computed
and output as normal metrics.
See the kernel commit exporting the events for details on the used
metrics.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/builtin-stat.c | 6 +-
tools/perf/util/stat-shadow.c | 164 +++++++++++++++++++++++++++++++++++++++++-
tools/perf/util/stat.c | 5 ++
tools/perf/util/stat.h | 8 ++-
4 files changed, 179 insertions(+), 4 deletions(-)
diff --git a/tools/perf/builtin-stat.c b/tools/perf/builtin-stat.c
index e10198c..6c44aae 100644
--- a/tools/perf/builtin-stat.c
+++ b/tools/perf/builtin-stat.c
@@ -859,7 +859,8 @@ static void printout(int id, int nr, struct perf_evsel *counter, double uval,
perf_stat__print_shadow_stats(counter, uval,
stat_config.aggr_mode == AGGR_GLOBAL ? 0 :
first_shadow_cpu(counter, id),
- &out);
+ &out,
+ topdown_run);
if (!metric_only) {
print_noise(counter, noise);
@@ -1048,7 +1049,8 @@ static void print_metric_headers(char *prefix)
os.evsel = counter;
perf_stat__print_shadow_stats(counter, 0,
0,
- &out);
+ &out,
+ topdown_run);
}
fputc('\n', stat_config.output);
}
diff --git a/tools/perf/util/stat-shadow.c b/tools/perf/util/stat-shadow.c
index 4d8f185..e977992 100644
--- a/tools/perf/util/stat-shadow.c
+++ b/tools/perf/util/stat-shadow.c
@@ -2,6 +2,7 @@
#include "evsel.h"
#include "stat.h"
#include "color.h"
+#include "pmu.h"
enum {
CTX_BIT_USER = 1 << 0,
@@ -28,6 +29,11 @@ static struct stats runtime_dtlb_cache_stats[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_cycles_in_tx_stats[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_transaction_stats[NUM_CTX][MAX_NR_CPUS];
static struct stats runtime_elision_stats[NUM_CTX][MAX_NR_CPUS];
+static struct stats runtime_topdown_total_slots[NUM_CTX][MAX_NR_CPUS];
+static struct stats runtime_topdown_slots_issued[NUM_CTX][MAX_NR_CPUS];
+static struct stats runtime_topdown_slots_retired[NUM_CTX][MAX_NR_CPUS];
+static struct stats runtime_topdown_fetch_bubbles[NUM_CTX][MAX_NR_CPUS];
+static struct stats runtime_topdown_recovery_bubbles[NUM_CTX][MAX_NR_CPUS];
struct stats walltime_nsecs_stats;
@@ -68,6 +74,11 @@ void perf_stat__reset_shadow_stats(void)
sizeof(runtime_transaction_stats));
memset(runtime_elision_stats, 0, sizeof(runtime_elision_stats));
memset(&walltime_nsecs_stats, 0, sizeof(walltime_nsecs_stats));
+ memset(runtime_topdown_total_slots, 0, sizeof(runtime_topdown_total_slots));
+ memset(runtime_topdown_slots_retired, 0, sizeof(runtime_topdown_slots_retired));
+ memset(runtime_topdown_slots_issued, 0, sizeof(runtime_topdown_slots_issued));
+ memset(runtime_topdown_fetch_bubbles, 0, sizeof(runtime_topdown_fetch_bubbles));
+ memset(runtime_topdown_recovery_bubbles, 0, sizeof(runtime_topdown_recovery_bubbles));
}
/*
@@ -90,6 +101,16 @@ void perf_stat__update_shadow_stats(struct perf_evsel *counter, u64 *count,
update_stats(&runtime_transaction_stats[ctx][cpu], count[0]);
else if (perf_stat_evsel__is(counter, ELISION_START))
update_stats(&runtime_elision_stats[ctx][cpu], count[0]);
+ else if (perf_stat_evsel__is(counter, TOPDOWN_TOTAL_SLOTS))
+ update_stats(&runtime_topdown_total_slots[ctx][cpu], count[0]);
+ else if (perf_stat_evsel__is(counter, TOPDOWN_SLOTS_ISSUED))
+ update_stats(&runtime_topdown_slots_issued[ctx][cpu], count[0]);
+ else if (perf_stat_evsel__is(counter, TOPDOWN_SLOTS_RETIRED))
+ update_stats(&runtime_topdown_slots_retired[ctx][cpu], count[0]);
+ else if (perf_stat_evsel__is(counter, TOPDOWN_FETCH_BUBBLES))
+ update_stats(&runtime_topdown_fetch_bubbles[ctx][cpu],count[0]);
+ else if (perf_stat_evsel__is(counter, TOPDOWN_RECOVERY_BUBBLES))
+ update_stats(&runtime_topdown_recovery_bubbles[ctx][cpu], count[0]);
else if (perf_evsel__match(counter, HARDWARE, HW_STALLED_CYCLES_FRONTEND))
update_stats(&runtime_stalled_cycles_front_stats[ctx][cpu], count[0]);
else if (perf_evsel__match(counter, HARDWARE, HW_STALLED_CYCLES_BACKEND))
@@ -289,9 +310,108 @@ static void print_ll_cache_misses(int cpu,
out->print_metric(out->ctx, color, "%7.2f%%", "of all LL-cache hits", ratio);
}
+/*
+ * High level "TopDown" CPU core pipe line bottleneck break down.
+ *
+ * Basic concept following
+ * Yasin, A Top Down Method for Performance analysis and Counter architecture
+ * ISPASS14
+ *
+ * The CPU pipeline is divided into 4 areas that can be bottlenecks:
+ *
+ * Frontend -> Backend -> Retiring
+ * BadSpeculation in addition means out of order execution that is thrown away
+ * (for example branch mispredictions)
+ * Frontend is instruction decoding.
+ * Backend is execution, like computation and accessing data in memory
+ * Retiring is good execution that is not directly bottlenecked
+ *
+ * The formulas are computed in slots.
+ * A slot is an entry in the pipeline each for the pipeline width
+ * (for example a 4-wide pipeline has 4 slots for each cycle)
+ *
+ * Formulas:
+ * BadSpeculation = ((SlotsIssued - SlotsRetired) + RecoveryBubbles) /
+ * TotalSlots
+ * Retiring = SlotsRetired / TotalSlots
+ * FrontendBound = FetchBubbles / TotalSlots
+ * BackendBound = 1.0 - BadSpeculation - Retiring - FrontendBound
+ *
+ * The kernel provides the mapping to the low level CPU events and any scaling
+ * needed for the CPU pipeline width, for example:
+ *
+ * TotalSlots = Cycles * 4
+ *
+ * The scaling factor is communicated in the sysfs unit.
+ *
+ * In some cases the CPU may not be able to measure all the formulas due to
+ * missing events. In this case multiple formulas are combined, as possible.
+ *
+ * With SMT the slots of thread siblings need to be combined to get meaningful
+ * results. This is implemented by the kernel forcing per-core mode with
+ * the .agg-per-core sysfs attribute.
+ *
+ * Full TopDown supports more levels to sub-divide each area: for example
+ * BackendBound into computing bound and memory bound. For now we only
+ * support Level 1 TopDown.
+ */
+
+static double td_total_slots(int ctx, int cpu)
+{
+ return avg_stats(&runtime_topdown_total_slots[ctx][cpu]);
+}
+
+static double td_bad_spec(int ctx, int cpu)
+{
+ double bad_spec = 0;
+ double total_slots;
+ double total;
+
+ total = avg_stats(&runtime_topdown_slots_issued[ctx][cpu]) -
+ avg_stats(&runtime_topdown_slots_retired[ctx][cpu]) +
+ avg_stats(&runtime_topdown_recovery_bubbles[ctx][cpu]);
+ total_slots = td_total_slots(ctx, cpu);
+ if (total_slots)
+ bad_spec = total / total_slots;
+ return bad_spec;
+}
+
+static double td_retiring(int ctx, int cpu)
+{
+ double retiring = 0;
+ double total_slots = td_total_slots(ctx, cpu);
+ double ret_slots = avg_stats(&runtime_topdown_slots_retired[ctx][cpu]);
+
+ if (total_slots)
+ retiring = ret_slots / total_slots;
+ return retiring;
+}
+
+static double td_fe_bound(int ctx, int cpu)
+{
+ double fe_bound = 0;
+ double total_slots = td_total_slots(ctx, cpu);
+ double fetch_bub = avg_stats(&runtime_topdown_fetch_bubbles[ctx][cpu]);
+
+ if (total_slots)
+ fe_bound = fetch_bub / total_slots;
+ return fe_bound;
+}
+
+static double td_be_bound(int ctx, int cpu)
+{
+ double sum = (td_fe_bound(ctx, cpu) +
+ td_bad_spec(ctx, cpu) +
+ td_retiring(ctx, cpu));
+ if (sum == 0)
+ return 0;
+ return 1.0 - sum;
+}
+
void perf_stat__print_shadow_stats(struct perf_evsel *evsel,
double avg, int cpu,
- struct perf_stat_output_ctx *out)
+ struct perf_stat_output_ctx *out,
+ int topdown_run)
{
void *ctxp = out->ctx;
print_metric_t print_metric = out->print_metric;
@@ -438,6 +558,48 @@ void perf_stat__print_shadow_stats(struct perf_evsel *evsel,
avg / ratio);
else
print_metric(ctxp, NULL, NULL, "CPUs utilized", 0);
+ } else if (perf_stat_evsel__is(evsel, TOPDOWN_FETCH_BUBBLES)) {
+ double fe_bound = td_fe_bound(ctx, cpu);
+
+ if (fe_bound > 0.2 || topdown_run > 1)
+ print_metric(ctxp, NULL, "%8.2f%%", "frontend bound",
+ fe_bound * 100.);
+ else
+ print_metric(ctxp, NULL, NULL, "frontend bound", 0);
+ } else if (perf_stat_evsel__is(evsel, TOPDOWN_SLOTS_RETIRED)) {
+ double retiring = td_retiring(ctx, cpu);
+
+ if (retiring > 0.7 || topdown_run > 1)
+ print_metric(ctxp, NULL, "%8.2f%%", "retiring",
+ retiring * 100.);
+ else
+ print_metric(ctxp, NULL, NULL, "retiring", 0);
+ } else if (perf_stat_evsel__is(evsel, TOPDOWN_RECOVERY_BUBBLES)) {
+ double bad_spec = td_bad_spec(ctx, cpu);
+
+ if (bad_spec > 0.1 || topdown_run > 1)
+ print_metric(ctxp, NULL, "%8.2f%%", "bad speculation",
+ bad_spec * 100.);
+ else
+ print_metric(ctxp, NULL, NULL, "bad speculation", 0);
+ } else if (perf_stat_evsel__is(evsel, TOPDOWN_SLOTS_ISSUED)) {
+ double be_bound = td_be_bound(ctx, cpu);
+ const char *name = "backend bound";
+ static int have_recovery_bubbles = -1;
+
+ /* In case the CPU does not support topdown-recovery-bubbles */
+ if (have_recovery_bubbles < 0)
+ have_recovery_bubbles = pmu_have_event("cpu",
+ "topdown-recovery-bubbles");
+ if (!have_recovery_bubbles)
+ name = "backend bound/bad spec";
+
+ if (td_total_slots(ctx, cpu) > 0 &&
+ (be_bound > 0.2 || topdown_run > 1))
+ print_metric(ctxp, NULL, "%8.2f%%", name,
+ be_bound * 100.);
+ else
+ print_metric(ctxp, NULL, NULL, name, 0);
} else if (runtime_nsecs_stats[cpu].n != 0) {
char unit = 'M';
char unit_buf[10];
diff --git a/tools/perf/util/stat.c b/tools/perf/util/stat.c
index 913d236..fdb34e3 100644
--- a/tools/perf/util/stat.c
+++ b/tools/perf/util/stat.c
@@ -79,6 +79,11 @@ static const char *id_str[PERF_STAT_EVSEL_ID__MAX] = {
ID(TRANSACTION_START, cpu/tx-start/),
ID(ELISION_START, cpu/el-start/),
ID(CYCLES_IN_TX_CP, cpu/cycles-ct/),
+ ID(TOPDOWN_TOTAL_SLOTS, topdown-total-slots),
+ ID(TOPDOWN_SLOTS_ISSUED, topdown-slots-issued),
+ ID(TOPDOWN_SLOTS_RETIRED, topdown-slots-retired),
+ ID(TOPDOWN_FETCH_BUBBLES, topdown-fetch-bubbles),
+ ID(TOPDOWN_RECOVERY_BUBBLES, topdown-recovery-bubbles),
};
#undef ID
diff --git a/tools/perf/util/stat.h b/tools/perf/util/stat.h
index f51d94e..0c26633 100644
--- a/tools/perf/util/stat.h
+++ b/tools/perf/util/stat.h
@@ -17,6 +17,11 @@ enum perf_stat_evsel_id {
PERF_STAT_EVSEL_ID__TRANSACTION_START,
PERF_STAT_EVSEL_ID__ELISION_START,
PERF_STAT_EVSEL_ID__CYCLES_IN_TX_CP,
+ PERF_STAT_EVSEL_ID__TOPDOWN_TOTAL_SLOTS,
+ PERF_STAT_EVSEL_ID__TOPDOWN_SLOTS_ISSUED,
+ PERF_STAT_EVSEL_ID__TOPDOWN_SLOTS_RETIRED,
+ PERF_STAT_EVSEL_ID__TOPDOWN_FETCH_BUBBLES,
+ PERF_STAT_EVSEL_ID__TOPDOWN_RECOVERY_BUBBLES,
PERF_STAT_EVSEL_ID__MAX,
};
@@ -83,7 +88,8 @@ struct perf_stat_output_ctx {
void perf_stat__print_shadow_stats(struct perf_evsel *evsel,
double avg, int cpu,
- struct perf_stat_output_ctx *out);
+ struct perf_stat_output_ctx *out,
+ int topdown_run);
void perf_evsel__reset_stat_priv(struct perf_evsel *evsel);
int perf_evsel__alloc_stat_priv(struct perf_evsel *evsel);
--
2.4.3
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2015-12-17 11:30 +0100 |
| Message-ID | <qGDYR-6sp-1@gated-at.bofh.it> |
| In reply to | #1292670 |
Andi, On Tue, Dec 15, 2015 at 4:54 PM, Andi Kleen <andi@firstfloor.org> wrote: > Note to reviewers: includes both tools and kernel patches. > The kernel patches are at the end. > > This patchkit adds support for TopDown measurements to perf stat > It applies on top of my earlier metrics patchkit, posted > separately, and the --metric-only patchkit (also > posted separately) > > TopDown is intended to replace the frontend cycles idle/ > backend cycles idle metrics in standard perf stat output. > These metrics are not reliable in many workloads, > due to out of order effects. > > This implements a new --topdown mode in perf stat > (similar to --transaction) that measures the pipe line > bottlenecks using standardized formulas. The measurement > can be all done with 5 counters (one fixed counter) > > The result are four metrics: > FrontendBound, BackendBound, BadSpeculation, Retiring > > that describe the CPU pipeline behavior on a high level. > > FrontendBound and BackendBound > BadSpeculation is a higher > > The full top down methology has many hierarchical metrics. > This implementation only supports level 1 which can be > collected without multiplexing. A full implementation > of top down on top of perf is available in pmu-tools toplev. > (http://github.com/andikleen/pmu-tools) > > The current version works on Intel Core CPUs starting > with Sandy Bridge, and Atom CPUs starting with Silvermont. > In principle the generic metrics should be also implementable > on other out of order CPUs. > > TopDown level 1 uses a set of abstracted metrics which > are generic to out of order CPU cores (although some > CPUs may not implement all of them): > > topdown-total-slots Available slots in the pipeline > topdown-slots-issued Slots issued into the pipeline > topdown-slots-retired Slots successfully retired > topdown-fetch-bubbles Pipeline gaps in the frontend > topdown-recovery-bubbles Pipeline gaps during recovery > from misspeculation > > These metrics then allow to compute four useful metrics: > FrontendBound, BackendBound, Retiring, BadSpeculation. > > The formulas to compute the metrics are generic, they > only change based on the availability on the abstracted > input values. > > The kernel declares the events supported by the current > CPU and perf stat then computes the formulas based on the > available metrics. > > > Example output: > > $ ./perf stat --topdown -a ./BC1s > > Performance counter stats for 'system wide': > > S0-C0 2 19650790 topdown-total-slots (100.00%) > S0-C0 2 4445680.00 topdown-fetch-bubbles # 22.62% frontend bound (100.00%) > S0-C0 2 1743552.00 topdown-slots-retired (100.00%) > S0-C0 2 622954 topdown-recovery-bubbles (100.00%) > S0-C0 2 2025498.00 topdown-slots-issued # 63.90% backend bound > S0-C1 2 16685216540 topdown-total-slots (100.00%) > S0-C1 2 962557931.00 topdown-fetch-bubbles (100.00%) > S0-C1 2 4175583320.00 topdown-slots-retired (100.00%) > S0-C1 2 1743329246 topdown-recovery-bubbles # 22.22% bad speculation (100.00%) > S0-C1 2 6138901193.50 topdown-slots-issued # 46.99% backend bound > I don't see how this output could be very useful. What matters is the percentage in the comments and not so much the raw counts because what is the unit? Same remark holds for the percentage. I think you need to explain or show that this is % of issue slots and not cycles. > 1.535832673 seconds time elapsed > > $ perf stat --topdown --topdown --metric-only -I 100 ./BC1s When I tried from your git tree the --metric-only option was not recognized. > 0.100576098 frontend bound retiring bad speculation backend bound > 0.100576098 8.83% 48.93% 35.24% 7.00% > 0.200800845 8.84% 48.49% 35.53% 7.13% > 0.300905983 8.73% 48.64% 35.58% 7.05% > ... > This kind of output is more meaningful and clearer for end-users based on my experience and you'd like it per-core possibly. > > > On Hyper Threaded CPUs Top Down computes metrics per core instead of per logical CPU. > In this case perf stat automatically enables --per-core mode and also requires > global mode (-a) and avoiding other filters (no cgroup mode) > > One side effect is that this may require root rights or a > kernel.perf_event_paranoid=-1 setting. > > On systems without Hyper Threading it can be used per process. > > Full tree available in > git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc perf/top-down-11 That is in the top-down-2 branch instead, I think. > > No changelog against previous version. There were lots of changes. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2015-12-17 15:10 +0100 |
| Message-ID | <qGHpM-py-3@gated-at.bofh.it> |
| In reply to | #1293739 |
On Thu, Dec 17, 2015 at 02:27:58AM -0800, Stephane Eranian wrote: > > S0-C1 2 4175583320.00 topdown-slots-retired (100.00%) > > S0-C1 2 1743329246 topdown-recovery-bubbles # 22.22% bad speculation (100.00%) > > S0-C1 2 6138901193.50 topdown-slots-issued # 46.99% backend bound > > > I don't see how this output could be very useful. What matters is the > percentage in the comments > and not so much the raw counts because what is the unit? Same remark > holds for the percentage. > I think you need to explain or show that this is % of issue slots and > not cycles. The events already say slots, not cycles. Except for recovery-bubbles. Could add -slots there too if you think it's helpful, although it would make the name very long and may not fit into the column anymore. > > > 1.535832673 seconds time elapsed > > > > $ perf stat --topdown --topdown --metric-only -I 100 ./BC1s > > When I tried from your git tree the --metric-only option was not recognized. See below. > > > 0.100576098 frontend bound retiring bad speculation backend bound > > 0.100576098 8.83% 48.93% 35.24% 7.00% > > 0.200800845 8.84% 48.49% 35.53% 7.13% > > 0.300905983 8.73% 48.64% 35.58% 7.05% > > ... > > > This kind of output is more meaningful and clearer for end-users based > on my experience > and you'd like it per-core possibly. Yes --metric-only is a lot clearer. per-core is supported and automatically enabled with SMT on. > > Full tree available in > > git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc perf/top-down-11 > > That is in the top-down-2 branch instead, I think. Sorry, typo The correct branch is perf/top-down-10 I also updated it now with the latest review feedback changes. top-down-2 is an really old branch that indeed didn't have metric-only. -Andi -- ak@linux.intel.com -- Speaking for myself only. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2015-12-18 00:40 +0100 |
| Message-ID | <qGQjo-6g4-25@gated-at.bofh.it> |
| In reply to | #1293914 |
On Thu, Dec 17, 2015 at 6:01 AM, Andi Kleen <andi@firstfloor.org> wrote: > On Thu, Dec 17, 2015 at 02:27:58AM -0800, Stephane Eranian wrote: >> > S0-C1 2 4175583320.00 topdown-slots-retired (100.00%) >> > S0-C1 2 1743329246 topdown-recovery-bubbles # 22.22% bad speculation (100.00%) >> > S0-C1 2 6138901193.50 topdown-slots-issued # 46.99% backend bound >> > >> I don't see how this output could be very useful. What matters is the >> percentage in the comments >> and not so much the raw counts because what is the unit? Same remark >> holds for the percentage. >> I think you need to explain or show that this is % of issue slots and >> not cycles. > > The events already say slots, not cycles. Except for recovery-bubbles. Could add > -slots there too if you think it's helpful, although it would make the > name very long and may not fit into the column anymore. > I would drop the default output, it is not useful. I would not add a --topdown option but instead a --metric option with arguments such that other metrics could be added later: $ perf stat --metrics topdown -I 1000 -a sleep 100 If you do this, you do not need the --metric-only option The double --topdown is confusing. Why force --per-core when HT is on. I know you you need to aggregate per core, but you could still display globally. And then if user requests --per-core, then display per core. Same if user specifies --per-socket. I know this requires some more plumbing inside perf but it would be clearer and simpler to interpret to users. One bug I found when testing is that if you do with HT-on: $ perf stat -a --topdown -I 1000 --metric-only sleep 100 Then you get data for frontend and backend but nothing for retiring or bad speculation. I suspect it is because you expect --metric-only to be used only when you have the double --topdown. That's why I think this double topdown is confusing. If you do as I suggest, it will be much simpler. >> >> > 1.535832673 seconds time elapsed >> > >> > $ perf stat --topdown --topdown --metric-only -I 100 ./BC1s >> >> When I tried from your git tree the --metric-only option was not recognized. > > See below. >> >> > 0.100576098 frontend bound retiring bad speculation backend bound >> > 0.100576098 8.83% 48.93% 35.24% 7.00% >> > 0.200800845 8.84% 48.49% 35.53% 7.13% >> > 0.300905983 8.73% 48.64% 35.58% 7.05% >> > ... >> > >> This kind of output is more meaningful and clearer for end-users based >> on my experience >> and you'd like it per-core possibly. > > Yes --metric-only is a lot clearer. > > per-core is supported and automatically enabled with SMT on. > >> > Full tree available in >> > git://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc perf/top-down-11 >> >> That is in the top-down-2 branch instead, I think. > > Sorry, typo > > The correct branch is perf/top-down-10 > > I also updated it now with the latest review feedback changes. > > top-down-2 is an really old branch that indeed didn't have metric-only. > > -Andi > > -- > ak@linux.intel.com -- Speaking for myself only. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2015-12-18 03:00 +0100 |
| Message-ID | <qGSuS-7xL-9@gated-at.bofh.it> |
| In reply to | #1294326 |
Thanks for testing. On Thu, Dec 17, 2015 at 03:31:30PM -0800, Stephane Eranian wrote: > I would not add a --topdown option but instead a --metric option with arguments > such that other metrics could be added later: > > $ perf stat --metrics topdown -I 1000 -a sleep 100 > > If you do this, you do not need the --metric-only option The --metric-only option is useful with other metrics too. For example to get concise (and plottable) IPC or TSX abort statistics. See the examples in the original commit. However could make --topdown default to --metric-only and add an option to turn it off. Yes that's probably a better default for more people, although some people could be annoyed by the wide output. > The double --topdown is confusing. Ok. I was thinking of changing it and adding an extra argument for the "ignore threshold" behavior. That would also make it more extensible if we ever add Level 2. > > Why force --per-core when HT is on. I know you you need to aggregate > per core, but > you could still display globally. And then if user requests > --per-core, then display per core. Global TopDown doesn't make much sense. Suppose you have two programs running on different cores, one frontend bound and one backend bound. What would the union of the two mean? And you may well end up with sums of ratios which are >100%. The only exception where it's useful is for the single threaded case (like the toplev --single-thread) option. However it is something ugly and difficult because the user would need to ensure that there is nothing active on the sibling thread. So I left it out. > Same if user specifies --per-socket. I know this requires some more > plumbing inside perf > but it would be clearer and simpler to interpret to users. Same problem as above. > > One bug I found when testing is that if you do with HT-on: > > $ perf stat -a --topdown -I 1000 --metric-only sleep 100 > Then you get data for frontend and backend but nothing for retiring or > bad speculation. You see all the columns, but no data in some? That's intended: the percentage is only printed when it crosses a threshold. That's part of the top down specificatio. > I suspect it is because you expect --metric-only to be used only when > you have the > double --topdown. That's why I think this double topdown is confusing. If you do > as I suggest, it will be much simpler. It works fine with single topdown as far as I can tell. -Andi -- ak@linux.intel.com -- Speaking for myself only. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Stephane Eranian <eranian@google.com> |
|---|---|
| Date | 2015-12-18 10:40 +0100 |
| Message-ID | <qGZG3-3Pt-31@gated-at.bofh.it> |
| In reply to | #1294392 |
Andi,
On Thu, Dec 17, 2015 at 5:55 PM, Andi Kleen <andi@firstfloor.org> wrote:
> Thanks for testing.
>
>
> On Thu, Dec 17, 2015 at 03:31:30PM -0800, Stephane Eranian wrote:
>> I would not add a --topdown option but instead a --metric option with arguments
>> such that other metrics could be added later:
>>
>> $ perf stat --metrics topdown -I 1000 -a sleep 100
>>
>> If you do this, you do not need the --metric-only option
>
> The --metric-only option is useful with other metrics too. For example
> to get concise (and plottable) IPC or TSX abort statistics. See the
> examples in the original commit.
>
> However could make --topdown default to --metric-only and add
> an option to turn it off. Yes that's probably a better default
> for more people, although some people could be annoyed by the
> wide output.
>
>> The double --topdown is confusing.
>
> Ok. I was thinking of changing it and adding an extra argument for
> the "ignore threshold" behavior. That would also make it more extensible
> if we ever add Level 2.
>
I think you should drop that implicit threshold data suppression
feature altogether.
See what I write below.
>>
>> Why force --per-core when HT is on. I know you you need to aggregate
>> per core, but
>> you could still display globally. And then if user requests
>> --per-core, then display per core.
>
> Global TopDown doesn't make much sense. Suppose you have two programs
> running on different cores, one frontend bound and one backend bound.
> What would the union of the two mean? And you may well end up
> with sums of ratios which are >100%.
>
How could that be if you consider that the machine is N-wide and not just 4-wide
anymore?
How what you are describing here is different when HT is off?
If you force --per-core with HT-on, then you need to force it too when
HT is off so that you get a similar per core breakdown. In the HT on
case, each Sx-Cy represents 2 threads, compared to 1 in the non HT
case.Right now, you have non-HT reporting global, HT reporting per-core.
That does not make much sense to me.
> The only exception where it's useful is for the single threaded
> case (like the toplev --single-thread) option. However it is something
> ugly and difficult because the user would need to ensure that there is
> nothing active on the sibling thread. So I left it out.
>
>> Same if user specifies --per-socket. I know this requires some more
>> plumbing inside perf
>> but it would be clearer and simpler to interpret to users.
>
> Same problem as above.
>
>>
>> One bug I found when testing is that if you do with HT-on:
>>
>> $ perf stat -a --topdown -I 1000 --metric-only sleep 100
>> Then you get data for frontend and backend but nothing for retiring or
>> bad speculation.
>
> You see all the columns, but no data in some?
>
yes, and I don't like that. It is confusing especially when you do not
know the threshold.
Why are you suppressing the 'retiring' data when it is at 25% (1/4 of
the maximum possible)
when I am running a simple noploop? 25% is a sign of underutilization,
that could be useful too.
Furthermore, it makes it harder to parse, including with the -x option
because some fields may
not be there. I would rather see all the values. In non -x mode, you
could use color to indicate
high/low thresholds (similar to perf report).
> That's intended: the percentage is only printed when it crosses a
> threshold. That's part of the top down specification.
>
I don't like that. I would rather see all the percentages.
My remark applies to non topdown metrics as well, such as IPC.
Clearly the IPC is awkward to use. You need to know you need to
measure cycles, instructions to get ipc with --metric-only. Again,
I would rather see:
$ perf stat --metrics ipc ....
$ perf stat --metrics topdown
It is more uniform and users do not have to worry about what events
to use to compute a metric. As an example, here is what you
could do (showing side by side metrics ipc and uops/cycles):
# perf stat -a --metric ipc,upc -I 1000 sleep 100
#======================================
# | ipc| upc
# | IPC %Peak| UPC
# | ^ ^ | ^
#======================================
1.006038929 0.21 5.16% 0.60
2.012169514 0.21 5.31% 0.60
3.018314389 0.20 5.08% 0.59
4.024430081 0.21 5.26% 0.60
>> I suspect it is because you expect --metric-only to be used only when
>> you have the
>> double --topdown. That's why I think this double topdown is confusing. If you do
>> as I suggest, it will be much simpler.
>
> It works fine with single topdown as far as I can tell.
>
>
> -Andi
> --
> ak@linux.intel.com -- Speaking for myself only.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2015-12-18 22:40 +0100 |
| Message-ID | <qHaUO-2Dm-17@gated-at.bofh.it> |
| In reply to | #1294616 |
On Fri, Dec 18, 2015 at 01:31:18AM -0800, Stephane Eranian wrote: > >> Why force --per-core when HT is on. I know you you need to aggregate > >> per core, but > >> you could still display globally. And then if user requests > >> --per-core, then display per core. > > > > Global TopDown doesn't make much sense. Suppose you have two programs > > running on different cores, one frontend bound and one backend bound. > > What would the union of the two mean? And you may well end up > > with sums of ratios which are >100%. > > > How could that be if you consider that the machine is N-wide and not just 4-wide > anymore? > > How what you are describing here is different when HT is off? I was talking about cores, not CPU threads. With global aggregation we would aggregate data from different cores, which is highly dubious for TopDown. CPU threads on a core are of course aggregated, that is why the patchkit forces --per-core with HT on. > If you force --per-core with HT-on, then you need to force it too when > HT is off so that you get a similar per core breakdown. In the HT on > case, each Sx-Cy represents 2 threads, compared to 1 in the non HT > case.Right now, you have non-HT reporting global, HT reporting per-core. > That does not make much sense to me. Ok. I guess can force --per-core in this case too. This would simplify things because can get rid of the agg-per-core attribute. > >> but it would be clearer and simpler to interpret to users. > > > > Same problem as above. > > > >> > >> One bug I found when testing is that if you do with HT-on: > >> > >> $ perf stat -a --topdown -I 1000 --metric-only sleep 100 > >> Then you get data for frontend and backend but nothing for retiring or > >> bad speculation. > > > > You see all the columns, but no data in some? > > > yes, and I don't like that. It is confusing especially when you do not > know the threshold. > Why are you suppressing the 'retiring' data when it is at 25% (1/4 of > the maximum possible) > when I am running a simple noploop? 25% is a sign of underutilization, > that could be useful too. It's what the TopDown specification uses and the paper describes. The thresholds are needed when you have more than one level because the lower levels become meaningless if their parents didn't cross the threshold. Otherwise you may report something that looks like a bottle neck, but isn't. Given there is currently only level 1 in the patchkit, but if we ever add more levels absolutely need thresholds. So it's better to have them from Day 1. Utilization should be reported separately. TopDown cannot give utilization because it doesn't know about idle time. I can report - for empty fields if it helps you. It's not clear to me why empty fields in CSV are a problem. I don't think colors are useful here, this would have the problem described above. > > > That's intended: the percentage is only printed when it crosses a > > threshold. That's part of the top down specification. > > > I don't like that. I would rather see all the percentages. > My remark applies to non topdown metrics as well, such as IPC. > Clearly the IPC is awkward to use. You need to know you need to > measure cycles, instructions to get ipc with --metric-only. Again, Well it's the default (perf stat --metric-only), or with -d*, and it works fine with --transaction too. If you think there should be more predefined sets of metrics that's fine for me too, but it would be a separate patch. -Andi -- ak@linux.intel.com -- Speaking for myself only. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web