Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1528644 > unrolled thread
| Started by | kan.liang@intel.com |
|---|---|
| First post | 2016-11-23 18:50 +0100 |
| Last post | 2016-11-24 05:30 +0100 |
| Articles | 20 on this page of 50 — 8 participants |
Back to article view | Back to linux.kernel
[PATCH 00/14] export perf overheads information kan.liang@intel.com - 2016-11-23 18:50 +0100
[PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Peter Zijlstra <peterz@infradead.org> - 2016-11-23 21:20 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Peter Zijlstra <peterz@infradead.org> - 2016-11-23 21:20 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:50 +0100
RE: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD "Liang, Kan" <kan.liang@intel.com> - 2016-11-24 14:50 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Peter Zijlstra <peterz@infradead.org> - 2016-11-24 15:00 +0100
RE: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD "Liang, Kan" <kan.liang@intel.com> - 2016-11-24 15:10 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Jiri Olsa <jolsa@redhat.com> - 2016-11-24 15:30 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Jiri Olsa <jolsa@redhat.com> - 2016-11-24 15:50 +0100
RE: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD "Liang, Kan" <kan.liang@intel.com> - 2016-11-24 15:50 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Andi Kleen <andi@firstfloor.org> - 2016-11-24 19:30 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Peter Zijlstra <peterz@infradead.org> - 2016-11-24 20:00 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Andi Kleen <andi@firstfloor.org> - 2016-11-24 20:10 +0100
Re: [PATCH 01/14] perf/x86: Introduce PERF_RECORD_OVERHEAD Peter Zijlstra <peterz@infradead.org> - 2016-11-24 20:10 +0100
[PATCH 11/14] perf tools: record write data overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 11/14] perf tools: record write data overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:10 +0100
Re: [PATCH 11/14] perf tools: record write data overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:20 +0100
[PATCH 04/14] perf/x86: output side-band events overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 04/14] perf/x86: output side-band events overhead Peter Zijlstra <peterz@infradead.org> - 2016-11-23 21:10 +0100
Re: [PATCH 04/14] perf/x86: output side-band events overhead Mark Rutland <mark.rutland@arm.com> - 2016-11-24 17:30 +0100
RE: [PATCH 04/14] perf/x86: output side-band events overhead "Liang, Kan" <kan.liang@intel.com> - 2016-11-24 20:50 +0100
[PATCH 07/14] perf tools: show multiplexing overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
[PATCH 08/14] perf tools: show side-band events overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
[PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type Jiri Olsa <jolsa@redhat.com> - 2016-11-23 23:40 +0100
Re: [PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:00 +0100
[PATCH 12/14] perf tools: record elapsed time kan.liang@intel.com - 2016-11-23 18:50 +0100
[PATCH 13/14] perf tools: warn on high overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 13/14] perf tools: warn on high overhead Andi Kleen <andi@firstfloor.org> - 2016-11-23 21:30 +0100
RE: [PATCH 13/14] perf tools: warn on high overhead "Liang, Kan" <kan.liang@intel.com> - 2016-11-23 23:10 +0100
[PATCH 14/14] perf script: show overhead events kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 14/14] perf script: show overhead events Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:30 +0100
Re: [PATCH 14/14] perf script: show overhead events Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:30 +0100
Re: [PATCH 14/14] perf script: show overhead events Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:40 +0100
Re: [PATCH 14/14] perf script: show overhead events Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:40 +0100
[PATCH 03/14] perf/x86: output multiplexing overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 03/14] perf/x86: output multiplexing overhead Peter Zijlstra <peterz@infradead.org> - 2016-11-23 21:10 +0100
RE: [PATCH 03/14] perf/x86: output multiplexing overhead "Liang, Kan" <kan.liang@intel.com> - 2016-11-23 21:20 +0100
[PATCH 10/14] perf tools: introduce PERF_RECORD_USER_OVERHEAD kan.liang@intel.com - 2016-11-23 18:50 +0100
[PATCH 06/14] perf tools: show NMI overhead kan.liang@intel.com - 2016-11-23 18:50 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:00 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:00 +0100
RE: [PATCH 06/14] perf tools: show NMI overhead "Liang, Kan" <kan.liang@intel.com> - 2016-11-24 14:40 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 16:30 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Namhyung Kim <namhyung@kernel.org> - 2016-11-25 00:30 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-25 00:50 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Andi Kleen <andi@firstfloor.org> - 2016-11-25 01:30 +0100
Re: [PATCH 06/14] perf tools: show NMI overhead Jiri Olsa <jolsa@redhat.com> - 2016-11-24 00:00 +0100
Re: [PATCH 00/14] export perf overheads information Ingo Molnar <mingo@kernel.org> - 2016-11-24 05:30 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Mark Rutland <mark.rutland@arm.com> |
|---|---|
| Date | 2016-11-24 17:30 +0100 |
| Subject | Re: [PATCH 04/14] perf/x86: output side-band events overhead |
| Message-ID | <sH54l-5b9-3@gated-at.bofh.it> |
| In reply to | #1528647 |
On Wed, Nov 23, 2016 at 04:44:42AM -0500, kan.liang@intel.com wrote:
> From: Kan Liang <kan.liang@intel.com>
>
> Iterating all events which need to receive side-band events also bring
> some overhead.
> Save the overhead information in task context or CPU context, whichever
> context is available.
Do we really want to expose this concept to userspace?
What if the implementation changes?
Thanks,
Mark.
> Signed-off-by: Kan Liang <kan.liang@intel.com>
> ---
> include/linux/perf_event.h | 2 ++
> include/uapi/linux/perf_event.h | 1 +
> kernel/events/core.c | 32 ++++++++++++++++++++++++++++----
> 3 files changed, 31 insertions(+), 4 deletions(-)
>
> diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
> index f72b97a..ec3cb7f 100644
> --- a/include/linux/perf_event.h
> +++ b/include/linux/perf_event.h
> @@ -764,6 +764,8 @@ struct perf_event_context {
> #endif
> void *task_ctx_data; /* pmu specific data */
> struct rcu_head rcu_head;
> +
> + struct perf_overhead_entry sb_overhead;
> };
>
> /*
> diff --git a/include/uapi/linux/perf_event.h b/include/uapi/linux/perf_event.h
> index 9124c7c..5e7c522 100644
> --- a/include/uapi/linux/perf_event.h
> +++ b/include/uapi/linux/perf_event.h
> @@ -994,6 +994,7 @@ struct perf_branch_entry {
> enum perf_record_overhead_type {
> PERF_NMI_OVERHEAD = 0,
> PERF_MUX_OVERHEAD,
> + PERF_SB_OVERHEAD,
>
> PERF_OVERHEAD_MAX,
> };
> diff --git a/kernel/events/core.c b/kernel/events/core.c
> index 9934059..51e9df7 100644
> --- a/kernel/events/core.c
> +++ b/kernel/events/core.c
> @@ -1829,9 +1829,15 @@ event_sched_out(struct perf_event *event,
> if (event->attr.exclusive || !cpuctx->active_oncpu)
> cpuctx->exclusive = 0;
>
> - if (log_overhead && cpuctx->mux_overhead.nr) {
> - cpuctx->mux_overhead.cpu = smp_processor_id();
> - perf_log_overhead(event, PERF_MUX_OVERHEAD, &cpuctx->mux_overhead);
> + if (log_overhead) {
> + if (cpuctx->mux_overhead.nr) {
> + cpuctx->mux_overhead.cpu = smp_processor_id();
> + perf_log_overhead(event, PERF_MUX_OVERHEAD, &cpuctx->mux_overhead);
> + }
> + if (ctx->sb_overhead.nr) {
> + ctx->sb_overhead.cpu = smp_processor_id();
> + perf_log_overhead(event, PERF_SB_OVERHEAD, &ctx->sb_overhead);
> + }
> }
>
> perf_pmu_enable(event->pmu);
> @@ -6133,6 +6139,14 @@ static void perf_iterate_sb_cpu(perf_iterate_f output, void *data)
> }
> }
>
> +static void
> +perf_caculate_sb_overhead(struct perf_event_context *ctx,
> + u64 time)
> +{
> + ctx->sb_overhead.nr++;
> + ctx->sb_overhead.time += time;
> +}
> +
> /*
> * Iterate all events that need to receive side-band events.
> *
> @@ -6143,9 +6157,12 @@ static void
> perf_iterate_sb(perf_iterate_f output, void *data,
> struct perf_event_context *task_ctx)
> {
> + struct perf_event_context *overhead_ctx = task_ctx;
> struct perf_event_context *ctx;
> + u64 start_clock, end_clock;
> int ctxn;
>
> + start_clock = perf_clock();
> rcu_read_lock();
> preempt_disable();
>
> @@ -6163,12 +6180,19 @@ perf_iterate_sb(perf_iterate_f output, void *data,
>
> for_each_task_context_nr(ctxn) {
> ctx = rcu_dereference(current->perf_event_ctxp[ctxn]);
> - if (ctx)
> + if (ctx) {
> perf_iterate_ctx(ctx, output, data, false);
> + if (!overhead_ctx)
> + overhead_ctx = ctx;
> + }
> }
> done:
> preempt_enable();
> rcu_read_unlock();
> +
> + end_clock = perf_clock();
> + if (overhead_ctx)
> + perf_caculate_sb_overhead(overhead_ctx, end_clock - start_clock);
> }
>
> /*
> --
> 2.5.5
>
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2016-11-24 20:50 +0100 |
| Subject | RE: [PATCH 04/14] perf/x86: output side-band events overhead |
| Message-ID | <sH8bU-76f-33@gated-at.bofh.it> |
| In reply to | #1529539 |
>
> On Wed, Nov 23, 2016 at 04:44:42AM -0500, kan.liang@intel.com wrote:
> > From: Kan Liang <kan.liang@intel.com>
> >
> > Iterating all events which need to receive side-band events also bring
> > some overhead.
> > Save the overhead information in task context or CPU context,
> > whichever context is available.
>
> Do we really want to expose this concept to userspace?
>
> What if the implementation changes?
The concept of side-band will be removed?
I thought we just use the rb-tree to replace the list.
I think no matter how do we implement it, we do need to calculate its
overhead, unless the concept is gone, or it merged with other overhead type.
Because based on my test, it brings big overhead on some cases.
Thanks,
Kan
>
> Thanks,
> Mark.
>
> > Signed-off-by: Kan Liang <kan.liang@intel.com>
> > ---
> > include/linux/perf_event.h | 2 ++
> > include/uapi/linux/perf_event.h | 1 +
> > kernel/events/core.c | 32 ++++++++++++++++++++++++++++----
> > 3 files changed, 31 insertions(+), 4 deletions(-)
> >
> > diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
> > index f72b97a..ec3cb7f 100644
> > --- a/include/linux/perf_event.h
> > +++ b/include/linux/perf_event.h
> > @@ -764,6 +764,8 @@ struct perf_event_context { #endif
> > void *task_ctx_data; /* pmu specific data
> */
> > struct rcu_head rcu_head;
> > +
> > + struct perf_overhead_entry sb_overhead;
> > };
> >
> > /*
> > diff --git a/include/uapi/linux/perf_event.h
> > b/include/uapi/linux/perf_event.h index 9124c7c..5e7c522 100644
> > --- a/include/uapi/linux/perf_event.h
> > +++ b/include/uapi/linux/perf_event.h
> > @@ -994,6 +994,7 @@ struct perf_branch_entry { enum
> > perf_record_overhead_type {
> > PERF_NMI_OVERHEAD = 0,
> > PERF_MUX_OVERHEAD,
> > + PERF_SB_OVERHEAD,
> >
> > PERF_OVERHEAD_MAX,
> > };
> > diff --git a/kernel/events/core.c b/kernel/events/core.c index
> > 9934059..51e9df7 100644
> > --- a/kernel/events/core.c
> > +++ b/kernel/events/core.c
> > @@ -1829,9 +1829,15 @@ event_sched_out(struct perf_event *event,
> > if (event->attr.exclusive || !cpuctx->active_oncpu)
> > cpuctx->exclusive = 0;
> >
> > - if (log_overhead && cpuctx->mux_overhead.nr) {
> > - cpuctx->mux_overhead.cpu = smp_processor_id();
> > - perf_log_overhead(event, PERF_MUX_OVERHEAD, &cpuctx-
> >mux_overhead);
> > + if (log_overhead) {
> > + if (cpuctx->mux_overhead.nr) {
> > + cpuctx->mux_overhead.cpu = smp_processor_id();
> > + perf_log_overhead(event, PERF_MUX_OVERHEAD,
> &cpuctx->mux_overhead);
> > + }
> > + if (ctx->sb_overhead.nr) {
> > + ctx->sb_overhead.cpu = smp_processor_id();
> > + perf_log_overhead(event, PERF_SB_OVERHEAD,
> &ctx->sb_overhead);
> > + }
> > }
> >
> > perf_pmu_enable(event->pmu);
> > @@ -6133,6 +6139,14 @@ static void perf_iterate_sb_cpu(perf_iterate_f
> output, void *data)
> > }
> > }
> >
> > +static void
> > +perf_caculate_sb_overhead(struct perf_event_context *ctx,
> > + u64 time)
> > +{
> > + ctx->sb_overhead.nr++;
> > + ctx->sb_overhead.time += time;
> > +}
> > +
> > /*
> > * Iterate all events that need to receive side-band events.
> > *
> > @@ -6143,9 +6157,12 @@ static void
> > perf_iterate_sb(perf_iterate_f output, void *data,
> > struct perf_event_context *task_ctx) {
> > + struct perf_event_context *overhead_ctx = task_ctx;
> > struct perf_event_context *ctx;
> > + u64 start_clock, end_clock;
> > int ctxn;
> >
> > + start_clock = perf_clock();
> > rcu_read_lock();
> > preempt_disable();
> >
> > @@ -6163,12 +6180,19 @@ perf_iterate_sb(perf_iterate_f output, void
> > *data,
> >
> > for_each_task_context_nr(ctxn) {
> > ctx = rcu_dereference(current->perf_event_ctxp[ctxn]);
> > - if (ctx)
> > + if (ctx) {
> > perf_iterate_ctx(ctx, output, data, false);
> > + if (!overhead_ctx)
> > + overhead_ctx = ctx;
> > + }
> > }
> > done:
> > preempt_enable();
> > rcu_read_unlock();
> > +
> > + end_clock = perf_clock();
> > + if (overhead_ctx)
> > + perf_caculate_sb_overhead(overhead_ctx, end_clock -
> start_clock);
> > }
> >
> > /*
> > --
> > 2.5.5
> >
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 07/14] perf tools: show multiplexing overhead |
| Message-ID | <sGJQe-7SF-21@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
Caculate the total multiplexing overhead on each CPU, and display them
in perf report
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/builtin-report.c | 8 ++++++--
tools/perf/util/event.h | 3 +++
tools/perf/util/machine.c | 5 +++++
tools/perf/util/session.c | 4 ++++
4 files changed, 18 insertions(+), 2 deletions(-)
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index b1437586..2515d7a 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -377,9 +377,13 @@ static int perf_evlist__tty_browse_hists(struct perf_evlist *evlist,
continue;
if (rep->cpu_list && !test_bit(cpu, rep->cpu_bitmap))
continue;
- fprintf(stdout, "#\tCPU %d: NMI#: %" PRIu64 " time: %" PRIu64 " ns\n",
- cpu, evlist->stats.total_nmi_overhead[cpu][0],
+ fprintf(stdout, "#\tCPU %d\n", cpu);
+ fprintf(stdout, "#\t\tNMI#: %" PRIu64 " time: %" PRIu64 " ns\n",
+ evlist->stats.total_nmi_overhead[cpu][0],
evlist->stats.total_nmi_overhead[cpu][1]);
+ fprintf(stdout, "#\t\tMultiplexing#: %" PRIu64 " time: %" PRIu64 " ns\n",
+ evlist->stats.total_mux_overhead[cpu][0],
+ evlist->stats.total_mux_overhead[cpu][1]);
}
fprintf(stdout, "#\n");
}
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index 7d40d54..70e2508 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -265,6 +265,8 @@ enum auxtrace_error_type {
*
* The total_nmi_overhead tells exactly the NMI handler overhead on each CPU.
* The total NMI# is stored in [0], while the accumulated time is in [1].
+ * The total_mux_overhead tells exactly the Multiplexing overhead on each CPU.
+ * The total rotate# is stored in [0], while the accumulated time is in [1].
*/
struct events_stats {
u64 total_period;
@@ -274,6 +276,7 @@ struct events_stats {
u64 total_aux_lost;
u64 total_invalid_chains;
u64 total_nmi_overhead[MAX_NR_CPUS][2];
+ u64 total_mux_overhead[MAX_NR_CPUS][2];
u32 nr_events[PERF_RECORD_HEADER_MAX];
u32 nr_non_filtered_samples;
u32 nr_lost_warned;
diff --git a/tools/perf/util/machine.c b/tools/perf/util/machine.c
index 58076f2..eca1f8b 100644
--- a/tools/perf/util/machine.c
+++ b/tools/perf/util/machine.c
@@ -563,6 +563,11 @@ int machine__process_overhead_event(struct machine *machine __maybe_unused,
event->overhead.entry.nr,
event->overhead.entry.time,
event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_MUX_OVERHEAD) {
+ dump_printf(" Multiplexing nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
} else {
dump_printf("\tUNSUPPORT OVERHEAD TYPE 0x%x!\n", event->overhead.type);
}
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index a79ab99..594fd5e 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -1218,6 +1218,10 @@ overhead_stats_update(struct perf_tool *tool,
evlist->stats.total_nmi_overhead[event->overhead.entry.cpu][0] += event->overhead.entry.nr;
evlist->stats.total_nmi_overhead[event->overhead.entry.cpu][1] += event->overhead.entry.time;
break;
+ case PERF_MUX_OVERHEAD:
+ evlist->stats.total_mux_overhead[event->overhead.entry.cpu][0] += event->overhead.entry.nr;
+ evlist->stats.total_mux_overhead[event->overhead.entry.cpu][1] += event->overhead.entry.time;
+ break;
default:
break;
}
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 08/14] perf tools: show side-band events overhead |
| Message-ID | <sGJQe-7SF-29@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
Caculate the total overhead from accessing side-band events handler
function, and display them in perf report
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/builtin-report.c | 3 +++
tools/perf/util/event.h | 5 +++++
tools/perf/util/machine.c | 5 +++++
tools/perf/util/session.c | 4 ++++
4 files changed, 17 insertions(+)
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 2515d7a..9c0a424 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -384,6 +384,9 @@ static int perf_evlist__tty_browse_hists(struct perf_evlist *evlist,
fprintf(stdout, "#\t\tMultiplexing#: %" PRIu64 " time: %" PRIu64 " ns\n",
evlist->stats.total_mux_overhead[cpu][0],
evlist->stats.total_mux_overhead[cpu][1]);
+ fprintf(stdout, "#\t\tSB#: %" PRIu64 " time: %" PRIu64 " ns\n",
+ evlist->stats.total_sb_overhead[cpu][0],
+ evlist->stats.total_sb_overhead[cpu][1]);
}
fprintf(stdout, "#\n");
}
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index 70e2508..3357529 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -267,6 +267,10 @@ enum auxtrace_error_type {
* The total NMI# is stored in [0], while the accumulated time is in [1].
* The total_mux_overhead tells exactly the Multiplexing overhead on each CPU.
* The total rotate# is stored in [0], while the accumulated time is in [1].
+ * The total_sb_overhead tells exactly the overhead to output side-band
+ * events on each CPU.
+ * The total number of accessing side-band events handler function is stored
+ * in [0], while the accumulated processing time is in [1].
*/
struct events_stats {
u64 total_period;
@@ -277,6 +281,7 @@ struct events_stats {
u64 total_invalid_chains;
u64 total_nmi_overhead[MAX_NR_CPUS][2];
u64 total_mux_overhead[MAX_NR_CPUS][2];
+ u64 total_sb_overhead[MAX_NR_CPUS][2];
u32 nr_events[PERF_RECORD_HEADER_MAX];
u32 nr_non_filtered_samples;
u32 nr_lost_warned;
diff --git a/tools/perf/util/machine.c b/tools/perf/util/machine.c
index eca1f8b..d8cde21 100644
--- a/tools/perf/util/machine.c
+++ b/tools/perf/util/machine.c
@@ -568,6 +568,11 @@ int machine__process_overhead_event(struct machine *machine __maybe_unused,
event->overhead.entry.nr,
event->overhead.entry.time,
event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_SB_OVERHEAD) {
+ dump_printf(" SB nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
} else {
dump_printf("\tUNSUPPORT OVERHEAD TYPE 0x%x!\n", event->overhead.type);
}
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index 594fd5e..e3aa9d7 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -1222,6 +1222,10 @@ overhead_stats_update(struct perf_tool *tool,
evlist->stats.total_mux_overhead[event->overhead.entry.cpu][0] += event->overhead.entry.nr;
evlist->stats.total_mux_overhead[event->overhead.entry.cpu][1] += event->overhead.entry.time;
break;
+ case PERF_SB_OVERHEAD:
+ evlist->stats.total_sb_overhead[event->overhead.entry.cpu][0] += event->overhead.entry.nr;
+ evlist->stats.total_sb_overhead[event->overhead.entry.cpu][1] += event->overhead.entry.time;
+ break;
default:
break;
}
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type |
| Message-ID | <sGJQe-7SF-27@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
The infrastructure to handle PERF_RECORD_OVERHEAD record type. A new
perf report option is also introduced as a knob to show the overhead
information.
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/include/uapi/linux/perf_event.h | 33 ++++++++++++++++++++++++++++++++
tools/perf/Documentation/perf-report.txt | 3 +++
tools/perf/builtin-report.c | 6 ++++++
tools/perf/util/event.c | 9 +++++++++
tools/perf/util/event.h | 11 +++++++++++
tools/perf/util/machine.c | 8 ++++++++
tools/perf/util/machine.h | 2 ++
tools/perf/util/session.c | 5 +++++
tools/perf/util/symbol.h | 3 ++-
tools/perf/util/tool.h | 1 +
10 files changed, 80 insertions(+), 1 deletion(-)
diff --git a/tools/include/uapi/linux/perf_event.h b/tools/include/uapi/linux/perf_event.h
index c66a485..5e7c522 100644
--- a/tools/include/uapi/linux/perf_event.h
+++ b/tools/include/uapi/linux/perf_event.h
@@ -862,6 +862,17 @@ enum perf_event_type {
*/
PERF_RECORD_SWITCH_CPU_WIDE = 15,
+ /*
+ * Records perf overhead
+ * struct {
+ * struct perf_event_header header;
+ * u32 type;
+ * struct perf_overhead_entry entry;
+ * struct sample_id sample_id;
+ * };
+ */
+ PERF_RECORD_OVERHEAD = 16,
+
PERF_RECORD_MAX, /* non-ABI */
};
@@ -980,4 +991,26 @@ struct perf_branch_entry {
reserved:44;
};
+enum perf_record_overhead_type {
+ PERF_NMI_OVERHEAD = 0,
+ PERF_MUX_OVERHEAD,
+ PERF_SB_OVERHEAD,
+
+ PERF_OVERHEAD_MAX,
+};
+
+/*
+ * single overhead record layout:
+ *
+ * cpu: The cpu which overhead occues
+ * nr: Times of overhead happens.
+ * E.g. for NMI, nr == times of NMI handler are called.
+ * time: Total overhead cost(ns)
+ */
+struct perf_overhead_entry {
+ __u32 cpu;
+ __u64 nr;
+ __u64 time;
+};
+
#endif /* _UAPI_LINUX_PERF_EVENT_H */
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 2d17462..fea8bea 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -412,6 +412,9 @@ include::itrace.txt[]
--hierarchy::
Enable hierarchical output.
+--show-overhead::
+ Show extra overhead which perf brings during monitoring
+
include::callchain-overhead-calculation.txt[]
SEE ALSO
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 3dfbfff..1416c39 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -368,6 +368,10 @@ static int perf_evlist__tty_browse_hists(struct perf_evlist *evlist,
struct perf_evsel *pos;
fprintf(stdout, "#\n# Total Lost Samples: %" PRIu64 "\n#\n", evlist->stats.total_lost_samples);
+ if (symbol_conf.show_overhead) {
+ fprintf(stdout, "# Overhead:\n");
+ fprintf(stdout, "#\n");
+ }
evlist__for_each_entry(evlist, pos) {
struct hists *hists = evsel__hists(pos);
const char *evname = perf_evsel__name(pos);
@@ -830,6 +834,8 @@ int cmd_report(int argc, const char **argv, const char *prefix __maybe_unused)
OPT_CALLBACK_DEFAULT(0, "stdio-color", NULL, "mode",
"'always' (default), 'never' or 'auto' only applicable to --stdio mode",
stdio__config_color, "always"),
+ OPT_BOOLEAN(0, "show-overhead", &symbol_conf.show_overhead,
+ "Show perf overhead"),
OPT_END()
};
struct perf_data_file file = {
diff --git a/tools/perf/util/event.c b/tools/perf/util/event.c
index 8ab0d7d..ca98c4c 100644
--- a/tools/perf/util/event.c
+++ b/tools/perf/util/event.c
@@ -31,6 +31,7 @@ static const char *perf_event__names[] = {
[PERF_RECORD_LOST_SAMPLES] = "LOST_SAMPLES",
[PERF_RECORD_SWITCH] = "SWITCH",
[PERF_RECORD_SWITCH_CPU_WIDE] = "SWITCH_CPU_WIDE",
+ [PERF_RECORD_OVERHEAD] = "OVERHEAD",
[PERF_RECORD_HEADER_ATTR] = "ATTR",
[PERF_RECORD_HEADER_EVENT_TYPE] = "EVENT_TYPE",
[PERF_RECORD_HEADER_TRACING_DATA] = "TRACING_DATA",
@@ -1056,6 +1057,14 @@ int perf_event__process_switch(struct perf_tool *tool __maybe_unused,
return machine__process_switch_event(machine, event);
}
+int perf_event__process_overhead(struct perf_tool *tool __maybe_unused,
+ union perf_event *event,
+ struct perf_sample *sample __maybe_unused,
+ struct machine *machine)
+{
+ return machine__process_overhead_event(machine, event);
+}
+
size_t perf_event__fprintf_mmap(union perf_event *event, FILE *fp)
{
return fprintf(fp, " %d/%d: [%#" PRIx64 "(%#" PRIx64 ") @ %#" PRIx64 "]: %c %s\n",
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index c735c53..d1b179b 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -480,6 +480,12 @@ struct time_conv_event {
u64 time_zero;
};
+struct perf_overhead {
+ struct perf_event_header header;
+ u32 type;
+ struct perf_overhead_entry entry;
+};
+
union perf_event {
struct perf_event_header header;
struct mmap_event mmap;
@@ -509,6 +515,7 @@ union perf_event {
struct stat_event stat;
struct stat_round_event stat_round;
struct time_conv_event time_conv;
+ struct perf_overhead overhead;
};
void perf_event__print_totals(void);
@@ -587,6 +594,10 @@ int perf_event__process_switch(struct perf_tool *tool,
union perf_event *event,
struct perf_sample *sample,
struct machine *machine);
+int perf_event__process_overhead(struct perf_tool *tool,
+ union perf_event *event,
+ struct perf_sample *sample,
+ struct machine *machine);
int perf_event__process_mmap(struct perf_tool *tool,
union perf_event *event,
struct perf_sample *sample,
diff --git a/tools/perf/util/machine.c b/tools/perf/util/machine.c
index 9b33bef..1101757 100644
--- a/tools/perf/util/machine.c
+++ b/tools/perf/util/machine.c
@@ -555,6 +555,12 @@ int machine__process_switch_event(struct machine *machine __maybe_unused,
return 0;
}
+int machine__process_overhead_event(struct machine *machine __maybe_unused,
+ union perf_event *event __maybe_unused)
+{
+ return 0;
+}
+
static void dso__adjust_kmod_long_name(struct dso *dso, const char *filename)
{
const char *dup_filename;
@@ -1536,6 +1542,8 @@ int machine__process_event(struct machine *machine, union perf_event *event,
case PERF_RECORD_SWITCH:
case PERF_RECORD_SWITCH_CPU_WIDE:
ret = machine__process_switch_event(machine, event); break;
+ case PERF_RECORD_OVERHEAD:
+ ret = machine__process_overhead_event(machine, event); break;
default:
ret = -1;
break;
diff --git a/tools/perf/util/machine.h b/tools/perf/util/machine.h
index 354de6e..ec2dd4d 100644
--- a/tools/perf/util/machine.h
+++ b/tools/perf/util/machine.h
@@ -97,6 +97,8 @@ int machine__process_itrace_start_event(struct machine *machine,
union perf_event *event);
int machine__process_switch_event(struct machine *machine,
union perf_event *event);
+int machine__process_overhead_event(struct machine *machine,
+ union perf_event *event);
int machine__process_mmap_event(struct machine *machine, union perf_event *event,
struct perf_sample *sample);
int machine__process_mmap2_event(struct machine *machine, union perf_event *event,
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index f268201..bc0bc21 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -373,6 +373,8 @@ void perf_tool__fill_defaults(struct perf_tool *tool)
tool->itrace_start = perf_event__process_itrace_start;
if (tool->context_switch == NULL)
tool->context_switch = perf_event__process_switch;
+ if (tool->overhead == NULL)
+ tool->overhead = perf_event__process_overhead;
if (tool->read == NULL)
tool->read = process_event_sample_stub;
if (tool->throttle == NULL)
@@ -786,6 +788,7 @@ static perf_event__swap_op perf_event__swap_ops[] = {
[PERF_RECORD_LOST_SAMPLES] = perf_event__all64_swap,
[PERF_RECORD_SWITCH] = perf_event__switch_swap,
[PERF_RECORD_SWITCH_CPU_WIDE] = perf_event__switch_swap,
+ [PERF_RECORD_OVERHEAD] = perf_event__all64_swap,
[PERF_RECORD_HEADER_ATTR] = perf_event__hdr_attr_swap,
[PERF_RECORD_HEADER_EVENT_TYPE] = perf_event__event_type_swap,
[PERF_RECORD_HEADER_TRACING_DATA] = perf_event__tracing_data_swap,
@@ -1267,6 +1270,8 @@ static int machines__deliver_event(struct machines *machines,
case PERF_RECORD_SWITCH:
case PERF_RECORD_SWITCH_CPU_WIDE:
return tool->context_switch(tool, event, sample, machine);
+ case PERF_RECORD_OVERHEAD:
+ return tool->overhead(tool, event, sample, machine);
default:
++evlist->stats.nr_unknown_events;
return -1;
diff --git a/tools/perf/util/symbol.h b/tools/perf/util/symbol.h
index 2d0a905..2d96bdb 100644
--- a/tools/perf/util/symbol.h
+++ b/tools/perf/util/symbol.h
@@ -117,7 +117,8 @@ struct symbol_conf {
show_ref_callgraph,
hide_unresolved,
raw_trace,
- report_hierarchy;
+ report_hierarchy,
+ show_overhead;
const char *vmlinux_name,
*kallsyms_name,
*source_prefix,
diff --git a/tools/perf/util/tool.h b/tools/perf/util/tool.h
index ac2590a..c5bbb34 100644
--- a/tools/perf/util/tool.h
+++ b/tools/perf/util/tool.h
@@ -47,6 +47,7 @@ struct perf_tool {
aux,
itrace_start,
context_switch,
+ overhead,
throttle,
unthrottle;
event_attr_op attr;
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-23 23:40 +0100 |
| Subject | Re: [PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type |
| Message-ID | <sGOmR-2pg-23@gated-at.bofh.it> |
| In reply to | #1528650 |
On Wed, Nov 23, 2016 at 04:44:43AM -0500, kan.liang@intel.com wrote:
SNIP
> +
> static void dso__adjust_kmod_long_name(struct dso *dso, const char *filename)
> {
> const char *dup_filename;
> @@ -1536,6 +1542,8 @@ int machine__process_event(struct machine *machine, union perf_event *event,
> case PERF_RECORD_SWITCH:
> case PERF_RECORD_SWITCH_CPU_WIDE:
> ret = machine__process_switch_event(machine, event); break;
> + case PERF_RECORD_OVERHEAD:
> + ret = machine__process_overhead_event(machine, event); break;
missing breaks
jirka
> default:
> ret = -1;
> break;
SNIP
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-24 00:00 +0100 |
| Subject | Re: [PATCH 05/14] perf tools: handle PERF_RECORD_OVERHEAD record type |
| Message-ID | <sGOGd-2vZ-9@gated-at.bofh.it> |
| In reply to | #1528813 |
On Wed, Nov 23, 2016 at 11:35:59PM +0100, Jiri Olsa wrote:
> On Wed, Nov 23, 2016 at 04:44:43AM -0500, kan.liang@intel.com wrote:
>
> SNIP
>
> > +
> > static void dso__adjust_kmod_long_name(struct dso *dso, const char *filename)
> > {
> > const char *dup_filename;
> > @@ -1536,6 +1542,8 @@ int machine__process_event(struct machine *machine, union perf_event *event,
> > case PERF_RECORD_SWITCH:
> > case PERF_RECORD_SWITCH_CPU_WIDE:
> > ret = machine__process_switch_event(machine, event); break;
> > + case PERF_RECORD_OVERHEAD:
> > + ret = machine__process_overhead_event(machine, event); break;
>
> missing breaks
ugh.. im blind.. sry :-\
jirka
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 12/14] perf tools: record elapsed time |
| Message-ID | <sGJQe-7SF-39@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
Record the elapsed time of perf record, and display it in perf report
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/builtin-record.c | 10 ++++++++++
tools/perf/builtin-report.c | 1 +
tools/perf/util/event.h | 3 +++
tools/perf/util/machine.c | 3 +++
tools/perf/util/session.c | 3 +++
5 files changed, 20 insertions(+)
diff --git a/tools/perf/builtin-record.c b/tools/perf/builtin-record.c
index 492058e..ea94e10 100644
--- a/tools/perf/builtin-record.c
+++ b/tools/perf/builtin-record.c
@@ -69,6 +69,7 @@ struct record {
bool switch_output;
unsigned long long samples;
struct write_overhead overhead[MAX_NR_CPUS];
+ u64 elapsed_time;
};
static u64 get_vnsecs(void)
@@ -866,6 +867,12 @@ static void perf_event__synth_overhead(struct record *rec, perf_event__handler_t
(void)process(&rec->tool, &event, NULL, NULL);
}
+
+ event.overhead.type = PERF_USER_ELAPSED_TIME;
+ event.overhead.entry.cpu = -1;
+ event.overhead.entry.nr = 1;
+ event.overhead.entry.time = rec->elapsed_time;
+ (void)process(&rec->tool, &event, NULL, NULL);
}
static int __cmd_record(struct record *rec, int argc, const char **argv)
@@ -1129,6 +1136,7 @@ static int __cmd_record(struct record *rec, int argc, const char **argv)
goto out_child;
}
+ rec->elapsed_time = get_nsecs() - rec->elapsed_time;
perf_event__synth_overhead(rec, process_synthesized_event);
if (!quiet)
@@ -1601,6 +1609,8 @@ int cmd_record(int argc, const char **argv, const char *prefix __maybe_unused)
# undef REASON
#endif
+ rec->elapsed_time = get_nsecs();
+
rec->evlist = perf_evlist__new();
if (rec->evlist == NULL)
return -ENOMEM;
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 9c0a424..de2a9b6 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -371,6 +371,7 @@ static int perf_evlist__tty_browse_hists(struct perf_evlist *evlist,
fprintf(stdout, "#\n# Total Lost Samples: %" PRIu64 "\n#\n", evlist->stats.total_lost_samples);
if (symbol_conf.show_overhead) {
+ fprintf(stdout, "# Elapsed time: %" PRIu64 " ns\n", evlist->stats.elapsed_time);
fprintf(stdout, "# Overhead:\n");
for (cpu = 0; cpu < session->header.env.nr_cpus_online; cpu++) {
if (!evlist->stats.total_nmi_overhead[cpu][0])
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index 9927cf9..ceb0968 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -275,6 +275,7 @@ enum auxtrace_error_type {
* The total_user_write_overhead tells exactly the overhead to write data in
* perf record.
* The total write# is stored in [0], while the accumulated time is in [1].
+ * The elapsed_time tells the elapsed time of perf record
*/
struct events_stats {
u64 total_period;
@@ -287,6 +288,7 @@ struct events_stats {
u64 total_mux_overhead[MAX_NR_CPUS][2];
u64 total_sb_overhead[MAX_NR_CPUS][2];
u64 total_user_write_overhead[MAX_NR_CPUS][2];
+ u64 elapsed_time;
u32 nr_events[PERF_RECORD_HEADER_MAX];
u32 nr_non_filtered_samples;
u32 nr_lost_warned;
@@ -497,6 +499,7 @@ struct time_conv_event {
u64 time_zero;
};
+#define PERF_USER_ELAPSED_TIME 200 /* above any possible overhead type */
enum perf_user_overhead_event_type { /* above any possible kernel type */
PERF_USER_OVERHEAD_TYPE_START = 100,
PERF_USER_WRITE_OVERHEAD = 100,
diff --git a/tools/perf/util/machine.c b/tools/perf/util/machine.c
index ce7a0ea..150071f 100644
--- a/tools/perf/util/machine.c
+++ b/tools/perf/util/machine.c
@@ -578,6 +578,9 @@ int machine__process_overhead_event(struct machine *machine __maybe_unused,
event->overhead.entry.nr,
event->overhead.entry.time,
event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_USER_ELAPSED_TIME) {
+ dump_printf(" Elapsed time: %llu\n",
+ event->overhead.entry.time);
} else {
dump_printf("\tUNSUPPORT OVERHEAD TYPE 0x%x!\n", event->overhead.type);
}
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index a72992b..e84808f 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -1231,6 +1231,9 @@ overhead_stats_update(struct perf_tool *tool,
evlist->stats.total_user_write_overhead[event->overhead.entry.cpu][0] += event->overhead.entry.nr;
evlist->stats.total_user_write_overhead[event->overhead.entry.cpu][1] += event->overhead.entry.time;
break;
+ case PERF_USER_ELAPSED_TIME:
+ evlist->stats.elapsed_time = event->overhead.entry.time;
+ break;
default:
break;
}
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 13/14] perf tools: warn on high overhead |
| Message-ID | <sGJQe-7SF-41@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
The rough overhead rate can be caculated by the sum of all kinds of
overhead / elapsed time.
If the overhead rate is higher than 10%, warning the user.
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/util/session.c | 26 ++++++++++++++++++++++++++
1 file changed, 26 insertions(+)
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index e84808f..decfc48 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -1559,6 +1559,30 @@ perf_session__warn_order(const struct perf_session *session)
ui__warning("%u out of order events recorded.\n", oe->nr_unordered_events);
}
+static void
+perf_session__warn_overhead(const struct perf_session *session)
+{
+ const struct events_stats *stats = &session->evlist->stats;
+ double overhead_rate;
+ u64 overhead;
+ int i;
+
+ for (i = 0; i < session->header.env.nr_cpus_online; i++) {
+ overhead = stats->total_nmi_overhead[i][1];
+ overhead += stats->total_mux_overhead[i][1];
+ overhead += stats->total_sb_overhead[i][1];
+ overhead += stats->total_user_write_overhead[i][1];
+
+ overhead_rate = (double)overhead / (double)stats->elapsed_time;
+
+ if (overhead_rate > 0.1) {
+ ui__warning("Perf overhead is high! The overhead rate is %3.2f%% on CPU %d\n\n"
+ "Please consider reducing the number of events, or increasing the period, or decrease the frequency.\n\n",
+ overhead_rate * 100.0, i);
+ }
+ }
+}
+
static void perf_session__warn_about_errors(const struct perf_session *session)
{
const struct events_stats *stats = &session->evlist->stats;
@@ -1632,6 +1656,8 @@ static void perf_session__warn_about_errors(const struct perf_session *session)
"Increase it by --proc-map-timeout\n",
stats->nr_proc_map_timeout);
}
+
+ perf_session__warn_overhead(session);
}
static int perf_session__flush_thread_stack(struct thread *thread,
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <andi@firstfloor.org> |
|---|---|
| Date | 2016-11-23 21:30 +0100 |
| Subject | Re: [PATCH 13/14] perf tools: warn on high overhead |
| Message-ID | <sGMl4-1a7-17@gated-at.bofh.it> |
| In reply to | #1528652 |
kan.liang@intel.com writes: > From: Kan Liang <kan.liang@intel.com> > > The rough overhead rate can be caculated by the sum of all kinds of > overhead / elapsed time. > If the overhead rate is higher than 10%, warning the user. Thinking about this more: this is comparing the cost of a single CPU to the total wall clock time. This isn't very good and can give confusing results with many cores. Perhaps we need two separate metrics here: - cost of perf record on its CPU (or later on if it gets multi threaded more multiple). Warn if this is >50% or so. - average perf collection overhead on a CPU. The 10% threshold here seems appropiate. -Andi
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2016-11-23 23:10 +0100 |
| Subject | RE: [PATCH 13/14] perf tools: warn on high overhead |
| Message-ID | <sGNTR-2bH-57@gated-at.bofh.it> |
| In reply to | #1528761 |
> > kan.liang@intel.com writes: > > > From: Kan Liang <kan.liang@intel.com> > > > > The rough overhead rate can be caculated by the sum of all kinds of > > overhead / elapsed time. > > If the overhead rate is higher than 10%, warning the user. > > Thinking about this more: this is comparing the cost of a single CPU to the > total wall clock time. This isn't very good and can give confusing results with > many cores. > > Perhaps we need two separate metrics here: > > - cost of perf record on its CPU (or later on if it gets multi threaded > more multiple). Warn if this is >50% or so. What's the formula for cost of perf record on its CPU? The cost only includes user space overhead or all overhead? What is the divisor? > - average perf collection overhead on a CPU. The 10% threshold here > seems appropiate. For the average, do you mean add all overheads among CPUs together and divide the CPU#? To calculate the rate, the divisor is wall clock time, right? Thanks, Kan
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 14/14] perf script: show overhead events |
| Message-ID | <sGJQe-7SF-49@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
Introduce a new option --show-overhead to show overhead events in perf
script
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/builtin-script.c | 36 ++++++++++++++++++++++++++++++++++++
tools/perf/util/event.c | 37 +++++++++++++++++++++++++++++++++++++
tools/perf/util/event.h | 1 +
3 files changed, 74 insertions(+)
diff --git a/tools/perf/builtin-script.c b/tools/perf/builtin-script.c
index e1daff3..76d9747 100644
--- a/tools/perf/builtin-script.c
+++ b/tools/perf/builtin-script.c
@@ -829,6 +829,7 @@ struct perf_script {
bool show_task_events;
bool show_mmap_events;
bool show_switch_events;
+ bool show_overhead;
bool allocated;
struct cpu_map *cpus;
struct thread_map *threads;
@@ -1264,6 +1265,37 @@ static int process_switch_event(struct perf_tool *tool,
return 0;
}
+static int process_overhead_event(struct perf_tool *tool,
+ union perf_event *event,
+ struct perf_sample *sample,
+ struct machine *machine)
+{
+ struct thread *thread;
+ struct perf_script *script = container_of(tool, struct perf_script, tool);
+ struct perf_session *session = script->session;
+ struct perf_evsel *evsel;
+
+ if (perf_event__process_switch(tool, event, sample, machine) < 0)
+ return -1;
+ if (sample) {
+ evsel = perf_evlist__id2evsel(session->evlist, sample->id);
+ thread = machine__findnew_thread(machine, sample->pid, sample->tid);
+ if (thread == NULL) {
+ pr_debug("problem processing OVERHEAD event, skipping it.\n");
+ return -1;
+ }
+
+ print_sample_start(sample, thread, evsel);
+ perf_event__fprintf(event, stdout);
+ thread__put(thread);
+ } else {
+ /* USER OVERHEAD event */
+ perf_event__fprintf(event, stdout);
+ }
+
+ return 0;
+}
+
static void sig_handler(int sig __maybe_unused)
{
session_done = 1;
@@ -1287,6 +1319,8 @@ static int __cmd_script(struct perf_script *script)
}
if (script->show_switch_events)
script->tool.context_switch = process_switch_event;
+ if (script->show_overhead)
+ script->tool.overhead = process_overhead_event;
ret = perf_session__process_events(script->session);
@@ -2172,6 +2206,8 @@ int cmd_script(int argc, const char **argv, const char *prefix __maybe_unused)
"Show the mmap events"),
OPT_BOOLEAN('\0', "show-switch-events", &script.show_switch_events,
"Show context switch events (if recorded)"),
+ OPT_BOOLEAN('\0', "show-overhead", &script.show_overhead,
+ "Show overhead events"),
OPT_BOOLEAN('f', "force", &file.force, "don't complain, do it"),
OPT_BOOLEAN(0, "ns", &nanosecs,
"Use 9 decimal places when displaying time"),
diff --git a/tools/perf/util/event.c b/tools/perf/util/event.c
index 6cd43c9..cd4f3aa 100644
--- a/tools/perf/util/event.c
+++ b/tools/perf/util/event.c
@@ -1190,6 +1190,39 @@ size_t perf_event__fprintf_switch(union perf_event *event, FILE *fp)
event->context_switch.next_prev_tid);
}
+size_t perf_event__fprintf_overhead(union perf_event *event, FILE *fp)
+{
+ size_t ret;
+
+ if (event->overhead.type == PERF_NMI_OVERHEAD) {
+ ret = fprintf(fp, " [NMI] nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_MUX_OVERHEAD) {
+ ret = fprintf(fp, " [MUX] nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_SB_OVERHEAD) {
+ ret = fprintf(fp, " [SB] nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_USER_WRITE_OVERHEAD) {
+ ret = fprintf(fp, " [USER WRITE] nr: %llu time: %llu cpu %u\n",
+ event->overhead.entry.nr,
+ event->overhead.entry.time,
+ event->overhead.entry.cpu);
+ } else if (event->overhead.type == PERF_USER_ELAPSED_TIME) {
+ ret = fprintf(fp, " [ELAPSED TIME] time: %llu\n",
+ event->overhead.entry.time);
+ } else {
+ ret = fprintf(fp, " unhandled!\n");
+ }
+ return ret;
+}
+
size_t perf_event__fprintf(union perf_event *event, FILE *fp)
{
size_t ret = fprintf(fp, "PERF_RECORD_%s",
@@ -1219,6 +1252,10 @@ size_t perf_event__fprintf(union perf_event *event, FILE *fp)
case PERF_RECORD_SWITCH_CPU_WIDE:
ret += perf_event__fprintf_switch(event, fp);
break;
+ case PERF_RECORD_OVERHEAD:
+ case PERF_RECORD_USER_OVERHEAD:
+ ret += perf_event__fprintf_overhead(event, fp);
+ break;
default:
ret += fprintf(fp, "\n");
}
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index ceb0968..36e295d 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -690,6 +690,7 @@ size_t perf_event__fprintf_switch(union perf_event *event, FILE *fp);
size_t perf_event__fprintf_thread_map(union perf_event *event, FILE *fp);
size_t perf_event__fprintf_cpu_map(union perf_event *event, FILE *fp);
size_t perf_event__fprintf(union perf_event *event, FILE *fp);
+size_t perf_event__fprintf_overhead(union perf_event *event, FILE *fp);
u64 kallsyms__get_function_start(const char *kallsyms_filename,
const char *symbol_name);
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-24 00:30 +0100 |
| Subject | Re: [PATCH 14/14] perf script: show overhead events |
| Message-ID | <sGP9f-2Vz-17@gated-at.bofh.it> |
| In reply to | #1528653 |
On Wed, Nov 23, 2016 at 04:44:52AM -0500, kan.liang@intel.com wrote:
> From: Kan Liang <kan.liang@intel.com>
>
> Introduce a new option --show-overhead to show overhead events in perf
> script
perf 7356 [001] 7292.203517: 482010 cycles:pp: ffffffff818e2150 _raw_spin_unlock_irqrestore+0x40 (/lib/modules/4.9.0-rc1+/build/vmlinux)
PERF_RECORD_USER_OVERHEAD [USER WRITE] nr: 1995 time: 14790661 cpu 1
PERF_RECORD_USER_OVERHEAD [ELAPSED TIME] time: 2721649901
perf 7356 [001] 7292.203766: 482076 cycles:pp: ffffffff81117f17 lock_release+0x37 (/lib/modules/4.9.0-rc1+/build/vmlinux)
I guess those 2 overhead events dont have context..?
should we make sure they are the last events?
thanks,
jirka
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-24 00:30 +0100 |
| Subject | Re: [PATCH 14/14] perf script: show overhead events |
| Message-ID | <sGP9f-2Vz-13@gated-at.bofh.it> |
| In reply to | #1528653 |
On Wed, Nov 23, 2016 at 04:44:52AM -0500, kan.liang@intel.com wrote: > From: Kan Liang <kan.liang@intel.com> > > Introduce a new option --show-overhead to show overhead events in perf > script please add exmaple output into changelog thanks, jirka
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-24 00:40 +0100 |
| Subject | Re: [PATCH 14/14] perf script: show overhead events |
| Message-ID | <sGPiW-2YE-5@gated-at.bofh.it> |
| In reply to | #1528653 |
On Wed, Nov 23, 2016 at 04:44:52AM -0500, kan.liang@intel.com wrote:
> From: Kan Liang <kan.liang@intel.com>
>
> Introduce a new option --show-overhead to show overhead events in perf
> script
>
> Signed-off-by: Kan Liang <kan.liang@intel.com>
> ---
> tools/perf/builtin-script.c | 36 ++++++++++++++++++++++++++++++++++++
> tools/perf/util/event.c | 37 +++++++++++++++++++++++++++++++++++++
> tools/perf/util/event.h | 1 +
> 3 files changed, 74 insertions(+)
>
> diff --git a/tools/perf/builtin-script.c b/tools/perf/builtin-script.c
> index e1daff3..76d9747 100644
> --- a/tools/perf/builtin-script.c
> +++ b/tools/perf/builtin-script.c
> @@ -829,6 +829,7 @@ struct perf_script {
> bool show_task_events;
> bool show_mmap_events;
> bool show_switch_events;
> + bool show_overhead;
> bool allocated;
> struct cpu_map *cpus;
> struct thread_map *threads;
> @@ -1264,6 +1265,37 @@ static int process_switch_event(struct perf_tool *tool,
> return 0;
> }
>
> +static int process_overhead_event(struct perf_tool *tool,
> + union perf_event *event,
> + struct perf_sample *sample,
> + struct machine *machine)
> +{
> + struct thread *thread;
> + struct perf_script *script = container_of(tool, struct perf_script, tool);
> + struct perf_session *session = script->session;
> + struct perf_evsel *evsel;
> +
> + if (perf_event__process_switch(tool, event, sample, machine) < 0)
> + return -1;
process_switch event? copy&paste error?
jirka
> + if (sample) {
> + evsel = perf_evlist__id2evsel(session->evlist, sample->id);
> + thread = machine__findnew_thread(machine, sample->pid, sample->tid);
> + if (thread == NULL) {
> + pr_debug("problem processing OVERHEAD event, skipping it.\n");
> + return -1;
> + }
> +
> + print_sample_start(sample, thread, evsel);
> + perf_event__fprintf(event, stdout);
> + thread__put(thread);
> + } else {
> + /* USER OVERHEAD event */
> + perf_event__fprintf(event, stdout);
> + }
> +
> + return 0;
> +}
> +
SNIP
[toc] | [prev] | [next] | [standalone]
| From | Jiri Olsa <jolsa@redhat.com> |
|---|---|
| Date | 2016-11-24 00:40 +0100 |
| Subject | Re: [PATCH 14/14] perf script: show overhead events |
| Message-ID | <sGPiW-2YE-29@gated-at.bofh.it> |
| In reply to | #1528653 |
On Wed, Nov 23, 2016 at 04:44:52AM -0500, kan.liang@intel.com wrote:
SNIP
> +}
> +
> static void sig_handler(int sig __maybe_unused)
> {
> session_done = 1;
> @@ -1287,6 +1319,8 @@ static int __cmd_script(struct perf_script *script)
> }
> if (script->show_switch_events)
> script->tool.context_switch = process_switch_event;
> + if (script->show_overhead)
> + script->tool.overhead = process_overhead_event;
>
> ret = perf_session__process_events(script->session);
>
> @@ -2172,6 +2206,8 @@ int cmd_script(int argc, const char **argv, const char *prefix __maybe_unused)
> "Show the mmap events"),
> OPT_BOOLEAN('\0', "show-switch-events", &script.show_switch_events,
> "Show context switch events (if recorded)"),
> + OPT_BOOLEAN('\0', "show-overhead", &script.show_overhead,
> + "Show overhead events"),
please add the '-events' suffix as for the mmap and task
thanks,
jirka
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 03/14] perf/x86: output multiplexing overhead |
| Message-ID | <sGJQe-7SF-53@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
Multiplexing overhead is one of the key overhead when the number of
events is more than available counters.
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
include/linux/perf_event.h | 2 ++
include/uapi/linux/perf_event.h | 1 +
kernel/events/core.c | 16 ++++++++++++++++
3 files changed, 19 insertions(+)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 632647f..f72b97a 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -793,6 +793,8 @@ struct perf_cpu_context {
struct list_head sched_cb_entry;
int sched_cb_usage;
+
+ struct perf_overhead_entry mux_overhead;
};
struct perf_output_handle {
diff --git a/include/uapi/linux/perf_event.h b/include/uapi/linux/perf_event.h
index 071323d..9124c7c 100644
--- a/include/uapi/linux/perf_event.h
+++ b/include/uapi/linux/perf_event.h
@@ -993,6 +993,7 @@ struct perf_branch_entry {
enum perf_record_overhead_type {
PERF_NMI_OVERHEAD = 0,
+ PERF_MUX_OVERHEAD,
PERF_OVERHEAD_MAX,
};
diff --git a/kernel/events/core.c b/kernel/events/core.c
index d82e6ca..9934059 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -1829,6 +1829,11 @@ event_sched_out(struct perf_event *event,
if (event->attr.exclusive || !cpuctx->active_oncpu)
cpuctx->exclusive = 0;
+ if (log_overhead && cpuctx->mux_overhead.nr) {
+ cpuctx->mux_overhead.cpu = smp_processor_id();
+ perf_log_overhead(event, PERF_MUX_OVERHEAD, &cpuctx->mux_overhead);
+ }
+
perf_pmu_enable(event->pmu);
}
@@ -3330,9 +3335,17 @@ static void rotate_ctx(struct perf_event_context *ctx)
list_rotate_left(&ctx->flexible_groups);
}
+static void
+perf_caculate_mux_overhead(struct perf_cpu_context *cpuctx, u64 time)
+{
+ cpuctx->mux_overhead.nr++;
+ cpuctx->mux_overhead.time += time;
+}
+
static int perf_rotate_context(struct perf_cpu_context *cpuctx)
{
struct perf_event_context *ctx = NULL;
+ u64 start_clock, end_clock;
int rotate = 0;
if (cpuctx->ctx.nr_events) {
@@ -3349,6 +3362,7 @@ static int perf_rotate_context(struct perf_cpu_context *cpuctx)
if (!rotate)
goto done;
+ start_clock = perf_clock();
perf_ctx_lock(cpuctx, cpuctx->task_ctx);
perf_pmu_disable(cpuctx->ctx.pmu);
@@ -3364,6 +3378,8 @@ static int perf_rotate_context(struct perf_cpu_context *cpuctx)
perf_pmu_enable(cpuctx->ctx.pmu);
perf_ctx_unlock(cpuctx, cpuctx->task_ctx);
+ end_clock = perf_clock();
+ perf_caculate_mux_overhead(cpuctx, end_clock - start_clock);
done:
return rotate;
--
2.5.5
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-11-23 21:10 +0100 |
| Subject | Re: [PATCH 03/14] perf/x86: output multiplexing overhead |
| Message-ID | <sGM1I-ZJ-29@gated-at.bofh.it> |
| In reply to | #1528654 |
On Wed, Nov 23, 2016 at 04:44:41AM -0500, kan.liang@intel.com wrote: > From: Kan Liang <kan.liang@intel.com> > > Multiplexing overhead is one of the key overhead when the number of > events is more than available counters. > > Signed-off-by: Kan Liang <kan.liang@intel.com> > --- > include/linux/perf_event.h | 2 ++ > include/uapi/linux/perf_event.h | 1 + > kernel/events/core.c | 16 ++++++++++++++++ > 3 files changed, 19 insertions(+) The subject says x86 specific, but its _all_ core code.
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2016-11-23 21:20 +0100 |
| Subject | RE: [PATCH 03/14] perf/x86: output multiplexing overhead |
| Message-ID | <sGMbn-16K-15@gated-at.bofh.it> |
| In reply to | #1528739 |
> > On Wed, Nov 23, 2016 at 04:44:41AM -0500, kan.liang@intel.com wrote: > > From: Kan Liang <kan.liang@intel.com> > > > > Multiplexing overhead is one of the key overhead when the number of > > events is more than available counters. > > > > Signed-off-by: Kan Liang <kan.liang@intel.com> > > --- > > include/linux/perf_event.h | 2 ++ > > include/uapi/linux/perf_event.h | 1 + > > kernel/events/core.c | 16 ++++++++++++++++ > > 3 files changed, 19 insertions(+) > > The subject says x86 specific, but its _all_ core code. Oh, Sorry. I will change the subject in next version. Thanks, Kan
[toc] | [prev] | [next] | [standalone]
| From | kan.liang@intel.com |
|---|---|
| Date | 2016-11-23 18:50 +0100 |
| Subject | [PATCH 10/14] perf tools: introduce PERF_RECORD_USER_OVERHEAD |
| Message-ID | <sGJQe-7SF-55@gated-at.bofh.it> |
| In reply to | #1528644 |
From: Kan Liang <kan.liang@intel.com>
User space perf tool also bring overhead. Introduce
PERF_RECORD_USER_OVERHEAD to track the overhead information.
Signed-off-by: Kan Liang <kan.liang@intel.com>
---
tools/perf/util/event.c | 1 +
tools/perf/util/event.h | 1 +
tools/perf/util/session.c | 4 ++++
3 files changed, 6 insertions(+)
diff --git a/tools/perf/util/event.c b/tools/perf/util/event.c
index ca98c4c..6cd43c9 100644
--- a/tools/perf/util/event.c
+++ b/tools/perf/util/event.c
@@ -48,6 +48,7 @@ static const char *perf_event__names[] = {
[PERF_RECORD_STAT_ROUND] = "STAT_ROUND",
[PERF_RECORD_EVENT_UPDATE] = "EVENT_UPDATE",
[PERF_RECORD_TIME_CONV] = "TIME_CONV",
+ [PERF_RECORD_USER_OVERHEAD] = "USER_OVERHEAD",
};
const char *perf_event__name(unsigned int id)
diff --git a/tools/perf/util/event.h b/tools/perf/util/event.h
index 3357529..1ef1a9d 100644
--- a/tools/perf/util/event.h
+++ b/tools/perf/util/event.h
@@ -237,6 +237,7 @@ enum perf_user_event_type { /* above any possible kernel type */
PERF_RECORD_STAT_ROUND = 77,
PERF_RECORD_EVENT_UPDATE = 78,
PERF_RECORD_TIME_CONV = 79,
+ PERF_RECORD_USER_OVERHEAD = 80,
PERF_RECORD_HEADER_MAX
};
diff --git a/tools/perf/util/session.c b/tools/perf/util/session.c
index e3aa9d7..27a5c8a 100644
--- a/tools/perf/util/session.c
+++ b/tools/perf/util/session.c
@@ -804,6 +804,7 @@ static perf_event__swap_op perf_event__swap_ops[] = {
[PERF_RECORD_STAT_ROUND] = perf_event__stat_round_swap,
[PERF_RECORD_EVENT_UPDATE] = perf_event__event_update_swap,
[PERF_RECORD_TIME_CONV] = perf_event__all64_swap,
+ [PERF_RECORD_USER_OVERHEAD] = perf_event__all64_swap,
[PERF_RECORD_HEADER_MAX] = NULL,
};
@@ -1382,6 +1383,9 @@ static s64 perf_session__process_user_event(struct perf_session *session,
case PERF_RECORD_TIME_CONV:
session->time_conv = event->time_conv;
return tool->time_conv(tool, event, session);
+ case PERF_RECORD_USER_OVERHEAD:
+ overhead_stats_update(tool, session->evlist, event);
+ return tool->overhead(tool, event, NULL, NULL);
default:
return -EINVAL;
}
--
2.5.5
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web