Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1371200 > unrolled thread
| Started by | Alexei Starovoitov <ast@fb.com> |
|---|---|
| First post | 2016-04-05 07:00 +0200 |
| Last post | 2016-04-09 00:20 +0200 |
| Articles | 5 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH net-next 0/8] allow bpf attach to tracepoints Alexei Starovoitov <ast@fb.com> - 2016-04-05 07:00 +0200
[PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs Alexei Starovoitov <ast@fb.com> - 2016-04-05 07:00 +0200
Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs Peter Zijlstra <peterz@infradead.org> - 2016-04-05 14:10 +0200
Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs Alexei Starovoitov <ast@fb.com> - 2016-04-05 19:50 +0200
Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs Steven Rostedt <rostedt@goodmis.org> - 2016-04-09 00:20 +0200
| From | Alexei Starovoitov <ast@fb.com> |
|---|---|
| Date | 2016-04-05 07:00 +0200 |
| Subject | [PATCH net-next 0/8] allow bpf attach to tracepoints |
| Message-ID | <rkrfP-652-3@gated-at.bofh.it> |
Hi Steven, Peter,
last time we discussed bpf+tracepoints it was a year ago [1] and the reason
we didn't proceed with that approach was that bpf would make arguments
arg1, arg2 to trace_xx(arg1, arg2) call to be exposed to bpf program
and that was considered unnecessary extension of abi. Back then I wanted
to avoid the cost of buffer alloc and field assign part in all
of the tracepoints, but looks like when optimized the cost is acceptable.
So this new apporach doesn't expose any new abi to bpf program.
The program is looking at tracepoint fields after they were copied
by perf_trace_xx() and described in /sys/kernel/debug/tracing/events/xxx/format
We made a tool [2] that takes arguments from /sys/.../format and works as:
$ tplist.py -v random:urandom_read
int got_bits;
int pool_left;
int input_left;
Then these fields can be copy-pasted into bpf program like:
struct urandom_read {
__u64 hidden_pad;
int got_bits;
int pool_left;
int input_left;
};
and the program can use it:
SEC("tracepoint/random/urandom_read")
int bpf_prog(struct urandom_read *ctx)
{
return ctx->pool_left > 0 ? 1 : 0;
}
This way the program can access tracepoint fields faster than
equivalent bpf+kprobe program, which is the main goal of these patches.
Patch 1 and 2 are simple changes in perf core side, please review.
I'd like to take the whole set via net-next tree, since the rest of
the patches might conflict with other bpf work going on in net-next
and we want to avoid cross-tree merge conflicts.
Patch 7 is an example of access to tracepoint fields from bpf prog.
Patch 8 is a micro benchmark for bpf+kprobe vs bpf+tracepoint.
Note that for actual tracing tools the user doesn't need to
run tplist.py and copy-paste fields manually. The tools do it
automatically. Like argdist tool [3] can be used as:
$ argdist -H 't:block:block_rq_complete():u32:nr_sector'
where 'nr_sector' is name of tracepoint field taken from
/sys/kernel/debug/tracing/events/block/block_rq_complete/format
and appropriate bpf program is generated on the fly.
[1] http://thread.gmane.org/gmane.linux.kernel.api/8127/focus=8165
[2] https://github.com/iovisor/bcc/blob/master/tools/tplist.py
[3] https://github.com/iovisor/bcc/blob/master/tools/argdist.py
Alexei Starovoitov (8):
perf: optimize perf_fetch_caller_regs
perf, bpf: allow bpf programs attach to tracepoints
bpf: register BPF_PROG_TYPE_TRACEPOINT program type
bpf: support bpf_get_stackid() and bpf_perf_event_output() in
tracepoint programs
bpf: sanitize bpf tracepoint access
samples/bpf: add tracepoint support to bpf loader
samples/bpf: tracepoint example
samples/bpf: add tracepoint vs kprobe performance tests
include/linux/bpf.h | 2 +
include/linux/perf_event.h | 2 -
include/linux/trace_events.h | 1 +
include/trace/perf.h | 18 +++-
include/uapi/linux/bpf.h | 1 +
kernel/bpf/stackmap.c | 2 +-
kernel/bpf/verifier.c | 6 +-
kernel/events/core.c | 21 ++++-
kernel/trace/bpf_trace.c | 85 ++++++++++++++++-
kernel/trace/trace_event_perf.c | 4 +
kernel/trace/trace_events.c | 18 ++++
samples/bpf/Makefile | 5 +
samples/bpf/bpf_load.c | 26 +++++-
samples/bpf/offwaketime_kern.c | 26 +++++-
samples/bpf/test_overhead_kprobe_kern.c | 41 ++++++++
samples/bpf/test_overhead_tp_kern.c | 36 +++++++
samples/bpf/test_overhead_user.c | 161 ++++++++++++++++++++++++++++++++
17 files changed, 432 insertions(+), 23 deletions(-)
create mode 100644 samples/bpf/test_overhead_kprobe_kern.c
create mode 100644 samples/bpf/test_overhead_tp_kern.c
create mode 100644 samples/bpf/test_overhead_user.c
--
2.8.0
[toc] | [next] | [standalone]
| From | Alexei Starovoitov <ast@fb.com> |
|---|---|
| Date | 2016-04-05 07:00 +0200 |
| Subject | [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs |
| Message-ID | <rkrfQ-652-21@gated-at.bofh.it> |
| In reply to | #1371200 |
avoid memset in perf_fetch_caller_regs, since it's the critical path of all tracepoints.
It's called from perf_sw_event_sched, perf_event_task_sched_in and all of perf_trace_##call
with this_cpu_ptr(&__perf_regs[..]) which are zero initialized by perpcu_alloc and
subsequent call to perf_arch_fetch_caller_regs initializes the same fields on all archs,
so we can safely drop memset from all of the above cases and move it into
perf_ftrace_function_call that calls it with stack allocated pt_regs.
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
---
include/linux/perf_event.h | 2 --
kernel/trace/trace_event_perf.c | 1 +
2 files changed, 1 insertion(+), 2 deletions(-)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index f291275ffd71..e89f7199c223 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -882,8 +882,6 @@ static inline void perf_arch_fetch_caller_regs(struct pt_regs *regs, unsigned lo
*/
static inline void perf_fetch_caller_regs(struct pt_regs *regs)
{
- memset(regs, 0, sizeof(*regs));
-
perf_arch_fetch_caller_regs(regs, CALLER_ADDR0);
}
diff --git a/kernel/trace/trace_event_perf.c b/kernel/trace/trace_event_perf.c
index 00df25fd86ef..7a68afca8249 100644
--- a/kernel/trace/trace_event_perf.c
+++ b/kernel/trace/trace_event_perf.c
@@ -316,6 +316,7 @@ perf_ftrace_function_call(unsigned long ip, unsigned long parent_ip,
BUILD_BUG_ON(ENTRY_SIZE > PERF_MAX_TRACE_SIZE);
+ memset(®s, 0, sizeof(regs));
perf_fetch_caller_regs(®s);
entry = perf_trace_buf_prepare(ENTRY_SIZE, TRACE_FN, NULL, &rctx);
--
2.8.0
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-05 14:10 +0200 |
| Subject | Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs |
| Message-ID | <rkxXX-36a-5@gated-at.bofh.it> |
| In reply to | #1371201 |
On Mon, Apr 04, 2016 at 09:52:47PM -0700, Alexei Starovoitov wrote: > avoid memset in perf_fetch_caller_regs, since it's the critical path of all tracepoints. > It's called from perf_sw_event_sched, perf_event_task_sched_in and all of perf_trace_##call > with this_cpu_ptr(&__perf_regs[..]) which are zero initialized by perpcu_alloc Its not actually allocated; but because its a static uninitialized variable we get .bss like behaviour and the initial value is copied to all CPUs when the per-cpu allocator thingy bootstraps SMP IIRC. > and > subsequent call to perf_arch_fetch_caller_regs initializes the same fields on all archs, > so we can safely drop memset from all of the above cases and Indeed. > move it into > perf_ftrace_function_call that calls it with stack allocated pt_regs. Hmm, is there a reason that's still on-stack instead of using the per-cpu thing, Steve? > Signed-off-by: Alexei Starovoitov <ast@kernel.org> In any case, Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@fb.com> |
|---|---|
| Date | 2016-04-05 19:50 +0200 |
| Subject | Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs |
| Message-ID | <rkDh1-7Hy-15@gated-at.bofh.it> |
| In reply to | #1371453 |
On 4/5/16 5:06 AM, Peter Zijlstra wrote: > On Mon, Apr 04, 2016 at 09:52:47PM -0700, Alexei Starovoitov wrote: >> avoid memset in perf_fetch_caller_regs, since it's the critical path of all tracepoints. >> It's called from perf_sw_event_sched, perf_event_task_sched_in and all of perf_trace_##call >> with this_cpu_ptr(&__perf_regs[..]) which are zero initialized by perpcu_alloc > > Its not actually allocated; but because its a static uninitialized > variable we get .bss like behaviour and the initial value is copied to > all CPUs when the per-cpu allocator thingy bootstraps SMP IIRC. yes, it's .bss-like in a special section. I think static percpu still goes through some fancy boot time init similar to dynamic. What I tried to emphasize that either static or dynamic percpu areas are guaranteed to be zero initialized. >> and >> subsequent call to perf_arch_fetch_caller_regs initializes the same fields on all archs, >> so we can safely drop memset from all of the above cases and > > Indeed. > >> move it into >> perf_ftrace_function_call that calls it with stack allocated pt_regs. > > Hmm, is there a reason that's still on-stack instead of using the > per-cpu thing, Steve? > >> Signed-off-by: Alexei Starovoitov <ast@kernel.org> > > In any case, > > Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org> Thanks for the quick review.
[toc] | [prev] | [next] | [standalone]
| From | Steven Rostedt <rostedt@goodmis.org> |
|---|---|
| Date | 2016-04-09 00:20 +0200 |
| Subject | Re: [PATCH net-next 1/8] perf: optimize perf_fetch_caller_regs |
| Message-ID | <rlMUW-2BG-1@gated-at.bofh.it> |
| In reply to | #1371453 |
On Tue, 5 Apr 2016 14:06:26 +0200 Peter Zijlstra <peterz@infradead.org> wrote: > On Mon, Apr 04, 2016 at 09:52:47PM -0700, Alexei Starovoitov wrote: > > avoid memset in perf_fetch_caller_regs, since it's the critical path of all tracepoints. > > It's called from perf_sw_event_sched, perf_event_task_sched_in and all of perf_trace_##call > > with this_cpu_ptr(&__perf_regs[..]) which are zero initialized by perpcu_alloc > > Its not actually allocated; but because its a static uninitialized > variable we get .bss like behaviour and the initial value is copied to > all CPUs when the per-cpu allocator thingy bootstraps SMP IIRC. > > > and > > subsequent call to perf_arch_fetch_caller_regs initializes the same fields on all archs, > > so we can safely drop memset from all of the above cases and > > Indeed. > > > move it into > > perf_ftrace_function_call that calls it with stack allocated pt_regs. > > Hmm, is there a reason that's still on-stack instead of using the > per-cpu thing, Steve? Well, what do you do when you are tracing with regs in an interrupt that already set the per cpu regs field? We could create our own per-cpu one as well, but then that would require checking which level we are in, as we can have one for normal context, one for softirq context, one for irq context and one for nmi context. -- Steve > > > Signed-off-by: Alexei Starovoitov <ast@kernel.org> > > In any case, > > Acked-by: Peter Zijlstra (Intel) <peterz@infradead.org>
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web