Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1246628 > unrolled thread
| Started by | Kaixu Xia <xiakaixu@huawei.com> |
|---|---|
| First post | 2015-10-14 14:40 +0200 |
| Last post | 2015-10-15 04:30 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH V2 0/2] bpf: enable/disable events stored in PERF_EVENT_ARRAY maps trace data output when perf sampling Kaixu Xia <xiakaixu@huawei.com> - 2015-10-14 14:40 +0200
[PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER Kaixu Xia <xiakaixu@huawei.com> - 2015-10-14 14:40 +0200
Re: [PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER Alexei Starovoitov <ast@plumgrid.com> - 2015-10-14 23:30 +0200
Re: [PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER xiakaixu <xiakaixu@huawei.com> - 2015-10-15 04:30 +0200
| From | Kaixu Xia <xiakaixu@huawei.com> |
|---|---|
| Date | 2015-10-14 14:40 +0200 |
| Subject | [PATCH V2 0/2] bpf: enable/disable events stored in PERF_EVENT_ARRAY maps trace data output when perf sampling |
| Message-ID | <qjtvA-70U-7@gated-at.bofh.it> |
Previous RFC patch url:
https://lkml.org/lkml/2015/10/12/135
changes in V2:
- rebase the whole patch set to net-next tree(4b418bf);
- remove the added flag perf_sample_disable in bpf_map;
- move the added fields in structure perf_event to proper place
to avoid cacheline miss;
- use counter based flag instead of 0/1 switcher in considering
of reentering events;
- use a single helper bpf_perf_event_sample_control() to enable/
disable events;
- implement a light-weight solution to control the trace data
output on current cpu;
- create a new ioctl PERF_EVENT_IOC_SET_ENABLER to enable/disable
a set of events;
Before this patch,
$ ./perf record -e cycles -a sleep 1
$ ./perf report --stdio
# To display the perf.data header info, please use --header/--header-only option
#
#
# Total Lost Samples: 0
#
# Samples: 643 of event 'cycles'
# Event count (approx.): 128313904
...
After this patch,
$ ./perf record -e pmux=cycles --event perf-bpf.o/my_cycles_map=pmux/ -a sleep 1
$ ./perf report --stdio
# To display the perf.data header info, please use --header/--header-only option
#
#
# Total Lost Samples: 0
#
# Samples: 25 of event 'cycles'
# Event count (approx.): 5788400
...
The bpf program example:
struct bpf_map_def SEC("maps") my_cycles_map = {
.type = BPF_MAP_TYPE_PERF_EVENT_ARRAY,
.key_size = sizeof(int),
.value_size = sizeof(u32),
.max_entries = 32,
};
SEC("enter=sys_write")
int bpf_prog_1(struct pt_regs *ctx)
{
bpf_perf_event_sample_control(&my_cycles_map, 32, 0);
return 0;
}
SEC("exit=sys_write%return")
int bpf_prog_2(struct pt_regs *ctx)
{
bpf_perf_event_sample_control(&my_cycles_map, 32, 1);
return 0;
}
Consider control sampling in function level, if we don't use the
PERF_EVENT_IOC_SET_ENABLER ioctl in perf user side, we must set
a high sample frequency to dump trace data.
Kaixu Xia (2):
bpf: control the trace data output on current cpu when perf sampling
bpf: control a set of perf events by creating a new ioctl
PERF_EVENT_IOC_SET_ENABLER
include/linux/perf_event.h | 2 ++
include/uapi/linux/bpf.h | 5 ++++
include/uapi/linux/perf_event.h | 4 +++-
kernel/bpf/verifier.c | 3 ++-
kernel/events/core.c | 53 +++++++++++++++++++++++++++++++++++++++++
kernel/trace/bpf_trace.c | 35 +++++++++++++++++++++++++++
6 files changed, 100 insertions(+), 2 deletions(-)
--
1.8.3.4
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Kaixu Xia <xiakaixu@huawei.com> |
|---|---|
| Date | 2015-10-14 14:40 +0200 |
| Subject | [PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER |
| Message-ID | <qjtvA-70U-21@gated-at.bofh.it> |
| In reply to | #1246628 |
This patch creates a new ioctl PERF_EVENT_IOC_SET_ENABLER to let
perf to select an event as 'enabler'. So we can set this 'enabler'
event to enable/disable a set of events. The event on CPU 0 is
treated as the 'enabler' event by default.
Signed-off-by: Kaixu Xia <xiakaixu@huawei.com>
---
include/linux/perf_event.h | 1 +
include/uapi/linux/perf_event.h | 1 +
kernel/events/core.c | 42 ++++++++++++++++++++++++++++++++++++++++-
kernel/trace/bpf_trace.c | 5 ++++-
4 files changed, 47 insertions(+), 2 deletions(-)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index dcbf7d5..bc9fe77 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -473,6 +473,7 @@ struct perf_event {
atomic_t event_limit;
atomic_t sample_disable;
+ atomic_t *p_sample_disable;
void (*destroy)(struct perf_event *);
struct rcu_head rcu_head;
diff --git a/include/uapi/linux/perf_event.h b/include/uapi/linux/perf_event.h
index a2b9dd7..3b4fb90 100644
--- a/include/uapi/linux/perf_event.h
+++ b/include/uapi/linux/perf_event.h
@@ -393,6 +393,7 @@ struct perf_event_attr {
#define PERF_EVENT_IOC_SET_FILTER _IOW('$', 6, char *)
#define PERF_EVENT_IOC_ID _IOR('$', 7, __u64 *)
#define PERF_EVENT_IOC_SET_BPF _IOW('$', 8, __u32)
+#define PERF_EVENT_IOC_SET_ENABLER _IO ('$', 9)
enum perf_event_ioc_flags {
PERF_IOC_FLAG_GROUP = 1U << 0,
diff --git a/kernel/events/core.c b/kernel/events/core.c
index 942351c..03d2594 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -4152,6 +4152,7 @@ static int perf_event_set_output(struct perf_event *event,
struct perf_event *output_event);
static int perf_event_set_filter(struct perf_event *event, void __user *arg);
static int perf_event_set_bpf_prog(struct perf_event *event, u32 prog_fd);
+static int perf_event_set_sample_enabler(struct perf_event *event, u32 enabler_fd);
static long _perf_ioctl(struct perf_event *event, unsigned int cmd, unsigned long arg)
{
@@ -4208,6 +4209,9 @@ static long _perf_ioctl(struct perf_event *event, unsigned int cmd, unsigned lon
case PERF_EVENT_IOC_SET_BPF:
return perf_event_set_bpf_prog(event, arg);
+ case PERF_EVENT_IOC_SET_ENABLER:
+ return perf_event_set_sample_enabler(event, arg);
+
default:
return -ENOTTY;
}
@@ -6337,7 +6341,7 @@ static int __perf_event_overflow(struct perf_event *event,
irq_work_queue(&event->pending);
}
- if (!atomic_read(&event->sample_disable))
+ if (!atomic_read(event->p_sample_disable))
return ret;
if (event->overflow_handler)
@@ -6989,6 +6993,35 @@ static int perf_event_set_bpf_prog(struct perf_event *event, u32 prog_fd)
return 0;
}
+static int perf_event_set_sample_enabler(struct perf_event *event, u32 enabler_fd)
+{
+ int ret;
+ struct fd enabler;
+ struct perf_event *enabler_event;
+
+ if (enabler_fd == -1)
+ return 0;
+
+ ret = perf_fget_light(enabler_fd, &enabler);
+ if (ret)
+ return ret;
+ enabler_event = enabler.file->private_data;
+ if (event == enabler_event) {
+ fdput(enabler);
+ return 0;
+ }
+
+ /* they must be on the same PMU*/
+ if (event->pmu != enabler_event->pmu) {
+ fdput(enabler);
+ return -EINVAL;
+ }
+
+ event->p_sample_disable = &enabler_event->sample_disable;
+ fdput(enabler);
+ return 0;
+}
+
static void perf_event_free_bpf_prog(struct perf_event *event)
{
struct bpf_prog *prog;
@@ -7023,6 +7056,11 @@ static int perf_event_set_bpf_prog(struct perf_event *event, u32 prog_fd)
return -ENOENT;
}
+static int perf_event_set_sample_enabler(struct perf_event *event, u32 group_fd)
+{
+ return -ENOENT;
+}
+
static void perf_event_free_bpf_prog(struct perf_event *event)
{
}
@@ -7718,6 +7756,8 @@ static void perf_event_check_sample_flag(struct perf_event *event)
atomic_set(&event->sample_disable, 0);
else
atomic_set(&event->sample_disable, 1);
+
+ event->p_sample_disable = &event->sample_disable;
}
/*
diff --git a/kernel/trace/bpf_trace.c b/kernel/trace/bpf_trace.c
index f261333..d012be3 100644
--- a/kernel/trace/bpf_trace.c
+++ b/kernel/trace/bpf_trace.c
@@ -221,9 +221,12 @@ static u64 bpf_perf_event_sample_control(u64 r1, u64 index, u64 flag, u64 r4, u6
struct bpf_array *array = container_of(map, struct bpf_array, map);
struct perf_event *event;
- if (unlikely(index >= array->map.max_entries))
+ if (unlikely(index > array->map.max_entries))
return -E2BIG;
+ if (index == array->map.max_entries)
+ index = 0;
+
event = (struct perf_event *)array->ptrs[index];
if (!event)
return -ENOENT;
--
1.8.3.4
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alexei Starovoitov <ast@plumgrid.com> |
|---|---|
| Date | 2015-10-14 23:30 +0200 |
| Subject | Re: [PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER |
| Message-ID | <qjBMu-2mZ-19@gated-at.bofh.it> |
| In reply to | #1246629 |
On 10/14/15 5:37 AM, Kaixu Xia wrote: > + event->p_sample_disable = &enabler_event->sample_disable; I don't like it as a concept and it's buggy implementation. What happens here when enabler is alive, but other event is destroyed? > --- a/kernel/trace/bpf_trace.c > +++ b/kernel/trace/bpf_trace.c > @@ -221,9 +221,12 @@ static u64 bpf_perf_event_sample_control(u64 r1, u64 index, u64 flag, u64 r4, u6 > struct bpf_array *array = container_of(map, struct bpf_array, map); > struct perf_event *event; > > - if (unlikely(index >= array->map.max_entries)) > + if (unlikely(index > array->map.max_entries)) > return -E2BIG; > > + if (index == array->map.max_entries) > + index = 0; what is this hack for ? Either use notification and user space disable or call bpf_perf_event_sample_control() manually for each cpu. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | xiakaixu <xiakaixu@huawei.com> |
|---|---|
| Date | 2015-10-15 04:30 +0200 |
| Subject | Re: [PATCH V2 2/2] bpf: control a set of perf events by creating a new ioctl PERF_EVENT_IOC_SET_ENABLER |
| Message-ID | <qjGsO-10W-9@gated-at.bofh.it> |
| In reply to | #1247198 |
δΊ 2015/10/15 5:28, Alexei Starovoitov ει: > On 10/14/15 5:37 AM, Kaixu Xia wrote: >> + event->p_sample_disable = &enabler_event->sample_disable; > > I don't like it as a concept and it's buggy implementation. > What happens here when enabler is alive, but other event is destroyed? > >> --- a/kernel/trace/bpf_trace.c >> +++ b/kernel/trace/bpf_trace.c >> @@ -221,9 +221,12 @@ static u64 bpf_perf_event_sample_control(u64 r1, u64 index, u64 flag, u64 r4, u6 >> struct bpf_array *array = container_of(map, struct bpf_array, map); >> struct perf_event *event; >> >> - if (unlikely(index >= array->map.max_entries)) >> + if (unlikely(index > array->map.max_entries)) >> return -E2BIG; >> >> + if (index == array->map.max_entries) >> + index = 0; > > what is this hack for ? > > Either use notification and user space disable or > call bpf_perf_event_sample_control() manually for each cpu. I will discard current implemention that controlling a set of perf events by the 'enabler' event. Call bpf_perf_event_sample_control() manually for each cpu is fine. Maybe we can add a loop to control all the events stored in maps by judging the index, OK? > > > > . > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web