Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1448160 > unrolled thread
| Started by | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| First post | 2016-07-21 23:20 +0200 |
| Last post | 2016-07-25 20:20 +0200 |
| Articles | 20 on this page of 60 — 7 participants |
Back to article view | Back to linux.kernel
[RFC PATCH v7 0/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-21 23:20 +0200
[RFC PATCH v7 2/7] tracing: instrument restartable sequences Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-21 23:20 +0200
[RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-21 23:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-07-26 01:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-26 05:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-03 14:30 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-03 18:50 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Christoph Lameter <cl@linux.com> - 2016-08-03 20:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-04 07:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Boqun Feng <boqun.feng@gmail.com> - 2016-08-04 06:30 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-04 07:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Boqun Feng <boqun.feng@gmail.com> - 2016-08-09 18:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 20:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-10 21:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 22:50 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-10 21:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Christoph Lameter <cl@linux.com> - 2016-08-03 20:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 22:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Christoph Lameter <cl@linux.com> - 2016-08-10 22:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Boqun Feng <boqun.feng@gmail.com> - 2016-07-27 17:10 +0200
[RFC 4/4] Restartable sequences: Add self-tests for PPC Boqun Feng <boqun.feng@gmail.com> - 2016-07-27 17:10 +0200
Re: [RFC 4/4] Restartable sequences: Add self-tests for PPC Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 05:10 +0200
Re: [RFC 4/4] Restartable sequences: Add self-tests for PPC Boqun Feng <boqun.feng@gmail.com> - 2016-07-28 06:50 +0200
[RFC v2] Restartable sequences: Add self-tests for PPC Boqun Feng <boqun.feng@gmail.com> - 2016-07-28 09:40 +0200
Re: [RFC v2] Restartable sequences: Add self-tests for PPC Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 16:10 +0200
Re: [RFC 4/4] Restartable sequences: Add self-tests for PPC Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 15:50 +0200
[RFC 1/4] rseq/param_test: Convert test_data_entry::count to intptr_t Boqun Feng <boqun.feng@gmail.com> - 2016-07-27 17:10 +0200
[RFC 3/4] Restartable sequences: Wire up powerpc system call Boqun Feng <boqun.feng@gmail.com> - 2016-07-27 17:10 +0200
Re: [RFC 3/4] Restartable sequences: Wire up powerpc system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 05:20 +0200
Re: [RFC 1/4] rseq/param_test: Convert test_data_entry::count to intptr_t Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 05:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-28 05:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-03 16:00 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-08-03 17:00 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Boqun Feng <boqun.feng@gmail.com> - 2016-08-03 17:50 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-07 17:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Boqun Feng <boqun.feng@gmail.com> - 2016-08-08 01:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-09 15:30 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-09 22:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-09 23:40 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 00:50 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 21:00 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 21:00 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-10 21:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-10 22:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 20:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 20:30 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-10 21:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 21:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-10 21:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 22:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-10 22:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-08-10 23:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Andy Lutomirski <luto@amacapital.net> - 2016-08-10 21:20 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-10 22:10 +0200
Re: [RFC PATCH v7 1/7] Restartable sequences system call Peter Zijlstra <peterz@infradead.org> - 2016-08-10 22:10 +0200
[RFC PATCH v7 6/7] Restartable sequences: wire up x86 32/64 system call Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-21 23:20 +0200
Re: [RFC PATCH v7 7/7] Restartable sequences: self-tests Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-24 05:10 +0200
Re: [RFC PATCH v7 7/7] Restartable sequences: self-tests Dave Watson <davejwatson@fb.com> - 2016-07-24 20:10 +0200
Re: [RFC PATCH v7 7/7] Restartable sequences: self-tests Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-25 18:50 +0200
Re: [RFC PATCH v7 7/7] Restartable sequences: self-tests Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2016-07-25 20:20 +0200
Page 1 of 3 [1] 2 3 Next page →
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-07-21 23:20 +0200 |
| Subject | [RFC PATCH v7 0/7] Restartable sequences system call |
| Message-ID | <rXtxT-4A4-3@gated-at.bofh.it> |
Hi, This is mostly a re-write of Paul Turner and Andrew Hunter's restartable critical sections (percpu atomics), which brings the following main benefits over Paul Turner's prior version (v2): - The ABI is now architecture-agnostic, and it requires fewer instruction on the user-space fast path, - Ported to ARM 32, in addition to cover x86 32/64. Adding support for new architectures is now trivial, - Progress is ensured by a fall-back to locking (purely userspace) when single-stepped by a debugger. This is v7, as it derives from my prior getcpu cache and thread local ABI patchsets. You will find benchmark results in the changelog of patch 1/7. Feedback is welcome! Thanks, Mathieu Mathieu Desnoyers (7): Restartable sequences system call tracing: instrument restartable sequences Restartable sequences: ARM 32 architecture support Restartable sequences: wire up ARM 32 system call Restartable sequences: x86 32/64 architecture support Restartable sequences: wire up x86 32/64 system call Restartable sequences: self-tests MAINTAINERS | 7 + arch/Kconfig | 7 + arch/arm/Kconfig | 1 + arch/arm/include/uapi/asm/unistd.h | 1 + arch/arm/kernel/calls.S | 1 + arch/arm/kernel/signal.c | 7 + arch/x86/Kconfig | 1 + arch/x86/entry/common.c | 1 + arch/x86/entry/syscalls/syscall_32.tbl | 1 + arch/x86/entry/syscalls/syscall_64.tbl | 1 + arch/x86/kernel/signal.c | 6 + fs/exec.c | 1 + include/linux/sched.h | 68 ++ include/trace/events/rseq.h | 60 ++ include/uapi/linux/Kbuild | 1 + include/uapi/linux/rseq.h | 85 +++ init/Kconfig | 13 + kernel/Makefile | 1 + kernel/fork.c | 2 + kernel/rseq.c | 243 +++++++ kernel/sched/core.c | 1 + kernel/sys_ni.c | 3 + tools/testing/selftests/rseq/.gitignore | 3 + tools/testing/selftests/rseq/Makefile | 13 + .../testing/selftests/rseq/basic_percpu_ops_test.c | 279 ++++++++ tools/testing/selftests/rseq/basic_test.c | 106 +++ tools/testing/selftests/rseq/param_test.c | 707 +++++++++++++++++++++ tools/testing/selftests/rseq/rseq.c | 200 ++++++ tools/testing/selftests/rseq/rseq.h | 449 +++++++++++++ 29 files changed, 2269 insertions(+) create mode 100644 include/trace/events/rseq.h create mode 100644 include/uapi/linux/rseq.h create mode 100644 kernel/rseq.c create mode 100644 tools/testing/selftests/rseq/.gitignore create mode 100644 tools/testing/selftests/rseq/Makefile create mode 100644 tools/testing/selftests/rseq/basic_percpu_ops_test.c create mode 100644 tools/testing/selftests/rseq/basic_test.c create mode 100644 tools/testing/selftests/rseq/param_test.c create mode 100644 tools/testing/selftests/rseq/rseq.c create mode 100644 tools/testing/selftests/rseq/rseq.h -- 2.1.4
[toc] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-07-21 23:20 +0200 |
| Subject | [RFC PATCH v7 2/7] tracing: instrument restartable sequences |
| Message-ID | <rXtxU-4A4-29@gated-at.bofh.it> |
| In reply to | #1448160 |
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
CC: Thomas Gleixner <tglx@linutronix.de>
CC: Paul Turner <pjt@google.com>
CC: Andrew Hunter <ahh@google.com>
CC: Peter Zijlstra <peterz@infradead.org>
CC: Andy Lutomirski <luto@amacapital.net>
CC: Andi Kleen <andi@firstfloor.org>
CC: Dave Watson <davejwatson@fb.com>
CC: Chris Lameter <cl@linux.com>
CC: Ingo Molnar <mingo@redhat.com>
CC: "H. Peter Anvin" <hpa@zytor.com>
CC: Ben Maurer <bmaurer@fb.com>
CC: Steven Rostedt <rostedt@goodmis.org>
CC: "Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
CC: Josh Triplett <josh@joshtriplett.org>
CC: Linus Torvalds <torvalds@linux-foundation.org>
CC: Andrew Morton <akpm@linux-foundation.org>
CC: Russell King <linux@arm.linux.org.uk>
CC: Catalin Marinas <catalin.marinas@arm.com>
CC: Will Deacon <will.deacon@arm.com>
CC: Michael Kerrisk <mtk.manpages@gmail.com>
CC: Boqun Feng <boqun.feng@gmail.com>
CC: linux-api@vger.kernel.org
---
include/trace/events/rseq.h | 60 +++++++++++++++++++++++++++++++++++++++++++++
kernel/rseq.c | 18 +++++++++++---
2 files changed, 75 insertions(+), 3 deletions(-)
create mode 100644 include/trace/events/rseq.h
diff --git a/include/trace/events/rseq.h b/include/trace/events/rseq.h
new file mode 100644
index 0000000..83fd31e
--- /dev/null
+++ b/include/trace/events/rseq.h
@@ -0,0 +1,60 @@
+#undef TRACE_SYSTEM
+#define TRACE_SYSTEM rseq
+
+#if !defined(_TRACE_RSEQ_H) || defined(TRACE_HEADER_MULTI_READ)
+#define _TRACE_RSEQ_H
+
+#include <linux/tracepoint.h>
+
+TRACE_EVENT(rseq_inc,
+
+ TP_PROTO(uint32_t event_counter, int ret),
+
+ TP_ARGS(event_counter, ret),
+
+ TP_STRUCT__entry(
+ __field(uint32_t, event_counter)
+ __field(int, ret)
+ ),
+
+ TP_fast_assign(
+ __entry->event_counter = event_counter;
+ __entry->ret = ret;
+ ),
+
+ TP_printk("event_counter=%u ret=%d",
+ __entry->event_counter, __entry->ret)
+);
+
+TRACE_EVENT(rseq_ip_fixup,
+
+ TP_PROTO(void __user *regs_ip, void __user *post_commit_ip,
+ void __user *abort_ip, uint32_t kevcount, int ret),
+
+ TP_ARGS(regs_ip, post_commit_ip, abort_ip, kevcount, ret),
+
+ TP_STRUCT__entry(
+ __field(void __user *, regs_ip)
+ __field(void __user *, post_commit_ip)
+ __field(void __user *, abort_ip)
+ __field(uint32_t, kevcount)
+ __field(int, ret)
+ ),
+
+ TP_fast_assign(
+ __entry->regs_ip = regs_ip;
+ __entry->post_commit_ip = post_commit_ip;
+ __entry->abort_ip = abort_ip;
+ __entry->kevcount = kevcount;
+ __entry->ret = ret;
+ ),
+
+ TP_printk("regs_ip=%p post_commit_ip=%p abort_ip=%p kevcount=%u ret=%d",
+ __entry->regs_ip, __entry->post_commit_ip, __entry->abort_ip,
+ __entry->kevcount, __entry->ret)
+);
+
+#endif /* _TRACE_SOCK_H */
+
+/* This part must be outside protection */
+#include <trace/define_trace.h>
diff --git a/kernel/rseq.c b/kernel/rseq.c
index e1c847b..cab326a 100644
--- a/kernel/rseq.c
+++ b/kernel/rseq.c
@@ -29,6 +29,9 @@
#include <linux/rseq.h>
#include <asm/ptrace.h>
+#define CREATE_TRACE_POINTS
+#include <trace/events/rseq.h>
+
/*
* Each restartable sequence assembly block defines a "struct rseq_cs"
* structure which describes the post_commit_ip address, and the
@@ -90,8 +93,12 @@
static int rseq_increment_event_counter(struct task_struct *t)
{
- if (__put_user(++t->rseq_event_counter,
- &t->rseq->u.e.event_counter))
+ int ret;
+
+ ret = __put_user(++t->rseq_event_counter,
+ &t->rseq->u.e.event_counter);
+ trace_rseq_inc(t->rseq_event_counter, ret);
+ if (ret)
return -1;
return 0;
}
@@ -134,8 +141,13 @@ static int rseq_ip_fixup(struct pt_regs *regs)
struct task_struct *t = current;
void __user *post_commit_ip = NULL;
void __user *abort_ip = NULL;
+ int ret;
- if (rseq_get_rseq_cs(t, &post_commit_ip, &abort_ip))
+ ret = rseq_get_rseq_cs(t, &post_commit_ip, &abort_ip);
+ trace_rseq_ip_fixup((void __user *)instruction_pointer(regs),
+ post_commit_ip, abort_ip, t->rseq_event_counter,
+ ret);
+ if (ret)
return -1;
/* Handle potentially being within a critical section. */
--
2.1.4
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-07-21 23:20 +0200 |
| Subject | [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <rXtxU-4A4-31@gated-at.bofh.it> |
| In reply to | #1448160 |
Expose a new system call allowing each thread to register one userspace
memory area to be used as an ABI between kernel and user-space for two
purposes: user-space restartable sequences and quick access to read the
current CPU number value from user-space.
* Restartable sequences (per-cpu atomics)
The restartable critical sections (percpu atomics) work has been started
by Paul Turner and Andrew Hunter. It lets the kernel handle restart of
critical sections. [1] [2] The re-implementation proposed here brings a
few simplifications to the ABI which facilitates porting to other
architectures and speeds up the user-space fast path. A locking-based
fall-back, purely implemented in user-space, is proposed here to deal
with debugger single-stepping. This fallback interacts with rseq_start()
and rseq_finish(), which force retries in response to concurrent
lock-based activity.
Here are benchmarks of counter increment in various scenarios compared
to restartable sequences:
ARMv7 Processor rev 4 (v7l)
Machine model: Cubietruck
Counter increment speed (ns/increment)
1 thread 2 threads
global increment (baseline) 6 N/A
percpu rseq increment 50 52
percpu rseq spinlock 94 94
global atomic increment 48 74 (__sync_add_and_fetch_4)
global atomic CAS 50 172 (__sync_val_compare_and_swap_4)
global pthread mutex 148 862
ARMv7 Processor rev 10 (v7l)
Machine model: Wandboard
Counter increment speed (ns/increment)
1 thread 4 threads
global increment (baseline) 7 N/A
percpu rseq increment 50 50
percpu rseq spinlock 82 84
global atomic increment 44 262 (__sync_add_and_fetch_4)
global atomic CAS 46 316 (__sync_val_compare_and_swap_4)
global pthread mutex 146 1400
x86-64 Intel(R) Xeon(R) CPU E5-2630 v3 @ 2.40GHz:
Counter increment speed (ns/increment)
1 thread 8 threads
global increment (baseline) 3.0 N/A
percpu rseq increment 3.6 3.8
percpu rseq spinlock 5.6 6.2
global LOCK; inc 8.0 166.4
global LOCK; cmpxchg 13.4 435.2
global pthread mutex 25.2 1363.6
* Reading the current CPU number
Speeding up reading the current CPU number on which the caller thread is
running is done by keeping the current CPU number up do date within the
cpu_id field of the memory area registered by the thread. This is done
by making scheduler migration set the TIF_NOTIFY_RESUME flag on the
current thread. Upon return to user-space, a notify-resume handler
updates the current CPU value within the registered user-space memory
area. User-space can then read the current CPU number directly from
memory.
Keeping the current cpu id in a memory area shared between kernel and
user-space is an improvement over current mechanisms available to read
the current CPU number, which has the following benefits over
alternative approaches:
- 35x speedup on ARM vs system call through glibc
- 20x speedup on x86 compared to calling glibc, which calls vdso
executing a "lsl" instruction,
- 14x speedup on x86 compared to inlined "lsl" instruction,
- Unlike vdso approaches, this cpu_id value can be read from an inline
assembly, which makes it a useful building block for restartable
sequences.
- The approach of reading the cpu id through memory mapping shared
between kernel and user-space is portable (e.g. ARM), which is not the
case for the lsl-based x86 vdso.
On x86, yet another possible approach would be to use the gs segment
selector to point to user-space per-cpu data. This approach performs
similarly to the cpu id cache, but it has two disadvantages: it is
not portable, and it is incompatible with existing applications already
using the gs segment selector for other purposes.
Benchmarking various approaches for reading the current CPU number:
ARMv7 Processor rev 4 (v7l)
Machine model: Cubietruck
- Baseline (empty loop): 8.4 ns
- Read CPU from rseq cpu_id: 16.7 ns
- Read CPU from rseq cpu_id (lazy register): 19.8 ns
- glibc 2.19-0ubuntu6.6 getcpu: 301.8 ns
- getcpu system call: 234.9 ns
x86-64 Intel(R) Xeon(R) CPU E5-2630 v3 @ 2.40GHz:
- Baseline (empty loop): 0.8 ns
- Read CPU from rseq cpu_id: 0.8 ns
- Read CPU from rseq cpu_id (lazy register): 0.8 ns
- Read using gs segment selector: 0.8 ns
- "lsl" inline assembly: 13.0 ns
- glibc 2.19-0ubuntu6 getcpu: 16.6 ns
- getcpu system call: 53.9 ns
- Speed
Running 10 runs of hackbench -l 100000 seems to indicate, contrary to
expectations, that enabling CONFIG_RSEQ slightly accelerates the
scheduler:
Configuration: 2 sockets * 8-core Intel(R) Xeon(R) CPU E5-2630 v3 @
2.40GHz (directly on hardware, hyperthreading disabled in BIOS, energy
saving disabled in BIOS, turboboost disabled in BIOS, cpuidle.off=1
kernel parameter), with a Linux v4.6 defconfig+localyesconfig,
restartable sequences series applied.
* CONFIG_RSEQ=n
avg.: 41.37 s
std.dev.: 0.36 s
* CONFIG_RSEQ=y
avg.: 40.46 s
std.dev.: 0.33 s
- Size
On x86-64, between CONFIG_RSEQ=n/y, the text size increase of vmlinux is
2855 bytes, and the data size increase of vmlinux is 1024 bytes.
* CONFIG_RSEQ=n
text data bss dec hex filename
9964559 4256280 962560 15183399 e7ae27 vmlinux.norseq
* CONFIG_RSEQ=y
text data bss dec hex filename
9967414 4257304 962560 15187278 e7bd4e vmlinux.rseq
[1] https://lwn.net/Articles/650333/
[2] http://www.linuxplumbersconf.org/2013/ocw/system/presentations/1695/original/LPC%20-%20PerCpu%20Atomics.pdf
Link: http://lkml.kernel.org/r/20151027235635.16059.11630.stgit@pjt-glaptop.roam.corp.google.com
Link: http://lkml.kernel.org/r/20150624222609.6116.86035.stgit@kitami.mtv.corp.google.com
Signed-off-by: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
CC: Thomas Gleixner <tglx@linutronix.de>
CC: Paul Turner <pjt@google.com>
CC: Andrew Hunter <ahh@google.com>
CC: Peter Zijlstra <peterz@infradead.org>
CC: Andy Lutomirski <luto@amacapital.net>
CC: Andi Kleen <andi@firstfloor.org>
CC: Dave Watson <davejwatson@fb.com>
CC: Chris Lameter <cl@linux.com>
CC: Ingo Molnar <mingo@redhat.com>
CC: "H. Peter Anvin" <hpa@zytor.com>
CC: Ben Maurer <bmaurer@fb.com>
CC: Steven Rostedt <rostedt@goodmis.org>
CC: "Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
CC: Josh Triplett <josh@joshtriplett.org>
CC: Linus Torvalds <torvalds@linux-foundation.org>
CC: Andrew Morton <akpm@linux-foundation.org>
CC: Russell King <linux@arm.linux.org.uk>
CC: Catalin Marinas <catalin.marinas@arm.com>
CC: Will Deacon <will.deacon@arm.com>
CC: Michael Kerrisk <mtk.manpages@gmail.com>
CC: Boqun Feng <boqun.feng@gmail.com>
CC: linux-api@vger.kernel.org
---
Changes since v1:
- Return -1, errno=EINVAL if cpu_cache pointer is not aligned on
sizeof(int32_t).
- Update man page to describe the pointer alignement requirements and
update atomicity guarantees.
- Add MAINTAINERS file GETCPU_CACHE entry.
- Remove dynamic memory allocation: go back to having a single
getcpu_cache entry per thread. Update documentation accordingly.
- Rebased on Linux 4.4.
Changes since v2:
- Introduce a "cmd" argument, along with an enum with GETCPU_CACHE_GET
and GETCPU_CACHE_SET. Introduce a uapi header linux/getcpu_cache.h
defining this enumeration.
- Split resume notifier architecture implementation from the system call
wire up in the following arch-specific patches.
- Man pages updates.
- Handle 32-bit compat pointers.
- Simplify handling of getcpu_cache GETCPU_CACHE_SET compiler barrier:
set the current cpu cache pointer before doing the cache update, and
set it back to NULL if the update fails. Setting it back to NULL on
error ensures that no resume notifier will trigger a SIGSEGV if a
migration happened concurrently.
Changes since v3:
- Fix __user annotations in compat code,
- Update memory ordering comments.
- Rebased on kernel v4.5-rc5.
Changes since v4:
- Inline getcpu_cache_fork, getcpu_cache_execve, and getcpu_cache_exit.
- Add new line between if() and switch() to improve readability.
- Added sched switch benchmarks (hackbench) and size overhead comparison
to change log.
Changes since v5:
- Rename "getcpu_cache" to "thread_local_abi", allowing to extend
this system call to cover future features such as restartable critical
sections. Generalizing this system call ensures that we can add
features similar to the cpu_id field within the same cache-line
without having to track one pointer per feature within the task
struct.
- Add a tlabi_nr parameter to the system call, thus allowing to extend
the ABI beyond the initial 64-byte structure by registering structures
with tlabi_nr greater than 0. The initial ABI structure is associated
with tlabi_nr 0.
- Rebased on kernel v4.5.
Changes since v6:
- Integrate "restartable sequences" v2 patchset from Paul Turner.
- Add handling of single-stepping purely in user-space, with a
fallback to locking after 2 rseq failures to ensure progress, and
by exposing a __rseq_table section to debuggers so they know where
to put breakpoints when dealing with rseq assembly blocks which
can be aborted at any point.
- make the code and ABI generic: porting the kernel implementation
simply requires to wire up the signal handler and return to user-space
hooks, and allocate the syscall number.
- extend testing with a fully configurable test program. See
param_spinlock_test -h for details.
- handling of rseq ENOSYS in user-space, also with a fallback
to locking.
- modify Paul Turner's rseq ABI to only require a single TLS store on
the user-space fast-path, removing the need to populate two additional
registers. This is made possible by introducing struct rseq_cs into
the ABI to describe a critical section start_ip, post_commit_ip, and
abort_ip.
- Rebased on kernel v4.7-rc7.
Man page associated:
RSEQ(2) Linux Programmer's Manual RSEQ(2)
NAME
rseq - Restartable sequences and cpu number cache
SYNOPSIS
#include <linux/rseq.h>
int rseq(struct rseq * rseq, int flags);
DESCRIPTION
The rseq() ABI accelerates user-space operations on per-cpu
data by defining a shared data structure ABI between each user-
space thread and the kernel.
The rseq argument is a pointer to the thread-local rseq struc‐
ture to be shared between kernel and user-space. A NULL rseq
value can be used to check whether rseq is registered for the
current thread.
The layout of struct rseq is as follows:
Structure alignment
This structure needs to be aligned on multiples of 64
bytes.
Structure size
This structure has a fixed size of 128 bytes.
Fields
cpu_id
Cache of the CPU number on which the calling thread is
running.
event_counter
Restartable sequences event_counter field.
rseq_cs
Restartable sequences rseq_cs field. Points to a struct
rseq_cs.
The layout of struct rseq_cs is as follows:
Structure alignment
This structure needs to be aligned on multiples of 64
bytes.
Structure size
This structure has a fixed size of 192 bytes.
Fields
start_ip
Instruction pointer address of the first instruction of
the sequence of consecutive assembly instructions.
post_commit_ip
Instruction pointer address after the last instruction
of the sequence of consecutive assembly instructions.
abort_ip
Instruction pointer address where to move the execution
flow in case of abort of the sequence of consecutive
assembly instructions.
The flags argument is currently unused and must be specified as
0.
Typically, a library or application will keep the rseq struc‐
ture in a thread-local storage variable, or other memory areas
belonging to each thread. It is recommended to perform volatile
reads of the thread-local cache to prevent the compiler from
doing load tearing. An alternative approach is to read each
field from inline assembly.
Each thread is responsible for registering its rseq structure.
Only one rseq structure address can be registered per thread.
Once set, the rseq address is idempotent for a given thread.
In a typical usage scenario, the thread registering the rseq
structure will be performing loads and stores from/to that
structure. It is however also allowed to read that structure
from other threads. The rseq field updates performed by the
kernel provide single-copy atomicity semantics, which guarantee
that other threads performing single-copy atomic reads of the
cpu number cache will always observe a consistent value.
Memory registered as rseq structure should never be deallocated
before the thread which registered it exits: specifically, it
should not be freed, and the library containing the registered
thread-local storage should not be dlclose'd. Violating this
constraint may cause a SIGSEGV signal to be delivered to the
thread.
Unregistration of associated rseq structure is implicitly per‐
formed when a thread or process exit.
RETURN VALUE
A return value of 0 indicates success. On error, -1 is
returned, and errno is set appropriately.
ERRORS
EINVAL Either flags is non-zero, or rseq contains an address
which is not appropriately aligned.
ENOSYS The rseq() system call is not implemented by this ker‐
nel.
EFAULT rseq is an invalid address.
EBUSY The rseq argument contains a non-NULL address which dif‐
fers from the memory location already registered for
this thread.
ENOENT The rseq argument is NULL, but no memory location is
currently registered for this thread.
VERSIONS
The rseq() system call was added in Linux 4.X (TODO).
CONFORMING TO
rseq() is Linux-specific.
EXAMPLE
The following code uses the rseq() system call to keep a
thread-local storage variable up to date with the current CPU
number, with a fallback on sched_getcpu(3) if the cache is not
available. For example simplicity, it is done in main(), but
multithreaded programs would need to invoke rseq() from each
program thread.
#define _GNU_SOURCE
#include <stdlib.h>
#include <stdio.h>
#include <unistd.h>
#include <stdint.h>
#include <sched.h>
#include <stddef.h>
#include <errno.h>
#include <string.h>
#include <sys/syscall.h>
#include <linux/rseq.h>
static __thread volatile struct rseq rseq_state = {
.u.e.cpu_id = -1,
};
static int
sys_rseq(volatile struct rseq *rseq_abi, int flags)
{
return syscall(__NR_rseq, rseq_abi, flags);
}
static int32_t
rseq_current_cpu_raw(void)
{
return rseq_state.u.e.cpu_id;
}
static int32_t
rseq_current_cpu(void)
{
int32_t cpu;
cpu = rseq_current_cpu_raw();
if (cpu < 0)
cpu = sched_getcpu();
return cpu;
}
static int
rseq_init_current_thread(void)
{
int rc;
rc = sys_rseq(&rseq_state, 0);
if (rc) {
fprintf(stderr, "Error: sys_rseq(...) failed(%d): %s\n",
errno, strerror(errno));
return -1;
}
return 0;
}
int
main(int argc, char **argv)
{
if (rseq_init_current_thread()) {
fprintf(stderr,
"Unable to initialize restartable sequences.\n");
fprintf(stderr, "Using sched_getcpu() as fallback.\n");
}
printf("Current CPU number: %d\n", rseq_current_cpu());
exit(EXIT_SUCCESS);
}
SEE ALSO
sched_getcpu(3)
Linux 2016-07-19 RSEQ(2)
---
MAINTAINERS | 7 ++
arch/Kconfig | 7 ++
fs/exec.c | 1 +
include/linux/sched.h | 68 ++++++++++++++
include/uapi/linux/Kbuild | 1 +
include/uapi/linux/rseq.h | 85 +++++++++++++++++
init/Kconfig | 13 +++
kernel/Makefile | 1 +
kernel/fork.c | 2 +
kernel/rseq.c | 231 ++++++++++++++++++++++++++++++++++++++++++++++
kernel/sched/core.c | 1 +
kernel/sys_ni.c | 3 +
12 files changed, 420 insertions(+)
create mode 100644 include/uapi/linux/rseq.h
create mode 100644 kernel/rseq.c
diff --git a/MAINTAINERS b/MAINTAINERS
index 1209323..daef027 100644
--- a/MAINTAINERS
+++ b/MAINTAINERS
@@ -5085,6 +5085,13 @@ M: Joe Perches <joe@perches.com>
S: Maintained
F: scripts/get_maintainer.pl
+RESTARTABLE SEQUENCES SUPPORT
+M: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
+L: linux-kernel@vger.kernel.org
+S: Supported
+F: kernel/rseq.c
+F: include/uapi/linux/rseq.h
+
GFS2 FILE SYSTEM
M: Steven Whitehouse <swhiteho@redhat.com>
M: Bob Peterson <rpeterso@redhat.com>
diff --git a/arch/Kconfig b/arch/Kconfig
index 1599629..2c23e26 100644
--- a/arch/Kconfig
+++ b/arch/Kconfig
@@ -242,6 +242,13 @@ config HAVE_REGS_AND_STACK_ACCESS_API
declared in asm/ptrace.h
For example the kprobes-based event tracer needs this API.
+config HAVE_RSEQ
+ bool
+ depends on HAVE_REGS_AND_STACK_ACCESS_API
+ help
+ This symbol should be selected by an architecture if it
+ supports an implementation of restartable sequences.
+
config HAVE_CLK
bool
help
diff --git a/fs/exec.c b/fs/exec.c
index 887c1c9..e912d87 100644
--- a/fs/exec.c
+++ b/fs/exec.c
@@ -1707,6 +1707,7 @@ static int do_execveat_common(int fd, struct filename *filename,
/* execve succeeded */
current->fs->in_exec = 0;
current->in_execve = 0;
+ rseq_execve(current);
acct_update_integrals(current);
task_numa_free(current);
free_bprm(bprm);
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 253538f..5c4b900 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -59,6 +59,7 @@ struct sched_param {
#include <linux/gfp.h>
#include <linux/magic.h>
#include <linux/cgroup-defs.h>
+#include <linux/rseq.h>
#include <asm/processor.h>
@@ -1918,6 +1919,10 @@ struct task_struct {
#ifdef CONFIG_MMU
struct task_struct *oom_reaper_list;
#endif
+#ifdef CONFIG_RSEQ
+ struct rseq __user *rseq;
+ uint32_t rseq_event_counter;
+#endif
/* CPU-specific state of this task */
struct thread_struct thread;
/*
@@ -3387,4 +3392,67 @@ void cpufreq_add_update_util_hook(int cpu, struct update_util_data *data,
void cpufreq_remove_update_util_hook(int cpu);
#endif /* CONFIG_CPU_FREQ */
+#ifdef CONFIG_RSEQ
+static inline void rseq_set_notify_resume(struct task_struct *t)
+{
+ if (t->rseq)
+ set_tsk_thread_flag(t, TIF_NOTIFY_RESUME);
+}
+void __rseq_handle_notify_resume(struct pt_regs *regs);
+static inline void rseq_handle_notify_resume(struct pt_regs *regs)
+{
+ if (current->rseq)
+ __rseq_handle_notify_resume(regs);
+}
+/*
+ * If parent process has a registered restartable sequences area, the
+ * child inherits. Only applies when forking a process, not a thread. In
+ * case a parent fork() in the middle of a restartable sequence, set the
+ * resume notifier to force the child to retry.
+ */
+static inline void rseq_fork(struct task_struct *t, unsigned long clone_flags)
+{
+ if (clone_flags & CLONE_THREAD) {
+ t->rseq = NULL;
+ t->rseq_event_counter = 0;
+ } else {
+ t->rseq = current->rseq;
+ t->rseq_event_counter = current->rseq_event_counter;
+ rseq_set_notify_resume(t);
+ }
+}
+static inline void rseq_execve(struct task_struct *t)
+{
+ t->rseq = NULL;
+ t->rseq_event_counter = 0;
+}
+static inline void rseq_sched_out(struct task_struct *t)
+{
+ rseq_set_notify_resume(t);
+}
+static inline void rseq_signal_deliver(struct pt_regs *regs)
+{
+ rseq_handle_notify_resume(regs);
+}
+#else
+static inline void rseq_set_notify_resume(struct task_struct *t)
+{
+}
+static inline void rseq_handle_notify_resume(struct pt_regs *regs)
+{
+}
+static inline void rseq_fork(struct task_struct *t, unsigned long clone_flags)
+{
+}
+static inline void rseq_execve(struct task_struct *t)
+{
+}
+static inline void rseq_sched_out(struct task_struct *t)
+{
+}
+static inline void rseq_signal_deliver(struct pt_regs *regs)
+{
+}
+#endif
+
#endif
diff --git a/include/uapi/linux/Kbuild b/include/uapi/linux/Kbuild
index 8bdae34..2e64fb8 100644
--- a/include/uapi/linux/Kbuild
+++ b/include/uapi/linux/Kbuild
@@ -403,6 +403,7 @@ header-y += tcp_metrics.h
header-y += telephony.h
header-y += termios.h
header-y += thermal.h
+header-y += rseq.h
header-y += time.h
header-y += times.h
header-y += timex.h
diff --git a/include/uapi/linux/rseq.h b/include/uapi/linux/rseq.h
new file mode 100644
index 0000000..3e79fa9
--- /dev/null
+++ b/include/uapi/linux/rseq.h
@@ -0,0 +1,85 @@
+#ifndef _UAPI_LINUX_RSEQ_H
+#define _UAPI_LINUX_RSEQ_H
+
+/*
+ * linux/rseq.h
+ *
+ * Restartable sequences system call API
+ *
+ * Copyright (c) 2015-2016 Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
+ *
+ * Permission is hereby granted, free of charge, to any person obtaining a copy
+ * of this software and associated documentation files (the "Software"), to deal
+ * in the Software without restriction, including without limitation the rights
+ * to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
+ * copies of the Software, and to permit persons to whom the Software is
+ * furnished to do so, subject to the following conditions:
+ *
+ * The above copyright notice and this permission notice shall be included in
+ * all copies or substantial portions of the Software.
+ *
+ * THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
+ * IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
+ * FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
+ * AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
+ * LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
+ * OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
+ * SOFTWARE.
+ */
+
+#ifdef __KERNEL__
+# include <linux/types.h>
+#else /* #ifdef __KERNEL__ */
+# include <stdint.h>
+#endif /* #else #ifdef __KERNEL__ */
+
+#include <asm/byteorder.h>
+
+#ifdef __LP64__
+# define RSEQ_FIELD_u32_u64(field) uint64_t field
+#elif defined(__BYTE_ORDER) ? \
+ __BYTE_ORDER == __BIG_ENDIAN : defined(__BIG_ENDIAN)
+# define RSEQ_FIELD_u32_u64(field) uint32_t _padding ## field, field
+#else
+# define RSEQ_FIELD_u32_u64(field) uint32_t field, _padding ## field
+#endif
+
+struct rseq_cs {
+ RSEQ_FIELD_u32_u64(start_ip);
+ RSEQ_FIELD_u32_u64(post_commit_ip);
+ RSEQ_FIELD_u32_u64(abort_ip);
+} __attribute__((aligned(sizeof(uint64_t))));
+
+struct rseq {
+ union {
+ struct {
+ /*
+ * Restartable sequences cpu_id field.
+ * Updated by the kernel, and read by user-space with
+ * single-copy atomicity semantics. Aligned on 32-bit.
+ * Negative values are reserved for user-space.
+ */
+ int32_t cpu_id;
+ /*
+ * Restartable sequences event_counter field.
+ * Updated by the kernel, and read by user-space with
+ * single-copy atomicity semantics. Aligned on 32-bit.
+ */
+ uint32_t event_counter;
+ } e;
+ /*
+ * On architectures with 64-bit aligned reads, both cpu_id and
+ * event_counter can be read with single-copy atomicity
+ * semantics.
+ */
+ uint64_t v;
+ } u;
+ /*
+ * Restartable sequences rseq_cs field.
+ * Updated by user-space, read by the kernel with
+ * single-copy atomicity semantics. Aligned on 64-bit.
+ */
+ RSEQ_FIELD_u32_u64(rseq_cs);
+} __attribute__((aligned(sizeof(uint64_t))));
+
+#endif /* _UAPI_LINUX_RSEQ_H */
diff --git a/init/Kconfig b/init/Kconfig
index c02d897..545b7ed 100644
--- a/init/Kconfig
+++ b/init/Kconfig
@@ -1653,6 +1653,19 @@ config MEMBARRIER
If unsure, say Y.
+config RSEQ
+ bool "Enable rseq() system call" if EXPERT
+ default y
+ depends on HAVE_RSEQ
+ help
+ Enable the restartable sequences system call. It provides a
+ user-space cache for the current CPU number value, which
+ speeds up getting the current CPU number from user-space,
+ as well as an ABI to speed up user-space operations on
+ per-CPU data.
+
+ If unsure, say Y.
+
config EMBEDDED
bool "Embedded system"
option allnoconfig_y
diff --git a/kernel/Makefile b/kernel/Makefile
index e2ec54e..4c6d8b5 100644
--- a/kernel/Makefile
+++ b/kernel/Makefile
@@ -112,6 +112,7 @@ obj-$(CONFIG_TORTURE_TEST) += torture.o
obj-$(CONFIG_MEMBARRIER) += membarrier.o
obj-$(CONFIG_HAS_IOMEM) += memremap.o
+obj-$(CONFIG_RSEQ) += rseq.o
$(obj)/configs.o: $(obj)/config_data.h
diff --git a/kernel/fork.c b/kernel/fork.c
index 4a7ec0c..cc7756b 100644
--- a/kernel/fork.c
+++ b/kernel/fork.c
@@ -1591,6 +1591,8 @@ static struct task_struct *copy_process(unsigned long clone_flags,
*/
copy_seccomp(p);
+ rseq_fork(p, clone_flags);
+
/*
* Process group and session signals need to be delivered to just the
* parent before the fork or both the parent and the child after the
diff --git a/kernel/rseq.c b/kernel/rseq.c
new file mode 100644
index 0000000..e1c847b
--- /dev/null
+++ b/kernel/rseq.c
@@ -0,0 +1,231 @@
+/*
+ * Restartable sequences system call
+ *
+ * Restartable sequences are a lightweight interface that allows
+ * user-level code to be executed atomically relative to scheduler
+ * preemption and signal delivery. Typically used for implementing
+ * per-cpu operations.
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ *
+ * This program is distributed in the hope that it will be useful,
+ * but WITHOUT ANY WARRANTY; without even the implied warranty of
+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
+ * GNU General Public License for more details.
+ *
+ * Copyright (C) 2015, Google, Inc.,
+ * Paul Turner <pjt@google.com> and Andrew Hunter <ahh@google.com>
+ * Copyright (C) 2015-2016, EfficiOS Inc.,
+ * Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
+ */
+
+#include <linux/sched.h>
+#include <linux/uaccess.h>
+#include <linux/syscalls.h>
+#include <linux/compat.h>
+#include <linux/rseq.h>
+#include <asm/ptrace.h>
+
+/*
+ * Each restartable sequence assembly block defines a "struct rseq_cs"
+ * structure which describes the post_commit_ip address, and the
+ * abort_ip address where the kernel should move the thread instruction
+ * pointer if a rseq critical section assembly block is preempted or if
+ * a signal is delivered on top of a rseq critical section assembly
+ * block. It also contains a start_ip, which is the address of the start
+ * of the rseq assembly block, which is useful to debuggers.
+ *
+ * The algorithm for a restartable sequence assembly block is as
+ * follows:
+ *
+ * rseq_start()
+ *
+ * 0. Userspace loads the current event counter value from the
+ * event_counter field of the registered struct rseq TLS area,
+ *
+ * rseq_finish()
+ *
+ * Steps [1]-[3] (inclusive) need to be a sequence of instructions in
+ * userspace that can handle being moved to the abort_ip between any
+ * of those instructions.
+ *
+ * The abort_ip address needs to be equal or above the post_commit_ip.
+ * Step [4] and the failure code step [F1] need to be at addresses
+ * equal or above the post_commit_ip.
+ *
+ * 1. Userspace stores the address of the struct rseq cs rseq
+ * assembly block descriptor into the rseq_cs field of the
+ * registered struct rseq TLS area.
+ *
+ * 2. Userspace tests to see whether the current event counter values
+ * match those loaded at [0]. Manually jumping to [F1] in case of
+ * a mismatch.
+ *
+ * Note that if we are preempted or interrupted by a signal
+ * after [1] and before post_commit_ip, then the kernel also
+ * performs the comparison performed in [2], and conditionally
+ * clears rseq_cs, then jumps us to abort_ip.
+ *
+ * 3. Userspace critical section final instruction before
+ * post_commit_ip is the commit. The critical section is
+ * self-terminating.
+ * [post_commit_ip]
+ *
+ * 4. Userspace clears the rseq_cs field of the struct rseq
+ * TLS area.
+ *
+ * 5. Return true.
+ *
+ * On failure at [2]:
+ *
+ * F1. Userspace clears the rseq_cs field of the struct rseq
+ * TLS area. Followed by step [F2].
+ *
+ * [abort_ip]
+ * F2. Return false.
+ */
+
+static int rseq_increment_event_counter(struct task_struct *t)
+{
+ if (__put_user(++t->rseq_event_counter,
+ &t->rseq->u.e.event_counter))
+ return -1;
+ return 0;
+}
+
+static int rseq_get_rseq_cs(struct task_struct *t,
+ void __user **post_commit_ip,
+ void __user **abort_ip)
+{
+ unsigned long ptr;
+ struct rseq_cs __user *rseq_cs;
+
+ if (__get_user(ptr, &t->rseq->rseq_cs))
+ return -1;
+ if (!ptr)
+ return 0;
+#ifdef CONFIG_COMPAT
+ if (in_compat_syscall()) {
+ rseq_cs = compat_ptr((compat_uptr_t)ptr);
+ if (get_user(ptr, &rseq_cs->post_commit_ip))
+ return -1;
+ *post_commit_ip = compat_ptr((compat_uptr_t)ptr);
+ if (get_user(ptr, &rseq_cs->abort_ip))
+ return -1;
+ *abort_ip = compat_ptr((compat_uptr_t)ptr);
+ return 0;
+ }
+#endif
+ rseq_cs = (struct rseq_cs __user *)ptr;
+ if (get_user(ptr, &rseq_cs->post_commit_ip))
+ return -1;
+ *post_commit_ip = (void __user *)ptr;
+ if (get_user(ptr, &rseq_cs->abort_ip))
+ return -1;
+ *abort_ip = (void __user *)ptr;
+ return 0;
+}
+
+static int rseq_ip_fixup(struct pt_regs *regs)
+{
+ struct task_struct *t = current;
+ void __user *post_commit_ip = NULL;
+ void __user *abort_ip = NULL;
+
+ if (rseq_get_rseq_cs(t, &post_commit_ip, &abort_ip))
+ return -1;
+
+ /* Handle potentially being within a critical section. */
+ if ((void __user *)instruction_pointer(regs) < post_commit_ip) {
+ /*
+ * We need to clear rseq_cs upon entry into a signal
+ * handler nested on top of a rseq assembly block, so
+ * the signal handler will not be fixed up if itself
+ * interrupted by a nested signal handler or preempted.
+ */
+ if (clear_user(&t->rseq->rseq_cs,
+ sizeof(t->rseq->rseq_cs)))
+ return -1;
+
+ /*
+ * We set this after potentially failing in
+ * clear_user so that the signal arrives at the
+ * faulting rip.
+ */
+ instruction_pointer_set(regs, (unsigned long)abort_ip);
+ }
+ return 0;
+}
+
+/*
+ * This resume handler should always be executed between any of:
+ * - preemption,
+ * - signal delivery,
+ * and return to user-space.
+ */
+void __rseq_handle_notify_resume(struct pt_regs *regs)
+{
+ struct task_struct *t = current;
+
+ if (unlikely(t->flags & PF_EXITING))
+ return;
+ if (!access_ok(VERIFY_WRITE, t->rseq, sizeof(*t->rseq)))
+ goto error;
+ if (__put_user(raw_smp_processor_id(), &t->rseq->u.e.cpu_id))
+ goto error;
+ if (rseq_increment_event_counter(t))
+ goto error;
+ if (rseq_ip_fixup(regs))
+ goto error;
+ return;
+
+error:
+ force_sig(SIGSEGV, t);
+}
+
+/*
+ * sys_rseq - setup restartable sequences for caller thread.
+ */
+SYSCALL_DEFINE2(rseq, struct rseq __user *, rseq, int, flags)
+{
+ if (unlikely(flags))
+ return -EINVAL;
+ if (!rseq) {
+ if (!current->rseq)
+ return -ENOENT;
+ return 0;
+ }
+
+ if (current->rseq) {
+ /*
+ * If rseq is already registered, check whether
+ * the provided address differs from the prior
+ * one.
+ */
+ if (current->rseq != rseq)
+ return -EBUSY;
+ } else {
+ /*
+ * If there was no rseq previously registered,
+ * we need to ensure the provided rseq is
+ * properly aligned and valid.
+ */
+ if (!IS_ALIGNED((unsigned long)rseq, sizeof(uint64_t)))
+ return -EINVAL;
+ if (!access_ok(VERIFY_WRITE, rseq, sizeof(*rseq)))
+ return -EFAULT;
+ current->rseq = rseq;
+ /*
+ * If rseq was previously inactive, and has just
+ * been registered, ensure the cpu_id and
+ * event_counter fields are updated before
+ * returning to user-space.
+ */
+ rseq_set_notify_resume(current);
+ }
+
+ return 0;
+}
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 51d7105..fbef0c3 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -2664,6 +2664,7 @@ prepare_task_switch(struct rq *rq, struct task_struct *prev,
{
sched_info_switch(rq, prev, next);
perf_event_task_sched_out(prev, next);
+ rseq_sched_out(prev);
fire_sched_out_preempt_notifiers(prev, next);
prepare_lock_switch(rq, next);
prepare_arch_switch(next);
diff --git a/kernel/sys_ni.c b/kernel/sys_ni.c
index 2c5e3a8..c653f78 100644
--- a/kernel/sys_ni.c
+++ b/kernel/sys_ni.c
@@ -250,3 +250,6 @@ cond_syscall(sys_execveat);
/* membarrier */
cond_syscall(sys_membarrier);
+
+/* restartable sequence */
+cond_syscall(sys_rseq);
--
2.1.4
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-07-26 01:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <rYXax-2Cq-1@gated-at.bofh.it> |
| In reply to | #1448163 |
On Thu, Jul 21, 2016 at 2:14 PM, Mathieu Desnoyers <mathieu.desnoyers@efficios.com> wrote: > Man page associated: > > RSEQ(2) Linux Programmer's Manual RSEQ(2) > > NAME > rseq - Restartable sequences and cpu number cache > > SYNOPSIS > #include <linux/rseq.h> > > int rseq(struct rseq * rseq, int flags); > > DESCRIPTION > The rseq() ABI accelerates user-space operations on per-cpu > data by defining a shared data structure ABI between each user- > space thread and the kernel. > > The rseq argument is a pointer to the thread-local rseq struc‐ > ture to be shared between kernel and user-space. A NULL rseq > value can be used to check whether rseq is registered for the > current thread. > > The layout of struct rseq is as follows: > > Structure alignment > This structure needs to be aligned on multiples of 64 > bytes. > > Structure size > This structure has a fixed size of 128 bytes. > > Fields > > cpu_id > Cache of the CPU number on which the calling thread is > running. > > event_counter > Restartable sequences event_counter field. That's an unhelpful description. > > rseq_cs > Restartable sequences rseq_cs field. Points to a struct > rseq_cs. Why is it a pointer? > > The layout of struct rseq_cs is as follows: > > Structure alignment > This structure needs to be aligned on multiples of 64 > bytes. > > Structure size > This structure has a fixed size of 192 bytes. > > Fields > > start_ip > Instruction pointer address of the first instruction of > the sequence of consecutive assembly instructions. > > post_commit_ip > Instruction pointer address after the last instruction > of the sequence of consecutive assembly instructions. > > abort_ip > Instruction pointer address where to move the execution > flow in case of abort of the sequence of consecutive > assembly instructions. > > The flags argument is currently unused and must be specified as > 0. > > Typically, a library or application will keep the rseq struc‐ > ture in a thread-local storage variable, or other memory areas "variable or other memory area" > belonging to each thread. It is recommended to perform volatile > reads of the thread-local cache to prevent the compiler from > doing load tearing. An alternative approach is to read each > field from inline assembly. I don't think the man page needs to tell people how to implement correct atomic loads. > > Each thread is responsible for registering its rseq structure. > Only one rseq structure address can be registered per thread. > Once set, the rseq address is idempotent for a given thread. "Idempotent" is a property that applies to an action, and the "rseq address" is not an action. I don't know what you're trying to say. > > In a typical usage scenario, the thread registering the rseq > structure will be performing loads and stores from/to that > structure. It is however also allowed to read that structure > from other threads. The rseq field updates performed by the > kernel provide single-copy atomicity semantics, which guarantee > that other threads performing single-copy atomic reads of the > cpu number cache will always observe a consistent value. s/single-copy/relaxed atomic/ perhaps? > > Memory registered as rseq structure should never be deallocated > before the thread which registered it exits: specifically, it > should not be freed, and the library containing the registered > thread-local storage should not be dlclose'd. Violating this > constraint may cause a SIGSEGV signal to be delivered to the > thread. That's an unfortunate constraint for threads that exit without help. > > Unregistration of associated rseq structure is implicitly per‐ > formed when a thread or process exit. exits. [...] Can you please document what this thing does prior to giving an example of how to use it. Hmm, here are the docs, sort of: > diff --git a/kernel/rseq.c b/kernel/rseq.c > new file mode 100644 > index 0000000..e1c847b > --- /dev/null > +++ b/kernel/rseq.c > +/* > + * Each restartable sequence assembly block defines a "struct rseq_cs" > + * structure which describes the post_commit_ip address, and the > + * abort_ip address where the kernel should move the thread instruction > + * pointer if a rseq critical section assembly block is preempted or if > + * a signal is delivered on top of a rseq critical section assembly > + * block. It also contains a start_ip, which is the address of the start > + * of the rseq assembly block, which is useful to debuggers. > + * > + * The algorithm for a restartable sequence assembly block is as > + * follows: > + * > + * rseq_start() > + * > + * 0. Userspace loads the current event counter value from the > + * event_counter field of the registered struct rseq TLS area, > + * > + * rseq_finish() > + * > + * Steps [1]-[3] (inclusive) need to be a sequence of instructions in > + * userspace that can handle being moved to the abort_ip between any > + * of those instructions. > + * > + * The abort_ip address needs to be equal or above the post_commit_ip. > + * Step [4] and the failure code step [F1] need to be at addresses > + * equal or above the post_commit_ip. > + * > + * 1. Userspace stores the address of the struct rseq cs rseq "struct rseq cs rseq" contains a typo. > + * assembly block descriptor into the rseq_cs field of the > + * registered struct rseq TLS area. > + * > + * 2. Userspace tests to see whether the current event counter values > + * match those loaded at [0]. Manually jumping to [F1] in case of > + * a mismatch. Grammar issues here. More importantly, you said "values", but you only described one value. > + * > + * Note that if we are preempted or interrupted by a signal > + * after [1] and before post_commit_ip, then the kernel also > + * performs the comparison performed in [2], and conditionally > + * clears rseq_cs, then jumps us to abort_ip. This is the first I've heard of rseq_cs being something that gets changed as a result of using this facility. What code sets it in the first place? I think you've also mentioned "preemption" and "migration". Which do you mean? > + * > + * 3. Userspace critical section final instruction before > + * post_commit_ip is the commit. The critical section is > + * self-terminating. > + * [post_commit_ip] > + * > + * 4. Userspace clears the rseq_cs field of the struct rseq > + * TLS area. > + * > + * 5. Return true. > + * > + * On failure at [2]: > + * A major issue I have with percpu critical sections or rseqs or whatever you want to call them is that, every time I talk to someone about them, there are a different set of requirements that they are supposed to satisfy. So: What problem does this solve? What are its atomicity properties? Under what conditions does it work? What assumptions does it make? What real-world operations become faster as a result of rseq (as opposed to just cpu number queries)? Why is it important for the kernel to do something special on every preemption? What "events" does "event_counter" count and why? If I'm understanding the intent of this code correctly (which is a big if), I think you're trying to do this: start a critical section; compute something; commit; if (commit worked) return; else try again; where "commit;" is a single instruction. The kernel guarantees that if the thread is preempted (or migrated, perhaps?) between the start and commit steps then commit will be forced to fail (or be skipped entirely). Because I don't understand what you're doing with this primitive, I can't really tell why you need to detect preemption as opposed to just migration. For example: would the following primitive solve the same problem? begin_dont_migrate_me() figure out what store to do to take the percpu lock; do that store; if (end_dont_migrate_me()) return; // oops, the kernel migrated us. retry. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-07-26 05:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <rZ0UN-53b-11@gated-at.bofh.it> |
| In reply to | #1450241 |
----- On Jul 25, 2016, at 7:02 PM, Andy Lutomirski luto@amacapital.net wrote: > On Thu, Jul 21, 2016 at 2:14 PM, Mathieu Desnoyers > <mathieu.desnoyers@efficios.com> wrote: >> Man page associated: >> >> RSEQ(2) Linux Programmer's Manual RSEQ(2) >> >> NAME >> rseq - Restartable sequences and cpu number cache >> >> SYNOPSIS >> #include <linux/rseq.h> >> >> int rseq(struct rseq * rseq, int flags); >> >> DESCRIPTION >> The rseq() ABI accelerates user-space operations on per-cpu >> data by defining a shared data structure ABI between each user- >> space thread and the kernel. >> >> The rseq argument is a pointer to the thread-local rseq struc‐ >> ture to be shared between kernel and user-space. A NULL rseq >> value can be used to check whether rseq is registered for the >> current thread. >> >> The layout of struct rseq is as follows: >> >> Structure alignment >> This structure needs to be aligned on multiples of 64 >> bytes. >> >> Structure size >> This structure has a fixed size of 128 bytes. >> >> Fields >> >> cpu_id >> Cache of the CPU number on which the calling thread is >> running. >> >> event_counter >> Restartable sequences event_counter field. > > That's an unhelpful description. Good point, how about: event_counter Counter guaranteed to be incremented when the current thread is preempted or when a signal is delivered to the current thread. In that same line of thoughts, I would reword cpu_id as: cpu_id Cache of the CPU number on which the current thread is running. > >> >> rseq_cs >> Restartable sequences rseq_cs field. Points to a struct >> rseq_cs. > > Why is it a pointer? Rewording like this should help understand: rseq_cs The rseq_cs field is a pointer to a struct rseq_cs. Is is NULL when no rseq assembly block critical section is active for the current thread. Setting it to point to a critical section descriptor (struct rseq_cs) marks the beginning of the critical section. It is cleared after the end of the critical section. > >> >> The layout of struct rseq_cs is as follows: >> >> Structure alignment >> This structure needs to be aligned on multiples of 64 >> bytes. >> >> Structure size >> This structure has a fixed size of 192 bytes. >> >> Fields >> >> start_ip >> Instruction pointer address of the first instruction of >> the sequence of consecutive assembly instructions. >> >> post_commit_ip >> Instruction pointer address after the last instruction >> of the sequence of consecutive assembly instructions. >> >> abort_ip >> Instruction pointer address where to move the execution >> flow in case of abort of the sequence of consecutive >> assembly instructions. >> >> The flags argument is currently unused and must be specified as >> 0. >> >> Typically, a library or application will keep the rseq struc‐ >> ture in a thread-local storage variable, or other memory areas > > "variable or other memory area" ok > >> belonging to each thread. It is recommended to perform volatile >> reads of the thread-local cache to prevent the compiler from >> doing load tearing. An alternative approach is to read each >> field from inline assembly. > > I don't think the man page needs to tell people how to implement > correct atomic loads. ok, I can remove the two previous sentences. > >> >> Each thread is responsible for registering its rseq structure. >> Only one rseq structure address can be registered per thread. >> Once set, the rseq address is idempotent for a given thread. > > "Idempotent" is a property that applies to an action, and the "rseq > address" is not an action. I don't know what you're trying to say. I mean there is only one address registered per thread, and it stays registered for the life-time of the thread. Perhaps I could say: "Once set, the rseq address never changes for a given thread." > >> >> In a typical usage scenario, the thread registering the rseq >> structure will be performing loads and stores from/to that >> structure. It is however also allowed to read that structure >> from other threads. The rseq field updates performed by the >> kernel provide single-copy atomicity semantics, which guarantee >> that other threads performing single-copy atomic reads of the >> cpu number cache will always observe a consistent value. > > s/single-copy/relaxed atomic/ perhaps? ok > >> >> Memory registered as rseq structure should never be deallocated >> before the thread which registered it exits: specifically, it >> should not be freed, and the library containing the registered >> thread-local storage should not be dlclose'd. Violating this >> constraint may cause a SIGSEGV signal to be delivered to the >> thread. > > That's an unfortunate constraint for threads that exit without help. I don't understand what you are pointing at here. I see this mostly as a constraint on the life-time of the library that holds the struct rseq TLS more than a constraint on the thread life-time. > >> >> Unregistration of associated rseq structure is implicitly per‐ >> formed when a thread or process exit. > > exits. ok > > [...] > > Can you please document what this thing does prior to giving an > example of how to use it. Good point, will do. (more comments on what can be added as documentation below) > > Hmm, here are the docs, sort of: > >> diff --git a/kernel/rseq.c b/kernel/rseq.c >> new file mode 100644 >> index 0000000..e1c847b >> --- /dev/null >> +++ b/kernel/rseq.c > >> +/* >> + * Each restartable sequence assembly block defines a "struct rseq_cs" >> + * structure which describes the post_commit_ip address, and the >> + * abort_ip address where the kernel should move the thread instruction >> + * pointer if a rseq critical section assembly block is preempted or if >> + * a signal is delivered on top of a rseq critical section assembly >> + * block. It also contains a start_ip, which is the address of the start >> + * of the rseq assembly block, which is useful to debuggers. >> + * >> + * The algorithm for a restartable sequence assembly block is as >> + * follows: >> + * >> + * rseq_start() >> + * >> + * 0. Userspace loads the current event counter value from the >> + * event_counter field of the registered struct rseq TLS area, >> + * >> + * rseq_finish() >> + * >> + * Steps [1]-[3] (inclusive) need to be a sequence of instructions in >> + * userspace that can handle being moved to the abort_ip between any >> + * of those instructions. >> + * >> + * The abort_ip address needs to be equal or above the post_commit_ip. >> + * Step [4] and the failure code step [F1] need to be at addresses >> + * equal or above the post_commit_ip. >> + * >> + * 1. Userspace stores the address of the struct rseq cs rseq > > "struct rseq cs rseq" contains a typo. should be "struct rseq_cs" > >> + * assembly block descriptor into the rseq_cs field of the >> + * registered struct rseq TLS area. >> + * >> + * 2. Userspace tests to see whether the current event counter values >> + * match those loaded at [0]. Manually jumping to [F1] in case of >> + * a mismatch. > > Grammar issues here. More importantly, you said "values", but you > only described one value. Indeed, values -> value, and those -> the value > >> + * >> + * Note that if we are preempted or interrupted by a signal >> + * after [1] and before post_commit_ip, then the kernel also >> + * performs the comparison performed in [2], and conditionally >> + * clears rseq_cs, then jumps us to abort_ip. > > This is the first I've heard of rseq_cs being something that gets > changed as a result of using this facility. What code sets it in the > first place? struct rseq_cs (the critical section descriptor) is statically declared, never changes. What I should clarify above is that the rseq_cs field of struct rseq gets cleared (not the struct rseq_cs per se). The struct rseq_cs field is initially at NULL, and is populated by the struct rseq_cs descriptor address when entering the critical section. It is set back to NULL right after exiting the critical section, through both the success and failure paths. > > I think you've also mentioned "preemption" and "migration". Which do you mean? We really care about preemption here. Every migration implies a preemption from a user-space perspective. If we would only care about keeping the CPU id up-to-date, hooking into migration would be enough. But since we want atomicity guarantees for restartable sequences, we need to hook into preemption. I should update the changelog of patch 1/7 to specify that we really do hook on preemption, even for the cpu_id update part. > >> + * >> + * 3. Userspace critical section final instruction before >> + * post_commit_ip is the commit. The critical section is >> + * self-terminating. >> + * [post_commit_ip] >> + * >> + * 4. Userspace clears the rseq_cs field of the struct rseq >> + * TLS area. >> + * >> + * 5. Return true. >> + * >> + * On failure at [2]: >> + * > > A major issue I have with percpu critical sections or rseqs or > whatever you want to call them is that, every time I talk to someone > about them, there are a different set of requirements that they are > supposed to satisfy. So: > > What problem does this solve? It allows user-space to perform update operations on per-cpu data without requiring heavy-weight atomic operations. > > What are its atomicity properties? Under what conditions does it > work? What assumptions does it make? Restartable sequences are atomic with respect to preemption (making it atomic with respect to other threads running on the same CPU), as well as signal delivery (user-space execution contexts nested over the same thread). It is suited for update operations on per-cpu data. It can be used on data structures shared between threads within a process, and on data structures shared between threads across different processes. > > What real-world operations become faster as a result of rseq (as > opposed to just cpu number queries)? A few examples of operations accelerated: - incrementing per-cpu counters, - per-cpu spin-lock, - per-cpu linked-lists (including memory allocator free-list), - per-cpu ring buffer, Perhaps others will have other operations in mind ? Note that compared to Paul Turner's patchset, I removed the percpu_cmpxchg and percpu_cmpxchg_check APIs from the test program rseq.h in user-space, because I found out that it was difficult to guarantee progress with those APIs. The do_rseq() approach, which does 2 attempts and falls back to locking, does provide progress guarantees even in the face of (unlikely) frequent migrations. > > Why is it important for the kernel to do something special on every preemption? This is how we can ensure that the entire critical section, consisting of both the C part and the assembly instruction sequence, will issue the commit instruction only if executed atomically with respect to other threads scheduled on the same CPU. > > What "events" does "event_counter" count and why? Technically, it increments each time a thread returns to user-space with the NOTIFY_RESUME thread flag set. We ensure to set this flag on preemption (out), as well as signal delivery. So it is guaranteed to increment when either of those events take place. It can however increment due to other kernel code setting TIF_NOTIFY_RESUME before returning to user-space. It is meant to allow user-space to detect preemption and signal delivery, not to count the exact number of such events. > > > If I'm understanding the intent of this code correctly (which is a big > if), I think you're trying to do this: > > start a critical section; > compute something; > commit; > if (commit worked) > return; > else > try again; > > where "commit;" is a single instruction. The kernel guarantees that > if the thread is preempted (or migrated, perhaps?) A thread needs to have been preempted in order to be migrated, so from a user-space perspective, detecting preemption is a super-set of detecting migration. We track preemption and signal delivery here. > between the start > and commit steps then commit will be forced to fail (or be skipped > entirely). Because I don't understand what you're doing with this > primitive, I can't really tell why you need to detect preemption as > opposed to just migration. > > For example: would the following primitive solve the same problem? > > begin_dont_migrate_me() > > figure out what store to do to take the percpu lock; > do that store; > > if (end_dont_migrate_me()) > return; > > // oops, the kernel migrated us. retry. First, prohibiting migration from user-space has been frowned upon by scheduler developers for a long time, and I doubt this mindset will change. But if we look at it from the point of view of letting user-space retry when it detects migration (rather than preemption), it would require that we use an atomic instruction (although without the lock prefix) as the commit instruction to ensure atomicity with respect to other threads running on the same CPU. Detecting preemption instead allows us to use a simple store instruction as the commit. Simple store instructions (e.g. mov) are faster than atomic instructions (e.g. xadd, cmpxchg...). Moreover, detecting migrations and using atomic instructions as commit is prone to ABA (e.g. free-list use-case) that are prevented by the restart on preemption or signal delivery. Thanks for looking into it! Mathieu > > > --Andy -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-08-03 14:30 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s23t8-2tz-17@gated-at.bofh.it> |
| In reply to | #1450341 |
On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote: > We really care about preemption here. Every migration implies a > preemption from a user-space perspective. If we would only care > about keeping the CPU id up-to-date, hooking into migration would be > enough. But since we want atomicity guarantees for restartable > sequences, we need to hook into preemption. > It allows user-space to perform update operations on per-cpu data without > requiring heavy-weight atomic operations. Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86. It is however on PPC and possibly other architectures, so in name of simplicity supporting only the one variant makes sense.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-08-03 18:50 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s27wJ-4Vz-13@gated-at.bofh.it> |
| In reply to | #1455783 |
On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote: > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote: >> We really care about preemption here. Every migration implies a >> preemption from a user-space perspective. If we would only care >> about keeping the CPU id up-to-date, hooking into migration would be >> enough. But since we want atomicity guarantees for restartable >> sequences, we need to hook into preemption. > >> It allows user-space to perform update operations on per-cpu data without >> requiring heavy-weight atomic operations. > > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86. > > It is however on PPC and possibly other architectures, so in name of > simplicity supporting only the one variant makes sense. > I wouldn't want to depend on CMPXCHG. But imagine we had primitives that were narrower than the full abort-on-preemption primitive. Specifically, suppose we had abort if (actual cpu != expected_cpu || *aptr != aval). We could do things like: expected_cpu = cpu; aval = NULL; // disarm for now begin(); aval = event_count[cpu] + 1; event_count[cpu] = aval; event_count[cpu]++; ... compute something ... // arm the rest of it aptr = &event_count[cpu]; if (*aptr != aval) goto fail; *thing_im_writing = value_i_computed; end(); The idea here is that we don't rely on the scheduler to increment the event count at all, which means that we get to determine the scope of what kinds of access conflicts we care about ourselves. This has an obvious downside: it's more complicated. It has several benefits, I think. It's debuggable without hassle (unless someone, accidentally or otherwise, sets aval incorrectly). It also allows much longer critical sections to work well, as merely being preempted in the middle won't cause an abort any more. So I'm hoping to understand whether we could make something like this work. This whole thing is roughly equivalent to abort-if-migrated plus an atomic "if (*aptr == aval) *b = c;" operation. (I think that, if this worked, we could improve it a bit by making the abort operation jump back to the "if (*aptr != aval) goto fail;" code, which should reduce the scope for error a bit and also reduces the need for extra code paths that only execute on an abort.)
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-08-03 20:40 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s29fc-6dM-21@gated-at.bofh.it> |
| In reply to | #1455893 |
On Wed, 3 Aug 2016, Andy Lutomirski wrote: > > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86. > > > > It is however on PPC and possibly other architectures, so in name of > > simplicity supporting only the one variant makes sense. > > > > I wouldn't want to depend on CMPXCHG. But imagine we had primitives > that were narrower than the full abort-on-preemption primitive. > Specifically, suppose we had abort if (actual cpu != expected_cpu || > *aptr != aval). We could do things like: > The latency issues that are addressed by restartable sequences require minimim instruction overhead. Lockless CMPXCHG is very important in that area and I would not simply remove it from consideration.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-08-04 07:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s2j4R-4sT-3@gated-at.bofh.it> |
| In reply to | #1455952 |
On Aug 3, 2016 11:31 AM, "Christoph Lameter" <cl@linux.com> wrote: > > On Wed, 3 Aug 2016, Andy Lutomirski wrote: > > > > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86. > > > > > > It is however on PPC and possibly other architectures, so in name of > > > simplicity supporting only the one variant makes sense. > > > > > > > I wouldn't want to depend on CMPXCHG. But imagine we had primitives > > that were narrower than the full abort-on-preemption primitive. > > Specifically, suppose we had abort if (actual cpu != expected_cpu || > > *aptr != aval). We could do things like: > > > > The latency issues that are addressed by restartable sequences require > minimim instruction overhead. Lockless CMPXCHG is very important in that > area and I would not simply remove it from consideration. What I mean is: I think the solution shouldn't depend on the x86-specific unlocked CMPXCHG instruction if it can be avoided.
[toc] | [prev] | [next] | [standalone]
| From | Boqun Feng <boqun.feng@gmail.com> |
|---|---|
| Date | 2016-08-04 06:30 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s2is9-3ZW-3@gated-at.bofh.it> |
| In reply to | #1455893 |
[Multipart message — attachments visible in raw view] — view raw
On Wed, Aug 03, 2016 at 09:37:57AM -0700, Andy Lutomirski wrote:
> On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote:
> >> We really care about preemption here. Every migration implies a
> >> preemption from a user-space perspective. If we would only care
> >> about keeping the CPU id up-to-date, hooking into migration would be
> >> enough. But since we want atomicity guarantees for restartable
> >> sequences, we need to hook into preemption.
> >
> >> It allows user-space to perform update operations on per-cpu data without
> >> requiring heavy-weight atomic operations.
> >
> > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86.
> >
> > It is however on PPC and possibly other architectures, so in name of
> > simplicity supporting only the one variant makes sense.
> >
>
> I wouldn't want to depend on CMPXCHG. But imagine we had primitives
> that were narrower than the full abort-on-preemption primitive.
> Specifically, suppose we had abort if (actual cpu != expected_cpu ||
> *aptr != aval). We could do things like:
>
> expected_cpu = cpu;
> aval = NULL; // disarm for now
> begin();
> aval = event_count[cpu] + 1;
> event_count[cpu] = aval;
> event_count[cpu]++;
This line is redundant, right? Because it will guarantee a failure even
in no-contention cases.
>
> ... compute something ...
>
> // arm the rest of it
> aptr = &event_count[cpu];
> if (*aptr != aval)
> goto fail;
>
> *thing_im_writing = value_i_computed;
> end();
>
> The idea here is that we don't rely on the scheduler to increment the
> event count at all, which means that we get to determine the scope of
> what kinds of access conflicts we care about ourselves.
>
If we increase the event count in userspace, how could we prevent two
userspace threads from racing on the event_count[cpu] field? For
example:
CPU 0
================
{event_count[0] is initially 0}
[Thread 1]
begin();
aval = event_count[cpu] + 1; // 1
(preempted)
[Thread 2]
begin();
aval = event_count[cpu] + 1; // 1, too
event_count[cpu] = aval; // event_count[0] is 1
(preempted)
[Thread 1]
event_count[cpu] = aval; // event_count[0] is 1
...
aptr = &event_count[cpu];
if (*aptr != aval) // false.
...
[Thread 2]
aptr = &event_count[cpu];
if (*aptr != aval) // false.
...
, in which case, both the critical sections are successful, and Thread 1
and Thread 2 will race on *thing_im_writing.
Am I missing your point here?
Regards,
Boqun
> This has an obvious downside: it's more complicated.
>
> It has several benefits, I think. It's debuggable without hassle
> (unless someone, accidentally or otherwise, sets aval incorrectly).
> It also allows much longer critical sections to work well, as merely
> being preempted in the middle won't cause an abort any more.
>
> So I'm hoping to understand whether we could make something like this
> work. This whole thing is roughly equivalent to abort-if-migrated
> plus an atomic "if (*aptr == aval) *b = c;" operation.
>
> (I think that, if this worked, we could improve it a bit by making the
> abort operation jump back to the "if (*aptr != aval) goto fail;" code,
> which should reduce the scope for error a bit and also reduces the
> need for extra code paths that only execute on an abort.)
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-08-04 07:20 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s2jex-4wK-15@gated-at.bofh.it> |
| In reply to | #1456148 |
On Wed, Aug 3, 2016 at 9:27 PM, Boqun Feng <boqun.feng@gmail.com> wrote:
> On Wed, Aug 03, 2016 at 09:37:57AM -0700, Andy Lutomirski wrote:
>> On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote:
>> >> We really care about preemption here. Every migration implies a
>> >> preemption from a user-space perspective. If we would only care
>> >> about keeping the CPU id up-to-date, hooking into migration would be
>> >> enough. But since we want atomicity guarantees for restartable
>> >> sequences, we need to hook into preemption.
>> >
>> >> It allows user-space to perform update operations on per-cpu data without
>> >> requiring heavy-weight atomic operations.
>> >
>> > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86.
>> >
>> > It is however on PPC and possibly other architectures, so in name of
>> > simplicity supporting only the one variant makes sense.
>> >
>>
>> I wouldn't want to depend on CMPXCHG. But imagine we had primitives
>> that were narrower than the full abort-on-preemption primitive.
>> Specifically, suppose we had abort if (actual cpu != expected_cpu ||
>> *aptr != aval). We could do things like:
>>
>> expected_cpu = cpu;
>> aval = NULL; // disarm for now
>> begin();
>> aval = event_count[cpu] + 1;
>> event_count[cpu] = aval;
>> event_count[cpu]++;
>
> This line is redundant, right? Because it will guarantee a failure even
> in no-contention cases.
>
>>
>> ... compute something ...
>>
>> // arm the rest of it
>> aptr = &event_count[cpu];
>> if (*aptr != aval)
>> goto fail;
>>
>> *thing_im_writing = value_i_computed;
>> end();
>>
>> The idea here is that we don't rely on the scheduler to increment the
>> event count at all, which means that we get to determine the scope of
>> what kinds of access conflicts we care about ourselves.
>>
>
> If we increase the event count in userspace, how could we prevent two
> userspace threads from racing on the event_count[cpu] field? For
> example:
>
> CPU 0
> ================
> {event_count[0] is initially 0}
>
> [Thread 1]
> begin();
> aval = event_count[cpu] + 1; // 1
>
> (preempted)
> [Thread 2]
> begin();
> aval = event_count[cpu] + 1; // 1, too
> event_count[cpu] = aval; // event_count[0] is 1
>
You're right :( This would work with an xadd instruction, but that's
very slow and doesn't exist on most architectures. It could also work
if we did:
aval = some_tls_value++;
where some_tls_value is set up such that no two threads could ever end
up with the same values (using high bits as thread ids, perhaps), but
that's messy. Maybe my idea is no good.
[toc] | [prev] | [next] | [standalone]
| From | Boqun Feng <boqun.feng@gmail.com> |
|---|---|
| Date | 2016-08-09 18:20 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4hV0-1tf-31@gated-at.bofh.it> |
| In reply to | #1456171 |
[Multipart message — attachments visible in raw view] — view raw
On Wed, Aug 03, 2016 at 10:03:32PM -0700, Andy Lutomirski wrote:
> On Wed, Aug 3, 2016 at 9:27 PM, Boqun Feng <boqun.feng@gmail.com> wrote:
> > On Wed, Aug 03, 2016 at 09:37:57AM -0700, Andy Lutomirski wrote:
> >> On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote:
> >> > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote:
> >> >> We really care about preemption here. Every migration implies a
> >> >> preemption from a user-space perspective. If we would only care
> >> >> about keeping the CPU id up-to-date, hooking into migration would be
> >> >> enough. But since we want atomicity guarantees for restartable
> >> >> sequences, we need to hook into preemption.
> >> >
> >> >> It allows user-space to perform update operations on per-cpu data without
> >> >> requiring heavy-weight atomic operations.
> >> >
> >> > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86.
> >> >
> >> > It is however on PPC and possibly other architectures, so in name of
> >> > simplicity supporting only the one variant makes sense.
> >> >
> >>
> >> I wouldn't want to depend on CMPXCHG. But imagine we had primitives
> >> that were narrower than the full abort-on-preemption primitive.
> >> Specifically, suppose we had abort if (actual cpu != expected_cpu ||
> >> *aptr != aval). We could do things like:
> >>
> >> expected_cpu = cpu;
> >> aval = NULL; // disarm for now
> >> begin();
> >> aval = event_count[cpu] + 1;
> >> event_count[cpu] = aval;
> >> event_count[cpu]++;
> >
> > This line is redundant, right? Because it will guarantee a failure even
> > in no-contention cases.
> >
> >>
> >> ... compute something ...
> >>
> >> // arm the rest of it
> >> aptr = &event_count[cpu];
> >> if (*aptr != aval)
> >> goto fail;
> >>
> >> *thing_im_writing = value_i_computed;
> >> end();
> >>
> >> The idea here is that we don't rely on the scheduler to increment the
> >> event count at all, which means that we get to determine the scope of
> >> what kinds of access conflicts we care about ourselves.
> >>
> >
> > If we increase the event count in userspace, how could we prevent two
> > userspace threads from racing on the event_count[cpu] field? For
> > example:
> >
> > CPU 0
> > ================
> > {event_count[0] is initially 0}
> >
> > [Thread 1]
> > begin();
> > aval = event_count[cpu] + 1; // 1
> >
> > (preempted)
> > [Thread 2]
> > begin();
> > aval = event_count[cpu] + 1; // 1, too
> > event_count[cpu] = aval; // event_count[0] is 1
> >
>
> You're right :( This would work with an xadd instruction, but that's
> very slow and doesn't exist on most architectures. It could also work
> if we did:
>
> aval = some_tls_value++;
>
> where some_tls_value is set up such that no two threads could ever end
> up with the same values (using high bits as thread ids, perhaps), but
> that's messy. Maybe my idea is no good.
This is a little more complex, plus I failed to find a way to do an
atomic "if (*aptr == aval) *b = c" in userspace ;-(
However, I'm thinking maybe we can use some tricks to avoid unnecessary
aborts-on-preemption.
First of all, I notice we haven't make any constraint on what kind of
memory objects could be "protected" by rseq critical sections yet. And I
think this is something we should decide before adding this feature into
kernel.
We can do some optimization if we have some constraints. For example, if
the memory objects inside the rseq critical sections could only be
modified by userspace programs, we therefore don't need to abort
immediately when userspace task -> kernel task context switch.
Further more, if the memory objects inside the rseq critical sections
could only be modified by userspace programs that have registered their
rseq structures, we don't need to abort immediately between the context
switches between two rseq-unregistered tasks or one rseq-registered
task and one rseq-unregistered task.
Instead, we do tricks as follow:
defining a percpu pointer in kernel:
DEFINE_PER_CPU(struct task_struct *, rseq_owner);
and a cpu field in struct task_struct:
struct task_struct {
...
#ifdef CONFIG_RSEQ
struct rseq __user *rseq;
uint32_t rseq_event_counter;
int rseq_cpu;
#endif
...
};
(task_struct::rseq_cpu should be initialized as -1.)
each time at sched out(in rseq_sched_out()), we do something like:
if (prev->rseq) {
raw_cpu_write(rseq_owner, prev);
prev->rseq_cpu = smp_processor_id();
}
each time sched in(in rseq_handle_notify_resume()), we do something
like:
if (current->rseq &&
(this_cpu_read(rseq_owner) != current ||
current->rseq_cpu != smp_processor_id()))
__rseq_handle_notify_resume(regs);
(Also need to modify rseq_signal_deliver() to call
__rseq_handle_notify_resume() directly).
I think this could save some unnecessary aborts-on-preemption, however,
TBH, I'm too sleepy to verify every corner case. Will recheck this
tomorrow.
Regards,
Boqun
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-08-10 20:20 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4GgG-79-47@gated-at.bofh.it> |
| In reply to | #1458909 |
----- On Aug 10, 2016, at 4:01 AM, Andy Lutomirski luto@amacapital.net wrote: > On Tue, Aug 9, 2016 at 9:13 AM, Boqun Feng <boqun.feng@gmail.com> wrote: <snip> > >> However, I'm thinking maybe we can use some tricks to avoid unnecessary >> aborts-on-preemption. >> >> First of all, I notice we haven't make any constraint on what kind of >> memory objects could be "protected" by rseq critical sections yet. And I >> think this is something we should decide before adding this feature into >> kernel. >> >> We can do some optimization if we have some constraints. For example, if >> the memory objects inside the rseq critical sections could only be >> modified by userspace programs, we therefore don't need to abort >> immediately when userspace task -> kernel task context switch. > > True, although trying to do a syscall in an rseq critical section > seems like a bad idea in general. The scenario above does not require the rseq critical section to perform an explicit system call. It can happen from simple timer-driven preemption of user-space. <snip> > > But do we need to protect MAP_SHARED objects? If not, maybe we could > only track context switches between different tasks sharing the same > mm. I have tracing use-cases involving MAP_SHARED objects for rseq: per-cpu buffers. Moreover, if you only track context switch between tasks with the same mm, you run into issues if you have: Process A Thread 1 (rseq) Thread 2 (rseq) Process B Thread 1 Scheduling: A.1 -> B.1 -> A.2 -> B.1 -> A.1 There is no scheduling between threads of the same process here, but the entire chain involves two threads of the same process accessing the same per-cpu data concurrently. Thanks, Mathieu > > --Andy -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-08-10 21:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4GgG-79-49@gated-at.bofh.it> |
| In reply to | #1458909 |
On Tue, Aug 9, 2016 at 9:13 AM, Boqun Feng <boqun.feng@gmail.com> wrote:
> On Wed, Aug 03, 2016 at 10:03:32PM -0700, Andy Lutomirski wrote:
>> On Wed, Aug 3, 2016 at 9:27 PM, Boqun Feng <boqun.feng@gmail.com> wrote:
>> > On Wed, Aug 03, 2016 at 09:37:57AM -0700, Andy Lutomirski wrote:
>> >> On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>> >> > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote:
>> >> >> We really care about preemption here. Every migration implies a
>> >> >> preemption from a user-space perspective. If we would only care
>> >> >> about keeping the CPU id up-to-date, hooking into migration would be
>> >> >> enough. But since we want atomicity guarantees for restartable
>> >> >> sequences, we need to hook into preemption.
>> >> >
>> >> >> It allows user-space to perform update operations on per-cpu data without
>> >> >> requiring heavy-weight atomic operations.
>> >> >
>> >> > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86.
>> >> >
>> >> > It is however on PPC and possibly other architectures, so in name of
>> >> > simplicity supporting only the one variant makes sense.
>> >> >
>> >>
>> >> I wouldn't want to depend on CMPXCHG. But imagine we had primitives
>> >> that were narrower than the full abort-on-preemption primitive.
>> >> Specifically, suppose we had abort if (actual cpu != expected_cpu ||
>> >> *aptr != aval). We could do things like:
>> >>
>> >> expected_cpu = cpu;
>> >> aval = NULL; // disarm for now
>> >> begin();
>> >> aval = event_count[cpu] + 1;
>> >> event_count[cpu] = aval;
>> >> event_count[cpu]++;
>> >
>> > This line is redundant, right? Because it will guarantee a failure even
>> > in no-contention cases.
>> >
>> >>
>> >> ... compute something ...
>> >>
>> >> // arm the rest of it
>> >> aptr = &event_count[cpu];
>> >> if (*aptr != aval)
>> >> goto fail;
>> >>
>> >> *thing_im_writing = value_i_computed;
>> >> end();
>> >>
>> >> The idea here is that we don't rely on the scheduler to increment the
>> >> event count at all, which means that we get to determine the scope of
>> >> what kinds of access conflicts we care about ourselves.
>> >>
>> >
>> > If we increase the event count in userspace, how could we prevent two
>> > userspace threads from racing on the event_count[cpu] field? For
>> > example:
>> >
>> > CPU 0
>> > ================
>> > {event_count[0] is initially 0}
>> >
>> > [Thread 1]
>> > begin();
>> > aval = event_count[cpu] + 1; // 1
>> >
>> > (preempted)
>> > [Thread 2]
>> > begin();
>> > aval = event_count[cpu] + 1; // 1, too
>> > event_count[cpu] = aval; // event_count[0] is 1
>> >
>>
>> You're right :( This would work with an xadd instruction, but that's
>> very slow and doesn't exist on most architectures. It could also work
>> if we did:
>>
>> aval = some_tls_value++;
>>
>> where some_tls_value is set up such that no two threads could ever end
>> up with the same values (using high bits as thread ids, perhaps), but
>> that's messy. Maybe my idea is no good.
>
> This is a little more complex, plus I failed to find a way to do an
> atomic "if (*aptr == aval) *b = c" in userspace ;-(
>
But the kernel might be able to help using something similar to this patchset.
> However, I'm thinking maybe we can use some tricks to avoid unnecessary
> aborts-on-preemption.
>
> First of all, I notice we haven't make any constraint on what kind of
> memory objects could be "protected" by rseq critical sections yet. And I
> think this is something we should decide before adding this feature into
> kernel.
>
> We can do some optimization if we have some constraints. For example, if
> the memory objects inside the rseq critical sections could only be
> modified by userspace programs, we therefore don't need to abort
> immediately when userspace task -> kernel task context switch.
True, although trying to do a syscall in an rseq critical section
seems like a bad idea in general.
>
> Further more, if the memory objects inside the rseq critical sections
> could only be modified by userspace programs that have registered their
> rseq structures, we don't need to abort immediately between the context
> switches between two rseq-unregistered tasks or one rseq-registered
> task and one rseq-unregistered task.
>
> Instead, we do tricks as follow:
>
> defining a percpu pointer in kernel:
>
> DEFINE_PER_CPU(struct task_struct *, rseq_owner);
>
> and a cpu field in struct task_struct:
>
> struct task_struct {
> ...
> #ifdef CONFIG_RSEQ
> struct rseq __user *rseq;
> uint32_t rseq_event_counter;
> int rseq_cpu;
> #endif
> ...
> };
>
> (task_struct::rseq_cpu should be initialized as -1.)
>
> each time at sched out(in rseq_sched_out()), we do something like:
>
> if (prev->rseq) {
> raw_cpu_write(rseq_owner, prev);
> prev->rseq_cpu = smp_processor_id();
> }
>
> each time sched in(in rseq_handle_notify_resume()), we do something
> like:
>
> if (current->rseq &&
> (this_cpu_read(rseq_owner) != current ||
> current->rseq_cpu != smp_processor_id()))
> __rseq_handle_notify_resume(regs);
>
> (Also need to modify rseq_signal_deliver() to call
> __rseq_handle_notify_resume() directly).
>
>
> I think this could save some unnecessary aborts-on-preemption, however,
> TBH, I'm too sleepy to verify every corner case. Will recheck this
> tomorrow.
Interesting. That could help a bit, although it would help less if
everyone started using rseq.
But do we need to protect MAP_SHARED objects? If not, maybe we could
only track context switches between different tasks sharing the same
mm.
--Andy
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-08-10 22:50 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4IBQ-1DK-25@gated-at.bofh.it> |
| In reply to | #1458909 |
----- On Aug 9, 2016, at 12:13 PM, Boqun Feng boqun.feng@gmail.com wrote:
<snip>
>
> However, I'm thinking maybe we can use some tricks to avoid unnecessary
> aborts-on-preemption.
>
> First of all, I notice we haven't make any constraint on what kind of
> memory objects could be "protected" by rseq critical sections yet. And I
> think this is something we should decide before adding this feature into
> kernel.
>
> We can do some optimization if we have some constraints. For example, if
> the memory objects inside the rseq critical sections could only be
> modified by userspace programs, we therefore don't need to abort
> immediately when userspace task -> kernel task context switch.
The rseq_owner per-cpu variable and rseq_cpu field in task_struct you
propose below would indeed take care of this scenario.
>
> Further more, if the memory objects inside the rseq critical sections
> could only be modified by userspace programs that have registered their
> rseq structures, we don't need to abort immediately between the context
> switches between two rseq-unregistered tasks or one rseq-registered
> task and one rseq-unregistered task.
>
> Instead, we do tricks as follow:
>
> defining a percpu pointer in kernel:
>
> DEFINE_PER_CPU(struct task_struct *, rseq_owner);
>
> and a cpu field in struct task_struct:
>
> struct task_struct {
> ...
> #ifdef CONFIG_RSEQ
> struct rseq __user *rseq;
> uint32_t rseq_event_counter;
> int rseq_cpu;
> #endif
> ...
> };
>
> (task_struct::rseq_cpu should be initialized as -1.)
>
> each time at sched out(in rseq_sched_out()), we do something like:
>
> if (prev->rseq) {
> raw_cpu_write(rseq_owner, prev);
> prev->rseq_cpu = smp_processor_id();
> }
>
> each time sched in(in rseq_handle_notify_resume()), we do something
> like:
>
> if (current->rseq &&
> (this_cpu_read(rseq_owner) != current ||
> current->rseq_cpu != smp_processor_id()))
> __rseq_handle_notify_resume(regs);
>
> (Also need to modify rseq_signal_deliver() to call
> __rseq_handle_notify_resume() directly).
>
>
> I think this could save some unnecessary aborts-on-preemption, however,
> TBH, I'm too sleepy to verify every corner case. Will recheck this
> tomorrow.
This adds extra fields to the task struct, per-cpu rseq_owner pointers,
and hooks into sched_in which are not needed otherwise, all this to
eliminate unneeded abort-on-preemption.
If we look at the single-stepping use-case, this means that gdb would
only be able to single-step applications as long as neither itself, nor
any of its libraries, use rseq. This seems to be quite fragile. I prefer
requiring rseq users to implement a fallback to locking which progresses
in every situation rather than adding complexity and overhead trying
lessen the odds of triggering the restart.
Simply lessening the odds of triggering the restart without a design that
ensures progress even in restart cases seems to make the lack-of-progress
problem just harder to debug when it will surface in real life.
Thanks,
Mathieu
>
> Regards,
> Boqun
--
Mathieu Desnoyers
EfficiOS Inc.
http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-08-10 21:40 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4Hw6-YV-19@gated-at.bofh.it> |
| In reply to | #1456171 |
On Wed, Aug 3, 2016 at 10:03 PM, Andy Lutomirski <luto@amacapital.net> wrote:
> On Wed, Aug 3, 2016 at 9:27 PM, Boqun Feng <boqun.feng@gmail.com> wrote:
>> On Wed, Aug 03, 2016 at 09:37:57AM -0700, Andy Lutomirski wrote:
>>> On Wed, Aug 3, 2016 at 5:27 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>>> > On Tue, Jul 26, 2016 at 03:02:19AM +0000, Mathieu Desnoyers wrote:
>>> >> We really care about preemption here. Every migration implies a
>>> >> preemption from a user-space perspective. If we would only care
>>> >> about keeping the CPU id up-to-date, hooking into migration would be
>>> >> enough. But since we want atomicity guarantees for restartable
>>> >> sequences, we need to hook into preemption.
>>> >
>>> >> It allows user-space to perform update operations on per-cpu data without
>>> >> requiring heavy-weight atomic operations.
>>> >
>>> > Well, a CMPXCHG without LOCK prefix isn't all that expensive on x86.
>>> >
>>> > It is however on PPC and possibly other architectures, so in name of
>>> > simplicity supporting only the one variant makes sense.
>>> >
>>>
>>> I wouldn't want to depend on CMPXCHG. But imagine we had primitives
>>> that were narrower than the full abort-on-preemption primitive.
>>> Specifically, suppose we had abort if (actual cpu != expected_cpu ||
>>> *aptr != aval). We could do things like:
>>>
>>> expected_cpu = cpu;
>>> aval = NULL; // disarm for now
>>> begin();
>>> aval = event_count[cpu] + 1;
>>> event_count[cpu] = aval;
>>> event_count[cpu]++;
>>
>> This line is redundant, right? Because it will guarantee a failure even
>> in no-contention cases.
>>
>>>
>>> ... compute something ...
>>>
>>> // arm the rest of it
>>> aptr = &event_count[cpu];
>>> if (*aptr != aval)
>>> goto fail;
>>>
>>> *thing_im_writing = value_i_computed;
>>> end();
>>>
>>> The idea here is that we don't rely on the scheduler to increment the
>>> event count at all, which means that we get to determine the scope of
>>> what kinds of access conflicts we care about ourselves.
>>>
>>
>> If we increase the event count in userspace, how could we prevent two
>> userspace threads from racing on the event_count[cpu] field? For
>> example:
>>
>> CPU 0
>> ================
>> {event_count[0] is initially 0}
>>
>> [Thread 1]
>> begin();
>> aval = event_count[cpu] + 1; // 1
>>
>> (preempted)
>> [Thread 2]
>> begin();
>> aval = event_count[cpu] + 1; // 1, too
>> event_count[cpu] = aval; // event_count[0] is 1
>>
>
> You're right :( This would work with an xadd instruction, but that's
> very slow and doesn't exist on most architectures. It could also work
> if we did:
Thinking about this slightly more, maybe it does work. We could use
basically the same mechanism to allow the kernel to restart if the
specific sequence:
aval = event_count[cpu] + 1
event_count[cpu] = avall
gets preempted by setting aptr = &event_count[cpu] and aval to
event_count[cpu], like this (although I might have screwed up any
number of small details):
aptr = &event_count[cpu];
barrier();
aval = event_count[cpu];
barrier();
tmp = aval + 1;
event_count[cpu] = tmp;
/* preemption here will cause an unnecessary retry, but that's okay */
aval = tmp;
--Andy
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-08-03 20:40 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s29fb-6dM-7@gated-at.bofh.it> |
| In reply to | #1450341 |
On Tue, 26 Jul 2016, Mathieu Desnoyers wrote: > > What problem does this solve? > > It allows user-space to perform update operations on per-cpu data without > requiring heavy-weight atomic operations. This is great but seems to indicate that such a facility would be better for kernel code instread of user space code. > First, prohibiting migration from user-space has been frowned upon > by scheduler developers for a long time, and I doubt this mindset will > change. Note that the task isolation patchset from Chris Metcalf does something that goes a long way towards this. If you set strict isolation mode then the kernel will terminate the process or notify you if the scheduler becomes involved. In some way we are getting that as a side effect. Also prohibiting migration is trivial form user space. Just do a taskset to a single cpu.
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2016-08-10 22:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4HZ7-1pa-3@gated-at.bofh.it> |
| In reply to | #1455950 |
----- On Aug 3, 2016, at 2:29 PM, Chris Lameter cl@linux.com wrote: > On Tue, 26 Jul 2016, Mathieu Desnoyers wrote: > >> > What problem does this solve? >> >> It allows user-space to perform update operations on per-cpu data without >> requiring heavy-weight atomic operations. > > > This is great but seems to indicate that such a facility would be better > for kernel code instread of user space code. It would be interesting to eventually investigate whether rseq is additionally useful for kernel code. It seems unrelated to its usefulness for user-space code though. Rseq for user-space only needs to hook into preemption and signal delivery, which doesn't seem to have measurable effects on overall performance. Doing rseq for kernel code would imply hooking into supplementary sites: - preemption of kernel code (for atomicity wrt other threads). This would replace preempt_disable()/preempt_enable() critical sections touching per-cpu data shared with other threads. We would have to do the event_counter increment and ip fixup directly in the sched_out hook when preempting kernel code. - possibly interrupt handlers (for atomicity wrt interrupts). This would replace local irq save/restore when touching per-cpu data shared with interrupt handlers. We would have to increment the event_counter and fixup on the pre-irq kernel frame. - possibly NMI handlers (for atomicity wrt NMIs). This would replace preempt/irq off protected local atomic operations on per-cpu data shared with NMIs. We would have to increment the event_counter and fixup on the pre-NMI kernel frame. Those supplementary hooks may add significant overall performance overhead, so careful benchmarking would be required to figure out if it's worth it. > >> First, prohibiting migration from user-space has been frowned upon >> by scheduler developers for a long time, and I doubt this mindset will >> change. > > Note that the task isolation patchset from Chris Metcalf does something > that goes a long way towards this. If you set strict isolation mode then > the kernel will terminate the process or notify you if the scheduler > becomes involved. In some way we are getting that as a side effect. AFAIU, what you propose here is doable at the application design level. We want to introduce rseq to speed up memory allocation, tracing, and other uses of per-cpu data without having to modify the design of each and every user-space applications out there. > Also prohibiting migration is trivial form user space. Just do a taskset > to a single cpu. This is also possible if you can redesign user-space applications, but not from a library perspective. Invoking system calls to change the affinity of a thread at each and every critical section would kill performance. Setting the affinity of a thread from a library on behalf of the application and leaving it affined requires changes to the application design. Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2016-08-10 22:40 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <s4Isb-1A3-63@gated-at.bofh.it> |
| In reply to | #1459760 |
On Wed, 10 Aug 2016, Mathieu Desnoyers wrote: > - preemption of kernel code (for atomicity wrt other threads). This would > replace preempt_disable()/preempt_enable() critical sections touching > per-cpu data shared with other threads. We would have to do the event_counter > increment and ip fixup directly in the sched_out hook when preempting > kernel code. What we would need is special handling when returning from a context switch so that we recognize in what type of code section we are in and continue execution at the proper retry site. This can be done by putting code into special sections or other methods that do not require additional coee. > - possibly interrupt handlers (for atomicity wrt interrupts). This would > replace local irq save/restore when touching per-cpu data shared with > interrupt handlers. We would have to increment the event_counter and > fixup on the pre-irq kernel frame. Same thing as before. Test if we are in a section by testing the return address and then maybe continue elsewhere. > Those supplementary hooks may add significant overall performance overhead, > so careful benchmarking would be required to figure out if it's worth it. We need a design that does not need these hooks. If we check the return IP address for a special range then we would not need those. Any hooks would bloat the code in such a way that the implementation would not be acceptable for the kernel code.
[toc] | [prev] | [next] | [standalone]
| From | Boqun Feng <boqun.feng@gmail.com> |
|---|---|
| Date | 2016-07-27 17:10 +0200 |
| Subject | Re: [RFC PATCH v7 1/7] Restartable sequences system call |
| Message-ID | <rZyD7-12s-15@gated-at.bofh.it> |
| In reply to | #1448163 |
[Multipart message — attachments visible in raw view] — view raw
Hi Mathieu,
On Thu, Jul 21, 2016 at 05:14:16PM -0400, Mathieu Desnoyers wrote:
> Expose a new system call allowing each thread to register one userspace
> memory area to be used as an ABI between kernel and user-space for two
> purposes: user-space restartable sequences and quick access to read the
> current CPU number value from user-space.
>
> * Restartable sequences (per-cpu atomics)
>
> The restartable critical sections (percpu atomics) work has been started
> by Paul Turner and Andrew Hunter. It lets the kernel handle restart of
> critical sections. [1] [2] The re-implementation proposed here brings a
> few simplifications to the ABI which facilitates porting to other
Agreed ;-)
> architectures and speeds up the user-space fast path. A locking-based
> fall-back, purely implemented in user-space, is proposed here to deal
> with debugger single-stepping. This fallback interacts with rseq_start()
> and rseq_finish(), which force retries in response to concurrent
> lock-based activity.
>
So I have enabled this on powerpc, thanks to your nice work to make
things easy for porting ;-)
A patchset will follow in-reply-to this email, which includes patches
enabling this on powerpc and a patch that improves the portability of
the selftests, which I think it's not necessary to be a standalone
patch, so it's OK to be merged into your patch #7.
I did some tests on 64bit little/big endian pSeries(guest) kernel with
selftest cases(64bit LE selftest on 64bit LE kernel, 64/32bit BE
selftest on 64bit BE kernel), things seemingly went well ;-)
Here are some benchmark results I got on a little endian guest with 64
VCPUs:
Benchmarking various approaches for reading the current CPU number:
Power8 PSeries Guest(64 VCPUs, the host has 16 cores, 128 hardware
threads):
- Baseline (empty loop): 1.56 ns
- Read CPU from rseq cpu_id: 1.56 ns
- Read CPU from rseq cpu_id (lazy register): 2.08 ns
- glibc 2.23-0ubuntu3 getcpu: 7.72 ns
- getcpu system call: 91.80 ns
Benchmarking various approaches for counter increment:
Power8 PSeries KVM Guest(64 VCPUs, the host has 16 cores, 128 hardware
threads):
Counter increment speed (ns/increment)
1 thread 2 threads 4 threads 8 threads 16 threads 32 threads
global increment (baseline) 6.5 N/A N/A N/A N/A N/A
percpu rseq increment 6.9 6.9 7.2 7.3 15.4 35.5
percpu rseq spinlock 19.0 18.9 19.4 19.4 35.5 71.8
global atomic increment 25.8 111.0 261.0 905.2 2319.5 4170.5 (__sync_add_and_fetch_4)
global atomic CAS 26.2 119.0 341.6 1183.0 3951.3 9312.5 (__sync_val_compare_and_swap_4)
global pthread mutex 40.0 238.1 644.0 2052.2 4272.5 8612.2
I surely need to run more tests for my patches in different
environments, and will try to adjust the patchset according to whatever
change you make(e.g. rseq_finish2) in the future.
(Add PPC maintainers in Cc)
Regards,
Boqun
> Here are benchmarks of counter increment in various scenarios compared
> to restartable sequences:
>
> ARMv7 Processor rev 4 (v7l)
> Machine model: Cubietruck
>
> Counter increment speed (ns/increment)
> 1 thread 2 threads
> global increment (baseline) 6 N/A
> percpu rseq increment 50 52
> percpu rseq spinlock 94 94
> global atomic increment 48 74 (__sync_add_and_fetch_4)
> global atomic CAS 50 172 (__sync_val_compare_and_swap_4)
> global pthread mutex 148 862
>
> ARMv7 Processor rev 10 (v7l)
> Machine model: Wandboard
>
> Counter increment speed (ns/increment)
> 1 thread 4 threads
> global increment (baseline) 7 N/A
> percpu rseq increment 50 50
> percpu rseq spinlock 82 84
> global atomic increment 44 262 (__sync_add_and_fetch_4)
> global atomic CAS 46 316 (__sync_val_compare_and_swap_4)
> global pthread mutex 146 1400
>
> x86-64 Intel(R) Xeon(R) CPU E5-2630 v3 @ 2.40GHz:
>
> Counter increment speed (ns/increment)
> 1 thread 8 threads
> global increment (baseline) 3.0 N/A
> percpu rseq increment 3.6 3.8
> percpu rseq spinlock 5.6 6.2
> global LOCK; inc 8.0 166.4
> global LOCK; cmpxchg 13.4 435.2
> global pthread mutex 25.2 1363.6
>
> * Reading the current CPU number
>
> Speeding up reading the current CPU number on which the caller thread is
> running is done by keeping the current CPU number up do date within the
> cpu_id field of the memory area registered by the thread. This is done
> by making scheduler migration set the TIF_NOTIFY_RESUME flag on the
> current thread. Upon return to user-space, a notify-resume handler
> updates the current CPU value within the registered user-space memory
> area. User-space can then read the current CPU number directly from
> memory.
>
> Keeping the current cpu id in a memory area shared between kernel and
> user-space is an improvement over current mechanisms available to read
> the current CPU number, which has the following benefits over
> alternative approaches:
>
> - 35x speedup on ARM vs system call through glibc
> - 20x speedup on x86 compared to calling glibc, which calls vdso
> executing a "lsl" instruction,
> - 14x speedup on x86 compared to inlined "lsl" instruction,
> - Unlike vdso approaches, this cpu_id value can be read from an inline
> assembly, which makes it a useful building block for restartable
> sequences.
> - The approach of reading the cpu id through memory mapping shared
> between kernel and user-space is portable (e.g. ARM), which is not the
> case for the lsl-based x86 vdso.
>
> On x86, yet another possible approach would be to use the gs segment
> selector to point to user-space per-cpu data. This approach performs
> similarly to the cpu id cache, but it has two disadvantages: it is
> not portable, and it is incompatible with existing applications already
> using the gs segment selector for other purposes.
>
> Benchmarking various approaches for reading the current CPU number:
>
> ARMv7 Processor rev 4 (v7l)
> Machine model: Cubietruck
> - Baseline (empty loop): 8.4 ns
> - Read CPU from rseq cpu_id: 16.7 ns
> - Read CPU from rseq cpu_id (lazy register): 19.8 ns
> - glibc 2.19-0ubuntu6.6 getcpu: 301.8 ns
> - getcpu system call: 234.9 ns
>
> x86-64 Intel(R) Xeon(R) CPU E5-2630 v3 @ 2.40GHz:
> - Baseline (empty loop): 0.8 ns
> - Read CPU from rseq cpu_id: 0.8 ns
> - Read CPU from rseq cpu_id (lazy register): 0.8 ns
> - Read using gs segment selector: 0.8 ns
> - "lsl" inline assembly: 13.0 ns
> - glibc 2.19-0ubuntu6 getcpu: 16.6 ns
> - getcpu system call: 53.9 ns
>
> - Speed
>
> Running 10 runs of hackbench -l 100000 seems to indicate, contrary to
> expectations, that enabling CONFIG_RSEQ slightly accelerates the
> scheduler:
>
> Configuration: 2 sockets * 8-core Intel(R) Xeon(R) CPU E5-2630 v3 @
> 2.40GHz (directly on hardware, hyperthreading disabled in BIOS, energy
> saving disabled in BIOS, turboboost disabled in BIOS, cpuidle.off=1
> kernel parameter), with a Linux v4.6 defconfig+localyesconfig,
> restartable sequences series applied.
>
[snip]
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web