Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1695159 > unrolled thread
| Started by | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| First post | 2017-07-25 00:00 +0200 |
| Last post | 2017-07-27 16:40 +0200 |
| Articles | 20 on this page of 56 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 00:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-25 06:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-25 15:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 18:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 19:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 19:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 21:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 21:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 22:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 23:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 00:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 20:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 20:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 22:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 23:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 03:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-27 14:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 11:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 12:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 10:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 17:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-26 11:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Page 1 of 3 [1] 2 3 Next page →
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 00:00 +0200 |
| Subject | [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option |
| Message-ID | <u6Tyq-2jq-15@gated-at.bofh.it> |
The sys_membarrier() system call has proven too slow for some use
cases, which has prompted users to instead rely on TLB shootdown.
Although TLB shootdown is much faster, it has the slight disadvantage
of not working at all on arm and arm64. This commit therefore adds
an expedited option to the sys_membarrier() system call.
Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
---
include/uapi/linux/membarrier.h | 11 +++++++++++
kernel/membarrier.c | 7 ++++++-
2 files changed, 17 insertions(+), 1 deletion(-)
diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
index e0b108bd2624..ba36d8a6be61 100644
--- a/include/uapi/linux/membarrier.h
+++ b/include/uapi/linux/membarrier.h
@@ -40,6 +40,16 @@
* (non-running threads are de facto in such a
* state). This covers threads from all processes
* running on the system. This command returns 0.
+ * @MEMBARRIER_CMD_SHARED_EXPEDITED: Execute a memory barrier on all
+ * running threads, but in an expedited fashion.
+ * Upon return from system call, the caller thread
+ * is ensured that all running threads have passed
+ * through a state where all memory accesses to
+ * user-space addresses match program order between
+ * entry to and return from the system call
+ * (non-running threads are de facto in such a
+ * state). This covers threads from all processes
+ * running on the system. This command returns 0.
*
* Command to be passed to the membarrier system call. The commands need to
* be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
@@ -48,6 +58,7 @@
enum membarrier_cmd {
MEMBARRIER_CMD_QUERY = 0,
MEMBARRIER_CMD_SHARED = (1 << 0),
+ MEMBARRIER_CMD_SHARED_EXPEDITED = (2 << 0),
};
#endif /* _UAPI_LINUX_MEMBARRIER_H */
diff --git a/kernel/membarrier.c b/kernel/membarrier.c
index 9f9284f37f8d..b749c39bb219 100644
--- a/kernel/membarrier.c
+++ b/kernel/membarrier.c
@@ -22,7 +22,8 @@
* Bitmask made from a "or" of all commands within enum membarrier_cmd,
* except MEMBARRIER_CMD_QUERY.
*/
-#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED)
+#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED | \
+ MEMBARRIER_CMD_SHARED_EXPEDITED)
/**
* sys_membarrier - issue memory barriers on a set of threads
@@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
if (num_online_cpus() > 1)
synchronize_sched();
return 0;
+ case MEMBARRIER_CMD_SHARED_EXPEDITED:
+ if (num_online_cpus() > 1)
+ synchronize_sched_expedited();
+ return 0;
default:
return -EINVAL;
}
--
2.5.2
[toc] | [next] | [standalone]
| From | Boqun Feng <boqun.feng@gmail.com> |
|---|---|
| Date | 2017-07-25 06:30 +0200 |
| Message-ID | <u6ZDP-6yq-3@gated-at.bofh.it> |
| In reply to | #1695159 |
[Multipart message — attachments visible in raw view] — view raw
On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote:
> The sys_membarrier() system call has proven too slow for some use
> cases, which has prompted users to instead rely on TLB shootdown.
> Although TLB shootdown is much faster, it has the slight disadvantage
> of not working at all on arm and arm64. This commit therefore adds
> an expedited option to the sys_membarrier() system call.
>
> Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> ---
> include/uapi/linux/membarrier.h | 11 +++++++++++
> kernel/membarrier.c | 7 ++++++-
> 2 files changed, 17 insertions(+), 1 deletion(-)
>
> diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
> index e0b108bd2624..ba36d8a6be61 100644
> --- a/include/uapi/linux/membarrier.h
> +++ b/include/uapi/linux/membarrier.h
> @@ -40,6 +40,16 @@
> * (non-running threads are de facto in such a
> * state). This covers threads from all processes
> * running on the system. This command returns 0.
> + * @MEMBARRIER_CMD_SHARED_EXPEDITED: Execute a memory barrier on all
> + * running threads, but in an expedited fashion.
> + * Upon return from system call, the caller thread
> + * is ensured that all running threads have passed
> + * through a state where all memory accesses to
> + * user-space addresses match program order between
> + * entry to and return from the system call
> + * (non-running threads are de facto in such a
> + * state). This covers threads from all processes
> + * running on the system. This command returns 0.
> *
> * Command to be passed to the membarrier system call. The commands need to
> * be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
> @@ -48,6 +58,7 @@
> enum membarrier_cmd {
> MEMBARRIER_CMD_QUERY = 0,
> MEMBARRIER_CMD_SHARED = (1 << 0),
> + MEMBARRIER_CMD_SHARED_EXPEDITED = (2 << 0),
Should this better be "(1 << 1)" ;-)
Regards,
Boqun
> };
>
> #endif /* _UAPI_LINUX_MEMBARRIER_H */
> diff --git a/kernel/membarrier.c b/kernel/membarrier.c
> index 9f9284f37f8d..b749c39bb219 100644
> --- a/kernel/membarrier.c
> +++ b/kernel/membarrier.c
> @@ -22,7 +22,8 @@
> * Bitmask made from a "or" of all commands within enum membarrier_cmd,
> * except MEMBARRIER_CMD_QUERY.
> */
> -#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED)
> +#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED | \
> + MEMBARRIER_CMD_SHARED_EXPEDITED)
>
> /**
> * sys_membarrier - issue memory barriers on a set of threads
> @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
> if (num_online_cpus() > 1)
> synchronize_sched();
> return 0;
> + case MEMBARRIER_CMD_SHARED_EXPEDITED:
> + if (num_online_cpus() > 1)
> + synchronize_sched_expedited();
> + return 0;
> default:
> return -EINVAL;
> }
> --
> 2.5.2
>
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 18:30 +0200 |
| Message-ID | <u7aSC-56H-15@gated-at.bofh.it> |
| In reply to | #1695369 |
On Tue, Jul 25, 2017 at 12:27:01PM +0800, Boqun Feng wrote:
> On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote:
> > The sys_membarrier() system call has proven too slow for some use
> > cases, which has prompted users to instead rely on TLB shootdown.
> > Although TLB shootdown is much faster, it has the slight disadvantage
> > of not working at all on arm and arm64. This commit therefore adds
> > an expedited option to the sys_membarrier() system call.
> >
> > Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> > ---
> > include/uapi/linux/membarrier.h | 11 +++++++++++
> > kernel/membarrier.c | 7 ++++++-
> > 2 files changed, 17 insertions(+), 1 deletion(-)
> >
> > diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
> > index e0b108bd2624..ba36d8a6be61 100644
> > --- a/include/uapi/linux/membarrier.h
> > +++ b/include/uapi/linux/membarrier.h
> > @@ -40,6 +40,16 @@
> > * (non-running threads are de facto in such a
> > * state). This covers threads from all processes
> > * running on the system. This command returns 0.
> > + * @MEMBARRIER_CMD_SHARED_EXPEDITED: Execute a memory barrier on all
> > + * running threads, but in an expedited fashion.
> > + * Upon return from system call, the caller thread
> > + * is ensured that all running threads have passed
> > + * through a state where all memory accesses to
> > + * user-space addresses match program order between
> > + * entry to and return from the system call
> > + * (non-running threads are de facto in such a
> > + * state). This covers threads from all processes
> > + * running on the system. This command returns 0.
> > *
> > * Command to be passed to the membarrier system call. The commands need to
> > * be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
> > @@ -48,6 +58,7 @@
> > enum membarrier_cmd {
> > MEMBARRIER_CMD_QUERY = 0,
> > MEMBARRIER_CMD_SHARED = (1 << 0),
> > + MEMBARRIER_CMD_SHARED_EXPEDITED = (2 << 0),
>
> Should this better be "(1 << 1)" ;-)
Same value, but yes, much more aligned with the intent. Good catch,
thank you, fixed!
Thanx, Paul
> Regards,
> Boqun
>
> > };
> >
> > #endif /* _UAPI_LINUX_MEMBARRIER_H */
> > diff --git a/kernel/membarrier.c b/kernel/membarrier.c
> > index 9f9284f37f8d..b749c39bb219 100644
> > --- a/kernel/membarrier.c
> > +++ b/kernel/membarrier.c
> > @@ -22,7 +22,8 @@
> > * Bitmask made from a "or" of all commands within enum membarrier_cmd,
> > * except MEMBARRIER_CMD_QUERY.
> > */
> > -#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED)
> > +#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED | \
> > + MEMBARRIER_CMD_SHARED_EXPEDITED)
> >
> > /**
> > * sys_membarrier - issue memory barriers on a set of threads
> > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
> > if (num_online_cpus() > 1)
> > synchronize_sched();
> > return 0;
> > + case MEMBARRIER_CMD_SHARED_EXPEDITED:
> > + if (num_online_cpus() > 1)
> > + synchronize_sched_expedited();
> > + return 0;
> > default:
> > return -EINVAL;
> > }
> > --
> > 2.5.2
> >
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-25 15:20 +0200 |
| Message-ID | <u77UJ-3gI-7@gated-at.bofh.it> |
| In reply to | #1695159 |
----- On Jul 24, 2017, at 5:58 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote:
> The sys_membarrier() system call has proven too slow for some use
> cases, which has prompted users to instead rely on TLB shootdown.
> Although TLB shootdown is much faster, it has the slight disadvantage
> of not working at all on arm and arm64. This commit therefore adds
> an expedited option to the sys_membarrier() system call.
Is this now possible because the synchronize_sched_expedited()
implementation does not require to send IPIs to all CPUS ? I
suspect that using tree srcu now solves this somehow, but can
you tell us a bit more about why it is now OK to expose this
to user-space ?
The commit message here does not explain why it is OK real-time
wise to expose this feature as a system call.
Thanks,
Mathieu
>
> Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> ---
> include/uapi/linux/membarrier.h | 11 +++++++++++
> kernel/membarrier.c | 7 ++++++-
> 2 files changed, 17 insertions(+), 1 deletion(-)
>
> diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
> index e0b108bd2624..ba36d8a6be61 100644
> --- a/include/uapi/linux/membarrier.h
> +++ b/include/uapi/linux/membarrier.h
> @@ -40,6 +40,16 @@
> * (non-running threads are de facto in such a
> * state). This covers threads from all processes
> * running on the system. This command returns 0.
> + * @MEMBARRIER_CMD_SHARED_EXPEDITED: Execute a memory barrier on all
> + * running threads, but in an expedited fashion.
> + * Upon return from system call, the caller thread
> + * is ensured that all running threads have passed
> + * through a state where all memory accesses to
> + * user-space addresses match program order between
> + * entry to and return from the system call
> + * (non-running threads are de facto in such a
> + * state). This covers threads from all processes
> + * running on the system. This command returns 0.
> *
> * Command to be passed to the membarrier system call. The commands need to
> * be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
> @@ -48,6 +58,7 @@
> enum membarrier_cmd {
> MEMBARRIER_CMD_QUERY = 0,
> MEMBARRIER_CMD_SHARED = (1 << 0),
> + MEMBARRIER_CMD_SHARED_EXPEDITED = (2 << 0),
> };
>
> #endif /* _UAPI_LINUX_MEMBARRIER_H */
> diff --git a/kernel/membarrier.c b/kernel/membarrier.c
> index 9f9284f37f8d..b749c39bb219 100644
> --- a/kernel/membarrier.c
> +++ b/kernel/membarrier.c
> @@ -22,7 +22,8 @@
> * Bitmask made from a "or" of all commands within enum membarrier_cmd,
> * except MEMBARRIER_CMD_QUERY.
> */
> -#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED)
> +#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED | \
> + MEMBARRIER_CMD_SHARED_EXPEDITED)
>
> /**
> * sys_membarrier - issue memory barriers on a set of threads
> @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
> if (num_online_cpus() > 1)
> synchronize_sched();
> return 0;
> + case MEMBARRIER_CMD_SHARED_EXPEDITED:
> + if (num_online_cpus() > 1)
> + synchronize_sched_expedited();
> + return 0;
> default:
> return -EINVAL;
> }
> --
> 2.5.2
--
Mathieu Desnoyers
EfficiOS Inc.
http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 18:50 +0200 |
| Message-ID | <u7bbZ-5du-19@gated-at.bofh.it> |
| In reply to | #1695722 |
On Tue, Jul 25, 2017 at 01:21:08PM +0000, Mathieu Desnoyers wrote:
> ----- On Jul 24, 2017, at 5:58 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote:
>
> > The sys_membarrier() system call has proven too slow for some use
> > cases, which has prompted users to instead rely on TLB shootdown.
> > Although TLB shootdown is much faster, it has the slight disadvantage
> > of not working at all on arm and arm64. This commit therefore adds
> > an expedited option to the sys_membarrier() system call.
>
> Is this now possible because the synchronize_sched_expedited()
> implementation does not require to send IPIs to all CPUS ? I
> suspect that using tree srcu now solves this somehow, but can
> you tell us a bit more about why it is now OK to expose this
> to user-space ?
I have gotten complaints from several users that sys_membarrier() is too
slow to be useful for them. So they are hacking around this problem by
unmapping a region of memory, thus getting the IPIs and memory barriers
on all CPUs, but with additional mm overhead. Plus this is non-portable,
and fragile with respect to reasonable optimizations, as was discussed
on LKML some time back:
https://marc.info/?l=linux-kernel&m=142619683526482
So we really need to make sys_membarrier() work for these users.
If we don't, we certainly will look quite silly criticizing their
use of invoking TLB shootdown via unmapping, now won't we?
Now back in 2015, expedited grace periods were horribly slow, but
I have optimized them to the point that it should be no worse than
TLB shootdown IPIs. Plus it is portable, and not subject to death
by optimization.
> The commit message here does not explain why it is OK real-time
> wise to expose this feature as a system call.
I figure that kernels providing that level of real-time response
will disable this, perhaps in a manner similar to that for NO_HZ_FULL.
Plus I intend to add your earlier IPI-all-threads-in-this-process
option, which will allow the people asking for this to do reasonable
testing.
Obviously, unless there are good test results and some level of user
enthusiasm, this patch goes nowhere.
Seem reasonable?
Thanx, Paul
> Thanks,
>
> Mathieu
>
>
> >
> > Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> > ---
> > include/uapi/linux/membarrier.h | 11 +++++++++++
> > kernel/membarrier.c | 7 ++++++-
> > 2 files changed, 17 insertions(+), 1 deletion(-)
> >
> > diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
> > index e0b108bd2624..ba36d8a6be61 100644
> > --- a/include/uapi/linux/membarrier.h
> > +++ b/include/uapi/linux/membarrier.h
> > @@ -40,6 +40,16 @@
> > * (non-running threads are de facto in such a
> > * state). This covers threads from all processes
> > * running on the system. This command returns 0.
> > + * @MEMBARRIER_CMD_SHARED_EXPEDITED: Execute a memory barrier on all
> > + * running threads, but in an expedited fashion.
> > + * Upon return from system call, the caller thread
> > + * is ensured that all running threads have passed
> > + * through a state where all memory accesses to
> > + * user-space addresses match program order between
> > + * entry to and return from the system call
> > + * (non-running threads are de facto in such a
> > + * state). This covers threads from all processes
> > + * running on the system. This command returns 0.
> > *
> > * Command to be passed to the membarrier system call. The commands need to
> > * be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
> > @@ -48,6 +58,7 @@
> > enum membarrier_cmd {
> > MEMBARRIER_CMD_QUERY = 0,
> > MEMBARRIER_CMD_SHARED = (1 << 0),
> > + MEMBARRIER_CMD_SHARED_EXPEDITED = (2 << 0),
> > };
> >
> > #endif /* _UAPI_LINUX_MEMBARRIER_H */
> > diff --git a/kernel/membarrier.c b/kernel/membarrier.c
> > index 9f9284f37f8d..b749c39bb219 100644
> > --- a/kernel/membarrier.c
> > +++ b/kernel/membarrier.c
> > @@ -22,7 +22,8 @@
> > * Bitmask made from a "or" of all commands within enum membarrier_cmd,
> > * except MEMBARRIER_CMD_QUERY.
> > */
> > -#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED)
> > +#define MEMBARRIER_CMD_BITMASK (MEMBARRIER_CMD_SHARED | \
> > + MEMBARRIER_CMD_SHARED_EXPEDITED)
> >
> > /**
> > * sys_membarrier - issue memory barriers on a set of threads
> > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
> > if (num_online_cpus() > 1)
> > synchronize_sched();
> > return 0;
> > + case MEMBARRIER_CMD_SHARED_EXPEDITED:
> > + if (num_online_cpus() > 1)
> > + synchronize_sched_expedited();
> > + return 0;
> > default:
> > return -EINVAL;
> > }
> > --
> > 2.5.2
>
> --
> Mathieu Desnoyers
> EfficiOS Inc.
> http://www.efficios.com
>
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-25 18:40 +0200 |
| Message-ID | <u7b2i-59O-19@gated-at.bofh.it> |
| In reply to | #1695159 |
On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote: > The sys_membarrier() system call has proven too slow for some use > cases, which has prompted users to instead rely on TLB shootdown. > Although TLB shootdown is much faster, it has the slight disadvantage > of not working at all on arm and arm64. This commit therefore adds > an expedited option to the sys_membarrier() system call. > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags) > if (num_online_cpus() > 1) > synchronize_sched(); > return 0; > + case MEMBARRIER_CMD_SHARED_EXPEDITED: > + if (num_online_cpus() > 1) > + synchronize_sched_expedited(); > + return 0; So you now give unprivileged userspace the means to IPI the entire machine? So what do we do when someone goes and does: for (;;) sys_membarrier(MEMBARRIER_CMD_SHARED_EXPEDITED, 0); on us?
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 18:50 +0200 |
| Message-ID | <u7bbY-5du-13@gated-at.bofh.it> |
| In reply to | #1695938 |
On Tue, Jul 25, 2017 at 06:33:18PM +0200, Peter Zijlstra wrote: > On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote: > > The sys_membarrier() system call has proven too slow for some use > > cases, which has prompted users to instead rely on TLB shootdown. > > Although TLB shootdown is much faster, it has the slight disadvantage > > of not working at all on arm and arm64. This commit therefore adds > > an expedited option to the sys_membarrier() system call. > > > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags) > > if (num_online_cpus() > 1) > > synchronize_sched(); > > return 0; > > + case MEMBARRIER_CMD_SHARED_EXPEDITED: > > + if (num_online_cpus() > 1) > > + synchronize_sched_expedited(); > > + return 0; > > So you now give unprivileged userspace the means to IPI the entire > machine? > > So what do we do when someone goes and does: > > for (;;) > sys_membarrier(MEMBARRIER_CMD_SHARED_EXPEDITED, 0); > > on us? The same thing that happens when they call munmap(). Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-25 19:10 +0200 |
| Message-ID | <u7bvk-5B8-5@gated-at.bofh.it> |
| In reply to | #1695944 |
On Tue, Jul 25, 2017 at 09:49:00AM -0700, Paul E. McKenney wrote: > On Tue, Jul 25, 2017 at 06:33:18PM +0200, Peter Zijlstra wrote: > > On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote: > > > The sys_membarrier() system call has proven too slow for some use > > > cases, which has prompted users to instead rely on TLB shootdown. > > > Although TLB shootdown is much faster, it has the slight disadvantage > > > of not working at all on arm and arm64. This commit therefore adds > > > an expedited option to the sys_membarrier() system call. > > > > > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags) > > > if (num_online_cpus() > 1) > > > synchronize_sched(); > > > return 0; > > > + case MEMBARRIER_CMD_SHARED_EXPEDITED: > > > + if (num_online_cpus() > 1) > > > + synchronize_sched_expedited(); > > > + return 0; > > > > So you now give unprivileged userspace the means to IPI the entire > > machine? > > > > So what do we do when someone goes and does: > > > > for (;;) > > sys_membarrier(MEMBARRIER_CMD_SHARED_EXPEDITED, 0); > > > > on us? > > The same thing that happens when they call munmap(). munmap() TLB invalidate is limited to those CPUs that actually ran threads of their process, while this is machine wide.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 19:20 +0200 |
| Message-ID | <u7bEZ-5Ep-5@gated-at.bofh.it> |
| In reply to | #1695952 |
On Tue, Jul 25, 2017 at 06:59:57PM +0200, Peter Zijlstra wrote: > On Tue, Jul 25, 2017 at 09:49:00AM -0700, Paul E. McKenney wrote: > > On Tue, Jul 25, 2017 at 06:33:18PM +0200, Peter Zijlstra wrote: > > > On Mon, Jul 24, 2017 at 02:58:16PM -0700, Paul E. McKenney wrote: > > > > The sys_membarrier() system call has proven too slow for some use > > > > cases, which has prompted users to instead rely on TLB shootdown. > > > > Although TLB shootdown is much faster, it has the slight disadvantage > > > > of not working at all on arm and arm64. This commit therefore adds > > > > an expedited option to the sys_membarrier() system call. > > > > > > > @@ -64,6 +65,10 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags) > > > > if (num_online_cpus() > 1) > > > > synchronize_sched(); > > > > return 0; > > > > + case MEMBARRIER_CMD_SHARED_EXPEDITED: > > > > + if (num_online_cpus() > 1) > > > > + synchronize_sched_expedited(); > > > > + return 0; > > > > > > So you now give unprivileged userspace the means to IPI the entire > > > machine? > > > > > > So what do we do when someone goes and does: > > > > > > for (;;) > > > sys_membarrier(MEMBARRIER_CMD_SHARED_EXPEDITED, 0); > > > > > > on us? > > > > The same thing that happens when they call munmap(). > > munmap() TLB invalidate is limited to those CPUs that actually ran > threads of their process, while this is machine wide. Or those CPUs running threads of any process mapping the underlying file or whatever. And in either case, this can span the whole machine. Plus there are a number of other ways for users to do on-demand full-system IPIs, including any number of ways to wake up large numbers of CPUs, including from unrelated processes. But I do plan to add another alternative that is limited to threads of the running process. I will be carrying both versions to enable those who have been bugging me about this to do testing. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-25 21:00 +0200 |
| Message-ID | <u7ddM-6sM-5@gated-at.bofh.it> |
| In reply to | #1695957 |
On Tue, Jul 25, 2017 at 10:17:01AM -0700, Paul E. McKenney wrote: > > munmap() TLB invalidate is limited to those CPUs that actually ran > > threads of their process, while this is machine wide. > > Or those CPUs running threads of any process mapping the underlying file > or whatever. That doesn't sound right. munmap() of a shared file only invalidates this process's map of it. Swapping a file page otoh will indeed touch the union of cpumasks over all processes mapping that page. > And in either case, this can span the whole machine. Plus > there are a number of other ways for users to do on-demand full-system > IPIs, including any number of ways to wake up large numbers of CPUs, > including from unrelated processes. Which are those? I thought we significantly reduced those with the nohz full work. Most IPI uses now first check if a CPU actually needs the IPI before sending it IIRC. > But I do plan to add another alternative that is limited to threads of > the running process. I will be carrying both versions to enable those > who have been bugging me about this to do testing. Sending IPIs to mm_cpumask() might be better than expedited, but I'm still hesitant. Just because people want it doesn't mean its a good idea. We need to weight this against the potential for abuse. People want userspace preempt disable, no matter how hard they want it, they're not getting it because its a completely crap idea.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 21:40 +0200 |
| Message-ID | <u7dQv-6W7-47@gated-at.bofh.it> |
| In reply to | #1696035 |
On Tue, Jul 25, 2017 at 08:53:20PM +0200, Peter Zijlstra wrote: > On Tue, Jul 25, 2017 at 10:17:01AM -0700, Paul E. McKenney wrote: > > > > munmap() TLB invalidate is limited to those CPUs that actually ran > > > threads of their process, while this is machine wide. > > > > Or those CPUs running threads of any process mapping the underlying file > > or whatever. > > That doesn't sound right. munmap() of a shared file only invalidates > this process's map of it. > > Swapping a file page otoh will indeed touch the union of cpumasks over > all processes mapping that page. There are a lot of variations, to be sure. For whatever it is worth, the original patch that started this uses mprotect(): https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0 > > And in either case, this can span the whole machine. Plus > > there are a number of other ways for users to do on-demand full-system > > IPIs, including any number of ways to wake up large numbers of CPUs, > > including from unrelated processes. > > Which are those? I thought we significantly reduced those with the nohz > full work. Most IPI uses now first check if a CPU actually needs the IPI > before sending it IIRC. If the task being awakened is higher priority than the task currently running on a given CPU, that CPU still gets an IPI, right? Or am I completely confused? > > But I do plan to add another alternative that is limited to threads of > > the running process. I will be carrying both versions to enable those > > who have been bugging me about this to do testing. > > Sending IPIs to mm_cpumask() might be better than expedited, but I'm > still hesitant. Just because people want it doesn't mean its a good > idea. We need to weight this against the potential for abuse. > > People want userspace preempt disable, no matter how hard they want it, > they're not getting it because its a completely crap idea. Unlike userspace preempt disable, in this case we get the abuse anyway via existing mechanisms, as in they are already being abused. If we provide a mechanism for this purpose, we at least have the potential for handling the abuse, for example: o "Defanging" sys_membarrier() on systems that are sensitive to latency. For example, this patch can be defanged by booting with the rcupdate.rcu_normal=1 kernel boot parameter, which causes requests for expedited grace periods to instead use normal grace periods. o Detecting and responding to abuse. For example, perhaps if there are more than (say) 50 expedited sys_membarrier()s within a given jiffy, the excess sys_membarrier()s are non-expedited. o Batching optimizations allow large number of concurrent requests to be handled with fewer grace periods -- and both normal and expedited grace periods already do exactly this. This horse is already out, so trying to shut the gate won't be effective. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-25 22:30 +0200 |
| Message-ID | <u7eCV-7wj-77@gated-at.bofh.it> |
| In reply to | #1696131 |
On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> There are a lot of variations, to be sure. For whatever it is worth,
> the original patch that started this uses mprotect():
>
> https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0
FWIW that will not work on s390 (and maybe others), they don't in fact
require IPIs for remote TLB invalidation.
> > Which are those? I thought we significantly reduced those with the nohz
> > full work. Most IPI uses now first check if a CPU actually needs the IPI
> > before sending it IIRC.
>
> If the task being awakened is higher priority than the task currently
> running on a given CPU, that CPU still gets an IPI, right? Or am I
> completely confused?
I was thinking of things like on_each_cpu_cond().
> Unlike userspace preempt disable, in this case we get the abuse anyway
> via existing mechanisms, as in they are already being abused. If we
> provide a mechanism for this purpose, we at least have the potential
> for handling the abuse, for example:
>
> o "Defanging" sys_membarrier() on systems that are sensitive to
> latency. For example, this patch can be defanged by booting
> with the rcupdate.rcu_normal=1 kernel boot parameter, which
> causes requests for expedited grace periods to instead use
> normal grace periods.
>
> o Detecting and responding to abuse. For example, perhaps if there
> are more than (say) 50 expedited sys_membarrier()s within a given
> jiffy, the excess sys_membarrier()s are non-expedited.
>
> o Batching optimizations allow large number of concurrent requests
> to be handled with fewer grace periods -- and both normal and
> expedited grace periods already do exactly this.
>
> This horse is already out, so trying to shut the gate won't be effective.
Well, I'm not sure there is an easy means of doing machine wide IPIs for
!root out there. This would be a first.
Something along the lines of:
void dummy(void *arg)
{
/* IPIs are assumed to be serializing */
}
void ipi_mm(struct mm_struct *mm)
{
cpumask_var_t cpus;
int cpu;
zalloc_cpumask_var(&cpus, GFP_KERNEL);
for_each_cpu(cpu, mm_cpumask(mm)) {
struct task_struct *p;
/*
* If the current task of @cpu isn't of this @mm, then
* it needs a context switch to become one, which will
* provide the ordering we require.
*/
rcu_read_lock();
p = task_rcu_dereference(&cpu_curr(cpu));
if (p && p->mm == mm)
__cpumask_set_cpu(cpu, cpus);
rcu_read_unlock();
}
on_each_cpu_mask(cpus, dummy, NULL, 1);
}
Would appear to be minimally invasive and only shoot at CPUs we're
currently running our process on, which greatly reduces the impact.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-25 23:20 +0200 |
| Message-ID | <u7fpf-86t-19@gated-at.bofh.it> |
| In reply to | #1696342 |
On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote:
> On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
>
> > There are a lot of variations, to be sure. For whatever it is worth,
> > the original patch that started this uses mprotect():
> >
> > https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0
>
> FWIW that will not work on s390 (and maybe others), they don't in fact
> require IPIs for remote TLB invalidation.
Nor will it for ARM. Nor (I think) for PowerPC. But that is in fact
what people are doing right now in real life. Hence my renewed interest
in sys_membarrier().
> > > Which are those? I thought we significantly reduced those with the nohz
> > > full work. Most IPI uses now first check if a CPU actually needs the IPI
> > > before sending it IIRC.
> >
> > If the task being awakened is higher priority than the task currently
> > running on a given CPU, that CPU still gets an IPI, right? Or am I
> > completely confused?
>
> I was thinking of things like on_each_cpu_cond().
Fair enough.
But it would not be hard for userspace code to force IPIs by repeatedly
awakening higher-priority threads that sleep immediately after being
awakened, right?
> > Unlike userspace preempt disable, in this case we get the abuse anyway
> > via existing mechanisms, as in they are already being abused. If we
> > provide a mechanism for this purpose, we at least have the potential
> > for handling the abuse, for example:
> >
> > o "Defanging" sys_membarrier() on systems that are sensitive to
> > latency. For example, this patch can be defanged by booting
> > with the rcupdate.rcu_normal=1 kernel boot parameter, which
> > causes requests for expedited grace periods to instead use
> > normal grace periods.
> >
> > o Detecting and responding to abuse. For example, perhaps if there
> > are more than (say) 50 expedited sys_membarrier()s within a given
> > jiffy, the excess sys_membarrier()s are non-expedited.
> >
> > o Batching optimizations allow large number of concurrent requests
> > to be handled with fewer grace periods -- and both normal and
> > expedited grace periods already do exactly this.
> >
> > This horse is already out, so trying to shut the gate won't be effective.
>
> Well, I'm not sure there is an easy means of doing machine wide IPIs for
> !root out there. This would be a first.
>
> Something along the lines of:
>
> void dummy(void *arg)
> {
> /* IPIs are assumed to be serializing */
> }
>
> void ipi_mm(struct mm_struct *mm)
> {
> cpumask_var_t cpus;
> int cpu;
>
> zalloc_cpumask_var(&cpus, GFP_KERNEL);
>
> for_each_cpu(cpu, mm_cpumask(mm)) {
> struct task_struct *p;
>
> /*
> * If the current task of @cpu isn't of this @mm, then
> * it needs a context switch to become one, which will
> * provide the ordering we require.
> */
> rcu_read_lock();
> p = task_rcu_dereference(&cpu_curr(cpu));
> if (p && p->mm == mm)
> __cpumask_set_cpu(cpu, cpus);
> rcu_read_unlock();
> }
>
> on_each_cpu_mask(cpus, dummy, NULL, 1);
> }
>
> Would appear to be minimally invasive and only shoot at CPUs we're
> currently running our process on, which greatly reduces the impact.
I am good with this approach as well, and I do very much like that it
avoids IPIing CPUs that aren't running our process (at least in the
common case). But don't we also need added memory ordering? It is
sort of OK to IPI a CPU that just now switched away from our process,
but not so good to miss IPIing a CPU that switched to our process just
a little before sys_membarrier().
I was intending to base this on the last few versions of a 2010 patch,
but maybe things have changed:
https://marc.info/?l=linux-kernel&m=126358017229620&w=2
https://marc.info/?l=linux-kernel&m=126436996014016&w=2
https://marc.info/?l=linux-kernel&m=126601479802978&w=2
https://marc.info/?l=linux-kernel&m=126970692903302&w=2
Discussion here:
https://marc.info/?l=linux-kernel&m=126349766324224&w=2
The discussion led to acquiring the runqueue locks, as there was
otherwise a need to add code to the scheduler fastpaths.
There was a desire to make this work automatically among multiple
processes sharing some memory, but I believe that in this case
the user is going to have to track the multiple processes anyway,
and so can simply do sys_membarrier from each:
https://marc.info/?l=linux-arch&m=126686393820832&w=2
Some architectures are less precise than others in tracking which
CPUs are running a given process due to ASIDs, though this is
thought to be a non-problem:
https://marc.info/?l=linux-arch&m=126716090413065&w=2
https://marc.info/?l=linux-arch&m=126716262815202&w=2
Thoughts?
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-26 00:00 +0200 |
| Message-ID | <u7g1Z-8nQ-41@gated-at.bofh.it> |
| In reply to | #1696502 |
On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
> On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote:
> > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> >
> > > There are a lot of variations, to be sure. For whatever it is worth,
> > > the original patch that started this uses mprotect():
> > >
> > > https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0
> >
> > FWIW that will not work on s390 (and maybe others), they don't in fact
> > require IPIs for remote TLB invalidation.
>
> Nor will it for ARM. Nor (I think) for PowerPC. But that is in fact
> what people are doing right now in real life. Hence my renewed interest
> in sys_membarrier().
People always do crazy stuff, but what surprised me is that such s patch
got merged in urcu even though its known broken for a number of
architectures.
> But it would not be hard for userspace code to force IPIs by repeatedly
> awakening higher-priority threads that sleep immediately after being
> awakened, right?
RT tasks are not readily available to !root, and the user might have
been constrained to a subset of available CPUs.
> > Well, I'm not sure there is an easy means of doing machine wide IPIs for
> > !root out there. This would be a first.
> >
> > Something along the lines of:
> >
> > void dummy(void *arg)
> > {
> > /* IPIs are assumed to be serializing */
> > }
> >
> > void ipi_mm(struct mm_struct *mm)
> > {
> > cpumask_var_t cpus;
> > int cpu;
> >
> > zalloc_cpumask_var(&cpus, GFP_KERNEL);
> >
> > for_each_cpu(cpu, mm_cpumask(mm)) {
> > struct task_struct *p;
> >
> > /*
> > * If the current task of @cpu isn't of this @mm, then
> > * it needs a context switch to become one, which will
> > * provide the ordering we require.
> > */
> > rcu_read_lock();
> > p = task_rcu_dereference(&cpu_curr(cpu));
> > if (p && p->mm == mm)
> > __cpumask_set_cpu(cpu, cpus);
> > rcu_read_unlock();
> > }
> >
> > on_each_cpu_mask(cpus, dummy, NULL, 1);
> > }
> >
> > Would appear to be minimally invasive and only shoot at CPUs we're
> > currently running our process on, which greatly reduces the impact.
>
> I am good with this approach as well, and I do very much like that it
> avoids IPIing CPUs that aren't running our process (at least in the
> common case). But don't we also need added memory ordering? It is
> sort of OK to IPI a CPU that just now switched away from our process,
> but not so good to miss IPIing a CPU that switched to our process just
> a little before sys_membarrier().
My thinking was that if we observe '!= mm' that CPU will have to do a
context switch in order to make it true. That context switch will
provide the ordering we're after so all is well.
Quite possible there's a hole in, but since I'm running on fumes someone
needs to spell it out for me :-)
> I was intending to base this on the last few versions of a 2010 patch,
> but maybe things have changed:
>
> https://marc.info/?l=linux-kernel&m=126358017229620&w=2
> https://marc.info/?l=linux-kernel&m=126436996014016&w=2
> https://marc.info/?l=linux-kernel&m=126601479802978&w=2
> https://marc.info/?l=linux-kernel&m=126970692903302&w=2
>
> Discussion here:
>
> https://marc.info/?l=linux-kernel&m=126349766324224&w=2
>
> The discussion led to acquiring the runqueue locks, as there was
> otherwise a need to add code to the scheduler fastpaths.
TL;DR.. that's far too much to trawl through.
> Some architectures are less precise than others in tracking which
> CPUs are running a given process due to ASIDs, though this is
> thought to be a non-problem:
>
> https://marc.info/?l=linux-arch&m=126716090413065&w=2
> https://marc.info/?l=linux-arch&m=126716262815202&w=2
>
> Thoughts?
Yes, there are architectures that only accumulate bits in mm_cpumask(),
with the additional check to see if the remote task belongs to our MM
this should be a non-issue.
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-26 00:40 +0200 |
| Message-ID | <u7gEG-op-25@gated-at.bofh.it> |
| In reply to | #1696624 |
----- On Jul 25, 2017, at 5:55 PM, Peter Zijlstra peterz@infradead.org wrote: > On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote: >> On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote: >> > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote: >> > >> > > There are a lot of variations, to be sure. For whatever it is worth, >> > > the original patch that started this uses mprotect(): >> > > >> > > https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0 >> > >> > FWIW that will not work on s390 (and maybe others), they don't in fact >> > require IPIs for remote TLB invalidation. >> >> Nor will it for ARM. Nor (I think) for PowerPC. But that is in fact >> what people are doing right now in real life. Hence my renewed interest >> in sys_membarrier(). > > People always do crazy stuff, but what surprised me is that such s patch > got merged in urcu even though its known broken for a number of > architectures. As maintainer of liburcu, I can certainly say that this patch never made it into liburcu master branch (official repo at git://git.liburcu.org/userspace-rcu.git). Paul is referring to a liburcu fork by a github user "msullivan", not the official tree. Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-26 00:50 +0200 |
| Message-ID | <u7gOl-s4-9@gated-at.bofh.it> |
| In reply to | #1696624 |
----- On Jul 25, 2017, at 5:55 PM, Peter Zijlstra peterz@infradead.org wrote:
> On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
>> On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote:
>> > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
[...]
>
>> But it would not be hard for userspace code to force IPIs by repeatedly
>> awakening higher-priority threads that sleep immediately after being
>> awakened, right?
>
> RT tasks are not readily available to !root, and the user might have
> been constrained to a subset of available CPUs.
>
>> > Well, I'm not sure there is an easy means of doing machine wide IPIs for
>> > !root out there. This would be a first.
>> >
>> > Something along the lines of:
>> >
>> > void dummy(void *arg)
>> > {
>> > /* IPIs are assumed to be serializing */
>> > }
>> >
>> > void ipi_mm(struct mm_struct *mm)
>> > {
>> > cpumask_var_t cpus;
>> > int cpu;
>> >
>> > zalloc_cpumask_var(&cpus, GFP_KERNEL);
>> >
>> > for_each_cpu(cpu, mm_cpumask(mm)) {
>> > struct task_struct *p;
>> >
>> > /*
>> > * If the current task of @cpu isn't of this @mm, then
>> > * it needs a context switch to become one, which will
>> > * provide the ordering we require.
>> > */
>> > rcu_read_lock();
>> > p = task_rcu_dereference(&cpu_curr(cpu));
>> > if (p && p->mm == mm)
>> > __cpumask_set_cpu(cpu, cpus);
>> > rcu_read_unlock();
>> > }
>> >
>> > on_each_cpu_mask(cpus, dummy, NULL, 1);
>> > }
>> >
>> > Would appear to be minimally invasive and only shoot at CPUs we're
>> > currently running our process on, which greatly reduces the impact.
>>
>> I am good with this approach as well, and I do very much like that it
>> avoids IPIing CPUs that aren't running our process (at least in the
>> common case). But don't we also need added memory ordering? It is
>> sort of OK to IPI a CPU that just now switched away from our process,
>> but not so good to miss IPIing a CPU that switched to our process just
>> a little before sys_membarrier().
>
> My thinking was that if we observe '!= mm' that CPU will have to do a
> context switch in order to make it true. That context switch will
> provide the ordering we're after so all is well.
>
> Quite possible there's a hole in, but since I'm running on fumes someone
> needs to spell it out for me :-)
>
>> I was intending to base this on the last few versions of a 2010 patch,
>> but maybe things have changed:
>>
>> https://marc.info/?l=linux-kernel&m=126358017229620&w=2
>> https://marc.info/?l=linux-kernel&m=126436996014016&w=2
>> https://marc.info/?l=linux-kernel&m=126601479802978&w=2
>> https://marc.info/?l=linux-kernel&m=126970692903302&w=2
>>
>> Discussion here:
>>
>> https://marc.info/?l=linux-kernel&m=126349766324224&w=2
>>
>> The discussion led to acquiring the runqueue locks, as there was
>> otherwise a need to add code to the scheduler fastpaths.
>
> TL;DR.. that's far too much to trawl through.
>
>> Some architectures are less precise than others in tracking which
>> CPUs are running a given process due to ASIDs, though this is
>> thought to be a non-problem:
>>
>> https://marc.info/?l=linux-arch&m=126716090413065&w=2
>> https://marc.info/?l=linux-arch&m=126716262815202&w=2
>>
>> Thoughts?
>
> Yes, there are architectures that only accumulate bits in mm_cpumask(),
> with the additional check to see if the remote task belongs to our MM
> this should be a non-issue.
This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag
for expedited process-local effect. This differs from the "SHARED" flag,
since the SHARED flag affects threads accessing memory mappings shared
across processes as well.
I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior
by iterating on all memory mappings mapped into the current process,
and build a cpumask based on the union of all mm masks encountered ?
Then we could send the IPI to all cpus belonging to that cpumask. Or
am I missing something obvious ?
Thanks,
Mathieu
--
Mathieu Desnoyers
EfficiOS Inc.
http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 02:10 +0200 |
| Message-ID | <u7i3L-1pC-1@gated-at.bofh.it> |
| In reply to | #1696647 |
On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote:
> ----- On Jul 25, 2017, at 5:55 PM, Peter Zijlstra peterz@infradead.org wrote:
>
> > On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
> >> On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote:
> >> > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> [...]
> >
> >> But it would not be hard for userspace code to force IPIs by repeatedly
> >> awakening higher-priority threads that sleep immediately after being
> >> awakened, right?
> >
> > RT tasks are not readily available to !root, and the user might have
> > been constrained to a subset of available CPUs.
> >
> >> > Well, I'm not sure there is an easy means of doing machine wide IPIs for
> >> > !root out there. This would be a first.
> >> >
> >> > Something along the lines of:
> >> >
> >> > void dummy(void *arg)
> >> > {
> >> > /* IPIs are assumed to be serializing */
> >> > }
> >> >
> >> > void ipi_mm(struct mm_struct *mm)
> >> > {
> >> > cpumask_var_t cpus;
> >> > int cpu;
> >> >
> >> > zalloc_cpumask_var(&cpus, GFP_KERNEL);
> >> >
> >> > for_each_cpu(cpu, mm_cpumask(mm)) {
> >> > struct task_struct *p;
> >> >
> >> > /*
> >> > * If the current task of @cpu isn't of this @mm, then
> >> > * it needs a context switch to become one, which will
> >> > * provide the ordering we require.
> >> > */
> >> > rcu_read_lock();
> >> > p = task_rcu_dereference(&cpu_curr(cpu));
> >> > if (p && p->mm == mm)
> >> > __cpumask_set_cpu(cpu, cpus);
> >> > rcu_read_unlock();
> >> > }
> >> >
> >> > on_each_cpu_mask(cpus, dummy, NULL, 1);
> >> > }
> >> >
> >> > Would appear to be minimally invasive and only shoot at CPUs we're
> >> > currently running our process on, which greatly reduces the impact.
> >>
> >> I am good with this approach as well, and I do very much like that it
> >> avoids IPIing CPUs that aren't running our process (at least in the
> >> common case). But don't we also need added memory ordering? It is
> >> sort of OK to IPI a CPU that just now switched away from our process,
> >> but not so good to miss IPIing a CPU that switched to our process just
> >> a little before sys_membarrier().
> >
> > My thinking was that if we observe '!= mm' that CPU will have to do a
> > context switch in order to make it true. That context switch will
> > provide the ordering we're after so all is well.
> >
> > Quite possible there's a hole in, but since I'm running on fumes someone
> > needs to spell it out for me :-)
> >
> >> I was intending to base this on the last few versions of a 2010 patch,
> >> but maybe things have changed:
> >>
> >> https://marc.info/?l=linux-kernel&m=126358017229620&w=2
> >> https://marc.info/?l=linux-kernel&m=126436996014016&w=2
> >> https://marc.info/?l=linux-kernel&m=126601479802978&w=2
> >> https://marc.info/?l=linux-kernel&m=126970692903302&w=2
> >>
> >> Discussion here:
> >>
> >> https://marc.info/?l=linux-kernel&m=126349766324224&w=2
> >>
> >> The discussion led to acquiring the runqueue locks, as there was
> >> otherwise a need to add code to the scheduler fastpaths.
> >
> > TL;DR.. that's far too much to trawl through.
> >
> >> Some architectures are less precise than others in tracking which
> >> CPUs are running a given process due to ASIDs, though this is
> >> thought to be a non-problem:
> >>
> >> https://marc.info/?l=linux-arch&m=126716090413065&w=2
> >> https://marc.info/?l=linux-arch&m=126716262815202&w=2
> >>
> >> Thoughts?
> >
> > Yes, there are architectures that only accumulate bits in mm_cpumask(),
> > with the additional check to see if the remote task belongs to our MM
> > this should be a non-issue.
>
> This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag
> for expedited process-local effect. This differs from the "SHARED" flag,
> since the SHARED flag affects threads accessing memory mappings shared
> across processes as well.
>
> I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior
> by iterating on all memory mappings mapped into the current process,
> and build a cpumask based on the union of all mm masks encountered ?
> Then we could send the IPI to all cpus belonging to that cpumask. Or
> am I missing something obvious ?
I suspect that something like this would work, but I agree with your 2010
self, who argued that this should be follow-on functionality. After all,
the user probably needs to be aware of who is sharing for other reasons,
and can then make each process do sys_membarrier().
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-26 09:50 +0200 |
| Message-ID | <u7peV-5RS-7@gated-at.bofh.it> |
| In reply to | #1696647 |
On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > for expedited process-local effect. This differs from the "SHARED" flag, > since the SHARED flag affects threads accessing memory mappings shared > across processes as well. > > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > by iterating on all memory mappings mapped into the current process, > and build a cpumask based on the union of all mm masks encountered ? > Then we could send the IPI to all cpus belonging to that cpumask. Or > am I missing something obvious ? I would readily object to such a beast. You far too quickly end up having to IPI everybody because of some stupid shared map or something (yes I know, normal DSOs are mapped private). The whole private_expedited is only palatable because we can only hinder our own threads (much).
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 17:50 +0200 |
| Message-ID | <u7wJr-2ak-1@gated-at.bofh.it> |
| In reply to | #1696850 |
On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: > On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > > for expedited process-local effect. This differs from the "SHARED" flag, > > since the SHARED flag affects threads accessing memory mappings shared > > across processes as well. > > > > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > > by iterating on all memory mappings mapped into the current process, > > and build a cpumask based on the union of all mm masks encountered ? > > Then we could send the IPI to all cpus belonging to that cpumask. Or > > am I missing something obvious ? > > I would readily object to such a beast. You far too quickly end up > having to IPI everybody because of some stupid shared map or something > (yes I know, normal DSOs are mapped private). Agreed, we should keep things simple to start with. The user can always invoke sys_membarrier() from each process. Thanx, Paul > The whole private_expedited is only palatable because we can only hinder > our own threads (much). >
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-26 20:00 +0200 |
| Message-ID | <u7yLh-3p3-25@gated-at.bofh.it> |
| In reply to | #1697315 |
----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag >> > for expedited process-local effect. This differs from the "SHARED" flag, >> > since the SHARED flag affects threads accessing memory mappings shared >> > across processes as well. >> > >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior >> > by iterating on all memory mappings mapped into the current process, >> > and build a cpumask based on the union of all mm masks encountered ? >> > Then we could send the IPI to all cpus belonging to that cpumask. Or >> > am I missing something obvious ? >> >> I would readily object to such a beast. You far too quickly end up >> having to IPI everybody because of some stupid shared map or something >> (yes I know, normal DSOs are mapped private). > > Agreed, we should keep things simple to start with. The user can always > invoke sys_membarrier() from each process. Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting per thread. For instance, we could add a new "ulimit" that would bound the number of expedited membarrier per thread that can be done per millisecond, and switch to synchronize_sched() whenever a thread goes beyond that limit for the rest of the time-slot. A RT system that really cares about not having userspace sending IPIs to all cpus could set the ulimit value to 0, which would always use synchronize_sched(). Thoughts ? Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web