Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1695159 > unrolled thread
| Started by | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| First post | 2017-07-25 00:00 +0200 |
| Last post | 2017-07-27 16:40 +0200 |
| Articles | 20 on this page of 56 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 00:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-25 06:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-25 15:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 18:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 19:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 19:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 21:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 21:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 22:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 23:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 00:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 20:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 20:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 22:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 23:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 03:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-27 14:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 11:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 12:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 10:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:10 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 17:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:30 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-26 11:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:00 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:40 +0200
Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 20:40 +0200 |
| Message-ID | <u7znY-3Tc-15@gated-at.bofh.it> |
| In reply to | #1697455 |
On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > > > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: > >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > >> > for expedited process-local effect. This differs from the "SHARED" flag, > >> > since the SHARED flag affects threads accessing memory mappings shared > >> > across processes as well. > >> > > >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > >> > by iterating on all memory mappings mapped into the current process, > >> > and build a cpumask based on the union of all mm masks encountered ? > >> > Then we could send the IPI to all cpus belonging to that cpumask. Or > >> > am I missing something obvious ? > >> > >> I would readily object to such a beast. You far too quickly end up > >> having to IPI everybody because of some stupid shared map or something > >> (yes I know, normal DSOs are mapped private). > > > > Agreed, we should keep things simple to start with. The user can always > > invoke sys_membarrier() from each process. > > Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > per thread. For instance, we could add a new "ulimit" that would bound the > number of expedited membarrier per thread that can be done per millisecond, > and switch to synchronize_sched() whenever a thread goes beyond that limit > for the rest of the time-slot. > > A RT system that really cares about not having userspace sending IPIs > to all cpus could set the ulimit value to 0, which would always use > synchronize_sched(). > > Thoughts ? The patch I posted reverts to synchronize_sched() in kernels booted with rcupdate.rcu_normal=1. ;-) But who is pushing for multiple-process sys_membarrier()? Everyone I have talked to is OK with it being local to the current process. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-26 22:40 +0200 |
| Message-ID | <u7Bg5-54j-3@gated-at.bofh.it> |
| In reply to | #1697475 |
----- On Jul 26, 2017, at 2:30 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: >> ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com >> wrote: >> >> > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: >> >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: >> >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag >> >> > for expedited process-local effect. This differs from the "SHARED" flag, >> >> > since the SHARED flag affects threads accessing memory mappings shared >> >> > across processes as well. >> >> > >> >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior >> >> > by iterating on all memory mappings mapped into the current process, >> >> > and build a cpumask based on the union of all mm masks encountered ? >> >> > Then we could send the IPI to all cpus belonging to that cpumask. Or >> >> > am I missing something obvious ? >> >> >> >> I would readily object to such a beast. You far too quickly end up >> >> having to IPI everybody because of some stupid shared map or something >> >> (yes I know, normal DSOs are mapped private). >> > >> > Agreed, we should keep things simple to start with. The user can always >> > invoke sys_membarrier() from each process. >> >> Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting >> per thread. For instance, we could add a new "ulimit" that would bound the >> number of expedited membarrier per thread that can be done per millisecond, >> and switch to synchronize_sched() whenever a thread goes beyond that limit >> for the rest of the time-slot. >> >> A RT system that really cares about not having userspace sending IPIs >> to all cpus could set the ulimit value to 0, which would always use >> synchronize_sched(). >> >> Thoughts ? > > The patch I posted reverts to synchronize_sched() in kernels booted with > rcupdate.rcu_normal=1. ;-) > > But who is pushing for multiple-process sys_membarrier()? Everyone I > have talked to is OK with it being local to the current process. I guess I'm probably the guilty one intending to do weird stuff in userspace ;) Here are my two use-cases: * a new multi-process liburcu flavor, useful if e.g. a set of processes are responsible for updating a shared memory data structure, and a separate set of processes read that data structure. The readers can be killed without ill effect on the other processes. The synchronization could be done by one multi-process liburcu flavor per reader process "group". * lttng-ust user-space ring buffers (shared across processes). Both rely on a shared memory mapping for communication between processes, and I would like to be able to issue a sys_membarrier targeting all CPUs that may currently touch the shared memory mapping. I don't really need a system-wide effect, but I would like to be able to target a shared memory mapping and efficiently do an expedited sys_membarrier on all cpus involved. With lttng-ust, the shared buffers can spawn across 1000+ processes, so asking each process to issue sys_membarrier would add lots of unneeded overhead, because this would issue lots of needless memory barriers. Thoughts ? Thanks, Mathieu -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 23:20 +0200 |
| Message-ID | <u7BSO-5zh-15@gated-at.bofh.it> |
| In reply to | #1697549 |
On Wed, Jul 26, 2017 at 08:37:23PM +0000, Mathieu Desnoyers wrote: > ----- On Jul 26, 2017, at 2:30 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > > > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > >> ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com > >> wrote: > >> > >> > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: > >> >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > >> >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > >> >> > for expedited process-local effect. This differs from the "SHARED" flag, > >> >> > since the SHARED flag affects threads accessing memory mappings shared > >> >> > across processes as well. > >> >> > > >> >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > >> >> > by iterating on all memory mappings mapped into the current process, > >> >> > and build a cpumask based on the union of all mm masks encountered ? > >> >> > Then we could send the IPI to all cpus belonging to that cpumask. Or > >> >> > am I missing something obvious ? > >> >> > >> >> I would readily object to such a beast. You far too quickly end up > >> >> having to IPI everybody because of some stupid shared map or something > >> >> (yes I know, normal DSOs are mapped private). > >> > > >> > Agreed, we should keep things simple to start with. The user can always > >> > invoke sys_membarrier() from each process. > >> > >> Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > >> per thread. For instance, we could add a new "ulimit" that would bound the > >> number of expedited membarrier per thread that can be done per millisecond, > >> and switch to synchronize_sched() whenever a thread goes beyond that limit > >> for the rest of the time-slot. > >> > >> A RT system that really cares about not having userspace sending IPIs > >> to all cpus could set the ulimit value to 0, which would always use > >> synchronize_sched(). > >> > >> Thoughts ? > > > > The patch I posted reverts to synchronize_sched() in kernels booted with > > rcupdate.rcu_normal=1. ;-) > > > > But who is pushing for multiple-process sys_membarrier()? Everyone I > > have talked to is OK with it being local to the current process. > > I guess I'm probably the guilty one intending to do weird stuff in userspace ;) > > Here are my two use-cases: > > * a new multi-process liburcu flavor, useful if e.g. a set of processes are > responsible for updating a shared memory data structure, and a separate set > of processes read that data structure. The readers can be killed without ill > effect on the other processes. The synchronization could be done by one > multi-process liburcu flavor per reader process "group". > > * lttng-ust user-space ring buffers (shared across processes). > > Both rely on a shared memory mapping for communication between processes, and > I would like to be able to issue a sys_membarrier targeting all CPUs that may > currently touch the shared memory mapping. > > I don't really need a system-wide effect, but I would like to be able to target > a shared memory mapping and efficiently do an expedited sys_membarrier on all > cpus involved. > > With lttng-ust, the shared buffers can spawn across 1000+ processes, so > asking each process to issue sys_membarrier would add lots of unneeded overhead, > because this would issue lots of needless memory barriers. > > Thoughts ? Dealing explicitly with 1000+ processes sounds like no picnic. It instead sounds like a job for synchronize_sched_expedited(). ;-) Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 03:50 +0200 |
| Message-ID | <u7G65-85D-1@gated-at.bofh.it> |
| In reply to | #1697566 |
On Wed, Jul 26, 2017 at 02:11:46PM -0700, Paul E. McKenney wrote: > On Wed, Jul 26, 2017 at 08:37:23PM +0000, Mathieu Desnoyers wrote: > > ----- On Jul 26, 2017, at 2:30 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > > > > > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > > >> ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com > > >> wrote: > > >> > > >> > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: > > >> >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > > >> >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > > >> >> > for expedited process-local effect. This differs from the "SHARED" flag, > > >> >> > since the SHARED flag affects threads accessing memory mappings shared > > >> >> > across processes as well. > > >> >> > > > >> >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > > >> >> > by iterating on all memory mappings mapped into the current process, > > >> >> > and build a cpumask based on the union of all mm masks encountered ? > > >> >> > Then we could send the IPI to all cpus belonging to that cpumask. Or > > >> >> > am I missing something obvious ? > > >> >> > > >> >> I would readily object to such a beast. You far too quickly end up > > >> >> having to IPI everybody because of some stupid shared map or something > > >> >> (yes I know, normal DSOs are mapped private). > > >> > > > >> > Agreed, we should keep things simple to start with. The user can always > > >> > invoke sys_membarrier() from each process. > > >> > > >> Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > > >> per thread. For instance, we could add a new "ulimit" that would bound the > > >> number of expedited membarrier per thread that can be done per millisecond, > > >> and switch to synchronize_sched() whenever a thread goes beyond that limit > > >> for the rest of the time-slot. > > >> > > >> A RT system that really cares about not having userspace sending IPIs > > >> to all cpus could set the ulimit value to 0, which would always use > > >> synchronize_sched(). > > >> > > >> Thoughts ? > > > > > > The patch I posted reverts to synchronize_sched() in kernels booted with > > > rcupdate.rcu_normal=1. ;-) > > > > > > But who is pushing for multiple-process sys_membarrier()? Everyone I > > > have talked to is OK with it being local to the current process. > > > > I guess I'm probably the guilty one intending to do weird stuff in userspace ;) > > > > Here are my two use-cases: > > > > * a new multi-process liburcu flavor, useful if e.g. a set of processes are > > responsible for updating a shared memory data structure, and a separate set > > of processes read that data structure. The readers can be killed without ill > > effect on the other processes. The synchronization could be done by one > > multi-process liburcu flavor per reader process "group". > > > > * lttng-ust user-space ring buffers (shared across processes). > > > > Both rely on a shared memory mapping for communication between processes, and > > I would like to be able to issue a sys_membarrier targeting all CPUs that may > > currently touch the shared memory mapping. > > > > I don't really need a system-wide effect, but I would like to be able to target > > a shared memory mapping and efficiently do an expedited sys_membarrier on all > > cpus involved. > > > > With lttng-ust, the shared buffers can spawn across 1000+ processes, so > > asking each process to issue sys_membarrier would add lots of unneeded overhead, > > because this would issue lots of needless memory barriers. > > > > Thoughts ? > > Dealing explicitly with 1000+ processes sounds like no picnic. It instead > sounds like a job for synchronize_sched_expedited(). ;-) Actually... Mathieu, does your use case require unprivileged access to sys_membarrier()? Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Mathieu Desnoyers <mathieu.desnoyers@efficios.com> |
|---|---|
| Date | 2017-07-27 14:40 +0200 |
| Message-ID | <u7Qf8-65e-29@gated-at.bofh.it> |
| In reply to | #1697653 |
----- On Jul 26, 2017, at 9:45 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > On Wed, Jul 26, 2017 at 02:11:46PM -0700, Paul E. McKenney wrote: >> On Wed, Jul 26, 2017 at 08:37:23PM +0000, Mathieu Desnoyers wrote: >> > ----- On Jul 26, 2017, at 2:30 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com >> > wrote: >> > >> > > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: >> > >> ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com >> > >> wrote: >> > >> >> > >> > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: >> > >> >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: >> > >> >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag >> > >> >> > for expedited process-local effect. This differs from the "SHARED" flag, >> > >> >> > since the SHARED flag affects threads accessing memory mappings shared >> > >> >> > across processes as well. >> > >> >> > >> > >> >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior >> > >> >> > by iterating on all memory mappings mapped into the current process, >> > >> >> > and build a cpumask based on the union of all mm masks encountered ? >> > >> >> > Then we could send the IPI to all cpus belonging to that cpumask. Or >> > >> >> > am I missing something obvious ? >> > >> >> >> > >> >> I would readily object to such a beast. You far too quickly end up >> > >> >> having to IPI everybody because of some stupid shared map or something >> > >> >> (yes I know, normal DSOs are mapped private). >> > >> > >> > >> > Agreed, we should keep things simple to start with. The user can always >> > >> > invoke sys_membarrier() from each process. >> > >> >> > >> Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting >> > >> per thread. For instance, we could add a new "ulimit" that would bound the >> > >> number of expedited membarrier per thread that can be done per millisecond, >> > >> and switch to synchronize_sched() whenever a thread goes beyond that limit >> > >> for the rest of the time-slot. >> > >> >> > >> A RT system that really cares about not having userspace sending IPIs >> > >> to all cpus could set the ulimit value to 0, which would always use >> > >> synchronize_sched(). >> > >> >> > >> Thoughts ? >> > > >> > > The patch I posted reverts to synchronize_sched() in kernels booted with >> > > rcupdate.rcu_normal=1. ;-) >> > > >> > > But who is pushing for multiple-process sys_membarrier()? Everyone I >> > > have talked to is OK with it being local to the current process. >> > >> > I guess I'm probably the guilty one intending to do weird stuff in userspace ;) >> > >> > Here are my two use-cases: >> > >> > * a new multi-process liburcu flavor, useful if e.g. a set of processes are >> > responsible for updating a shared memory data structure, and a separate set >> > of processes read that data structure. The readers can be killed without ill >> > effect on the other processes. The synchronization could be done by one >> > multi-process liburcu flavor per reader process "group". >> > >> > * lttng-ust user-space ring buffers (shared across processes). >> > >> > Both rely on a shared memory mapping for communication between processes, and >> > I would like to be able to issue a sys_membarrier targeting all CPUs that may >> > currently touch the shared memory mapping. >> > >> > I don't really need a system-wide effect, but I would like to be able to target >> > a shared memory mapping and efficiently do an expedited sys_membarrier on all >> > cpus involved. >> > >> > With lttng-ust, the shared buffers can spawn across 1000+ processes, so >> > asking each process to issue sys_membarrier would add lots of unneeded overhead, >> > because this would issue lots of needless memory barriers. >> > >> > Thoughts ? >> >> Dealing explicitly with 1000+ processes sounds like no picnic. It instead >> sounds like a job for synchronize_sched_expedited(). ;-) > > Actually... > > Mathieu, does your use case require unprivileged access to sys_membarrier()? Unfortunately, yes, it does require sys_membarrier to be used from non-root both for lttng-ust and liburcu multi-process. And as Peter pointed out, stuff like containers complicates things even for the root case. Thanks, Mathieu > > Thanx, Paul -- Mathieu Desnoyers EfficiOS Inc. http://www.efficios.com
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 16:50 +0200 |
| Message-ID | <u7SgW-7jy-19@gated-at.bofh.it> |
| In reply to | #1697947 |
On Thu, Jul 27, 2017 at 12:39:36PM +0000, Mathieu Desnoyers wrote: > ----- On Jul 26, 2017, at 9:45 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com wrote: > > > On Wed, Jul 26, 2017 at 02:11:46PM -0700, Paul E. McKenney wrote: > >> On Wed, Jul 26, 2017 at 08:37:23PM +0000, Mathieu Desnoyers wrote: > >> > ----- On Jul 26, 2017, at 2:30 PM, Paul E. McKenney paulmck@linux.vnet.ibm.com > >> > wrote: > >> > > >> > > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > >> > >> ----- On Jul 26, 2017, at 11:42 AM, Paul E. McKenney paulmck@linux.vnet.ibm.com > >> > >> wrote: > >> > >> > >> > >> > On Wed, Jul 26, 2017 at 09:46:56AM +0200, Peter Zijlstra wrote: > >> > >> >> On Tue, Jul 25, 2017 at 10:50:13PM +0000, Mathieu Desnoyers wrote: > >> > >> >> > This would implement a MEMBARRIER_CMD_PRIVATE_EXPEDITED (or such) flag > >> > >> >> > for expedited process-local effect. This differs from the "SHARED" flag, > >> > >> >> > since the SHARED flag affects threads accessing memory mappings shared > >> > >> >> > across processes as well. > >> > >> >> > > >> > >> >> > I wonder if we could create a MEMBARRIER_CMD_SHARED_EXPEDITED behavior > >> > >> >> > by iterating on all memory mappings mapped into the current process, > >> > >> >> > and build a cpumask based on the union of all mm masks encountered ? > >> > >> >> > Then we could send the IPI to all cpus belonging to that cpumask. Or > >> > >> >> > am I missing something obvious ? > >> > >> >> > >> > >> >> I would readily object to such a beast. You far too quickly end up > >> > >> >> having to IPI everybody because of some stupid shared map or something > >> > >> >> (yes I know, normal DSOs are mapped private). > >> > >> > > >> > >> > Agreed, we should keep things simple to start with. The user can always > >> > >> > invoke sys_membarrier() from each process. > >> > >> > >> > >> Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > >> > >> per thread. For instance, we could add a new "ulimit" that would bound the > >> > >> number of expedited membarrier per thread that can be done per millisecond, > >> > >> and switch to synchronize_sched() whenever a thread goes beyond that limit > >> > >> for the rest of the time-slot. > >> > >> > >> > >> A RT system that really cares about not having userspace sending IPIs > >> > >> to all cpus could set the ulimit value to 0, which would always use > >> > >> synchronize_sched(). > >> > >> > >> > >> Thoughts ? > >> > > > >> > > The patch I posted reverts to synchronize_sched() in kernels booted with > >> > > rcupdate.rcu_normal=1. ;-) > >> > > > >> > > But who is pushing for multiple-process sys_membarrier()? Everyone I > >> > > have talked to is OK with it being local to the current process. > >> > > >> > I guess I'm probably the guilty one intending to do weird stuff in userspace ;) > >> > > >> > Here are my two use-cases: > >> > > >> > * a new multi-process liburcu flavor, useful if e.g. a set of processes are > >> > responsible for updating a shared memory data structure, and a separate set > >> > of processes read that data structure. The readers can be killed without ill > >> > effect on the other processes. The synchronization could be done by one > >> > multi-process liburcu flavor per reader process "group". > >> > > >> > * lttng-ust user-space ring buffers (shared across processes). > >> > > >> > Both rely on a shared memory mapping for communication between processes, and > >> > I would like to be able to issue a sys_membarrier targeting all CPUs that may > >> > currently touch the shared memory mapping. > >> > > >> > I don't really need a system-wide effect, but I would like to be able to target > >> > a shared memory mapping and efficiently do an expedited sys_membarrier on all > >> > cpus involved. > >> > > >> > With lttng-ust, the shared buffers can spawn across 1000+ processes, so > >> > asking each process to issue sys_membarrier would add lots of unneeded overhead, > >> > because this would issue lots of needless memory barriers. > >> > > >> > Thoughts ? > >> > >> Dealing explicitly with 1000+ processes sounds like no picnic. It instead > >> sounds like a job for synchronize_sched_expedited(). ;-) > > > > Actually... > > > > Mathieu, does your use case require unprivileged access to sys_membarrier()? > > Unfortunately, yes, it does require sys_membarrier to be used from non-root > both for lttng-ust and liburcu multi-process. And as Peter pointed out, stuff > like containers complicates things even for the root case. Hey, I was hoping! ;-) Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 12:30 +0200 |
| Message-ID | <u7Odj-4Tl-5@gated-at.bofh.it> |
| In reply to | #1697475 |
On Wed, Jul 26, 2017 at 11:30:32AM -0700, Paul E. McKenney wrote: > The patch I posted reverts to synchronize_sched() in kernels booted with > rcupdate.rcu_normal=1. ;-) So boot parameters are no solution and are only slightly better than compile time switches. What if you have a machine that runs workloads that want both options? partitioning and containers are somewhat popular these days and system wide tunables don't work for them.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 17:00 +0200 |
| Message-ID | <u7SqB-7mN-11@gated-at.bofh.it> |
| In reply to | #1697873 |
On Thu, Jul 27, 2017 at 12:24:22PM +0200, Peter Zijlstra wrote: > On Wed, Jul 26, 2017 at 11:30:32AM -0700, Paul E. McKenney wrote: > > The patch I posted reverts to synchronize_sched() in kernels booted with > > rcupdate.rcu_normal=1. ;-) > > So boot parameters are no solution and are only slightly better than > compile time switches. > > What if you have a machine that runs workloads that want both options? > partitioning and containers are somewhat popular these days and system > wide tunables don't work for them. Agreed, these are disadvantages. And again, I will nevertheless be carrying some variant of this patch until something better is in place, where that something has been shown to satisfy the needs of the people requesting this feature. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 11:00 +0200 |
| Message-ID | <u7MOd-3UB-1@gated-at.bofh.it> |
| In reply to | #1697455 |
On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > per thread. For instance, we could add a new "ulimit" that would bound the > number of expedited membarrier per thread that can be done per millisecond, > and switch to synchronize_sched() whenever a thread goes beyond that limit > for the rest of the time-slot. You forgot to ask yourself how you could abuse this.. just spawn more threads. Per-thread limits are nearly useless, because spawning new threads is cheap. > A RT system that really cares about not having userspace sending IPIs > to all cpus could set the ulimit value to 0, which would always use > synchronize_sched(). > > Thoughts ? So I really don't like SHARED_EXPEDITED, and your use-cases (from later emails) makes me think sys_membarrier() should have a pointer argument to identify the shared mapping. But even then, iterating the rmap for something that has 1000+ maps isn't going to be nice or fast, even in kernel space. Another crazy idea is using madvise() for this. The new MADV_MEMBAR could revoke PROT_WRITE and PROT_READ for all extant PTEs. Then the tasks attempting access will fault and the fault handler can figure out if it still needs to issue a MB or not before reinstating the PTE. That is fully contained to the tasks actually having that map, and doesn't perturb anybody else.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 12:20 +0200 |
| Message-ID | <u7O3E-4Qj-19@gated-at.bofh.it> |
| In reply to | #1697814 |
On Thu, Jul 27, 2017 at 10:53:12AM +0200, Peter Zijlstra wrote: > Another crazy idea is using madvise() for this. The new MADV_MEMBAR > could revoke PROT_WRITE and PROT_READ for all extant PTEs. Then the > tasks attempting access will fault and the fault handler can figure out > if it still needs to issue a MB or not before reinstating the PTE. Slight oversight is that page-tables are per process and we need a MB per thread, so the fault handler can simply issue a private sys_membarrier() to spread the joy around the local process. > That is fully contained to the tasks actually having that map, and > doesn't perturb anybody else.
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2017-07-27 12:30 +0200 |
| Message-ID | <u7Odj-4Tl-11@gated-at.bofh.it> |
| In reply to | #1697814 |
On Thu, Jul 27, 2017 at 10:53:12AM +0200, Peter Zijlstra wrote: > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > > > Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > > per thread. For instance, we could add a new "ulimit" that would bound the > > number of expedited membarrier per thread that can be done per millisecond, > > and switch to synchronize_sched() whenever a thread goes beyond that limit > > for the rest of the time-slot. > > You forgot to ask yourself how you could abuse this.. just spawn more > threads. > > Per-thread limits are nearly useless, because spawning new threads is > cheap. > > > A RT system that really cares about not having userspace sending IPIs > > to all cpus could set the ulimit value to 0, which would always use > > synchronize_sched(). > > > > Thoughts ? > > So I really don't like SHARED_EXPEDITED, and your use-cases (from later > emails) makes me think sys_membarrier() should have a pointer argument > to identify the shared mapping. > > But even then, iterating the rmap for something that has 1000+ maps > isn't going to be nice or fast, even in kernel space. > > Another crazy idea is using madvise() for this. The new MADV_MEMBAR > could revoke PROT_WRITE and PROT_READ for all extant PTEs. Then the > tasks attempting access will fault and the fault handler can figure out > if it still needs to issue a MB or not before reinstating the PTE. If you did that, wouldn't you need to leave the faulting entries intact until all CPUs with the mm scheduled have either faulted or context-switched? We don't have per-cpu permission bits in the page tables, unfortunately. Will
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 15:20 +0200 |
| Message-ID | <u7QRQ-6zj-3@gated-at.bofh.it> |
| In reply to | #1697814 |
On Thu, Jul 27, 2017 at 10:53:12AM +0200, Peter Zijlstra wrote: > On Wed, Jul 26, 2017 at 06:01:15PM +0000, Mathieu Desnoyers wrote: > > > Another alternative for a MEMBARRIER_CMD_SHARED_EXPEDITED would be rate-limiting > > per thread. For instance, we could add a new "ulimit" that would bound the > > number of expedited membarrier per thread that can be done per millisecond, > > and switch to synchronize_sched() whenever a thread goes beyond that limit > > for the rest of the time-slot. > > You forgot to ask yourself how you could abuse this.. just spawn more > threads. > > Per-thread limits are nearly useless, because spawning new threads is > cheap. Agreed -- any per-thread limit has to be some portion/fraction of a global limit. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 02:00 +0200 |
| Message-ID | <u7hU5-15V-3@gated-at.bofh.it> |
| In reply to | #1696624 |
On Tue, Jul 25, 2017 at 11:55:10PM +0200, Peter Zijlstra wrote:
> On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
> > On Tue, Jul 25, 2017 at 10:24:51PM +0200, Peter Zijlstra wrote:
> > > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> > >
> > > > There are a lot of variations, to be sure. For whatever it is worth,
> > > > the original patch that started this uses mprotect():
> > > >
> > > > https://github.com/msullivan/userspace-rcu/commit/04656b468d418efbc5d934ab07954eb8395a7ab0
> > >
> > > FWIW that will not work on s390 (and maybe others), they don't in fact
> > > require IPIs for remote TLB invalidation.
> >
> > Nor will it for ARM. Nor (I think) for PowerPC. But that is in fact
> > what people are doing right now in real life. Hence my renewed interest
> > in sys_membarrier().
>
> People always do crazy stuff, but what surprised me is that such s patch
> got merged in urcu even though its known broken for a number of
> architectures.
It did not get merged into urcu. It is instead used directly by a
number of people for a number of concurrent algorithms.
> > But it would not be hard for userspace code to force IPIs by repeatedly
> > awakening higher-priority threads that sleep immediately after being
> > awakened, right?
>
> RT tasks are not readily available to !root, and the user might have
> been constrained to a subset of available CPUs.
So non-idle non-nohz CPUs never get IPIed for wakeups of SCHED_OTHER
threads?
> > > Well, I'm not sure there is an easy means of doing machine wide IPIs for
> > > !root out there. This would be a first.
> > >
> > > Something along the lines of:
> > >
> > > void dummy(void *arg)
> > > {
> > > /* IPIs are assumed to be serializing */
> > > }
> > >
> > > void ipi_mm(struct mm_struct *mm)
> > > {
> > > cpumask_var_t cpus;
> > > int cpu;
> > >
> > > zalloc_cpumask_var(&cpus, GFP_KERNEL);
> > >
> > > for_each_cpu(cpu, mm_cpumask(mm)) {
> > > struct task_struct *p;
> > >
> > > /*
> > > * If the current task of @cpu isn't of this @mm, then
> > > * it needs a context switch to become one, which will
> > > * provide the ordering we require.
> > > */
> > > rcu_read_lock();
> > > p = task_rcu_dereference(&cpu_curr(cpu));
> > > if (p && p->mm == mm)
> > > __cpumask_set_cpu(cpu, cpus);
> > > rcu_read_unlock();
> > > }
> > >
> > > on_each_cpu_mask(cpus, dummy, NULL, 1);
> > > }
> > >
> > > Would appear to be minimally invasive and only shoot at CPUs we're
> > > currently running our process on, which greatly reduces the impact.
> >
> > I am good with this approach as well, and I do very much like that it
> > avoids IPIing CPUs that aren't running our process (at least in the
> > common case). But don't we also need added memory ordering? It is
> > sort of OK to IPI a CPU that just now switched away from our process,
> > but not so good to miss IPIing a CPU that switched to our process just
> > a little before sys_membarrier().
>
> My thinking was that if we observe '!= mm' that CPU will have to do a
> context switch in order to make it true. That context switch will
> provide the ordering we're after so all is well.
>
> Quite possible there's a hole in, but since I'm running on fumes someone
> needs to spell it out for me :-)
This would be the https://marc.info/?l=linux-kernel&m=126349766324224&w=2
URL below.
Which might or might not still be applicable.
> > I was intending to base this on the last few versions of a 2010 patch,
> > but maybe things have changed:
> >
> > https://marc.info/?l=linux-kernel&m=126358017229620&w=2
> > https://marc.info/?l=linux-kernel&m=126436996014016&w=2
> > https://marc.info/?l=linux-kernel&m=126601479802978&w=2
> > https://marc.info/?l=linux-kernel&m=126970692903302&w=2
> >
> > Discussion here:
> >
> > https://marc.info/?l=linux-kernel&m=126349766324224&w=2
> >
> > The discussion led to acquiring the runqueue locks, as there was
> > otherwise a need to add code to the scheduler fastpaths.
>
> TL;DR.. that's far too much to trawl through.
So we re-derive it from first principles instead? ;-)
> > Some architectures are less precise than others in tracking which
> > CPUs are running a given process due to ASIDs, though this is
> > thought to be a non-problem:
> >
> > https://marc.info/?l=linux-arch&m=126716090413065&w=2
> > https://marc.info/?l=linux-arch&m=126716262815202&w=2
> >
> > Thoughts?
>
> Yes, there are architectures that only accumulate bits in mm_cpumask(),
> with the additional check to see if the remote task belongs to our MM
> this should be a non-issue.
Makes sense.
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-26 09:50 +0200 |
| Message-ID | <u7peW-5RS-17@gated-at.bofh.it> |
| In reply to | #1696669 |
On Tue, Jul 25, 2017 at 04:59:36PM -0700, Paul E. McKenney wrote: > On Tue, Jul 25, 2017 at 11:55:10PM +0200, Peter Zijlstra wrote: > > People always do crazy stuff, but what surprised me is that such s patch > > got merged in urcu even though its known broken for a number of > > architectures. > > It did not get merged into urcu. It is instead used directly by a > number of people for a number of concurrent algorithms. Yah, Mathieu also already pointed that out. It seems I really cannot deal with github well -- that website always terminally confuses me. > > > But it would not be hard for userspace code to force IPIs by repeatedly > > > awakening higher-priority threads that sleep immediately after being > > > awakened, right? > > > > RT tasks are not readily available to !root, and the user might have > > been constrained to a subset of available CPUs. > > So non-idle non-nohz CPUs never get IPIed for wakeups of SCHED_OTHER > threads? Sure, but SCHED_OTHER auto throttles in that if there's anything else to run, you get to wait. So you can't generate an IPI storm with it. Also, again, we can be limited to a subset of CPUs. > > My thinking was that if we observe '!= mm' that CPU will have to do a > > context switch in order to make it true. That context switch will > > provide the ordering we're after so all is well. > > > > Quite possible there's a hole in, but since I'm running on fumes someone > > needs to spell it out for me :-) > > This would be the https://marc.info/?l=linux-kernel&m=126349766324224&w=2 > URL below. > > Which might or might not still be applicable. I think we actually have those two smp_mb()'s around the rq->curr assignment. we have smp_mb__before_spinlock(), which per the argument here: https://lkml.kernel.org/r/20170607162013.755917928@infradead.org is actually a full MB, irrespective of that weird smp_wmb() definition we have now. And we have switch_mm() on the other side. > > > I was intending to base this on the last few versions of a 2010 patch, > > > but maybe things have changed: > > > > > > https://marc.info/?l=linux-kernel&m=126358017229620&w=2 > > > https://marc.info/?l=linux-kernel&m=126436996014016&w=2 > > > https://marc.info/?l=linux-kernel&m=126601479802978&w=2 > > > https://marc.info/?l=linux-kernel&m=126970692903302&w=2 > > > > > > Discussion here: > > > > > > https://marc.info/?l=linux-kernel&m=126349766324224&w=2 > > > > > > The discussion led to acquiring the runqueue locks, as there was > > > otherwise a need to add code to the scheduler fastpaths. > > > > TL;DR.. that's far too much to trawl through. > > So we re-derive it from first principles instead? ;-) Yep, that's what I usually do anyway, who knows what kind of crazy our younger selves were up to ;-)
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-26 17:50 +0200 |
| Message-ID | <u7wJs-2ak-21@gated-at.bofh.it> |
| In reply to | #1696853 |
On Wed, Jul 26, 2017 at 09:41:28AM +0200, Peter Zijlstra wrote: > On Tue, Jul 25, 2017 at 04:59:36PM -0700, Paul E. McKenney wrote: > > On Tue, Jul 25, 2017 at 11:55:10PM +0200, Peter Zijlstra wrote: > > > > People always do crazy stuff, but what surprised me is that such s patch > > > got merged in urcu even though its known broken for a number of > > > architectures. > > > > It did not get merged into urcu. It is instead used directly by a > > number of people for a number of concurrent algorithms. > > Yah, Mathieu also already pointed that out. It seems I really cannot > deal with github well -- that website always terminally confuses me. > > > > > But it would not be hard for userspace code to force IPIs by repeatedly > > > > awakening higher-priority threads that sleep immediately after being > > > > awakened, right? > > > > > > RT tasks are not readily available to !root, and the user might have > > > been constrained to a subset of available CPUs. > > > > So non-idle non-nohz CPUs never get IPIed for wakeups of SCHED_OTHER > > threads? > > Sure, but SCHED_OTHER auto throttles in that if there's anything else to > run, you get to wait. So you can't generate an IPI storm with it. Also, > again, we can be limited to a subset of CPUs. OK, what is its auto-throttle policy? One round of IPIs per jiffy or some such? Does this auto-throttling also apply if the user is running a CPU-bound SCHED_BATCH or SCHED_IDLE task on each CPU, and periodically waking up one of a large group of SCHED_OTHER tasks, where the SCHED_OTHER tasks immediately sleep upon being awakened? > > > My thinking was that if we observe '!= mm' that CPU will have to do a > > > context switch in order to make it true. That context switch will > > > provide the ordering we're after so all is well. > > > > > > Quite possible there's a hole in, but since I'm running on fumes someone > > > needs to spell it out for me :-) > > > > This would be the https://marc.info/?l=linux-kernel&m=126349766324224&w=2 > > URL below. > > > > Which might or might not still be applicable. > > I think we actually have those two smp_mb()'s around the rq->curr > assignment. > > we have smp_mb__before_spinlock(), which per the argument here: > > https://lkml.kernel.org/r/20170607162013.755917928@infradead.org > > is actually a full MB, irrespective of that weird smp_wmb() definition > we have now. And we have switch_mm() on the other side. OK, and the rq->curr assignment is in common code, correct? Does this allow the IPI-only-requesting-process approach to live entirely within common code? The 2010 email thread ended up with sys_membarrier() acquiring the runqueue lock for each CPU, because doing otherwise meant adding code to the scheduler fastpath. Don't we still need to do this? https://marc.info/?l=linux-kernel&m=126341138408407&w=2 https://marc.info/?l=linux-kernel&m=126349766324224&w=2 > > > > I was intending to base this on the last few versions of a 2010 patch, > > > > but maybe things have changed: > > > > > > > > https://marc.info/?l=linux-kernel&m=126358017229620&w=2 > > > > https://marc.info/?l=linux-kernel&m=126436996014016&w=2 > > > > https://marc.info/?l=linux-kernel&m=126601479802978&w=2 > > > > https://marc.info/?l=linux-kernel&m=126970692903302&w=2 > > > > > > > > Discussion here: > > > > > > > > https://marc.info/?l=linux-kernel&m=126349766324224&w=2 > > > > > > > > The discussion led to acquiring the runqueue locks, as there was > > > > otherwise a need to add code to the scheduler fastpaths. > > > > > > TL;DR.. that's far too much to trawl through. > > > > So we re-derive it from first principles instead? ;-) > > Yep, that's what I usually do anyway, who knows what kind of crazy our > younger selves were up to ;-) In my experience, it ends up being a type of crazy worth ignoring only if I don't ignore it. ;-) Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 10:40 +0200 |
| Message-ID | <u7MuR-3NX-9@gated-at.bofh.it> |
| In reply to | #1697319 |
On Wed, Jul 26, 2017 at 08:41:10AM -0700, Paul E. McKenney wrote:
> On Wed, Jul 26, 2017 at 09:41:28AM +0200, Peter Zijlstra wrote:
> > On Tue, Jul 25, 2017 at 04:59:36PM -0700, Paul E. McKenney wrote:
> > Sure, but SCHED_OTHER auto throttles in that if there's anything else to
> > run, you get to wait. So you can't generate an IPI storm with it. Also,
> > again, we can be limited to a subset of CPUs.
>
> OK, what is its auto-throttle policy? One round of IPIs per jiffy or
> some such?
No. Its called wakeup latency :-) Your SCHED_OTHER task will not get to
insta-run all the time. If there are other tasks already running, we'll
not IPI unless it should preempt.
If its idle, nobody cares..
> Does this auto-throttling also apply if the user is running a CPU-bound
> SCHED_BATCH or SCHED_IDLE task on each CPU, and periodically waking up
> one of a large group of SCHED_OTHER tasks, where the SCHED_OTHER tasks
> immediately sleep upon being awakened?
SCHED_BATCH is even more likely to suffer wakeup latency since it will
never preempt anything.
> OK, and the rq->curr assignment is in common code, correct? Does this
> allow the IPI-only-requesting-process approach to live entirely within
> common code?
That is the idea.
> The 2010 email thread ended up with sys_membarrier() acquiring the
> runqueue lock for each CPU,
Yes, that's something I'm not happy with. Machine wide banging of that
lock will be a performance no-no.
> because doing otherwise meant adding code to the scheduler fastpath.
And that's obviously another thing I'm not happy with either.
> Don't we still need to do this?
>
> https://marc.info/?l=linux-kernel&m=126341138408407&w=2
> https://marc.info/?l=linux-kernel&m=126349766324224&w=2
I don't know.. those seem focussed on mm_cpumask() and we can't use that
per Will's email.
So I think we need to think anew on this, start from the ground up.
What is missing for this:
static void ipi_mb(void *info)
{
smp_mb(); // IPIs should be serializing but paranoid
}
sys_membarrier()
{
smp_mb(); // because sysenter isn't an unconditional mb
for_each_online_cpu(cpu) {
struct task_struct *p;
rcu_read_lock();
p = task_rcu_dereference(&cpu_curr(cpu));
if (p && p->mm == current->mm)
__set_bit(cpus, cpu);
rcu_read_unlock();
}
on_cpu_cpu_mask(cpus, ipi_mb, NULL, 1); // does local smp_mb() too
}
VS
__schedule()
{
spin_lock(&rq->lock);
smp_mb__after_spinlock(); // really full mb implied
/* lots */
if (likely(prev != next)) {
rq->curr = next;
context_switch() {
switch_mm();
switch_to();
// neither need imply a barrier
spin_unlock(&rq->lock);
}
}
}
So I think we need either switch_mm() or switch_to() to imply a full
barrier for this to work, otherwise we get:
CPU0 CPU1
lock rq->lock
mb
rq->curr = A
unlock rq->lock
lock rq->lock
mb
sys_membarrier()
mb
for_each_online_cpu()
p = A
// no match no IPI
mb
rq->curr = B
unlock rq->lock
And that's bad, because now CPU0 doesn't have an MB happening _after_
sys_membarrier() if B matches.
So without audit, I only know of PPC and Alpha not having a barrier in
either switch_*().
x86 obviously has barriers all over the place, arm has a super duper
heavy barrier in switch_to().
One option might be to resurrect spin_unlock_wait(), although to use
that here is really ugly too, but it would avoid thrashing the
rq->lock.
I think it'd end up having to look like:
rq = cpu_rq(cpu);
again:
rcu_read_lock()
p = task_rcu_dereference(&rq->curr);
if (p) {
raw_spin_unlock_wait(&rq->lock);
q = task_rcu_dereference(&rq->curr);
if (q != p) {
rcu_read_unlock();
goto again;
}
}
...
which is just about as horrible as it looks.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 15:10 +0200 |
| Message-ID | <u7QIb-6w8-23@gated-at.bofh.it> |
| In reply to | #1697803 |
On Thu, Jul 27, 2017 at 10:30:03AM +0200, Peter Zijlstra wrote:
> On Wed, Jul 26, 2017 at 08:41:10AM -0700, Paul E. McKenney wrote:
> > On Wed, Jul 26, 2017 at 09:41:28AM +0200, Peter Zijlstra wrote:
> > > On Tue, Jul 25, 2017 at 04:59:36PM -0700, Paul E. McKenney wrote:
>
> > > Sure, but SCHED_OTHER auto throttles in that if there's anything else to
> > > run, you get to wait. So you can't generate an IPI storm with it. Also,
> > > again, we can be limited to a subset of CPUs.
> >
> > OK, what is its auto-throttle policy? One round of IPIs per jiffy or
> > some such?
>
> No. Its called wakeup latency :-) Your SCHED_OTHER task will not get to
> insta-run all the time. If there are other tasks already running, we'll
> not IPI unless it should preempt.
>
> If its idle, nobody cares..
So it does IPI immediately sometimes.
> > Does this auto-throttling also apply if the user is running a CPU-bound
> > SCHED_BATCH or SCHED_IDLE task on each CPU, and periodically waking up
> > one of a large group of SCHED_OTHER tasks, where the SCHED_OTHER tasks
> > immediately sleep upon being awakened?
>
> SCHED_BATCH is even more likely to suffer wakeup latency since it will
> never preempt anything.
Ahem. In this scenario, SCHED_BATCH is already running on a the CPU in
question, and a SCHED_OTHER task is awakened from some other CPU.
Do we IPI in that case.
> > OK, and the rq->curr assignment is in common code, correct? Does this
> > allow the IPI-only-requesting-process approach to live entirely within
> > common code?
>
> That is the idea.
>
> > The 2010 email thread ended up with sys_membarrier() acquiring the
> > runqueue lock for each CPU,
>
> Yes, that's something I'm not happy with. Machine wide banging of that
> lock will be a performance no-no.
Regardless of whether or not we need to acquire runqueue locks, IPI,
or whatever other distasteful operation, we should also be able to
throttle and batch operations to at least some extent.
> > because doing otherwise meant adding code to the scheduler fastpath.
>
> And that's obviously another thing I'm not happy with either.
Nor should you or anyone be.
> > Don't we still need to do this?
> >
> > https://marc.info/?l=linux-kernel&m=126341138408407&w=2
> > https://marc.info/?l=linux-kernel&m=126349766324224&w=2
>
> I don't know.. those seem focussed on mm_cpumask() and we can't use that
> per Will's email.
>
> So I think we need to think anew on this, start from the ground up.
Probably from several points in the ground, but OK...
> What is missing for this:
>
> static void ipi_mb(void *info)
> {
> smp_mb(); // IPIs should be serializing but paranoid
> }
>
>
> sys_membarrier()
> {
> smp_mb(); // because sysenter isn't an unconditional mb
>
> for_each_online_cpu(cpu) {
> struct task_struct *p;
>
> rcu_read_lock();
> p = task_rcu_dereference(&cpu_curr(cpu));
> if (p && p->mm == current->mm)
> __set_bit(cpus, cpu);
> rcu_read_unlock();
> }
>
> on_cpu_cpu_mask(cpus, ipi_mb, NULL, 1); // does local smp_mb() too
> }
>
> VS
>
> __schedule()
> {
> spin_lock(&rq->lock);
> smp_mb__after_spinlock(); // really full mb implied
>
> /* lots */
>
> if (likely(prev != next)) {
>
> rq->curr = next;
>
> context_switch() {
> switch_mm();
> switch_to();
> // neither need imply a barrier
>
> spin_unlock(&rq->lock);
> }
> }
> }
>
>
>
>
> So I think we need either switch_mm() or switch_to() to imply a full
> barrier for this to work, otherwise we get:
>
> CPU0 CPU1
>
>
> lock rq->lock
> mb
>
> rq->curr = A
>
> unlock rq->lock
>
> lock rq->lock
> mb
>
> sys_membarrier()
>
> mb
>
> for_each_online_cpu()
> p = A
> // no match no IPI
>
> mb
> rq->curr = B
>
> unlock rq->lock
>
>
> And that's bad, because now CPU0 doesn't have an MB happening _after_
> sys_membarrier() if B matches.
Yes, this looks somewhat similar to the scenario that Mathieu pointed out
back in 2010: https://marc.info/?l=linux-kernel&m=126349766324224&w=2
> So without audit, I only know of PPC and Alpha not having a barrier in
> either switch_*().
>
> x86 obviously has barriers all over the place, arm has a super duper
> heavy barrier in switch_to().
Agreed, if we are going to rely on ->mm, we need ordering on assignment
to it.
> One option might be to resurrect spin_unlock_wait(), although to use
> that here is really ugly too, but it would avoid thrashing the
> rq->lock.
>
> I think it'd end up having to look like:
>
> rq = cpu_rq(cpu);
> again:
> rcu_read_lock()
> p = task_rcu_dereference(&rq->curr);
> if (p) {
> raw_spin_unlock_wait(&rq->lock);
> q = task_rcu_dereference(&rq->curr);
> if (q != p) {
> rcu_read_unlock();
> goto again;
> }
> }
> ...
>
> which is just about as horrible as it looks.
It does indeed look a bit suboptimal.
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 15:50 +0200 |
| Message-ID | <u7RkR-6KK-11@gated-at.bofh.it> |
| In reply to | #1697975 |
On Thu, Jul 27, 2017 at 06:08:16AM -0700, Paul E. McKenney wrote: > > No. Its called wakeup latency :-) Your SCHED_OTHER task will not get to > > insta-run all the time. If there are other tasks already running, we'll > > not IPI unless it should preempt. > > > > If its idle, nobody cares.. > > So it does IPI immediately sometimes. > > > > Does this auto-throttling also apply if the user is running a CPU-bound > > > SCHED_BATCH or SCHED_IDLE task on each CPU, and periodically waking up > > > one of a large group of SCHED_OTHER tasks, where the SCHED_OTHER tasks > > > immediately sleep upon being awakened? > > > > SCHED_BATCH is even more likely to suffer wakeup latency since it will > > never preempt anything. > > Ahem. In this scenario, SCHED_BATCH is already running on a the CPU in > question, and a SCHED_OTHER task is awakened from some other CPU. > > Do we IPI in that case. So I'm a bit confused as to where you're trying to go with this. I'm saying that if there are other users of our CPU, we can't significantly disturb them with IPIs. Yes, we'll sometimes IPI in order to do a preemption on wakeup. But we cannot always win that. There is no wakeup triggered starvation case. If you create a thread per CPU and have them insta sleep after wakeup, and then keep prodding them to wakeup. They will disturb things less than if they were while(1); loops. If the machine is otherwise idle, nobody cares. If there are other tasks on the system, the IPI rate is limited to the wakeup latency of you tasks. And any of this is limited to the CPUs we're allowed to run in the first place. So yes, the occasional IPI happens, but if there's other tasks, we can't disturb them more than we could with while(1); tasks. OTOH, something like: while(1) synchronize_sched_expedited(); as per your proposed patch, will spray IPIs to all CPUs and at high rates.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-27 16:40 +0200 |
| Message-ID | <u7S7g-7g8-35@gated-at.bofh.it> |
| In reply to | #1697997 |
On Thu, Jul 27, 2017 at 07:32:05AM -0700, Paul E. McKenney wrote: > > as per your proposed patch, will spray IPIs to all CPUs and at high > > rates. > > OK, I have updated my patch to do throttling. But not respect cpusets. Which is the other important point. The scheduler based IPIs are limited to where we are allowed to place tasks (as are the TLB ones for that matter, for the exact same reason).
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-27 16:50 +0200 |
| Message-ID | <u7SgW-7jy-15@gated-at.bofh.it> |
| In reply to | #1698037 |
On Thu, Jul 27, 2017 at 04:36:24PM +0200, Peter Zijlstra wrote: > On Thu, Jul 27, 2017 at 07:32:05AM -0700, Paul E. McKenney wrote: > > > as per your proposed patch, will spray IPIs to all CPUs and at high > > > rates. > > > > OK, I have updated my patch to do throttling. > > But not respect cpusets. Which is the other important point. > > The scheduler based IPIs are limited to where we are allowed to place > tasks (as are the TLB ones for that matter, for the exact same reason). Yes, these are disadvantages, no argument there. But I will nevertheless be carrying some variant of this patch until something better exists, including that something being tested and found satisfactory by the various people asking for this. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web