Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1695159 > unrolled thread

[PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option

Started by"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
First post2017-07-25 00:00 +0200
Last post2017-07-27 16:40 +0200
Articles 16 on this page of 56 — 5 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 00:00 +0200
    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-25 06:30 +0200
      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:30 +0200
    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-25 15:20 +0200
      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 18:40 +0200
      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 18:50 +0200
        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 19:10 +0200
          Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 19:20 +0200
            Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 21:00 +0200
              Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 21:40 +0200
                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-25 22:30 +0200
                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-25 23:20 +0200
                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 00:00 +0200
                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:40 +0200
                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 00:50 +0200
                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:10 +0200
                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
                          Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
                            Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 20:00 +0200
                              Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 20:40 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-26 22:40 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 23:20 +0200
                                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 03:50 +0200
                                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Mathieu Desnoyers <mathieu.desnoyers@efficios.com> - 2017-07-27 14:40 +0200
                                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:30 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
                              Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 11:00 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 12:30 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:20 +0200
                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 02:00 +0200
                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-26 09:50 +0200
                          Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
                            Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 10:40 +0200
                              Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:10 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:50 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:40 +0200
                                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:50 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 16:00 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 17:30 +0200
                                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:00 +0200
                                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:20 +0200
                                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:30 +0200
                                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200
                                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-27 16:50 +0200
                                        Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Boqun Feng <boqun.feng@gmail.com> - 2017-07-27 16:50 +0200
                                          Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 17:00 +0200
                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Will Deacon <will.deacon@arm.com> - 2017-07-26 11:40 +0200
                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-26 17:50 +0200
                Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 12:20 +0200
                  Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 15:00 +0200
                    Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option Peter Zijlstra <peterz@infradead.org> - 2017-07-27 15:40 +0200
                      Re: [PATCH tip/core/rcu 4/5] sys_membarrier: Add expedited option "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-27 16:40 +0200

Page 3 of 3 — ← Prev page 1 2 [3]


#1698041

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 16:40 +0200
Message-ID<u7S7g-7g8-37@gated-at.bofh.it>
In reply to#1697997
On Thu, Jul 27, 2017 at 03:49:08PM +0200, Peter Zijlstra wrote:
> On Thu, Jul 27, 2017 at 06:08:16AM -0700, Paul E. McKenney wrote:
> 
> > > No. Its called wakeup latency :-) Your SCHED_OTHER task will not get to
> > > insta-run all the time. If there are other tasks already running, we'll
> > > not IPI unless it should preempt.
> > > 
> > > If its idle, nobody cares..
> > 
> > So it does IPI immediately sometimes.
> > 
> > > > Does this auto-throttling also apply if the user is running a CPU-bound
> > > > SCHED_BATCH or SCHED_IDLE task on each CPU, and periodically waking up
> > > > one of a large group of SCHED_OTHER tasks, where the SCHED_OTHER tasks
> > > > immediately sleep upon being awakened?
> > > 
> > > SCHED_BATCH is even more likely to suffer wakeup latency since it will
> > > never preempt anything.
> > 
> > Ahem.  In this scenario, SCHED_BATCH is already running on a the CPU in
> > question, and a SCHED_OTHER task is awakened from some other CPU.
> > 
> > Do we IPI in that case.
> 
> So I'm a bit confused as to where you're trying to go with this.
> 
> I'm saying that if there are other users of our CPU, we can't
> significantly disturb them with IPIs.
> 
> Yes, we'll sometimes IPI in order to do a preemption on wakeup. But we
> cannot always win that. There is no wakeup triggered starvation case.
> 
> If you create a thread per CPU and have them insta sleep after wakeup,
> and then keep prodding them to wakeup. They will disturb things less
> than if they were while(1); loops.
> 
> If the machine is otherwise idle, nobody cares.
> 
> If there are other tasks on the system, the IPI rate is limited to the
> wakeup latency of you tasks.
> 
> 
> And any of this is limited to the CPUs we're allowed to run in the first
> place.
> 
> 
> So yes, the occasional IPI happens, but if there's other tasks, we can't
> disturb them more than we could with while(1); tasks.
> 
> 
> OTOH, something like:
> 
> 	while(1)
> 		synchronize_sched_expedited();
> 
> as per your proposed patch, will spray IPIs to all CPUs and at high
> rates.

OK, I have updated my patch to do throttling.

							Thanx, Paul

------------------------------------------------------------------------

commit 4cd5253094b6d7f9501e21e13aa4e2e78e8a70cd
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date:   Tue Jul 18 13:53:32 2017 -0700

    sys_membarrier: Add expedited option
    
    The sys_membarrier() system call has proven too slow for some use cases,
    which has prompted users to instead rely on TLB shootdown.  Although TLB
    shootdown is much faster, it has the slight disadvantage of not working
    at all on arm and arm64 and also of being vulnerable to reasonable
    optimizations that might skip some IPIs.  However, the Linux kernel
    does not currrently provide a reasonable alternative, so it is hard to
    criticize these users from doing what works for them on a given piece
    of hardware at a given time.
    
    This commit therefore adds an expedited option to the sys_membarrier()
    system call, thus providing a faster mechanism that is portable and
    is not subject to death by optimization.  Note that if more than one
    MEMBARRIER_CMD_SHARED_EXPEDITED sys_membarrier() call happens within
    the same jiffy, all but the first will use synchronize_sched() instead
    of synchronize_sched_expedited().
    
    Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
    [ paulmck: Fix code style issue pointed out by Boqun Feng. ]
    Tested-by: Avi Kivity <avi@scylladb.com>
    Cc: Maged Michael <maged.michael@gmail.com>
    Cc: Andrew Hunter <ahh@google.com>
    Cc: Geoffrey Romer <gromer@google.com>

diff --git a/include/uapi/linux/membarrier.h b/include/uapi/linux/membarrier.h
index e0b108bd2624..5720386d0904 100644
--- a/include/uapi/linux/membarrier.h
+++ b/include/uapi/linux/membarrier.h
@@ -40,6 +40,16 @@
  *                          (non-running threads are de facto in such a
  *                          state). This covers threads from all processes
  *                          running on the system. This command returns 0.
+ * @MEMBARRIER_CMD_SHARED_EXPEDITED:  Execute a memory barrier on all
+ *			    running threads, but in an expedited fashion.
+ *                          Upon return from system call, the caller thread
+ *                          is ensured that all running threads have passed
+ *                          through a state where all memory accesses to
+ *                          user-space addresses match program order between
+ *                          entry to and return from the system call
+ *                          (non-running threads are de facto in such a
+ *                          state). This covers threads from all processes
+ *                          running on the system. This command returns 0.
  *
  * Command to be passed to the membarrier system call. The commands need to
  * be a single bit each, except for MEMBARRIER_CMD_QUERY which is assigned to
@@ -48,6 +58,7 @@
 enum membarrier_cmd {
 	MEMBARRIER_CMD_QUERY = 0,
 	MEMBARRIER_CMD_SHARED = (1 << 0),
+	MEMBARRIER_CMD_SHARED_EXPEDITED = (1 << 1),
 };
 
 #endif /* _UAPI_LINUX_MEMBARRIER_H */
diff --git a/kernel/membarrier.c b/kernel/membarrier.c
index 9f9284f37f8d..587e3bbfae7e 100644
--- a/kernel/membarrier.c
+++ b/kernel/membarrier.c
@@ -22,7 +22,8 @@
  * Bitmask made from a "or" of all commands within enum membarrier_cmd,
  * except MEMBARRIER_CMD_QUERY.
  */
-#define MEMBARRIER_CMD_BITMASK	(MEMBARRIER_CMD_SHARED)
+#define MEMBARRIER_CMD_BITMASK	(MEMBARRIER_CMD_SHARED |		\
+				 MEMBARRIER_CMD_SHARED_EXPEDITED)
 
 /**
  * sys_membarrier - issue memory barriers on a set of threads
@@ -64,6 +65,20 @@ SYSCALL_DEFINE2(membarrier, int, cmd, int, flags)
 		if (num_online_cpus() > 1)
 			synchronize_sched();
 		return 0;
+	case MEMBARRIER_CMD_SHARED_EXPEDITED:
+		if (num_online_cpus() > 1) {
+			static unsigned long lastexp;
+			unsigned long j;
+
+			j = jiffies;
+			if (READ_ONCE(lastexp) == j) {
+				synchronize_sched();
+				WRITE_ONCE(lastexp, j);
+			} else {
+				synchronize_sched_expedited();
+			}
+		}
+		return 0;
 	default:
 		return -EINVAL;
 	}

[toc] | [prev] | [next] | [standalone]


#1698005

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-27 16:00 +0200
Message-ID<u7Ruy-6NU-17@gated-at.bofh.it>
In reply to#1697975
On Thu, Jul 27, 2017 at 06:08:16AM -0700, Paul E. McKenney wrote:

> > So I think we need either switch_mm() or switch_to() to imply a full
> > barrier for this to work, otherwise we get:
> > 
> >   CPU0				CPU1
> > 
> > 
> >   lock rq->lock
> >   mb
> > 
> >   rq->curr = A
> > 
> >   unlock rq->lock
> > 
> >   lock rq->lock
> >   mb
> > 
> > 				sys_membarrier()
> > 
> > 				mb
> > 
> > 				for_each_online_cpu()
> > 				  p = A
> > 				  // no match no IPI
> > 
> > 				mb
> >   rq->curr = B
> > 
> >   unlock rq->lock
> > 
> > 
> > And that's bad, because now CPU0 doesn't have an MB happening _after_
> > sys_membarrier() if B matches.
> 
> Yes, this looks somewhat similar to the scenario that Mathieu pointed out
> back in 2010: https://marc.info/?l=linux-kernel&m=126349766324224&w=2

Yes. Minus the mm_cpumask() worries.

> > So without audit, I only know of PPC and Alpha not having a barrier in
> > either switch_*().
> > 
> > x86 obviously has barriers all over the place, arm has a super duper
> > heavy barrier in switch_to().
> 
> Agreed, if we are going to rely on ->mm, we need ordering on assignment
> to it.

Right, Boqun provided this reordering to show the problem:

  CPU0                                CPU1
 
 
  <in process X>
  lock rq->lock
  mb
 
  rq->curr = A
 
  unlock rq->lock
 
  <switch to process A>
 
  lock rq->lock
  mb
  read Y(reordered)<---+
                       |        store to Y
                       |
                       |        sys_membarrier()
                       |
                       |        mb
                       |
                       |        for_each_online_cpu()
                       |          p = A
                       |          // no match no IPI
                       |
                       |        mb
                       |
                       |        store to X
  rq->curr = B         |
                       |
  unlock rq->lock      |
  <switch to B>        |
  read X               |
                       |
  read Y --------------+

[toc] | [prev] | [next] | [standalone]


#1698102

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-27 17:30 +0200
Message-ID<u7STF-7Md-27@gated-at.bofh.it>
In reply to#1698005
Hi Nick,

See below,

On Thu, Jul 27, 2017 at 03:56:10PM +0200, Peter Zijlstra wrote:
> On Thu, Jul 27, 2017 at 06:08:16AM -0700, Paul E. McKenney wrote:
> 
> > > So I think we need either switch_mm() or switch_to() to imply a full
> > > barrier for this to work, otherwise we get:
> > > 
> > >   CPU0				CPU1
> > > 
> > > 
> > >   lock rq->lock
> > >   mb
> > > 
> > >   rq->curr = A
> > > 
> > >   unlock rq->lock
> > > 
> > >   lock rq->lock
> > >   mb
> > > 
> > > 				sys_membarrier()
> > > 
> > > 				mb
> > > 
> > > 				for_each_online_cpu()
> > > 				  p = A
> > > 				  // no match no IPI
> > > 
> > > 				mb
> > >   rq->curr = B
> > > 
> > >   unlock rq->lock
> > > 
> > > 
> > > And that's bad, because now CPU0 doesn't have an MB happening _after_
> > > sys_membarrier() if B matches.
> > 
> > Yes, this looks somewhat similar to the scenario that Mathieu pointed out
> > back in 2010: https://marc.info/?l=linux-kernel&m=126349766324224&w=2
> 
> Yes. Minus the mm_cpumask() worries.
> 
> > > So without audit, I only know of PPC and Alpha not having a barrier in
> > > either switch_*().
> > > 
> > > x86 obviously has barriers all over the place, arm has a super duper
> > > heavy barrier in switch_to().
> > 
> > Agreed, if we are going to rely on ->mm, we need ordering on assignment
> > to it.
> 
> Right, Boqun provided this reordering to show the problem:
> 
>   CPU0                                CPU1
>  
>  
>   <in process X>
>   lock rq->lock
>   mb
>  
>   rq->curr = A
>  
>   unlock rq->lock
>  
>   <switch to process A>
>  
>   lock rq->lock
>   mb
>   read Y(reordered)<---+
>                        |        store to Y
>                        |
>                        |        sys_membarrier()
>                        |
>                        |        mb
>                        |
>                        |        for_each_online_cpu()
>                        |          p = A
>                        |          // no match no IPI
>                        |
>                        |        mb
>                        |
>                        |        store to X
>   rq->curr = B         |
>                        |
>   unlock rq->lock      |
>   <switch to B>        |
>   read X               |
>                        |
>   read Y --------------+

In order to make this work we need either switch_to() or switch_mm() to
provide smp_mb(). Now you're recently taken that out on PPC and I'm
thinking you're not keen to have to put it back in.

Mathieu was wondering if placing it in switch_mm() would be less onerous
on performance, thinking that address space changes are more expensive
in any case, seeing how they have a tail of cache and translation
misses. I'm thinking you're not happy either way :-)

Opinions?

[toc] | [prev] | [next] | [standalone]


#1698007

FromBoqun Feng <boqun.feng@gmail.com>
Date2017-07-27 16:00 +0200
Message-ID<u7Ruy-6NU-21@gated-at.bofh.it>
In reply to#1697975

[Multipart message — attachments visible in raw view] — view raw

Hi Paul,

I have a side question out of curiosity:

How does synchronize_sched() work properly for sys_membarrier()?

sys_membarrier() requires every other CPU does a smp_mb() before it
returns, and I know synchronize_sched() will wait until all CPUs running
a kernel thread do a context-switch, which has a smp_mb(). However, I
believe sched flavor RCU treat CPU running a user thread as a quiesent
state, so synchronize_sched() could return without that CPU does a
context switch. 

So why could we use synchronize_sched() for sys_membarrier()?

In particular, could the following happens?

	CPU 0:				CPU 1:
	=========================	==========================
	<in user space>			<in user space>
					{read Y}(reordered) <------------------------------+
	store Y;			                                                   |
					read X; --------------------------------------+    |
	sys_membarrier():		<timer interrupt>                             |    |
	  synchronize_sched();		  update_process_times(user): //user == true  |    |
	  				    rcu_check_callbacks(usr):                 |    |
					      if (user || ..) {                       |    |
					        rcu_sched_qs()                        |    |
						...                                   |    |
						<report quesient state in softirq>    |    |
					<return to user space>                        |    |
					read Y; --------------------------------------+----+
	store X;			                                              |
					{read X}(reordered) <-------------------------+

I assume the timer interrupt handler, which interrupts a user space and
reports a quiesent state for sched flavor RCU, may not have a smp_mb()
in some code path.

I may miss something subtle, but it just not very obvious how
synchronize_sched() will guarantee a remote CPU running in userspace to
do a smp_mb() before it returns, this is at least not in RCU
requirements, right?

Regards,
Boqun

[toc] | [prev] | [next] | [standalone]


#1698014

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 16:20 +0200
Message-ID<u7RNU-79J-7@gated-at.bofh.it>
In reply to#1698007
On Thu, Jul 27, 2017 at 09:55:51PM +0800, Boqun Feng wrote:
> Hi Paul,
> 
> I have a side question out of curiosity:
> 
> How does synchronize_sched() work properly for sys_membarrier()?
> 
> sys_membarrier() requires every other CPU does a smp_mb() before it
> returns, and I know synchronize_sched() will wait until all CPUs running
> a kernel thread do a context-switch, which has a smp_mb(). However, I
> believe sched flavor RCU treat CPU running a user thread as a quiesent
> state, so synchronize_sched() could return without that CPU does a
> context switch. 
> 
> So why could we use synchronize_sched() for sys_membarrier()?
> 
> In particular, could the following happens?
> 
> 	CPU 0:				CPU 1:
> 	=========================	==========================
> 	<in user space>			<in user space>
> 					{read Y}(reordered) <------------------------------+
> 	store Y;			                                                   |
> 					read X; --------------------------------------+    |
> 	sys_membarrier():		<timer interrupt>                             |    |
> 	  synchronize_sched();		  update_process_times(user): //user == true  |    |
> 	  				    rcu_check_callbacks(usr):                 |    |
> 					      if (user || ..) {                       |    |
> 					        rcu_sched_qs()                        |    |
> 						...                                   |    |
> 						<report quesient state in softirq>    |    |

The reporting of the quiescent state will acquire the leaf rcu_node
structure's lock, with an smp_mb__after_unlock_lock(), which will
one way or another be a full memory barrier.  So the reorderings
cannot happen.

Unless I am missing something subtle.  ;-)

						Thanx, Paul

> 					<return to user space>                        |    |
> 					read Y; --------------------------------------+----+
> 	store X;			                                              |
> 					{read X}(reordered) <-------------------------+
> 
> I assume the timer interrupt handler, which interrupts a user space and
> reports a quiesent state for sched flavor RCU, may not have a smp_mb()
> in some code path.
> 
> I may miss something subtle, but it just not very obvious how
> synchronize_sched() will guarantee a remote CPU running in userspace to
> do a smp_mb() before it returns, this is at least not in RCU
> requirements, right?
> 
> Regards,
> Boqun

[toc] | [prev] | [next] | [standalone]


#1698020

FromBoqun Feng <boqun.feng@gmail.com>
Date2017-07-27 16:30 +0200
Message-ID<u7RXA-7d4-15@gated-at.bofh.it>
In reply to#1698014

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 27, 2017 at 07:16:33AM -0700, Paul E. McKenney wrote:
> On Thu, Jul 27, 2017 at 09:55:51PM +0800, Boqun Feng wrote:
> > Hi Paul,
> > 
> > I have a side question out of curiosity:
> > 
> > How does synchronize_sched() work properly for sys_membarrier()?
> > 
> > sys_membarrier() requires every other CPU does a smp_mb() before it
> > returns, and I know synchronize_sched() will wait until all CPUs running
> > a kernel thread do a context-switch, which has a smp_mb(). However, I
> > believe sched flavor RCU treat CPU running a user thread as a quiesent
> > state, so synchronize_sched() could return without that CPU does a
> > context switch. 
> > 
> > So why could we use synchronize_sched() for sys_membarrier()?
> > 
> > In particular, could the following happens?
> > 
> > 	CPU 0:				CPU 1:
> > 	=========================	==========================
> > 	<in user space>			<in user space>
> > 					{read Y}(reordered) <------------------------------+
> > 	store Y;			                                                   |
> > 					read X; --------------------------------------+    |
> > 	sys_membarrier():		<timer interrupt>                             |    |
> > 	  synchronize_sched();		  update_process_times(user): //user == true  |    |
> > 	  				    rcu_check_callbacks(usr):                 |    |
> > 					      if (user || ..) {                       |    |
> > 					        rcu_sched_qs()                        |    |
> > 						...                                   |    |
> > 						<report quesient state in softirq>    |    |
> 
> The reporting of the quiescent state will acquire the leaf rcu_node
> structure's lock, with an smp_mb__after_unlock_lock(), which will
> one way or another be a full memory barrier.  So the reorderings
> cannot happen.
> 
> Unless I am missing something subtle.  ;-)
> 

Well, smp_mb__after_unlock_lock() in ARM64 is a no-op, and ARM64's lock
doesn't provide a smp_mb().

So my point is more like: synchronize_sched() happens to be a
sys_membarrier() because of some implementation detail, and if some day
we come up with a much cheaper way to implement sched flavor
RCU(hopefully!), synchronize_sched() may be not good for the job. So at
least, we'd better document this somewhere?

Regards,
Boqun

> 						Thanx, Paul
> 
> > 					<return to user space>                        |    |
> > 					read Y; --------------------------------------+----+
> > 	store X;			                                              |
> > 					{read X}(reordered) <-------------------------+
> > 
> > I assume the timer interrupt handler, which interrupts a user space and
> > reports a quiesent state for sched flavor RCU, may not have a smp_mb()
> > in some code path.
> > 
> > I may miss something subtle, but it just not very obvious how
> > synchronize_sched() will guarantee a remote CPU running in userspace to
> > do a smp_mb() before it returns, this is at least not in RCU
> > requirements, right?
> > 
> > Regards,
> > Boqun
> 
> 

[toc] | [prev] | [next] | [standalone]


#1698043

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 16:40 +0200
Message-ID<u7S7g-7g8-45@gated-at.bofh.it>
In reply to#1698020
On Thu, Jul 27, 2017 at 10:29:55PM +0800, Boqun Feng wrote:
> On Thu, Jul 27, 2017 at 07:16:33AM -0700, Paul E. McKenney wrote:
> > On Thu, Jul 27, 2017 at 09:55:51PM +0800, Boqun Feng wrote:
> > > Hi Paul,
> > > 
> > > I have a side question out of curiosity:
> > > 
> > > How does synchronize_sched() work properly for sys_membarrier()?
> > > 
> > > sys_membarrier() requires every other CPU does a smp_mb() before it
> > > returns, and I know synchronize_sched() will wait until all CPUs running
> > > a kernel thread do a context-switch, which has a smp_mb(). However, I
> > > believe sched flavor RCU treat CPU running a user thread as a quiesent
> > > state, so synchronize_sched() could return without that CPU does a
> > > context switch. 
> > > 
> > > So why could we use synchronize_sched() for sys_membarrier()?
> > > 
> > > In particular, could the following happens?
> > > 
> > > 	CPU 0:				CPU 1:
> > > 	=========================	==========================
> > > 	<in user space>			<in user space>
> > > 					{read Y}(reordered) <------------------------------+
> > > 	store Y;			                                                   |
> > > 					read X; --------------------------------------+    |
> > > 	sys_membarrier():		<timer interrupt>                             |    |
> > > 	  synchronize_sched();		  update_process_times(user): //user == true  |    |
> > > 	  				    rcu_check_callbacks(usr):                 |    |
> > > 					      if (user || ..) {                       |    |
> > > 					        rcu_sched_qs()                        |    |
> > > 						...                                   |    |
> > > 						<report quesient state in softirq>    |    |
> > 
> > The reporting of the quiescent state will acquire the leaf rcu_node
> > structure's lock, with an smp_mb__after_unlock_lock(), which will
> > one way or another be a full memory barrier.  So the reorderings
> > cannot happen.
> > 
> > Unless I am missing something subtle.  ;-)
> > 
> 
> Well, smp_mb__after_unlock_lock() in ARM64 is a no-op, and ARM64's lock
> doesn't provide a smp_mb().
> 
> So my point is more like: synchronize_sched() happens to be a
> sys_membarrier() because of some implementation detail, and if some day
> we come up with a much cheaper way to implement sched flavor
> RCU(hopefully!), synchronize_sched() may be not good for the job. So at
> least, we'd better document this somewhere?

Last I heard, ARM's unlock/lock acted as a full barrier.  Will?

Please see the synchronize_sched() comment header for the documentation
you are asking for.  And the "Memory-Barrier Guarantees" section of
Documentation/RCU/Design/Requirements/Requirements.html.

							Thanx, Paul

> Regards,
> Boqun
> 
> > 						Thanx, Paul
> > 
> > > 					<return to user space>                        |    |
> > > 					read Y; --------------------------------------+----+
> > > 	store X;			                                              |
> > > 					{read X}(reordered) <-------------------------+
> > > 
> > > I assume the timer interrupt handler, which interrupts a user space and
> > > reports a quiesent state for sched flavor RCU, may not have a smp_mb()
> > > in some code path.
> > > 
> > > I may miss something subtle, but it just not very obvious how
> > > synchronize_sched() will guarantee a remote CPU running in userspace to
> > > do a smp_mb() before it returns, this is at least not in RCU
> > > requirements, right?
> > > 
> > > Regards,
> > > Boqun
> > 
> > 

[toc] | [prev] | [next] | [standalone]


#1698046

FromWill Deacon <will.deacon@arm.com>
Date2017-07-27 16:50 +0200
Message-ID<u7SgW-7jy-7@gated-at.bofh.it>
In reply to#1698043
On Thu, Jul 27, 2017 at 07:36:58AM -0700, Paul E. McKenney wrote:
> On Thu, Jul 27, 2017 at 10:29:55PM +0800, Boqun Feng wrote:
> > On Thu, Jul 27, 2017 at 07:16:33AM -0700, Paul E. McKenney wrote:
> > > On Thu, Jul 27, 2017 at 09:55:51PM +0800, Boqun Feng wrote:
> > > > Hi Paul,
> > > > 
> > > > I have a side question out of curiosity:
> > > > 
> > > > How does synchronize_sched() work properly for sys_membarrier()?
> > > > 
> > > > sys_membarrier() requires every other CPU does a smp_mb() before it
> > > > returns, and I know synchronize_sched() will wait until all CPUs running
> > > > a kernel thread do a context-switch, which has a smp_mb(). However, I
> > > > believe sched flavor RCU treat CPU running a user thread as a quiesent
> > > > state, so synchronize_sched() could return without that CPU does a
> > > > context switch. 
> > > > 
> > > > So why could we use synchronize_sched() for sys_membarrier()?
> > > > 
> > > > In particular, could the following happens?
> > > > 
> > > > 	CPU 0:				CPU 1:
> > > > 	=========================	==========================
> > > > 	<in user space>			<in user space>
> > > > 					{read Y}(reordered) <------------------------------+
> > > > 	store Y;			                                                   |
> > > > 					read X; --------------------------------------+    |
> > > > 	sys_membarrier():		<timer interrupt>                             |    |
> > > > 	  synchronize_sched();		  update_process_times(user): //user == true  |    |
> > > > 	  				    rcu_check_callbacks(usr):                 |    |
> > > > 					      if (user || ..) {                       |    |
> > > > 					        rcu_sched_qs()                        |    |
> > > > 						...                                   |    |
> > > > 						<report quesient state in softirq>    |    |
> > > 
> > > The reporting of the quiescent state will acquire the leaf rcu_node
> > > structure's lock, with an smp_mb__after_unlock_lock(), which will
> > > one way or another be a full memory barrier.  So the reorderings
> > > cannot happen.
> > > 
> > > Unless I am missing something subtle.  ;-)
> > > 
> > 
> > Well, smp_mb__after_unlock_lock() in ARM64 is a no-op, and ARM64's lock
> > doesn't provide a smp_mb().
> > 
> > So my point is more like: synchronize_sched() happens to be a
> > sys_membarrier() because of some implementation detail, and if some day
> > we come up with a much cheaper way to implement sched flavor
> > RCU(hopefully!), synchronize_sched() may be not good for the job. So at
> > least, we'd better document this somewhere?
> 
> Last I heard, ARM's unlock/lock acted as a full barrier.  Will?

Yeah, should do. unlock is release, lock is acquire and we're RCsc.

Will

[toc] | [prev] | [next] | [standalone]


#1698049

FromBoqun Feng <boqun.feng@gmail.com>
Date2017-07-27 16:50 +0200
Message-ID<u7SgW-7jy-17@gated-at.bofh.it>
In reply to#1698043

[Multipart message — attachments visible in raw view] — view raw

On Thu, Jul 27, 2017 at 07:36:58AM -0700, Paul E. McKenney wrote:
> > > 
> > > The reporting of the quiescent state will acquire the leaf rcu_node
> > > structure's lock, with an smp_mb__after_unlock_lock(), which will
> > > one way or another be a full memory barrier.  So the reorderings
> > > cannot happen.
> > > 
> > > Unless I am missing something subtle.  ;-)
> > > 
> > 
> > Well, smp_mb__after_unlock_lock() in ARM64 is a no-op, and ARM64's lock
> > doesn't provide a smp_mb().
> > 
> > So my point is more like: synchronize_sched() happens to be a
> > sys_membarrier() because of some implementation detail, and if some day
> > we come up with a much cheaper way to implement sched flavor
> > RCU(hopefully!), synchronize_sched() may be not good for the job. So at
> > least, we'd better document this somewhere?
> 
> Last I heard, ARM's unlock/lock acted as a full barrier.  Will?
> 
> Please see the synchronize_sched() comment header for the documentation
> you are asking for.  And the "Memory-Barrier Guarantees" section of
> Documentation/RCU/Design/Requirements/Requirements.html.
> 

All those barrier guarantees are subject to a RCU read-side critical
section with a synchonize_*(), IIRC, for example:

 * On systems with more than one CPU, when synchronize_sched() returns,
 * each CPU is guaranteed to have executed a full memory barrier since the
 * end of its last RCU-sched read-side critical section whose beginning
 * preceded the call to synchronize_sched().  In addition, each CPU having

, which is not the case for a quiesent state without a read-side
critical section(i.e. non-context-switch quiesent state for sched Flavor)

I've read those requirements and could not find one to explain why there
will be a full barrier emitted in an interrupted user-space program.

Regards,
Boqun

> 							Thanx, Paul
> 
> > Regards,
> > Boqun
> > 
> > > 						Thanx, Paul
> > > 
> > > > 					<return to user space>                        |    |
> > > > 					read Y; --------------------------------------+----+
> > > > 	store X;			                                              |
> > > > 					{read X}(reordered) <-------------------------+
> > > > 
> > > > I assume the timer interrupt handler, which interrupts a user space and
> > > > reports a quiesent state for sched flavor RCU, may not have a smp_mb()
> > > > in some code path.
> > > > 
> > > > I may miss something subtle, but it just not very obvious how
> > > > synchronize_sched() will guarantee a remote CPU running in userspace to
> > > > do a smp_mb() before it returns, this is at least not in RCU
> > > > requirements, right?
> > > > 
> > > > Regards,
> > > > Boqun
> > > 
> > > 
> 
> 

[toc] | [prev] | [next] | [standalone]


#1698066

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 17:00 +0200
Message-ID<u7SqB-7mN-13@gated-at.bofh.it>
In reply to#1698049
On Thu, Jul 27, 2017 at 10:47:03PM +0800, Boqun Feng wrote:
> On Thu, Jul 27, 2017 at 07:36:58AM -0700, Paul E. McKenney wrote:
> > > > 
> > > > The reporting of the quiescent state will acquire the leaf rcu_node
> > > > structure's lock, with an smp_mb__after_unlock_lock(), which will
> > > > one way or another be a full memory barrier.  So the reorderings
> > > > cannot happen.
> > > > 
> > > > Unless I am missing something subtle.  ;-)
> > > > 
> > > 
> > > Well, smp_mb__after_unlock_lock() in ARM64 is a no-op, and ARM64's lock
> > > doesn't provide a smp_mb().
> > > 
> > > So my point is more like: synchronize_sched() happens to be a
> > > sys_membarrier() because of some implementation detail, and if some day
> > > we come up with a much cheaper way to implement sched flavor
> > > RCU(hopefully!), synchronize_sched() may be not good for the job. So at
> > > least, we'd better document this somewhere?
> > 
> > Last I heard, ARM's unlock/lock acted as a full barrier.  Will?
> > 
> > Please see the synchronize_sched() comment header for the documentation
> > you are asking for.  And the "Memory-Barrier Guarantees" section of
> > Documentation/RCU/Design/Requirements/Requirements.html.
> > 
> 
> All those barrier guarantees are subject to a RCU read-side critical
> section with a synchonize_*(), IIRC, for example:
> 
>  * On systems with more than one CPU, when synchronize_sched() returns,
>  * each CPU is guaranteed to have executed a full memory barrier since the
>  * end of its last RCU-sched read-side critical section whose beginning
>  * preceded the call to synchronize_sched().  In addition, each CPU having
> 
> , which is not the case for a quiesent state without a read-side
> critical section(i.e. non-context-switch quiesent state for sched Flavor)
> 
> I've read those requirements and could not find one to explain why there
> will be a full barrier emitted in an interrupted user-space program.

What you are forgetting is that for synchronize_sched(), any region of
code with preemption disabled is an RCU-sched read-side critical section.

							Thanx, Paul

> Regards,
> Boqun
> 
> > 							Thanx, Paul
> > 
> > > Regards,
> > > Boqun
> > > 
> > > > 						Thanx, Paul
> > > > 
> > > > > 					<return to user space>                        |    |
> > > > > 					read Y; --------------------------------------+----+
> > > > > 	store X;			                                              |
> > > > > 					{read X}(reordered) <-------------------------+
> > > > > 
> > > > > I assume the timer interrupt handler, which interrupts a user space and
> > > > > reports a quiesent state for sched flavor RCU, may not have a smp_mb()
> > > > > in some code path.
> > > > > 
> > > > > I may miss something subtle, but it just not very obvious how
> > > > > synchronize_sched() will guarantee a remote CPU running in userspace to
> > > > > do a smp_mb() before it returns, this is at least not in RCU
> > > > > requirements, right?
> > > > > 
> > > > > Regards,
> > > > > Boqun
> > > > 
> > > > 
> > 
> > 

[toc] | [prev] | [next] | [standalone]


#1696926

FromWill Deacon <will.deacon@arm.com>
Date2017-07-26 11:40 +0200
Message-ID<u7qXo-6Yw-13@gated-at.bofh.it>
In reply to#1696502
On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
> Some architectures are less precise than others in tracking which
> CPUs are running a given process due to ASIDs, though this is
> thought to be a non-problem:
> 
> 	https://marc.info/?l=linux-arch&m=126716090413065&w=2
> 	https://marc.info/?l=linux-arch&m=126716262815202&w=2
> 
> Thoughts?

On arm64, we *never* touch mm_cpumask, so it will always be empty. The only
thing we could potentially use it for is non-broadcast TLB invalidation if
the mm was only active on a single CPU, but when I implemented that we
pretty much never hit that case. I was hoping fork()+exec() would trigger it
(e.g. scripts), but the scheduler treats both of those as rebalancing points
so the temporary ASID before the exec always has two CPUs set in its mask.

Will

[toc] | [prev] | [next] | [standalone]


#1697316

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-26 17:50 +0200
Message-ID<u7wJr-2ak-5@gated-at.bofh.it>
In reply to#1696926
On Wed, Jul 26, 2017 at 10:36:46AM +0100, Will Deacon wrote:
> On Tue, Jul 25, 2017 at 02:19:26PM -0700, Paul E. McKenney wrote:
> > Some architectures are less precise than others in tracking which
> > CPUs are running a given process due to ASIDs, though this is
> > thought to be a non-problem:
> > 
> > 	https://marc.info/?l=linux-arch&m=126716090413065&w=2
> > 	https://marc.info/?l=linux-arch&m=126716262815202&w=2
> > 
> > Thoughts?
> 
> On arm64, we *never* touch mm_cpumask, so it will always be empty. The only
> thing we could potentially use it for is non-broadcast TLB invalidation if
> the mm was only active on a single CPU, but when I implemented that we
> pretty much never hit that case. I was hoping fork()+exec() would trigger it
> (e.g. scripts), but the scheduler treats both of those as rebalancing points
> so the temporary ASID before the exec always has two CPUs set in its mask.

OK, so it sounds like we should ignore ->mm_cpumask except possibly
for an architecture-specific optimization, and rely solely on ->mm
in the absence of such optimizations, right?

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1697871

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-27 12:20 +0200
Message-ID<u7O3E-4Qj-21@gated-at.bofh.it>
In reply to#1696131
On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> This horse is already out, so trying to shut the gate won't be effective.

So I'm not convinced it is. The mprotect() hack isn't portable as we've
established and on x86 where it does work, it doesn't (much) perturb
tasks not related to our process because we keep a tight mm_cpumask().

And if there are other (unpriv.) means of spraying IPIs around, we
should most certainly look at fixing those, not just shrug and make
matters worse.

[toc] | [prev] | [next] | [standalone]


#1697965

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 15:00 +0200
Message-ID<u7Qyv-6dA-37@gated-at.bofh.it>
In reply to#1697871
On Thu, Jul 27, 2017 at 12:14:26PM +0200, Peter Zijlstra wrote:
> On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> > This horse is already out, so trying to shut the gate won't be effective.
> 
> So I'm not convinced it is. The mprotect() hack isn't portable as we've
> established and on x86 where it does work, it doesn't (much) perturb
> tasks not related to our process because we keep a tight mm_cpumask().

Wrong.  People are using it today, portable or not.  If we want them
to stop using it, we need to give them an alternative.  Period.

> And if there are other (unpriv.) means of spraying IPIs around, we
> should most certainly look at fixing those, not just shrug and make
> matters worse.

We need to keep an open mind for a bit.  If there was a trivial
solution, we would have implemented it back in 2010.

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1697993

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-27 15:40 +0200
Message-ID<u7Rbc-6Gi-19@gated-at.bofh.it>
In reply to#1697965
On Thu, Jul 27, 2017 at 05:56:59AM -0700, Paul E. McKenney wrote:
> On Thu, Jul 27, 2017 at 12:14:26PM +0200, Peter Zijlstra wrote:
> > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> > > This horse is already out, so trying to shut the gate won't be effective.
> > 
> > So I'm not convinced it is. The mprotect() hack isn't portable as we've
> > established and on x86 where it does work, it doesn't (much) perturb
> > tasks not related to our process because we keep a tight mm_cpumask().
> 
> Wrong.  People are using it today, portable or not.  If we want them
> to stop using it, we need to give them an alternative.  Period.

What's wrong? The mprotect() hack isn't portable, nor does it perturb
other users much.

I would much rather they use this than your
synchronize_sched_expedited() thing. Its much better behaved. And
they'll run into pain the moment they start using ARM,PPC,S390,etc..

[toc] | [prev] | [next] | [standalone]


#1698025

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-07-27 16:40 +0200
Message-ID<u7S7f-7g8-11@gated-at.bofh.it>
In reply to#1697993
On Thu, Jul 27, 2017 at 03:37:29PM +0200, Peter Zijlstra wrote:
> On Thu, Jul 27, 2017 at 05:56:59AM -0700, Paul E. McKenney wrote:
> > On Thu, Jul 27, 2017 at 12:14:26PM +0200, Peter Zijlstra wrote:
> > > On Tue, Jul 25, 2017 at 12:36:12PM -0700, Paul E. McKenney wrote:
> > > > This horse is already out, so trying to shut the gate won't be effective.
> > > 
> > > So I'm not convinced it is. The mprotect() hack isn't portable as we've
> > > established and on x86 where it does work, it doesn't (much) perturb
> > > tasks not related to our process because we keep a tight mm_cpumask().
> > 
> > Wrong.  People are using it today, portable or not.  If we want them
> > to stop using it, we need to give them an alternative.  Period.
> 
> What's wrong? The mprotect() hack isn't portable, nor does it perturb
> other users much.
> 
> I would much rather they use this than your
> synchronize_sched_expedited() thing. Its much better behaved. And
> they'll run into pain the moment they start using ARM,PPC,S390,etc..

What is wrong is that we currently don't provide them a reasonable
alternative.

							Thanx, Paul

[toc] | [prev] | [standalone]


Page 3 of 3 — ← Prev page 1 2 [3]

Back to top | Article view | linux.kernel


csiph-web