Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1324594 > unrolled thread
| Started by | Michal Hocko <mhocko@kernel.org> |
|---|---|
| First post | 2016-02-02 21:30 +0100 |
| Last post | 2016-02-02 21:30 +0100 |
| Articles | 11 — 2 participants |
Back to article view | Back to linux.kernel
[RFC 0/12] introduce down_write_killable for rw_semaphore Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 03/12] locking, rwsem: introduce basis for down_write_killable Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 02/12] locking, rwsem: drop explicit memory barriers Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 05/12] ia64, rwsem: provide __down_write_killable Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 10/12] x86, rwsem: simplify __down_write Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
Re: [RFC 10/12] x86, rwsem: simplify __down_write Ingo Molnar <mingo@kernel.org> - 2016-02-03 09:20 +0100
Re: [RFC 10/12] x86, rwsem: simplify __down_write Michal Hocko <mhocko@kernel.org> - 2016-02-03 13:20 +0100
[RFC 06/12] s390, rwsem: provide __down_write_killable Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 04/12] alpha, rwsem: provide __down_write_killable Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 01/12] locking, rwsem: get rid of __down_write_nested Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
[RFC 09/12] xtensa, rwsem: provide __down_write_killable Michal Hocko <mhocko@kernel.org> - 2016-02-02 21:30 +0100
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 0/12] introduce down_write_killable for rw_semaphore |
| Message-ID | <qXPKi-8sh-3@gated-at.bofh.it> |
Hi, the following patchset implements a killable variant of write lock for rw_semaphore. My usecase is to turn as many mmap_sem write users to use a killable variant which will be helpful for the oom_reaper [1] to asynchronously tear down the oom victim address space which requires mmap_sem for read. This will reduce a likelihood of OOM livelocks caused by oom victim being stuck on a lock or other resource which prevents it to reach its exit path and release the memory. I haven't implemented the killable variant of the read lock because I do not have any usecase for this API. The patchset is organized as follows. - Patch 1 is a trivial cleanup - Patch 2, I belive, shouldn't introduce any functional changes as per Documentation/memory-barriers.txt. - Patch 3 is the preparatory work and necessary infrastructure for down_write_killable. It implements generic __down_write_killable and prepares the write lock slow path to bail out earlier when told so - Patch 4-9 are implementing arch specific __down_write_killable. One patch per architecture. I haven't even tried to compile test anything but sparch which uses CONFIG_RWSEM_GENERIC_SPINLOCK in allnoconfig. Those shold be mostly trivial. - One exception is x86 which replaces the current implementation of __down_write with the generic one to make easier to read and get rid of one level of indirection to the slow path. More on that in patch 10. I do not have any problems to drop patch 10 and rework 11 to the current inline asm but I think the easier code would be better. - finally patch 11 implements down_write_killable and ties everything together. I am not really an expert on lockdep so I hope I got it right. Many of arch specific patches are basically same and I can squash them into one patch if this is preferred but I thought that one patch per arch is preferable. My patch to change mmap_sem write users to killable form is not part of the series because it is not finished yet but I guess it is not really necessary for the RFC. The API is used in the same way as mutex_lock_killable. I have tested on x86 with OOM situations with high mmap_sem contention (basically many parallel page faults racing with many parallel mmap/munmap tight loops) so the waiters for the write locks are routinely interrupted by SIGKILL. Patches should apply cleanly on both Linus and next tree. Any feedback is highly appreciated. --- [1] http://lkml.kernel.org/r/1452094975-551-1-git-send-email-mhocko@kernel.org
[toc] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 03/12] locking, rwsem: introduce basis for down_write_killable |
| Message-ID | <qXPKj-8sh-19@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
Introduce a generic implementation necessary for down_write_killable.
This is a trivial extension of the already existing down_write call
which can be interrupted by SIGKILL. This patch doesn't provide
down_write_killable yet because arches have to provide the necessary
pieces before.
rwsem_down_write_failed which is a generic slow path for the
write lock is extended to allow a task state and renamed to
__rwsem_down_write_failed_state. The return value is either a valid
semaphore pointer or ERR_PTR(-EINTR).
rwsem_down_write_failed_killable is exported as a new way to wait for
the lock and be killable.
For rwsem-spinlock implementation the current __down_write it updated
in a similar way as __rwsem_down_write_failed_state except it doesn't
need new exports just visible __down_write_killable.
Architectures which are not using the generic rwsem implementation are
supposed to provide their __down_write_killable implementation and
use rwsem_down_write_failed_killable for the slow path.
Signed-off-by: Michal Hocko <mhocko@suse.com>
fold me "locking, rwsem: introduce basis for down_write_killable"
---
include/asm-generic/rwsem.h | 13 +++++++++++++
include/linux/rwsem-spinlock.h | 1 +
include/linux/rwsem.h | 2 ++
kernel/locking/rwsem-spinlock.c | 23 +++++++++++++++++++++--
kernel/locking/rwsem-xadd.c | 31 +++++++++++++++++++++++++------
5 files changed, 62 insertions(+), 8 deletions(-)
diff --git a/include/asm-generic/rwsem.h b/include/asm-generic/rwsem.h
index b8d8a6cf4ca8..b3f3aebbb994 100644
--- a/include/asm-generic/rwsem.h
+++ b/include/asm-generic/rwsem.h
@@ -63,6 +63,19 @@ static inline void __down_write(struct rw_semaphore *sem)
rwsem_down_write_failed(sem);
}
+static inline int __down_write_killable(struct rw_semaphore *sem)
+{
+ long tmp;
+ int ret;
+
+ tmp = atomic_long_add_return_acquire(RWSEM_ACTIVE_WRITE_BIAS,
+ (atomic_long_t *)&sem->count);
+ if (unlikely(tmp != RWSEM_ACTIVE_WRITE_BIAS))
+ if (IS_ERR(rwsem_down_write_failed_killable(sem)))
+ return -EINTR;
+ return 0;
+}
+
static inline int __down_write_trylock(struct rw_semaphore *sem)
{
long tmp;
diff --git a/include/linux/rwsem-spinlock.h b/include/linux/rwsem-spinlock.h
index a733a5467e6c..ae0528b834cd 100644
--- a/include/linux/rwsem-spinlock.h
+++ b/include/linux/rwsem-spinlock.h
@@ -34,6 +34,7 @@ struct rw_semaphore {
extern void __down_read(struct rw_semaphore *sem);
extern int __down_read_trylock(struct rw_semaphore *sem);
extern void __down_write(struct rw_semaphore *sem);
+extern int __must_check __down_write_killable(struct rw_semaphore *sem);
extern int __down_write_trylock(struct rw_semaphore *sem);
extern void __up_read(struct rw_semaphore *sem);
extern void __up_write(struct rw_semaphore *sem);
diff --git a/include/linux/rwsem.h b/include/linux/rwsem.h
index 8f498cdde280..7d7ae029dac5 100644
--- a/include/linux/rwsem.h
+++ b/include/linux/rwsem.h
@@ -14,6 +14,7 @@
#include <linux/list.h>
#include <linux/spinlock.h>
#include <linux/atomic.h>
+#include <linux/err.h>
#ifdef CONFIG_RWSEM_SPIN_ON_OWNER
#include <linux/osq_lock.h>
#endif
@@ -43,6 +44,7 @@ struct rw_semaphore {
extern struct rw_semaphore *rwsem_down_read_failed(struct rw_semaphore *sem);
extern struct rw_semaphore *rwsem_down_write_failed(struct rw_semaphore *sem);
+extern struct rw_semaphore *rwsem_down_write_failed_killable(struct rw_semaphore *sem);
extern struct rw_semaphore *rwsem_wake(struct rw_semaphore *);
extern struct rw_semaphore *rwsem_downgrade_wake(struct rw_semaphore *sem);
diff --git a/kernel/locking/rwsem-spinlock.c b/kernel/locking/rwsem-spinlock.c
index bab26104a5d0..d1d04ca10d0e 100644
--- a/kernel/locking/rwsem-spinlock.c
+++ b/kernel/locking/rwsem-spinlock.c
@@ -191,11 +191,12 @@ int __down_read_trylock(struct rw_semaphore *sem)
/*
* get a write lock on the semaphore
*/
-void __sched __down_write(struct rw_semaphore *sem)
+int __sched __down_write_state(struct rw_semaphore *sem, int state)
{
struct rwsem_waiter waiter;
struct task_struct *tsk;
unsigned long flags;
+ int ret = 0;
raw_spin_lock_irqsave(&sem->wait_lock, flags);
@@ -215,16 +216,34 @@ void __sched __down_write(struct rw_semaphore *sem)
*/
if (sem->count == 0)
break;
- set_task_state(tsk, TASK_UNINTERRUPTIBLE);
+ set_task_state(tsk, state);
raw_spin_unlock_irqrestore(&sem->wait_lock, flags);
schedule();
+ if (signal_pending_state(state, current)) {
+ ret = -EINTR;
+ raw_spin_lock_irqsave(&sem->wait_lock, flags);
+ goto out;
+ }
raw_spin_lock_irqsave(&sem->wait_lock, flags);
}
/* got the lock */
sem->count = -1;
+out:
list_del(&waiter.list);
raw_spin_unlock_irqrestore(&sem->wait_lock, flags);
+
+ return ret;
+}
+
+void __sched __down_write(struct rw_semaphore *sem)
+{
+ __down_write_state(sem, TASK_UNINTERRUPTIBLE);
+}
+
+int __sched __down_write_killable(struct rw_semaphore *sem)
+{
+ return __down_write_state(sem, TASK_KILLABLE);
}
/*
diff --git a/kernel/locking/rwsem-xadd.c b/kernel/locking/rwsem-xadd.c
index a4d4de05b2d1..5cec34f1ad6f 100644
--- a/kernel/locking/rwsem-xadd.c
+++ b/kernel/locking/rwsem-xadd.c
@@ -433,12 +433,13 @@ static inline bool rwsem_has_spinner(struct rw_semaphore *sem)
/*
* Wait until we successfully acquire the write lock
*/
-__visible
-struct rw_semaphore __sched *rwsem_down_write_failed(struct rw_semaphore *sem)
+static inline struct rw_semaphore *
+__rwsem_down_write_failed_state(struct rw_semaphore *sem, int state)
{
long count;
bool waiting = true; /* any queued threads before us */
struct rwsem_waiter waiter;
+ struct rw_semaphore *ret = sem;
/* undo write bias from down_write operation, stop active locking */
count = rwsem_atomic_update(-RWSEM_ACTIVE_WRITE_BIAS, sem);
@@ -478,7 +479,7 @@ struct rw_semaphore __sched *rwsem_down_write_failed(struct rw_semaphore *sem)
count = rwsem_atomic_update(RWSEM_WAITING_BIAS, sem);
/* wait until we successfully acquire the lock */
- set_current_state(TASK_UNINTERRUPTIBLE);
+ set_current_state(state);
while (true) {
if (rwsem_try_write_lock(count, sem))
break;
@@ -487,20 +488,38 @@ struct rw_semaphore __sched *rwsem_down_write_failed(struct rw_semaphore *sem)
/* Block until there are no active lockers. */
do {
schedule();
- set_current_state(TASK_UNINTERRUPTIBLE);
+ if (signal_pending_state(state, current)) {
+ raw_spin_lock_irq(&sem->wait_lock);
+ ret = ERR_PTR(-EINTR);
+ goto out;
+ }
+ set_current_state(state);
} while ((count = sem->count) & RWSEM_ACTIVE_MASK);
raw_spin_lock_irq(&sem->wait_lock);
}
__set_current_state(TASK_RUNNING);
-
+out:
list_del(&waiter.list);
raw_spin_unlock_irq(&sem->wait_lock);
- return sem;
+ return ret;
+}
+
+__visible struct rw_semaphore * __sched
+rwsem_down_write_failed(struct rw_semaphore *sem)
+{
+ return __rwsem_down_write_failed_state(sem, TASK_UNINTERRUPTIBLE);
}
EXPORT_SYMBOL(rwsem_down_write_failed);
+__visible struct rw_semaphore * __sched
+rwsem_down_write_failed_killable(struct rw_semaphore *sem)
+{
+ return __rwsem_down_write_failed_state(sem, TASK_KILLABLE);
+}
+EXPORT_SYMBOL(rwsem_down_write_failed_killable);
+
/*
* handle waking up a waiter on the semaphore
* - up_read/up_write has decremented the active part of count if we come here
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 02/12] locking, rwsem: drop explicit memory barriers |
| Message-ID | <qXPKj-8sh-21@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
sh and xtensa seem to be the only architectures which use explicit
memory barriers for rw_semaphore operations even though they are not
really needed because there is the full memory barrier is always implied
by atomic_{inc,dec,add,sub}_return resp. cmpxchg. Remove them.
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/sh/include/asm/rwsem.h | 14 ++------------
arch/xtensa/include/asm/rwsem.h | 14 ++------------
2 files changed, 4 insertions(+), 24 deletions(-)
diff --git a/arch/sh/include/asm/rwsem.h b/arch/sh/include/asm/rwsem.h
index a5104bebd1eb..f6c951c7a875 100644
--- a/arch/sh/include/asm/rwsem.h
+++ b/arch/sh/include/asm/rwsem.h
@@ -24,9 +24,7 @@
*/
static inline void __down_read(struct rw_semaphore *sem)
{
- if (atomic_inc_return((atomic_t *)(&sem->count)) > 0)
- smp_wmb();
- else
+ if (atomic_inc_return((atomic_t *)(&sem->count)) <= 0)
rwsem_down_read_failed(sem);
}
@@ -37,7 +35,6 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
while ((tmp = sem->count) >= 0) {
if (tmp == cmpxchg(&sem->count, tmp,
tmp + RWSEM_ACTIVE_READ_BIAS)) {
- smp_wmb();
return 1;
}
}
@@ -53,9 +50,7 @@ static inline void __down_write(struct rw_semaphore *sem)
tmp = atomic_add_return(RWSEM_ACTIVE_WRITE_BIAS,
(atomic_t *)(&sem->count));
- if (tmp == RWSEM_ACTIVE_WRITE_BIAS)
- smp_wmb();
- else
+ if (tmp != RWSEM_ACTIVE_WRITE_BIAS)
rwsem_down_write_failed(sem);
}
@@ -65,7 +60,6 @@ static inline int __down_write_trylock(struct rw_semaphore *sem)
tmp = cmpxchg(&sem->count, RWSEM_UNLOCKED_VALUE,
RWSEM_ACTIVE_WRITE_BIAS);
- smp_wmb();
return tmp == RWSEM_UNLOCKED_VALUE;
}
@@ -76,7 +70,6 @@ static inline void __up_read(struct rw_semaphore *sem)
{
int tmp;
- smp_wmb();
tmp = atomic_dec_return((atomic_t *)(&sem->count));
if (tmp < -1 && (tmp & RWSEM_ACTIVE_MASK) == 0)
rwsem_wake(sem);
@@ -87,7 +80,6 @@ static inline void __up_read(struct rw_semaphore *sem)
*/
static inline void __up_write(struct rw_semaphore *sem)
{
- smp_wmb();
if (atomic_sub_return(RWSEM_ACTIVE_WRITE_BIAS,
(atomic_t *)(&sem->count)) < 0)
rwsem_wake(sem);
@@ -108,7 +100,6 @@ static inline void __downgrade_write(struct rw_semaphore *sem)
{
int tmp;
- smp_wmb();
tmp = atomic_add_return(-RWSEM_WAITING_BIAS, (atomic_t *)(&sem->count));
if (tmp < 0)
rwsem_downgrade_wake(sem);
@@ -119,7 +110,6 @@ static inline void __downgrade_write(struct rw_semaphore *sem)
*/
static inline int rwsem_atomic_update(int delta, struct rw_semaphore *sem)
{
- smp_mb();
return atomic_add_return(delta, (atomic_t *)(&sem->count));
}
diff --git a/arch/xtensa/include/asm/rwsem.h b/arch/xtensa/include/asm/rwsem.h
index 249619e7e7f2..593483f6e1ff 100644
--- a/arch/xtensa/include/asm/rwsem.h
+++ b/arch/xtensa/include/asm/rwsem.h
@@ -29,9 +29,7 @@
*/
static inline void __down_read(struct rw_semaphore *sem)
{
- if (atomic_add_return(1,(atomic_t *)(&sem->count)) > 0)
- smp_wmb();
- else
+ if (atomic_add_return(1,(atomic_t *)(&sem->count)) <= 0)
rwsem_down_read_failed(sem);
}
@@ -42,7 +40,6 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
while ((tmp = sem->count) >= 0) {
if (tmp == cmpxchg(&sem->count, tmp,
tmp + RWSEM_ACTIVE_READ_BIAS)) {
- smp_wmb();
return 1;
}
}
@@ -58,9 +55,7 @@ static inline void __down_write(struct rw_semaphore *sem)
tmp = atomic_add_return(RWSEM_ACTIVE_WRITE_BIAS,
(atomic_t *)(&sem->count));
- if (tmp == RWSEM_ACTIVE_WRITE_BIAS)
- smp_wmb();
- else
+ if (tmp != RWSEM_ACTIVE_WRITE_BIAS)
rwsem_down_write_failed(sem);
}
@@ -70,7 +65,6 @@ static inline int __down_write_trylock(struct rw_semaphore *sem)
tmp = cmpxchg(&sem->count, RWSEM_UNLOCKED_VALUE,
RWSEM_ACTIVE_WRITE_BIAS);
- smp_wmb();
return tmp == RWSEM_UNLOCKED_VALUE;
}
@@ -81,7 +75,6 @@ static inline void __up_read(struct rw_semaphore *sem)
{
int tmp;
- smp_wmb();
tmp = atomic_sub_return(1,(atomic_t *)(&sem->count));
if (tmp < -1 && (tmp & RWSEM_ACTIVE_MASK) == 0)
rwsem_wake(sem);
@@ -92,7 +85,6 @@ static inline void __up_read(struct rw_semaphore *sem)
*/
static inline void __up_write(struct rw_semaphore *sem)
{
- smp_wmb();
if (atomic_sub_return(RWSEM_ACTIVE_WRITE_BIAS,
(atomic_t *)(&sem->count)) < 0)
rwsem_wake(sem);
@@ -113,7 +105,6 @@ static inline void __downgrade_write(struct rw_semaphore *sem)
{
int tmp;
- smp_wmb();
tmp = atomic_add_return(-RWSEM_WAITING_BIAS, (atomic_t *)(&sem->count));
if (tmp < 0)
rwsem_downgrade_wake(sem);
@@ -124,7 +115,6 @@ static inline void __downgrade_write(struct rw_semaphore *sem)
*/
static inline int rwsem_atomic_update(int delta, struct rw_semaphore *sem)
{
- smp_mb();
return atomic_add_return(delta, (atomic_t *)(&sem->count));
}
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 05/12] ia64, rwsem: provide __down_write_killable |
| Message-ID | <qXPKj-8sh-31@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
Introduce ___down_write for the fast path and reuse it for __down_write
resp. __down_write_killable each using the respective generic slow path
(rwsem_down_write_failed resp. rwsem_down_write_failed_killable).
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/ia64/include/asm/rwsem.h | 22 +++++++++++++++++++---
1 file changed, 19 insertions(+), 3 deletions(-)
diff --git a/arch/ia64/include/asm/rwsem.h b/arch/ia64/include/asm/rwsem.h
index 3027e7516d85..5e78cb40d9df 100644
--- a/arch/ia64/include/asm/rwsem.h
+++ b/arch/ia64/include/asm/rwsem.h
@@ -49,8 +49,8 @@ __down_read (struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void
-__down_write (struct rw_semaphore *sem)
+static inline long
+___down_write (struct rw_semaphore *sem)
{
long old, new;
@@ -59,10 +59,26 @@ __down_write (struct rw_semaphore *sem)
new = old + RWSEM_ACTIVE_WRITE_BIAS;
} while (cmpxchg_acq(&sem->count, old, new) != old);
- if (old != 0)
+ return old;
+}
+
+static inline void
+__down_write (struct rw_semaphore *sem)
+{
+ if (___down_write(sem))
rwsem_down_write_failed(sem);
}
+static inline int
+__down_write_killable (struct rw_semaphore *sem)
+{
+ if (___down_write(sem))
+ if (IS_ERR(rwsem_down_write_failed_killable(sem)))
+ return -EINTR;
+
+ return 0;
+}
+
/*
* unlock after reading
*/
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 10/12] x86, rwsem: simplify __down_write |
| Message-ID | <qXPKj-8sh-35@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
x86 implementation of __down_write is using inline asm to optimize the
code flow. This however requires that it has go over an additional hop
for the slow path call_rwsem_down_write_failed which has to
save_common_regs/restore_common_regs to preserve the calling convention.
This, however doesn't add much because the fast path only saves one
register push/pop (rdx) when compared to the generic implementation:
Before:
0000000000000019 <down_write>:
19: e8 00 00 00 00 callq 1e <down_write+0x5>
1e: 55 push %rbp
1f: 48 ba 01 00 00 00 ff movabs $0xffffffff00000001,%rdx
26: ff ff ff
29: 48 89 f8 mov %rdi,%rax
2c: 48 89 e5 mov %rsp,%rbp
2f: f0 48 0f c1 10 lock xadd %rdx,(%rax)
34: 85 d2 test %edx,%edx
36: 74 05 je 3d <down_write+0x24>
38: e8 00 00 00 00 callq 3d <down_write+0x24>
3d: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax
44: 00 00
46: 5d pop %rbp
47: 48 89 47 38 mov %rax,0x38(%rdi)
4b: c3 retq
After:
0000000000000019 <down_write>:
19: e8 00 00 00 00 callq 1e <down_write+0x5>
1e: 55 push %rbp
1f: 48 b8 01 00 00 00 ff movabs $0xffffffff00000001,%rax
26: ff ff ff
29: 48 89 e5 mov %rsp,%rbp
2c: 53 push %rbx
2d: 48 89 fb mov %rdi,%rbx
30: f0 48 0f c1 07 lock xadd %rax,(%rdi)
35: 48 85 c0 test %rax,%rax
38: 74 05 je 3f <down_write+0x26>
3a: e8 00 00 00 00 callq 3f <down_write+0x26>
3f: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax
46: 00 00
48: 48 89 43 38 mov %rax,0x38(%rbx)
4c: 5b pop %rbx
4d: 5d pop %rbp
4e: c3 retq
This doesn't seem to justify the code obfuscation and complexity. Use
the generic implementation instead.
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/x86/include/asm/rwsem.h | 17 +++++------------
arch/x86/lib/rwsem.S | 9 ---------
2 files changed, 5 insertions(+), 21 deletions(-)
diff --git a/arch/x86/include/asm/rwsem.h b/arch/x86/include/asm/rwsem.h
index d79a218675bc..1b5e89b3643d 100644
--- a/arch/x86/include/asm/rwsem.h
+++ b/arch/x86/include/asm/rwsem.h
@@ -102,18 +102,11 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
static inline void __down_write(struct rw_semaphore *sem)
{
long tmp;
- asm volatile("# beginning down_write\n\t"
- LOCK_PREFIX " xadd %1,(%2)\n\t"
- /* adds 0xffff0001, returns the old value */
- " test " __ASM_SEL(%w1,%k1) "," __ASM_SEL(%w1,%k1) "\n\t"
- /* was the active mask 0 before? */
- " jz 1f\n"
- " call call_rwsem_down_write_failed\n"
- "1:\n"
- "# ending down_write"
- : "+m" (sem->count), "=d" (tmp)
- : "a" (sem), "1" (RWSEM_ACTIVE_WRITE_BIAS)
- : "memory", "cc");
+
+ tmp = atomic_long_add_return(RWSEM_ACTIVE_WRITE_BIAS,
+ (atomic_long_t *)&sem->count);
+ if (unlikely(tmp != RWSEM_ACTIVE_WRITE_BIAS))
+ rwsem_down_write_failed(sem);
}
/*
diff --git a/arch/x86/lib/rwsem.S b/arch/x86/lib/rwsem.S
index 40027db99140..ea5c7c177483 100644
--- a/arch/x86/lib/rwsem.S
+++ b/arch/x86/lib/rwsem.S
@@ -57,7 +57,6 @@
* is also the input argument to these helpers)
*
* The following can clobber %rdx because the asm clobbers it:
- * call_rwsem_down_write_failed
* call_rwsem_wake
* but %rdi, %rsi, %rcx, %r8-r11 always need saving.
*/
@@ -93,14 +92,6 @@ ENTRY(call_rwsem_down_read_failed)
ret
ENDPROC(call_rwsem_down_read_failed)
-ENTRY(call_rwsem_down_write_failed)
- save_common_regs
- movq %rax,%rdi
- call rwsem_down_write_failed
- restore_common_regs
- ret
-ENDPROC(call_rwsem_down_write_failed)
-
ENTRY(call_rwsem_wake)
/* do nothing if still outstanding active readers */
__ASM_HALF_SIZE(dec) %__ASM_HALF_REG(dx)
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2016-02-03 09:20 +0100 |
| Subject | Re: [RFC 10/12] x86, rwsem: simplify __down_write |
| Message-ID | <qY0Po-7NH-11@gated-at.bofh.it> |
| In reply to | #1324600 |
* Michal Hocko <mhocko@kernel.org> wrote: > From: Michal Hocko <mhocko@suse.com> > > x86 implementation of __down_write is using inline asm to optimize the > code flow. This however requires that it has go over an additional hop > for the slow path call_rwsem_down_write_failed which has to > save_common_regs/restore_common_regs to preserve the calling convention. > This, however doesn't add much because the fast path only saves one > register push/pop (rdx) when compared to the generic implementation: > > Before: > 0000000000000019 <down_write>: > 19: e8 00 00 00 00 callq 1e <down_write+0x5> > 1e: 55 push %rbp > 1f: 48 ba 01 00 00 00 ff movabs $0xffffffff00000001,%rdx > 26: ff ff ff > 29: 48 89 f8 mov %rdi,%rax > 2c: 48 89 e5 mov %rsp,%rbp > 2f: f0 48 0f c1 10 lock xadd %rdx,(%rax) > 34: 85 d2 test %edx,%edx > 36: 74 05 je 3d <down_write+0x24> > 38: e8 00 00 00 00 callq 3d <down_write+0x24> > 3d: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax > 44: 00 00 > 46: 5d pop %rbp > 47: 48 89 47 38 mov %rax,0x38(%rdi) > 4b: c3 retq > > After: > 0000000000000019 <down_write>: > 19: e8 00 00 00 00 callq 1e <down_write+0x5> > 1e: 55 push %rbp > 1f: 48 b8 01 00 00 00 ff movabs $0xffffffff00000001,%rax > 26: ff ff ff > 29: 48 89 e5 mov %rsp,%rbp > 2c: 53 push %rbx > 2d: 48 89 fb mov %rdi,%rbx > 30: f0 48 0f c1 07 lock xadd %rax,(%rdi) > 35: 48 85 c0 test %rax,%rax > 38: 74 05 je 3f <down_write+0x26> > 3a: e8 00 00 00 00 callq 3f <down_write+0x26> > 3f: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax > 46: 00 00 > 48: 48 89 43 38 mov %rax,0x38(%rbx) > 4c: 5b pop %rbx > 4d: 5d pop %rbp > 4e: c3 retq I'm not convinced about the removal of this optimization at all. > This doesn't seem to justify the code obfuscation and complexity. Use > the generic implementation instead. > > Signed-off-by: Michal Hocko <mhocko@suse.com> > --- > arch/x86/include/asm/rwsem.h | 17 +++++------------ > arch/x86/lib/rwsem.S | 9 --------- > 2 files changed, 5 insertions(+), 21 deletions(-) Turn the argument around, would we be willing to save two instructions off the fast path of a commonly used locking construct, with such a simple optimization: > arch/x86/include/asm/rwsem.h | 17 ++++++++++++----- > arch/x86/lib/rwsem.S | 9 +++++++++ > 2 files changed, 21 insertions(+), 5 deletions(-) ? Yes! So, if you want to remove the assembly code - can we achieve that without hurting the generated fast path, using the compiler? Thanks, Ingo
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-03 13:20 +0100 |
| Subject | Re: [RFC 10/12] x86, rwsem: simplify __down_write |
| Message-ID | <qY4zD-1Ok-1@gated-at.bofh.it> |
| In reply to | #1325024 |
On Wed 03-02-16 09:10:16, Ingo Molnar wrote: > > * Michal Hocko <mhocko@kernel.org> wrote: > > > From: Michal Hocko <mhocko@suse.com> > > > > x86 implementation of __down_write is using inline asm to optimize the > > code flow. This however requires that it has go over an additional hop > > for the slow path call_rwsem_down_write_failed which has to > > save_common_regs/restore_common_regs to preserve the calling convention. > > This, however doesn't add much because the fast path only saves one > > register push/pop (rdx) when compared to the generic implementation: > > > > Before: > > 0000000000000019 <down_write>: > > 19: e8 00 00 00 00 callq 1e <down_write+0x5> > > 1e: 55 push %rbp > > 1f: 48 ba 01 00 00 00 ff movabs $0xffffffff00000001,%rdx > > 26: ff ff ff > > 29: 48 89 f8 mov %rdi,%rax > > 2c: 48 89 e5 mov %rsp,%rbp > > 2f: f0 48 0f c1 10 lock xadd %rdx,(%rax) > > 34: 85 d2 test %edx,%edx > > 36: 74 05 je 3d <down_write+0x24> > > 38: e8 00 00 00 00 callq 3d <down_write+0x24> > > 3d: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax > > 44: 00 00 > > 46: 5d pop %rbp > > 47: 48 89 47 38 mov %rax,0x38(%rdi) > > 4b: c3 retq > > > > After: > > 0000000000000019 <down_write>: > > 19: e8 00 00 00 00 callq 1e <down_write+0x5> > > 1e: 55 push %rbp > > 1f: 48 b8 01 00 00 00 ff movabs $0xffffffff00000001,%rax > > 26: ff ff ff > > 29: 48 89 e5 mov %rsp,%rbp > > 2c: 53 push %rbx > > 2d: 48 89 fb mov %rdi,%rbx > > 30: f0 48 0f c1 07 lock xadd %rax,(%rdi) > > 35: 48 85 c0 test %rax,%rax > > 38: 74 05 je 3f <down_write+0x26> > > 3a: e8 00 00 00 00 callq 3f <down_write+0x26> > > 3f: 65 48 8b 04 25 00 00 mov %gs:0x0,%rax > > 46: 00 00 > > 48: 48 89 43 38 mov %rax,0x38(%rbx) > > 4c: 5b pop %rbx > > 4d: 5d pop %rbp > > 4e: c3 retq > > I'm not convinced about the removal of this optimization at all. OK, fair enough. As I've mentioned in the cover letter I do not really insist on this patch. I just found the current code too ugly to live without a good reason because down_write is a call so saving one push/pop seems like really negligible to the call itself. Moreover this is a write lock which is expected to be heavier. It is the read path which is expected to be light and contention (slow path) is expected on the write lock. That being said, if you really believe that the current code is easier to maintain then I will not pursue this patch. The rest doesn't really depend on it. I will just respin the follow up x86 specifi __down_write_killable to follow the same code convention. [...] > So, if you want to remove the assembly code - can we achieve that without hurting > the generated fast path, using the compiler? One way would be to do the same thing as mutex does and do the fast path as an inline. This could bloat the kernel and require some additional changes to allow arch specific reimplementations though so I didn't want to go that path. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 06/12] s390, rwsem: provide __down_write_killable |
| Message-ID | <qXPKj-8sh-37@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
Introduce ___down_write for the fast path and reuse it for __down_write
resp. __down_write_killable each using the respective generic slow path
(rwsem_down_write_failed resp. rwsem_down_write_failed_killable).
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/s390/include/asm/rwsem.h | 19 +++++++++++++++++--
1 file changed, 17 insertions(+), 2 deletions(-)
diff --git a/arch/s390/include/asm/rwsem.h b/arch/s390/include/asm/rwsem.h
index 8e52e72f3efc..8e7b2b7e10f3 100644
--- a/arch/s390/include/asm/rwsem.h
+++ b/arch/s390/include/asm/rwsem.h
@@ -90,7 +90,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void __down_write(struct rw_semaphore *sem)
+static inline long ___down_write(struct rw_semaphore *sem)
{
signed long old, new, tmp;
@@ -104,10 +104,25 @@ static inline void __down_write(struct rw_semaphore *sem)
: "=&d" (old), "=&d" (new), "=Q" (sem->count)
: "Q" (sem->count), "m" (tmp)
: "cc", "memory");
- if (old != 0)
+
+ return old;
+}
+
+static inline void __down_write(struct rw_semaphore *sem)
+{
+ if (___down_write(sem))
rwsem_down_write_failed(sem);
}
+static inline int __down_write_killable(struct rw_semaphore *sem)
+{
+ if (___down_write(sem))
+ if (IS_ERR(rwsem_down_write_failed_killable(sem)))
+ return -EINTR;
+
+ return 0;
+}
+
/*
* trylock for writing -- returns 1 if successful, 0 if contention
*/
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 04/12] alpha, rwsem: provide __down_write_killable |
| Message-ID | <qXPKk-8sh-39@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
Introduce ___down_write for the fast path and reuse it for __down_write
resp. __down_write_killable each using the respective generic slow path
(rwsem_down_write_failed resp. rwsem_down_write_failed_killable).
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/alpha/include/asm/rwsem.h | 18 ++++++++++++++++--
1 file changed, 16 insertions(+), 2 deletions(-)
diff --git a/arch/alpha/include/asm/rwsem.h b/arch/alpha/include/asm/rwsem.h
index a83bbea62c67..0131a7058778 100644
--- a/arch/alpha/include/asm/rwsem.h
+++ b/arch/alpha/include/asm/rwsem.h
@@ -63,7 +63,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
return res >= 0 ? 1 : 0;
}
-static inline void __down_write(struct rw_semaphore *sem)
+static inline long ___down_write(struct rw_semaphore *sem)
{
long oldcount;
#ifndef CONFIG_SMP
@@ -83,10 +83,24 @@ static inline void __down_write(struct rw_semaphore *sem)
:"=&r" (oldcount), "=m" (sem->count), "=&r" (temp)
:"Ir" (RWSEM_ACTIVE_WRITE_BIAS), "m" (sem->count) : "memory");
#endif
- if (unlikely(oldcount))
+ return oldcount;
+}
+
+static inline void __down_write(struct rw_semaphore *sem)
+{
+ if (unlikely(___down_write(sem)))
rwsem_down_write_failed(sem);
}
+static inline int __down_write_killable(struct rw_semaphore *sem)
+{
+ if (unlikely(___down_write(sem)))
+ if (IS_ERR(rwsem_down_write_failed_killable(sem)))
+ return -EINTR;
+
+ return 0;
+}
+
/*
* trylock for writing -- returns 1 if successful, 0 if contention
*/
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 01/12] locking, rwsem: get rid of __down_write_nested |
| Message-ID | <qXPKk-8sh-43@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
This is no longer used anywhere and all callers (__down_write) use
0 as a subclass. Ditch __down_write_nested to make the code easier
to follow.
This shouldn't introduce any functional change.
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/s390/include/asm/rwsem.h | 7 +------
arch/sh/include/asm/rwsem.h | 5 -----
arch/sparc/include/asm/rwsem.h | 7 +------
arch/x86/include/asm/rwsem.h | 7 +------
include/asm-generic/rwsem.h | 7 +------
include/linux/rwsem-spinlock.h | 1 -
kernel/locking/rwsem-spinlock.c | 7 +------
7 files changed, 5 insertions(+), 36 deletions(-)
diff --git a/arch/s390/include/asm/rwsem.h b/arch/s390/include/asm/rwsem.h
index 4b43ee7e6776..8e52e72f3efc 100644
--- a/arch/s390/include/asm/rwsem.h
+++ b/arch/s390/include/asm/rwsem.h
@@ -90,7 +90,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
+static inline void __down_write(struct rw_semaphore *sem)
{
signed long old, new, tmp;
@@ -108,11 +108,6 @@ static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
rwsem_down_write_failed(sem);
}
-static inline void __down_write(struct rw_semaphore *sem)
-{
- __down_write_nested(sem, 0);
-}
-
/*
* trylock for writing -- returns 1 if successful, 0 if contention
*/
diff --git a/arch/sh/include/asm/rwsem.h b/arch/sh/include/asm/rwsem.h
index edab57265293..a5104bebd1eb 100644
--- a/arch/sh/include/asm/rwsem.h
+++ b/arch/sh/include/asm/rwsem.h
@@ -114,11 +114,6 @@ static inline void __downgrade_write(struct rw_semaphore *sem)
rwsem_downgrade_wake(sem);
}
-static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
-{
- __down_write(sem);
-}
-
/*
* implement exchange and add functionality
*/
diff --git a/arch/sparc/include/asm/rwsem.h b/arch/sparc/include/asm/rwsem.h
index 069bf4d663a1..e5a0d575bc7f 100644
--- a/arch/sparc/include/asm/rwsem.h
+++ b/arch/sparc/include/asm/rwsem.h
@@ -45,7 +45,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
+static inline void __down_write(struct rw_semaphore *sem)
{
long tmp;
@@ -55,11 +55,6 @@ static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
rwsem_down_write_failed(sem);
}
-static inline void __down_write(struct rw_semaphore *sem)
-{
- __down_write_nested(sem, 0);
-}
-
static inline int __down_write_trylock(struct rw_semaphore *sem)
{
long tmp;
diff --git a/arch/x86/include/asm/rwsem.h b/arch/x86/include/asm/rwsem.h
index cad82c9c2fde..d79a218675bc 100644
--- a/arch/x86/include/asm/rwsem.h
+++ b/arch/x86/include/asm/rwsem.h
@@ -99,7 +99,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
+static inline void __down_write(struct rw_semaphore *sem)
{
long tmp;
asm volatile("# beginning down_write\n\t"
@@ -116,11 +116,6 @@ static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
: "memory", "cc");
}
-static inline void __down_write(struct rw_semaphore *sem)
-{
- __down_write_nested(sem, 0);
-}
-
/*
* trylock for writing -- returns 1 if successful, 0 if contention
*/
diff --git a/include/asm-generic/rwsem.h b/include/asm-generic/rwsem.h
index d6d5dc98d7da..b8d8a6cf4ca8 100644
--- a/include/asm-generic/rwsem.h
+++ b/include/asm-generic/rwsem.h
@@ -53,7 +53,7 @@ static inline int __down_read_trylock(struct rw_semaphore *sem)
/*
* lock for writing
*/
-static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
+static inline void __down_write(struct rw_semaphore *sem)
{
long tmp;
@@ -63,11 +63,6 @@ static inline void __down_write_nested(struct rw_semaphore *sem, int subclass)
rwsem_down_write_failed(sem);
}
-static inline void __down_write(struct rw_semaphore *sem)
-{
- __down_write_nested(sem, 0);
-}
-
static inline int __down_write_trylock(struct rw_semaphore *sem)
{
long tmp;
diff --git a/include/linux/rwsem-spinlock.h b/include/linux/rwsem-spinlock.h
index 561e8615528d..a733a5467e6c 100644
--- a/include/linux/rwsem-spinlock.h
+++ b/include/linux/rwsem-spinlock.h
@@ -34,7 +34,6 @@ struct rw_semaphore {
extern void __down_read(struct rw_semaphore *sem);
extern int __down_read_trylock(struct rw_semaphore *sem);
extern void __down_write(struct rw_semaphore *sem);
-extern void __down_write_nested(struct rw_semaphore *sem, int subclass);
extern int __down_write_trylock(struct rw_semaphore *sem);
extern void __up_read(struct rw_semaphore *sem);
extern void __up_write(struct rw_semaphore *sem);
diff --git a/kernel/locking/rwsem-spinlock.c b/kernel/locking/rwsem-spinlock.c
index 3a5048572065..bab26104a5d0 100644
--- a/kernel/locking/rwsem-spinlock.c
+++ b/kernel/locking/rwsem-spinlock.c
@@ -191,7 +191,7 @@ int __down_read_trylock(struct rw_semaphore *sem)
/*
* get a write lock on the semaphore
*/
-void __sched __down_write_nested(struct rw_semaphore *sem, int subclass)
+void __sched __down_write(struct rw_semaphore *sem)
{
struct rwsem_waiter waiter;
struct task_struct *tsk;
@@ -227,11 +227,6 @@ void __sched __down_write_nested(struct rw_semaphore *sem, int subclass)
raw_spin_unlock_irqrestore(&sem->wait_lock, flags);
}
-void __sched __down_write(struct rw_semaphore *sem)
-{
- __down_write_nested(sem, 0);
-}
-
/*
* trylock for writing -- returns 1 if successful, 0 if contention
*/
--
2.7.0
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-02-02 21:30 +0100 |
| Subject | [RFC 09/12] xtensa, rwsem: provide __down_write_killable |
| Message-ID | <qXPKk-8sh-47@gated-at.bofh.it> |
| In reply to | #1324594 |
From: Michal Hocko <mhocko@suse.com>
which is uses the same fast path as __down_write except it falls back to
rwsem_down_write_failed_killable slow path and return -EINTR if killed.
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
arch/xtensa/include/asm/rwsem.h | 13 +++++++++++++
1 file changed, 13 insertions(+)
diff --git a/arch/xtensa/include/asm/rwsem.h b/arch/xtensa/include/asm/rwsem.h
index 593483f6e1ff..6283823b8040 100644
--- a/arch/xtensa/include/asm/rwsem.h
+++ b/arch/xtensa/include/asm/rwsem.h
@@ -59,6 +59,19 @@ static inline void __down_write(struct rw_semaphore *sem)
rwsem_down_write_failed(sem);
}
+static inline int __down_write_killable(struct rw_semaphore *sem)
+{
+ int tmp;
+
+ tmp = atomic_add_return(RWSEM_ACTIVE_WRITE_BIAS,
+ (atomic_t *)(&sem->count));
+ if (tmp != RWSEM_ACTIVE_WRITE_BIAS)
+ if (IS_ERR(rwsem_down_write_failed_killable(sem)))
+ return -EINTR;
+
+ return 0;
+}
+
static inline int __down_write_trylock(struct rw_semaphore *sem)
{
int tmp;
--
2.7.0
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web