Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1458548 > unrolled thread
| Started by | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| First post | 2016-08-09 11:40 +0200 |
| Last post | 2016-08-09 13:40 +0200 |
| Articles | 3 — 2 participants |
Back to article view | Back to linux.kernel
[PATCH RESEND v4] locking/pvqspinlock: Fix double hash race Wanpeng Li <kernellwp@gmail.com> - 2016-08-09 11:40 +0200
Re: [PATCH RESEND v4] locking/pvqspinlock: Fix double hash race Peter Zijlstra <peterz@infradead.org> - 2016-08-09 12:50 +0200
Re: [PATCH RESEND v4] locking/pvqspinlock: Fix double hash race Wanpeng Li <kernellwp@gmail.com> - 2016-08-09 13:40 +0200
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-08-09 11:40 +0200 |
| Subject | [PATCH RESEND v4] locking/pvqspinlock: Fix double hash race |
| Message-ID | <s4bFT-5HW-9@gated-at.bofh.it> |
From: Wanpeng Li <wanpeng.li@hotmail.com>
When the lock holder vCPU is racing with the queue head vCPU:
lock holder vCPU queue head vCPU
===================== ==================
node->locked = 1;
<preemption> READ_ONCE(node->locked)
... pv_wait_head_or_lock():
SPIN_THRESHOLD loop;
pv_hash();
lock->locked = _Q_SLOW_VAL;
node->state = vcpu_hashed;
pv_kick_node():
cmpxchg(node->state,
vcpu_halted, vcpu_hashed);
lock->locked = _Q_SLOW_VAL;
pv_hash();
With preemption at the right moment, it is possible that both the
lock holder and queue head vCPUs can be racing to set node->state
which can result in hash entry race. Making sure the state is never
set to vcpu_halted will prevent this racing from happening.
This patch fix it by setting vcpu_hashed after we did all hash thing.
Acked-by: Waiman Long <Waiman.Long@hpe.com>
Reviewed-by: Davidlohr Bueso <dave@stgolabs.net>
Reviewed-by: Pan Xinhui <xinhui.pan@linux.vnet.ibm.com>
Cc: Peter Zijlstra (Intel) <peterz@infradead.org>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: Waiman Long <Waiman.Long@hpe.com>
Cc: Davidlohr Bueso <dave@stgolabs.net>
Signed-off-by: Wanpeng Li <wanpeng.li@hotmail.com>
---
v3 -> v4:
* update patch subject
* add code comments
v2 -> v3:
* fix typo in patch description
v1 -> v2:
* adjust patch description
kernel/locking/qspinlock_paravirt.h | 23 ++++++++++++++++++++++-
1 file changed, 22 insertions(+), 1 deletion(-)
diff --git a/kernel/locking/qspinlock_paravirt.h b/kernel/locking/qspinlock_paravirt.h
index 21ede57..ca96db4 100644
--- a/kernel/locking/qspinlock_paravirt.h
+++ b/kernel/locking/qspinlock_paravirt.h
@@ -450,7 +450,28 @@ pv_wait_head_or_lock(struct qspinlock *lock, struct mcs_spinlock *node)
goto gotlock;
}
}
- WRITE_ONCE(pn->state, vcpu_halted);
+ /*
+ * lock holder vCPU queue head vCPU
+ * ---------------- ---------------
+ * node->locked = 1;
+ * <preemption> READ_ONCE(node->locked)
+ * ... pv_wait_head_or_lock():
+ * SPIN_THRESHOLD loop;
+ * pv_hash();
+ * lock->locked = _Q_SLOW_VAL;
+ * node->state = vcpu_hashed;
+ * pv_kick_node():
+ * cmpxchg(node->state,
+ * vcpu_halted, vcpu_hashed);
+ * lock->locked = _Q_SLOW_VAL;
+ * pv_hash();
+ *
+ * With preemption at the right moment, it is possible that both the
+ * lock holder and queue head vCPUs can be racing to set node->state.
+ * Making sure the state is never set to vcpu_halted will prevent this
+ * racing from happening.
+ */
+ WRITE_ONCE(pn->state, vcpu_hashed);
qstat_inc(qstat_pv_wait_head, true);
qstat_inc(qstat_pv_wait_again, waitcnt);
pv_wait(&l->locked, _Q_SLOW_VAL);
--
2.1.0
[toc] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-08-09 12:50 +0200 |
| Message-ID | <s4cLD-6of-1@gated-at.bofh.it> |
| In reply to | #1458548 |
On Tue, Aug 09, 2016 at 05:37:47PM +0800, Wanpeng Li wrote: > From: Wanpeng Li <wanpeng.li@hotmail.com> > > When the lock holder vCPU is racing with the queue head vCPU: > > lock holder vCPU queue head vCPU > ===================== ================== > > node->locked = 1; > <preemption> READ_ONCE(node->locked) > ... pv_wait_head_or_lock(): > SPIN_THRESHOLD loop; > pv_hash(); > lock->locked = _Q_SLOW_VAL; > node->state = vcpu_hashed; > pv_kick_node(): > cmpxchg(node->state, > vcpu_halted, vcpu_hashed); > lock->locked = _Q_SLOW_VAL; > pv_hash(); So here the example is 'wrong' in that it doesn't illustrate the fail case, namely having vcpu_halted win while we're hashed. > +++ b/kernel/locking/qspinlock_paravirt.h > @@ -450,7 +450,28 @@ pv_wait_head_or_lock(struct qspinlock *lock, struct mcs_spinlock *node) > goto gotlock; > } > } > - WRITE_ONCE(pn->state, vcpu_halted); > + /* > + * lock holder vCPU queue head vCPU > + * ---------------- --------------- > + * node->locked = 1; > + * <preemption> READ_ONCE(node->locked) > + * ... pv_wait_head_or_lock(): > + * SPIN_THRESHOLD loop; > + * pv_hash(); > + * lock->locked = _Q_SLOW_VAL; > + * node->state = vcpu_hashed; > + * pv_kick_node(): > + * cmpxchg(node->state, > + * vcpu_halted, vcpu_hashed); > + * lock->locked = _Q_SLOW_VAL; > + * pv_hash(); > + * > + * With preemption at the right moment, it is possible that both the > + * lock holder and queue head vCPUs can be racing to set node->state. > + * Making sure the state is never set to vcpu_halted will prevent this > + * racing from happening. > + */ > + WRITE_ONCE(pn->state, vcpu_hashed); > qstat_inc(qstat_pv_wait_head, true); > qstat_inc(qstat_pv_wait_again, waitcnt); > pv_wait(&l->locked, _Q_SLOW_VAL); And I completely fail to see the point of this comment, still. Yes, if we would have used vcpu_halted, there would be a problem, but we don't so there isn't. What does this comment tell us about the current code? In any case, I have an older version (possibly v1) queued up. That fixes the bug.
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-08-09 13:40 +0200 |
| Message-ID | <s4dy2-6VH-33@gated-at.bofh.it> |
| In reply to | #1458595 |
2016-08-09 18:49 GMT+08:00 Peter Zijlstra <peterz@infradead.org>: [...] > In any case, I have an older version (possibly v1) queued up. That fixes > the bug. Thanks. :) Regards, Wanpeng Li
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web