Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1622375 > unrolled thread

Re: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks

Started byAdam Wallis <awallis@codeaurora.org>
First post2017-04-12 19:10 +0200
Last post2017-04-28 17:40 +0200
Articles 3 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks Adam Wallis <awallis@codeaurora.org> - 2017-04-12 19:10 +0200
    Re: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks Will Deacon <will.deacon@arm.com> - 2017-04-24 15:40 +0200
    Re: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks Will Deacon <will.deacon@arm.com> - 2017-04-28 17:40 +0200

#1622375 — Re: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks

FromAdam Wallis <awallis@codeaurora.org>
Date2017-04-12 19:10 +0200
SubjectRe: [RFC PATCH 0/3] arm64: queued spinlocks and rw-locks
Message-ID<tvtWh-7dn-5@gated-at.bofh.it>
On 4/10/2017 5:35 PM, Yury Norov wrote:
> The patch of Jan Glauber enables queued spinlocks on arm64. I rebased it on
> latest kernel sources, and added a couple of fixes to headers to apply it 
> smoothly.
> 
> Though, locktourture test shows significant performance degradation in the
> acquisition of rw-lock for read on qemu:
> 
>                           Before           After
> spin_lock-torture:      38957034        37076367         -4.83
> rw_lock-torture W:       5369471        18971957        253.33
> rw_lock-torture R:       6413179         3668160        -42.80
> 

On our 48 core QDF2400 part, I am seeing huge improvements with these patches on
the torture tests. The improvements go up even further when I apply Jason Low's
MCS Spinlock patch: https://lkml.org/lkml/2016/4/20/725

> I'm  not much experienced in locking, and so wonder how it's possible that
> simple switching to generic queued rw-lock causes so significant performance
> degradation, while in theory it should improve it. Even more, on x86 there
> are no such problems probably.
> 
> I also think that patches 1 and 2 are correct and useful, and should be applied
> anyway.
> 
> Any comments appreciated.
> 
> Yury.
> 

I will be happy to tests these patches more thoroughly after you get some
additional comments/feedback.

> Jan Glauber (1):
>   arm64/locking: qspinlocks and qrwlocks support
> 
> Yury Norov (2):
>   kernel/locking: #include <asm/spinlock.h> in qrwlock.c
>   asm-generic: don't #include <linux/atomic.h> in qspinlock_types.h
> 
>  arch/arm64/Kconfig                      |  2 ++
>  arch/arm64/include/asm/qrwlock.h        |  7 +++++++
>  arch/arm64/include/asm/qspinlock.h      | 20 ++++++++++++++++++++
>  arch/arm64/include/asm/spinlock.h       | 12 ++++++++++++
>  arch/arm64/include/asm/spinlock_types.h | 14 +++++++++++---
>  include/asm-generic/qspinlock.h         |  1 +
>  include/asm-generic/qspinlock_types.h   |  8 --------
>  kernel/locking/qrwlock.c                |  1 +
>  8 files changed, 54 insertions(+), 11 deletions(-)
>  create mode 100644 arch/arm64/include/asm/qrwlock.h
>  create mode 100644 arch/arm64/include/asm/qspinlock.h
> 

Thanks

-- 
Adam Wallis
Qualcomm Datacenter Technologies as an affiliate of Qualcomm Technologies, Inc.
Qualcomm Technologies, Inc. is a member of the Code Aurora Forum,
a Linux Foundation Collaborative Project.

[toc] | [next] | [standalone]


#1629566

FromWill Deacon <will.deacon@arm.com>
Date2017-04-24 15:40 +0200
Message-ID<tzMnD-7Bg-5@gated-at.bofh.it>
In reply to#1622375
On Wed, Apr 12, 2017 at 01:04:55PM -0400, Adam Wallis wrote:
> On 4/10/2017 5:35 PM, Yury Norov wrote:
> > The patch of Jan Glauber enables queued spinlocks on arm64. I rebased it on
> > latest kernel sources, and added a couple of fixes to headers to apply it 
> > smoothly.
> > 
> > Though, locktourture test shows significant performance degradation in the
> > acquisition of rw-lock for read on qemu:
> > 
> >                           Before           After
> > spin_lock-torture:      38957034        37076367         -4.83
> > rw_lock-torture W:       5369471        18971957        253.33
> > rw_lock-torture R:       6413179         3668160        -42.80
> > 
> 
> On our 48 core QDF2400 part, I am seeing huge improvements with these patches on
> the torture tests. The improvements go up even further when I apply Jason Low's
> MCS Spinlock patch: https://lkml.org/lkml/2016/4/20/725

Does the QDF2400 implement the large system extensions? If so, how do the
queued lock implementations compare to the LSE-based ticket locks?

Will

[toc] | [prev] | [next] | [standalone]


#1632976

FromWill Deacon <will.deacon@arm.com>
Date2017-04-28 17:40 +0200
Message-ID<tBg9Y-Qc-19@gated-at.bofh.it>
In reply to#1622375
On Thu, Apr 13, 2017 at 01:33:09PM +0300, Yury Norov wrote:
> On Wed, Apr 12, 2017 at 01:04:55PM -0400, Adam Wallis wrote:
> > On 4/10/2017 5:35 PM, Yury Norov wrote:
> > > The patch of Jan Glauber enables queued spinlocks on arm64. I rebased it on
> > > latest kernel sources, and added a couple of fixes to headers to apply it 
> > > smoothly.
> > > 
> > > Though, locktourture test shows significant performance degradation in the
> > > acquisition of rw-lock for read on qemu:
> > > 
> > >                           Before           After
> > > spin_lock-torture:      38957034        37076367         -4.83
> > > rw_lock-torture W:       5369471        18971957        253.33
> > > rw_lock-torture R:       6413179         3668160        -42.80
> > > 
> > 
> > On our 48 core QDF2400 part, I am seeing huge improvements with these patches on
> > the torture tests. The improvements go up even further when I apply Jason Low's
> > MCS Spinlock patch: https://lkml.org/lkml/2016/4/20/725
> 
> It sounds great. So performance issue is looking like my local
> problem, most probably because I ran tests on Qemu VM.
> 
> I don't see any problems with this series, other than performance,
> and if it looks fine now, I think it's good enough for upstream.

I would still like to understand why you see such a significant performance
degradation, and whether or not you also see that on native hardware (i.e.
without Qemu involved).

Will

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web