Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1305571 > unrolled thread
| Started by | "Michael S. Tsirkin" <mst@redhat.com> |
|---|---|
| First post | 2016-01-10 15:20 +0100 |
| Last post | 2016-01-12 14:00 +0100 |
| Articles | 20 on this page of 86 — 7 participants |
Back to article view | Back to linux.kernel
[PATCH v3 00/41] arch: barrier cleanup + barriers for virt "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 06/41] s390: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 13/41] x86: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
Re: [PATCH v3 13/41] x86: reuse asm-generic/barrier.h Thomas Gleixner <tglx@linutronix.de> - 2016-01-12 15:20 +0100
[PATCH v3 08/41] arm: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 21/41] mips: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 03/41] ia64: rename nop->iosapic_nop "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 07/41] sparc: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 15/41] powerpc: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 12/41] x86/um: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 22/41] s390: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 11/41] mips: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-12 02:20 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-12 09:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-12 11:00 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-12 10:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-12 11:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-12 11:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Will Deacon <will.deacon@arm.com> - 2016-01-12 12:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-12 21:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-12 22:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-13 01:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Will Deacon <will.deacon@arm.com> - 2016-01-13 11:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-13 20:10 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-13 21:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-13 22:00 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Will Deacon <will.deacon@arm.com> - 2016-01-14 13:10 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 18:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 20:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-14 21:20 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 21:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-14 21:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 21:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 22:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 22:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 23:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-15 00:10 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 21:20 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 21:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 22:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 23:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-13 23:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-14 10:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Will Deacon <will.deacon@arm.com> - 2016-01-14 13:20 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 20:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 21:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 22:10 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-14 22:30 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-14 22:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-15 00:00 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-15 00:40 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-15 01:50 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> - 2016-01-15 02:10 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-15 10:00 +0100
Re: [v3,11/41] mips: reuse asm-generic/barrier.h Peter Zijlstra <peterz@infradead.org> - 2016-01-15 10:20 +0100
[PATCH v3 17/41] arm: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 14/41] asm-generic: add __smp_xxx wrappers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 09/41] arm64: reuse asm-generic/barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 18/41] blackfin: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 16/41] arm64: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 20/41] metag: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:20 +0100
[PATCH v3 35/41] checkpatch: check for __smp outside barrier.h "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 38/41] xen/io: use virt_xxx barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 36/41] checkpatch: add virt barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 26/41] xtensa: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 28/41] asm-generic: implement virt_xxx memory barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 37/41] xenbus: use virt_xxx barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 24/41] sparc: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 40/41] s390: use generic memory barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 25/41] tile: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 29/41] Revert "virtio_ring: Update weak barriers to use dma_wmb/rmb" "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 34/41] checkpatch.pl: add missing memory barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 41/41] s390: more efficient smp barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 33/41] virtio_ring: use virt_store_mb "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 01/41] lcoking/barriers, arch: Use smp barriers in smp_store_release() "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
Re: [PATCH v3 01/41] lcoking/barriers, arch: Use smp barriers in smp_store_release() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-01-12 17:30 +0100
Re: [PATCH v3 01/41] lcoking/barriers, arch: Use smp barriers in smp_store_release() "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-12 19:50 +0100
[PATCH v3 30/41] virtio_ring: update weak barriers to use virt_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 32/41] sh: move xchg_cmpxchg to a header by itself "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 23/41] sh: define __smp_xxx, fix smp_store_mb for !SMP "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
[PATCH v3 27/41] x86: define __smp_xxx "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
Re: [PATCH v3 27/41] x86: define __smp_xxx Thomas Gleixner <tglx@linutronix.de> - 2016-01-12 15:20 +0100
[PATCH v3 39/41] xen/events: use virt_xxx barriers "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
Re: [PATCH v3 39/41] xen/events: use virt_xxx barriers David Vrabel <david.vrabel@citrix.com> - 2016-01-11 12:20 +0100
[PATCH v3 31/41] sh: support 1 and 2 byte xchg "Michael S. Tsirkin" <mst@redhat.com> - 2016-01-10 15:30 +0100
Re: [PATCH v3 00/41] arch: barrier cleanup + barriers for virt Peter Zijlstra <peterz@infradead.org> - 2016-01-12 14:00 +0100
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-01-12 22:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQeZd-6Mv-33@gated-at.bofh.it> |
| In reply to | #1307820 |
On Tue, Jan 12, 2016 at 12:45:14PM -0800, Leonid Yegoshin wrote:
> (I try to answer on multiple mails in one)
>
> First of all, it seems like some generic notes should be given here:
>
> 1. Generic MIPS "SYNC" (aka "SYNC 0") instruction is a very heavy in some
> CPUs. On that CPUs it basically kills pipelines in each CPU, can do a
> special memory/IO bus transaction (similar to "fence") and hold a system
> until all R/W is completed. It is like Big Kernel Lock but worse. So, the
> move to SMP_* kind of barriers is needed to improve performance, especially
> on newest CPUs with long pipelines.
The MIPS SYNC isn't any worse than the PPC SYNC, x86 MFENCE or arm DSB
SY, yes they're heavy, so what.
> 2. MIPS Arch document may be misleading because words "ordering" and
> "completion" means different from Linux, the SYNC instruction description is
> written for HW engineers. I wrote that in a separate patch of the same
> patchset - http://patchwork.linux-mips.org/patch/10505/ "MIPS: R6: Use
> lightweight SYNC instruction in smp_* memory barriers":
Did you actually say anything here?
> >This instructions were specifically designed to work for smp_*() sort of
> >memory barriers in MIPS R2/R3/R5 and R6.
> >
> >Unfortunately, it's description is very cryptic and is done in HW engineering
> >style which prevents use of it by SW.
>
> 3. I bother MIPS Arch team long time until I completely understood that MIPS
> SYNC_WMB, SYNC_MB, SYNC_RMB, SYNC_RELEASE and SYNC_ACQUIRE do an exactly
> that is required in Documentation/memory-barriers.txt
Ha! and you think that document covers all the really fun details?
In particular we're very much all 'confused' about the various notions
of transitivity and what barriers imply how much of it.
> In Peter Zijlstra mail:
>
> >1) you do not make such things selectable; either the hardware needs
> >them or it doesn't. If it does you_must_ use them, however unlikely.
> It is selectable only for MIPS R2 but not MIPS R6. The reason is - most of
> MIPS R2 CPUs have short pipeline and that SYNC is just waste of CPU
> resource, especially taking into account that "lightweight syncs" are
> converted to a heavy "SYNC 0" in many of that CPUs. However the latest
> MIPS/Imagination CPU have a pipeline long enough to hit a problem - absence
> of SYNC at LL/SC inside atomics, barriers etc.
What ?! Are you saying that because R2 has short pipelines its unlikely
to hit the reordering issues and we can omit barriers?
> >And reading the MIPS64 v6.04 instruction set manual, I think 0x11/0x12
> >are_NOT_ transitive and therefore cannot be used to implement the
> >smp_mb__{before,after} stuff.
> >
> >That is, in MIPS speak, those SYNC types are Ordering Barriers, not
> >Completion Barriers.
>
> Please see above, point 2.
That did not in fact enlighten things. Are they transitive/multi-copy
atomic or not?
(and here Will will go into great detail on the differences between the
two and make our collective brains explode :-)
> >That is, currently all architectures -- with exception of PPC -- have
> >RCsc locks, but using these non-transitive things will get you RCpc
> >locks.
> >
> >So yes, MIPS can go RCpc for its locks and share the burden of pain with
> >PPC, but that needs to be a very concious decision.
>
> I don't understand that - I tried hard but I can't find any word like
> "RCsc", "RCpc" in Documents/ directory. Web search goes nowhere, of course.
From: lkml.kernel.org/r/20150828153921.GF19282@twins.programming.kicks-ass.net
Yes, the difference between RCpc and RCsc is in the meaning of RELEASE +
ACQUIRE. With RCsc that implies a full memory barrier, with RCpc it does
not.
Currently PowerPC is the only arch that (can, and) does RCpc and gives a
weaker RELEASE + ACQUIRE. Only the CPU who did the ACQUIRE is guaranteed
to see the stores of the CPU which did the RELEASE in order.
As it stands, RCU is the only _known_ codebase where this matters, but
we did in fact write code for a fair number of years 'assuming' RELEASE
+ ACQUIRE was a full barrier, so who knows what else is out there.
RCsc - release consistency sequential consistency
RCpc - release consistency processor consistency
https://en.wikipedia.org/wiki/Processor_consistency
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-13 01:30 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQhu1-8P-1@gated-at.bofh.it> |
| In reply to | #1307854 |
On 01/12/2016 01:40 PM, Peter Zijlstra wrote:
>
>> It is selectable only for MIPS R2 but not MIPS R6. The reason is - most of
>> MIPS R2 CPUs have short pipeline and that SYNC is just waste of CPU
>> resource, especially taking into account that "lightweight syncs" are
>> converted to a heavy "SYNC 0" in many of that CPUs. However the latest
>> MIPS/Imagination CPU have a pipeline long enough to hit a problem - absence
>> of SYNC at LL/SC inside atomics, barriers etc.
> What ?! Are you saying that because R2 has short pipelines its unlikely
> to hit the reordering issues and we can omit barriers?
It was my guess to explain - why barriers was not included originally.
You can check with Ralf, he knows more about that time MIPS Linux code.
I bother with this more than 2 years and I just try to solve that issue
- in recent CPUs the load after LL/SC synchronization instruction loop
can get ahead of SC for sure, it was tested.
>
>>> And reading the MIPS64 v6.04 instruction set manual, I think 0x11/0x12
>>> are_NOT_ transitive and therefore cannot be used to implement the
>>> smp_mb__{before,after} stuff.
>>>
>>> That is, in MIPS speak, those SYNC types are Ordering Barriers, not
>>> Completion Barriers.
>> Please see above, point 2.
> That did not in fact enlighten things. Are they transitive/multi-copy
> atomic or not?
Peter Zijlstra recently wrote: "In particular we're very much all
'confused' about the various notions of transitivity". I am actually
confused too and need some examples here.
>
> (and here Will will go into great detail on the differences between the
> two and make our collective brains explode :-)
>
>>> That is, currently all architectures -- with exception of PPC -- have
>>> RCsc locks, but using these non-transitive things will get you RCpc
>>> locks.
>>>
>>> So yes, MIPS can go RCpc for its locks and share the burden of pain with
>>> PPC, but that needs to be a very concious decision.
>> I don't understand that - I tried hard but I can't find any word like
>> "RCsc", "RCpc" in Documents/ directory. Web search goes nowhere, of course.
> From: lkml.kernel.org/r/20150828153921.GF19282@twins.programming.kicks-ass.net
>
> Yes, the difference between RCpc and RCsc is in the meaning of RELEASE +
> ACQUIRE. With RCsc that implies a full memory barrier, with RCpc it does
> not.
MIPS Arch starting from R2 requires that. If some CPU can't, it should
execute a full "SYNC 0" instead, which is a full memory barrier.
>
> Currently PowerPC is the only arch that (can, and) does RCpc and gives a
> weaker RELEASE + ACQUIRE. Only the CPU who did the ACQUIRE is guaranteed
> to see the stores of the CPU which did the RELEASE in order.
Yes, it was a goal for SYNC_ACQUIRE and SYNC_RELEASE.
Caveats:
- "Full memory barrier" on MIPS means - full barrier for any device
in coherent domain. In MIPS Tech/Imagination Tech MIPS-based CPU it is
"for any device connected to CM or IOCU + directly connected memory".
- It is not applied to instruction fetch. However, I-Cache flushes
and SYNCI are consistent with that. There is also hazard barrier
instructions to clear CPU pipeline to some extent - to help with this
limitation.
I don't think that these caveats prevent a correct Acquire/Release semantic.
- Leonid.
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-01-13 11:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQra1-6Tj-11@gated-at.bofh.it> |
| In reply to | #1307820 |
On Tue, Jan 12, 2016 at 12:45:14PM -0800, Leonid Yegoshin wrote: > >The issue I have with the SYNC description in the text above is that it > >describes the single CPU (program order) and the dual-CPU (confusingly > >named global order) cases, but then doesn't generalise any further. That > >means we can't sensibly reason about transitivity properties when a third > >agent is involved. For example, the WRC+sync+addr test: > > > > > >P0: > >Wx = 1 > > > >P1: > >Rx == 1 > >SYNC > >Wy = 1 > > > >P2: > >Ry == 1 > ><address dep> > >Rx = 0 > > > > > >I can't find anything to forbid that, given the text. The main problem > >is having the SYNC on P1 affect the write by P0. > > As I understand that test, the visibility of P0: W[x] = 1 is identical to P1 > and P2 here. If P1 got X before SYNC and write to Y after SYNC then > instruction source register dependency tracking in P2 prevents a speculative > load of X before P2 obtains Y from the same place as P0/P1 and calculate > address of X. If some load of X in P2 happens before address dependency > calculation it's result is discarded. I don't think the address dependency is enough on its own. By that reasoning, the following variant (WRC+addr+addr) would work too: P0: Wx = 1 P1: Rx == 1 <address dep> Wy = 1 P2: Ry == 1 <address dep> Rx = 0 So are you saying that this is also forbidden? Imagine that P0 and P1 are two threads that share a store buffer. What then? > Yes, you can't find that in MIPS SYNC instruction description, it is more > likely in CM (Coherence Manager) area. I just pointed our arch team member > responsible for documents and he will think how to explain that. I tried grepping the linked documents for "coherence manager" but couldn't find anything. Is the description you refer to available anywhere? Will
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-13 20:10 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQyXU-41l-9@gated-at.bofh.it> |
| In reply to | #1308275 |
On 01/13/2016 02:45 AM, Will Deacon wrote: > On Tue, Jan 12, 2016 at 12:45:14PM -0800, Leonid Yegoshin wrote: >> > I don't think the address dependency is enough on its own. By that > reasoning, the following variant (WRC+addr+addr) would work too: > > > P0: > Wx = 1 > > P1: > Rx == 1 > <address dep> > Wy = 1 > > P2: > Ry == 1 > <address dep> > Rx = 0 > > > So are you saying that this is also forbidden? > Imagine that P0 and P1 are two threads that share a store buffer. What > then? > I ask HW team about it but I have a question - has it any relationship with replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? You use any barrier or do not use it and I just voice an intention to use a more efficient instruction instead of bold hummer (SYNC instruction). If you don't use any barrier here then it is a different issue. May be it has sense to return back to original issue? - Leonid
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-01-13 21:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQAwG-50b-31@gated-at.bofh.it> |
| In reply to | #1308737 |
On Wed, Jan 13, 2016 at 11:02:35AM -0800, Leonid Yegoshin wrote: > I ask HW team about it but I have a question - has it any relationship with > replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? Of course. If you cannot explain the semantics of the primitives you introduce, how can we judge the patch. This barrier business is hard enough as it is, but magic unexplained hardware makes it impossible. Rest assured, you (MIPS) isn't the first (nor likely the last) to go through all this. We've had these discussions (and to a certain extend are still having them) for x86, PPC, Alpha, ARM, etc.. Any every time new barriers instructions get introduced we had better have a full and comprehensive explanation to go along with them.
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-13 22:00 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQAGm-54j-13@gated-at.bofh.it> |
| In reply to | #1308806 |
On 01/13/2016 12:48 PM, Peter Zijlstra wrote: > On Wed, Jan 13, 2016 at 11:02:35AM -0800, Leonid Yegoshin wrote: > >> I ask HW team about it but I have a question - has it any relationship with >> replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? > Of course. If you cannot explain the semantics of the primitives you > introduce, how can we judge the patch. > > You missed a point - it is a question about replacement of SYNC with lightweight primitives. It is NOT a question about multithread system behavior without any SYNC. The answer on a latest Will's question lies in different area. - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-01-14 13:10 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQOT0-6Lb-19@gated-at.bofh.it> |
| In reply to | #1308812 |
On Wed, Jan 13, 2016 at 12:58:22PM -0800, Leonid Yegoshin wrote: > On 01/13/2016 12:48 PM, Peter Zijlstra wrote: > >On Wed, Jan 13, 2016 at 11:02:35AM -0800, Leonid Yegoshin wrote: > > > >>I ask HW team about it but I have a question - has it any relationship with > >>replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? > >Of course. If you cannot explain the semantics of the primitives you > >introduce, how can we judge the patch. > > > > > You missed a point - it is a question about replacement of SYNC with > lightweight primitives. It is NOT a question about multithread system > behavior without any SYNC. The answer on a latest Will's question lies in > different area. The reason we (Peter and I) care about this isn't because we enjoy being obstructive. It's because there is a whole load of core (i.e. portable) kernel code that is written to the *kernel* memory model. For example, the scheduler, RCU, mutex implementations, perf, drivers, you name it. Consequently, it's important that the architecture back-ends implement these portable primitives (e.g. smp_mb()) in a way that satisfies the kernel memory model so that core code doesn't need to worry about the underlying architecture for synchronisation purposes. You could turn around and say "but if MIPS gets it wrong, then that's MIPS's problem", but actually not having a general understanding of the ordering guarantees provided by each architecture makes it very difficult for us to extend the kernel memory model in such a way that it can be implemented efficiently across the board *and* relied upon by core code. The virtio patch at the start of the thread doesn't particularly concern me. It's the other patches you linked to that implement acquire/release that have me worried. Will
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-01-14 18:40 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQU2l-1Oc-11@gated-at.bofh.it> |
| In reply to | #1309212 |
On Thu, Jan 14, 2016 at 12:04:45PM +0000, Will Deacon wrote: > On Wed, Jan 13, 2016 at 12:58:22PM -0800, Leonid Yegoshin wrote: > > On 01/13/2016 12:48 PM, Peter Zijlstra wrote: > > >On Wed, Jan 13, 2016 at 11:02:35AM -0800, Leonid Yegoshin wrote: > > > > > >>I ask HW team about it but I have a question - has it any relationship with > > >>replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? > > >Of course. If you cannot explain the semantics of the primitives you > > >introduce, how can we judge the patch. > > > > > > > > You missed a point - it is a question about replacement of SYNC with > > lightweight primitives. It is NOT a question about multithread system > > behavior without any SYNC. The answer on a latest Will's question lies in > > different area. > > The reason we (Peter and I) care about this isn't because we enjoy being > obstructive. It's because there is a whole load of core (i.e. portable) > kernel code that is written to the *kernel* memory model. For example, > the scheduler, RCU, mutex implementations, perf, drivers, you name it. > > Consequently, it's important that the architecture back-ends implement > these portable primitives (e.g. smp_mb()) in a way that satisfies the > kernel memory model so that core code doesn't need to worry about the > underlying architecture for synchronisation purposes. You could turn > around and say "but if MIPS gets it wrong, then that's MIPS's problem", > but actually not having a general understanding of the ordering guarantees > provided by each architecture makes it very difficult for us to extend > the kernel memory model in such a way that it can be implemented > efficiently across the board *and* relied upon by core code. What Will said! Yes, you can cut corners within MIPS architecture-specific code, but primitives that are used in the core kernel really do need to work as expected. Thanx, Paul > The virtio patch at the start of the thread doesn't particularly concern > me. It's the other patches you linked to that implement acquire/release > that have me worried. > > Will >
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-14 20:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQW4a-3gr-25@gated-at.bofh.it> |
| In reply to | #1309529 |
On 01/14/2016 08:16 AM, Paul E. McKenney wrote: > On Thu, Jan 14, 2016 at 12:04:45PM +0000, Will Deacon wrote: >> On Wed, Jan 13, 2016 at 12:58:22PM -0800, Leonid Yegoshin wrote: >>> On 01/13/2016 12:48 PM, Peter Zijlstra wrote: >>>> On Wed, Jan 13, 2016 at 11:02:35AM -0800, Leonid Yegoshin wrote: >>>> >>>>> I ask HW team about it but I have a question - has it any relationship with >>>>> replacing MIPS SYNC with lightweight SYNCs (SYNC_WMB etc)? >>>> Of course. If you cannot explain the semantics of the primitives you >>>> introduce, how can we judge the patch. >>>> >>>> >>> You missed a point - it is a question about replacement of SYNC with >>> lightweight primitives. It is NOT a question about multithread system >>> behavior without any SYNC. The answer on a latest Will's question lies in >>> different area. > What Will said! > > Yes, you can cut corners within MIPS architecture-specific code, > but primitives that are used in the core kernel really do need to > work as expected. > > Thanx, Paul > > Absolutelly! Please use SYNC - right now it is not. An the only point - please use an appropriate SYNC_* barriers instead of heavy bold hammer. That stuff was design explicitly to support the requirements of Documentation/memory-barriers.txt It is easy - just use smp_acquire instead of plain smp_mb insmp_load_acquire, at least for MIPS. - Leonid. - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-01-14 21:20 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQWxc-3FW-19@gated-at.bofh.it> |
| In reply to | #1309610 |
On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: > An the only point - please use an appropriate SYNC_* barriers instead of > heavy bold hammer. That stuff was design explicitly to support the > requirements of Documentation/memory-barriers.txt That's madness. That document changes from version to version as to what we _think_ the actual hardware does. It is _NOT_ a specification. You cannot design hardware from that. Its incomplete and fails to specify a bunch of things. It not a mathematically sound definition of a memory model. Please stop referring to that document for what a particular barrier _should_ do. Explain what MIPS does, so we can attempt to integrate this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve upon our understanding of hardware and improve the Linux memory model.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-01-14 21:40 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQWQy-3OB-9@gated-at.bofh.it> |
| In reply to | #1309629 |
On Thu, Jan 14, 2016 at 09:15:13PM +0100, Peter Zijlstra wrote: > On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: > > An the only point - please use an appropriate SYNC_* barriers instead of > > heavy bold hammer. That stuff was design explicitly to support the > > requirements of Documentation/memory-barriers.txt > > That's madness. That document changes from version to version as to what > we _think_ the actual hardware does. It is _NOT_ a specification. There is work in progress on a specification, but please don't hold your breath. And I am not as optimistic as I might be about any formal specification keeping up with the Linux kernel or with the hardware that it supports. But it seems worth a good try. > You cannot design hardware from that. Its incomplete and fails to > specify a bunch of things. It not a mathematically sound definition of a > memory model. > > Please stop referring to that document for what a particular barrier > _should_ do. Explain what MIPS does, so we can attempt to integrate > this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve > upon our understanding of hardware and improve the Linux memory model. Please! Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-01-14 21:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQX0e-3S8-5@gated-at.bofh.it> |
| In reply to | #1309629 |
On Thu, Jan 14, 2016 at 09:15:13PM +0100, Peter Zijlstra wrote:
> On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote:
> > An the only point - please use an appropriate SYNC_* barriers instead of
> > heavy bold hammer. That stuff was design explicitly to support the
> > requirements of Documentation/memory-barriers.txt
>
> That's madness. That document changes from version to version as to what
> we _think_ the actual hardware does. It is _NOT_ a specification.
>
> You cannot design hardware from that. Its incomplete and fails to
> specify a bunch of things. It not a mathematically sound definition of a
> memory model.
>
> Please stop referring to that document for what a particular barrier
> _should_ do. Explain what MIPS does, so we can attempt to integrate
> this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve
> upon our understanding of hardware and improve the Linux memory model.
That is, if you'd managed to read that file at the right point in time,
you might have through we'd be OK with requiring a barrier for
control dependencies.
We got rid of that mistake. It was based on a flawed reading of the
Alpha docs. See: 105ff3cbf225 ("atomic: remove all traces of
READ_ONCE_CTRL() and atomic*_read_ctrl()")
Similarly, while the document goes to great length to explain the
read_barrier_depends thing, nobody actually thinks its a brilliant idea
to have. Ideally we'd kill the thing the moment we drop Alpha support.
Again, memory-barriers.txt is _NOT_, I repeat, _NOT_ a hardware spec, it
is not even a recommendation. It are our best effort (but flawed)
scribbles of what we think is makes sense given the huge amount of
actual hardware we have to run on.
As to the ACQUIRE/RELEASE semantics, ARM64 actually has
multi-copy-atomic acquire/release (as does ia64, although in reality it
doesn't actually have acquire/release). PPC otoh does _NOT_ have this,
and is currently the only arch to suffer RCpc locks.
Now for a long long time we assumed our locks were RCsc, and we've
written code assuming UNLOCK x + LOCK y was in fact a full barrier with
transitiviy. Then we figured out PPC didn't actually match that. RCU is
the only piece of code we _know_ relied on that, but there might be more
out there...
So we document, for new code, that UNLOCK+LOCK isn't a MB, while at the
same time we lobby PPC to stick a full barrier in and get rid of this
stuff.
Nobody really likes RCpc locks, esp. given the history we have of
assuming RCsc.
The current document allowing for RCpc is not an endorsement thereof.
Ideally we'd _NOT_ have to worry about that. We can do without these
head-aches.
So again, stop referring to our document as a spec. Also please don't
make MIPS push the limits of weak memory models, we really can do
without the pain.
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-14 21:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQX0f-3S8-23@gated-at.bofh.it> |
| In reply to | #1309629 |
On 01/14/2016 12:15 PM, Peter Zijlstra wrote: > On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: >> An the only point - please use an appropriate SYNC_* barriers instead of >> heavy bold hammer. That stuff was design explicitly to support the >> requirements of Documentation/memory-barriers.txt > That's madness. That document changes from version to version as to what > we _think_ the actual hardware does. It is _NOT_ a specification. > > You cannot design hardware from that. Its incomplete and fails to > specify a bunch of things. It not a mathematically sound definition of a > memory model. > > Please stop referring to that document for what a particular barrier > _should_ do. Explain what MIPS does, so we can attempt to integrate > this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve > upon our understanding of hardware and improve the Linux memory model. I am afraid I can't help you here. It is very complicated stuff and a model is actually doesn't fit your assumptions about CPUs well without some simplifications which are based on what you want to have. I say that SYNC_ACQUIRE/etc follows what you expect for smp_acquire etc (basing on that document). And at least two CPU models were tested with my patches (see it in LMO) for that last year and that instructions are implemented now in engineering kernel. If you have something else in mind, you can ask me. But I prefer to do not deviate too much from Documentation/memory-barriers.txt, for exam - if it asks to have memory barrier somewhere, then I assume the code should have it, and please - don't ask me a test which violates the current version of document recommendations. For a moment I don't see a significant changes in this document for MIPS Arch at least 1.5 year, and the only significant point is that MIPS CPU Arch doesn't have yet smp_read_barrier_depends() and smp_rmb() should be used instead. - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-01-14 22:40 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQXMC-4pB-9@gated-at.bofh.it> |
| In reply to | #1309645 |
On Thu, Jan 14, 2016 at 12:46:43PM -0800, Leonid Yegoshin wrote: > On 01/14/2016 12:15 PM, Peter Zijlstra wrote: > >On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: > >>An the only point - please use an appropriate SYNC_* barriers instead of > >>heavy bold hammer. That stuff was design explicitly to support the > >>requirements of Documentation/memory-barriers.txt > >That's madness. That document changes from version to version as to what > >we _think_ the actual hardware does. It is _NOT_ a specification. > > > >You cannot design hardware from that. Its incomplete and fails to > >specify a bunch of things. It not a mathematically sound definition of a > >memory model. > > > >Please stop referring to that document for what a particular barrier > >_should_ do. Explain what MIPS does, so we can attempt to integrate > >this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve > >upon our understanding of hardware and improve the Linux memory model. > > I am afraid I can't help you here. It is very complicated stuff and > a model is actually doesn't fit your assumptions about CPUs well > without some simplifications which are based on what you want to > have. > > I say that SYNC_ACQUIRE/etc follows what you expect for smp_acquire > etc (basing on that document). And at least two CPU models were > tested with my patches (see it in LMO) for that last year and that > instructions are implemented now in engineering kernel. > > If you have something else in mind, you can ask me. But I prefer to > do not deviate too much from Documentation/memory-barriers.txt, for > exam - if it asks to have memory barrier somewhere, then I assume > the code should have it, and please - don't ask me a test which > violates the current version of document recommendations. > > For a moment I don't see a significant changes in this document for > MIPS Arch at least 1.5 year, and the only significant point is that > MIPS CPU Arch doesn't have yet smp_read_barrier_depends() and > smp_rmb() should be used instead. Is SYNC_ACQUIRE a memory-barrier instruction that orders prior loads against later loads and stores? If so, and if MIPS does not do ordering based on address and data dependencies, I suggest making read_barrier_depends() be a SYNC_ACQUIRE rather than SYNC_RMB. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-14 22:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQXWj-4xw-53@gated-at.bofh.it> |
| In reply to | #1309664 |
On 01/14/2016 01:34 PM, Paul E. McKenney wrote: > On Thu, Jan 14, 2016 at 12:46:43PM -0800, Leonid Yegoshin wrote: >> On 01/14/2016 12:15 PM, Peter Zijlstra wrote: >>> On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: >>>> An the only point - please use an appropriate SYNC_* barriers instead of >>>> heavy bold hammer. That stuff was design explicitly to support the >>>> requirements of Documentation/memory-barriers.txt >>> That's madness. That document changes from version to version as to what >>> we _think_ the actual hardware does. It is _NOT_ a specification. >>> >>> You cannot design hardware from that. Its incomplete and fails to >>> specify a bunch of things. It not a mathematically sound definition of a >>> memory model. >>> >>> Please stop referring to that document for what a particular barrier >>> _should_ do. Explain what MIPS does, so we can attempt to integrate >>> this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve >>> upon our understanding of hardware and improve the Linux memory model. >> I am afraid I can't help you here. It is very complicated stuff and >> a model is actually doesn't fit your assumptions about CPUs well >> without some simplifications which are based on what you want to >> have. >> >> I say that SYNC_ACQUIRE/etc follows what you expect for smp_acquire >> etc (basing on that document). And at least two CPU models were >> tested with my patches (see it in LMO) for that last year and that >> instructions are implemented now in engineering kernel. >> >> If you have something else in mind, you can ask me. But I prefer to >> do not deviate too much from Documentation/memory-barriers.txt, for >> exam - if it asks to have memory barrier somewhere, then I assume >> the code should have it, and please - don't ask me a test which >> violates the current version of document recommendations. >> >> For a moment I don't see a significant changes in this document for >> MIPS Arch at least 1.5 year, and the only significant point is that >> MIPS CPU Arch doesn't have yet smp_read_barrier_depends() and >> smp_rmb() should be used instead. > Is SYNC_ACQUIRE a memory-barrier instruction that orders prior loads > against later loads and stores? Yes, it is in MD00087 (table 6.6 of document Ver 6.04) - https://imgtec.com/?do-download=4302 > If so, and if MIPS does not do > ordering based on address and data dependencies, I suggest making > read_barrier_depends() be a SYNC_ACQUIRE rather than SYNC_RMB. I understood that, after I see the example of using it. Please consider to add that into Documentation/memory-barriers.txt (it is not easy to find that this barrier is used for shared WRITE basing on shared pointer), it would be helpful. - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-01-14 23:30 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQYz0-5b5-3@gated-at.bofh.it> |
| In reply to | #1309683 |
On Thu, Jan 14, 2016 at 01:45:44PM -0800, Leonid Yegoshin wrote: > On 01/14/2016 01:34 PM, Paul E. McKenney wrote: > >On Thu, Jan 14, 2016 at 12:46:43PM -0800, Leonid Yegoshin wrote: > >>On 01/14/2016 12:15 PM, Peter Zijlstra wrote: > >>>On Thu, Jan 14, 2016 at 11:42:02AM -0800, Leonid Yegoshin wrote: > >>>>An the only point - please use an appropriate SYNC_* barriers instead of > >>>>heavy bold hammer. That stuff was design explicitly to support the > >>>>requirements of Documentation/memory-barriers.txt > >>>That's madness. That document changes from version to version as to what > >>>we _think_ the actual hardware does. It is _NOT_ a specification. > >>> > >>>You cannot design hardware from that. Its incomplete and fails to > >>>specify a bunch of things. It not a mathematically sound definition of a > >>>memory model. > >>> > >>>Please stop referring to that document for what a particular barrier > >>>_should_ do. Explain what MIPS does, so we can attempt to integrate > >>>this knowledge with our knowledge of PPC/ARM/Alpha/x86/etc. and improve > >>>upon our understanding of hardware and improve the Linux memory model. > >>I am afraid I can't help you here. It is very complicated stuff and > >>a model is actually doesn't fit your assumptions about CPUs well > >>without some simplifications which are based on what you want to > >>have. > >> > >>I say that SYNC_ACQUIRE/etc follows what you expect for smp_acquire > >>etc (basing on that document). And at least two CPU models were > >>tested with my patches (see it in LMO) for that last year and that > >>instructions are implemented now in engineering kernel. > >> > >>If you have something else in mind, you can ask me. But I prefer to > >>do not deviate too much from Documentation/memory-barriers.txt, for > >>exam - if it asks to have memory barrier somewhere, then I assume > >>the code should have it, and please - don't ask me a test which > >>violates the current version of document recommendations. > >> > >>For a moment I don't see a significant changes in this document for > >>MIPS Arch at least 1.5 year, and the only significant point is that > >>MIPS CPU Arch doesn't have yet smp_read_barrier_depends() and > >>smp_rmb() should be used instead. > > >Is SYNC_ACQUIRE a memory-barrier instruction that orders prior loads > >against later loads and stores? > > Yes, it is in MD00087 (table 6.6 of document Ver 6.04) - > https://imgtec.com/?do-download=4302 OK, it does look like it should work. Of course, if you can rely on straight address/data dependencies, that would be even better. > > If so, and if MIPS does not do > >ordering based on address and data dependencies, I suggest making > >read_barrier_depends() be a SYNC_ACQUIRE rather than SYNC_RMB. > > I understood that, after I see the example of using it. > Please consider to add that into Documentation/memory-barriers.txt > (it is not easy to find that this barrier is used for shared WRITE > basing on shared pointer), it would be helpful. Actually, the Linux kernel doesn't have an acquire barrier, just an smp_load_acquire(). Or did someone sneak one in while I wasn't looking? ;-) Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-15 00:10 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQZbI-5Fp-19@gated-at.bofh.it> |
| In reply to | #1309707 |
On 01/14/2016 02:24 PM, Paul E. McKenney wrote: > Actually, the Linux kernel doesn't have an acquire barrier, just an > smp_load_acquire(). Or did someone sneak one in while I wasn't looking? That was an exactly starting point for this discussion. This patch just pulls out from MIPS files smp_load_acquire() and smp_store_release(). However, I put into LMO half year ago the patch http://patchwork.linux-mips.org/patch/10506/ which replaces a generic smp_mb with MIPS specific smp_release/acquire in that functions. This patch also fixes use of SYNCs barriers in spin_locks/atomics/bitops for Imagination MIPS CPUs too - it is just absent now for any Imagination MIPS CPUs! Michael later pointed me that it can be returned back with his series of patches but discussion was already here. - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-14 21:20 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQWxb-3FW-9@gated-at.bofh.it> |
| In reply to | #1309212 |
On 01/14/2016 04:04 AM, Will Deacon wrote: > Consequently, it's important that the architecture back-ends implement > these portable primitives (e.g. smp_mb()) in a way that satisfies the > kernel memory model so that core code doesn't need to worry about the > underlying architecture for synchronisation purposes. It seems you don't listen me. I said multiple times - MIPS implementation of SYNC_RMB/SYNC_WMB/SYNC_MB/SYNC_ACQUIRE/SYNC_RELEASE instructions matches the description of smp_rmb/smp_wmb/smp_mb/sync_acquire/sync_release from Documentation/memory-barriers.txt file. What else do you want from me - RTL or microArch design for that? - Leonid.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-01-14 21:50 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQX0d-3S8-1@gated-at.bofh.it> |
| In reply to | #1309626 |
On Thu, Jan 14, 2016 at 12:12:53PM -0800, Leonid Yegoshin wrote: > On 01/14/2016 04:04 AM, Will Deacon wrote: > >Consequently, it's important that the architecture back-ends > >implement these portable primitives (e.g. smp_mb()) in a way that > >satisfies the kernel memory model so that core code doesn't need > >to worry about the underlying architecture for synchronisation > >purposes. > > It seems you don't listen me. I said multiple times - MIPS > implementation of > SYNC_RMB/SYNC_WMB/SYNC_MB/SYNC_ACQUIRE/SYNC_RELEASE instructions > matches the description of > smp_rmb/smp_wmb/smp_mb/sync_acquire/sync_release from > Documentation/memory-barriers.txt file. > > What else do you want from me - RTL or microArch design for that? I suspect that it is more likely that we are talking past each other. This stuff is subtle and although we have better ways of talking about it than (say) ten years ago, it is subtle. Two ways of talking about it are herd and ppcmem. The overview of ppcmem (AKA armmem and cppmem) is here: https://www.cl.cam.ac.uk/~pes20/ppcmem/help.html The intro to herd is here: http://arxiv.org/pdf/1308.6810v5.pdf It may be downloaded here: http://diy.inria.fr/herd/ As a very rough rule of thumb, herd is faster and easier to use and ppcmem is more precise. So SYNC_RMB is intended to implement smp_rmb(), correct? You could use SYNC_ACQUIRE() to implement read_barrier_depends() and smp_read_barrier_depends(), but SYNC_RMB probably does not suffice. The reason for this is that smp_read_barrier_depends() must order the pointer load against any subsequent read or write through a dereference of that pointer. For example: p = READ_ONCE(gp); smp_rmb(); r1 = p->a; /* ordered by smp_rmb(). */ p->b = 42; /* NOT ordered by smp_rmb(), BUG!!! */ r2 = x; /* ordered by smp_rmb(), but doesn't need to be. */ In contrast: p = READ_ONCE(gp); smp_read_barrier_depends(); r1 = p->a; /* ordered by smp_read_barrier_depends(). */ p->b = 42; /* ordered by smp_read_barrier_depends(). */ r2 = x; /* not ordered by smp_read_barrier_depends(), which is OK. */ Again, if your hardware maintains local ordering for address and data dependencies, you can have read_barrier_depends() and smp_read_barrier_depends() be no-ops like they are for most architectures. Does that help? Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Leonid Yegoshin <Leonid.Yegoshin@imgtec.com> |
|---|---|
| Date | 2016-01-14 22:30 +0100 |
| Subject | Re: [v3,11/41] mips: reuse asm-generic/barrier.h |
| Message-ID | <qQXCX-4lN-21@gated-at.bofh.it> |
| In reply to | #1309643 |
On 01/14/2016 12:48 PM, Paul E. McKenney wrote: > > So SYNC_RMB is intended to implement smp_rmb(), correct? Yes. > > You could use SYNC_ACQUIRE() to implement read_barrier_depends() and > smp_read_barrier_depends(), but SYNC_RMB probably does not suffice. If smp_read_barrier_depends() is used to separate not only two reads but read pointer and WRITE basing on that pointer (example below) - yes. I just doesn't see any example of this in famous Documentation/memory-barriers.txt and had no chance to know what you use it in this way too. > The reason for this is that smp_read_barrier_depends() must order the > pointer load against any subsequent read or write through a dereference > of that pointer. I can't see that requirement anywhere in Documents directory. I mean - the words "write through a dereference of that pointer" or similar for smp_read_barrier_depends. > For example: > > p = READ_ONCE(gp); > smp_rmb(); > r1 = p->a; /* ordered by smp_rmb(). */ > p->b = 42; /* NOT ordered by smp_rmb(), BUG!!! */ > r2 = x; /* ordered by smp_rmb(), but doesn't need to be. */ > > In contrast: > > p = READ_ONCE(gp); > smp_read_barrier_depends(); > r1 = p->a; /* ordered by smp_read_barrier_depends(). */ > p->b = 42; /* ordered by smp_read_barrier_depends(). */ > r2 = x; /* not ordered by smp_read_barrier_depends(), which is OK. */ > > Again, if your hardware maintains local ordering for address > and data dependencies, you can have read_barrier_depends() and > smp_read_barrier_depends() be no-ops like they are for most > architectures. It is not so simple, I mean "local ordering for address and data dependencies". Local ordering is NOT enough. It happens that current MIPS R6 doesn't require in your example smp_read_barrier_depends() but in discussion it comes out that it may not. Because without smp_read_barrier_depends() your example can be a part of Will's WRC+addr+addr and we found some design which easily can bump into this test. And that design actually performs "local ordering for address and data dependencies" too. - Leonid.
[toc] | [prev] | [next] | [standalone]
Page 2 of 5 — ← Prev page 1 [2] 3 4 5 Next page →
Back to top | Article view | linux.kernel
csiph-web