Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1260937 > unrolled thread

Re: [PATCH 3/4] x86,asm: Re-work smp_store_mb()

Started byDavidlohr Bueso <dave@stgolabs.net>
First post2015-11-02 21:20 +0100
Last post2015-11-03 02:40 +0100
Articles 3 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH 3/4] x86,asm: Re-work smp_store_mb() Davidlohr Bueso <dave@stgolabs.net> - 2015-11-02 21:20 +0100
    Re: [PATCH 3/4] x86,asm: Re-work smp_store_mb() Linus Torvalds <torvalds@linux-foundation.org> - 2015-11-03 01:10 +0100
      Re: [PATCH 3/4] x86,asm: Re-work smp_store_mb() Davidlohr Bueso <dave@stgolabs.net> - 2015-11-03 02:40 +0100

#1260937 — Re: [PATCH 3/4] x86,asm: Re-work smp_store_mb()

FromDavidlohr Bueso <dave@stgolabs.net>
Date2015-11-02 21:20 +0100
SubjectRe: [PATCH 3/4] x86,asm: Re-work smp_store_mb()
Message-ID<qqtKa-1nc-21@gated-at.bofh.it>
On Tue, 27 Oct 2015, Peter Zijlstra wrote:

>On Wed, Oct 28, 2015 at 06:33:56AM +0900, Linus Torvalds wrote:
>> On Wed, Oct 28, 2015 at 4:53 AM, Davidlohr Bueso <dave@stgolabs.net> wrote:
>> >
>> > Note that this might affect callers that could/would rely on the
>> > atomicity semantics, but there are no guarantees of that for
>> > smp_store_mb() mentioned anywhere, plus most archs use this anyway.
>> > Thus we continue to be consistent with the memory-barriers.txt file,
>> > and more importantly, maintain the semantics of the smp_ nature.
>>
>
>> So with this patch, the whole thing becomes pointless, I feel. (Ok, so
>> it may have been pointless before too, but at least before this patch
>> it generated special code, now it doesn't). So why carry it along at
>> all?
>
>So I suppose this boils down to if: XCHG ends up being cheaper than
>MOV+FENCE.

So I ran some experiments on an IvyBridge (2.8GHz) and the cost of XCHG is
constantly cheaper (by at least half the latency) than MFENCE. While there
was a decent amount of variation, this difference remained rather constant.

Then again, I'm not sure this matters. Thoughts?

Thanks,
Davidlohr
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1261070

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2015-11-03 01:10 +0100
Message-ID<qqxkK-3Fq-13@gated-at.bofh.it>
In reply to#1260937
On Mon, Nov 2, 2015 at 12:15 PM, Davidlohr Bueso <dave@stgolabs.net> wrote:
>
> So I ran some experiments on an IvyBridge (2.8GHz) and the cost of XCHG is
> constantly cheaper (by at least half the latency) than MFENCE. While there
> was a decent amount of variation, this difference remained rather constant.

Mind testing "lock addq $0,0(%rsp)" instead of mfence? That's what we
use on old cpu's without one (ie 32-bit).

I'm not actually convinced that mfence is necessarily a good idea. I
could easily see it being microcode, for example.

At least on my Haswell, the "lock addq" is pretty much exactly half
the cost of "mfence".

                     Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1261112

FromDavidlohr Bueso <dave@stgolabs.net>
Date2015-11-03 02:40 +0100
Message-ID<qqyJQ-4pO-1@gated-at.bofh.it>
In reply to#1261070
On Mon, 02 Nov 2015, Linus Torvalds wrote:

>On Mon, Nov 2, 2015 at 12:15 PM, Davidlohr Bueso <dave@stgolabs.net> wrote:
>>
>> So I ran some experiments on an IvyBridge (2.8GHz) and the cost of XCHG is
>> constantly cheaper (by at least half the latency) than MFENCE. While there
>> was a decent amount of variation, this difference remained rather constant.
>
>Mind testing "lock addq $0,0(%rsp)" instead of mfence? That's what we
>use on old cpu's without one (ie 32-bit).

I'm getting results very close to xchg.

>I'm not actually convinced that mfence is necessarily a good idea. I
>could easily see it being microcode, for example.

Interesting.

>
>At least on my Haswell, the "lock addq" is pretty much exactly half
>the cost of "mfence".

Ok, his coincides with my results on IvB.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web