Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #25546 > unrolled thread

RAFTS-like optimiser; progress report

Started byAlex McDonald <blog@rivadpm.com>
First post2013-09-06 14:12 -0700
Last post2013-09-07 03:32 -0700
Articles 20 on this page of 63 — 10 participants

Back to article view | Back to comp.lang.forth


Contents

  RAFTS-like optimiser; progress report Alex McDonald <blog@rivadpm.com> - 2013-09-06 14:12 -0700
    Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-06 16:54 -0500
      Re: RAFTS-like optimiser; progress report Alex McDonald <blog@rivadpm.com> - 2013-09-06 16:02 -0700
        Re: RAFTS-like optimiser; progress report Bernd Paysan <bernd.paysan@gmx.de> - 2013-09-07 03:01 +0200
          Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-07 03:21 -0500
            Re: RAFTS-like optimiser; progress report albert@spenarnc.xs4all.nl (Albert van der Horst) - 2013-09-07 19:56 +0000
      Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-07 15:06 +0000
        Re: RAFTS-like optimiser; progress report "Alex McDonald" <blog@rivadpm.com> - 2013-09-07 19:01 +0100
          Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-09 12:23 +0000
            Re: RAFTS-like optimiser; progress report "Alex McDonald" <blog@rivadpm.com> - 2013-09-10 12:01 +0100
              Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-10 12:45 +0000
                Re: RAFTS-like optimiser; progress report "Alex McDonald" <blog@rivadpm.com> - 2013-09-10 16:33 +0100
                  Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-11 15:02 +0000
                    Re: RAFTS-like optimiser; progress report "Alex McDonald" <blog@rivadpm.com> - 2013-09-11 17:44 +0100
                      Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-13 11:58 +0000
                        Re: RAFTS-like optimiser; progress report "Rod Pemberton" <dont_use_email@nohavenotit.com> - 2013-09-14 22:47 -0400
                          Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-15 16:05 +0000
        Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-07 16:08 -0500
          Re: RAFTS-like optimiser; progress report mhx@iae.nl - 2013-09-08 03:21 -0700
            Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-08 06:42 -0500
              Re: RAFTS-like optimiser; progress report mhx@iae.nl - 2013-09-08 04:51 -0700
                Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-08 07:02 -0500
                  Re: RAFTS-like optimiser; progress report Paul Rubin <no.email@nospam.invalid> - 2013-09-08 10:08 -0700
                    Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-08 12:47 -0500
                      Re: RAFTS-like optimiser; progress report Paul Rubin <no.email@nospam.invalid> - 2013-09-08 11:10 -0700
                        Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-08 15:16 -0500
                          Re: RAFTS-like optimiser; progress report Paul Rubin <no.email@nospam.invalid> - 2013-09-08 16:08 -0700
                            Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-09 04:57 -0500
          Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-09 12:41 +0000
            Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-09 08:56 -0500
              Re: RAFTS-like optimiser; progress report albert@spenarnc.xs4all.nl (Albert van der Horst) - 2013-09-09 14:20 +0000
                Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-09 09:27 -0500
                  Re: RAFTS-like optimiser; progress report "Rod Pemberton" <dont_use_email@nohavenotit.com> - 2013-09-10 04:31 -0400
                    Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-10 05:11 -0500
              Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-09 15:08 +0000
                Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-09 10:45 -0500
                  Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-11 15:21 +0000
                    Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-11 12:04 -0500
                      Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-11 17:46 +0000
                        Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-11 14:45 -0500
                          Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-13 13:00 +0000
                            Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-13 10:46 -0500
                            Re: RAFTS-like optimiser; progress report stephenXXX@mpeforth.com (Stephen Pelc) - 2013-09-13 17:44 +0000
                              Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-14 15:31 +0000
                                Re: RAFTS-like optimiser; progress report Paul Rubin <no.email@nospam.invalid> - 2013-09-14 10:33 -0700
                                  Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-14 12:39 -0500
                                    Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-15 15:05 +0000
                                      Re: RAFTS-like optimiser; progress report Alex McDonald <blog@rivadpm.com> - 2013-09-15 08:32 -0700
                                        Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-15 13:20 -0500
                                      Re: RAFTS-like optimiser; progress report stephenXXX@mpeforth.com (Stephen Pelc) - 2013-09-15 16:54 +0000
                                        Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-15 17:02 +0000
                                      Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-15 12:54 -0500
                                        Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-16 09:21 +0000
                                          Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-16 08:39 -0500
                                            Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-16 15:48 +0000
                                              Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-16 15:03 -0500
                                                Re: RAFTS-like optimiser; progress report anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-09-17 10:15 +0000
                                                  Re: RAFTS-like optimiser; progress report "Rod Pemberton" <dont_use_email@nohavenotit.com> - 2013-09-17 17:24 -0400
                                              Re: RAFTS-like optimiser; progress report stephenXXX@mpeforth.com (Stephen Pelc) - 2013-09-17 13:10 +0000
              Re: RAFTS-like optimiser; progress report "Rod Pemberton" <dont_use_email@nohavenotit.com> - 2013-09-10 04:30 -0400
                Re: RAFTS-like optimiser; progress report Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-09-10 05:14 -0500
    Re: RAFTS-like optimiser; progress report mhx@iae.nl - 2013-09-07 03:22 -0700
    Re: RAFTS-like optimiser; progress report mhx@iae.nl - 2013-09-07 03:32 -0700

Page 2 of 4 — ← Prev page 1 [2] 3 4  Next page →


#25573

Frommhx@iae.nl
Date2013-09-08 04:51 -0700
Message-ID<09402346-cb08-4c1d-9c2f-83502c1468f4@googlegroups.com>
In reply to#25572
On Sunday, September 8, 2013 1:42:05 PM UTC+2, Andrew Haley wrote:
> mhx@iae.nl wrote:
[..]
>>> Well, then you have to wonder about visiblity to other threads.  Do
>>> you really want every store to have a store barrier to ensure that it
>>> actually becomes visible to other threads?  There's a performance
>>> penalty to pay.  There are no realy easy answers. 
>>
>> Which Forth has this problem?
> 
> Any Forth that does dead store elimination.

You mentioned threads in a Forth context. Which Forth did you mean?
What threads model are you thinking of (that would have [special]
problems)?

-marcel

-marcel

[toc] | [prev] | [next] | [standalone]


#25574

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-08 07:02 -0500
Message-ID<RPidnbnRBbXs-rHPnZ2dnUVZ_uudnZ2d@supernews.com>
In reply to#25573
mhx@iae.nl wrote:
> On Sunday, September 8, 2013 1:42:05 PM UTC+2, Andrew Haley wrote:
>> mhx@iae.nl wrote:
> [..]
>>>> Well, then you have to wonder about visiblity to other threads.  Do
>>>> you really want every store to have a store barrier to ensure that it
>>>> actually becomes visible to other threads?  There's a performance
>>>> penalty to pay.  There are no realy easy answers. 
>>>
>>> Which Forth has this problem?
>> 
>> Any Forth that does dead store elimination.
> 
> You mentioned threads in a Forth context.

Yes.

> Which Forth did you mean?

Any Forth that does dead store elimination.  I don't think I can be
any clearer about it.

> What threads model are you thinking of (that would have [special]
> problems)?

Any model of threads running on a shared-memory multiprocessor;
there's nothing spcific to any Forth here.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25578

FromPaul Rubin <no.email@nospam.invalid>
Date2013-09-08 10:08 -0700
Message-ID<7xsixfnsvx.fsf@ruckus.brouhaha.com>
In reply to#25574
> Any model of threads running on a shared-memory multiprocessor;
> there's nothing spcific to any Forth here.

Does anyone actually use Forth in environments like that?

[toc] | [prev] | [next] | [standalone]


#25579

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-08 12:47 -0500
Message-ID<HOydndCfjd-7JbHPnZ2dnUVZ_qmdnZ2d@supernews.com>
In reply to#25578
Paul Rubin <no.email@nospam.invalid> wrote:
>> Any model of threads running on a shared-memory multiprocessor;
>> there's nothing spcific to any Forth here.
> 
> Does anyone actually use Forth in environments like that?

Yes.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25580

FromPaul Rubin <no.email@nospam.invalid>
Date2013-09-08 11:10 -0700
Message-ID<7xob83dw1e.fsf@ruckus.brouhaha.com>
In reply to#25579
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>> Any model of threads running on a shared-memory multiprocessor;
>> Does anyone actually use Forth in environments like that?
> Yes.

Yikes ;-).  I'd want to use your STM library from a while back.  Would
it work?  Would it help?  I guess you'd use special STM update
operations rather than ordinary stores, so that would sidestep the issue
of memory barrier words.

[toc] | [prev] | [next] | [standalone]


#25583

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-08 15:16 -0500
Message-ID<haadnV0KaO6RRrHPnZ2dnUVZ_oqdnZ2d@supernews.com>
In reply to#25580
Paul Rubin <no.email@nospam.invalid> wrote:
> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>> Any model of threads running on a shared-memory multiprocessor;
>>> Does anyone actually use Forth in environments like that?
>> Yes.
> 
> Yikes ;-).  I'd want to use your STM library from a while back.  Would
> it work?

Certainly,

> Would it help?

Maybe.  But Forth has been a concurrnt language since the 1970s, so I
have to admit that it's not necessary.  :-)

Andrew.

[toc] | [prev] | [next] | [standalone]


#25584

FromPaul Rubin <no.email@nospam.invalid>
Date2013-09-08 16:08 -0700
Message-ID<7xk3iqewsw.fsf@ruckus.brouhaha.com>
In reply to#25583
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>> STM
> Maybe.  But Forth has been a concurrnt language since the 1970s, so I
> have to admit that it's not necessary.  :-)

I thought Forth traditionally used cooperative multitasking.  I didn't
realize preemptive tasking or actual parallel multiprocessing Forths had
a significant presence.  They (just like anything comparable) would have
been susceptible to the usual concurrency hazards that STM is supposed
to sidestep.  Yes people wrote that sort of code in the 1970s but they
had a terrible time of it.

[toc] | [prev] | [next] | [standalone]


#25585

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-09 04:57 -0500
Message-ID<isydnalKqY0QBrDPnZ2dnUVZ_t6dnZ2d@supernews.com>
In reply to#25584
Paul Rubin <no.email@nospam.invalid> wrote:
> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>> STM
>> Maybe.  But Forth has been a concurrnt language since the 1970s, so I
>> have to admit that it's not necessary.  :-)
> 
> I thought Forth traditionally used cooperative multitasking.

Indeed it did.

> I didn't realize preemptive tasking or actual parallel
> multiprocessing Forths had a significant presence.

That's every Forth on a PC these days and most of them for the last
decade or so.

> They (just like anything comparable) would have been susceptible to
> the usual concurrency hazards that STM is supposed to sidestep.

Indeed.  So everyone has to cope with it in their own way.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25592

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-09-09 12:41 +0000
Message-ID<2013Sep9.144131@mips.complang.tuwien.ac.at>
In reply to#25559
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>Alex McDonald <blog@rivadpm.com> wrote:
>>>> 2. Elimination of redundant stores, such as 10 y1 ! 20 y1 !
>>>
>>>And at that point you'll have to decide what to do about volatile.
>> 
>> On the first order I would not let the compiler eliminate any @s and
>> !s.  If the programmer wants to do that he can do it in the source code.
>
>Well, then you have to wonder about visiblity to other threads.  Do
>you really want every store to have a store barrier to ensure that it
>actually becomes visible to other threads?  There's a performance
>penalty to pay.

My view is that what some processor makers are doing wrt shared memory
is totally impractical for general consumption.  It's the
supercomputer mindset: Let's do minimal hardware and leave the
headaches to the programmers.  This approach has failed to work for
most programmers repeatedly (e.g., look at the Cell), and it will fail
when it comes to shared memory, too.  I see two possible outcomes:

1) Hardware will become capable enough to support sequential
consistency or something close enough to it (and the barriers will be
cheap, because they are noops).  That's a common development in
hardware: First a new feature appears and is designed for minimizing
hardware, later it is implemented in a more usable way.  Examples:
trap barrier (trapb) on Alpha; alignment restrictions on SSE and
SSE-based SIMD extensions.

2) Communicating between threads with shared memory will stay a topic
for specialists only.  Mere mortals will use some higher-level
libraries to communicate between threads.  The specialists will use
the mechanisms the hardware provides (like barriers) instead of using
C-inspired abstractions like volatile.

What does that mean for the question you raise?  What should a Forth
system implementor do?

If we bet on outcome 1, we just generate a barrier after
every @ and !, and on good hardware, these barriers will cost little.

If we bet on outcome 2, we don't generate barriers automatically, but
provide ways for specialists to use them explicitly.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#25594

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-09 08:56 -0500
Message-ID<DKydncrylYLrTrDPnZ2dnUVZ_gWdnZ2d@supernews.com>
In reply to#25592
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>>Alex McDonald <blog@rivadpm.com> wrote:
>>>>> 2. Elimination of redundant stores, such as 10 y1 ! 20 y1 !
>>>>
>>>>And at that point you'll have to decide what to do about volatile.
>>> 
>>> On the first order I would not let the compiler eliminate any @s and
>>> !s.  If the programmer wants to do that he can do it in the source code.
>>
>>Well, then you have to wonder about visiblity to other threads.  Do
>>you really want every store to have a store barrier to ensure that it
>>actually becomes visible to other threads?  There's a performance
>>penalty to pay.
> 
> My view is that what some processor makers are doing wrt shared
> memory is totally impractical for general consumption.  It's the
> supercomputer mindset: Let's do minimal hardware and leave the
> headaches to the programmers.  This approach has failed to work for
> most programmers repeatedly (e.g., look at the Cell), and it will
> fail when it comes to shared memory, too.

Well, it's not failed yet.

> I see two possible outcomes:
> 
> 1) Hardware will become capable enough to support sequential
> consistency or something close enough to it (and the barriers will be
> cheap, because they are noops).  That's a common development in
> hardware: First a new feature appears and is designed for minimizing
> hardware, later it is implemented in a more usable way.  Examples:
> trap barrier (trapb) on Alpha; alignment restrictions on SSE and
> SSE-based SIMD extensions.
> 
> 2) Communicating between threads with shared memory will stay a topic
> for specialists only.  Mere mortals will use some higher-level
> libraries to communicate between threads.  The specialists will use
> the mechanisms the hardware provides (like barriers) instead of using
> C-inspired abstractions like volatile.

Or some combination of the two.

Java's volatile is, in my view, the best model to follow: volatile
accesses generate barriers, non-volatile accesses don't.  C's volatile
is useless.

> What does that mean for the question you raise?  What should a Forth
> system implementor do?
> 
> If we bet on outcome 1, we just generate a barrier after
> every @ and !, and on good hardware, these barriers will cost little.
> 
> If we bet on outcome 2, we don't generate barriers automatically, but
> provide ways for specialists to use them explicitly.

I bet on 1.5 .

Andrew.

[toc] | [prev] | [next] | [standalone]


#25595

Fromalbert@spenarnc.xs4all.nl (Albert van der Horst)
Date2013-09-09 14:20 +0000
Message-ID<522dd918$0$3195$e4fe514c@dreader36.news.xs4all.nl>
In reply to#25594
In article <DKydncrylYLrTrDPnZ2dnUVZ_gWdnZ2d@supernews.com>,
Andrew Haley  <andrew29@littlepinkcloud.invalid> wrote:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>>>Alex McDonald <blog@rivadpm.com> wrote:
>>>>>> 2. Elimination of redundant stores, such as 10 y1 ! 20 y1 !
>>>>>
>>>>>And at that point you'll have to decide what to do about volatile.
>>>>
>>>> On the first order I would not let the compiler eliminate any @s and
>>>> !s.  If the programmer wants to do that he can do it in the source code.
>>>
>>>Well, then you have to wonder about visiblity to other threads.  Do
>>>you really want every store to have a store barrier to ensure that it
>>>actually becomes visible to other threads?  There's a performance
>>>penalty to pay.
>>
>> My view is that what some processor makers are doing wrt shared
>> memory is totally impractical for general consumption.  It's the
>> supercomputer mindset: Let's do minimal hardware and leave the
>> headaches to the programmers.  This approach has failed to work for
>> most programmers repeatedly (e.g., look at the Cell), and it will
>> fail when it comes to shared memory, too.
>
>Well, it's not failed yet.
>
>> I see two possible outcomes:
>>
>> 1) Hardware will become capable enough to support sequential
>> consistency or something close enough to it (and the barriers will be
>> cheap, because they are noops).  That's a common development in
>> hardware: First a new feature appears and is designed for minimizing
>> hardware, later it is implemented in a more usable way.  Examples:
>> trap barrier (trapb) on Alpha; alignment restrictions on SSE and
>> SSE-based SIMD extensions.
>>
>> 2) Communicating between threads with shared memory will stay a topic
>> for specialists only.  Mere mortals will use some higher-level
>> libraries to communicate between threads.  The specialists will use
>> the mechanisms the hardware provides (like barriers) instead of using
>> C-inspired abstractions like volatile.
>
>Or some combination of the two.
>
>Java's volatile is, in my view, the best model to follow: volatile
>accesses generate barriers, non-volatile accesses don't.  C's volatile
>is useless.

Interesting. Can you elaborate? I would like to have an idea in order
to apply it to Forth.

>Andrew.

Groetjes Albert
-- 
Albert van der Horst, UTRECHT,THE NETHERLANDS
Economic growth -- being exponential -- ultimately falters.
albert@spe&ar&c.xs4all.nl &=n http://home.hccnet.nl/a.w.m.van.der.horst

[toc] | [prev] | [next] | [standalone]


#25596

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-09 09:27 -0500
Message-ID<JJ6dnRXb79d0R7DPnZ2dnUVZ_jSdnZ2d@supernews.com>
In reply to#25595
Albert van der Horst <albert@spenarnc.xs4all.nl> wrote:
> In article <DKydncrylYLrTrDPnZ2dnUVZ_gWdnZ2d@supernews.com>,
> Andrew Haley  <andrew29@littlepinkcloud.invalid> wrote:
>>
>>Java's volatile is, in my view, the best model to follow: volatile
>>accesses generate barriers, non-volatile accesses don't.  C's volatile
>>is useless.
> 
> Interesting. Can you elaborate? I would like to have an idea in order
> to apply it to Forth.

There are many pages about this, but here is a good summary:

http://software.intel.com/en-us/blogs/2007/11/30/volatile-almost-useless-for-multi-threaded-programming

Andrew.

[toc] | [prev] | [next] | [standalone]


#25605

From"Rod Pemberton" <dont_use_email@nohavenotit.com>
Date2013-09-10 04:31 -0400
Message-ID<op.w26smeks0e5s1z@->
In reply to#25596
On Mon, 09 Sep 2013 10:27:53 -0400, Andrew Haley  
<andrew29@littlepinkcloud.invalid> wrote:

> Albert van der Horst <albert@spenarnc.xs4all.nl> wrote:
>> In article <DKydncrylYLrTrDPnZ2dnUVZ_gWdnZ2d@supernews.com>,
>> Andrew Haley  <andrew29@littlepinkcloud.invalid> wrote:

>>> Java's volatile is, in my view, the best model to follow:
>>> volatile accesses generate barriers, non-volatile accesses
>>> don't.  C's volatile> is useless.
>>
>> Interesting. Can you elaborate? I would like to have an idea
>> in order to apply it to Forth.
>
> There are many pages about this, but here is a good summary:
>
> [link]
>

Are you sure you're using "volatile" for the correct purpose?
The article you cite in another post is for multi-threading.
C is single-threaded by design ...  "volatile" is not a catch-all
for anything outside C's scope, like locking, semaphores, etc.
It's just used to ensure read or write to a memory mapped device.

The article cited lists three uses of volatile.

The 2nd, being the "externally modified", is standard use for
volatile.  This is generally for memory mapped devices.

The 3rd, for the "signal handler" could be considered to be the
same issue as the 2nd, if the data is updated due to a hardware
interrupt, or updated by code other than the C application, e.g.,
 from the operating system.

The 1st, preserving a variable in setjmp()'s scope, seems rather
strange to me.  "volatile" will prevent optimization, not necessarily
preserve a value.  setjmp() usually just saves the register state
at that point into an array.  Restoring the registers via longjmp()
will usually reset the stack frame as well since the stack pointers
are commonly stored in registers, and most C's use a stack.
longjmp() doesn't unroll anything else.  Execution resumes after
setjmp().  In this case, I can only guess that this is used to
prevent the compiler from using a register, instead of the normal
use of ensuring a read or write to memory.  If the variable was in
a register, it'd be overwritten by the register restore via longjmp().


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#25606

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-10 05:11 -0500
Message-ID<s9SdnX31VoLGbbPPnZ2dnUVZ_tOdnZ2d@supernews.com>
In reply to#25605
Rod Pemberton <dont_use_email@nohavenotit.com> wrote:
> On Mon, 09 Sep 2013 10:27:53 -0400, Andrew Haley  
> <andrew29@littlepinkcloud.invalid> wrote:
> 
>> Albert van der Horst <albert@spenarnc.xs4all.nl> wrote:
>>> In article <DKydncrylYLrTrDPnZ2dnUVZ_gWdnZ2d@supernews.com>,
>>> Andrew Haley  <andrew29@littlepinkcloud.invalid> wrote:
> 
>>>> Java's volatile is, in my view, the best model to follow:
>>>> volatile accesses generate barriers, non-volatile accesses
>>>> don't.  C's volatile> is useless.
>>>
>>> Interesting. Can you elaborate? I would like to have an idea
>>> in order to apply it to Forth.
>>
>> There are many pages about this, but here is a good summary:
>>
>> [link]
>>
> 
> Are you sure you're using "volatile" for the correct purpose?
> The article you cite in another post is for multi-threading.

Which, bizarrely, is the subject under discussion.  Who'd have though
it?

> C is single-threaded by design ... 

Hmmm.

> "volatile" is not a catch-all for anything outside C's scope, like
> locking, semaphores, etc.  It's just used to ensure read or write to
> a memory mapped device.

That's right.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25598

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-09-09 15:08 +0000
Message-ID<2013Sep9.170811@mips.complang.tuwien.ac.at>
In reply to#25594
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>>>Alex McDonald <blog@rivadpm.com> wrote:
>>>>>> 2. Elimination of redundant stores, such as 10 y1 ! 20 y1 !
>>>>>
>>>>>And at that point you'll have to decide what to do about volatile.
>>>> 
>>>> On the first order I would not let the compiler eliminate any @s and
>>>> !s.  If the programmer wants to do that he can do it in the source code.
>>>
>>>Well, then you have to wonder about visiblity to other threads.  Do
>>>you really want every store to have a store barrier to ensure that it
>>>actually becomes visible to other threads?  There's a performance
>>>penalty to pay.
>> 
>> My view is that what some processor makers are doing wrt shared
>> memory is totally impractical for general consumption.  It's the
>> supercomputer mindset: Let's do minimal hardware and leave the
>> headaches to the programmers.  This approach has failed to work for
>> most programmers repeatedly (e.g., look at the Cell), and it will
>> fail when it comes to shared memory, too.
>
>Well, it's not failed yet.

I would say it is failing all the time.  Most programs (including
mine, apart from pipelines) are still single-threaded; there are some
multi-threaded programs, but stuff like lockless shared-memory access
is still black magic that hardly anybody does; AFAIK there's not even
a textbook for that stuff.

>> I see two possible outcomes:
>> 
>> 1) Hardware will become capable enough to support sequential
>> consistency or something close enough to it (and the barriers will be
>> cheap, because they are noops).  That's a common development in
>> hardware: First a new feature appears and is designed for minimizing
>> hardware, later it is implemented in a more usable way.  Examples:
>> trap barrier (trapb) on Alpha; alignment restrictions on SSE and
>> SSE-based SIMD extensions.
>> 
>> 2) Communicating between threads with shared memory will stay a topic
>> for specialists only.  Mere mortals will use some higher-level
>> libraries to communicate between threads.  The specialists will use
>> the mechanisms the hardware provides (like barriers) instead of using
>> C-inspired abstractions like volatile.
>
>Or some combination of the two.

Yes, it is likely, that, if processor makers make their hardware
better, the number of programmers who can implement lockless code
correctly rises from <1000 to maybe 100000, which will still be a
minority of programmers, and most programmers will still use
higher-level libraries.

And I think that the processor makers will make their hardware better,
because they want to sell multi-cores, and the more programmers are
out there that can make good use of them, the more applications will
there be that are an argument for buying them.

>> What does that mean for the question you raise?  What should a Forth
>> system implementor do?
>> 
>> If we bet on outcome 1, we just generate a barrier after
>> every @ and !, and on good hardware, these barriers will cost little.
>> 
>> If we bet on outcome 2, we don't generate barriers automatically, but
>> provide ways for specialists to use them explicitly.
>
>I bet on 1.5 .

The Forth implementation generates barriers for half of the
accesses?-)

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#25599

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-09 10:45 -0500
Message-ID<_vmdnVjAG8aecLDPnZ2dnUVZ_oudnZ2d@supernews.com>
In reply to#25598
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>>>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>>>>Alex McDonald <blog@rivadpm.com> wrote:
>>>>>>> 2. Elimination of redundant stores, such as 10 y1 ! 20 y1 !
>>>>>>
>>>>>>And at that point you'll have to decide what to do about volatile.
>>>>> 
>>>>> On the first order I would not let the compiler eliminate any @s and
>>>>> !s.  If the programmer wants to do that he can do it in the source code.
>>>>
>>>>Well, then you have to wonder about visibility to other threads.  Do
>>>>you really want every store to have a store barrier to ensure that it
>>>>actually becomes visible to other threads?  There's a performance
>>>>penalty to pay.
>>> 
>>> My view is that what some processor makers are doing wrt shared
>>> memory is totally impractical for general consumption.  It's the
>>> supercomputer mindset: Let's do minimal hardware and leave the
>>> headaches to the programmers.  This approach has failed to work for
>>> most programmers repeatedly (e.g., look at the Cell), and it will
>>> fail when it comes to shared memory, too.
>>
>>Well, it's not failed yet.
> 
> I would say it is failing all the time.  Most programs (including
> mine, apart from pipelines) are still single-threaded; there are some
> multi-threaded programs,

Indeed there are.  It depends on the quality of language support for
threads: some programming languages support it well, and programs are
commonly multi-threaded, and libraries use multiple threads when
they're there.

> but stuff like lockless shared-memory access is still black magic
> that hardly anybody does; AFAIK there's not even a textbook for that
> stuff.

The Art of Multiprocessor Programming, Maurice Herlihy, Nir Shavit,
MK 2008, 2012.

I suppose all this depends on whether you see the glass as half full
or half empty.  Programming with multiple threads is what I do all the
time, so it's the world I know.

>>> I see two possible outcomes:
>>> 
>>> 1) Hardware will become capable enough to support sequential
>>> consistency or something close enough to it (and the barriers will be
>>> cheap, because they are noops).  That's a common development in
>>> hardware: First a new feature appears and is designed for minimizing
>>> hardware, later it is implemented in a more usable way.  Examples:
>>> trap barrier (trapb) on Alpha; alignment restrictions on SSE and
>>> SSE-based SIMD extensions.
>>> 
>>> 2) Communicating between threads with shared memory will stay a topic
>>> for specialists only.  Mere mortals will use some higher-level
>>> libraries to communicate between threads.  The specialists will use
>>> the mechanisms the hardware provides (like barriers) instead of using
>>> C-inspired abstractions like volatile.
>>
>>Or some combination of the two.
> 
> Yes, it is likely, that, if processor makers make their hardware
> better, the number of programmers who can implement lockless code
> correctly rises from <1000 to maybe 100000, which will still be a
> minority of programmers, and most programmers will still use
> higher-level libraries.

The hardware isn't going to change wrt memory consistency, from what
I've seen of newish designs.  What we are going to get instead,
though, is hardware support for transactions.  Then you don't need to
support memory consistency where it isn't necessary.

> And I think that the processor makers will make their hardware
> better, because they want to sell multi-cores, and the more
> programmers are out there that can make good use of them, the more
> applications will there be that are an argument for buying them.

For some value of "better", yes, but not in the way you're suggesting.
Maintaining a consistent view of memory all the time across all cores
is too expensive.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25629

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-09-11 15:21 +0000
Message-ID<2013Sep11.172119@mips.complang.tuwien.ac.at>
In reply to#25599
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Yes, it is likely, that, if processor makers make their hardware
>> better, the number of programmers who can implement lockless code
>> correctly rises from <1000 to maybe 100000, which will still be a
>> minority of programmers, and most programmers will still use
>> higher-level libraries.
>
>The hardware isn't going to change wrt memory consistency, from what
>I've seen of newish designs.  What we are going to get instead,
>though, is hardware support for transactions.  Then you don't need to
>support memory consistency where it isn't necessary.
>
>> And I think that the processor makers will make their hardware
>> better, because they want to sell multi-cores, and the more
>> programmers are out there that can make good use of them, the more
>> applications will there be that are an argument for buying them.
>
>For some value of "better", yes, but not in the way you're suggesting.
>Maintaining a consistent view of memory all the time across all cores
>is too expensive.

Precise exceptions are too expensive, so Alpha has the (originally
expensive) TRAPB instruction.  Eventually they implemented precise
exceptions, and TRAPB became a noop.  Byte and 16-bit accesses are too
expensive, so Alpha did not implement them.  Eventually they were
added to Alpha.  Cache-coherence shared memory for more than ~16 CPUs
is too expensive, so supercomputers are using private memory nodes
(called "massively parallel processors").  Then directory-based cache
coherence (NUMA) was developed, and last I heard up to 1024 CPUs were
possible in a shared-memory machine (maybe more these days), and even
supercomputers used that.

I am confident that the hardware people can make hardware that has a
reasonable consistency model if they put their minds to it.

The weakly consistent models do not cut it, though; they may be ok for
some academics who have a nice puzzle to solve for writing a paper on
a small problem, but for producing software they are too complex.
Telling the programmers: "Your code will be correct if you use
barriers everywhere.  Your code will be dog-slow if you use barriers
everywhere." does not lead to correct and fast programs.

You may be right, and the hardware guys may find a different way to
give programmers an understandable way to program with shared memory,
and if that way is cheaper to implement than a reasonable consistency
model, then yes, the reasonable consistency model will be too
expensive.  But if the hardware guys tell us: "Your code will be
correct if you use transactions everywhere.  Your code will be
dog-slow if you use transactions everywhere.", then transactions won't
cut it, either.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#25632

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-11 12:04 -0500
Message-ID<kbudnRdw6fElP63PnZ2dnUVZ_gmdnZ2d@supernews.com>
In reply to#25629
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:

> I am confident that the hardware people can make hardware that has a
> reasonable consistency model if they put their minds to it.

And I am confident that they won't.  We'll see.

> The weakly consistent models do not cut it, though; they may be ok
> for some academics who have a nice puzzle to solve for writing a
> paper on a small problem, but for producing software they are too
> complex.

This isn't academic anything, though: it's processors like ARM, which
is arguably more mainstream than anything x86.

The surprising thing is that multi-threaded programs seem to work just
fine on ARM, so I assume that people are using locking correctly.

> Telling the programmers: "Your code will be correct if you use
> barriers everywhere.  Your code will be dog-slow if you use barriers
> everywhere." does not lead to correct and fast programs.
> 
> You may be right, and the hardware guys may find a different way to
> give programmers an understandable way to program with shared
> memory, and if that way is cheaper to implement than a reasonable
> consistency model, then yes, the reasonable consistency model will
> be too expensive.  But if the hardware guys tell us: "Your code will
> be correct if you use transactions everywhere.  Your code will be
> dog-slow if you use transactions everywhere.", then transactions
> won't cut it, either.

Sure, but in-cache hardware transactions aren't slow, so that argument
doesn't apply.

Andrew.

[toc] | [prev] | [next] | [standalone]


#25633

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-09-11 17:46 +0000
Message-ID<2013Sep11.194637@mips.complang.tuwien.ac.at>
In reply to#25632
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>
>> I am confident that the hardware people can make hardware that has a
>> reasonable consistency model if they put their minds to it.
>
>And I am confident that they won't.  We'll see.
>
>> The weakly consistent models do not cut it, though; they may be ok
>> for some academics who have a nice puzzle to solve for writing a
>> paper on a small problem, but for producing software they are too
>> complex.
>
>This isn't academic anything, though: it's processors like ARM, which
>is arguably more mainstream than anything x86.

Mainstream when it comes to shared-memory processing?  How many ARM
servers with 2 sockets have been sold?  How many with 4?  How many
with more?

And ARM gets away with lots of nonsense.  They don't even have a
properly specified architecture (at least I could not find the
specification when I tried some time ago).

>Sure, but in-cache hardware transactions aren't slow, so that argument
>doesn't apply.

On ARMs, i.e., single-CPU-chip machines?  Sure, but proper consistency
would not be slow on that, either.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#25637

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-09-11 14:45 -0500
Message-ID<z4mdnVWD7MH6Va3PnZ2dnUVZ_rKdnZ2d@supernews.com>
In reply to#25633
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>
>>> The weakly consistent models do not cut it, though; they may be ok
>>> for some academics who have a nice puzzle to solve for writing a
>>> paper on a small problem, but for producing software they are too
>>> complex.
>>
>>This isn't academic anything, though: it's processors like ARM, which
>>is arguably more mainstream than anything x86.
> 
> Mainstream when it comes to shared-memory processing?

Sure: all those processors in phones.

> How many ARM servers with 2 sockets have been sold?  How many with
> 4?  How many with more?

What does this have to do with servers or sockets?  It applies just
the same with multi-core processors.

> And ARM gets away with lots of nonsense.  They don't even have a
> properly specified architecture (at least I could not find the
> specification when I tried some time ago).

I don't know why you had that problem: it's at least as well-specified
as anything else I've seen.

>>Sure, but in-cache hardware transactions aren't slow, so that argument
>>doesn't apply.
> 
> On ARMs, i.e., single-CPU-chip machines?

ARM doesn't have that yet; neither does Intel.  It's only SPARC and
IBM at present, as far as the mainstream goes.

Andrew.

[toc] | [prev] | [next] | [standalone]


Page 2 of 4 — ← Prev page 1 [2] 3 4  Next page →

Back to top | Article view | comp.lang.forth


csiph-web