Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #28595 > unrolled thread

Forth reinvention

Started bydambere@web.de
First post2014-02-20 01:19 -0800
Last post2014-03-08 05:31 -0800
Articles 20 on this page of 73 — 24 participants

Back to article view | Back to comp.lang.forth


Contents

  Forth reinvention dambere@web.de - 2014-02-20 01:19 -0800
    Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-02-20 05:07 -0500
      Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-20 14:56 +0000
        Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-02-20 16:16 -0500
          Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-20 23:25 +0000
            Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-02-20 20:42 -0500
              Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-21 12:41 +0000
    Re: Forth reinvention Hans Bezemer <the.beez.speaks@gmail.com> - 2014-02-20 11:18 +0100
      Re: Forth reinvention dambere@web.de - 2014-02-20 02:36 -0800
        Re: Forth reinvention Hans Bezemer <the.beez.speaks@gmail.com> - 2014-02-20 17:35 +0100
          Re: Forth reinvention dambere@web.de - 2014-02-20 11:34 -0800
            Re: Forth reinvention "Elizabeth D. Rather" <erather@forth.com> - 2014-02-20 09:53 -1000
            Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-02-20 12:23 -0800
              Re: Forth reinvention dambere@web.de - 2014-02-20 14:00 -0800
                Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-02-20 14:23 -0800
                  Re: Forth reinvention dambere@web.de - 2014-02-22 02:21 -0800
                    Re: Forth reinvention AKK <akk@nospam.org> - 2014-02-22 12:02 +0100
                Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-21 04:38 -0600
            Re: Forth reinvention Mark Wills <markrobertwills@yahoo.co.uk> - 2014-02-20 12:35 -0800
              Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-02-20 13:38 -0800
    Re: Forth reinvention m.a.m.hendrix@tue.nl - 2014-02-20 03:59 -0800
      Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-20 15:11 +0000
        Re: Forth reinvention m.a.m.hendrix@tue.nl - 2014-02-20 07:41 -0800
    Re: Forth reinvention Richard Owlett <rowlett@pcnetinc.com> - 2014-02-20 06:38 -0600
      Re: Forth reinvention dambere@web.de - 2014-02-20 12:03 -0800
    Re: Forth reinvention Julian Fondren <julian.fondren@gmail.com> - 2014-02-20 07:08 -0800
    Re: Forth reinvention "Elizabeth D. Rather" <erather@forth.com> - 2014-02-20 09:37 -1000
      Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-20 20:50 +0000
        Re: Forth reinvention "Elizabeth D. Rather" <erather@forth.com> - 2014-02-20 15:39 -1000
          Re: Forth reinvention albert@spenarnc.xs4all.nl (Albert van der Horst) - 2014-02-21 10:50 +0000
          Re: Forth reinvention Paul E Bennett <Paul_E.Bennett@topmail.co.uk> - 2014-02-21 11:10 +0000
            Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-21 13:50 +0000
          Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-21 05:46 -0600
            Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-21 12:55 +0000
            Re: Forth reinvention anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2014-02-21 15:11 +0000
              Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-22 09:23 -0600
          Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-21 12:38 +0000
          Re: Forth reinvention anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2014-02-22 13:29 +0000
            Re: Forth reinvention "Elizabeth D. Rather" <erather@forth.com> - 2014-02-22 07:52 -1000
              Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-22 19:06 +0000
              Re: Forth reinvention anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2014-02-23 13:40 +0000
      Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-02-20 20:40 -0500
        Re: Forth reinvention "Elizabeth D. Rather" <erather@forth.com> - 2014-02-20 18:43 -1000
          Re: Forth reinvention Hans Bezemer <the.beez.speaks@gmail.com> - 2014-02-21 13:16 +0100
        Re: Forth reinvention stephenXXX@mpeforth.com (Stephen Pelc) - 2014-02-21 10:49 +0000
      Re: Forth reinvention anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2014-02-21 14:50 +0000
        Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-21 16:28 +0000
          Re: Forth reinvention anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2014-02-22 12:54 +0000
            Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-02-22 10:27 -0500
            Re: Forth reinvention "Alex McDonald" <blog@rivadpm.com> - 2014-02-22 17:49 +0000
        Re: Forth reinvention Bernd Paysan <bernd.paysan@gmx.de> - 2014-02-21 20:51 +0100
          Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-22 09:39 -0600
            Re: Forth reinvention Bernd Paysan <bernd.paysan@gmx.de> - 2014-02-23 02:40 +0100
              Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-02-22 19:24 -0800
                Re: Forth reinvention Bernd Paysan <bernd.paysan@gmx.de> - 2014-02-23 22:48 +0100
              Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-23 04:39 -0600
                Re: Forth reinvention Bernd Paysan <bernd.paysan@gmx.de> - 2014-02-23 22:46 +0100
                  Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-02-23 14:26 -0800
                    Re: Forth reinvention Bernd Paysan <bernd.paysan@gmx.de> - 2014-02-24 02:44 +0100
                      Re: Forth reinvention Spam@ControlQ.com - 2014-03-03 12:43 -0500
                      Re: Forth reinvention Paul Rubin <no.email@nospam.invalid> - 2014-03-03 10:06 -0800
                  Re: Forth reinvention Andrew Haley <andrew29@littlepinkcloud.invalid> - 2014-02-24 03:59 -0600
    Re: Forth reinvention mike73900@gmail.com - 2014-03-03 14:04 -0800
      Re: Forth reinvention mhx@iae.nl - 2014-03-05 06:59 -0800
        Re: Forth reinvention Matthias Koch <matthias.koch@hot.uni-hannover.de> - 2014-03-05 16:33 +0100
        Re: Forth reinvention Mark Wills <markwills1970@gmail.com> - 2014-03-05 09:08 -0800
          Re: Forth reinvention mhx@iae.nl - 2014-03-05 10:47 -0800
    Re: Forth reinvention AKK <akk@nospam.org> - 2014-03-07 07:30 +0100
      Re: Forth reinvention Lars Brinkhoff <lars.spam@nocrew.org> - 2014-03-07 07:55 +0100
      Re: Forth reinvention albert@spenarnc.xs4all.nl (Albert van der Horst) - 2014-03-07 09:29 +0000
        Re: Forth reinvention AKK <akk@nospam.org> - 2014-03-07 12:04 +0100
        Re: Forth reinvention "Rod Pemberton" <dont_use_email@xnothavet.cqm> - 2014-03-07 17:03 -0500
          Re: Forth reinvention Mark Wills <markwills1970@gmail.com> - 2014-03-08 05:31 -0800

Page 3 of 4 — ← Prev page 1 2 [3] 4  Next page →


#28712

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2014-02-23 13:40 +0000
Message-ID<2014Feb23.144040@mips.complang.tuwien.ac.at>
In reply to#28692
"Elizabeth D. Rather" <erather@forth.com> writes:
>On 2/22/14 3:29 AM, Anton Ertl wrote:
>> "Elizabeth D. Rather" <erather@forth.com> writes:
>>> There seem to be lots of multi-core architectures out there nowadays,
>>> but I haven't heard of a "new language" in which everything has to be
>>> re-written in order to run on them.
>>
>> I have heard of many such languages.
>>
>>> If programs written in C, C++, Java,
>>> or whatever are now suddenly running parallel,
>>
>> They are not.  If they have been written multi-threaded, and the
>> threads don't wait too much on each other, they run in parallel,
>> otherwise they just run sequentially.
>
>Where is the "decision" to run them parallel made? At the language 
>level, or the OS level?

The programmer decides whether and how the application is divided into
threads and where these threads wait on each other; he then codes this
into the program with more or less language support.  Then the
run-time environment decides whether it can actually make use of the
parallelism exposed by the program, and may run the different threads
on different cores.  In Java the multi-threading support is present in
the language, while in C it is present through APIs like POSIX threads
that are usually seen as part of the OS.  But in any case, a
sequential program does not suddenly run in parallel (apart from maybe
a few easy cases).

>The style of multitasking in Forth has very 
>little interdependence between threads.

The traditional style of multi-tasking in Forth is cooperative
multitasking.  There each task can rely on the code pieces between two
PAUSEs (and other task-switching words like I/O words) being processed
without another task intervening.  It's hard or impossible for the
compiler and run-time to know whether such a piece of code is
independent from another such piece of code, so these tasks cannot be
processed in parallel (at least I don't see a practical way).  Even if
there is little interdependence, the compiler and run-time system have
no way of knowing that.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#28632

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2014-02-20 20:40 -0500
Message-ID<op.xblyxvfp5zc71u@localhost>
In reply to#28616
On Thu, 20 Feb 2014 14:37:54 -0500, Elizabeth D. Rather  
<erather@forth.com> wrote:
> On 2/19/14 11:19 PM, dambere@web.de wrote:

>> Personally I found RPN an elegant, mathematical notation both solving
>> inconsistencies of traditional infix notations and ease parsing.
>
> RPN is way overstated as an impediment to Forth, because it's an obvious
> feature that's easy to point to. But I really don't think it's a serious
> impediment.

The HP calculators used RPN.

It wasn't an impediment, but more a hindrance or nuisance.

People structure the problems the way they find them easiest to solve.

E.g., after using scientific calculators and software versions for many  
years,
I find it difficult to solve problems on them without a "1/X" button.  I  
recently
realized this after installing Linux, where the Galculator calculator had  
no
"1/X" button.  The only way to solve without "1/X" is to save to memory,  
clear,
enter 1, divide, recall, next operation, recall other saved results, ...   
I.e.,
not having "1/X" is not an impediment, but a nuisance or hindrance.  One  
button
as compared to numerous operations.  Fortunately, there is a patch to fix  
that, at
least for GTK 2 implementations of Galculator, but not for the newer GTK 3  
version ...

> As Hans so succinctly said, stacks are used in many compilers. By
> exposing them (and other internal features) Forth puts more power in the
> hands of programmers. That's a good thing.

Or, it can result in more errors since the user must be more cautious.

It's true stacks are used in many compilers, but they're mostly used as
fast temporary storage for register overflows (too few registers) instead  
of
saving registers to slower memory, and for passing of parameters to  
functions
or procedures.  So, I don't see how these would equate to Forth.  I.e.,  
other
languages use a stack differently than Forth.  Forth is stack-based, but
most microprocessor instruction sets are register based.  I.e., design
conflict between Forth's way of doing things and ability of the processor.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#28636

From"Elizabeth D. Rather" <erather@forth.com>
Date2014-02-20 18:43 -1000
Message-ID<U9ednTiqA9_LRJvOnZ2dnUVZ_sudnZ2d@supernews.com>
In reply to#28632
On 2/20/14 3:40 PM, Rod Pemberton wrote:
> On Thu, 20 Feb 2014 14:37:54 -0500, Elizabeth D. Rather
> <erather@forth.com> wrote:
...
>> As Hans so succinctly said, stacks are used in many compilers. By
>> exposing them (and other internal features) Forth puts more power in the
>> hands of programmers. That's a good thing.
>
> Or, it can result in more errors since the user must be more cautious.

"With great power comes great responsibility." Forth comes without seat 
belts and air bags. But it's very powerful. It's a tool for good 
programmers who are willing to accept responsibility in order to get the 
power and convenience.

Cheers,
Elizabeth

-- 
==================================================
Elizabeth D. Rather   (US & Canada)   800-55-FORTH
FORTH Inc.                         +1 310.999.6784
5959 West Century Blvd. Suite 700
Los Angeles, CA 90045
http://www.forth.com

"Forth-based products and Services for real-time
applications since 1973."
==================================================

[toc] | [prev] | [next] | [standalone]


#28653

FromHans Bezemer <the.beez.speaks@gmail.com>
Date2014-02-21 13:16 +0100
Message-ID<53074399$0$2831$e4fe514c@news2.news.xs4all.nl>
In reply to#28636
Elizabeth D. Rather wrote:

>> Or, it can result in more errors since the user must be more cautious.
> 
> "With great power comes great responsibility." Forth comes without seat
> belts and air bags. But it's very powerful. It's a tool for good
> programmers who are willing to accept responsibility in order to get the
> power and convenience.
I couldn't agree more! I always say to friends who don't know Forth "With C
you can shoot yourself in the foot, with Forth your blow your head straight
off".

If you want a "safe" environment, there is lots more to "improve" on Forth:
no type checking, no parameter checking, even no checking whether the
execution token (function pointer) you execute or address you pop from the
return stack holds a sane value.

In the old days (WAY back in the previous century) even the OS didn't
prevent you from anything, so you blew up your whole system - and
consequently had to reboot. But we're civilized now.. ;-)

But then again all that has a lot of advantages as well:

- Since you're responsible, the compiler doesn't whine on "loss of
significant digits", "not the right type" or anything. I don't like to
punch all kinds of blabla all over the place just to silence compiler
warnings;

- Since there no types you don't have to crack your head on "how do I
declare a pointer to a pointer to an array of functions that take doubles
and returns a pointer to a pointer to an array of pointers to functions
that take a pointer to a matrix of pointers to integers".

- Return of multiple values or variadic functions are a no-brainer.

- The DOES> allows you to change the run time behavior of datatypes, so they
can take care of themselves, e.g. a table that does all lookup by itself.
Sure, you can wrap it into a function, but since it's there, there's always
that little guy in your head that says "Can't I do this smarter". Not to
mention avoiding cluttering your symbol table.

And one more tip here: if you need "blocks", {} that is, you're doing it
wrong! In Forth you don't have pages and pages per word. If I have a ten
line word I'm already wondering where I went wrong. One or two loop
constructs per word is more than enough. E.g. compare my implementation of
getopts() to a C-one (plenty of those around).

I'll probably get fried over this (it's the c.l.f way), first of all because
it isn't ANS, but maybe it gets even better ;-) But it's just to prove a
point:

---8<---
argn value option-index                \ first invalid option
                                       \ argument in next ARGS
: +argument                            ( n1 n2 -- n1+1 n2 a2 n3 )
  swap 1+ tuck argn <                  \ is there still an argument?
  if
    over args                          \ if valid, return argument
  else                                 \ else issue message
    ." Missing argument to option -" dup emit cr 0 dup
  then                                 \ and return empty string
;
                                       \ execute option in table
: get-option                           ( c a --)
  2 num-key row                        \ if found execute option
  if cell+ @c execute drop else drop ." Invalid option -" emit cr then
;                                      \ if not, issue message
                                       ( n1 a1 n2 -- argn a1 n2)
: end-arguments 2>r to option-index argn 2r> ;
                                       \ check a single argument
: check-argument                       ( a -- a)
  >r dup args over c@ [char] - =       \ if it is a valid option
  if                                   \ then evaluate it
    begin chop 2dup 0> while c@ dup while r@ get-option repeat drop
  else
    end-arguments                      \ else terminate scanning
  then 2drop 1+ r>                     \ all scanned, increment index
;
                                       \ get an optional argument
: get-argument                         ( n1 c a1 n2 -- n1 c a2 n3 a3 n4)
  >r chop dup 0> if r> -rot else 2drop r> +argument then 0 dup 2swap
;
                                       ( xt --)
: get-options 1 swap begin over argn < while check-argument repeat drop
drop ;
---8<---

BTW, the only externals are end-arguments, get-argument and get-options.

Hans Bezemer

[toc] | [prev] | [next] | [standalone]


#28645

FromstephenXXX@mpeforth.com (Stephen Pelc)
Date2014-02-21 10:49 +0000
Message-ID<53072e0e.717702068@news.demon.co.uk>
In reply to#28632
On Thu, 20 Feb 2014 20:40:33 -0500, "Rod Pemberton"
<dont_use_email@xnohavenotit.cnm> wrote:

>Forth is stack-based, but
>most microprocessor instruction sets are register based.  I.e., design
>conflict between Forth's way of doing things and ability of the processor.

It's not a conflict, just a difference. The VFX code generator was
designed after reading a lot of the classical compiler literature.
The only algorithm that is specific to Forth is the "canonical stack
shuffle" at various boundaries, and I'm told that similar
algorithms are used in modern C compilers.

Wen we finished the first VFX compiler, we at MPE were surprised
by the performance gains. So much so that we stopped believing in
the benefits of silicon stack machines other than for deterministic
interrupt handling.

Stephen

-- 
Stephen Pelc, stephenXXX@mpeforth.com
MicroProcessor Engineering Ltd - More Real, Less Time
133 Hill Lane, Southampton SO15 5AF, England
tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691
web: http://www.mpeforth.com - free VFX Forth downloads

[toc] | [prev] | [next] | [standalone]


#28659

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2014-02-21 14:50 +0000
Message-ID<2014Feb21.155029@mips.complang.tuwien.ac.at>
In reply to#28616
"Elizabeth D. Rather" <erather@forth.com> writes:
>> - Automatic parallelization:
>> The trend goes to multi core architectures. A programming language that is relieving the programmer from the not inconsiderable task of paralleling programs will necessarily be attractive.
>
>Parallelization is an implementation strategy, not a language feature.

Automatic parallelization would make it an implementation strategy.
It works for some application areas, but not in general, and is quite
complex even when it works.

Therefore we need language features if we want to parallelize our
programs.

But do we want to parallelize our programs?  The highest expected
speedup is by the number of cores, i.e., typically 4.  For many
programs there is lower-hanging fruit that gives a speedup by a factor
of 4, so the case for parallelization does not seem strong to me.  It
only applies to the few programs where the lower-hanging fruit has
already been picked.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#28663

From"Alex McDonald" <blog@rivadpm.com>
Date2014-02-21 16:28 +0000
Message-ID<le7usa$l9g$1@dont-email.me>
In reply to#28659
on 21/02/2014 14:50:29,  wrote:
> "Elizabeth D. Rather" <erather@forth.com> writes:
>>> - Automatic parallelization:
>>> The trend goes to multi core architectures. A programming language that is relieving the programmer 
from the not inconsiderable task of paralleling programs will necessarily
be attractive. >>
>>Parallelization is an implementation strategy, not a language feature.
> 
> Automatic parallelization would make it an implementation strategy.
> It works for some application areas, but not in general, and is quite
> complex even when it works.
> 
> Therefore we need language features if we want to parallelize our
> programs.
> 
> But do we want to parallelize our programs?  The highest expected
> speedup is by the number of cores, i.e., typically 4.  For many
> programs there is lower-hanging fruit that gives a speedup by a factor
> of 4, so the case for parallelization does not seem strong to me. It
> only applies to the few programs where the lower-hanging fruit has
> already been picked.
> 
> - anton

Your example is a task split N ways; that's just one of many possible
descriptions of a "parallel program". Not all parallelisation types will
be desirable, possible or require language features, but stating that the
case for it "doesn't seem strong" may only be true in this specific
example.

And then, only if you can demonstrate that there are lower hanging fruit,
by which I take it you mean easier ways to programmatically get the 4x
speedup. I would be surprised if you could point to an example that
didn't involve an observation as to the general stupidity of programmers
who select inappropriate and dreadful algorithms.

[toc] | [prev] | [next] | [standalone]


#28677

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2014-02-22 12:54 +0000
Message-ID<2014Feb22.135404@mips.complang.tuwien.ac.at>
In reply to#28663
"Alex McDonald" <blog@rivadpm.com> writes:
>on 21/02/2014 14:50:29,  wrote:
>> But do we want to parallelize our programs?  The highest expected
>> speedup is by the number of cores, i.e., typically 4.  For many
>> programs there is lower-hanging fruit that gives a speedup by a factor
>> of 4, so the case for parallelization does not seem strong to me. It
>> only applies to the few programs where the lower-hanging fruit has
>> already been picked.
>> 
>> - anton
>
>Your example is a task split N ways; that's just one of many possible
>descriptions of a "parallel program". Not all parallelisation types will
>be desirable, possible or require language features, but stating that the
>case for it "doesn't seem strong" may only be true in this specific
>example.

I did not give an example.  And I have no idea what you you mean with
'many possible descriptions of a "parallel program"'.

>And then, only if you can demonstrate that there are lower hanging fruit,
>by which I take it you mean easier ways to programmatically get the 4x
>speedup. I would be surprised if you could point to an example that
>didn't involve an observation as to the general stupidity of programmers
>who select inappropriate and dreadful algorithms.

In my course on efficient programs the students often optimize a
program I give them, and in several cases I have given them a program
written by someone else, and speedups by a factor of 4 or more were
not rare.  Were the original programmers stupid?  I don't think so.
In any case, even for relatively small programs, there is a lot of
sequential performance that can be gained, and I expect that there is
more in larger programs.

With regard to stupidity: In 2009 I chose a program for converting
uids into user names (on a Unix system), which is a performance
bottleneck for ls -l on servers with many users.  Some Unix variants
therefore go away from the good old principle of just editing the
passwd file (which contains the mapping), and require to call some
system-specific program afterwards.

In this case, I think the programmers were stupid.  There are lots of
ways to do better without requiring to call that program explicitly;
among them is to call that program automatically when such a mapping
is needed and the passwd file is newer than the mapping file.  A group
of my students that used such an approach got a speedup by a factor of
600 when the cache file is up-to-date (the usual case).

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2013: http://www.euroforth.org/ef13/

[toc] | [prev] | [next] | [standalone]


#28684

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2014-02-22 10:27 -0500
Message-ID<op.xbovvo065zc71u@localhost>
In reply to#28677
On Sat, 22 Feb 2014 07:54:04 -0500, Anton Ertl  
<anton@mips.complang.tuwien.ac.at> wrote:
> "Alex McDonald" <blog@rivadpm.com> writes:

>> And then, only if you can demonstrate that there are lower hanging  
>> fruit,
>> by which I take it you mean easier ways to programmatically get the 4x
>> speedup. I would be surprised if you could point to an example that
>> didn't involve an observation as to the general stupidity of programmers
>> who select inappropriate and dreadful algorithms.
>
> In my course on efficient programs the students often optimize a
> program I give them, and in several cases I have given them a program
> written by someone else, and speedups by a factor of 4 or more were
> not rare.  Were the original programmers stupid?  I don't think so.
> In any case, even for relatively small programs, there is a lot of
> sequential performance that can be gained, and I expect that there is
> more in larger programs.
>
> With regard to stupidity: In 2009 I chose a program for converting
> uids into user names (on a Unix system), which is a performance
> bottleneck for ls -l on servers with many users.  Some Unix variants
> therefore go away from the good old principle of just editing the
> passwd file (which contains the mapping), and require to call some
> system-specific program afterwards.
>
> In this case, I think the programmers were stupid.  There are lots of
> ways to do better without requiring to call that program explicitly;
> among them is to call that program automatically when such a mapping
> is needed and the passwd file is newer than the mapping file.  A group
> of my students that used such an approach got a speedup by a factor of
> 600 when the cache file is up-to-date (the usual case).

Or, you could have "that program" which provides "such a mapping"
be called automatically by the Unix tools which update the password file
when a new user account is created.  Then, "such a mapping" is always
up-to-date, and then there is never a need to "call that program
explicitly".  But, the on-demand or "demand paging" style solution you
posted here works too.

When I started to read the part of the reply on converting uids to
user names, I thought you were going to mention how they unexpectedly
used a more efficient method of integer-to-string conversion or a new
type of hashing.  I.e., anti-climatic.  It seems other things have
become more important over time than speed, e.g., security, convenience.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#28691

From"Alex McDonald" <blog@rivadpm.com>
Date2014-02-22 17:49 +0000
Message-ID<leanur$1ha$1@dont-email.me>
In reply to#28677
on 22/02/2014 12:54:00,  wrote:
> "Alex McDonald" <blog@rivadpm.com> writes:
>>on 21/02/2014 14:50:29,  wrote:
>>> But do we want to parallelize our programs?  The highest expected
>>> speedup is by the number of cores, i.e., typically 4.  For many
>>> programs there is lower-hanging fruit that gives a speedup by a factor
>>> of 4, so the case for parallelization does not seem strong to me. It
>>> only applies to the few programs where the lower-hanging fruit has
>>> already been picked.
>>>
>>> - anton
>>
>>Your example is a task split N ways; that's just one of many possible
>>descriptions of a "parallel program". Not all parallelisation types will
>>be desirable, possible or require language features, but stating that the
>>case for it "doesn't seem strong" may only be true in this specific
>>example.
> 
> I did not give an example. 

Not explicitly; but since you specifically refer to task parallelisation
across cores (and I further assume you mean on a single system), I take
that as an example of that specific technique.


> And I have no idea what you you mean with
> 'many possible descriptions of a "parallel program"'.

SIMD for example. Or distributed execution. Parallel execution across
tightly coupled cores is not the only solution.

> 
>>And then, only if you can demonstrate that there are lower hanging fruit,
>>by which I take it you mean easier ways to programmatically get the 4x
>>speedup. I would be surprised if you could point to an example that
>>didn't involve an observation as to the general stupidity of programmers
>>who select inappropriate and dreadful algorithms.
> 
> In my course on efficient programs the students often optimize a
> program I give them, and in several cases I have given them a program
> written by someone else, and speedups by a factor of 4 or more were
> not rare.  Were the original programmers stupid?  I don't think so.
> In any case, even for relatively small programs, there is a lot of
> sequential performance that can be gained, and I expect that there is
> more in larger programs.
> 
> With regard to stupidity: In 2009 I chose a program for converting
> uids into user names (on a Unix system), which is a performance
> bottleneck for ls -l on servers with many users.  Some Unix variants
> therefore go away from the good old principle of just editing the
> passwd file (which contains the mapping), and require to call some
> system-specific program afterwards.
> 
> In this case, I think the programmers were stupid.  There are lots of
> ways to do better without requiring to call that program explicitly;
> among them is to call that program automatically when such a mapping
> is needed and the passwd file is newer than the mapping file. A group
> of my students that used such an approach got a speedup by a factor of
> 600 when the cache file is up-to-date (the usual case).
> 
> - anton

[toc] | [prev] | [next] | [standalone]


#28668

FromBernd Paysan <bernd.paysan@gmx.de>
Date2014-02-21 20:51 +0100
Message-ID<le8an9$fi$1@online.de>
In reply to#28659
Anton Ertl wrote:
> But do we want to parallelize our programs?  The highest expected
> speedup is by the number of cores, i.e., typically 4.  For many
> programs there is lower-hanging fruit that gives a speedup by a factor
> of 4, so the case for parallelization does not seem strong to me.  It
> only applies to the few programs where the lower-hanging fruit has
> already been picked.

One very typical problem with these multi-core CPUs is that communication 
overhead is quite high.  I've experimented with parallelism in net2o, one 
task would do the Unix socket stuff, and another task would do the rest of 
the protocol, including encryption (which does cost performance, about 10 
cycles per byte).  However, the result was that the communication overhead 
between the two tasks was big enough to make the parallel solution slower 
than the do-it-all-on-one-core simple sequential solution.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#28688

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2014-02-22 09:39 -0600
Message-ID<BJ-dnTMtgeYtWZXOnZ2dnUVZ_t-dnZ2d@supernews.com>
In reply to#28668
Bernd Paysan <bernd.paysan@gmx.de> wrote:
> 
> One very typical problem with these multi-core CPUs is that
> communication overhead is quite high.  I've experimented with
> parallelism in net2o, one task would do the Unix socket stuff, and
> another task would do the rest of the protocol, including encryption
> (which does cost performance, about 10 cycles per byte).  However,
> the result was that the communication overhead between the two tasks
> was big enough to make the parallel solution slower than the
> do-it-all-on-one-core simple sequential solution.

Hmmm, that's a little surprising.  I'm guessing that an L3 cache hit
in another core is maybe 75 cycles.  At 10 cycles per byte that's the
time it takes to encrypt 7 bytes, but a cache line is 64 bytes in
size.  There's some signalling overhead, but it still sounds
worthwhile, but only just.  I suppose the problem is the latency
signalling the worker thread to start.

Andrew.

[toc] | [prev] | [next] | [standalone]


#28699

FromBernd Paysan <bernd.paysan@gmx.de>
Date2014-02-23 02:40 +0100
Message-ID<lebji3$1qg$1@online.de>
In reply to#28688
Andrew Haley wrote:

> Bernd Paysan <bernd.paysan@gmx.de> wrote:
>> 
>> One very typical problem with these multi-core CPUs is that
>> communication overhead is quite high.  I've experimented with
>> parallelism in net2o, one task would do the Unix socket stuff, and
>> another task would do the rest of the protocol, including encryption
>> (which does cost performance, about 10 cycles per byte).  However,
>> the result was that the communication overhead between the two tasks
>> was big enough to make the parallel solution slower than the
>> do-it-all-on-one-core simple sequential solution.
> 
> Hmmm, that's a little surprising.  I'm guessing that an L3 cache hit
> in another core is maybe 75 cycles.  At 10 cycles per byte that's the
> time it takes to encrypt 7 bytes, but a cache line is 64 bytes in
> size.

And a network packet has 1024 bytes, or 10k cycles to decrypt it.  The L3 
transfer speed would be about 1 cycle per byte, and certainly worth it.

> There's some signalling overhead, but it still sounds
> worthwhile, but only just.  I suppose the problem is the latency
> signalling the worker thread to start.

Yes.  That goes through the OS, and takes ages.  There's a good reason why 
the Transputer had its links in hardware.  In any case, the OS overhead here 
is so big that the time it takes Linux to get a simple UDP packet from one 
process to another is the same amount of time it takes net2o to do the 
entire stack including encryption (as user task).  And encryption is 80% or 
90% of that.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#28701

FromPaul Rubin <no.email@nospam.invalid>
Date2014-02-22 19:24 -0800
Message-ID<7xios6h5o6.fsf@ruckus.brouhaha.com>
In reply to#28699
Bernd Paysan <bernd.paysan@gmx.de> writes:
>> I suppose the problem is the latency signalling the worker thread to
>> start.
> Yes.  That goes through the OS, and takes ages.  

GHC somehow deals with this using shared memory and work-stealing,
without going through the OS except at initial startup when the OS
threads are launched.  I wonder if something like that could be
practical for net2o.

[toc] | [prev] | [next] | [standalone]


#28722

FromBernd Paysan <bernd.paysan@gmx.de>
Date2014-02-23 22:48 +0100
Message-ID<ledqas$n6j$2@online.de>
In reply to#28701
Paul Rubin wrote:

> Bernd Paysan <bernd.paysan@gmx.de> writes:
>>> I suppose the problem is the latency signalling the worker thread to
>>> start.
>> Yes.  That goes through the OS, and takes ages.
> 
> GHC somehow deals with this using shared memory and work-stealing,
> without going through the OS except at initial startup when the OS
> threads are launched.  I wonder if something like that could be
> practical for net2o.

For high-speed connections, probably, for low-speed, I don't want to have a 
thread busy all the time.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#28710

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2014-02-23 04:39 -0600
Message-ID<Uq2dnT2Ha9VtUpTOnZ2dnUVZ_s-dnZ2d@supernews.com>
In reply to#28699
Bernd Paysan <bernd.paysan@gmx.de> wrote:
> Andrew Haley wrote:
> 
>> Bernd Paysan <bernd.paysan@gmx.de> wrote:
>>> 
>>> One very typical problem with these multi-core CPUs is that
>>> communication overhead is quite high.  I've experimented with
>>> parallelism in net2o, one task would do the Unix socket stuff, and
>>> another task would do the rest of the protocol, including encryption
>>> (which does cost performance, about 10 cycles per byte).  However,
>>> the result was that the communication overhead between the two tasks
>>> was big enough to make the parallel solution slower than the
>>> do-it-all-on-one-core simple sequential solution.
>> 
>> Hmmm, that's a little surprising.  I'm guessing that an L3 cache hit
>> in another core is maybe 75 cycles.  At 10 cycles per byte that's the
>> time it takes to encrypt 7 bytes, but a cache line is 64 bytes in
>> size.
> 
> And a network packet has 1024 bytes, or 10k cycles to decrypt it.
> The L3 transfer speed would be about 1 cycle per byte, and certainly
> worth it.

Right.

>> There's some signalling overhead, but it still sounds worthwhile,
>> but only just.  I suppose the problem is the latency signalling the
>> worker thread to start.
>
> Yes.  That goes through the OS, and takes ages.  There's a good
> reason why the Transputer had its links in hardware.  In any case,
> the OS overhead here is so big that the time it takes Linux to get a
> simple UDP packet from one process to another is the same amount of
> time it takes net2o to do the entire stack including encryption (as
> user task).  And encryption is 80% or 90% of that.

Buy why on Earth would you use the network stack when doing this high-
speed communication?  Why involve the OS at all, lat alone UDP?
Surely that's far too complex a protocol for a job like this.  If you
can hit the cache of another core in 75 cycles, why not just do so?

Andrew.

[toc] | [prev] | [next] | [standalone]


#28721

FromBernd Paysan <bernd.paysan@gmx.de>
Date2014-02-23 22:46 +0100
Message-ID<ledq81$n6j$1@online.de>
In reply to#28710
Andrew Haley wrote:
> Buy why on Earth would you use the network stack when doing this high-
> speed communication?  Why involve the OS at all, lat alone UDP?

That was just a comparison.  Sending a byte via a Unix fifo from one thread 
to the other takes about the same amount of time.

> Surely that's far too complex a protocol for a job like this.  If you
> can hit the cache of another core in 75 cycles, why not just do so?

Because the hardware doesn't give me a "start/stop thread" communication 
mean.  As this is a network stack, the decrypt thread would have no idea how 
often it needs to decrypt a packet.  A possible solution is to busy-wait on 
a shared variable for some time, and if that fails, resort to OS 
communication.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#28723

FromPaul Rubin <no.email@nospam.invalid>
Date2014-02-23 14:26 -0800
Message-ID<7x38j9ea9b.fsf@ruckus.brouhaha.com>
In reply to#28721
Bernd Paysan <bernd.paysan@gmx.de> writes:
> As this is a network stack, the decrypt thread would have no idea how
> often it needs to decrypt a packet.  A possible solution is to
> busy-wait on a shared variable for some time, and if that fails,
> resort to OS communication.

If it's POSIX threads you'd probably use a condition variable.  That
should be a lot faster than socket communication.  If you're going to
use sockets for IPC on the same machine, Unix domain sockets are
probably faster than UDP and have various other advantages as well.

[toc] | [prev] | [next] | [standalone]


#28870

FromBernd Paysan <bernd.paysan@gmx.de>
Date2014-02-24 02:44 +0100
Message-ID<lee85i$k0p$1@online.de>
In reply to#28723
Paul Rubin wrote:

> Bernd Paysan <bernd.paysan@gmx.de> writes:
>> As this is a network stack, the decrypt thread would have no idea how
>> often it needs to decrypt a packet.  A possible solution is to
>> busy-wait on a shared variable for some time, and if that fails,
>> resort to OS communication.
> 
> If it's POSIX threads you'd probably use a condition variable.  That
> should be a lot faster than socket communication.  If you're going to
> use sockets for IPC on the same machine, Unix domain sockets are
> probably faster than UDP and have various other advantages as well.

Yes, I thought so, and therefore I'm using an Unix fifo for thread-
communication.  It's not faster than the UDP socket, which was a surprise 
for me.  I can try with the condition variable.  There however is a reason 
why I'm using a Unix fifo:  you can wait on that and on any other file 
descriptor in parallel.  I frequently have the case that a thread waits for 
a file or socket *and* for signals from other threads, and whatever comes 
first, is handled appropriately.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#28871

FromSpam@ControlQ.com
Date2014-03-03 12:43 -0500
Message-ID<alpine.BSF.2.00.1403031241520.55061@yoko.controlq.com>
In reply to#28870
On Mon, 24 Feb 2014, Bernd Paysan wrote:

> Date: Mon, 24 Feb 2014 02:44:18 +0100
> From: Bernd Paysan <bernd.paysan@gmx.de>
> Newsgroups: comp.lang.forth
> Subject: Re: Forth reinvention
> 
> Paul Rubin wrote:
>
>> Bernd Paysan <bernd.paysan@gmx.de> writes:
>>> As this is a network stack, the decrypt thread would have no idea how
>>> often it needs to decrypt a packet.  A possible solution is to
>>> busy-wait on a shared variable for some time, and if that fails,
>>> resort to OS communication.
>>
>> If it's POSIX threads you'd probably use a condition variable.  That
>> should be a lot faster than socket communication.  If you're going to
>> use sockets for IPC on the same machine, Unix domain sockets are
>> probably faster than UDP and have various other advantages as well.
>
> Yes, I thought so, and therefore I'm using an Unix fifo for thread-
> communication.  It's not faster than the UDP socket, which was a surprise
> for me.  I can try with the condition variable.  There however is a reason
> why I'm using a Unix fifo:  you can wait on that and on any other file
> descriptor in parallel.  I frequently have the case that a thread waits for
> a file or socket *and* for signals from other threads, and whatever comes
> first, is handled appropriately.
>
>

Before you roll your own, perhaps you should look at nanomsg ... a BSD 
licensed replacement for 0MQ which supports various communication models.

 	nanomsg.org

Cheers,
Rob.

[toc] | [prev] | [next] | [standalone]


Page 3 of 4 — ← Prev page 1 2 [3] 4  Next page →

Back to top | Article view | comp.lang.forth


csiph-web