Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #19200 > unrolled thread

About stack access profundity

Started byPablo Hugo Reda <pabloreda@gmail.com>
First post2013-01-27 12:11 -0800
Last post2013-02-03 15:49 -0800
Articles 20 on this page of 46 — 10 participants

Back to article view | Back to comp.lang.forth


Contents

  About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-27 12:11 -0800
    Re: About stack access profundity humptydumpty <ouatubi@gmail.com> - 2013-01-28 00:59 -0800
      Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-28 07:17 -0800
    Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-28 03:10 -0600
      Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-28 07:23 -0800
        Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-28 10:46 -0600
          Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 08:20 +0000
            Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 03:25 -0600
              Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 09:47 +0000
                Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 04:08 -0600
                  Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:21 +0000
                    Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 09:57 -0600
                      Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 17:01 +0000
                        Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 12:22 -0600
                          Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 17:40 +0000
                            Re: About stack access profundity Paul Rubin <no.email@nospam.invalid> - 2013-01-30 11:49 -0800
                              Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-02-04 16:39 +0000
                            Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-30 16:32 -0600
                              Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-31 16:02 +0000
                    Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-29 17:06 +0000
                      Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 17:58 +0000
                        Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 12:29 -0600
                          Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 17:22 +0000
                        Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-29 20:17 +0100
                          Re: About stack access profundity mhx@iae.nl (Marcel Hendrix) - 2013-01-29 21:44 +0200
                            Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-29 22:14 +0100
                          Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 16:29 +0000
                            Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-30 09:58 -0800
                          Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-31 12:32 +0000
                            Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-31 15:37 +0100
                        Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-29 23:12 -0800
            Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-29 10:08 +0000
              Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-29 07:10 -0800
              Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:17 +0000
        Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-29 00:57 -0800
    Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-28 01:56 -0800
    Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:30 +0000
      Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-29 09:06 -0800
    Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 06:22 -0800
      Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-02-03 07:25 -0800
        Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 15:51 -0800
          Re: About stack access profundity Coos Haak <chforth@hccnet.nl> - 2013-02-04 01:12 +0100
      Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-02-03 23:10 -0800
        Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-02-04 03:26 -0600
    Re: About stack access profundity humptydumpty <ouatubi@gmail.com> - 2013-02-03 11:26 -0800
      Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 15:49 -0800

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#19256

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-01-29 17:58 +0000
Message-ID<2013Jan29.185800@mips.complang.tuwien.ac.at>
In reply to#19254
stephenXXX@mpeforth.com (Stephen Pelc) writes:
>What happens is that PICKs produce fetches indexed from the data
>stack pointer,

VFX is better than you give it credit for:

variable A
variable B
variable C
: bla A @ B @ 1 pick + swap drop C ! ;
see bla

shows

( 080BF3A0    8B153C240A08 )          MOV       EDX, [080A243C]
( 080BF3A6    031540240A08 )          ADD       EDX, [080A2440]
( 080BF3AC    891544240A08 )          MOV       [080A2444], EDX
( 080BF3B2    C3 )                    NEXT,

i.e., no fetch indexed from the data stack pointer (no reference to
the data stack at all).  VFX recognizes that PICK accesses a stack
element in a register and optimizes it away.

>When we rewrote our PowerView embedded GUI to pass structures rather
>than than keep graphics coordinates on the stack, the code (for ARM
>and Cortex) became shorter and faster. In our experience, your
>assertion does not hold bcause the use of structures considerably
>reduces the stack traffic.

I would have to look at the concrete code (before and after the
change) to give a proper comment on that.  But if all that happens is
that you replace a "5 PICK" (which should produce a register reference
or, if there are not enough registers, a memory reference to the
memory part of the stack) with something like "DUP .X @", "OVER .X @",
"R@ .X @" (which should produce at least one memory reference, for the
@), the use of structures should not be shorter and faster.

I think that the limited scope of VFXs register allocator reduces the
benefit of stack references, but they still should not hurt.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#19258

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2013-01-29 12:29 -0600
Message-ID<iaedncGLNcloiZXMnZ2dnUVZ_qadnZ2d@supernews.com>
In reply to#19256
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
> stephenXXX@mpeforth.com (Stephen Pelc) writes:
> 
>>When we rewrote our PowerView embedded GUI to pass structures rather
>>than than keep graphics coordinates on the stack, the code (for ARM
>>and Cortex) became shorter and faster. In our experience, your
>>assertion does not hold bcause the use of structures considerably
>>reduces the stack traffic.
> 
> I would have to look at the concrete code (before and after the
> change) to give a proper comment on that.  But if all that happens is
> that you replace a "5 PICK" (which should produce a register reference
> or, if there are not enough registers, a memory reference to the
> memory part of the stack) with something like "DUP .X @", "OVER .X @",
> "R@ .X @" (which should produce at least one memory reference, for the
> @), the use of structures should not be shorter and faster.

It's unlikely to be just "5 PICK", though.  There will be writes too,
and that's either POKE (aargh) or lots of stack thrashing to get the
data into position.

Andrew.

[toc] | [prev] | [next] | [standalone]


#19287

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-01-30 17:22 +0000
Message-ID<2013Jan30.182257@mips.complang.tuwien.ac.at>
In reply to#19258
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> stephenXXX@mpeforth.com (Stephen Pelc) writes:
>> 
>>>When we rewrote our PowerView embedded GUI to pass structures rather
>>>than than keep graphics coordinates on the stack, the code (for ARM
>>>and Cortex) became shorter and faster. In our experience, your
>>>assertion does not hold bcause the use of structures considerably
>>>reduces the stack traffic.
>> 
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that.  But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.
>
>It's unlikely to be just "5 PICK", though.  There will be writes too,

Well, in Pablo Reda's code, and that's what we are talking about,
there were only PICKs (for the memory variant, @s), no STICKs/POKEs,
or ROLLs.

>and that's either POKE (aargh) or lots of stack thrashing to get the
>data into position.

Yes, if Stephen's code did that, that might explain the larger code.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#19260

FromBernd Paysan <bernd.paysan@gmx.de>
Date2013-01-29 20:17 +0100
Message-ID<1950087.W2tCE9ddpV@sunwukong.fritz.box>
In reply to#19256
Anton Ertl wrote:
> I would have to look at the concrete code (before and after the
> change) to give a proper comment on that.  But if all that happens is
> that you replace a "5 PICK" (which should produce a register reference
> or, if there are not enough registers, a memory reference to the
> memory part of the stack) with something like "DUP .X @", "OVER .X @",
> "R@ .X @" (which should produce at least one memory reference, for the
> @), the use of structures should not be shorter and faster.

Not convinced.  VFX doesn't do that too well:

begin-structure point  ok-2 
field: .x  ok-2 
field: .y  ok-2 
end-structure  ok
: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ;  ok
see test  
TEST 
( 080BC940    53 )                    PUSH      EBX
( 080BC941    8B1424 )                MOV       EDX, [ESP]
( 080BC944    8B4A04 )                MOV       ECX, [EDX+04]
( 080BC947    030B )                  ADD       ECX, 0 [EBX]
( 080BC949    8B0424 )                MOV       EAX, [ESP]
( 080BC94C    8D6DF4 )                LEA       EBP, [EBP+-0C]
( 080BC94F    894D00 )                MOV       [EBP], ECX
( 080BC952    8B4A04 )                MOV       ECX, [EDX+04]
( 080BC955    894D04 )                MOV       [EBP+04], ECX
( 080BC958    8B13 )                  MOV       EDX, 0 [EBX]
( 080BC95A    895508 )                MOV       [EBP+08], EDX
( 080BC95D    8BD8 )                  MOV       EBX, EAX
( 080BC95F    8B5500 )                MOV       EDX, [EBP]
( 080BC962    8913 )                  MOV       0 [EBX], EDX
( 080BC964    8B5D08 )                MOV       EBX, [EBP+08]
( 080BC967    2B5D04 )                SUB       EBX, [EBP+04]
( 080BC96A    5A )                    POP       EDX
( 080BC96B    895A04 )                MOV       [EDX+04], EBX
( 080BC96E    8B5D0C )                MOV       EBX, [EBP+0C]
( 080BC971    8D6D10 )                LEA       EBP, [EBP+10]
( 080BC974    C3 )                    NEXT,
( 53 bytes, 21 instructions )

That's 21 instructions, clearly not what I would have written by hand.

Compare that to bigForth, using the current object pointer OOP:

debugging class point  ok
cell var .x  ok
cell var .y  ok
how:   ok
public: : test >o .x @ .y @ 2dup + .x ! - .y ! o> ;  ok
disw test  Adresse : 268670656
100396C0: push    EDI               57
100396C1: mov     EDI,EAX           8BF8
100396C3: lodsd                     AD
100396C4: xchg    ESP,ESI           87F4
100396C6: push    EAX               50
100396C7: push    DWORD PTR $04[EDI]
                                    FF7704
100396CA: mov     EAX,$08[EDI]      8B4708
100396CD: mov     EDX,[ESP]         8B1424
100396D0: push    EAX               50
100396D1: add     EAX,EDX           03C2
100396D3: push    EAX               50
100396D4: pop     DWORD PTR $04[EDI]
                                    8F4704
100396D7: pop     EAX               58
100396D8: pop     EDX               5A
100396D9: xchg    EDX,EAX           92
100396DA: sub     EAX,EDX           2BC2
100396DC: push    EAX               50
100396DD: pop     DWORD PTR $08[EDI]
                                    8F4708
100396E0: pop     EAX               58
100396E1: xchg    ESP,ESI           87F4
100396E3: pop     EDI               5F
100396E4: ret                       C3

22 instructions, room for improvement, because bigForth isn't an 
analytical compiler.  What I would expect is that apart from the struct 
memory accesses, everything would fit into the registers.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#19266

Frommhx@iae.nl (Marcel Hendrix)
Date2013-01-29 21:44 +0200
Message-ID<63591406028434@frunobulax.edu>
In reply to#19260
Bernd Paysan <bernd.paysan@gmx.de> writes Re: About stack access profundity

> Anton Ertl wrote:
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that.  But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.

> Not convinced.  VFX doesn't do that too well:
[..]

iForth64 ...

FORTH> begin-structure point  ok
[3]FORTH> field: .x  ok
[3]FORTH> field: .y  ok
[3]FORTH> end-structure  ok
FORTH>   ok
FORTH> : test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ;  ok
FORTH> see test
Flags: TOKENIZE, ANSI
: test  >R R@ .x @ R@ .y @ 2DUP + R@ .x ! - R> .y ! ;  ok
FORTH> ' test idis
$01404300  : [trashed]
$0140430A  pop           rbx
$0140430B  mov           rdi, [rbx] qword
$0140430E  add           rdi, [rbx 8 +] qword
$01404312  mov           rax, [rbx] qword
$01404315  mov           rdx, [rbx 8 +] qword
$01404319  mov           [rbx] qword, rdi
$0140431C  sub           rax, rdx
$0140431F  mov           [rbx 8 +] qword, rax
$01404323  ;
FORTH> create ape 2 cells allot  ok
FORTH> : tt ape test ;  ok
FORTH> see tt
Flags: TOKENIZE, ANSI
: tt  ape test ;  ok
FORTH> ' tt idis
$01404BC0  : [trashed]
$01404BCA  mov           rbx, $01404780 qword-offset
$01404BD1  add           rbx, $01404788 qword-offset
$01404BD8  mov           rdi, $01404780 qword-offset
$01404BDF  mov           rax, $01404788 qword-offset
$01404BE6  mov           $01404780 qword-offset, rbx
$01404BED  sub           rdi, rax
$01404BF0  mov           $01404788 qword-offset, rdi
$01404BF7  ;

-marcel

[toc] | [prev] | [next] | [standalone]


#19267

FromBernd Paysan <bernd.paysan@gmx.de>
Date2013-01-29 22:14 +0100
Message-ID<4134598.R9bZG8GDba@sunwukong.fritz.box>
In reply to#19266
Marcel Hendrix wrote:
> iForth64 ...

Great!  That's pretty close to what I would write by hand.

Hand-code (let's assume rax is tos):

mov    rbx, [rax]
mov    rcx, [rax+8]
lea    rdx, [rbx+rcx]
sub    rbx, rcx
mov    [rax], rdx
mov    [rax+8], rbx

Approach: Don't load values twice, though on x86, you have quite a lot 
of load units.  Use lea for add when you need a three operand add.

Not sure why your compiler generates

mov           rax, [rbx] qword
mov           rdx, [rbx 8 +] qword
sub           rax, rdx

instead of

mov           rax, [rbx] qword
sub           rax, [rbx 8 +] qword

> FORTH> begin-structure point  ok
> [3]FORTH> field: .x  ok
> [3]FORTH> field: .y  ok
> [3]FORTH> end-structure  ok
> FORTH>   ok
> FORTH> : test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y !
> ;  ok FORTH> see test
> Flags: TOKENIZE, ANSI
> : test  >R R@ .x @ R@ .y @ 2DUP + R@ .x ! - R> .y ! ;  ok
> FORTH> ' test idis
> $01404300  : [trashed]
> $0140430A  pop           rbx
> $0140430B  mov           rdi, [rbx] qword
> $0140430E  add           rdi, [rbx 8 +] qword
> $01404312  mov           rax, [rbx] qword
> $01404315  mov           rdx, [rbx 8 +] qword
> $01404319  mov           [rbx] qword, rdi
> $0140431C  sub           rax, rdx
> $0140431F  mov           [rbx 8 +] qword, rax
> $01404323  ;
> FORTH> create ape 2 cells allot  ok
> FORTH> : tt ape test ;  ok
> FORTH> see tt
> Flags: TOKENIZE, ANSI
> : tt  ape test ;  ok
> FORTH> ' tt idis
> $01404BC0  : [trashed]
> $01404BCA  mov           rbx, $01404780 qword-offset
> $01404BD1  add           rbx, $01404788 qword-offset
> $01404BD8  mov           rdi, $01404780 qword-offset
> $01404BDF  mov           rax, $01404788 qword-offset
> $01404BE6  mov           $01404780 qword-offset, rbx
> $01404BED  sub           rdi, rax
> $01404BF0  mov           $01404788 qword-offset, rdi
> $01404BF7  ;
> 
> -marcel
-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#19285

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-01-30 16:29 +0000
Message-ID<2013Jan30.172942@mips.complang.tuwien.ac.at>
In reply to#19260
Bernd Paysan <bernd.paysan@gmx.de> writes:
>Anton Ertl wrote:
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that.  But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.
>
>Not convinced.

Of what are you are not convinced?

>  VFX doesn't do that too well:
>
>begin-structure point  ok-2 
>field: .x  ok-2 
>field: .y  ok-2 
>end-structure  ok
>: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ;  ok
>see test  
>TEST 
>( 080BC940    53 )                    PUSH      EBX
>( 080BC941    8B1424 )                MOV       EDX, [ESP]
>( 080BC944    8B4A04 )                MOV       ECX, [EDX+04]
>( 080BC947    030B )                  ADD       ECX, 0 [EBX]
>( 080BC949    8B0424 )                MOV       EAX, [ESP]
>( 080BC94C    8D6DF4 )                LEA       EBP, [EBP+-0C]
>( 080BC94F    894D00 )                MOV       [EBP], ECX
>( 080BC952    8B4A04 )                MOV       ECX, [EDX+04]
>( 080BC955    894D04 )                MOV       [EBP+04], ECX
>( 080BC958    8B13 )                  MOV       EDX, 0 [EBX]
>( 080BC95A    895508 )                MOV       [EBP+08], EDX
>( 080BC95D    8BD8 )                  MOV       EBX, EAX
>( 080BC95F    8B5500 )                MOV       EDX, [EBP]
>( 080BC962    8913 )                  MOV       0 [EBX], EDX
>( 080BC964    8B5D08 )                MOV       EBX, [EBP+08]
>( 080BC967    2B5D04 )                SUB       EBX, [EBP+04]
>( 080BC96A    5A )                    POP       EDX
>( 080BC96B    895A04 )                MOV       [EDX+04], EBX
>( 080BC96E    8B5D0C )                MOV       EBX, [EBP+0C]
>( 080BC971    8D6D10 )                LEA       EBP, [EBP+10]
>( 080BC974    C3 )                    NEXT,
>( 53 bytes, 21 instructions )
>
>That's 21 instructions, clearly not what I would have written by hand.

Yes, so VFX is not as great as we might like, but that does not tell
us anything about whether it does better for PICKing or for @ing code.

But let's try it.  I wanted to use the original example
<61d42b93-0ac8-4df3-8cef-6c1cc059d0ef@googlegroups.com> for this, but
it's unclear to me what it does, in particular the line

	4 <? ( drop line 2drop line ; ) drop

and PX and PY.

So I fell back to the good old rectangle example:

begin-structure point
field: point-x
field: point-y
end-structure

defer line ( p1 p2 -- )
defer make-point ( x y -- p )
defer free-point ( p -- )

: line-line ( p1 p2 p3 -- )
  \ draw a line between p1 and p2
  \  and one between p2 and p3
  over line line ;    

: rect-mem ( ll ur -- )
  over point-x @ over point-y @ make-point
  ( ll ur ul )
  >r 2dup r@ swap line-line r> free-point
  over point-y @ over point-x @ swap make-point
  ( ll ur lr )
  >r 2dup r@ swap line-line r> free-point ;

defer line-stack ( x1 y1 x2 y2 -- )

: rect-local {: x1 y1 x2 y2 -- :}
  x1 y1 x1 y2 line-stack
  x1 y2 x2 y2 line-stack
  x2 y2 x2 y1 line-stack
  x2 y1 x1 y1 line-stack ;

: rect-stack ( x1 y1 x2 y2 -- )
    3 pick 3 pick over   3 pick line-stack
    3 pick over   3 pick over   line-stack
    over   over   over   3 pick line-stack
    over   3 pick 3 pick 3 pick line-stack
    2drop 2drop ;

see rect-mem
see rect-local
see rect-stack

and SEEing the results showed:

RECT-MEM 
...
( 147 bytes, 46 instructions )

RECT-LOCAL 
...
( 166 bytes, 54 instructions )

RECT-STACK 
...
( 121 bytes, 37 instructions )

If you want a rect-mem2 that calls line-stack and is thus closer to
rect-stack and rect-local in what it does internally, feel free to
post it and I'll run it through VFX.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#19288

FromPablo Hugo Reda <pabloreda@gmail.com>
Date2013-01-30 09:58 -0800
Message-ID<db5cb315-6eb1-4a2e-96e0-352dc58103ba@googlegroups.com>
In reply to#19285
>  but
> 
> it's unclear to me what it does, in particular the line
> 
> 
> 
> 	4 <? ( drop line 2drop line ; ) drop
> 
> 
> 
> and PX and PY.
> 
> 

sorry for have a forth dialect, not at ans-forth
I take the ideas from colorforth, I not use STATE, not use DOES>, etc.


PX and PY are variables, the initial point for the curve, LINE draw a line and update the PX and PY.

4 <? (..

go inside (..) when the Top of stack is <4 (and consume 4)

here is the code generated by the compiler, in FASM syntax.
I remove the comments because have 200 lines with this.
"uso" is the stack profundity and "dD" is the stack variation
------------------------------------------------------------
w24: ; ::: sp-dist ::: uso:-4 dD:1
mov ebx,dword [esi+8]
sub ebx,dword [esi]
mov edx,ebx
sar edx,31
add ebx,edx
xor ebx,edx
mov ecx,dword [esi+4]
sub ecx,eax
mov edx,ecx
sar edx,31
add ecx,edx
xor ecx,edx
add ebx,ecx
lea esi,[esi-4]
mov dword [esi],eax
mov eax,ebx
ret

w25: ; ::: sp-c ::: uso:-4 dD:2
mov ebx,dword [esi]
sal ebx,1
mov ecx,dword [esi+8]
add ecx,ebx
add ecx,dword [w9]
sar ecx,$2
mov ebx,eax
sal ebx,1
mov edx,dword [esi+4]
add edx,ebx
add edx,dword [wA]
sar edx,$2
lea esi,[esi-8]
mov dword [esi+4],eax
mov dword [esi],ecx
mov eax,edx
ret

w26: ; ::: spl ::: uso:-4 dD:-4
call w25
call w24
cmp eax,$4
jge _46
lea esi,[esi+4]
mov eax,dword [esi-4]
call w23
lea esi,[esi+8]
mov eax,dword [esi-4]
jmp w23
_46:
push dword [esi]
push dword [esi+4]
mov eax,dword [esi+20]
add eax,dword [esi+12]
sar eax,1
mov ebx,dword [esi+16]
add ebx,dword [esi+8]
sar ebx,1
mov ecx,dword [esi+8]
add ecx,dword [wA]
sar ecx,1
mov edx,dword [esi+12]
add edx,dword [w9]
sar edx,1
pop edi
mov [esi+8],eax
pop eax
lea esi,[esi-4]
xchg dword [esi+12],ebx
mov dword [esi+8],edi
mov dword [esi+4],eax
mov dword [esi],edx
mov dword [esi+16],ebx
mov eax,ecx
call w26
jmp w26
----------------------------------------------------

sure Hendrix and others compiler is better but I not finish. have more ideas for improvement.

I don't have the solution, but thank's for all for comment.


[toc] | [prev] | [next] | [standalone]


#19310

FromstephenXXX@mpeforth.com (Stephen Pelc)
Date2013-01-31 12:32 +0000
Message-ID<510a62ea.510135172@192.168.0.50>
In reply to#19260
On Tue, 29 Jan 2013 20:17:54 +0100, Bernd Paysan <bernd.paysan@gmx.de>
wrote:

>Anton Ertl wrote:
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that.  But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.
>
>Not convinced.  VFX doesn't do that too well:
>
>begin-structure point  ok-2 
>field: .x  ok-2 
>field: .y  ok-2 
>end-structure  ok
>: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ;  ok
>see test  
>TEST 
>( 080BC940    53 )                    PUSH      EBX
>( 080BC941    8B1424 )                MOV       EDX, [ESP]
>( 080BC944    8B4A04 )                MOV       ECX, [EDX+04]
>( 080BC947    030B )                  ADD       ECX, 0 [EBX]
>( 080BC949    8B0424 )                MOV       EAX, [ESP]
>( 080BC94C    8D6DF4 )                LEA       EBP, [EBP+-0C]
>( 080BC94F    894D00 )                MOV       [EBP], ECX
>( 080BC952    8B4A04 )                MOV       ECX, [EDX+04]
>( 080BC955    894D04 )                MOV       [EBP+04], ECX
>( 080BC958    8B13 )                  MOV       EDX, 0 [EBX]
>( 080BC95A    895508 )                MOV       [EBP+08], EDX
>( 080BC95D    8BD8 )                  MOV       EBX, EAX
>( 080BC95F    8B5500 )                MOV       EDX, [EBP]
>( 080BC962    8913 )                  MOV       0 [EBX], EDX
>( 080BC964    8B5D08 )                MOV       EBX, [EBP+08]
>( 080BC967    2B5D04 )                SUB       EBX, [EBP+04]
>( 080BC96A    5A )                    POP       EDX
>( 080BC96B    895A04 )                MOV       [EDX+04], EBX
>( 080BC96E    8B5D0C )                MOV       EBX, [EBP+0C]
>( 080BC971    8D6D10 )                LEA       EBP, [EBP+10]
>( 080BC974    C3 )                    NEXT,
>( 53 bytes, 21 instructions )
>
>That's 21 instructions, clearly not what I would have written by hand.

The code from the LEA up to but not including the SUB is a stack
shuffle (we ran out of registers) followed by a store. Although it's
possible to do it without the shuffle, VFX's "rip up and retry"
rules are not clever enough here.

Marcel's iForth version is for 64 bit code with 16 registers. This
means more registers, so no shuffle, and because there are more
registers, some can be used for the return stack. Nice code.

Stephen

-- 
Stephen Pelc, stephenXXX@mpeforth.com
MicroProcessor Engineering Ltd - More Real, Less Time
133 Hill Lane, Southampton SO15 5AF, England
tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691
web: http://www.mpeforth.com - free VFX Forth downloads

[toc] | [prev] | [next] | [standalone]


#19314

FromBernd Paysan <bernd.paysan@gmx.de>
Date2013-01-31 15:37 +0100
Message-ID<3427167.9fDfFC78xt@sunwukong.fritz.box>
In reply to#19310
Stephen Pelc wrote:
> The code from the LEA up to but not including the SUB is a stack
> shuffle (we ran out of registers) followed by a store. Although it's
> possible to do it without the shuffle, VFX's "rip up and retry"
> rules are not clever enough here.

You need four free registers (including TOS) to fit this code in.  In 
bigForth, I have three free registers, because I have SP, RP, UP, OP 
(object pointer) and an index for loops in a register.  VFX doesn't have 
an OP, and doesn't waste another register for the index, so it should be 
possible to fit it into the available registers.  The main thing you 
need to do is to try hard to eliminate redundancy - all active values 
should only occupy one single register, unless a copy is dearly needed.

> Marcel's iForth version is for 64 bit code with 16 registers. This
> means more registers, so no shuffle, and because there are more
> registers, some can be used for the return stack. Nice code.

Yes, x64 is a much better target, because you have 8 really free 
registers (in x86, there is no register without a special role, though 
EBX is only used in xlat, and xlat is really superfluous - however, you 
can argue that x86 has 8 special purpose registers, and no general 
purpose register at all).

This is also impacting C compilers.  The string instructions use up ECX, 
ESI, and EDI, multiplication uses up EAX and EDX; -fomit-frame-pointer 
gives you EBP, so you have EBP and EBX, no more.  In Gforth, we 
therefore moved all string operations into real subroutines (even though 
that makes them a bit slower), and only then the C compiler is able to 
fit in the most important things.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#19276

FromMark Wills <forthfreak@gmail.com>
Date2013-01-29 23:12 -0800
Message-ID<680dad03-965e-4357-b2fc-97a0783deeca@w7g2000yqo.googlegroups.com>
In reply to#19256
On Jan 29, 5:58 pm, an...@mips.complang.tuwien.ac.at (Anton Ertl)
wrote:
> stephen...@mpeforth.com (Stephen Pelc) writes:
> >What happens is that PICKs produce fetches indexed from the data
> >stack pointer,
>
> VFX is better than you give it credit for:
>
> variable A
> variable B
> variable C
> : bla A @ B @ 1 pick + swap drop C ! ;
> see bla
>
> shows
>
> ( 080BF3A0    8B153C240A08 )          MOV       EDX, [080A243C]
> ( 080BF3A6    031540240A08 )          ADD       EDX, [080A2440]
> ( 080BF3AC    891544240A08 )          MOV       [080A2444], EDX
> ( 080BF3B2    C3 )                    NEXT,
>
> i.e., no fetch indexed from the data stack pointer (no reference to
> the data stack at all).  VFX recognizes that PICK accesses a stack
> element in a register and optimizes it away.
>
> >When we rewrote our PowerView embedded GUI to pass structures rather
> >than than keep graphics coordinates on the stack, the code (for ARM
> >and Cortex) became shorter and faster. In our experience, your
> >assertion does not hold bcause the use of structures considerably
> >reduces the stack traffic.
>
> I would have to look at the concrete code (before and after the
> change) to give a proper comment on that.  But if all that happens is
> that you replace a "5 PICK" (which should produce a register reference
> or, if there are not enough registers, a memory reference to the
> memory part of the stack) with something like "DUP .X @", "OVER .X @",
> "R@ .X @" (which should produce at least one memory reference, for the
> @), the use of structures should not be shorter and faster.
>
> I think that the limited scope of VFXs register allocator reduces the
> benefit of stack references, but they still should not hurt.
>
> - anton
> --
> M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
> comp.lang.forth FAQs:http://www.complang.tuwien.ac.at/forth/faq/toc.html
>      New standard:http://www.forth200x.org/forth200x.html
>    EuroForth 2012:http://www.euroforth.org/ef12/

Exactly the point I was making on another thread. An obsessive
programmer might second guess the compiler, and instead of writing
plain, vanilla, straight-ahead code, come up with something that he
*perceives* to be faster, but instead, it's slower because the
optimiser can't identify what the programmers real intention was.

Unless you are using a dumb ITC compiler, just write the code.

[toc] | [prev] | [next] | [standalone]


#19243

FromstephenXXX@mpeforth.com (Stephen Pelc)
Date2013-01-29 10:08 +0000
Message-ID<51079f62.329007481@192.168.0.50>
In reply to#19236
On Tue, 29 Jan 2013 08:20:47 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:

>There's a good reason why decent compilers will produce worse code for
>@ than for PICK.  Whether the worseness is "much" or not can be
>debated forever without resolution, so let's not go there.

Please explain this.

Stephen

-- 
Stephen Pelc, stephenXXX@mpeforth.com
MicroProcessor Engineering Ltd - More Real, Less Time
133 Hill Lane, Southampton SO15 5AF, England
tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691
web: http://www.mpeforth.com - free VFX Forth downloads

[toc] | [prev] | [next] | [standalone]


#19246

FromPablo Hugo Reda <pabloreda@gmail.com>
Date2013-01-29 07:10 -0800
Message-ID<60c59889-21bf-446e-beb8-303e5af3132d@googlegroups.com>
In reply to#19243
> 
> >There's a good reason why decent compilers will produce worse code for
> 
> >@ than for PICK.  Whether the worseness is "much" or not can be
> 
> >debated forever without resolution, so let's not go there.
> 

stack operations disappears in the compiled code (for me implementation)
but memory access not (for now)

[toc] | [prev] | [next] | [standalone]


#19247

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-01-29 15:17 +0000
Message-ID<2013Jan29.161718@mips.complang.tuwien.ac.at>
In reply to#19243
stephenXXX@mpeforth.com (Stephen Pelc) writes:
>On Tue, 29 Jan 2013 08:20:47 GMT, anton@mips.complang.tuwien.ac.at
>(Anton Ertl) wrote:
>
>>There's a good reason why decent compilers will produce worse code for
>>@ than for PICK.  Whether the worseness is "much" or not can be
>>debated forever without resolution, so let's not go there.
>
>Please explain this.

A decent compiler will produce a load-from-memory for a @, while stack
items can be kept in registers and thus don't require a load (PICK and
ROLL with non-constant u are an exception, but that's not what the OP
did.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#19237

FromMark Wills <forthfreak@gmail.com>
Date2013-01-29 00:57 -0800
Message-ID<c696cb3b-d0cc-4151-806c-a731a96c4715@w3g2000yqj.googlegroups.com>
In reply to#19222
On Jan 28, 3:23 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote:
> > I think the mistake you're making is keeping everything on the stack.
>
> > As you say, objects deep on the stack are hare to access.
>
> > Instead, use a structure with named offsets for your co-ordinates and
>
> > vectors.  You'll find the code is much easier to read and write.
>
> the stack if the more fast access and I need here because the recursion use this estructure
>
> > Pass pointers, rather than the parameters directly?
>
> but with address need a @ and ! to load store values.
>
> when the compiler optimice the best place is the stack
>
> look the cuadric bezier
> -------------------------------------------------------------------
> :sp-dist | x y xe ye -- x y xe ye dd
>         pick3 pick2 - abs pick3 pick2 - abs + ;
>
> :sp-c | fx fy cx cy -- fx fy cx cy xn yn  ; xn=(cx*2+fx+px)/4
>         pick3 pick2 2* + px + 2 >>
>         pick3 pick2 2* + py + 2 >> ;
>
> :spl | fx fy cx cy --
>         sp-c sp-dist
>         4 <? ( drop line 2drop line ; ) drop
>         >r >r
>         pick3 pick2 + 2/ pick3 pick2 + 2/               | fx fy cx cy c2 c2
>         2swap                                           | fx fy c2 c2 cx cy
>         py + 2/ swap px + 2/ swap                       | fx fy c2 c2 c1 c1
>         r> r> 2swap
>         spl spl ;
>
> -------------------------------------------------------------------

This code is crying out for locals! Use locals. All your PICKs will go
away :-)

[toc] | [prev] | [next] | [standalone]


#19212

FromMark Wills <forthfreak@gmail.com>
Date2013-01-28 01:56 -0800
Message-ID<7168762f-2b01-4827-929a-6725775caf6d@l13g2000yqe.googlegroups.com>
In reply to#19200
On Jan 27, 8:11 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote:
> For two definition who make the same, the less distance in stack access is more optimal, need less register when compiler, etc. Is convenient make definition with less profundity access en the stack
> .
> I choose limit the word PICK and make it static, I have only PICK2,PICK3,PICK4 and no more.
> Whell, until now, not need more, but now I found a word who need more PICK's.
>
> I reeimplement the Quadratic Bezier Curve word and try to draw Cubic Bezier Curves now.
> I use recursion, only integers and add,shift and compare.
>
> Because I use recursion, I need 6 parameters (3 2D points) and the rutine access to 8 position (I need PICK8 definition).
>
> I think about workaround the problem:
>
> if compress the 2dvector for use one place, the word only need PICK4, work in the current system!
> but this add the time for change representation (2 cells to 1 cell..) or extact (address to x y..).
>
> but when PICK8 the solution is more direct.
>
> I need advice..
> thank's

Pass pointers, rather than the parameters directly?

[toc] | [prev] | [next] | [standalone]


#19249

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2013-01-29 15:30 +0000
Message-ID<2013Jan29.163037@mips.complang.tuwien.ac.at>
In reply to#19200
I published a paper that discusses various ways to deal with needing
to deal with too much data at once.

@InProceedings{ertl11euroforth,
  author =       {M. Anton Ertl},
  title =        {Ways to Reduce the Stack Depth},
  crossref =     {euroforth11},
  pages =        {36--41},
  url =          {http://www.complang.tuwien.ac.at/papers/ertl11euroforth.ps.gz},
  url2 =         {http://www.complang.tuwien.ac.at/anton/euroforth/ef11/papers/ertl.pdf},
  OPTnote =      {not refereed},
  abstract =     {Having to deal with many different data can lead to
                  problems in Forth: The data stack is the preferred
                  place to store data; on the other hand, dealing with
                  too many data stack items is cumbersome and usually
                  bad style. This paper presents and discusses ways to
                  unburden the data stack; some of them are used
                  widely, others are almost unknown or new.}
}

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#19253

FromPablo Hugo Reda <pabloreda@gmail.com>
Date2013-01-29 09:06 -0800
Message-ID<f6b5ce4f-4548-4683-a2a4-3a40b9ff8917@googlegroups.com>
In reply to#19249
El martes, 29 de enero de 2013 12:30:37 UTC-3, Anton Ertl  escribió:
> I published a paper that discusses various ways to deal with needing
> 
> to deal with too much data at once.
> 
> 
> 
> @InProceedings{ertl11euroforth,
> 
>   author =       {M. Anton Ertl},
> 
>   title =        {Ways to Reduce the Stack Depth},
> 
>   crossref =     {euroforth11},
> 
>   pages =        {36--41},
> 
>   url =          {http://www.complang.tuwien.ac.at/papers/ertl11euroforth.ps.gz},
> 
>   url2 =         {http://www.complang.tuwien.ac.at/anton/euroforth/ef11/papers/ertl.pdf},
> 
>   OPTnote =      {not refereed},
> 
>   abstract =     {Having to deal with many different data can lead to
> 
>                   problems in Forth: The data stack is the preferred
> 
>                   place to store data; on the other hand, dealing with
> 
>                   too many data stack items is cumbersome and usually
> 
>                   bad style. This paper presents and discusses ways to
> 
>                   unburden the data stack; some of them are used
> 
>                   widely, others are almost unknown or new.}
> 
> }
> 
> 
> 
> - anton
> 
> -- 
> 
> M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
> 
> comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
> 
>      New standard: http://www.forth200x.org/forth200x.html
> 
>    EuroForth 2012: http://www.euroforth.org/ef12/

Thank's Anton

reading...

[toc] | [prev] | [next] | [standalone]


#19388

FromPablo Hugo Reda <pabloreda@gmail.com>
Date2013-02-03 06:22 -0800
Message-ID<722c944e-7658-4603-a7ab-e696573ac34d@googlegroups.com>
In reply to#19200
Well, at last

using some vars and reorganize I get a version without more picks
there is the test version

#x1 #y1 #x2 #y2 #bx #by

:curve3 | xf yf x2 y2 x1 y1
	pick3 pick2 + 2/ pick3 pick2 + 2/ 'by ! 'bx !
	'y1 ! 'x1 !
	pick3 pick2 + 2/ pick3 pick2 + 2/ 2swap
	'y2 ! 'x2 !
	over bx + 2/ over by + 2/
	over x2 - abs over y2 - abs + >r
	px x1 + 2/ py y1 + 2/
	over bx + 2/ over by + 2/
	over x1 - abs over y1 - abs + >r
	2swap >r >r
	pick3 pick2 + 2/ pick3 pick2 + 2/
	2swap r> r>
	r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line )( drop curve3 )
	r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line ; )
	drop curve3 ;

[toc] | [prev] | [next] | [standalone]


#19392

FromMark Wills <forthfreak@gmail.com>
Date2013-02-03 07:25 -0800
Message-ID<5b5748f3-5a9c-4768-a28a-476b11263282@5g2000yqz.googlegroups.com>
In reply to#19388
On Feb 3, 2:22 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote:
> Well, at last
>
> using some vars and reorganize I get a version without more picks
> there is the test version
>
> #x1 #y1 #x2 #y2 #bx #by
>
> :curve3 | xf yf x2 y2 x1 y1
>         pick3 pick2 + 2/ pick3 pick2 + 2/ 'by ! 'bx !
>         'y1 ! 'x1 !
>         pick3 pick2 + 2/ pick3 pick2 + 2/ 2swap
>         'y2 ! 'x2 !
>         over bx + 2/ over by + 2/
>         over x2 - abs over y2 - abs + >r
>         px x1 + 2/ py y1 + 2/
>         over bx + 2/ over by + 2/
>         over x1 - abs over y1 - abs + >r
>         2swap >r >r
>         pick3 pick2 + 2/ pick3 pick2 + 2/
>         2swap r> r>
>         r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line )( drop curve3 )
>         r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line ; )
>         drop curve3 ;

This looks truly horrible to me! Utterly unreadable :-\

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.lang.forth


csiph-web