Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #19200 > unrolled thread
| Started by | Pablo Hugo Reda <pabloreda@gmail.com> |
|---|---|
| First post | 2013-01-27 12:11 -0800 |
| Last post | 2013-02-03 15:49 -0800 |
| Articles | 20 on this page of 46 — 10 participants |
Back to article view | Back to comp.lang.forth
About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-27 12:11 -0800
Re: About stack access profundity humptydumpty <ouatubi@gmail.com> - 2013-01-28 00:59 -0800
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-28 07:17 -0800
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-28 03:10 -0600
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-28 07:23 -0800
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-28 10:46 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 08:20 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 03:25 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 09:47 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 04:08 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:21 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 09:57 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 17:01 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 12:22 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 17:40 +0000
Re: About stack access profundity Paul Rubin <no.email@nospam.invalid> - 2013-01-30 11:49 -0800
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-02-04 16:39 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-30 16:32 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-31 16:02 +0000
Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-29 17:06 +0000
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 17:58 +0000
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-01-29 12:29 -0600
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 17:22 +0000
Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-29 20:17 +0100
Re: About stack access profundity mhx@iae.nl (Marcel Hendrix) - 2013-01-29 21:44 +0200
Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-29 22:14 +0100
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-30 16:29 +0000
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-30 09:58 -0800
Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-31 12:32 +0000
Re: About stack access profundity Bernd Paysan <bernd.paysan@gmx.de> - 2013-01-31 15:37 +0100
Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-29 23:12 -0800
Re: About stack access profundity stephenXXX@mpeforth.com (Stephen Pelc) - 2013-01-29 10:08 +0000
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-29 07:10 -0800
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:17 +0000
Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-29 00:57 -0800
Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-01-28 01:56 -0800
Re: About stack access profundity anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2013-01-29 15:30 +0000
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-01-29 09:06 -0800
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 06:22 -0800
Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-02-03 07:25 -0800
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 15:51 -0800
Re: About stack access profundity Coos Haak <chforth@hccnet.nl> - 2013-02-04 01:12 +0100
Re: About stack access profundity Mark Wills <forthfreak@gmail.com> - 2013-02-03 23:10 -0800
Re: About stack access profundity Andrew Haley <andrew29@littlepinkcloud.invalid> - 2013-02-04 03:26 -0600
Re: About stack access profundity humptydumpty <ouatubi@gmail.com> - 2013-02-03 11:26 -0800
Re: About stack access profundity Pablo Hugo Reda <pabloreda@gmail.com> - 2013-02-03 15:49 -0800
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2013-01-29 17:58 +0000 |
| Message-ID | <2013Jan29.185800@mips.complang.tuwien.ac.at> |
| In reply to | #19254 |
stephenXXX@mpeforth.com (Stephen Pelc) writes:
>What happens is that PICKs produce fetches indexed from the data
>stack pointer,
VFX is better than you give it credit for:
variable A
variable B
variable C
: bla A @ B @ 1 pick + swap drop C ! ;
see bla
shows
( 080BF3A0 8B153C240A08 ) MOV EDX, [080A243C]
( 080BF3A6 031540240A08 ) ADD EDX, [080A2440]
( 080BF3AC 891544240A08 ) MOV [080A2444], EDX
( 080BF3B2 C3 ) NEXT,
i.e., no fetch indexed from the data stack pointer (no reference to
the data stack at all). VFX recognizes that PICK accesses a stack
element in a register and optimizes it away.
>When we rewrote our PowerView embedded GUI to pass structures rather
>than than keep graphics coordinates on the stack, the code (for ARM
>and Cortex) became shorter and faster. In our experience, your
>assertion does not hold bcause the use of structures considerably
>reduces the stack traffic.
I would have to look at the concrete code (before and after the
change) to give a proper comment on that. But if all that happens is
that you replace a "5 PICK" (which should produce a register reference
or, if there are not enough registers, a memory reference to the
memory part of the stack) with something like "DUP .X @", "OVER .X @",
"R@ .X @" (which should produce at least one memory reference, for the
@), the use of structures should not be shorter and faster.
I think that the limited scope of VFXs register allocator reduces the
benefit of stack references, but they still should not hurt.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2013-01-29 12:29 -0600 |
| Message-ID | <iaedncGLNcloiZXMnZ2dnUVZ_qadnZ2d@supernews.com> |
| In reply to | #19256 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > stephenXXX@mpeforth.com (Stephen Pelc) writes: > >>When we rewrote our PowerView embedded GUI to pass structures rather >>than than keep graphics coordinates on the stack, the code (for ARM >>and Cortex) became shorter and faster. In our experience, your >>assertion does not hold bcause the use of structures considerably >>reduces the stack traffic. > > I would have to look at the concrete code (before and after the > change) to give a proper comment on that. But if all that happens is > that you replace a "5 PICK" (which should produce a register reference > or, if there are not enough registers, a memory reference to the > memory part of the stack) with something like "DUP .X @", "OVER .X @", > "R@ .X @" (which should produce at least one memory reference, for the > @), the use of structures should not be shorter and faster. It's unlikely to be just "5 PICK", though. There will be writes too, and that's either POKE (aargh) or lots of stack thrashing to get the data into position. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2013-01-30 17:22 +0000 |
| Message-ID | <2013Jan30.182257@mips.complang.tuwien.ac.at> |
| In reply to | #19258 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> stephenXXX@mpeforth.com (Stephen Pelc) writes:
>>
>>>When we rewrote our PowerView embedded GUI to pass structures rather
>>>than than keep graphics coordinates on the stack, the code (for ARM
>>>and Cortex) became shorter and faster. In our experience, your
>>>assertion does not hold bcause the use of structures considerably
>>>reduces the stack traffic.
>>
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that. But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.
>
>It's unlikely to be just "5 PICK", though. There will be writes too,
Well, in Pablo Reda's code, and that's what we are talking about,
there were only PICKs (for the memory variant, @s), no STICKs/POKEs,
or ROLLs.
>and that's either POKE (aargh) or lots of stack thrashing to get the
>data into position.
Yes, if Stephen's code did that, that might explain the larger code.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2013-01-29 20:17 +0100 |
| Message-ID | <1950087.W2tCE9ddpV@sunwukong.fritz.box> |
| In reply to | #19256 |
Anton Ertl wrote:
> I would have to look at the concrete code (before and after the
> change) to give a proper comment on that. But if all that happens is
> that you replace a "5 PICK" (which should produce a register reference
> or, if there are not enough registers, a memory reference to the
> memory part of the stack) with something like "DUP .X @", "OVER .X @",
> "R@ .X @" (which should produce at least one memory reference, for the
> @), the use of structures should not be shorter and faster.
Not convinced. VFX doesn't do that too well:
begin-structure point ok-2
field: .x ok-2
field: .y ok-2
end-structure ok
: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ; ok
see test
TEST
( 080BC940 53 ) PUSH EBX
( 080BC941 8B1424 ) MOV EDX, [ESP]
( 080BC944 8B4A04 ) MOV ECX, [EDX+04]
( 080BC947 030B ) ADD ECX, 0 [EBX]
( 080BC949 8B0424 ) MOV EAX, [ESP]
( 080BC94C 8D6DF4 ) LEA EBP, [EBP+-0C]
( 080BC94F 894D00 ) MOV [EBP], ECX
( 080BC952 8B4A04 ) MOV ECX, [EDX+04]
( 080BC955 894D04 ) MOV [EBP+04], ECX
( 080BC958 8B13 ) MOV EDX, 0 [EBX]
( 080BC95A 895508 ) MOV [EBP+08], EDX
( 080BC95D 8BD8 ) MOV EBX, EAX
( 080BC95F 8B5500 ) MOV EDX, [EBP]
( 080BC962 8913 ) MOV 0 [EBX], EDX
( 080BC964 8B5D08 ) MOV EBX, [EBP+08]
( 080BC967 2B5D04 ) SUB EBX, [EBP+04]
( 080BC96A 5A ) POP EDX
( 080BC96B 895A04 ) MOV [EDX+04], EBX
( 080BC96E 8B5D0C ) MOV EBX, [EBP+0C]
( 080BC971 8D6D10 ) LEA EBP, [EBP+10]
( 080BC974 C3 ) NEXT,
( 53 bytes, 21 instructions )
That's 21 instructions, clearly not what I would have written by hand.
Compare that to bigForth, using the current object pointer OOP:
debugging class point ok
cell var .x ok
cell var .y ok
how: ok
public: : test >o .x @ .y @ 2dup + .x ! - .y ! o> ; ok
disw test Adresse : 268670656
100396C0: push EDI 57
100396C1: mov EDI,EAX 8BF8
100396C3: lodsd AD
100396C4: xchg ESP,ESI 87F4
100396C6: push EAX 50
100396C7: push DWORD PTR $04[EDI]
FF7704
100396CA: mov EAX,$08[EDI] 8B4708
100396CD: mov EDX,[ESP] 8B1424
100396D0: push EAX 50
100396D1: add EAX,EDX 03C2
100396D3: push EAX 50
100396D4: pop DWORD PTR $04[EDI]
8F4704
100396D7: pop EAX 58
100396D8: pop EDX 5A
100396D9: xchg EDX,EAX 92
100396DA: sub EAX,EDX 2BC2
100396DC: push EAX 50
100396DD: pop DWORD PTR $08[EDI]
8F4708
100396E0: pop EAX 58
100396E1: xchg ESP,ESI 87F4
100396E3: pop EDI 5F
100396E4: ret C3
22 instructions, room for improvement, because bigForth isn't an
analytical compiler. What I would expect is that apart from the struct
memory accesses, everything would fit into the registers.
--
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | mhx@iae.nl (Marcel Hendrix) |
|---|---|
| Date | 2013-01-29 21:44 +0200 |
| Message-ID | <63591406028434@frunobulax.edu> |
| In reply to | #19260 |
Bernd Paysan <bernd.paysan@gmx.de> writes Re: About stack access profundity > Anton Ertl wrote: >> I would have to look at the concrete code (before and after the >> change) to give a proper comment on that. But if all that happens is >> that you replace a "5 PICK" (which should produce a register reference >> or, if there are not enough registers, a memory reference to the >> memory part of the stack) with something like "DUP .X @", "OVER .X @", >> "R@ .X @" (which should produce at least one memory reference, for the >> @), the use of structures should not be shorter and faster. > Not convinced. VFX doesn't do that too well: [..] iForth64 ... FORTH> begin-structure point ok [3]FORTH> field: .x ok [3]FORTH> field: .y ok [3]FORTH> end-structure ok FORTH> ok FORTH> : test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ; ok FORTH> see test Flags: TOKENIZE, ANSI : test >R R@ .x @ R@ .y @ 2DUP + R@ .x ! - R> .y ! ; ok FORTH> ' test idis $01404300 : [trashed] $0140430A pop rbx $0140430B mov rdi, [rbx] qword $0140430E add rdi, [rbx 8 +] qword $01404312 mov rax, [rbx] qword $01404315 mov rdx, [rbx 8 +] qword $01404319 mov [rbx] qword, rdi $0140431C sub rax, rdx $0140431F mov [rbx 8 +] qword, rax $01404323 ; FORTH> create ape 2 cells allot ok FORTH> : tt ape test ; ok FORTH> see tt Flags: TOKENIZE, ANSI : tt ape test ; ok FORTH> ' tt idis $01404BC0 : [trashed] $01404BCA mov rbx, $01404780 qword-offset $01404BD1 add rbx, $01404788 qword-offset $01404BD8 mov rdi, $01404780 qword-offset $01404BDF mov rax, $01404788 qword-offset $01404BE6 mov $01404780 qword-offset, rbx $01404BED sub rdi, rax $01404BF0 mov $01404788 qword-offset, rdi $01404BF7 ; -marcel
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2013-01-29 22:14 +0100 |
| Message-ID | <4134598.R9bZG8GDba@sunwukong.fritz.box> |
| In reply to | #19266 |
Marcel Hendrix wrote: > iForth64 ... Great! That's pretty close to what I would write by hand. Hand-code (let's assume rax is tos): mov rbx, [rax] mov rcx, [rax+8] lea rdx, [rbx+rcx] sub rbx, rcx mov [rax], rdx mov [rax+8], rbx Approach: Don't load values twice, though on x86, you have quite a lot of load units. Use lea for add when you need a three operand add. Not sure why your compiler generates mov rax, [rbx] qword mov rdx, [rbx 8 +] qword sub rax, rdx instead of mov rax, [rbx] qword sub rax, [rbx 8 +] qword > FORTH> begin-structure point ok > [3]FORTH> field: .x ok > [3]FORTH> field: .y ok > [3]FORTH> end-structure ok > FORTH> ok > FORTH> : test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! > ; ok FORTH> see test > Flags: TOKENIZE, ANSI > : test >R R@ .x @ R@ .y @ 2DUP + R@ .x ! - R> .y ! ; ok > FORTH> ' test idis > $01404300 : [trashed] > $0140430A pop rbx > $0140430B mov rdi, [rbx] qword > $0140430E add rdi, [rbx 8 +] qword > $01404312 mov rax, [rbx] qword > $01404315 mov rdx, [rbx 8 +] qword > $01404319 mov [rbx] qword, rdi > $0140431C sub rax, rdx > $0140431F mov [rbx 8 +] qword, rax > $01404323 ; > FORTH> create ape 2 cells allot ok > FORTH> : tt ape test ; ok > FORTH> see tt > Flags: TOKENIZE, ANSI > : tt ape test ; ok > FORTH> ' tt idis > $01404BC0 : [trashed] > $01404BCA mov rbx, $01404780 qword-offset > $01404BD1 add rbx, $01404788 qword-offset > $01404BD8 mov rdi, $01404780 qword-offset > $01404BDF mov rax, $01404788 qword-offset > $01404BE6 mov $01404780 qword-offset, rbx > $01404BED sub rdi, rax > $01404BF0 mov $01404788 qword-offset, rdi > $01404BF7 ; > > -marcel -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2013-01-30 16:29 +0000 |
| Message-ID | <2013Jan30.172942@mips.complang.tuwien.ac.at> |
| In reply to | #19260 |
Bernd Paysan <bernd.paysan@gmx.de> writes:
>Anton Ertl wrote:
>> I would have to look at the concrete code (before and after the
>> change) to give a proper comment on that. But if all that happens is
>> that you replace a "5 PICK" (which should produce a register reference
>> or, if there are not enough registers, a memory reference to the
>> memory part of the stack) with something like "DUP .X @", "OVER .X @",
>> "R@ .X @" (which should produce at least one memory reference, for the
>> @), the use of structures should not be shorter and faster.
>
>Not convinced.
Of what are you are not convinced?
> VFX doesn't do that too well:
>
>begin-structure point ok-2
>field: .x ok-2
>field: .y ok-2
>end-structure ok
>: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ; ok
>see test
>TEST
>( 080BC940 53 ) PUSH EBX
>( 080BC941 8B1424 ) MOV EDX, [ESP]
>( 080BC944 8B4A04 ) MOV ECX, [EDX+04]
>( 080BC947 030B ) ADD ECX, 0 [EBX]
>( 080BC949 8B0424 ) MOV EAX, [ESP]
>( 080BC94C 8D6DF4 ) LEA EBP, [EBP+-0C]
>( 080BC94F 894D00 ) MOV [EBP], ECX
>( 080BC952 8B4A04 ) MOV ECX, [EDX+04]
>( 080BC955 894D04 ) MOV [EBP+04], ECX
>( 080BC958 8B13 ) MOV EDX, 0 [EBX]
>( 080BC95A 895508 ) MOV [EBP+08], EDX
>( 080BC95D 8BD8 ) MOV EBX, EAX
>( 080BC95F 8B5500 ) MOV EDX, [EBP]
>( 080BC962 8913 ) MOV 0 [EBX], EDX
>( 080BC964 8B5D08 ) MOV EBX, [EBP+08]
>( 080BC967 2B5D04 ) SUB EBX, [EBP+04]
>( 080BC96A 5A ) POP EDX
>( 080BC96B 895A04 ) MOV [EDX+04], EBX
>( 080BC96E 8B5D0C ) MOV EBX, [EBP+0C]
>( 080BC971 8D6D10 ) LEA EBP, [EBP+10]
>( 080BC974 C3 ) NEXT,
>( 53 bytes, 21 instructions )
>
>That's 21 instructions, clearly not what I would have written by hand.
Yes, so VFX is not as great as we might like, but that does not tell
us anything about whether it does better for PICKing or for @ing code.
But let's try it. I wanted to use the original example
<61d42b93-0ac8-4df3-8cef-6c1cc059d0ef@googlegroups.com> for this, but
it's unclear to me what it does, in particular the line
4 <? ( drop line 2drop line ; ) drop
and PX and PY.
So I fell back to the good old rectangle example:
begin-structure point
field: point-x
field: point-y
end-structure
defer line ( p1 p2 -- )
defer make-point ( x y -- p )
defer free-point ( p -- )
: line-line ( p1 p2 p3 -- )
\ draw a line between p1 and p2
\ and one between p2 and p3
over line line ;
: rect-mem ( ll ur -- )
over point-x @ over point-y @ make-point
( ll ur ul )
>r 2dup r@ swap line-line r> free-point
over point-y @ over point-x @ swap make-point
( ll ur lr )
>r 2dup r@ swap line-line r> free-point ;
defer line-stack ( x1 y1 x2 y2 -- )
: rect-local {: x1 y1 x2 y2 -- :}
x1 y1 x1 y2 line-stack
x1 y2 x2 y2 line-stack
x2 y2 x2 y1 line-stack
x2 y1 x1 y1 line-stack ;
: rect-stack ( x1 y1 x2 y2 -- )
3 pick 3 pick over 3 pick line-stack
3 pick over 3 pick over line-stack
over over over 3 pick line-stack
over 3 pick 3 pick 3 pick line-stack
2drop 2drop ;
see rect-mem
see rect-local
see rect-stack
and SEEing the results showed:
RECT-MEM
...
( 147 bytes, 46 instructions )
RECT-LOCAL
...
( 166 bytes, 54 instructions )
RECT-STACK
...
( 121 bytes, 37 instructions )
If you want a rect-mem2 that calls line-stack and is thus closer to
rect-stack and rect-local in what it does internally, feel free to
post it and I'll run it through VFX.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Pablo Hugo Reda <pabloreda@gmail.com> |
|---|---|
| Date | 2013-01-30 09:58 -0800 |
| Message-ID | <db5cb315-6eb1-4a2e-96e0-352dc58103ba@googlegroups.com> |
| In reply to | #19285 |
> but > > it's unclear to me what it does, in particular the line > > > > 4 <? ( drop line 2drop line ; ) drop > > > > and PX and PY. > > sorry for have a forth dialect, not at ans-forth I take the ideas from colorforth, I not use STATE, not use DOES>, etc. PX and PY are variables, the initial point for the curve, LINE draw a line and update the PX and PY. 4 <? (.. go inside (..) when the Top of stack is <4 (and consume 4) here is the code generated by the compiler, in FASM syntax. I remove the comments because have 200 lines with this. "uso" is the stack profundity and "dD" is the stack variation ------------------------------------------------------------ w24: ; ::: sp-dist ::: uso:-4 dD:1 mov ebx,dword [esi+8] sub ebx,dword [esi] mov edx,ebx sar edx,31 add ebx,edx xor ebx,edx mov ecx,dword [esi+4] sub ecx,eax mov edx,ecx sar edx,31 add ecx,edx xor ecx,edx add ebx,ecx lea esi,[esi-4] mov dword [esi],eax mov eax,ebx ret w25: ; ::: sp-c ::: uso:-4 dD:2 mov ebx,dword [esi] sal ebx,1 mov ecx,dword [esi+8] add ecx,ebx add ecx,dword [w9] sar ecx,$2 mov ebx,eax sal ebx,1 mov edx,dword [esi+4] add edx,ebx add edx,dword [wA] sar edx,$2 lea esi,[esi-8] mov dword [esi+4],eax mov dword [esi],ecx mov eax,edx ret w26: ; ::: spl ::: uso:-4 dD:-4 call w25 call w24 cmp eax,$4 jge _46 lea esi,[esi+4] mov eax,dword [esi-4] call w23 lea esi,[esi+8] mov eax,dword [esi-4] jmp w23 _46: push dword [esi] push dword [esi+4] mov eax,dword [esi+20] add eax,dword [esi+12] sar eax,1 mov ebx,dword [esi+16] add ebx,dword [esi+8] sar ebx,1 mov ecx,dword [esi+8] add ecx,dword [wA] sar ecx,1 mov edx,dword [esi+12] add edx,dword [w9] sar edx,1 pop edi mov [esi+8],eax pop eax lea esi,[esi-4] xchg dword [esi+12],ebx mov dword [esi+8],edi mov dword [esi+4],eax mov dword [esi],edx mov dword [esi+16],ebx mov eax,ecx call w26 jmp w26 ---------------------------------------------------- sure Hendrix and others compiler is better but I not finish. have more ideas for improvement. I don't have the solution, but thank's for all for comment.
[toc] | [prev] | [next] | [standalone]
| From | stephenXXX@mpeforth.com (Stephen Pelc) |
|---|---|
| Date | 2013-01-31 12:32 +0000 |
| Message-ID | <510a62ea.510135172@192.168.0.50> |
| In reply to | #19260 |
On Tue, 29 Jan 2013 20:17:54 +0100, Bernd Paysan <bernd.paysan@gmx.de> wrote: >Anton Ertl wrote: >> I would have to look at the concrete code (before and after the >> change) to give a proper comment on that. But if all that happens is >> that you replace a "5 PICK" (which should produce a register reference >> or, if there are not enough registers, a memory reference to the >> memory part of the stack) with something like "DUP .X @", "OVER .X @", >> "R@ .X @" (which should produce at least one memory reference, for the >> @), the use of structures should not be shorter and faster. > >Not convinced. VFX doesn't do that too well: > >begin-structure point ok-2 >field: .x ok-2 >field: .y ok-2 >end-structure ok >: test ( addr -- ) >r r@ .x @ r@ .y @ 2dup + r@ .x ! - r> .y ! ; ok >see test >TEST >( 080BC940 53 ) PUSH EBX >( 080BC941 8B1424 ) MOV EDX, [ESP] >( 080BC944 8B4A04 ) MOV ECX, [EDX+04] >( 080BC947 030B ) ADD ECX, 0 [EBX] >( 080BC949 8B0424 ) MOV EAX, [ESP] >( 080BC94C 8D6DF4 ) LEA EBP, [EBP+-0C] >( 080BC94F 894D00 ) MOV [EBP], ECX >( 080BC952 8B4A04 ) MOV ECX, [EDX+04] >( 080BC955 894D04 ) MOV [EBP+04], ECX >( 080BC958 8B13 ) MOV EDX, 0 [EBX] >( 080BC95A 895508 ) MOV [EBP+08], EDX >( 080BC95D 8BD8 ) MOV EBX, EAX >( 080BC95F 8B5500 ) MOV EDX, [EBP] >( 080BC962 8913 ) MOV 0 [EBX], EDX >( 080BC964 8B5D08 ) MOV EBX, [EBP+08] >( 080BC967 2B5D04 ) SUB EBX, [EBP+04] >( 080BC96A 5A ) POP EDX >( 080BC96B 895A04 ) MOV [EDX+04], EBX >( 080BC96E 8B5D0C ) MOV EBX, [EBP+0C] >( 080BC971 8D6D10 ) LEA EBP, [EBP+10] >( 080BC974 C3 ) NEXT, >( 53 bytes, 21 instructions ) > >That's 21 instructions, clearly not what I would have written by hand. The code from the LEA up to but not including the SUB is a stack shuffle (we ran out of registers) followed by a store. Although it's possible to do it without the shuffle, VFX's "rip up and retry" rules are not clever enough here. Marcel's iForth version is for 64 bit code with 16 registers. This means more registers, so no shuffle, and because there are more registers, some can be used for the return stack. Nice code. Stephen -- Stephen Pelc, stephenXXX@mpeforth.com MicroProcessor Engineering Ltd - More Real, Less Time 133 Hill Lane, Southampton SO15 5AF, England tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691 web: http://www.mpeforth.com - free VFX Forth downloads
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2013-01-31 15:37 +0100 |
| Message-ID | <3427167.9fDfFC78xt@sunwukong.fritz.box> |
| In reply to | #19310 |
Stephen Pelc wrote: > The code from the LEA up to but not including the SUB is a stack > shuffle (we ran out of registers) followed by a store. Although it's > possible to do it without the shuffle, VFX's "rip up and retry" > rules are not clever enough here. You need four free registers (including TOS) to fit this code in. In bigForth, I have three free registers, because I have SP, RP, UP, OP (object pointer) and an index for loops in a register. VFX doesn't have an OP, and doesn't waste another register for the index, so it should be possible to fit it into the available registers. The main thing you need to do is to try hard to eliminate redundancy - all active values should only occupy one single register, unless a copy is dearly needed. > Marcel's iForth version is for 64 bit code with 16 registers. This > means more registers, so no shuffle, and because there are more > registers, some can be used for the return stack. Nice code. Yes, x64 is a much better target, because you have 8 really free registers (in x86, there is no register without a special role, though EBX is only used in xlat, and xlat is really superfluous - however, you can argue that x86 has 8 special purpose registers, and no general purpose register at all). This is also impacting C compilers. The string instructions use up ECX, ESI, and EDI, multiplication uses up EAX and EDX; -fomit-frame-pointer gives you EBP, so you have EBP and EBX, no more. In Gforth, we therefore moved all string operations into real subroutines (even though that makes them a bit slower), and only then the C compiler is able to fit in the most important things. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2013-01-29 23:12 -0800 |
| Message-ID | <680dad03-965e-4357-b2fc-97a0783deeca@w7g2000yqo.googlegroups.com> |
| In reply to | #19256 |
On Jan 29, 5:58 pm, an...@mips.complang.tuwien.ac.at (Anton Ertl) wrote: > stephen...@mpeforth.com (Stephen Pelc) writes: > >What happens is that PICKs produce fetches indexed from the data > >stack pointer, > > VFX is better than you give it credit for: > > variable A > variable B > variable C > : bla A @ B @ 1 pick + swap drop C ! ; > see bla > > shows > > ( 080BF3A0 8B153C240A08 ) MOV EDX, [080A243C] > ( 080BF3A6 031540240A08 ) ADD EDX, [080A2440] > ( 080BF3AC 891544240A08 ) MOV [080A2444], EDX > ( 080BF3B2 C3 ) NEXT, > > i.e., no fetch indexed from the data stack pointer (no reference to > the data stack at all). VFX recognizes that PICK accesses a stack > element in a register and optimizes it away. > > >When we rewrote our PowerView embedded GUI to pass structures rather > >than than keep graphics coordinates on the stack, the code (for ARM > >and Cortex) became shorter and faster. In our experience, your > >assertion does not hold bcause the use of structures considerably > >reduces the stack traffic. > > I would have to look at the concrete code (before and after the > change) to give a proper comment on that. But if all that happens is > that you replace a "5 PICK" (which should produce a register reference > or, if there are not enough registers, a memory reference to the > memory part of the stack) with something like "DUP .X @", "OVER .X @", > "R@ .X @" (which should produce at least one memory reference, for the > @), the use of structures should not be shorter and faster. > > I think that the limited scope of VFXs register allocator reduces the > benefit of stack references, but they still should not hurt. > > - anton > -- > M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html > comp.lang.forth FAQs:http://www.complang.tuwien.ac.at/forth/faq/toc.html > New standard:http://www.forth200x.org/forth200x.html > EuroForth 2012:http://www.euroforth.org/ef12/ Exactly the point I was making on another thread. An obsessive programmer might second guess the compiler, and instead of writing plain, vanilla, straight-ahead code, come up with something that he *perceives* to be faster, but instead, it's slower because the optimiser can't identify what the programmers real intention was. Unless you are using a dumb ITC compiler, just write the code.
[toc] | [prev] | [next] | [standalone]
| From | stephenXXX@mpeforth.com (Stephen Pelc) |
|---|---|
| Date | 2013-01-29 10:08 +0000 |
| Message-ID | <51079f62.329007481@192.168.0.50> |
| In reply to | #19236 |
On Tue, 29 Jan 2013 08:20:47 GMT, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote: >There's a good reason why decent compilers will produce worse code for >@ than for PICK. Whether the worseness is "much" or not can be >debated forever without resolution, so let's not go there. Please explain this. Stephen -- Stephen Pelc, stephenXXX@mpeforth.com MicroProcessor Engineering Ltd - More Real, Less Time 133 Hill Lane, Southampton SO15 5AF, England tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691 web: http://www.mpeforth.com - free VFX Forth downloads
[toc] | [prev] | [next] | [standalone]
| From | Pablo Hugo Reda <pabloreda@gmail.com> |
|---|---|
| Date | 2013-01-29 07:10 -0800 |
| Message-ID | <60c59889-21bf-446e-beb8-303e5af3132d@googlegroups.com> |
| In reply to | #19243 |
> > >There's a good reason why decent compilers will produce worse code for > > >@ than for PICK. Whether the worseness is "much" or not can be > > >debated forever without resolution, so let's not go there. > stack operations disappears in the compiled code (for me implementation) but memory access not (for now)
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2013-01-29 15:17 +0000 |
| Message-ID | <2013Jan29.161718@mips.complang.tuwien.ac.at> |
| In reply to | #19243 |
stephenXXX@mpeforth.com (Stephen Pelc) writes:
>On Tue, 29 Jan 2013 08:20:47 GMT, anton@mips.complang.tuwien.ac.at
>(Anton Ertl) wrote:
>
>>There's a good reason why decent compilers will produce worse code for
>>@ than for PICK. Whether the worseness is "much" or not can be
>>debated forever without resolution, so let's not go there.
>
>Please explain this.
A decent compiler will produce a load-from-memory for a @, while stack
items can be kept in registers and thus don't require a load (PICK and
ROLL with non-constant u are an exception, but that's not what the OP
did.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2013-01-29 00:57 -0800 |
| Message-ID | <c696cb3b-d0cc-4151-806c-a731a96c4715@w3g2000yqj.googlegroups.com> |
| In reply to | #19222 |
On Jan 28, 3:23 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote: > > I think the mistake you're making is keeping everything on the stack. > > > As you say, objects deep on the stack are hare to access. > > > Instead, use a structure with named offsets for your co-ordinates and > > > vectors. You'll find the code is much easier to read and write. > > the stack if the more fast access and I need here because the recursion use this estructure > > > Pass pointers, rather than the parameters directly? > > but with address need a @ and ! to load store values. > > when the compiler optimice the best place is the stack > > look the cuadric bezier > ------------------------------------------------------------------- > :sp-dist | x y xe ye -- x y xe ye dd > pick3 pick2 - abs pick3 pick2 - abs + ; > > :sp-c | fx fy cx cy -- fx fy cx cy xn yn ; xn=(cx*2+fx+px)/4 > pick3 pick2 2* + px + 2 >> > pick3 pick2 2* + py + 2 >> ; > > :spl | fx fy cx cy -- > sp-c sp-dist > 4 <? ( drop line 2drop line ; ) drop > >r >r > pick3 pick2 + 2/ pick3 pick2 + 2/ | fx fy cx cy c2 c2 > 2swap | fx fy c2 c2 cx cy > py + 2/ swap px + 2/ swap | fx fy c2 c2 c1 c1 > r> r> 2swap > spl spl ; > > ------------------------------------------------------------------- This code is crying out for locals! Use locals. All your PICKs will go away :-)
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2013-01-28 01:56 -0800 |
| Message-ID | <7168762f-2b01-4827-929a-6725775caf6d@l13g2000yqe.googlegroups.com> |
| In reply to | #19200 |
On Jan 27, 8:11 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote: > For two definition who make the same, the less distance in stack access is more optimal, need less register when compiler, etc. Is convenient make definition with less profundity access en the stack > . > I choose limit the word PICK and make it static, I have only PICK2,PICK3,PICK4 and no more. > Whell, until now, not need more, but now I found a word who need more PICK's. > > I reeimplement the Quadratic Bezier Curve word and try to draw Cubic Bezier Curves now. > I use recursion, only integers and add,shift and compare. > > Because I use recursion, I need 6 parameters (3 2D points) and the rutine access to 8 position (I need PICK8 definition). > > I think about workaround the problem: > > if compress the 2dvector for use one place, the word only need PICK4, work in the current system! > but this add the time for change representation (2 cells to 1 cell..) or extact (address to x y..). > > but when PICK8 the solution is more direct. > > I need advice.. > thank's Pass pointers, rather than the parameters directly?
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2013-01-29 15:30 +0000 |
| Message-ID | <2013Jan29.163037@mips.complang.tuwien.ac.at> |
| In reply to | #19200 |
I published a paper that discusses various ways to deal with needing
to deal with too much data at once.
@InProceedings{ertl11euroforth,
author = {M. Anton Ertl},
title = {Ways to Reduce the Stack Depth},
crossref = {euroforth11},
pages = {36--41},
url = {http://www.complang.tuwien.ac.at/papers/ertl11euroforth.ps.gz},
url2 = {http://www.complang.tuwien.ac.at/anton/euroforth/ef11/papers/ertl.pdf},
OPTnote = {not refereed},
abstract = {Having to deal with many different data can lead to
problems in Forth: The data stack is the preferred
place to store data; on the other hand, dealing with
too many data stack items is cumbersome and usually
bad style. This paper presents and discusses ways to
unburden the data stack; some of them are used
widely, others are almost unknown or new.}
}
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Pablo Hugo Reda <pabloreda@gmail.com> |
|---|---|
| Date | 2013-01-29 09:06 -0800 |
| Message-ID | <f6b5ce4f-4548-4683-a2a4-3a40b9ff8917@googlegroups.com> |
| In reply to | #19249 |
El martes, 29 de enero de 2013 12:30:37 UTC-3, Anton Ertl escribió:
> I published a paper that discusses various ways to deal with needing
>
> to deal with too much data at once.
>
>
>
> @InProceedings{ertl11euroforth,
>
> author = {M. Anton Ertl},
>
> title = {Ways to Reduce the Stack Depth},
>
> crossref = {euroforth11},
>
> pages = {36--41},
>
> url = {http://www.complang.tuwien.ac.at/papers/ertl11euroforth.ps.gz},
>
> url2 = {http://www.complang.tuwien.ac.at/anton/euroforth/ef11/papers/ertl.pdf},
>
> OPTnote = {not refereed},
>
> abstract = {Having to deal with many different data can lead to
>
> problems in Forth: The data stack is the preferred
>
> place to store data; on the other hand, dealing with
>
> too many data stack items is cumbersome and usually
>
> bad style. This paper presents and discusses ways to
>
> unburden the data stack; some of them are used
>
> widely, others are almost unknown or new.}
>
> }
>
>
>
> - anton
>
> --
>
> M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
>
> comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
>
> New standard: http://www.forth200x.org/forth200x.html
>
> EuroForth 2012: http://www.euroforth.org/ef12/
Thank's Anton
reading...
[toc] | [prev] | [next] | [standalone]
| From | Pablo Hugo Reda <pabloreda@gmail.com> |
|---|---|
| Date | 2013-02-03 06:22 -0800 |
| Message-ID | <722c944e-7658-4603-a7ab-e696573ac34d@googlegroups.com> |
| In reply to | #19200 |
Well, at last using some vars and reorganize I get a version without more picks there is the test version #x1 #y1 #x2 #y2 #bx #by :curve3 | xf yf x2 y2 x1 y1 pick3 pick2 + 2/ pick3 pick2 + 2/ 'by ! 'bx ! 'y1 ! 'x1 ! pick3 pick2 + 2/ pick3 pick2 + 2/ 2swap 'y2 ! 'x2 ! over bx + 2/ over by + 2/ over x2 - abs over y2 - abs + >r px x1 + 2/ py y1 + 2/ over bx + 2/ over by + 2/ over x1 - abs over y1 - abs + >r 2swap >r >r pick3 pick2 + 2/ pick3 pick2 + 2/ 2swap r> r> r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line )( drop curve3 ) r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line ; ) drop curve3 ;
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2013-02-03 07:25 -0800 |
| Message-ID | <5b5748f3-5a9c-4768-a28a-476b11263282@5g2000yqz.googlegroups.com> |
| In reply to | #19388 |
On Feb 3, 2:22 pm, Pablo Hugo Reda <pablor...@gmail.com> wrote: > Well, at last > > using some vars and reorganize I get a version without more picks > there is the test version > > #x1 #y1 #x2 #y2 #bx #by > > :curve3 | xf yf x2 y2 x1 y1 > pick3 pick2 + 2/ pick3 pick2 + 2/ 'by ! 'bx ! > 'y1 ! 'x1 ! > pick3 pick2 + 2/ pick3 pick2 + 2/ 2swap > 'y2 ! 'x2 ! > over bx + 2/ over by + 2/ > over x2 - abs over y2 - abs + >r > px x1 + 2/ py y1 + 2/ > over bx + 2/ over by + 2/ > over x1 - abs over y1 - abs + >r > 2swap >r >r > pick3 pick2 + 2/ pick3 pick2 + 2/ > 2swap r> r> > r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line )( drop curve3 ) > r> 3 <? ( drop 4drop 2dup 'py ! 'px ! line ; ) > drop curve3 ; This looks truly horrible to me! Utterly unreadable :-\
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.lang.forth
csiph-web