Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #135490
| From | peter <peter.noreply@tin.it> |
|---|---|
| Newsgroups | comp.lang.forth |
| Subject | Re: F>R, FR@ and FR> Float to return stack words |
| Date | 2026-08-30 10:06 +0200 |
| Organization | A noiseless patient Spider |
| Message-ID | <20260830100637.000013f4@tin.it> (permalink) |
| References | <20260828000428.00007e53@tin.it> |
On Fri, 28 Aug 2026 00:04:28 +0200
peter <peter.noreply@tin.it> wrote:
>
> I recently wanted to adapt compelx-kahan.fs, the complex implementation
> by Julian V. Noble and David N. Williams, to the capabilities of
> ntf64/lxf64.
>
> One thing I wanted to get away from was the Baden-style float locals,
> fa fb .. and to-fa to-fb .. and the use of temporary buffers.
>
> I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
> This works well as both stacks are 64 bits. The effect was very good.
>
>
> : z* ( f: x y u v -- x*u-y*v x*v+y*u)
> zdup f+ f>r ( f: x y u v r: u+v)
> fover f- f>r ( f: x y u r: u+v v-u)
> frot fdup fr> ( f: y u x x v-u r: u+v)
> f* ( f: y u x x*'v-u' r: u+v)
> fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
> frot fdup fr> ( f: u x y y u+v r: x*'v-u')
> f* f>r ( f: u x y r: x*'v-u' y*'u+v')
> f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
> f+ ; ( f: x*u-y*v x*v+y*u)
>
>
> There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer
>
> The revised version gives
>
> seea Z*
> 0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
> 0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
> 0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
> 0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8]
> 0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
> 0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10]
> 0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
> 0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
> 0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
> 0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
> 0x428A88 4D8D6D10 lea r13, [r13+0x10]
> 0x428A8C C3 ret
> 61 bytes, 12 instructions
> ok
>
> Half the size and all memory loads/stores gone.
>
I went ahead and implemented float locals.
Now I can write using JVN formula
: Z* {f: x y u v :}
x y f+ u f* fdup u v f+ y f* f- fswap v u f- x f* f+ ;
this compiles to
seea Z*
0x42A6E0 C4C17B104D08 vmovsd xmm1, qword ptr [r13+0x8]
0x42A6E6 C4C173584D10 vaddsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A6EC C4C173594D00 vmulsd xmm1, xmm1, qword ptr [r13]
0x42A6F2 C4C17B585500 vaddsd xmm2, xmm0, qword ptr [r13]
0x42A6F8 C4C16B595508 vmulsd xmm2, xmm2, qword ptr [r13+0x8]
0x42A6FE C5F35CDA vsubsd xmm3, xmm1, xmm2
0x42A702 C4C17B5C5500 vsubsd xmm2, xmm0, qword ptr [r13]
0x42A708 C4C16B595510 vmulsd xmm2, xmm2, qword ptr [r13+0x10]
0x42A70E C5EB58D1 vaddsd xmm2, xmm2, xmm1
0x42A712 C5EB10C2 vmovsd xmm0, xmm2, xmm2
0x42A716 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
0x42A71C 4D8D6D10 lea r13, [r13+0x10]
0x42A720 C3 ret
65 bytes, 13 instructions
4 bytes and 1 instruction more
If I do the school book formual I get
: Z* {f: x y u v :}
x u f* y v f* f- x v f* y u f* f+ ;
Which give more compact code
seea Z*
0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
0x42A766 4D8D6D10 lea r13, [r13+0x10]
0x42A76A C3 ret
59 bytes, 12 instructions
4 multiplications 2 add/sub
I tried to measure the execution times. This shows no difference!
: z1 0 do zdup zdup z* zdrop loop ;
fpi 345e4 timer-reset 100000000 z1 .elapsed 179 ms elapsed ok
fpi 345e4 timer-reset 100000000 z2 .elapsed 176 ms elapsed ok
fpi 345e4 timer-reset 100000000 z3 .elapsed 177 ms elapsed ok
I had to start with the complex number already on the stack!
If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
BR
Peter
> An alternative I considered was to introduce float locals. It would probably
> have been more complicated. The result, in the best case, would be equal in
> the generated code. The problem with locals is the scope, that is to the end
> of the word. Of course the source would be more readable!
>
> I might still implement float locals.
>
> What do you have in your system?
>
> BR
> Peter
>
Back to comp.lang.forth | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-28 00:04 +0200
Re: F>R, FR@ and FR> Float to return stack words dxf <dxforth@gmail.com> - 2026-08-28 14:44 +1000
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-28 06:50 +0000
Re: F>R, FR@ and FR> Float to return stack words marcel hendrix <mhx@iae.nl> - 2026-08-28 19:35 +0200
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 10:06 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 13:31 +0000
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 15:58 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 15:12 +0000
csiph-web