Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #135446 > unrolled thread
| Started by | peter <peter.noreply@tin.it> |
|---|---|
| First post | 2026-08-28 00:04 +0200 |
| Last post | 2026-08-30 15:12 +0000 |
| Articles | 8 — 4 participants |
Back to article view | Back to comp.lang.forth
F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-28 00:04 +0200
Re: F>R, FR@ and FR> Float to return stack words dxf <dxforth@gmail.com> - 2026-08-28 14:44 +1000
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-28 06:50 +0000
Re: F>R, FR@ and FR> Float to return stack words marcel hendrix <mhx@iae.nl> - 2026-08-28 19:35 +0200
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 10:06 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 13:31 +0000
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 15:58 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 15:12 +0000
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-08-28 00:04 +0200 |
| Subject | F>R, FR@ and FR> Float to return stack words |
| Message-ID | <20260828000428.00007e53@tin.it> |
I recently wanted to adapt compelx-kahan.fs, the complex implementation
by Julian V. Noble and David N. Williams, to the capabilities of
ntf64/lxf64.
One thing I wanted to get away from was the Baden-style float locals,
fa fb .. and to-fa to-fb .. and the use of temporary buffers.
I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
This works well as both stacks are 64 bits. The effect was very good.
Taking Z* as an example:
Original Z* using 2 temporary buffers
: z* (f: x y u v -- x*u-y*v x*v+y*u)
(
Uses the algorithm followed by jvn:
[x+iy]*[u+iv] = [[x+y]*u - y*[u+v]] + i[[x+y]*u + x*[v-u]]
Requires 3 multiplications and 5 additions.
)
zdup f+ (f: x y u v u+v)
[ fnoname ] literal f! (f: x y u v)
fover f- (f: x y u v-u)
[ fnoname float+ ] literal f! (f: x y u)
frot fdup (f: y u x x)
[ fnoname float+ ] literal f@ (f: y u x x v-u)
f*
[ fnoname float+ ] literal f! (f: y u x)
frot fdup (f: u x y y)
[ fnoname ] literal f@ (f: u x y y u+v)
f*
[ fnoname ] literal f! (f: u x y)
f+ f* fdup (f: u*[x+y] u*[x+y])
[ fnoname ] literal f@ f- (f: u*[x+y] x*u-y*v)
fswap
[ fnoname float+ ] literal f@ (f: x*u-y*v u*[x+y] x*[v-u])
f+ ; (f: x*u-y*v x*v+y*u)
With more or less direct exchange to using the return stack
I got
: z* ( f: x y u v -- x*u-y*v x*v+y*u)
zdup f+ f>r ( f: x y u v r: u+v)
fover f- f>r ( f: x y u r: u+v v-u)
frot fdup fr> ( f: y u x x v-u r: u+v)
f* ( f: y u x x*'v-u' r: u+v)
fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
frot fdup fr> ( f: u x y y u+v r: x*'v-u')
f* f>r ( f: u x y r: x*'v-u' y*'u+v')
f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
f+ ; ( f: x*u-y*v x*v+y*u)
The original give the following assembler code
seea Z*
0x4291A0 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
0x4291A6 C5FB110DD2965D00 vmovsd qword ptr [rel 0xA02880], xmm1
0x4291AE C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
0x4291B4 C5FB1105CC965D00 vmovsd qword ptr [rel 0xA02888], xmm0
0x4291BC C5FB1005C4965D00 vmovsd xmm0, qword ptr [rel 0xA02888]
0x4291C4 C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
0x4291CA C5FB1105B6965D00 vmovsd qword ptr [rel 0xA02888], xmm0
0x4291D2 C5FB1005A6965D00 vmovsd xmm0, qword ptr [rel 0xA02880]
0x4291DA C4C17B594508 vmulsd xmm0, xmm0, qword ptr [r13+0x8]
0x4291E0 C5FB110598965D00 vmovsd qword ptr [rel 0xA02880], xmm0
0x4291E8 C4C17B104508 vmovsd xmm0, qword ptr [r13+0x8]
0x4291EE C4C17B584510 vaddsd xmm0, xmm0, qword ptr [r13+0x10]
0x4291F4 C4C17B594500 vmulsd xmm0, xmm0, qword ptr [r13]
0x4291FA C5FB100D7E965D00 vmovsd xmm1, qword ptr [rel 0xA02880]
0x429202 C5FB5CD1 vsubsd xmm2, xmm0, xmm1
0x429206 C5FB100D7A965D00 vmovsd xmm1, qword ptr [rel 0xA02888]
0x42920E C5F358C8 vaddsd xmm1, xmm1, xmm0
0x429212 C5F310C1 vmovsd xmm0, xmm1, xmm1
0x429216 C4C17B115510 vmovsd qword ptr [r13+0x10], xmm2
0x42921C 4D8D6D10 lea r13, [r13+0x10]
0x429220 C3 ret
129 bytes, 21 instructions
ok
There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer
The revised version gives
seea Z*
0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8]
0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10]
0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
0x428A88 4D8D6D10 lea r13, [r13+0x10]
0x428A8C C3 ret
61 bytes, 12 instructions
ok
Half the size and all memory loads/stores gone.
With the returns stack use I can also do things like
: +ulp f>r r> 1+ >r fr> ; ok
0e +ulp e. 5e-324 ok
0e +ulp fh. 0x0.0000000000001p-1022 ok
seea +ulp
0x42A260 C4E1F97EC0 vmovq rax, xmm0
0x42A265 48FFC0 inc rax
0x42A268 C4E1F96EC0 vmovq xmm0, rax
0x42A26D C3 ret
14 bytes, 4 instructions
ok
An alternative I considered was to introduce float locals. It would probably
have been more complicated. The result, in the best case, would be equal in
the generated code. The problem with locals is the scope, that is to the end
of the word. Of course the source would be more readable!
I might still implement float locals.
What do you have in your system?
BR
Peter
[toc] | [next] | [standalone]
| From | dxf <dxforth@gmail.com> |
|---|---|
| Date | 2026-08-28 14:44 +1000 |
| Message-ID | <6a911235$1@news.ausics.net> |
| In reply to | #135446 |
On 28/08/2026 8:04 am, peter wrote: > ... > What do you have in your system? Because return-stacking is more complicated, I use fvariable where I can. Where I can't, I have the following def's to fall back on. I've not played enough with Baden-type flocals to get a sense of what they're like in practice. The varying HERE (PAD, HOLD etc) might be an issue. \ x86 code F>R ( r -- ) 1 floats # bp sub bp push ' f! ) jmp end-code code FR> ( -- r ) bp push c: f@ ;c 1 floats # bp add next end-code \ 8080 code F>R ( r -- ) rpp lhld -1 floats d lxi d dad rpp shld h push ' f! jmp end-code code FR> ( -- r ) rpp lhld h push c: f@ ;c rpp lhld 1 floats d lxi d dad rpp shld next end-code
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2026-08-28 06:50 +0000 |
| Message-ID | <2026Aug28.085007@mips.complang.tuwien.ac.at> |
| In reply to | #135446 |
peter <peter.noreply@tin.it> writes:
>: z* ( f: x y u v -- x*u-y*v x*v+y*u)
> zdup f+ f>r ( f: x y u v r: u+v)
> fover f- f>r ( f: x y u r: u+v v-u)
> frot fdup fr> ( f: y u x x v-u r: u+v)
> f* ( f: y u x x*'v-u' r: u+v)
> fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
> frot fdup fr> ( f: u x y y u+v r: x*'v-u')
> f* f>r ( f: u x y r: x*'v-u' y*'u+v')
> f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
> f+ ; ( f: x*u-y*v x*v+y*u)
[...]
>seea Z*
>0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
>0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
>0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
>0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8]
>0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
>0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10]
>0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
>0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
>0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
>0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
>0x428A88 4D8D6D10 lea r13, [r13+0x10]
>0x428A8C C3 ret
>61 bytes, 12 instructions
Nice.
>An alternative I considered was to introduce float locals. It would probably
>have been more complicated. The result, in the best case, would be equal in
>the generated code. The problem with locals is the scope, that is to the end
>of the word. Of course the source would be more readable!
>
>I might still implement float locals.
>
>What do you have in your system?
Gforth has had FP locals since 1994. We recently added F>R, FR> and
FR@. The reason why we did not have them was that the return-stack
pointer is not necessarily aligned for FP loads and stores, so on many
hardware platforms of the 1990s F>R and FR> would have been expensive
operations.
With 64-bit cells becoming mainstream, 64-bit floats (what we use in
Gforth), and the death of most architectures that do not support
unaligned accesses, the balance has changed in the meantime.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
[toc] | [prev] | [next] | [standalone]
| From | marcel hendrix <mhx@iae.nl> |
|---|---|
| Date | 2026-08-28 19:35 +0200 |
| Message-ID | <116sgt6$ll12$1@dont-email.me> |
| In reply to | #135446 |
On 8/28/2026 12:04 AM, peter wrote: > > I recently wanted to adapt compelx-kahan.fs, the complex implementation > by Julian V. Noble and David N. Williams, to the capabilities of > ntf64/lxf64. > > One thing I wanted to get away from was the Baden-style float locals, > fa fb .. and to-fa to-fb .. and the use of temporary buffers. > > I replaced them with F>R, FR@ and FR>. Moving floats to the return stack. > This works well as both stacks are 64 bits. The effect was very good. > ... > What do you have in your system? > > BR > Peter Using the FPU: FORTH> see z* Flags: ANSI $0135F340 : Z* $0135F34A fpop, $0135F354 sub rsi, #16 b# $0135F358 fstp [rsi] tbyte $0135F35A fpop, $0135F364 sub rsi, #16 b# $0135F368 fstp [rsi] tbyte $0135F36A fpop, $0135F374 sub rsi, #16 b# $0135F378 fstp [rsi] tbyte $0135F37A fpop, $0135F384 sub rsi, #16 b# $0135F388 fstp [rsi] tbyte $0135F38A fld [rsi] tbyte $0135F38C fld [rsi #32 +] tbyte $0135F38F fmulp ST(1), ST $0135F391 fld [rsi #16 +] tbyte $0135F394 fld [rsi #48 +] tbyte $0135F397 fmulp ST(1), ST $0135F399 fsubp ST(1), ST $0135F39B fld [rsi] tbyte $0135F39D fld [rsi #48 +] tbyte $0135F3A0 fmulp ST(1), ST $0135F3A2 fld [rsi #16 +] tbyte $0135F3A5 fld [rsi #32 +] tbyte $0135F3A8 fmulp ST(1), ST $0135F3AA faddp ST(1), ST $0135F3AC lea r13, [r13 #-32 +] qword $0135F3B0 fxch ST(2) $0135F3B2 fstp [r13 #16 +] tbyte $0135F3B6 fstp [r13 0 +] tbyte $0135F3BA add rsi, #64 b# $0135F3BE ; 116 bytes. -marcel
[toc] | [prev] | [next] | [standalone]
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-08-30 10:06 +0200 |
| Message-ID | <20260830100637.000013f4@tin.it> |
| In reply to | #135446 |
On Fri, 28 Aug 2026 00:04:28 +0200
peter <peter.noreply@tin.it> wrote:
>
> I recently wanted to adapt compelx-kahan.fs, the complex implementation
> by Julian V. Noble and David N. Williams, to the capabilities of
> ntf64/lxf64.
>
> One thing I wanted to get away from was the Baden-style float locals,
> fa fb .. and to-fa to-fb .. and the use of temporary buffers.
>
> I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
> This works well as both stacks are 64 bits. The effect was very good.
>
>
> : z* ( f: x y u v -- x*u-y*v x*v+y*u)
> zdup f+ f>r ( f: x y u v r: u+v)
> fover f- f>r ( f: x y u r: u+v v-u)
> frot fdup fr> ( f: y u x x v-u r: u+v)
> f* ( f: y u x x*'v-u' r: u+v)
> fr> fswap f>r f>r ( f: y u x r: x*'v-u' u+v)
> frot fdup fr> ( f: u x y y u+v r: x*'v-u')
> f* f>r ( f: u x y r: x*'v-u' y*'u+v')
> f+ f* fdup fr> f- fswap fr> ( f: x*u-y*v u*[x+y] x*[v-u])
> f+ ; ( f: x*u-y*v x*v+y*u)
>
>
> There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer
>
> The revised version gives
>
> seea Z*
> 0x428A50 C4C17B584D00 vaddsd xmm1, xmm0, qword ptr [r13]
> 0x428A56 C4C17B5C4500 vsubsd xmm0, xmm0, qword ptr [r13]
> 0x428A5C C4C17B594510 vmulsd xmm0, xmm0, qword ptr [r13+0x10]
> 0x428A62 C4C173594D08 vmulsd xmm1, xmm1, qword ptr [r13+0x8]
> 0x428A68 C4C17B105508 vmovsd xmm2, qword ptr [r13+0x8]
> 0x428A6E C4C16B585510 vaddsd xmm2, xmm2, qword ptr [r13+0x10]
> 0x428A74 C4C16B595500 vmulsd xmm2, xmm2, qword ptr [r13]
> 0x428A7A C5EB5CD9 vsubsd xmm3, xmm2, xmm1
> 0x428A7E C5FB58C2 vaddsd xmm0, xmm0, xmm2
> 0x428A82 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
> 0x428A88 4D8D6D10 lea r13, [r13+0x10]
> 0x428A8C C3 ret
> 61 bytes, 12 instructions
> ok
>
> Half the size and all memory loads/stores gone.
>
I went ahead and implemented float locals.
Now I can write using JVN formula
: Z* {f: x y u v :}
x y f+ u f* fdup u v f+ y f* f- fswap v u f- x f* f+ ;
this compiles to
seea Z*
0x42A6E0 C4C17B104D08 vmovsd xmm1, qword ptr [r13+0x8]
0x42A6E6 C4C173584D10 vaddsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A6EC C4C173594D00 vmulsd xmm1, xmm1, qword ptr [r13]
0x42A6F2 C4C17B585500 vaddsd xmm2, xmm0, qword ptr [r13]
0x42A6F8 C4C16B595508 vmulsd xmm2, xmm2, qword ptr [r13+0x8]
0x42A6FE C5F35CDA vsubsd xmm3, xmm1, xmm2
0x42A702 C4C17B5C5500 vsubsd xmm2, xmm0, qword ptr [r13]
0x42A708 C4C16B595510 vmulsd xmm2, xmm2, qword ptr [r13+0x10]
0x42A70E C5EB58D1 vaddsd xmm2, xmm2, xmm1
0x42A712 C5EB10C2 vmovsd xmm0, xmm2, xmm2
0x42A716 C4C17B115D10 vmovsd qword ptr [r13+0x10], xmm3
0x42A71C 4D8D6D10 lea r13, [r13+0x10]
0x42A720 C3 ret
65 bytes, 13 instructions
4 bytes and 1 instruction more
If I do the school book formual I get
: Z* {f: x y u v :}
x u f* y v f* f- x v f* y u f* f+ ;
Which give more compact code
seea Z*
0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
0x42A766 4D8D6D10 lea r13, [r13+0x10]
0x42A76A C3 ret
59 bytes, 12 instructions
4 multiplications 2 add/sub
I tried to measure the execution times. This shows no difference!
: z1 0 do zdup zdup z* zdrop loop ;
fpi 345e4 timer-reset 100000000 z1 .elapsed 179 ms elapsed ok
fpi 345e4 timer-reset 100000000 z2 .elapsed 176 ms elapsed ok
fpi 345e4 timer-reset 100000000 z3 .elapsed 177 ms elapsed ok
I had to start with the complex number already on the stack!
If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
BR
Peter
> An alternative I considered was to introduce float locals. It would probably
> have been more complicated. The result, in the best case, would be equal in
> the generated code. The problem with locals is the scope, that is to the end
> of the word. Of course the source would be more readable!
>
> I might still implement float locals.
>
> What do you have in your system?
>
> BR
> Peter
>
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2026-08-30 13:31 +0000 |
| Message-ID | <2026Aug30.153156@mips.complang.tuwien.ac.at> |
| In reply to | #135490 |
peter <peter.noreply@tin.it> writes:
>If I do the school book formual I get
>
>: Z* {f: x y u v :}
> x u f* y v f* f- x v f* y u f* f+ ;
>
>
>Which give more compact code
>
>seea Z*
>0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
>0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
>0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
>0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
>0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
>0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
>0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
>0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
>0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
>0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
>0x42A766 4D8D6D10 lea r13, [r13+0x10]
>0x42A76A C3 ret
>59 bytes, 12 instructions
Cool!
>I tried to measure the execution times. This shows no difference!
Which indicates that the bottleneck of your benchmark is not in the z*
code.
>I had to start with the complex number already on the stack!
>If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
How do you implement FP literals? One technique is to call FLIT, have
the FP literal behind the call, and let FLIT load from its return
address, then change its return address. That results in a branch
misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
technique do not explain 100* slowdown.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
[toc] | [prev] | [next] | [standalone]
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-08-30 15:58 +0200 |
| Message-ID | <20260830155856.00001689@tin.it> |
| In reply to | #135492 |
On Sun, 30 Aug 2026 13:31:56 GMT
anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
> peter <peter.noreply@tin.it> writes:
> >If I do the school book formual I get
> >
> >: Z* {f: x y u v :}
> > x u f* y v f* f- x v f* y u f* f+ ;
> >
> >
> >Which give more compact code
> >
> >seea Z*
> >0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
> >0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
> >0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
> >0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
> >0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
> >0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
> >0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
> >0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
> >0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
> >0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
> >0x42A766 4D8D6D10 lea r13, [r13+0x10]
> >0x42A76A C3 ret
> >59 bytes, 12 instructions
>
> Cool!
>
> >I tried to measure the execution times. This shows no difference!
>
> Which indicates that the bottleneck of your benchmark is not in the z*
> code.
>
> >I had to start with the complex number already on the stack!
> >If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
>
> How do you implement FP literals? One technique is to call FLIT, have
> the FP literal behind the call, and let FLIT load from its return
> address, then change its return address. That results in a branch
> misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
> technique do not explain 100* slowdown.
I found the main problem. It was FPI that was implemented as a call
to the C math library I use. It goes thru several calls saves all regs
and switches the stack to the original (correctly aligned).
Now it is defined as
$4009´21FB´5444´2D18 pad ! pad f@ FCONSTANT FPI
This compiles to the token code
see fpi macro
Address OP Instruction
0x720300 81 182D4454FB210940 FLIT 0x400921FB54442D18
0x720309 25 RET
10 bytes, 2 instructions
It is a macro so will get inlined
The assembler code becomes
seea fpi
0x423AD0 48B8182D4454FB210940 mov rax, 0x400921FB54442D18
0x423ADA C4E1F96EC8 vmovq xmm1, rax
0x423ADF C4C17B1145F8 vmovsd qword ptr [r13-0x8], xmm0
0x423AE5 C5F310C1 vmovsd xmm0, xmm1, xmm1
0x423AE9 4D8D6DF8 lea r13, [r13-0x8]
0x423AED C3 ret
30 bytes, 6 instructions
ok
2 instruction to load the litteral in a xmm register
It is a bit long but I do not think it can be better
The other possible solution would be to store litterals in memmory
and load from there. I think this is what Aarch64 compilers do
Peter
> - anton
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2026-08-30 15:12 +0000 |
| Message-ID | <2026Aug30.171200@mips.complang.tuwien.ac.at> |
| In reply to | #135493 |
peter <peter.noreply@tin.it> writes:
>The other possible solution would be to store litterals in memmory
>and load from there. I think this is what Aarch64 compilers do
That's also what gcc does on AMD64:
[/tmp:171587] cat xxx.c
double f(double x)
{
return x+1234.0;
}
[/tmp:171588] gcc -O -S xxx.c
[/tmp:171589] cat xxx.s
.file "xxx.c"
.text
.globl f
.type f, @function
f:
.LFB0:
.cfi_startproc
addsd .LC0(%rip), %xmm0
ret
.cfi_endproc
.LFE0:
.size f, .-f
.section .rodata.cst8,"aM",@progbits,8
.align 8
.LC0:
.long 0
.long 1083394048
.ident "GCC: (Debian 12.2.0-14) 12.2.0"
.section .note.GNU-stack,"",@progbits
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
[toc] | [prev] | [standalone]
Back to top | Article view | comp.lang.forth
csiph-web