Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #135446 > unrolled thread

F>R, FR@ and FR> Float to return stack words

Started bypeter <peter.noreply@tin.it>
First post2026-08-28 00:04 +0200
Last post2026-08-30 15:12 +0000
Articles 8 — 4 participants

Back to article view | Back to comp.lang.forth


Contents

  F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-28 00:04 +0200
    Re: F>R, FR@ and FR> Float to return stack words dxf <dxforth@gmail.com> - 2026-08-28 14:44 +1000
    Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-28 06:50 +0000
    Re: F>R, FR@ and FR> Float to return stack words marcel hendrix <mhx@iae.nl> - 2026-08-28 19:35 +0200
    Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 10:06 +0200
      Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 13:31 +0000
        Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 15:58 +0200
          Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 15:12 +0000

#135446 — F>R, FR@ and FR> Float to return stack words

Frompeter <peter.noreply@tin.it>
Date2026-08-28 00:04 +0200
SubjectF>R, FR@ and FR> Float to return stack words
Message-ID<20260828000428.00007e53@tin.it>
I recently wanted to adapt compelx-kahan.fs, the complex implementation
by Julian V. Noble and David N. Williams, to the capabilities of 
ntf64/lxf64.

One thing I wanted to get away from was the Baden-style float locals,
fa fb .. and to-fa to-fb .. and the use of temporary buffers.

I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
This works well as both stacks are 64 bits. The effect was very good.

Taking Z* as an example:

Original Z* using 2 temporary buffers

: z*    (f: x y u v -- x*u-y*v  x*v+y*u)
(
Uses the algorithm followed by jvn:
  [x+iy]*[u+iv] = [[x+y]*u - y*[u+v]] + i[[x+y]*u + x*[v-u]]
Requires 3 multiplications and 5 additions.
)
  zdup f+                               (f: x y u v u+v)
  [ fnoname ] literal f!                (f: x y u v)
  fover f-                              (f: x y u v-u)
  [ fnoname float+ ] literal f!         (f: x y u)
  frot fdup                             (f: y u x x)
  [ fnoname float+ ] literal f@         (f: y u x x v-u)
  f*
  [ fnoname float+ ] literal f!         (f: y u x)
  frot fdup                             (f: u x y y)
  [ fnoname ] literal f@                (f: u x y y u+v)
  f*
  [ fnoname ] literal f!                (f: u x y)
  f+  f* fdup                           (f: u*[x+y] u*[x+y])
  [ fnoname ] literal f@ f-             (f: u*[x+y] x*u-y*v)
  fswap
  [ fnoname float+ ] literal f@         (f: x*u-y*v u*[x+y] x*[v-u])
  f+ ;                                  (f: x*u-y*v x*v+y*u)

With more or less direct exchange to using the return stack
I got

: z*    ( f: x y u v -- x*u-y*v  x*v+y*u)
  zdup f+ f>r                           ( f: x y u v  r: u+v)
  fover f- f>r                          ( f: x y u  r: u+v v-u)
  frot fdup fr>                         ( f: y u x x v-u r: u+v)
  f*                                    ( f: y u x x*'v-u' r: u+v)
  fr> fswap f>r f>r                     ( f: y u x r: x*'v-u' u+v)
  frot fdup fr>                         ( f: u x y y u+v r: x*'v-u')
  f* f>r                                ( f: u x y r: x*'v-u' y*'u+v')
  f+ f* fdup fr> f- fswap fr>           ( f: x*u-y*v u*[x+y] x*[v-u])
  f+ ;                                  ( f: x*u-y*v x*v+y*u)

The original give the following assembler code

seea Z*
0x4291A0  C4C17B584D00         vaddsd xmm1, xmm0, qword ptr [r13]
0x4291A6  C5FB110DD2965D00     vmovsd qword ptr [rel 0xA02880], xmm1
0x4291AE  C4C17B5C4500         vsubsd xmm0, xmm0, qword ptr [r13]
0x4291B4  C5FB1105CC965D00     vmovsd qword ptr [rel 0xA02888], xmm0
0x4291BC  C5FB1005C4965D00     vmovsd xmm0, qword ptr [rel 0xA02888]
0x4291C4  C4C17B594510         vmulsd xmm0, xmm0, qword ptr [r13+0x10]
0x4291CA  C5FB1105B6965D00     vmovsd qword ptr [rel 0xA02888], xmm0
0x4291D2  C5FB1005A6965D00     vmovsd xmm0, qword ptr [rel 0xA02880]
0x4291DA  C4C17B594508         vmulsd xmm0, xmm0, qword ptr [r13+0x8]
0x4291E0  C5FB110598965D00     vmovsd qword ptr [rel 0xA02880], xmm0
0x4291E8  C4C17B104508         vmovsd xmm0, qword ptr [r13+0x8]
0x4291EE  C4C17B584510         vaddsd xmm0, xmm0, qword ptr [r13+0x10]
0x4291F4  C4C17B594500         vmulsd xmm0, xmm0, qword ptr [r13]
0x4291FA  C5FB100D7E965D00     vmovsd xmm1, qword ptr [rel 0xA02880]
0x429202  C5FB5CD1             vsubsd xmm2, xmm0, xmm1
0x429206  C5FB100D7A965D00     vmovsd xmm1, qword ptr [rel 0xA02888]
0x42920E  C5F358C8             vaddsd xmm1, xmm1, xmm0
0x429212  C5F310C1             vmovsd xmm0, xmm1, xmm1
0x429216  C4C17B115510         vmovsd qword ptr [r13+0x10], xmm2
0x42921C  4D8D6D10             lea r13, [r13+0x10]
0x429220  C3                   ret
129 bytes, 21 instructions
ok

There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer

The revised version gives

seea Z*
0x428A50  C4C17B584D00         vaddsd xmm1, xmm0, qword ptr [r13]
0x428A56  C4C17B5C4500         vsubsd xmm0, xmm0, qword ptr [r13]
0x428A5C  C4C17B594510         vmulsd xmm0, xmm0, qword ptr [r13+0x10]
0x428A62  C4C173594D08         vmulsd xmm1, xmm1, qword ptr [r13+0x8]
0x428A68  C4C17B105508         vmovsd xmm2, qword ptr [r13+0x8]
0x428A6E  C4C16B585510         vaddsd xmm2, xmm2, qword ptr [r13+0x10]
0x428A74  C4C16B595500         vmulsd xmm2, xmm2, qword ptr [r13]
0x428A7A  C5EB5CD9             vsubsd xmm3, xmm2, xmm1
0x428A7E  C5FB58C2             vaddsd xmm0, xmm0, xmm2
0x428A82  C4C17B115D10         vmovsd qword ptr [r13+0x10], xmm3
0x428A88  4D8D6D10             lea r13, [r13+0x10]
0x428A8C  C3                   ret
61 bytes, 12 instructions
ok

Half the size and all memory loads/stores gone.

With the returns stack use I can also do things like
    
: +ulp f>r r> 1+ >r fr> ;  ok

0e +ulp e.  5e-324  ok
0e +ulp fh.  0x0.0000000000001p-1022  ok

seea +ulp
0x42A260  C4E1F97EC0          vmovq rax, xmm0
0x42A265  48FFC0              inc rax
0x42A268  C4E1F96EC0          vmovq xmm0, rax
0x42A26D  C3                  ret
14 bytes, 4 instructions
ok

An alternative I considered was to introduce float locals. It would probably 
have been more complicated. The result, in the best case, would be equal in
the generated code. The problem with locals is the scope, that is to the end 
of the word. Of course the source would be more readable!

I might still implement float locals.

What do you have in your system?

BR
Peter

[toc] | [next] | [standalone]


#135448

Fromdxf <dxforth@gmail.com>
Date2026-08-28 14:44 +1000
Message-ID<6a911235$1@news.ausics.net>
In reply to#135446
On 28/08/2026 8:04 am, peter wrote:
> ...
> What do you have in your system?

Because return-stacking is more complicated, I use fvariable where I can.
Where I can't, I have the following def's to fall back on.  I've not
played enough with Baden-type flocals to get a sense of what they're like
in practice.  The varying HERE (PAD, HOLD etc) might be an issue.

\ x86
code F>R ( r -- )
  1 floats # bp sub  bp push  ' f! ) jmp
end-code

code FR> ( -- r )
  bp push  c: f@ ;c  1 floats # bp add  next
end-code


\ 8080
code F>R ( r -- )
  rpp lhld  -1 floats d lxi  d dad  rpp shld
  h push  ' f! jmp
end-code

code FR> ( -- r )
  rpp lhld  h push  c: f@ ;c  rpp lhld
  1 floats d lxi  d dad  rpp shld  next
end-code

[toc] | [prev] | [next] | [standalone]


#135449

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2026-08-28 06:50 +0000
Message-ID<2026Aug28.085007@mips.complang.tuwien.ac.at>
In reply to#135446
peter <peter.noreply@tin.it> writes:
>: z*    ( f: x y u v -- x*u-y*v  x*v+y*u)
>  zdup f+ f>r                           ( f: x y u v  r: u+v)
>  fover f- f>r                          ( f: x y u  r: u+v v-u)
>  frot fdup fr>                         ( f: y u x x v-u r: u+v)
>  f*                                    ( f: y u x x*'v-u' r: u+v)
>  fr> fswap f>r f>r                     ( f: y u x r: x*'v-u' u+v)
>  frot fdup fr>                         ( f: u x y y u+v r: x*'v-u')
>  f* f>r                                ( f: u x y r: x*'v-u' y*'u+v')
>  f+ f* fdup fr> f- fswap fr>           ( f: x*u-y*v u*[x+y] x*[v-u])
>  f+ ;                                  ( f: x*u-y*v x*v+y*u)
[...]
>seea Z*
>0x428A50  C4C17B584D00         vaddsd xmm1, xmm0, qword ptr [r13]
>0x428A56  C4C17B5C4500         vsubsd xmm0, xmm0, qword ptr [r13]
>0x428A5C  C4C17B594510         vmulsd xmm0, xmm0, qword ptr [r13+0x10]
>0x428A62  C4C173594D08         vmulsd xmm1, xmm1, qword ptr [r13+0x8]
>0x428A68  C4C17B105508         vmovsd xmm2, qword ptr [r13+0x8]
>0x428A6E  C4C16B585510         vaddsd xmm2, xmm2, qword ptr [r13+0x10]
>0x428A74  C4C16B595500         vmulsd xmm2, xmm2, qword ptr [r13]
>0x428A7A  C5EB5CD9             vsubsd xmm3, xmm2, xmm1
>0x428A7E  C5FB58C2             vaddsd xmm0, xmm0, xmm2
>0x428A82  C4C17B115D10         vmovsd qword ptr [r13+0x10], xmm3
>0x428A88  4D8D6D10             lea r13, [r13+0x10]
>0x428A8C  C3                   ret
>61 bytes, 12 instructions

Nice.

>An alternative I considered was to introduce float locals. It would probably 
>have been more complicated. The result, in the best case, would be equal in
>the generated code. The problem with locals is the scope, that is to the end 
>of the word. Of course the source would be more readable!
>
>I might still implement float locals.
>
>What do you have in your system?

Gforth has had FP locals since 1994.  We recently added F>R, FR> and
FR@.  The reason why we did not have them was that the return-stack
pointer is not necessarily aligned for FP loads and stores, so on many
hardware platforms of the 1990s F>R and FR> would have been expensive
operations.

With 64-bit cells becoming mainstream, 64-bit floats (what we use in
Gforth), and the death of most architectures that do not support
unaligned accesses, the balance has changed in the meantime.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html

[toc] | [prev] | [next] | [standalone]


#135462

Frommarcel hendrix <mhx@iae.nl>
Date2026-08-28 19:35 +0200
Message-ID<116sgt6$ll12$1@dont-email.me>
In reply to#135446
On 8/28/2026 12:04 AM, peter wrote:
> 
> I recently wanted to adapt compelx-kahan.fs, the complex implementation
> by Julian V. Noble and David N. Williams, to the capabilities of
> ntf64/lxf64.
> 
> One thing I wanted to get away from was the Baden-style float locals,
> fa fb .. and to-fa to-fb .. and the use of temporary buffers.
> 
> I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
> This works well as both stacks are 64 bits. The effect was very good.
>
... > What do you have in your system?
> 
> BR
> Peter

Using the FPU:
FORTH> see z*
Flags: ANSI
$0135F340  : Z*
$0135F34A  fpop,
$0135F354  sub           rsi, #16 b#
$0135F358  fstp          [rsi] tbyte
$0135F35A  fpop,
$0135F364  sub           rsi, #16 b#
$0135F368  fstp          [rsi] tbyte
$0135F36A  fpop,
$0135F374  sub           rsi, #16 b#
$0135F378  fstp          [rsi] tbyte
$0135F37A  fpop,
$0135F384  sub           rsi, #16 b#
$0135F388  fstp          [rsi] tbyte
$0135F38A  fld           [rsi] tbyte
$0135F38C  fld           [rsi #32 +] tbyte
$0135F38F  fmulp         ST(1), ST
$0135F391  fld           [rsi #16 +] tbyte
$0135F394  fld           [rsi #48 +] tbyte
$0135F397  fmulp         ST(1), ST
$0135F399  fsubp         ST(1), ST
$0135F39B  fld           [rsi] tbyte
$0135F39D  fld           [rsi #48 +] tbyte
$0135F3A0  fmulp         ST(1), ST
$0135F3A2  fld           [rsi #16 +] tbyte
$0135F3A5  fld           [rsi #32 +] tbyte
$0135F3A8  fmulp         ST(1), ST
$0135F3AA  faddp         ST(1), ST
$0135F3AC  lea           r13, [r13 #-32 +] qword
$0135F3B0  fxch          ST(2)
$0135F3B2  fstp          [r13 #16 +] tbyte
$0135F3B6  fstp          [r13 0 +] tbyte
$0135F3BA  add           rsi, #64 b#
$0135F3BE  ;
116 bytes.

-marcel

[toc] | [prev] | [next] | [standalone]


#135490

Frompeter <peter.noreply@tin.it>
Date2026-08-30 10:06 +0200
Message-ID<20260830100637.000013f4@tin.it>
In reply to#135446
On Fri, 28 Aug 2026 00:04:28 +0200
peter <peter.noreply@tin.it> wrote:

> 
> I recently wanted to adapt compelx-kahan.fs, the complex implementation
> by Julian V. Noble and David N. Williams, to the capabilities of 
> ntf64/lxf64.
> 
> One thing I wanted to get away from was the Baden-style float locals,
> fa fb .. and to-fa to-fb .. and the use of temporary buffers.
> 
> I replaced them with F>R, FR@ and FR>. Moving floats to the return stack.
> This works well as both stacks are 64 bits. The effect was very good.
> 

> 
> : z*    ( f: x y u v -- x*u-y*v  x*v+y*u)
>   zdup f+ f>r                           ( f: x y u v  r: u+v)
>   fover f- f>r                          ( f: x y u  r: u+v v-u)
>   frot fdup fr>                         ( f: y u x x v-u r: u+v)
>   f*                                    ( f: y u x x*'v-u' r: u+v)
>   fr> fswap f>r f>r                     ( f: y u x r: x*'v-u' u+v)
>   frot fdup fr>                         ( f: u x y y u+v r: x*'v-u')
>   f* f>r                                ( f: u x y r: x*'v-u' y*'u+v')
>   f+ f* fdup fr> f- fswap fr>           ( f: x*u-y*v u*[x+y] x*[v-u])
>   f+ ;                                  ( f: x*u-y*v x*v+y*u)
> 

> 
> There are 3 multiplications, 5 add/sub and 8 loads/stores to the buffer
> 
> The revised version gives
> 
> seea Z*
> 0x428A50  C4C17B584D00         vaddsd xmm1, xmm0, qword ptr [r13]
> 0x428A56  C4C17B5C4500         vsubsd xmm0, xmm0, qword ptr [r13]
> 0x428A5C  C4C17B594510         vmulsd xmm0, xmm0, qword ptr [r13+0x10]
> 0x428A62  C4C173594D08         vmulsd xmm1, xmm1, qword ptr [r13+0x8]
> 0x428A68  C4C17B105508         vmovsd xmm2, qword ptr [r13+0x8]
> 0x428A6E  C4C16B585510         vaddsd xmm2, xmm2, qword ptr [r13+0x10]
> 0x428A74  C4C16B595500         vmulsd xmm2, xmm2, qword ptr [r13]
> 0x428A7A  C5EB5CD9             vsubsd xmm3, xmm2, xmm1
> 0x428A7E  C5FB58C2             vaddsd xmm0, xmm0, xmm2
> 0x428A82  C4C17B115D10         vmovsd qword ptr [r13+0x10], xmm3
> 0x428A88  4D8D6D10             lea r13, [r13+0x10]
> 0x428A8C  C3                   ret
> 61 bytes, 12 instructions
> ok
> 
> Half the size and all memory loads/stores gone.
> 

I went ahead and implemented float locals. 
Now I can write using JVN formula

: Z*  {f: x y u v :}  
	x y f+ u f* fdup u v f+ y f* f- fswap v u f- x f* f+ ;

this compiles to

seea Z*
0x42A6E0  C4C17B104D08         vmovsd xmm1, qword ptr [r13+0x8]
0x42A6E6  C4C173584D10         vaddsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A6EC  C4C173594D00         vmulsd xmm1, xmm1, qword ptr [r13]
0x42A6F2  C4C17B585500         vaddsd xmm2, xmm0, qword ptr [r13]
0x42A6F8  C4C16B595508         vmulsd xmm2, xmm2, qword ptr [r13+0x8]
0x42A6FE  C5F35CDA             vsubsd xmm3, xmm1, xmm2
0x42A702  C4C17B5C5500         vsubsd xmm2, xmm0, qword ptr [r13]
0x42A708  C4C16B595510         vmulsd xmm2, xmm2, qword ptr [r13+0x10]
0x42A70E  C5EB58D1             vaddsd xmm2, xmm2, xmm1
0x42A712  C5EB10C2             vmovsd xmm0, xmm2, xmm2
0x42A716  C4C17B115D10         vmovsd qword ptr [r13+0x10], xmm3
0x42A71C  4D8D6D10             lea r13, [r13+0x10]
0x42A720  C3                   ret
65 bytes, 13 instructions
4 bytes and 1 instruction more

If I do the school book formual I get

: Z*  {f: x y u v :} 
	x u f* y v f* f- x v f* y u f* f+ ;


Which give more compact code

seea Z*
0x42A730  C4C17B104D00         vmovsd xmm1, qword ptr [r13]
0x42A736  C4C173594D10         vmulsd xmm1, xmm1, qword ptr [r13+0x10]
0x42A73C  C4C17B595508         vmulsd xmm2, xmm0, qword ptr [r13+0x8]
0x42A742  C5F35CCA             vsubsd xmm1, xmm1, xmm2
0x42A746  C4C17B595510         vmulsd xmm2, xmm0, qword ptr [r13+0x10]
0x42A74C  C4C17B105D00         vmovsd xmm3, qword ptr [r13]
0x42A752  C4C163595D08         vmulsd xmm3, xmm3, qword ptr [r13+0x8]
0x42A758  C5E358DA             vaddsd xmm3, xmm3, xmm2
0x42A75C  C5E310C3             vmovsd xmm0, xmm3, xmm3
0x42A760  C4C17B114D10         vmovsd qword ptr [r13+0x10], xmm1
0x42A766  4D8D6D10             lea r13, [r13+0x10]
0x42A76A  C3                   ret
59 bytes, 12 instructions

4 multiplications 2 add/sub

I tried to measure the execution times. This shows no difference!

: z1 0 do zdup zdup z* zdrop  loop ; 

fpi 345e4 timer-reset 100000000 z1 .elapsed  179 ms elapsed  ok
fpi 345e4 timer-reset 100000000 z2 .elapsed  176 ms elapsed  ok
fpi 345e4 timer-reset 100000000 z3 .elapsed  177 ms elapsed  ok

I had to start with the complex number already on the stack!
If I put fpi 345e4 inside the loop elapsed time increase with 100 times!

BR
Peter

> An alternative I considered was to introduce float locals. It would probably 
> have been more complicated. The result, in the best case, would be equal in
> the generated code. The problem with locals is the scope, that is to the end 
> of the word. Of course the source would be more readable!
> 
> I might still implement float locals.
> 
> What do you have in your system?
> 
> BR
> Peter
> 

[toc] | [prev] | [next] | [standalone]


#135492

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2026-08-30 13:31 +0000
Message-ID<2026Aug30.153156@mips.complang.tuwien.ac.at>
In reply to#135490
peter <peter.noreply@tin.it> writes:
>If I do the school book formual I get
>
>: Z*  {f: x y u v :} 
>	x u f* y v f* f- x v f* y u f* f+ ;
>
>
>Which give more compact code
>
>seea Z*
>0x42A730  C4C17B104D00         vmovsd xmm1, qword ptr [r13]
>0x42A736  C4C173594D10         vmulsd xmm1, xmm1, qword ptr [r13+0x10]
>0x42A73C  C4C17B595508         vmulsd xmm2, xmm0, qword ptr [r13+0x8]
>0x42A742  C5F35CCA             vsubsd xmm1, xmm1, xmm2
>0x42A746  C4C17B595510         vmulsd xmm2, xmm0, qword ptr [r13+0x10]
>0x42A74C  C4C17B105D00         vmovsd xmm3, qword ptr [r13]
>0x42A752  C4C163595D08         vmulsd xmm3, xmm3, qword ptr [r13+0x8]
>0x42A758  C5E358DA             vaddsd xmm3, xmm3, xmm2
>0x42A75C  C5E310C3             vmovsd xmm0, xmm3, xmm3
>0x42A760  C4C17B114D10         vmovsd qword ptr [r13+0x10], xmm1
>0x42A766  4D8D6D10             lea r13, [r13+0x10]
>0x42A76A  C3                   ret
>59 bytes, 12 instructions

Cool!

>I tried to measure the execution times. This shows no difference!

Which indicates that the bottleneck of your benchmark is not in the z*
code.

>I had to start with the complex number already on the stack!
>If I put fpi 345e4 inside the loop elapsed time increase with 100 times!

How do you implement FP literals?  One technique is to call FLIT, have
the FP literal behind the call, and let FLIT load from its return
address, then change its return address.  That results in a branch
misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
technique do not explain 100* slowdown.

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html

[toc] | [prev] | [next] | [standalone]


#135493

Frompeter <peter.noreply@tin.it>
Date2026-08-30 15:58 +0200
Message-ID<20260830155856.00001689@tin.it>
In reply to#135492
On Sun, 30 Aug 2026 13:31:56 GMT
anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:

> peter <peter.noreply@tin.it> writes:
> >If I do the school book formual I get
> >
> >: Z*  {f: x y u v :} 
> >	x u f* y v f* f- x v f* y u f* f+ ;
> >
> >
> >Which give more compact code
> >
> >seea Z*
> >0x42A730  C4C17B104D00         vmovsd xmm1, qword ptr [r13]
> >0x42A736  C4C173594D10         vmulsd xmm1, xmm1, qword ptr [r13+0x10]
> >0x42A73C  C4C17B595508         vmulsd xmm2, xmm0, qword ptr [r13+0x8]
> >0x42A742  C5F35CCA             vsubsd xmm1, xmm1, xmm2
> >0x42A746  C4C17B595510         vmulsd xmm2, xmm0, qword ptr [r13+0x10]
> >0x42A74C  C4C17B105D00         vmovsd xmm3, qword ptr [r13]
> >0x42A752  C4C163595D08         vmulsd xmm3, xmm3, qword ptr [r13+0x8]
> >0x42A758  C5E358DA             vaddsd xmm3, xmm3, xmm2
> >0x42A75C  C5E310C3             vmovsd xmm0, xmm3, xmm3
> >0x42A760  C4C17B114D10         vmovsd qword ptr [r13+0x10], xmm1
> >0x42A766  4D8D6D10             lea r13, [r13+0x10]
> >0x42A76A  C3                   ret
> >59 bytes, 12 instructions
> 
> Cool!
> 
> >I tried to measure the execution times. This shows no difference!
> 
> Which indicates that the bottleneck of your benchmark is not in the z*
> code.
> 
> >I had to start with the complex number already on the stack!
> >If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
> 
> How do you implement FP literals?  One technique is to call FLIT, have
> the FP literal behind the call, and let FLIT load from its return
> address, then change its return address.  That results in a branch
> misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
> technique do not explain 100* slowdown.

I found the main problem. It was FPI that was implemented as a call
to the C math library I use. It goes thru several calls saves all regs
and switches the stack to the original (correctly aligned).

Now it is defined as 

$4009´21FB´5444´2D18 pad ! pad f@ FCONSTANT FPI

This compiles to the token code

see fpi macro
Address   OP                         Instruction
0x720300  81 182D4454FB210940        FLIT  0x400921FB54442D18
0x720309  25                         RET
10 bytes, 2 instructions

It is a macro so will get inlined
The assembler code becomes
seea fpi
0x423AD0  48B8182D4454FB210940 mov rax, 0x400921FB54442D18
0x423ADA  C4E1F96EC8           vmovq xmm1, rax
0x423ADF  C4C17B1145F8         vmovsd qword ptr [r13-0x8], xmm0
0x423AE5  C5F310C1             vmovsd xmm0, xmm1, xmm1
0x423AE9  4D8D6DF8             lea r13, [r13-0x8]
0x423AED  C3                   ret
30 bytes, 6 instructions
ok

2 instruction to load the litteral in a xmm register
It is a bit long but I do not think it can be better

The other possible solution would be to store litterals in memmory
and load from there. I think this is  what Aarch64 compilers do

Peter
 
> - anton

[toc] | [prev] | [next] | [standalone]


#135494

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2026-08-30 15:12 +0000
Message-ID<2026Aug30.171200@mips.complang.tuwien.ac.at>
In reply to#135493
peter <peter.noreply@tin.it> writes:
>The other possible solution would be to store litterals in memmory
>and load from there. I think this is  what Aarch64 compilers do

That's also what gcc does on AMD64:

[/tmp:171587] cat xxx.c
double f(double x)
{
  return x+1234.0;
}
[/tmp:171588] gcc -O -S xxx.c
[/tmp:171589] cat xxx.s
        .file   "xxx.c"
        .text
        .globl  f
        .type   f, @function
f:
.LFB0:
        .cfi_startproc
        addsd   .LC0(%rip), %xmm0
        ret
        .cfi_endproc
.LFE0:
        .size   f, .-f
        .section        .rodata.cst8,"aM",@progbits,8
        .align 8
.LC0:
        .long   0
        .long   1083394048
        .ident  "GCC: (Debian 12.2.0-14) 12.2.0"
        .section        .note.GNU-stack,"",@progbits

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: https://forth-standard.org/
EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html

[toc] | [prev] | [standalone]


Back to top | Article view | comp.lang.forth


csiph-web