Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #135493
| From | peter <peter.noreply@tin.it> |
|---|---|
| Newsgroups | comp.lang.forth |
| Subject | Re: F>R, FR@ and FR> Float to return stack words |
| Date | 2026-08-30 15:58 +0200 |
| Organization | A noiseless patient Spider |
| Message-ID | <20260830155856.00001689@tin.it> (permalink) |
| References | <20260828000428.00007e53@tin.it> <20260830100637.000013f4@tin.it> <2026Aug30.153156@mips.complang.tuwien.ac.at> |
On Sun, 30 Aug 2026 13:31:56 GMT
anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote:
> peter <peter.noreply@tin.it> writes:
> >If I do the school book formual I get
> >
> >: Z* {f: x y u v :}
> > x u f* y v f* f- x v f* y u f* f+ ;
> >
> >
> >Which give more compact code
> >
> >seea Z*
> >0x42A730 C4C17B104D00 vmovsd xmm1, qword ptr [r13]
> >0x42A736 C4C173594D10 vmulsd xmm1, xmm1, qword ptr [r13+0x10]
> >0x42A73C C4C17B595508 vmulsd xmm2, xmm0, qword ptr [r13+0x8]
> >0x42A742 C5F35CCA vsubsd xmm1, xmm1, xmm2
> >0x42A746 C4C17B595510 vmulsd xmm2, xmm0, qword ptr [r13+0x10]
> >0x42A74C C4C17B105D00 vmovsd xmm3, qword ptr [r13]
> >0x42A752 C4C163595D08 vmulsd xmm3, xmm3, qword ptr [r13+0x8]
> >0x42A758 C5E358DA vaddsd xmm3, xmm3, xmm2
> >0x42A75C C5E310C3 vmovsd xmm0, xmm3, xmm3
> >0x42A760 C4C17B114D10 vmovsd qword ptr [r13+0x10], xmm1
> >0x42A766 4D8D6D10 lea r13, [r13+0x10]
> >0x42A76A C3 ret
> >59 bytes, 12 instructions
>
> Cool!
>
> >I tried to measure the execution times. This shows no difference!
>
> Which indicates that the bottleneck of your benchmark is not in the z*
> code.
>
> >I had to start with the complex number already on the stack!
> >If I put fpi 345e4 inside the loop elapsed time increase with 100 times!
>
> How do you implement FP literals? One technique is to call FLIT, have
> the FP literal behind the call, and let FLIT load from its return
> address, then change its return address. That results in a branch
> misprediction (~20 cycles) for every FLIT, but even 2 FLITs with this
> technique do not explain 100* slowdown.
I found the main problem. It was FPI that was implemented as a call
to the C math library I use. It goes thru several calls saves all regs
and switches the stack to the original (correctly aligned).
Now it is defined as
$4009´21FB´5444´2D18 pad ! pad f@ FCONSTANT FPI
This compiles to the token code
see fpi macro
Address OP Instruction
0x720300 81 182D4454FB210940 FLIT 0x400921FB54442D18
0x720309 25 RET
10 bytes, 2 instructions
It is a macro so will get inlined
The assembler code becomes
seea fpi
0x423AD0 48B8182D4454FB210940 mov rax, 0x400921FB54442D18
0x423ADA C4E1F96EC8 vmovq xmm1, rax
0x423ADF C4C17B1145F8 vmovsd qword ptr [r13-0x8], xmm0
0x423AE5 C5F310C1 vmovsd xmm0, xmm1, xmm1
0x423AE9 4D8D6DF8 lea r13, [r13-0x8]
0x423AED C3 ret
30 bytes, 6 instructions
ok
2 instruction to load the litteral in a xmm register
It is a bit long but I do not think it can be better
The other possible solution would be to store litterals in memmory
and load from there. I think this is what Aarch64 compilers do
Peter
> - anton
Back to comp.lang.forth | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-28 00:04 +0200
Re: F>R, FR@ and FR> Float to return stack words dxf <dxforth@gmail.com> - 2026-08-28 14:44 +1000
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-28 06:50 +0000
Re: F>R, FR@ and FR> Float to return stack words marcel hendrix <mhx@iae.nl> - 2026-08-28 19:35 +0200
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 10:06 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 13:31 +0000
Re: F>R, FR@ and FR> Float to return stack words peter <peter.noreply@tin.it> - 2026-08-30 15:58 +0200
Re: F>R, FR@ and FR> Float to return stack words anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2026-08-30 15:12 +0000
csiph-web