Path: csiph.com!eternal-september.org!feeder.eternal-september.org!nntp.eternal-september.org!.POSTED!not-for-mail From: peter Newsgroups: comp.lang.forth Subject: Re: Forth on ARM64 Date: Sat, 5 Sep 2026 09:54:19 +0200 Organization: A noiseless patient Spider Lines: 107 Message-ID: <20260905095419.0000461b@tin.it> References: <117cmk0$1st8p$1@paganini.bofh.team> <20260904073311.00003526@tin.it> <117fgge$27441$1@paganini.bofh.team> MIME-Version: 1.0 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Injection-Date: Sat, 05 Sep 2026 07:54:19 +0000 (UTC) Injection-Info: dont-email.me; logging-data="1369339"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX1+6ytU/NhJLtYv2APmpWHNZae14MAflMew="; posting-host="dd2646d3236745580f2aa81941843108" Cancel-Lock: sha1:mkfFDucGdLqq8SfBSw6g8U1+xRM= sha256:0axMMWpHsHM4Zs/jebYDc+yNQ8KlKwH/8Ur1m/GpsEc= sha1:4VHfJTnSiwhqQ4HoADR1Z09EjOI= sha256:Rh0aud+AuF7mZ3cC58rSZaakHbgJSHUDxXzMF0R+z5k= X-Newsreader: Claws Mail 4.4.0 (GTK 3.24.51; x86_64-w64-mingw32) Xref: csiph.com comp.lang.forth:135564 On Fri, 4 Sep 2026 22:25:20 -0000 (UTC) antispam@fricas.org (Waldek Hebisch) wrote: > peter wrote: > > On Thu, 3 Sep 2026 20:51:14 -0000 (UTC) > > antispam@fricas.org (Waldek Hebisch) wrote: > >=20 > >> Under Linux on ARM64 trying to set machine stack pointer to > >> value which is not divisible by 16 leads to error. AFAICS this > >> means that in default setting machine stack pointer is not > >> usable as as Forth user stack pointer or return stack pointer. > >> I wonder what Forth implementation do? Do they use different > >> registers as user and return stack pointer? Maybe they use > >> machine stack pointer for control and locals? Or maybe some > >> system magic removes the restriction? > >>=20 > >=20 > > Here is the register assignments for the token VM I wrote for=20 > > ARM64 for lxf 64 > >=20 > > /* > > VM8 assembler based aarch64 vm for lxf64 Forth > > Copyright 2020 Peter F=E4lth > >=20 > > Register usage > > X19 ip vm instruction pointer > > x20 TOP top of stack cached in x20 > > x21 sp vm stack pointer > > x22 rp vm return stack pointer > > x23 fp vm floating point stack pointer > > x24 lp vm local stack pointer > > x25 idx loop index of innermost loop > > x26 limit loop limit of innermost loop > > x27 address of jump table > > d8 FTOP top of float stack cached in d8 > >=20 > > sequence to nest to next opcode is RELOAD > >=20 > > ldrb w0, [x19], 1 load opcode byte at ip,= advance ip by 1 > > ldr x2, [x27, x0, lsl 3] load address of machine= code from jmptable+opcode*8 > > br x2 jump to next machine co= de > >=20 > > */ > >=20 > > There are just 2 calls in the hole VM, in these cases 2 registers are=20 > > pushed to maintain 16 byte alignment. > >=20 > > You can avoid the 16 byte alignment by using a register other then sp > > for the processor stack. My tests showed this code to be about 30% slow= er > > in execution speed. >=20 > What do you compare? Token threaded code to traditional threaded > code using memory addresses? >=20 > On ARM64 calls have range +-128MB relative. So if the program code > does not exceed 128MB, then subroutine threaded code is smaller > than traditional threaded code and probably quite a bit faster. > Drawback is that non-leaf words need to push and pop return > address from the machine stack, using 16 bytes of stack space. >=20 I have rerun some tests on my old RPi4. Now I do not get any difference in speed in using sp as stack pointer or another register. I might remember wrong as it was several years ago i did test this last time. I test a fib defined as code so threading type should not influence. code fib2 START:=09 stp x29, x30, [sp, -16]! cmp x20, 2 b.ge L1 mov x20, 1 b END L1:=09 sub x10, x20, 1 str x20, [x21, -8]! mov x20, x10 bl START ldr x9, [x21] sub x9, x9, 2 str x20, [x21] mov x20, x9 bl START ldr x8, [x21], 8 add x20, x20, x8 END:=09 ldp x29, x30, [sp], 16 ret end-code=09 note that this is a wrong fib as the fib in onebench.fs used by gforth have= the same problem. I wanted the same number of iterations. My system works in 3 steps. 1. Compile forth source for token threaded VM 2. Compile the token code to native assembler code 3. Assemble with a built in assembler to native code For ARM64 only the first step is implemented Peter