Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #135543 > unrolled thread

Forth on ARM64

Started byantispam@fricas.org (Waldek Hebisch)
First post2026-09-03 20:51 +0000
Last post2026-09-05 09:54 +0200
Articles 10 — 4 participants

Back to article view | Back to comp.lang.forth


Contents

  Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-03 20:51 +0000
    Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 07:33 +0200
      Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-04 10:21 +0200
        Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 21:19 +0000
          Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 23:51 +0200
          Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-05 13:11 +0200
          Re: Forth on ARM64 Paul Rubin <no.email@nospam.invalid> - 2026-09-05 15:49 -0700
            Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-05 23:39 +0000
      Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 22:25 +0000
        Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-05 09:54 +0200

#135543 — Forth on ARM64

Fromantispam@fricas.org (Waldek Hebisch)
Date2026-09-03 20:51 +0000
SubjectForth on ARM64
Message-ID<117cmk0$1st8p$1@paganini.bofh.team>
Under Linux on ARM64 trying to set machine stack pointer to
value which is not divisible by 16 leads to error.  AFAICS this
means that in default setting machine stack pointer is not
usable as as Forth user stack pointer or return stack pointer.
I wonder what Forth implementation do?  Do they use different
registers as user and return stack pointer?  Maybe they use
machine stack pointer for control and locals?  Or maybe some
system magic removes the restriction?

-- 
                              Waldek Hebisch

[toc] | [next] | [standalone]


#135550

Frompeter <peter.noreply@tin.it>
Date2026-09-04 07:33 +0200
Message-ID<20260904073311.00003526@tin.it>
In reply to#135543
On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
antispam@fricas.org (Waldek Hebisch) wrote:

> Under Linux on ARM64 trying to set machine stack pointer to
> value which is not divisible by 16 leads to error.  AFAICS this
> means that in default setting machine stack pointer is not
> usable as as Forth user stack pointer or return stack pointer.
> I wonder what Forth implementation do?  Do they use different
> registers as user and return stack pointer?  Maybe they use
> machine stack pointer for control and locals?  Or maybe some
> system magic removes the restriction?
> 

Here is the register assignments for the token VM I wrote for 
ARM64 for lxf 64

/*
VM8 assembler based aarch64 vm for lxf64 Forth
Copyright 2020 Peter Fälth

Register usage
	X19	ip	vm instruction pointer
    	x20	TOP	top of stack cached in x20
	x21	sp	vm stack pointer
    	x22	rp	vm return stack pointer
	x23	fp	vm floating point stack pointer
	x24	lp	vm local stack pointer
	x25	idx	loop index of innermost loop
    	x26 	limit 	loop limit of innermost loop
	x27	address of jump table
	d8	FTOP	top of float stack cached in d8

sequence to nest to next opcode is RELOAD

	ldrb	w0, [x19], 1			load opcode byte at ip, advance ip by 1
	ldr	x2, [x27, x0, lsl 3]		load address of machine code from jmptable+opcode*8
	br	x2				jump to next machine code

*/

There are just 2 calls in the hole VM, in these cases 2 registers are 
pushed to maintain 16 byte alignment.

You can avoid the 16 byte alignment by using a register other then sp
for the processor stack. My tests showed this code to be about 30% slower
in execution speed.

BR
Peter

[toc] | [prev] | [next] | [standalone]


#135551

Fromalbert@spenarnc.xs4all.nl
Date2026-09-04 10:21 +0200
Message-ID<nnd$3b980ba8$57c49a3b@950882b8bd21bd72>
In reply to#135550
In article <20260904073311.00003526@tin.it>,
peter  <peter.noreply@tin.it> wrote:
>On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
>antispam@fricas.org (Waldek Hebisch) wrote:
>
>> Under Linux on ARM64 trying to set machine stack pointer to
>> value which is not divisible by 16 leads to error.  AFAICS this
>> means that in default setting machine stack pointer is not
>> usable as as Forth user stack pointer or return stack pointer.
>> I wonder what Forth implementation do?  Do they use different
>> registers as user and return stack pointer?  Maybe they use
>> machine stack pointer for control and locals?  Or maybe some
>> system magic removes the restriction?
>>
>
>Here is the register assignments for the token VM I wrote for
>ARM64 for lxf 64
>
>/*
>VM8 assembler based aarch64 vm for lxf64 Forth
>Copyright 2020 Peter Fälth
>
>Register usage
>       X19     ip      vm instruction pointer
>       x20     TOP     top of stack cached in x20
>       x21     sp      vm stack pointer
>       x22     rp      vm return stack pointer
>       x23     fp      vm floating point stack pointer
>       x24     lp      vm local stack pointer
>       x25     idx     loop index of innermost loop
>       x26     limit   loop limit of innermost loop
>       x27     address of jump table
>       d8      FTOP    top of float stack cached in d8
>
>sequence to nest to next opcode is RELOAD
>
>       ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
>       ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
>       br      x2                              jump to next machine code
>
>*/
>
>There are just 2 calls in the hole VM, in these cases 2 registers are
>pushed to maintain 16 byte alignment.
>
>You can avoid the 16 byte alignment by using a register other then sp
>for the processor stack. My tests showed this code to be about 30% slower
>in execution speed.

The problem you have is unrelated to linux, but more a c-compatibility.

With my assembler Forth I have no problem in linux.
All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc.
The stack pointer SP plays no role in the whole Forth, it is arbitrarily
mapped to R13. I could map SPO to x21 equally well.

[In thumb there seems to be a an SP but I do not do thumb.]

There is a similar problem in DLL calls for Windows system.
I have a brilliant solution, but the margin is to small to
contain it.

You could consult :
https://github.com/albertvanderhorst/ciforth/wiki

>
>BR
>Peter

Groetjes Albert
>
>
-- 
The Chinese government is satisfied with its military superiority over USA.
The next 5 year plan has as primary goal to advance life expectancy
over 80 years, like Western Europe.

[toc] | [prev] | [next] | [standalone]


#135559

Fromantispam@fricas.org (Waldek Hebisch)
Date2026-09-04 21:19 +0000
Message-ID<117fcl7$26t23$1@paganini.bofh.team>
In reply to#135551
albert@spenarnc.xs4all.nl wrote:
> In article <20260904073311.00003526@tin.it>,
> peter  <peter.noreply@tin.it> wrote:
>>On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
>>antispam@fricas.org (Waldek Hebisch) wrote:
>>
>>> Under Linux on ARM64 trying to set machine stack pointer to
>>> value which is not divisible by 16 leads to error.  AFAICS this
>>> means that in default setting machine stack pointer is not
>>> usable as as Forth user stack pointer or return stack pointer.
>>> I wonder what Forth implementation do?  Do they use different
>>> registers as user and return stack pointer?  Maybe they use
>>> machine stack pointer for control and locals?  Or maybe some
>>> system magic removes the restriction?
>>>
>>
>>Here is the register assignments for the token VM I wrote for
>>ARM64 for lxf 64
>>
>>/*
>>VM8 assembler based aarch64 vm for lxf64 Forth
>>Copyright 2020 Peter Fälth
>>
>>Register usage
>>       X19     ip      vm instruction pointer
>>       x20     TOP     top of stack cached in x20
>>       x21     sp      vm stack pointer
>>       x22     rp      vm return stack pointer
>>       x23     fp      vm floating point stack pointer
>>       x24     lp      vm local stack pointer
>>       x25     idx     loop index of innermost loop
>>       x26     limit   loop limit of innermost loop
>>       x27     address of jump table
>>       d8      FTOP    top of float stack cached in d8
>>
>>sequence to nest to next opcode is RELOAD
>>
>>       ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
>>       ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
>>       br      x2                              jump to next machine code
>>
>>*/
>>
>>There are just 2 calls in the hole VM, in these cases 2 registers are
>>pushed to maintain 16 byte alignment.
>>
>>You can avoid the 16 byte alignment by using a register other then sp
>>for the processor stack. My tests showed this code to be about 30% slower
>>in execution speed.
> 
> The problem you have is unrelated to linux, but more a c-compatibility.
> 
> With my assembler Forth I have no problem in linux.
> All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc.
> The stack pointer SP plays no role in the whole Forth, it is arbitrarily
> mapped to R13. I could map SPO to x21 equally well.

I wrote "machine stack pointer" (or if you prefer register number 31)
because it is special on ARM64.  Of course, Forth can use a different
register, but not using machine stack pointer is a waste.  C
compatibility makes this waste more painful, because there are
only 11 registers not touched by C.  For traditional Forth
implementations this is enough, but I am looking at generating
machine code.

-- 
                              Waldek Hebisch

[toc] | [prev] | [next] | [standalone]


#135560

Frompeter <peter.noreply@tin.it>
Date2026-09-04 23:51 +0200
Message-ID<20260904235129.000057b0@tin.it>
In reply to#135559
On Fri, 4 Sep 2026 21:19:37 -0000 (UTC)
antispam@fricas.org (Waldek Hebisch) wrote:

> albert@spenarnc.xs4all.nl wrote:
> > In article <20260904073311.00003526@tin.it>,
> > peter  <peter.noreply@tin.it> wrote:
> >>On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
> >>antispam@fricas.org (Waldek Hebisch) wrote:
> >>
> >>> Under Linux on ARM64 trying to set machine stack pointer to
> >>> value which is not divisible by 16 leads to error.  AFAICS this
> >>> means that in default setting machine stack pointer is not
> >>> usable as as Forth user stack pointer or return stack pointer.
> >>> I wonder what Forth implementation do?  Do they use different
> >>> registers as user and return stack pointer?  Maybe they use
> >>> machine stack pointer for control and locals?  Or maybe some
> >>> system magic removes the restriction?
> >>>
> >>
> >>Here is the register assignments for the token VM I wrote for
> >>ARM64 for lxf 64
> >>
> >>/*
> >>VM8 assembler based aarch64 vm for lxf64 Forth
> >>Copyright 2020 Peter Fälth
> >>
> >>Register usage
> >>       X19     ip      vm instruction pointer
> >>       x20     TOP     top of stack cached in x20
> >>       x21     sp      vm stack pointer
> >>       x22     rp      vm return stack pointer
> >>       x23     fp      vm floating point stack pointer
> >>       x24     lp      vm local stack pointer
> >>       x25     idx     loop index of innermost loop
> >>       x26     limit   loop limit of innermost loop
> >>       x27     address of jump table
> >>       d8      FTOP    top of float stack cached in d8
> >>
> >>sequence to nest to next opcode is RELOAD
> >>
> >>       ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
> >>       ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
> >>       br      x2                              jump to next machine code
> >>
> >>*/
> >>
> >>There are just 2 calls in the hole VM, in these cases 2 registers are
> >>pushed to maintain 16 byte alignment.
> >>
> >>You can avoid the 16 byte alignment by using a register other then sp
> >>for the processor stack. My tests showed this code to be about 30% slower
> >>in execution speed.
> > 
> > The problem you have is unrelated to linux, but more a c-compatibility.
> > 
> > With my assembler Forth I have no problem in linux.
> > All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc.
> > The stack pointer SP plays no role in the whole Forth, it is arbitrarily
> > mapped to R13. I could map SPO to x21 equally well.
> 
> I wrote "machine stack pointer" (or if you prefer register number 31)
> because it is special on ARM64.  Of course, Forth can use a different
> register, but not using machine stack pointer is a waste.  C
> compatibility makes this waste more painful, because there are
> only 11 registers not touched by C.  For traditional Forth
> implementations this is enough, but I am looking at generating
> machine code.
> 
I will also in the future make a code generator for ARM64. 
I have one for X64 now!

my idea is to use

stp	x29, x30, [sp, -16]!
and
ldp	x29, x30, [sp], 16

at the start and end of words that has call in them

>r r> r@ can be done with
stp	xzr, x20, [sp, -16]!
and
ldp	xzr, x20, [sp], 16

This keeps everything aligned at 16 bytes

I will use all registers for my code generator. To handle C library 
compatibility I will have a gate that all C calls pass that saves 
needed registers.
This is how I handle it on X64 now. It works great.
X64 also has a requirement of 16 byte alignment of the return stack.
This is not enforced by the CPU, but it can bite you for specific
opcodes that require 16 byte alignment. It is worse as Windows and Linux
have different opinions on when the stack should be aligned, before or
after the call?
I handle that now by not aligning in code I generate and instead switch to
a properly aligned stack at the call gate.

BR
Peter

[toc] | [prev] | [next] | [standalone]


#135567

Fromalbert@spenarnc.xs4all.nl
Date2026-09-05 13:11 +0200
Message-ID<nnd$221f4dce$35e3b01f@9c6c99f017e8c6d5>
In reply to#135559
In article <117fcl7$26t23$1@paganini.bofh.team>,
Waldek Hebisch <antispam@fricas.org> wrote:
>albert@spenarnc.xs4all.nl wrote:
>I wrote "machine stack pointer" (or if you prefer register number 31)
>because it is special on ARM64.  Of course, Forth can use a different
>register, but not using machine stack pointer is a waste.  C
>compatibility makes this waste more painful, because there are
>only 11 registers not touched by C.  For traditional Forth
>implementations this is enough, but I am looking at generating
>machine code.

I do not buy that I can use only 11 registers not touched by C.
I can easily save the 3 registers needed for Forth, if I want
to call a C-routine. So I freely use all of them.

>                              Waldek Hebisch
-- 
The Chinese government is satisfied with its military superiority over USA.
The next 5 year plan has as primary goal to advance life expectancy
over 80 years, like Western Europe.

[toc] | [prev] | [next] | [standalone]


#135572

FromPaul Rubin <no.email@nospam.invalid>
Date2026-09-05 15:49 -0700
Message-ID<875x0jl3f6.fsf@nightsong.com>
In reply to#135559
antispam@fricas.org (Waldek Hebisch) writes:
> there are only 11 registers not touched by C.

I missed some parts of this thread but why would a C compiler leave 11
registers untouched?  Usually a compiler with register allocation will
use all the registers, unless you ask it to reserve some for other
purposes.

[toc] | [prev] | [next] | [standalone]


#135575

Fromantispam@fricas.org (Waldek Hebisch)
Date2026-09-05 23:39 +0000
Message-ID<117i97f$2h3f8$1@paganini.bofh.team>
In reply to#135572
Paul Rubin <no.email@nospam.invalid> wrote:
> antispam@fricas.org (Waldek Hebisch) writes:
>> there are only 11 registers not touched by C.
> 
> I missed some parts of this thread but why would a C compiler leave 11
> registers untouched?  Usually a compiler with register allocation will
> use all the registers, unless you ask it to reserve some for other
> purposes.

What I wrote was a shortcut.  Full version is: when you call a routine
in C, C compiler will ensure that 11 registers are preserved.  That
is calling convention.  And yes, unless register is reserved for
some global use C compiler may use it.  Which means that any
register that is not designated as preserved may be clobbered by
C code.  There is possible exception for global registers like pointer
to thread data, but changing value of such register is likely to
cause trouble for C code, so it is unwise to use them for different
purpose.  And on ARM64 there seem to be no such register among general
registers.

-- 
                              Waldek Hebisch

[toc] | [prev] | [next] | [standalone]


#135561

Fromantispam@fricas.org (Waldek Hebisch)
Date2026-09-04 22:25 +0000
Message-ID<117fgge$27441$1@paganini.bofh.team>
In reply to#135550
peter <peter.noreply@tin.it> wrote:
> On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
> antispam@fricas.org (Waldek Hebisch) wrote:
> 
>> Under Linux on ARM64 trying to set machine stack pointer to
>> value which is not divisible by 16 leads to error.  AFAICS this
>> means that in default setting machine stack pointer is not
>> usable as as Forth user stack pointer or return stack pointer.
>> I wonder what Forth implementation do?  Do they use different
>> registers as user and return stack pointer?  Maybe they use
>> machine stack pointer for control and locals?  Or maybe some
>> system magic removes the restriction?
>> 
> 
> Here is the register assignments for the token VM I wrote for 
> ARM64 for lxf 64
> 
> /*
> VM8 assembler based aarch64 vm for lxf64 Forth
> Copyright 2020 Peter Fälth
> 
> Register usage
>         X19     ip      vm instruction pointer
>         x20     TOP     top of stack cached in x20
>         x21     sp      vm stack pointer
>         x22     rp      vm return stack pointer
>         x23     fp      vm floating point stack pointer
>         x24     lp      vm local stack pointer
>         x25     idx     loop index of innermost loop
>         x26     limit   loop limit of innermost loop
>         x27     address of jump table
>         d8      FTOP    top of float stack cached in d8
> 
> sequence to nest to next opcode is RELOAD
> 
>         ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
>         ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
>         br      x2                              jump to next machine code
> 
> */
> 
> There are just 2 calls in the hole VM, in these cases 2 registers are 
> pushed to maintain 16 byte alignment.
> 
> You can avoid the 16 byte alignment by using a register other then sp
> for the processor stack. My tests showed this code to be about 30% slower
> in execution speed.

What do you compare?  Token threaded code to traditional threaded
code using memory addresses?

On ARM64 calls have range +-128MB relative.  So if the program code
does not exceed 128MB, then subroutine threaded code is smaller
than traditional threaded code and probably quite a bit faster.
Drawback is that non-leaf words need to push and pop return
address from the machine stack, using 16 bytes of stack space.

-- 
                              Waldek Hebisch

[toc] | [prev] | [next] | [standalone]


#135564

Frompeter <peter.noreply@tin.it>
Date2026-09-05 09:54 +0200
Message-ID<20260905095419.0000461b@tin.it>
In reply to#135561
On Fri, 4 Sep 2026 22:25:20 -0000 (UTC)
antispam@fricas.org (Waldek Hebisch) wrote:

> peter <peter.noreply@tin.it> wrote:
> > On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
> > antispam@fricas.org (Waldek Hebisch) wrote:
> > 
> >> Under Linux on ARM64 trying to set machine stack pointer to
> >> value which is not divisible by 16 leads to error.  AFAICS this
> >> means that in default setting machine stack pointer is not
> >> usable as as Forth user stack pointer or return stack pointer.
> >> I wonder what Forth implementation do?  Do they use different
> >> registers as user and return stack pointer?  Maybe they use
> >> machine stack pointer for control and locals?  Or maybe some
> >> system magic removes the restriction?
> >> 
> > 
> > Here is the register assignments for the token VM I wrote for 
> > ARM64 for lxf 64
> > 
> > /*
> > VM8 assembler based aarch64 vm for lxf64 Forth
> > Copyright 2020 Peter Fälth
> > 
> > Register usage
> >         X19     ip      vm instruction pointer
> >         x20     TOP     top of stack cached in x20
> >         x21     sp      vm stack pointer
> >         x22     rp      vm return stack pointer
> >         x23     fp      vm floating point stack pointer
> >         x24     lp      vm local stack pointer
> >         x25     idx     loop index of innermost loop
> >         x26     limit   loop limit of innermost loop
> >         x27     address of jump table
> >         d8      FTOP    top of float stack cached in d8
> > 
> > sequence to nest to next opcode is RELOAD
> > 
> >         ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
> >         ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
> >         br      x2                              jump to next machine code
> > 
> > */
> > 
> > There are just 2 calls in the hole VM, in these cases 2 registers are 
> > pushed to maintain 16 byte alignment.
> > 
> > You can avoid the 16 byte alignment by using a register other then sp
> > for the processor stack. My tests showed this code to be about 30% slower
> > in execution speed.
> 
> What do you compare?  Token threaded code to traditional threaded
> code using memory addresses?
> 
> On ARM64 calls have range +-128MB relative.  So if the program code
> does not exceed 128MB, then subroutine threaded code is smaller
> than traditional threaded code and probably quite a bit faster.
> Drawback is that non-leaf words need to push and pop return
> address from the machine stack, using 16 bytes of stack space.
> 

I have rerun some tests on my old RPi4. Now I do not get any difference
in speed in using sp as stack pointer or another register. I might remember
wrong as it was several years ago i did test this last time.
I test a fib defined as code so threading type should not influence.

code fib2
START:	
	stp	x29, x30, [sp, -16]!
	cmp	x20, 2
	b.ge	L1
	mov	x20, 1
	b	END
L1:	
	sub	x10, x20, 1
	str	x20, [x21, -8]!
	mov	x20, x10
	bl	START
	ldr	x9, [x21]
	sub	x9, x9, 2
	str	x20, [x21]
	mov	x20, x9
	bl	START
	ldr	x8, [x21], 8
	add	x20, x20, x8
END:	
	ldp	x29, x30, [sp], 16
	ret
end-code	

note that this is a wrong fib as the fib in onebench.fs used by gforth have the
same problem. I wanted the same number of iterations.

My system works in 3 steps.
1. Compile forth source for token threaded VM
2. Compile the token code to native assembler code
3. Assemble with a built in assembler to native code

For ARM64 only the first step is implemented

Peter

[toc] | [prev] | [standalone]


Back to top | Article view | comp.lang.forth


csiph-web