Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #135564

Re: Forth on ARM64

From peter <peter.noreply@tin.it>
Newsgroups comp.lang.forth
Subject Re: Forth on ARM64
Date 2026-09-05 09:54 +0200
Organization A noiseless patient Spider
Message-ID <20260905095419.0000461b@tin.it> (permalink)
References <117cmk0$1st8p$1@paganini.bofh.team> <20260904073311.00003526@tin.it> <117fgge$27441$1@paganini.bofh.team>

Show all headers | View raw


On Fri, 4 Sep 2026 22:25:20 -0000 (UTC)
antispam@fricas.org (Waldek Hebisch) wrote:

> peter <peter.noreply@tin.it> wrote:
> > On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
> > antispam@fricas.org (Waldek Hebisch) wrote:
> > 
> >> Under Linux on ARM64 trying to set machine stack pointer to
> >> value which is not divisible by 16 leads to error.  AFAICS this
> >> means that in default setting machine stack pointer is not
> >> usable as as Forth user stack pointer or return stack pointer.
> >> I wonder what Forth implementation do?  Do they use different
> >> registers as user and return stack pointer?  Maybe they use
> >> machine stack pointer for control and locals?  Or maybe some
> >> system magic removes the restriction?
> >> 
> > 
> > Here is the register assignments for the token VM I wrote for 
> > ARM64 for lxf 64
> > 
> > /*
> > VM8 assembler based aarch64 vm for lxf64 Forth
> > Copyright 2020 Peter Fälth
> > 
> > Register usage
> >         X19     ip      vm instruction pointer
> >         x20     TOP     top of stack cached in x20
> >         x21     sp      vm stack pointer
> >         x22     rp      vm return stack pointer
> >         x23     fp      vm floating point stack pointer
> >         x24     lp      vm local stack pointer
> >         x25     idx     loop index of innermost loop
> >         x26     limit   loop limit of innermost loop
> >         x27     address of jump table
> >         d8      FTOP    top of float stack cached in d8
> > 
> > sequence to nest to next opcode is RELOAD
> > 
> >         ldrb    w0, [x19], 1                    load opcode byte at ip, advance ip by 1
> >         ldr     x2, [x27, x0, lsl 3]            load address of machine code from jmptable+opcode*8
> >         br      x2                              jump to next machine code
> > 
> > */
> > 
> > There are just 2 calls in the hole VM, in these cases 2 registers are 
> > pushed to maintain 16 byte alignment.
> > 
> > You can avoid the 16 byte alignment by using a register other then sp
> > for the processor stack. My tests showed this code to be about 30% slower
> > in execution speed.
> 
> What do you compare?  Token threaded code to traditional threaded
> code using memory addresses?
> 
> On ARM64 calls have range +-128MB relative.  So if the program code
> does not exceed 128MB, then subroutine threaded code is smaller
> than traditional threaded code and probably quite a bit faster.
> Drawback is that non-leaf words need to push and pop return
> address from the machine stack, using 16 bytes of stack space.
> 

I have rerun some tests on my old RPi4. Now I do not get any difference
in speed in using sp as stack pointer or another register. I might remember
wrong as it was several years ago i did test this last time.
I test a fib defined as code so threading type should not influence.

code fib2
START:	
	stp	x29, x30, [sp, -16]!
	cmp	x20, 2
	b.ge	L1
	mov	x20, 1
	b	END
L1:	
	sub	x10, x20, 1
	str	x20, [x21, -8]!
	mov	x20, x10
	bl	START
	ldr	x9, [x21]
	sub	x9, x9, 2
	str	x20, [x21]
	mov	x20, x9
	bl	START
	ldr	x8, [x21], 8
	add	x20, x20, x8
END:	
	ldp	x29, x30, [sp], 16
	ret
end-code	

note that this is a wrong fib as the fib in onebench.fs used by gforth have the
same problem. I wanted the same number of iterations.

My system works in 3 steps.
1. Compile forth source for token threaded VM
2. Compile the token code to native assembler code
3. Assemble with a built in assembler to native code

For ARM64 only the first step is implemented

Peter

Back to comp.lang.forth | Previous | Next — Previous in thread | Find similar | Unroll thread


Thread

Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-03 20:51 +0000
  Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 07:33 +0200
    Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-04 10:21 +0200
      Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 21:19 +0000
        Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 23:51 +0200
        Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-05 13:11 +0200
        Re: Forth on ARM64 Paul Rubin <no.email@nospam.invalid> - 2026-09-05 15:49 -0700
          Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-05 23:39 +0000
    Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 22:25 +0000
      Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-05 09:54 +0200

csiph-web