Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #135543 > unrolled thread
| Started by | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| First post | 2026-09-03 20:51 +0000 |
| Last post | 2026-09-05 09:54 +0200 |
| Articles | 10 — 4 participants |
Back to article view | Back to comp.lang.forth
Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-03 20:51 +0000
Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 07:33 +0200
Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-04 10:21 +0200
Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 21:19 +0000
Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-04 23:51 +0200
Re: Forth on ARM64 albert@spenarnc.xs4all.nl - 2026-09-05 13:11 +0200
Re: Forth on ARM64 Paul Rubin <no.email@nospam.invalid> - 2026-09-05 15:49 -0700
Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-05 23:39 +0000
Re: Forth on ARM64 antispam@fricas.org (Waldek Hebisch) - 2026-09-04 22:25 +0000
Re: Forth on ARM64 peter <peter.noreply@tin.it> - 2026-09-05 09:54 +0200
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-09-03 20:51 +0000 |
| Subject | Forth on ARM64 |
| Message-ID | <117cmk0$1st8p$1@paganini.bofh.team> |
Under Linux on ARM64 trying to set machine stack pointer to
value which is not divisible by 16 leads to error. AFAICS this
means that in default setting machine stack pointer is not
usable as as Forth user stack pointer or return stack pointer.
I wonder what Forth implementation do? Do they use different
registers as user and return stack pointer? Maybe they use
machine stack pointer for control and locals? Or maybe some
system magic removes the restriction?
--
Waldek Hebisch
[toc] | [next] | [standalone]
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-09-04 07:33 +0200 |
| Message-ID | <20260904073311.00003526@tin.it> |
| In reply to | #135543 |
On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
antispam@fricas.org (Waldek Hebisch) wrote:
> Under Linux on ARM64 trying to set machine stack pointer to
> value which is not divisible by 16 leads to error. AFAICS this
> means that in default setting machine stack pointer is not
> usable as as Forth user stack pointer or return stack pointer.
> I wonder what Forth implementation do? Do they use different
> registers as user and return stack pointer? Maybe they use
> machine stack pointer for control and locals? Or maybe some
> system magic removes the restriction?
>
Here is the register assignments for the token VM I wrote for
ARM64 for lxf 64
/*
VM8 assembler based aarch64 vm for lxf64 Forth
Copyright 2020 Peter Fälth
Register usage
X19 ip vm instruction pointer
x20 TOP top of stack cached in x20
x21 sp vm stack pointer
x22 rp vm return stack pointer
x23 fp vm floating point stack pointer
x24 lp vm local stack pointer
x25 idx loop index of innermost loop
x26 limit loop limit of innermost loop
x27 address of jump table
d8 FTOP top of float stack cached in d8
sequence to nest to next opcode is RELOAD
ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
br x2 jump to next machine code
*/
There are just 2 calls in the hole VM, in these cases 2 registers are
pushed to maintain 16 byte alignment.
You can avoid the 16 byte alignment by using a register other then sp
for the processor stack. My tests showed this code to be about 30% slower
in execution speed.
BR
Peter
[toc] | [prev] | [next] | [standalone]
| From | albert@spenarnc.xs4all.nl |
|---|---|
| Date | 2026-09-04 10:21 +0200 |
| Message-ID | <nnd$3b980ba8$57c49a3b@950882b8bd21bd72> |
| In reply to | #135550 |
In article <20260904073311.00003526@tin.it>, peter <peter.noreply@tin.it> wrote: >On Thu, 3 Sep 2026 20:51:14 -0000 (UTC) >antispam@fricas.org (Waldek Hebisch) wrote: > >> Under Linux on ARM64 trying to set machine stack pointer to >> value which is not divisible by 16 leads to error. AFAICS this >> means that in default setting machine stack pointer is not >> usable as as Forth user stack pointer or return stack pointer. >> I wonder what Forth implementation do? Do they use different >> registers as user and return stack pointer? Maybe they use >> machine stack pointer for control and locals? Or maybe some >> system magic removes the restriction? >> > >Here is the register assignments for the token VM I wrote for >ARM64 for lxf 64 > >/* >VM8 assembler based aarch64 vm for lxf64 Forth >Copyright 2020 Peter Fälth > >Register usage > X19 ip vm instruction pointer > x20 TOP top of stack cached in x20 > x21 sp vm stack pointer > x22 rp vm return stack pointer > x23 fp vm floating point stack pointer > x24 lp vm local stack pointer > x25 idx loop index of innermost loop > x26 limit loop limit of innermost loop > x27 address of jump table > d8 FTOP top of float stack cached in d8 > >sequence to nest to next opcode is RELOAD > > ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1 > ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8 > br x2 jump to next machine code > >*/ > >There are just 2 calls in the hole VM, in these cases 2 registers are >pushed to maintain 16 byte alignment. > >You can avoid the 16 byte alignment by using a register other then sp >for the processor stack. My tests showed this code to be about 30% slower >in execution speed. The problem you have is unrelated to linux, but more a c-compatibility. With my assembler Forth I have no problem in linux. All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc. The stack pointer SP plays no role in the whole Forth, it is arbitrarily mapped to R13. I could map SPO to x21 equally well. [In thumb there seems to be a an SP but I do not do thumb.] There is a similar problem in DLL calls for Windows system. I have a brilliant solution, but the margin is to small to contain it. You could consult : https://github.com/albertvanderhorst/ciforth/wiki > >BR >Peter Groetjes Albert > > -- The Chinese government is satisfied with its military superiority over USA. The next 5 year plan has as primary goal to advance life expectancy over 80 years, like Western Europe.
[toc] | [prev] | [next] | [standalone]
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-09-04 21:19 +0000 |
| Message-ID | <117fcl7$26t23$1@paganini.bofh.team> |
| In reply to | #135551 |
albert@spenarnc.xs4all.nl wrote:
> In article <20260904073311.00003526@tin.it>,
> peter <peter.noreply@tin.it> wrote:
>>On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
>>antispam@fricas.org (Waldek Hebisch) wrote:
>>
>>> Under Linux on ARM64 trying to set machine stack pointer to
>>> value which is not divisible by 16 leads to error. AFAICS this
>>> means that in default setting machine stack pointer is not
>>> usable as as Forth user stack pointer or return stack pointer.
>>> I wonder what Forth implementation do? Do they use different
>>> registers as user and return stack pointer? Maybe they use
>>> machine stack pointer for control and locals? Or maybe some
>>> system magic removes the restriction?
>>>
>>
>>Here is the register assignments for the token VM I wrote for
>>ARM64 for lxf 64
>>
>>/*
>>VM8 assembler based aarch64 vm for lxf64 Forth
>>Copyright 2020 Peter Fälth
>>
>>Register usage
>> X19 ip vm instruction pointer
>> x20 TOP top of stack cached in x20
>> x21 sp vm stack pointer
>> x22 rp vm return stack pointer
>> x23 fp vm floating point stack pointer
>> x24 lp vm local stack pointer
>> x25 idx loop index of innermost loop
>> x26 limit loop limit of innermost loop
>> x27 address of jump table
>> d8 FTOP top of float stack cached in d8
>>
>>sequence to nest to next opcode is RELOAD
>>
>> ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
>> ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
>> br x2 jump to next machine code
>>
>>*/
>>
>>There are just 2 calls in the hole VM, in these cases 2 registers are
>>pushed to maintain 16 byte alignment.
>>
>>You can avoid the 16 byte alignment by using a register other then sp
>>for the processor stack. My tests showed this code to be about 30% slower
>>in execution speed.
>
> The problem you have is unrelated to linux, but more a c-compatibility.
>
> With my assembler Forth I have no problem in linux.
> All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc.
> The stack pointer SP plays no role in the whole Forth, it is arbitrarily
> mapped to R13. I could map SPO to x21 equally well.
I wrote "machine stack pointer" (or if you prefer register number 31)
because it is special on ARM64. Of course, Forth can use a different
register, but not using machine stack pointer is a waste. C
compatibility makes this waste more painful, because there are
only 11 registers not touched by C. For traditional Forth
implementations this is enough, but I am looking at generating
machine code.
--
Waldek Hebisch
[toc] | [prev] | [next] | [standalone]
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-09-04 23:51 +0200 |
| Message-ID | <20260904235129.000057b0@tin.it> |
| In reply to | #135559 |
On Fri, 4 Sep 2026 21:19:37 -0000 (UTC) antispam@fricas.org (Waldek Hebisch) wrote: > albert@spenarnc.xs4all.nl wrote: > > In article <20260904073311.00003526@tin.it>, > > peter <peter.noreply@tin.it> wrote: > >>On Thu, 3 Sep 2026 20:51:14 -0000 (UTC) > >>antispam@fricas.org (Waldek Hebisch) wrote: > >> > >>> Under Linux on ARM64 trying to set machine stack pointer to > >>> value which is not divisible by 16 leads to error. AFAICS this > >>> means that in default setting machine stack pointer is not > >>> usable as as Forth user stack pointer or return stack pointer. > >>> I wonder what Forth implementation do? Do they use different > >>> registers as user and return stack pointer? Maybe they use > >>> machine stack pointer for control and locals? Or maybe some > >>> system magic removes the restriction? > >>> > >> > >>Here is the register assignments for the token VM I wrote for > >>ARM64 for lxf 64 > >> > >>/* > >>VM8 assembler based aarch64 vm for lxf64 Forth > >>Copyright 2020 Peter Fälth > >> > >>Register usage > >> X19 ip vm instruction pointer > >> x20 TOP top of stack cached in x20 > >> x21 sp vm stack pointer > >> x22 rp vm return stack pointer > >> x23 fp vm floating point stack pointer > >> x24 lp vm local stack pointer > >> x25 idx loop index of innermost loop > >> x26 limit loop limit of innermost loop > >> x27 address of jump table > >> d8 FTOP top of float stack cached in d8 > >> > >>sequence to nest to next opcode is RELOAD > >> > >> ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1 > >> ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8 > >> br x2 jump to next machine code > >> > >>*/ > >> > >>There are just 2 calls in the hole VM, in these cases 2 registers are > >>pushed to maintain 16 byte alignment. > >> > >>You can avoid the 16 byte alignment by using a register other then sp > >>for the processor stack. My tests showed this code to be about 30% slower > >>in execution speed. > > > > The problem you have is unrelated to linux, but more a c-compatibility. > > > > With my assembler Forth I have no problem in linux. > > All system calls are done via SVC and filling registers X0, X1 X2 X3 X4 etc. > > The stack pointer SP plays no role in the whole Forth, it is arbitrarily > > mapped to R13. I could map SPO to x21 equally well. > > I wrote "machine stack pointer" (or if you prefer register number 31) > because it is special on ARM64. Of course, Forth can use a different > register, but not using machine stack pointer is a waste. C > compatibility makes this waste more painful, because there are > only 11 registers not touched by C. For traditional Forth > implementations this is enough, but I am looking at generating > machine code. > I will also in the future make a code generator for ARM64. I have one for X64 now! my idea is to use stp x29, x30, [sp, -16]! and ldp x29, x30, [sp], 16 at the start and end of words that has call in them >r r> r@ can be done with stp xzr, x20, [sp, -16]! and ldp xzr, x20, [sp], 16 This keeps everything aligned at 16 bytes I will use all registers for my code generator. To handle C library compatibility I will have a gate that all C calls pass that saves needed registers. This is how I handle it on X64 now. It works great. X64 also has a requirement of 16 byte alignment of the return stack. This is not enforced by the CPU, but it can bite you for specific opcodes that require 16 byte alignment. It is worse as Windows and Linux have different opinions on when the stack should be aligned, before or after the call? I handle that now by not aligning in code I generate and instead switch to a properly aligned stack at the call gate. BR Peter
[toc] | [prev] | [next] | [standalone]
| From | albert@spenarnc.xs4all.nl |
|---|---|
| Date | 2026-09-05 13:11 +0200 |
| Message-ID | <nnd$221f4dce$35e3b01f@9c6c99f017e8c6d5> |
| In reply to | #135559 |
In article <117fcl7$26t23$1@paganini.bofh.team>, Waldek Hebisch <antispam@fricas.org> wrote: >albert@spenarnc.xs4all.nl wrote: >I wrote "machine stack pointer" (or if you prefer register number 31) >because it is special on ARM64. Of course, Forth can use a different >register, but not using machine stack pointer is a waste. C >compatibility makes this waste more painful, because there are >only 11 registers not touched by C. For traditional Forth >implementations this is enough, but I am looking at generating >machine code. I do not buy that I can use only 11 registers not touched by C. I can easily save the 3 registers needed for Forth, if I want to call a C-routine. So I freely use all of them. > Waldek Hebisch -- The Chinese government is satisfied with its military superiority over USA. The next 5 year plan has as primary goal to advance life expectancy over 80 years, like Western Europe.
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2026-09-05 15:49 -0700 |
| Message-ID | <875x0jl3f6.fsf@nightsong.com> |
| In reply to | #135559 |
antispam@fricas.org (Waldek Hebisch) writes: > there are only 11 registers not touched by C. I missed some parts of this thread but why would a C compiler leave 11 registers untouched? Usually a compiler with register allocation will use all the registers, unless you ask it to reserve some for other purposes.
[toc] | [prev] | [next] | [standalone]
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-09-05 23:39 +0000 |
| Message-ID | <117i97f$2h3f8$1@paganini.bofh.team> |
| In reply to | #135572 |
Paul Rubin <no.email@nospam.invalid> wrote:
> antispam@fricas.org (Waldek Hebisch) writes:
>> there are only 11 registers not touched by C.
>
> I missed some parts of this thread but why would a C compiler leave 11
> registers untouched? Usually a compiler with register allocation will
> use all the registers, unless you ask it to reserve some for other
> purposes.
What I wrote was a shortcut. Full version is: when you call a routine
in C, C compiler will ensure that 11 registers are preserved. That
is calling convention. And yes, unless register is reserved for
some global use C compiler may use it. Which means that any
register that is not designated as preserved may be clobbered by
C code. There is possible exception for global registers like pointer
to thread data, but changing value of such register is likely to
cause trouble for C code, so it is unwise to use them for different
purpose. And on ARM64 there seem to be no such register among general
registers.
--
Waldek Hebisch
[toc] | [prev] | [next] | [standalone]
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-09-04 22:25 +0000 |
| Message-ID | <117fgge$27441$1@paganini.bofh.team> |
| In reply to | #135550 |
peter <peter.noreply@tin.it> wrote:
> On Thu, 3 Sep 2026 20:51:14 -0000 (UTC)
> antispam@fricas.org (Waldek Hebisch) wrote:
>
>> Under Linux on ARM64 trying to set machine stack pointer to
>> value which is not divisible by 16 leads to error. AFAICS this
>> means that in default setting machine stack pointer is not
>> usable as as Forth user stack pointer or return stack pointer.
>> I wonder what Forth implementation do? Do they use different
>> registers as user and return stack pointer? Maybe they use
>> machine stack pointer for control and locals? Or maybe some
>> system magic removes the restriction?
>>
>
> Here is the register assignments for the token VM I wrote for
> ARM64 for lxf 64
>
> /*
> VM8 assembler based aarch64 vm for lxf64 Forth
> Copyright 2020 Peter Fälth
>
> Register usage
> X19 ip vm instruction pointer
> x20 TOP top of stack cached in x20
> x21 sp vm stack pointer
> x22 rp vm return stack pointer
> x23 fp vm floating point stack pointer
> x24 lp vm local stack pointer
> x25 idx loop index of innermost loop
> x26 limit loop limit of innermost loop
> x27 address of jump table
> d8 FTOP top of float stack cached in d8
>
> sequence to nest to next opcode is RELOAD
>
> ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1
> ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8
> br x2 jump to next machine code
>
> */
>
> There are just 2 calls in the hole VM, in these cases 2 registers are
> pushed to maintain 16 byte alignment.
>
> You can avoid the 16 byte alignment by using a register other then sp
> for the processor stack. My tests showed this code to be about 30% slower
> in execution speed.
What do you compare? Token threaded code to traditional threaded
code using memory addresses?
On ARM64 calls have range +-128MB relative. So if the program code
does not exceed 128MB, then subroutine threaded code is smaller
than traditional threaded code and probably quite a bit faster.
Drawback is that non-leaf words need to push and pop return
address from the machine stack, using 16 bytes of stack space.
--
Waldek Hebisch
[toc] | [prev] | [next] | [standalone]
| From | peter <peter.noreply@tin.it> |
|---|---|
| Date | 2026-09-05 09:54 +0200 |
| Message-ID | <20260905095419.0000461b@tin.it> |
| In reply to | #135561 |
On Fri, 4 Sep 2026 22:25:20 -0000 (UTC) antispam@fricas.org (Waldek Hebisch) wrote: > peter <peter.noreply@tin.it> wrote: > > On Thu, 3 Sep 2026 20:51:14 -0000 (UTC) > > antispam@fricas.org (Waldek Hebisch) wrote: > > > >> Under Linux on ARM64 trying to set machine stack pointer to > >> value which is not divisible by 16 leads to error. AFAICS this > >> means that in default setting machine stack pointer is not > >> usable as as Forth user stack pointer or return stack pointer. > >> I wonder what Forth implementation do? Do they use different > >> registers as user and return stack pointer? Maybe they use > >> machine stack pointer for control and locals? Or maybe some > >> system magic removes the restriction? > >> > > > > Here is the register assignments for the token VM I wrote for > > ARM64 for lxf 64 > > > > /* > > VM8 assembler based aarch64 vm for lxf64 Forth > > Copyright 2020 Peter Fälth > > > > Register usage > > X19 ip vm instruction pointer > > x20 TOP top of stack cached in x20 > > x21 sp vm stack pointer > > x22 rp vm return stack pointer > > x23 fp vm floating point stack pointer > > x24 lp vm local stack pointer > > x25 idx loop index of innermost loop > > x26 limit loop limit of innermost loop > > x27 address of jump table > > d8 FTOP top of float stack cached in d8 > > > > sequence to nest to next opcode is RELOAD > > > > ldrb w0, [x19], 1 load opcode byte at ip, advance ip by 1 > > ldr x2, [x27, x0, lsl 3] load address of machine code from jmptable+opcode*8 > > br x2 jump to next machine code > > > > */ > > > > There are just 2 calls in the hole VM, in these cases 2 registers are > > pushed to maintain 16 byte alignment. > > > > You can avoid the 16 byte alignment by using a register other then sp > > for the processor stack. My tests showed this code to be about 30% slower > > in execution speed. > > What do you compare? Token threaded code to traditional threaded > code using memory addresses? > > On ARM64 calls have range +-128MB relative. So if the program code > does not exceed 128MB, then subroutine threaded code is smaller > than traditional threaded code and probably quite a bit faster. > Drawback is that non-leaf words need to push and pop return > address from the machine stack, using 16 bytes of stack space. > I have rerun some tests on my old RPi4. Now I do not get any difference in speed in using sp as stack pointer or another register. I might remember wrong as it was several years ago i did test this last time. I test a fib defined as code so threading type should not influence. code fib2 START: stp x29, x30, [sp, -16]! cmp x20, 2 b.ge L1 mov x20, 1 b END L1: sub x10, x20, 1 str x20, [x21, -8]! mov x20, x10 bl START ldr x9, [x21] sub x9, x9, 2 str x20, [x21] mov x20, x9 bl START ldr x8, [x21], 8 add x20, x20, x8 END: ldp x29, x30, [sp], 16 ret end-code note that this is a wrong fib as the fib in onebench.fs used by gforth have the same problem. I wanted the same number of iterations. My system works in 3 steps. 1. Compile forth source for token threaded VM 2. Compile the token code to native assembler code 3. Assemble with a built in assembler to native code For ARM64 only the first step is implemented Peter
[toc] | [prev] | [standalone]
Back to top | Article view | comp.lang.forth
csiph-web