Path: csiph.com!eternal-september.org!feeder.eternal-september.org!nntp.eternal-september.org!.POSTED!not-for-mail From: Kragen Javier Sitaker Newsgroups: comp.lang.forth,comp.arch,alt.lang.asm Subject: interrupt handling latency and weird register use (was Re: OT: Epic RISC-V rant) Date: Tue, 01 Sep 2026 17:00:23 -0300 Organization: Primarily biological and memetic Lines: 120 Message-ID: <87ecfcagiw.fsf_-_@debian> References: <87qzk0nejz.fsf@nightsong.com> <116d89d$27baf$1@paganini.bofh.team> MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 8bit Injection-Date: Tue, 01 Sep 2026 20:02:46 +0000 (UTC) Injection-Info: dont-email.me; logging-data="2260601"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX19cKneZSVyj8hzzTDFRh/wX"; posting-host="fa8aa933a678f631da317852f835fb16" User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/28.2 (gnu/linux) Cancel-Lock: sha1:hA791SX3XgB9MjI/VaWXMtlmyVU= sha1:M9nniHq+95ujD8LJMHrE8yvlPeY= sha256:SeUKujAJA1HWXhD7KfnwZg0ASgkIcuKhUDBgz2Ks2oo= sha1:lqdkpJ9Ir00uEUw78cs3JmGE9Rg= sha256:Uxnh+B/pxtZ5RqKJlThOJbXOV+TU1++ytJBpvvg60qs= Xref: csiph.com comp.lang.forth:135510 comp.arch:117718 alt.lang.asm:8890 antispam@fricas.org (Waldek Hebisch) writes: > Paul Rubin wrote: >> https://dmitry.gr/?r=06.%20Thoughts&proj=12.%20RV > > Some comments. > > 1) Interrupt latency claim is half-truth. First, normal RISC-V has > 32 registers while ARM has 16. If you can do with 16 registers > divide register set into 2 parts, use one part for normal code > and the other for interrupt handler. That way you will get few > cycle overhead, impossible feat for Cortex M. So it is really > a tradeoff; do you want to have modest overhead for pretty > typical use case, or do you for very low overhead in cases that > need it and higher overhead for typical cases? This is a really good point, and one I should have thought of. You can reserve some registers as “FIQ” registers and only use them inside interrupt handlers, depsite the absence of an architectural FIQ mechanism (which was present on the ARM2 but excised in the Cortex-M). (Note that most existing RISC-V cores like the QingKe core used in the CH32V003 have their own, incompatible, interrupt-handling mechanisms. Generally these are FIQ-like, enabling lower latency than the standard RISC-V mechanism but with a limited depth of nested interrupts.) You can write the “FIQ” handler itself in assembly, since if it needs to be more than about 16 instructions long you might as well switch to the standard ABI, but you also need to compile the rest of your application with the “FIQ registers” reserved, including any system libraries you might be using, such as an integer division subroutine. This is true whether you’re using Forth, or C, or any other language. 8 registers is probably enough for the FIQ handler, so with RV32I you have 24 left over for the rest of the code. Reserving some registers is a very small modification to most compilers (those that do some kind of register allocation), and Forth compilers are fairly small, but making any modification to GCC or LLVM is a bit daunting, even writing a new “machine definition” file. I had previously looked for an existing way to do this with GCC and given up, but it turns out that it’s fairly simple; GCC has an `-ffixed-reg` option, where `reg` is the name of the register to reserve. See . So, in theory, you ought to be able to reserve x24 through x31 for “FIQ” handlers by compiling all your C source code with something like gcc -ffixed-x24 -ffixed-x25 -ffixed-x26 -ffixed-x27 -ffixed-x28 \ -ffixed-x29 -ffixed-x30 -ffixed-x31 I have verified that GCC does indeed accept these command-line options, and that it respects -ffixed-a5, but I haven’t written code with sufficient register pressure to provoke GCC to try to use x24 (s8) even without this option. So, it’s documented to work, but it’s kind of a niche feature, and I haven’t tried it in practice. You might be tempted to try to use GCC’s global register variables for this: register int *foo asm ("r12"); However, the documentation specifically says that you should use `-ffixed-reg` instead, because global register variables will not work for the interrupt-handling case: >>> Similarly, it is not safe to access the global register variables >>> from signal handlers or from more than one thread of control. Unless >>> you recompile them specially for the task at hand, the system >>> library routines may temporarily use the register for other things. >>> Furthermore, since the register is not reserved exclusively for the >>> variable, accessing it from handlers of asynchronous signals may >>> observe unrelated temporary values residing in the register. For this particular case, another possibility may be to compile for RV32E, which only has 16 architectural general-purpose registers but is otherwise compatible with RV32I. I had looked for how to achieve this three years ago for a project called Monokokko, which is a five-machine-instruction-long cooperative-multitasking OS for ARM: .thumb_func yield: push {r4-r9, r11, lr} @ save all callee-saved regs except r10 str sp, [r10], #4 @ save stack pointer in current task ldr r10, [r10] @ load pointer to next task ldr sp, [r10] @ switch to next task's stack pop {r4-r9, r11, pc} @ return into yielded context there Now I know how to solve the problem! So I can write Monokokko tasks in C now! > 5) Optionality. I do not like it and it is probably biggest > problem of RISC-V. OTOH to have any chance of success RISC-V > need a buy-in from several independent parties. I suspect that > the only practical way to get consensus from varied parties is > by making most of specification optional. For CPU vendors and computer architecture researchers, I suspect this is the biggest selling point of RISC-V: they can add experimental vector extensions without waiting for them to be ratified, they can add FIQ-style low-latency interrupt handling, they can implement their own memory protection mechanisms, they can add zero-overhead loops, etc. As I understand it, Berkeley researchers’ nightmarish negotiations with ARM to get a license to perform such experiments on ARM cores was the hell from which RISC-V came in the first place. It also means that bad decisions by RISC-V International have much less impact on licensees. The most obvious example, to my mind, is something Grinberg’s critique doesn’t even mention — it’s having the divide instruction in the M extension. Hardware multiplication is crucial for all kinds of applications, including software-defined radio and other DSP, image processing, 3-D rendering, and neural networks. By contrast, hardware division is a minor advantage but rarely appears in inner loops, and has been omitted from historical architectures including the Cray-1, Cray-2, and ARM. Kragen