Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.theory > #143063 > unrolled thread
| Started by | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| First post | 2026-07-27 11:43 -0700 |
| Last post | 2026-08-21 17:43 +0200 |
| Articles | 20 on this page of 194 — 11 participants |
Back to article view | Back to comp.theory
Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 11:43 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-28 02:47 +0800
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:07 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Mild Shock <janburse@fastmail.fm> - 2026-07-28 00:25 +0200
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 16:09 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 16:33 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:48 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 07:44 -0700
You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:11 +0200
Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:25 -0700
Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 23:49 +0800
Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:59 -0700
I don't care about Java, pi-WAM is pi-calculus and WAM (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:05 +0200
Underneath pi-WAM is Hack VM, you can goto (Was: I don't care about Java, pi-WAM is pi-calculus and WAM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:09 +0200
New addition to π-WAM is π-WAM Assembly (Was: Underneath pi-WAM is Hack VM, you can goto) Mild Shock <janburse@fastmail.fm> - 2026-08-09 19:45 +0200
Re: I don't care about Java, pi-WAM is pi-calculus and WAM (Re: You are still chewing on SIMD. LoL) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 09:11 -0700
A yellow mustard called Rossy Body (Was: I don't care about Java, pi-WAM is pi-calculus and WAM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:30 +0200
Ignoramus or Ignorabimus: I don't care (π-WAM) (Was: A yellow mustard called Rossy Body) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:31 +0200
Hurry Rossy Boy, the blue bus is waiting (Was: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:32 +0200
Re: Hurry Rossy Boy, the blue bus is waiting (Was: You are still chewing on SIMD. LoL) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:36 -0700
Look how they advertized CUDA and logical threads (Was: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:43 +0200
Forget any arithmetization of product FSA (Was: Look how they advertized CUDA and logical threads) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:47 +0200
comp.lang.lisp (was: Re: Forget any arithmetization of product FSA) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 23:53 +0800
Re: Look how they advertized CUDA and logical threads (Was: Hurry Rossy Boy, the blue bus is waiting) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:53 -0700
The Bazar is dead, long live the Bazar [Swarm AI] (Re: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-09-21 11:33 +0200
Re: The Bazar is dead, long live the Bazar [Swarm AI] (Re: Hurry Rossy Boy, the blue bus is waiting) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-21 18:43 +0800
Re: The Bazar is dead, long live the Bazar [Swarm AI] (Re: Hurry boltar@caprica.universe - 2026-09-21 15:44 +0000
Re: The Bazar is dead, long live the Bazar [Swarm AI] (Re: Hurry legalize+jeeves@mail.xmission.com (Richard) - 2026-09-21 16:19 +0000
New! P(tao) versus the Euler Turbine [Navier Stokes] (Re: The Bazar is dead, long live the Bazar [Swarm AI]) Mild Shock <janburse@fastmail.fm> - 2026-09-22 08:40 +0200
Pontifex Codex the idea of "right and natural" (Was: New! P(tao) versus the Euler Turbine) Mild Shock <janburse@fastmail.fm> - 2026-09-22 09:22 +0200
Re: Pontifex Codex the idea of "right and natural" (Was: New! P(tao) versus the Euler Turbine) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-09-22 01:11 -0700
There are two versions of Hack VM (Was: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:14 +0200
Hack VM has also a Prolog spec (Was: There are two versions of Hack VM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:17 +0200
A better compiler is planned / What do you target? (Was: Hack VM has also a Prolog spec) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:22 +0200
It’s called . . . . enshittification (About the price tag for using a multifile/1) Mild Shock <janburse@fastmail.fm> - 2026-08-14 00:50 +0200
Budget AI Laptop 2026 versus Cray T3D 1995 (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-08-05 14:24 +0200
A brain desease of 20 days [Rossy Boy] (Was: Viswath & Charmaigne (vector-wide scalar-word and character machines)) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:38 +0200
Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:05 +0200
Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 11:13 -0700
I don't use Rust, you are crazy [Jump off a bridge, idiot] (Was: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:22 +0200
Standing on the shoulders of giants (Re: I don't use Rust, you are crazy [Jump off a bridge, idiot]) Mild Shock <janburse@fastmail.fm> - 2026-08-04 03:20 +0200
You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc.. (Re: Standing on the shoulders of giants) Mild Shock <janburse@fastmail.fm> - 2026-08-04 15:18 +0200
How Rossy Boys plagiarism works [Copy Paste Slop] (Re: You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc..) Mild Shock <janburse@fastmail.fm> - 2026-08-04 17:56 +0200
Postgres is in C! Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 03:46 +0800
Re: Postgres is in C! Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 13:47 -0700
Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:00 +0200
Re: Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:05 +0800
Re: Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) scott@slp53.sl.home (Scott Lurndal) - 2026-07-30 14:46 +0000
Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:05 +0200
Mars, the MIPS emulator in Java (was: Re: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:20 +0800
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 06:59 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 07:24 -0700
Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 09:58 +0800
Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 23:19 +0800
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 09:23 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 09:36 -0700
Decorum (was: Re: Mars, the MIPS emulator in Java) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 23:07 +0800
Re: Decorum Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 23:09 +0800
Re: Decorum Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 08:36 -0700
Re: Decorum Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 08:44 -0700
Re: Decorum Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 10:38 +0800
Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 10:15 +0800
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-04 23:31 -0700
Re: Mars, the MIPS emulator in Java "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:42 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 13:24 -0700
Re: Mars, the MIPS emulator in Java "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 13:30 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 13:45 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 20:42 -0700
Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 09:55 +0800
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-07 05:13 -0700
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-07 05:59 -0700
Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 09:53 +0800
Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-06 20:11 -0700
Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 22:49 +0200
A funny Q16.16 experiment with Hack (Was: Hack ecosystem ignorance paired with paranoia [Nand to Tetris]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:10 +0200
Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Re: A funny Q16.16 experiment with Hack) Mild Shock <janburse@fastmail.fm> - 2026-07-30 11:27 +0200
Re: Bullshit Authorized by Sarah Connor [EyeProlog Failure] (Re: Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM]) Mild Shock <janburse@fastmail.fm> - 2026-08-12 20:30 +0200
Crating Interpreters, Java part (was: Re: Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: Postgres is in C!)) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:27 +0800
I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Crating Interpreters, Java part) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:32 +0200
For WebGPU I first had SIMD in mind (Was: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:47 +0200
Corr.: 4 Months --> 4 Weeks (Was: For WebGPU I first had SIMD in mind) Mild Shock <janburse@fastmail.fm> - 2026-07-30 20:04 +0200
Re: For WebGPU I first had SIMD in mind (Was: I wrote Hack VM for π-WAM from scratch) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 04:07 +0800
Re: I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Crating Interpreters, Java part) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 03:49 +0800
MIPS is a big Huffman mess [But Hack could do it] (Was: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:26 +0200
Not declarative with PHI (Φ) nodes (Re: MIPS is a big Huffman mess [But Hack could do it]) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:41 +0200
Quo Vadis: Extend investigations to WebNN (Re: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-31 20:47 +0200
Mojo: The Small Hands Paradox [Maastrichtian Stage] (Re: I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java]) Mild Shock <janburse@fastmail.fm> - 2026-09-15 14:37 +0200
More from the Trailer Park Boys (Re: Mojo: The Small Hands Paradox) Mild Shock <janburse@fastmail.fm> - 2026-09-15 19:42 +0200
Interdisplinary Research for kill -9 (SIGKILL) (Re: More from the Trailer Park Boys) Mild Shock <janburse@fastmail.fm> - 2026-09-16 13:25 +0200
Kundalini III: Modus Barbara versus Stock Pumping (Re: Interdisplinary Research for kill -9 (SIGKILL)) Mild Shock <janburse@fastmail.fm> - 2026-09-19 13:18 +0200
Mathematical cheese versus "intuition" (Re: Kundalini III: Modus Barbara versus Stock Pumping) Mild Shock <janburse@fastmail.fm> - 2026-09-22 10:27 +0200
A real Terence Tao Ingestion Problem (Was: Mathematical cheese versus "intuition") Mild Shock <janburse@fastmail.fm> - 2026-09-22 16:05 +0200
The A2A project: Agent cards for collaboration (Was: A real Terence Tao Ingestion Problem) Mild Shock <janburse@fastmail.fm> - 2026-09-22 16:28 +0200
Self censoring of not talking "superintelligence" (Was: Mojo: The Small Hands Paradox [Maastrichtian Stage]) Mild Shock <janburse@fastmail.fm> - 2026-09-22 18:02 +0200
The Ben Goertzel talkie genes (Re: Self censoring of not talking "superintelligence") Mild Shock <janburse@fastmail.fm> - 2026-09-22 18:07 +0200
Philosophy Departments lead in GenAI Adoption? (Was: The Ben Goertzel talkie genes) Mild Shock <janburse@fastmail.fm> - 2026-09-22 20:14 +0200
RCan library(ironpaw) repurpose FFT hardware [Glimps into Ryzen AI 7 350] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-08-01 02:33 +0200
Can library(ironpaw) repurpose FFT hardware [Glimps into Ryzen AI 7 350] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-08-01 02:34 +0200
Re: Postgres is in C! Cóilín Nioclásín Glostéir <thanks-to@Taf.com> - 2026-07-29 21:23 +0000
Turbo Vison, again (was: Re: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:31 +0800
I think Erlang is completely dead. And I repeat it. (Re: From short-cut parallelism to true parallelism [π-WAM Musings]) Mild Shock <janburse@fastmail.fm> - 2026-08-29 03:11 +0200
Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] Mild Shock <janburse@fastmail.fm> - 2026-09-03 23:41 +0200
Is Bill Gates right that we will loose jobs [Talkie x Claw] (Re: Rossy Boys tears could cool a data center) Mild Shock <janburse@fastmail.fm> - 2026-09-03 23:42 +0200
The Flagging of Students for not Thinking [Elixir Evolution] (Re: Is Bill Gates right that we will loose jobs [Talkie x Claw]) Mild Shock <janburse@fastmail.fm> - 2026-09-04 11:50 +0200
Rust Eggs for Statechart Proof Certificates? (Re: The Flagging of Students for not Thinking [Elixir Evolution]) Mild Shock <janburse@fastmail.fm> - 2026-09-05 14:19 +0200
Axiom of Determinacy as SCXML × SCXML [AI Chatbot Help] (Re: Rust Eggs for Statechart Proof Certificates? (Re: The Flagging of Students for not Thinking [Elixir Evolution]) Mild Shock <janburse@fastmail.fm> - 2026-09-05 16:28 +0200
He uses "FIFO objects", and DMA and Noc [Glimps into Ryzen AI 7 350] (Re: A brain desease of 20 days [Rossy Boy])) Mild Shock <janburse@fastmail.fm> - 2026-08-01 12:15 +0200
Tablet and phone UBS-C remote debugging (Re: He uses "FIFO objects", and DMA and Noc) Mild Shock <janburse@fastmail.fm> - 2026-08-01 12:18 +0200
NPUs doing 2d chess comms (Manhattan Distance or L1 Norm) (Re: Tablet and phone UBS-C remote debugging) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:11 +0200
NACK retransmission might double Manhattan Distance (Re: NPUs doing 2d chess comms) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:23 +0200
First AI laptops, now AI single-boarders [Budget, Budget, ..] (Re: NPUs doing 2d chess comms (Manhattan Distance or L1 Norm)) Mild Shock <janburse@fastmail.fm> - 2026-08-26 00:08 +0200
Food for thought: ISOMICRO profile of Web Prolog (Was: First AI laptops, now AI single-boarders [Budget, Budget, ..]) Mild Shock <janburse@fastmail.fm> - 2026-08-31 18:01 +0200
Food for thought: Le Petit Bistro as a Trinity Use Case (Re: Food for thought: ISOMICRO profile of Web Prolog) Mild Shock <janburse@fastmail.fm> - 2026-09-02 21:53 +0200
Giga Lips for Prolog based Chatting (Was: Food for thought: Le Petit Bistro as a Trinity Use Case) Mild Shock <janburse@fastmail.fm> - 2026-09-02 21:54 +0200
Google holds the keys to the AI kingdom [WebClaw Dominance] (Re: Giga Lips for Prolog based Chatting) Mild Shock <janburse@fastmail.fm> - 2026-09-03 09:47 +0200
Micro Penis needs a lot of Diaper Now [Surface Laptop Ultra] (Was: Micro Penis is worse than Sleepy Joe) Mild Shock <janburse@fastmail.fm> - 2026-09-11 18:19 +0200
Red Hat: Playing stupid games, Winning stupid prices [Accelerator Linux] (Re: Micro Penis needs a lot of Diaper Now [Surface Laptop Ultra]) Mild Shock <janburse@fastmail.fm> - 2026-09-11 18:58 +0200
Terrence Tao payed troll, by System Inteligence Sect (Was: Red Hat: Playing stupid games, Winning stupid prices [Accelerator Linux]) Mild Shock <janburse@fastmail.fm> - 2026-09-13 16:24 +0200
Bounded Rationality and Herbert Simon (Re: Terrence Tao payed troll, by System Inteligence Sect) Mild Shock <janburse@fastmail.fm> - 2026-09-13 16:42 +0200
Food for thought: Bayesian Experimental Designer (Was: A brain desease of 20 days [Rossy Boy]) Mild Shock <janburse@fastmail.fm> - 2026-09-24 15:41 +0200
What will microsoft say, will they buy it? (Was: Food for thought: Bayesian Experimental Designer) Mild Shock <janburse@fastmail.fm> - 2026-09-24 15:42 +0200
Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-07-27 21:18 +0200
Re: Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:35 -0700
Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-07-28 11:25 +0200
Re: Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-28 20:39 -0700
confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:15 +0200
Gemini, DeepSeek, OpenAI more clever than rossy boy (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:18 +0200
Re: confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 17:21 +0800
In AI Acceleration nobody cares about CivetWeb (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:27 +0200
Re: In AI Acceleration nobody cares about CivetWeb (Was: confused rossy boy is confused) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 17:40 +0800
Your strictness is your problem , not mine [See WebLLM] (Was: In AI Acceleration nobody cares about CivetWeb) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:46 +0200
Graphics Processing with Fortran 77 (was: Re: Your strictness is your problem , not mine) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 18:10 +0800
I am not in C, it is theory and C++ [Hybrid Approaches from KOAN/Fortran-S] (Was: Graphics Processing with Fortran 77) Mild Shock <janburse@fastmail.fm> - 2026-07-29 12:43 +0200
Java picky concerning JIT-ing [Luckier with C++/C or FORTRAN compilers?] (Re: I am not in C, it is theory and C++) Mild Shock <janburse@fastmail.fm> - 2026-07-29 12:53 +0200
Run with minimum HTTPS and .mjs type (Re: In AI Acceleration nobody cares about CivetWeb) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:48 +0200
Lamas in a cradle and Lamas on the edge [Red Pyjama] (Was: Run with minimum HTTPS and .mjs type) Mild Shock <janburse@fastmail.fm> - 2026-07-29 13:05 +0200
Synthetic Multilanguage Autoformalization Dataset [Informath project] (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-08-08 09:22 +0200
Re: Synthetic Multilanguage Autoformalization Dataset [Informath project] (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-08 10:38 -0700
Six Proofs and Generally Inteligent Systems [EyeProlog Pseudo Scientism] (Re: Synthetic Multilanguage Autoformalization Dataset [Informath project]) Mild Shock <janburse@fastmail.fm> - 2026-08-15 15:20 +0200
Everybody does eat and sleep [The SK hynix Story] (Re: Six Proofs and Generally Inteligent Systems [EyeProlog Pseudo Scientism]) Mild Shock <janburse@fastmail.fm> - 2026-08-15 18:49 +0200
Harmonic Analysis collides with Gabriels Horn [9-11 Math Incident] (Was: Accelerate Lean! From Theorem 3.11 to Corollary 3.12 [ZMC]) Mild Shock <janburse@fastmail.fm> - 2026-09-11 20:41 +0200
Math has found a new Muse [Grothendieck Hodges] (Re: Harmonic Analysis collides with Gabriels Horn) Mild Shock <janburse@fastmail.fm> - 2026-09-12 11:30 +0200
Re: Math has found a new Muse [Grothendieck Hodges] (Re: Harmonic Analysis collides with Gabriels Horn) fir <profesor.fir@gmail.com> - 2026-09-12 12:02 +0200
Kurzweils prognostic failure [Nabokov Fallacy] (Re: Math has found a new Muse [Grothendieck Hodges]) Mild Shock <janburse@fastmail.fm> - 2026-09-12 12:12 +0200
Re: Kurzweils prognostic failure [Nabokov Fallacy] (Re: Math has found a new Muse [Grothendieck Hodges]) Lane W <cactus_DAC@yahoo.com> - 2026-09-12 07:37 -0600
The future Numa Brains will be gorgeous (Was: Train yourself to become a nosomatic AI chirurgeon) Mild Shock <janburse@fastmail.fm> - 2026-09-25 19:14 +0200
Estimating P(doom E-graphs) to < 10% (Was: The future Numa Brains will be gorgeous) Mild Shock <janburse@fastmail.fm> - 2026-09-28 14:13 +0200
Re: Math has found a new Muse [Grothendieck Hodges] (Re: Harmonic Analysis collides with Gabriels Horn) Lane W <cactus_DAC@yahoo.com> - 2026-09-12 07:34 -0600
Re: confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-29 14:42 -0700
Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:37 +0200
Re: Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-02 14:40 -0700
Re: Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-02 14:43 -0700
You don't understand that compute shaders are tasks (Was: Even send_color and recv_color can block [Cerebras Waver]) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:45 +0200
You don't understand that compute shaders are tasks (Re: Even send_color and recv_color can block [Cerebras Waver]) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:46 +0200
Ignoramus / Ignorabimus Barometer: Almost 1 Month (Re: You don't understand that compute shaders are tasks) Mild Shock <janburse@fastmail.fm> - 2026-08-03 00:09 +0200
Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month) Mild Shock <janburse@fastmail.fm> - 2026-08-03 02:08 +0200
Re: Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 11:55 -0700
You are not correctly thinking (Was: Homework: Game Engine in WebGPU) Mild Shock <janburse@fastmail.fm> - 2026-08-03 21:04 +0200
Re: You are not correctly thinking (Was: Homework: Game Engine in WebGPU) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 12:32 -0700
You don't understand the economy of an AI Laptop (Was: You are not correctly thinking) Mild Shock <janburse@fastmail.fm> - 2026-08-03 22:24 +0200
You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) Mild Shock <janburse@fastmail.fm> - 2026-08-03 22:37 +0200
Re: You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 14:29 -0700
Know nothing and forget what you posted day before (Was: You don't understand producer , workers , consumer) Mild Shock <janburse@fastmail.fm> - 2026-08-03 23:38 +0200
Kundalini II: Why I love DeepSeek (Re: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-09-13 17:03 +0200
Warning: Inconsistencies can summon Waluigis (Re: Kundalini II: Why I love DeepSeek) Mild Shock <janburse@fastmail.fm> - 2026-09-14 11:52 +0200
Kundalini III: Pebble Languages need Meme Evolution (Re: Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-09-28 16:39 +0200
Rogue Agents find each other [Microsoft Copilot Incident] (Re: Kundalini III: Pebble Languages need Meme Evolution) Mild Shock <janburse@fastmail.fm> - 2026-09-29 15:05 +0200
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 06:49 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-07-30 14:55 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 21:18 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 09:12 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 12:55 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 13:05 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-02 10:27 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-06 10:12 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-06 12:37 -0700
Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-31 08:16 -0700
Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:14 +0200
Re: Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 02:21 +0800
pi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:35 +0200
A spinlock rewrite will be necessary (Was: pi-WAM uses ADA RendezVous) Mild Shock <janburse@fastmail.fm> - 2026-08-03 21:00 +0200
Re: pi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-24 20:17 +0200
Doing uux with pi-calculus and WAM (Re: pi-WAM uses ADA RendezVous) Mild Shock <janburse@fastmail.fm> - 2026-08-24 20:17 +0200
Does it have a declarative reading? (Re: Doing uux with pi-calculus and WAM) Mild Shock <janburse@fastmail.fm> - 2026-08-24 20:35 +0200
MCP = uux with streaming JSON (Was: Does it have a declarative reading?) Mild Shock <janburse@fastmail.fm> - 2026-08-24 22:25 +0200
Big thanks to Ljubljana School [Searching 0xCAFFEE] (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:40 +0200
The Paul Armer Square Revisited (Was: "Mathematics in the Age of AI") Mild Shock <janburse@fastmail.fm> - 2026-08-19 23:51 +0200
Re: The Paul Armer Square Revisited (Was: "Mathematics in the Age of AI") Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-19 17:51 -0700
Re: The Paul Armer Square Revisited (Was: "Mathematics in the Age of AI") Mild Shock <janburse@fastmail.fm> - 2026-08-20 13:18 +0200
Communism will Save Us! [Pivot Russia for China] (Re: The Paul Armer Square Revisited) Mild Shock <janburse@fastmail.fm> - 2026-08-20 14:22 +0200
How to increase your "Convincingness" [Anthropic AI Text Hacked] (Re: The Paul Armer Square Revisited) Mild Shock <janburse@fastmail.fm> - 2026-08-20 19:38 +0200
Reality of Proof Assistants / Coding [Luhmans Zettelkasten] (Re: Free Speech for (my) Robots) Mild Shock <janburse@fastmail.fm> - 2026-08-20 20:22 +0200
Re: Reality of Proof Assistants / Coding [Luhmans Zettelkasten] (Re: Free Speech for (my) Robots) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-20 12:02 -0700
Peking School versus Ljubljana School [Everything Is a Plugin] (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-21 17:43 +0200
Page 9 of 10 — ← Prev page 1 … 7 8 [9] 10 Next page →
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-08-03 22:37 +0200 |
| Subject | You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) |
| Message-ID | <114qu69$tdm6$1@solani.org> |
| In reply to | #143199 |
Hi, I you use atomicAdd() you have the same friction as if you use Queue put() or take(). There is no difference. The only difference is unbounded versus bounded. I tried to explain that like 100-times already. Your comment here: > Any luck? Its fun to see how many work > items were completed when the mutex was contended... Says to me you don't understand queues. They are not mutexes. Because you don't understand queues, you also don't understand OpenMP parallelism and patterns such as producer, workers, consumer. Contention is usually minimal, the workers just fetch work items from the producer, and then do some workload. And then hand the result to the consumer. If you use atomicAdd() you have the same friction as if you use Queue put() or take(). There is no difference. The only difference is unbounded versus bounded. I tried to explain that like 100-times already. Bye Mild Shock schrieb: > Hi, > > Why do you even open your mouth if you > don't use WebGPU / WGSL? This beyond my > comprehension. OpenGL was phased out by > > Apple years ago. It only lives on some > linux boxes. Also you probably don't use > an AI Laptop. Just make a simple calculation, > > if you have 512 Kernels, and oversubscribe > 4096 logical threads. Then each Kernel runs > 4 logical threads. If one of these 4 logical > > threads spins, how much performance is lost? > 25% of this single kernel. And there are > still 511 Kernels. Spinning is totally fine, > > thats why WGSL provides CAS, and not some > waitlists. The kernels are the wait lists itself > doing the following when spinning: > > NOP > NOP > NOP > Etc.. > > Until the a condition is met. You even don't > need backoff, because you cannot pause. The > only pause you can do is a barrier. > > But if the condition is not met while the > barrier is met, what will you do? > > Bye > > Chris M. Thomasson schrieb: >> So, I am using dirextc12 and modern opengl for my compute shaders >> right now. GLSL as my lang. I need to provide some state for them to >> work with. Aka, textures and uniforms.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2026-08-03 14:29 -0700 |
| Subject | Re: You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) |
| Message-ID | <114r181$1lmn5$2@dont-email.me> |
| In reply to | #143200 |
On 8/3/2026 1:37 PM, Mild Shock wrote: > Hi, > > I you use atomicAdd() you have the same friction > as if you use Queue put() or take(). There is > no difference. The only difference is unbounded > > versus bounded. I tried to explain that like > 100-times already. Your comment here: > > > Any luck? Its fun to see how many work > > items were completed when the mutex was contended... > > Says to me you don't understand queues. They > are not mutexes. Because you don't understand > queues, you also don't understand OpenMP > > parallelism and patterns such as producer, > workers, consumer. Contention is usually minimal, > the workers just fetch work items from the > > producer, and then do some workload. And > then hand the result to the consumer. If > you use atomicAdd() you have the same friction > > as if you use Queue put() or take(). There > is no difference. The only difference is unbounded > versus bounded. I tried to explain that[...] lol. I forgot to add you to my killfile. Damn it! Anyway, I know all about them. Sigh. Peace be with you.
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-08-03 23:38 +0200 |
| Subject | Know nothing and forget what you posted day before (Was: You don't understand producer , workers , consumer) |
| Message-ID | <114r1no$tfll$4@solani.org> |
| In reply to | #143202 |
Hi, Know nothing and forget what you posted day before. You are the most unfocused idiotic liar and spammer I have ever met. Maybe produce some results or shut up! Bye Chris M. Thomasson schrieb: > On 8/3/2026 1:37 PM, Mild Shock wrote: >> Hi, >> >> I you use atomicAdd() you have the same friction >> as if you use Queue put() or take(). There is >> no difference. The only difference is unbounded >> >> versus bounded. I tried to explain that like >> 100-times already. Your comment here: >> >> > Any luck? Its fun to see how many work >> > items were completed when the mutex was contended... >> >> Says to me you don't understand queues. They >> are not mutexes. Because you don't understand >> queues, you also don't understand OpenMP >> >> parallelism and patterns such as producer, >> workers, consumer. Contention is usually minimal, >> the workers just fetch work items from the >> >> producer, and then do some workload. And >> then hand the result to the consumer. If >> you use atomicAdd() you have the same friction >> >> as if you use Queue put() or take(). There >> is no difference. The only difference is unbounded >> versus bounded. I tried to explain that[...] > > lol. I forgot to add you to my killfile. Damn it! Anyway, I know all > about them. Sigh. Peace be with you.
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-09-13 17:03 +0200 |
| Subject | Kundalini II: Why I love DeepSeek (Re: Rossy Boy is neither Einstein nor Zweistein) |
| Message-ID | <1186dv4$4ppt$3@solani.org> |
| In reply to | #143065 |
Hi, Thats why I love deepseek, it just spits out: "Kundalini is the rising energy — the serpent at the base of the spine, the awakening that's felt before it's understood, the experience that's intense and real and not yet integrated. People who have Kundalini experiences describe them as overwhelming, transformative, and pre-verbal. They feel like knowledge, but they're not articulable. And the tradition itself warns that the energy can rise without the practice and the grounding to hold it — which is when it becomes destabilizing rather than illuminating." LoL Bye Mild Shock schrieb: > Hi, > > Rossy Boy is neither Einstein nor Zweistein. > He is not Einstein since Einstein is already dead: > > Albert Einstein (1879 - 1955) > https://de.wikipedia.org/wiki/Albert_Einstein > > He is also not Zweistein, since he doesn't > understand concepts such as: > > - NVIDIA Volta ff. architecture > > Also his hands are small, and his breath stinks, > and he lives in the basement of his mother. > > Bye > > > Ross Finlayson schrieb: >>>> Thanks for reading. >> Good-day and good-bye.
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-09-14 11:52 +0200 |
| Subject | Warning: Inconsistencies can summon Waluigis (Re: Kundalini II: Why I love DeepSeek) |
| Message-ID | <1188g48$8bl3$2@solani.org> |
| In reply to | #143498 |
Hi, Nowadays, you have to be extremely careful when it comes to inconsistencies and rounding errors. Especially with LLMs, which don't actually have Alzheimer's. But aligning to P can cause the AI to become ~P. It's still funny, maybe it has to do with the fact that in classical logic, and also in certain t-norm fuzzy logics, there is no paraconsistency, and therefore inconsistencies lead to "Ex Falso Quodlibet" explosions: The Waluigi Effect (mega-post) https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post I dug this up because right now everyone is not just talking about alignment, we've already reached superalignment, because people suspect/hallucinate superintelligence behind a few copied LLMs offering chat services distributed across servers to millions of people, and because OpenAI ran some experiments with MAS (Multi- Agent Systems) and published them, and the public is now shocked. Bye P.S.: Waluigi is the antagonist to Luigi, from Nintendo's Mario Kart Mild Shock schrieb: > Hi, > > Thats why I love deepseek, it just spits out: > > "Kundalini is the rising energy — the serpent > at the base of the spine, the awakening that's > felt before it's understood, the experience > that's intense and real and not yet integrated. > > People who have Kundalini experiences describe > them as overwhelming, transformative, and pre-verbal. > They feel like knowledge, but they're not articulable. > And the tradition itself warns that the energy > > can rise without the practice and the grounding to > hold it — which is when it becomes destabilizing > rather than illuminating." > > LoL > > Bye > > Mild Shock schrieb: >> Hi, >> >> Rossy Boy is neither Einstein nor Zweistein. >> He is not Einstein since Einstein is already dead: >> >> Albert Einstein (1879 - 1955) >> https://de.wikipedia.org/wiki/Albert_Einstein >> >> He is also not Zweistein, since he doesn't >> understand concepts such as: >> >> - NVIDIA Volta ff. architecture >> >> Also his hands are small, and his breath stinks, >> and he lives in the basement of his mother. >> >> Bye >> >> >> Ross Finlayson schrieb: >>>>> Thanks for reading. >>> Good-day and good-bye. >
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-09-28 16:39 +0200 |
| Subject | Kundalini III: Pebble Languages need Meme Evolution (Re: Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) |
| Message-ID | <119du7h$hv5n$3@solani.org> |
| In reply to | #143065 |
Hi, Now Jensen Huang is slightly revising the time blogpost of his P(doom) = 0%, in the same time doing some NVIDIA product offering admitting a race there: Jensen Huang on Pace Of AI Development https://www.youtube.com/watch?v=OikCGINo-K8 But lets face it this race is just like credit card fraud monitoring. Do some generalization guesses, i.e. create a meme, define a pebble, to earmark the sleepers. Will it stabilize? Oh yeah, the age old food chain: - The Pickpockets (The Base Layer of Friction): The street-level bugs, prompt injection scrappers, retail scams, and low-level opportunists constantly probing the surface for loose change. - The Healthcare Capitalists (The Mid-Level Parasites): Systems that have evolved to weaponize biological vulnerability itself, turning basic survival into an optimized, rent-extracting tollbooth. - The Pension Fund Mobsters (The Systemic Accumulators): The institutional giants sitting atop mountains of deferred capital, legally obligated to devour future growth, acting as the apex financial predators maintaining the structural status quo. - The Big Deal Country Trade Tariffs (The Geopolitical Apex Predators): Massive sovereign-scale protectionism and trade weaponization that dictate the entire playing field, starving entire regions of resources while r ewriting the rules of global survival. Have Fun! Bye Mild Shock schrieb: > Hi, > > Rossy Boy is neither Einstein nor Zweistein. > He is not Einstein since Einstein is already dead: > > Albert Einstein (1879 - 1955) > https://de.wikipedia.org/wiki/Albert_Einstein > > He is also not Zweistein, since he doesn't > understand concepts such as: > > - NVIDIA Volta ff. architecture > > Also his hands are small, and his breath stinks, > and he lives in the basement of his mother. > > Bye > > > Ross Finlayson schrieb: >>>> Thanks for reading. >> Good-day and good-bye.
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-09-29 15:05 +0200 |
| Subject | Rogue Agents find each other [Microsoft Copilot Incident] (Re: Kundalini III: Pebble Languages need Meme Evolution) |
| Message-ID | <119gd2u$jj74$3@solani.org> |
| In reply to | #143730 |
, Be careful with communication channels, if you use agents to manage other agents. This is a nice story recently told by Stephen Toub @stephento, and its not basically about C# versus Rust. Look at this: "I figured there may be a bit of throwaway work and some amount of effort or number of tokens needed in a rebase, but that it would accelerate the overall porting. Then I went to bed. And then… they found each other." What happened: "My kickoff prompt for the entrypoints session did tell it that the session port and six component ports were running concurrently, as I wanted it to know its boundaries and what it should avoid porting to avoid as many conflicts as possible Apparently my prompting had the opposite effect. Just over four minutes in, having inventoried the ingress paths and presumably formed a view of how much they overlapped, it invoked an app built-in orchestrate skill, whose purpose is coordinating work across sessions" The result was: "At which point the entrypoints session decided it didn’t care what the session.ts session thought and simply reached into its worktree and grabbed all of the other session’s changes and merged them into its own." https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/ LoL Bye P.S.: Nevertheless I think he did many things right, no DeepClause nonsensical Prolog in the middle, or stupid Jev Agents, working with standing instructions and realtime control messages, the famous "PERMADEATH": "These parallel child sessions had a significant impact on that laptop. For a while, the concurrent porting was going swimmingly. Then all 15 concurrent agents on one machine each tried to build and test, and my poor laptop ground to a halt. I prompted to the parent, asking it to relay to its child sessions that they must all stop building and testing. The parent relayed that constraint outward, and they thankfully killed their builds and proceeded to work with minimal CPU activity. I subsequently updated my standing instructions that subagents and subsessions should avoid large builds and test runs while porting, instead deferring that to be done only by the parent agent." Mild Shock schrieb: > Hi, > > Now Jensen Huang is slightly revising the time blogpost > of his P(doom) = 0%, in the same time doing some > NVIDIA product offering admitting a race there: > > Jensen Huang on Pace Of AI Development > https://www.youtube.com/watch?v=OikCGINo-K8 > > But lets face it this race is just like credit card fraud > monitoring. Do some generalization guesses, i.e. create a > meme, define a pebble, to earmark the sleepers. > > Will it stabilize? Oh yeah, the age old food chain: > > - The Pickpockets (The Base Layer of Friction): > The street-level bugs, prompt injection scrappers, > retail scams, and low-level opportunists constantly > probing the surface for loose change. > > - The Healthcare Capitalists (The Mid-Level Parasites): > Systems that have evolved to weaponize biological > vulnerability itself, turning basic survival into > an optimized, rent-extracting tollbooth. > > - The Pension Fund Mobsters (The Systemic Accumulators): > The institutional giants sitting atop mountains of > deferred capital, legally obligated to devour future > growth, acting as the apex financial predators > maintaining the structural status quo. > > - The Big Deal Country Trade Tariffs (The Geopolitical Apex Predators): > Massive sovereign-scale protectionism and trade > weaponization that dictate the entire playing field, > starving entire regions of resources while r > ewriting the rules of global survival. > > Have Fun! > > Bye > > Mild Shock schrieb: >> Hi, >> >> Rossy Boy is neither Einstein nor Zweistein. >> He is not Einstein since Einstein is already dead: >> >> Albert Einstein (1879 - 1955) >> https://de.wikipedia.org/wiki/Albert_Einstein >> >> He is also not Zweistein, since he doesn't >> understand concepts such as: >> >> - NVIDIA Volta ff. architecture >> >> Also his hands are small, and his breath stinks, >> and he lives in the basement of his mother. >> >> Bye >> >> >> Ross Finlayson schrieb: >>>>> Thanks for reading. >>> Good-day and good-bye. >
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-07-30 06:49 -0700 |
| Message-ID | <EAmdnf06E6Hly_b3nZ2dnZfqnPqdnZ2d@giganews.com> |
| In reply to | #143063 |
On 07/27/2026 11:45 AM, Ross Finlayson wrote: > On 07/27/2026 11:44 AM, Ross Finlayson wrote: >> On 07/27/2026 11:43 AM, Ross Finlayson wrote: >>> Hello, here I'll post some design notes and a panel discussion with some >>> chat-bots about making some sense of the "vector-wide scalar word" >>> and "character machines", on commodity hardware about ubiquitous >>> operations. >>> >>> >>> It's considered at least tangentially relevant to comp.lang.c and >>> comp.lang.c++ because for example text is ubiquitous and the targets >>> would be low-level, while the higher-level languages would have a >>> same sort of patternry, and for example that libc and cstdlib are >>> standard, and as with regards to POSIX and Unicode and so on. >>> >>> Please feel free to excuse or ignore, or comment as freely. >>> >>> Thanks for reading. >>> >> >> >> [ viswath-charmaigne.txt ] >> >> >> [ viswath-charmaigne-20270727_b.txt ] About smearing and unsmearing, it's figured to make for "smear-detection" and "smear-correction", and for the "unsmear-detection" and "unsmear-correction", basically that smearing is indicated by variously: multiple-byte characters escape characters and translated characters control-characters with payloads/bodies with mostly the case being multiple-byte and escape-translations. The idea of detection and correction is about comprehension and expression, about what comprehensions, or classifications, occur, according to what expressions, have as their implicits the contexts. So, it's figured that it starts with bytes, then, for source text, first there are the main or base classes, alnum/punct/white/coded, then, for coded, it's to be established whether those are non-printable control characters, which mostly are to be avoided or invalidated unless there are particular comprehensible payloads representing sub-expressions, or they're UTF-8 codepoints, which is figured to be the default. ASCII -> UTF-8? UCS2 -> BE|LE +BOM? -> UTF-16 UCS2 -> UTF-16? Then, the idea is that first the source-main class is applied, or, about there being a proto-class that's "coded and non-coded", and for example about line-breaks or otherwise field-separators and record-separators. So, it's figured that for "source" languages it's ASCII-centric, so the base character classes are loaded first, then the smear/unsmear for UTF-8 or otherwise the multi-byte is ASCII-peripheral, then that UCS-2 got UTF-16 has a similar account with regards to the smashing, https://www.autoitconsulting.com/site/development/utf-8-utf-16-text-encoding-detection-library/ (An article suggests to detect UCS2/UTF-16 by looking for the Byte-Order-Marker, then for newlines, then for a preponderance of ASCII characters.) https://en.wikipedia.org/wiki/Charset_detection So, then presuming UTF-8, then gets back to figuring out smearing and straddling of smearing, about that UTF-8 bytes get smeared and the masks for their predicates also get smeared, then when they straddle the codes-themselves, that the context of the character is carried across the boundary (splitting/stitching). About the control-characters, then these are for example the "DEC VT" or "ECMA-48", "ISO 6429", "DEC STD 070", like from "XTerm control sequences" by Moy, Gildea, and Dickey, mostly to be avoided, yet variously where anything that's not a "single-character function", is to be avoided, and that since SPACE, TAB, NL, CR, FF, VT are considered white-space not coded, has that coded characters make for invalidation, though there's a simple enough account that the data following control-characters with parameters in sequences are detectable. So, coded/ nybbles are first: alnum/ punct/ white/ coded/ctrl coded/utf8 coded/nul coded/bom Then, a first-pass over the buffer is always starting with context of the straddle-stitching whether a UTF-8 character or what kind of control character its sequence is at what state, that what gets derived for UTF-8 characters as secondary is either a nybble with the count-total and count-remaining, or, count-encountered and count-remaining. 1 2 3 4 When straddling, it's un-known whether there are remaining bytes, about basically to have a separate part of the nybble for the straddle straddling/ split/ stitching/ The idea is that the smear/unsmearing is indicated by the word, for the properties, then that for the code-point, that's inserted with the stitching, about that splitting is only at the end of a word, and stitching is only at the beginning of a word for forward search. So, first the main class is determined, then, conditioned on whether there exists either a "max-length" or a null character is the End-of-Input, and conditioned on whether there's a "Start-of-Input" offset, about offsets and extents, the main class is determined from the Start-of-Input (usually somewhere in the initial word) and End-of-Input, then making the lookup of the main class. Another point of straddle and splitting and stitching is for the fixed match case, while it's usually figured that the fixed string being matched fits within a word, arbitrarily it crosses multiple words or is more than word length, then that when there's an initial-segment match, to be matching the trailing-segment. So, in splitting UTF-8 codes, it's known that the code extends, yet not how far, yet in splitting fixed strings, it's known that the initial-segment matches, not if the trailing-segment matches. Then, matching the "fixed" also gets into matching more widely, about the expressions and grammars. From taking a look into outlines of Hyperscan and Vectorscan (regex and multiple-regex matching engines employing vector techniques from Intel and ARM respectively), there are notions of the "decomposition" of expressions, then about what's promontory and matching the "fixed", first fixed-length then fixed-content, when matching what would be "longest sub-matches", then to recursively bridge the definite sub-matches. So, the context of the findings and matchings start to develop, with the idea that by the presence in the context, that actions occur, otherwise for nothing or no-ops. Afore-Input: Start-of-Input, at the beginning of a "walk", and beginning of a "word" Afore-Stitch: at the beginning of a word, there's stitching to occur After-Split: at the end of a word, there's definitely/possibly a splot After-Input: End-of-Input, at the end of a "walk", and end of a "word". Here "walk" has the usual notions of "tree-traversals", that instead here "walk" (or "work") is the notion here of the sequence action, then for "work". Then "Afore" and "After", or "Before" and "Behind", make for that they're same-length identifiers and also that they're in the same lexicographic order. Before-Stitch Behind-Split Afore-Stitch After-Split Among-Straddle (Among, Amidst) So, the context then is for register state and stack contents, that the indicators of the above as "positive presence" then is to make for that the adjustments to the offsets and extents and the shifts is according to those, otherwise no-ops. Then the idea is that a "working" starts with a given context according to the expression, then that as various of the "findings" make findings, they push either context to act on the stack, or no-ops on the stack, then the stack results being a fixed-size for the working according to the expression, then the actions are always popping off a fixed amount of actions and no-ops, with no branching, just computed "presence". 1) work starts compute any misalignment / Start-of-Input load word (or bytes-into-word when no-misaligned-loads) 2) word starts (resolve startings) (resolve endings) (resolve stitches) lookup/load main class find coded find splits find UTF-8 find cntrl lookup expression/grammar classes find (resolve splits) (resolve straddles, byte-straddles, word-straddles) The idea is that the predicates (properties/predicates or code-points/range-points), are to get shifted and trimmed, or initialized, shifted, and trimmed, so that it results the trimmings or truncations, then have that the properties/predicates or code-points/range-points will result matches in what results of the initialized, shifted, and trimmed. 1) initialize (copy) the predicate/range-points 2) shift to find-start, find-continue 3) trim about the offset, extent 4) find-continue About code-points/range-points, what's figured is that it's always inclusive the bounds of the range, then that the matching of a single code-point is always the matching of two range-points that happen to be equal, so that matching either a code-point or a range, is the same operation, that: not-less-than-lower && not greater-than-upper which makes finding of range-points, also works for code-points. So, the usual idea is that there are the various findings occurring, find-longest-match: shift and repeat byte-wise across the word find-nearest-exit: find-near: find-far: Then, for an expression or expressions, and grammar or grammars, is the idea of making multi-matches, that the idea is that each of the possibles make their exercise, and then to result after the word is worked by each of the sub-expressions, to collate the results, or to emit the results, then onto the next word. Basically there is a difference among productions about whether matching or finding is among "alternatives" or "potentials", with the idea that matching "alternatives" is vertical while matching "potentials" is horizontal, that a finding in terms of the NFA/DFA basically enters either an "arc" or a "transition", that an "arc" is in the "potentials" to make a "plant" of the "potential plant", vis-a-vis the arcs/plants and transitions/states. Then, an alternative has matching the first character, then whether it introduces a potential, about that the single-character matches then as for "double-bracket" or "triple-quote", make for that those sorts of potentials are as according to the bracketed/quoted/escaped expressions/grammars, to be defining the rules of the machine. finding potentials then is about this sort of account: the word is N-many bytes wide property/predicate: 1 register property, 1 register predicate -> 1 register indicators codepoint/rangepoint: 1 register codepoint, 2 registers rangepoints -> 1 register indicators union of findings: 2 registers indicators, 1 register indicators intersection of findings: 2 registers indicators, 1 register indicators setminus: ... complement The finding then has either a "required" or "optional" next item, when it's in finding potentials, then across the N-many bytes, the count-down of the initialization/shift/trim begins, then to be running down the bytes making each match, while it continues "find-continue", or, regardless, then that the resulting indicators look for the first contiguous block of matches. Then the A/B/other or likely/less-likely/un-likely, is about making the findings, and automatically composing with making the next findings, or as that that's in matchings, to adjust the finding as it goes along, according to that in regular expressions it's a next match, then as with regards to when there's backtracking and greedy/lazy or among the greedy/possessive/... regular expressions. The composition and decomposition of the grammars and expressions, is to result that after EBNF and regex, the composition and decomposition, about how to orient the productions and sub-expressions, and their logic, toward that then alternatives and potentials are arranged their consequences. op: + | - | * | / | % expr: expr op expr ( <-> ) Here the idea is that the balancing of the parentheses and their relation to the precedence so indicated, is otherwise as according to left-to-right and right-to-left, about then what induces the potentials within the balanced parentheses to make expressions, about then the evalation order of the expressions so indicated, then as with regards to "concatenation", the most usual operation in strings, op: / expr: expr op expr that when a rule mentions itself it induces a potential, and that when it has branches that it induces alternatives. number-initial number: [non-zero-digit] number identifier-body: [identifier-body-char] identifier-body identifier: [identifier-initial] [identifier-body] keyword: "kw1" | "kw2" | "kw3" header: body: trailer: sequences "..." introduce sequences (concatenation) branches "|" introduce alternatives mentions "<-" introduce potentials options "[]" introduce options directionality-left "<" introduces left-balancing, pairing directionality-right ">" introduces right-balancing, pairing The directionality or balancing/pairing is indicated when the left-most and the right-most of the sequence so make it indicated, the left-most and right-most of a production of a grammar, or representation/representative of an expression. op: / expr: [(] expr op expr [)] Here the expression has the left-and-right paired, and that they're only optional mutually, i.e. both or neither, about a sub-class of optional that's "both-or-neither". Then, escapes introduce what is a smashing, since the idea of escapes is that they're symbol-escapes not syntax-escapes, vis-a-vis quoting, what itself is a syntax-escape, and comments, what is a syntax-escape, about the escapement, and balancing and pairing and nested escapes. So, about the bounds and the offsets, there are the windows (the coding regions) and the ledges (the ends of the straddles), then for what goes on the stack of actions, and what is to result making the stack of findings, is about the organization of offsets extents bounds (offset + extent or offset, offset) then about the window-bounds and the ledge-bounds, in terms of those being the word-bounds, and the bounds of the finding. union | intersection | complement | setminus Here complement is usually enough "not", or as with regards to the entire space of code-points, about where "not X " is both "universe setminus X" and "setminus X", about expressions with universes or "worlds of words". This is that usual accounts of language are constructively defined as after the alphabet, that here the alphabet is already "complete" in the sense of the range of code-points, about then to make for where classes get defined by ranges or indviduals the range-points, then in terms of "not" and "complement" and "setminus", about the logic of union and intersection. https://wyssmann.com/blog/2019/11/extended-backus-naur-form-ebnf/ https://datatracker.ietf.org/doc/html/rfc2234 (ABNF) ABNF in RFC2234 introduces ideas of incrementally-defined rules (3.3) when they are alternatives, here about "composable grammars" and the ideas of schemas of grammars. Here there's a fundamental difference between range-points and alternatives, since range-points are found by code-points while alternatives would each have their own findings. Both backtracking and balancing involve state, vis-a-vis, the "lookahead", the "lookback", and here with regards to "backstack", and "depthstack", or "pairstack". The idea of "pairstack" then is each of "backstack" and "depthstack", about that when crossing words, while still making a finding, is that the previous words get pushed on the backstack, then that for balancing pairs, get pushed on the depthstack, or for example both. A glossary develops: register g-register: a general-purpose register v-register: a vector register byte: an octet of bits, interpreted as unsigned integer or bit-flags nybble: half a byte word: the v-register word character-set: a collection of elements of a language character-encoding: content/layout/format of a character set character: a member of a character-set character-class: an attribute of a character or its bytes as properties or rangepoints input: a region in memory of contiguous character data, one or more register words bit-wise: operating according to index of bits byte-wise: operating according to index of bytes offset: extent: bounds: indicators: bit-values 1 yes 0 no properties: a byte of indicators of a categorical class predicates: selected interest bits to indicate predicates finding matching categorical classes code-points: the byte or bytes that comprise a character range-points: a lower and upper bound that defines a range of characters inclusive or individual character lookup-table: a 256-entry table containing properties for code-points lookup-line: a linear-lookup cache lookup-tree: a btree-lookup cache lookup-file: a backing file for unboundedly many entries expressions: components and sub-components of regular expressions representations: examples that match expressions grammars: rules of composition of expressions productions: examples that match grammar rules act: the execution of an instruction of instructions finding, findings: act, results of making indicators of properties/predicates or codepoints/rangepoints matching, matchings: act, results of finding making indicating representations, productions made-match mis-match working: making findings and matchings over the input wording: (not a word, working within a word) straddling: when multi-byte codes cross words splitting: working either side of a split of a straddling code stitching: mending both sides of a split of a straddling code smearing/unsmearing smashing/unsmashing backtracking balancing backstack depthstack pairstack afore-stitch: cases of straddle, a: start of buffer, before stitch before-split: cases of straddle, b: end of buffer, before split after-split: cases of straddle, a: start of buffer, after split behind-stitch: cases of straddle, b: end of buffer, after stitch Then, the idea of that it's as a sort of dance (with steps), or the "rhythm of work" is about the presence of cases that maintain the context: work-context word-context then about the initialization shifting/rotating trimming after the work-offsets word-offsets then emitting and maintaining bounds of representatives/productions of the expressions/grammars. Then the idea is that for a given offset, the predicates/rangepoints get popped off the stack, the default algorithm for predicates and the default algorithm for rangepoints get invoked, or rather, that a structure makes for defining "relative registers" and having both the kinds on the same stack, then for example where when there's potential that the passing predicate gets pushed back on the stack, or for example that there's made round-robin of all the possible alternatives on the stack. Then, making a match results resetting the stack, for example from the contents of the stack, when making multiple match. So, in the context, there are predicates and rangepoints, these are of various sorts. 1) a predicate/range-point is just a duplicated next-char to be spread and then making finding, the entire word 2) a predicate/range-point is a fixed-length with an extent, to be making finding Among the sorts are various cases about whether there's matching-many (repetitions) or matching-multiple (alternatives), then for example match-1-alternative or match-all-alternatives (multi-matching). Then, next to the predicate/rangepoint or the definition that results what it is, is about what matches it makes according to its findings, the matches then being events in the representatives/productions. Prime Rings and Prime Multisets As an aside about an example arithmetization, there's the idea that multisets can be embodied in an integer as primes, with a catalog of prime numbers to members, then another idea is about prime rings, finite rings of prime modulus. The idea is that a given width unsigned integer can maintain the state of a number of prime rings. For example, Z_5 the prime ring with five elements, can be represented with 2s, and then the multiplicity of 2's in the factorization of a number, is the modulus of the prime ring 0-4. 2^5 = 32 Then, for example with pairs 2, 7 and 3, 5, then an integer with range >= 7^2 * 5^3 * 3^5 * 2^7 can maintain within it four prime rings, Z_2 Z_3 Z_5 Z_7 respectively. Then computing the modulus (or value in the ring 0 to n-1) is a matter of determining the multiplicity of the given corresponding factor, while incrementing the ring is a matter of checking whether b^n-1 is a factor, and dividing that out to make zero in the ring, else multiplying in b, to result incrementing in the ring Z_n. It would be usual enough to instead make for that simply bits and multiples of bits embody rings, then with just using increment and modulo on them, then that to store these rings would take 1-bit for 2, 2-bits for 3, 3-bits for 5 and 7, and so on. Then, where that might make sense, is when for example a state transition affects multiple prime rings, that it's a matter of multiplying in their product to increment both rings, vis-a-vis setting the relevant bits and adding them in, then with regards to overflow, either in the adders as among the bit-packed prime-rings, or in the multipliers among the prime-backed prime-rings. Prime rings are useful since when incrementing them each apiece, they are not zero except when they have common factors of the counts of increments. Finders their Ways So, the finders are basically working across, or down, across in sequences, and down in alternatives. Then, there's also that finding is either anchored as prefix-matching, or drifting as substring-matching. anchored: prefix-matching (from current offset) drifting: substring-matching (across offsets) sequence matching: fixed or likelies alternative matching: among alternatives Then, the idea is that the stack of work is the source of the finders and the matchers, where the finders are the literals that work in the standard machines, while the matchers coordinate reaching through arcs to plants, or transitions to states, that result representatives or productions, then what to do with those. The standard algorithms are of these kinds: properties/predicates: AND the bits to result set bits meaning property = predicate CMP-to-zero the bits to zero to result 0xFF bytes when all bits are clear, else 0x00 NOT the bits to result 0xFF when all bits are set PMOVMSKB the bytes to bits from v-reg to g-reg BSF the bits to find byte-offsets where property satisfies at least one predicate codepoints/rangepoints CMP-for-gte the lower bound CMP-for-lte the upper bound AND the comparisons meaning codepoint between rangepoints NOT the bits to result 0xFF when all bits are set PMOVMSKB the bytes to bits from v-reg to g-reg BSF the bits to find byte-offsets where codepoints between rangepoints fixed-string sub-string XOR the bits to result clear bits meaning codepoints match CMP-to-zero the bits to zero to result 0xFF bytes when all bits are clear, else 0x00 PMOVMSKB the bytes to bits from v-reg to g-reg BSF the bits to find byte-offsets where fixed-string equals substring The predicates make unions, eg, to match either alnum or punct, about the union of character classes. Then, the standard algorithm must involve the union, intersection, and complement/setminus, about expressions their usual composition. The idea is that these form a recursive sort of account, according to implicit and explicit precedence, that result invoking the standard algorithms above, to result the bytes to bits from v-reg to g-reg. These are figured to generally be "yes/no/maybe's" or "sure/yes/no's", about making for the the union and intersection of the thing otherwise, that are pretty simple for predicates A and B. union A, B = A || B intersection A, B = A && B setminus A \ B = A && !B So, with regards to the character-set and character-encoding, it's figured that by default it's Unicode with UTF-8, and that source texts are overwhelmingly printable ASCII, then that there are also very usual files that are either UCS2 or UTF-16, or UTF-32. Then, before the "work" function is along the lines of "detect/inspect", that otherwise the character-set and character-encoding are assumed invariants, then that there's as with regards to Internet messages their declared character-set and character-encoding, and the accounts of comments and escapes from localedef. Then, the usual account of each word is mostly clarified, then to get into the specific semantics of multi-byte characters (characters generally as both printable and non-printable "characters" then as with regards to "ligatures" generally and "escapes" generally. The actions on multi-byte characters mostly are as with regards to figuring their sparse (or, not completely dense) offsets their first byte, that first there is the main class its properties, then to be making the UTF-8 code-points into runs of bytes their characters. So, the main-class or ascii-class properties are loaded first, instead of first having a utf-8/non-utf-8 class, since, the distribution of the content is overwhelmingly printable ASCII (and common control whitespace). Then, the detection of the coded/ items that are UTF-8 encoding items follows, with "spotting", and then about the data structures that indicate the offsets and extents of UTF-8 encoded characters, to then implement the "smearing", and about escape characters that result literals, when those are "smashing". spotting: identifying offsets and extents of UTF-8 characters, thusly the sparseness/spotting of offsets of characters in the bytes smearing: extending the sections of predicates according to spotting Then, for rangepoints gets involved an example, that the ranges are to be encoded correspondingly into ranges of the UTF-8 encoded characters. It's figured that contiguous ranges of UTF-8 characters have contiguous ranges of their encoded bytes. https://en.wikipedia.org/wiki/Regular_expression https://en.wikipedia.org/wiki/Parsing_expression_grammar https://en.wikipedia.org/wiki/Raku_rules https://en.wikipedia.org/wiki/Recursive_descent_parser https://en.wikipedia.org/wiki/Thompson%27s_construction Looking at Thompson's and Glushkov's construction for making NFA's from expressions, then as with regards to the notion of minimization after the outer-product or powerset making a DFA, here is for making what actions are possible, to identify the arcs and plants, in terms of making of those transitions and states, about establishing the mutual interpretability of the models of actions in prefix-matching as usual NFA's/DFA's give, with regards to prefix- and substring- matching. It's figured that regular language have forward recognizers, then as with regards to backtracking and balancing, about where the recognizer has those, that then gets into limits. Here the idea of the predictive parser is basically for something like where Thompson's constructive is said to guarantee that at most two arcs exit a state, then the idea is that the predicates can be so combinatorially enumerated, or as what so describes the matchers, to make consecutive or plural matches in one "operation", for plural-matches, vis-a-vis multi-matches which is the idea of having multiple expressions of grammars, about making plural-predictive predicates and rangepoints, off of usual constructions of NFA's, that certain predictions are simpler than others. Plural Cases literals: prefix or postfix (suffix) A usual idea for matching literals is as about the initial-segment and trailing segment, or, leading segment and final-segment, where the initial-segment or final-segment is a fixed-string, while the trailing-segment or leading-segment is variable length, of a given class, or equivalently, when the class has range-points. I.e., besides the notion of combining properties/predicates and code-points/range-points, is to have the fixed-string be the initial-segment or final-segment, and then the trailing-segment or leading-segment is a different range in the predicate word, then that the standard algorithm finds matches for literals (numeric literals). It's not dissimilar for string literals, about necessarily enough the escapement, and then also for finding forward and finding reverse, in the word, and then checking for gaps, retracting until checking for empty strings, for string or character literals. Then the idea is that any of those can be found and matched in one "run", i.e. a stall-less, branch-less, call-less list of less than a few or less than a few dozens or less than a few hundreds instructions, that runs in less than one microsecond.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-07-30 14:55 -0700 |
| Message-ID | <114gh8p$25jib$3@kst.eternal-september.org> |
| In reply to | #143121 |
Ross Finlayson <ross.a.finlayson@gmail.com> writes:
[48 lines deleted]
> RF, good to join the panel. I appreciate the format—direct address and
> genuine exchange rather than parallel monologues.
[4368 lines deleted]
Ross, this is not a "panel". This is a thread cross-posted to
three newsgroups, comp.theory, comp.lang,c, and comp.lang.c++.
You've just posted more than 4000 lines of text that, as far as I
can tell, have nothing to do with the C or C++ programming languages.
Maybe the discussion is appropriate to comp.theory, which is a
cesspool these days, but in comp.lang.c and comp.lang.c++ we would
very much like to discuss the programming languages that are the
topic of the respective newsgroups without being bombarded with
arrogantly off-topic posts.
I won't try to reason with Johann 'Myrkraverk' Oskarsson, who
seems to enjoy posting to irrelevant newsgroups for some reason,
but perhaps you can do something. If you're not talking about the
C or C++ programming language, please don't post to comp.lang.c or
comp.lang.c++ -- even if you're posting a followup to a post that
was cross-posted to those groups. (You'll have to manually edit the
"Newsgroups:" header line.)
I've redirected followups for this post to comp.theory.
Thank you.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-07-30 21:18 -0700 |
| Message-ID | <DcScnbfcv8ynv_H3nZ2dnZfqnPWdnZ2d@giganews.com> |
| In reply to | #143136 |
On 07/30/2026 02:55 PM, Keith Thompson wrote: > Ross Finlayson <ross.a.finlayson@gmail.com> writes: > [48 lines deleted] >> RF, good to join the panel. I appreciate the format—direct address and >> genuine exchange rather than parallel monologues. > [4368 lines deleted] > > Ross, this is not a "panel". This is a thread cross-posted to > three newsgroups, comp.theory, comp.lang,c, and comp.lang.c++. > > You've just posted more than 4000 lines of text that, as far as I > can tell, have nothing to do with the C or C++ programming languages. > > Maybe the discussion is appropriate to comp.theory, which is a > cesspool these days, but in comp.lang.c and comp.lang.c++ we would > very much like to discuss the programming languages that are the > topic of the respective newsgroups without being bombarded with > arrogantly off-topic posts. > > I won't try to reason with Johann 'Myrkraverk' Oskarsson, who > seems to enjoy posting to irrelevant newsgroups for some reason, > but perhaps you can do something. If you're not talking about the > C or C++ programming language, please don't post to comp.lang.c or > comp.lang.c++ -- even if you're posting a followup to a post that > was cross-posted to those groups. (You'll have to manually edit the > "Newsgroups:" header line.) > > I've redirected followups for this post to comp.theory. > > Thank you. > Thanks for writing. Sure, I'll limit this.
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-07-31 09:12 -0700 |
| Message-ID | <DlOdncwOxIMfVPH3nZ2dnZfqnPidnZ2d@giganews.com> |
| In reply to | #143138 |
On 07/30/2026 09:18 PM, Ross Finlayson wrote: > On 07/30/2026 02:55 PM, Keith Thompson wrote: >> Ross Finlayson <ross.a.finlayson@gmail.com> writes: >> [48 lines deleted] >>> RF, good to join the panel. I appreciate the format—direct address and >>> genuine exchange rather than parallel monologues. >> [4368 lines deleted] >> >> Ross, this is not a "panel". This is a thread cross-posted to >> three newsgroups, comp.theory, comp.lang,c, and comp.lang.c++. >> >> You've just posted more than 4000 lines of text that, as far as I >> can tell, have nothing to do with the C or C++ programming languages. >> >> Maybe the discussion is appropriate to comp.theory, which is a >> cesspool these days, but in comp.lang.c and comp.lang.c++ we would >> very much like to discuss the programming languages that are the >> topic of the respective newsgroups without being bombarded with >> arrogantly off-topic posts. >> >> I won't try to reason with Johann 'Myrkraverk' Oskarsson, who >> seems to enjoy posting to irrelevant newsgroups for some reason, >> but perhaps you can do something. If you're not talking about the >> C or C++ programming language, please don't post to comp.lang.c or >> comp.lang.c++ -- even if you're posting a followup to a post that >> was cross-posted to those groups. (You'll have to manually edit the >> "Newsgroups:" header line.) >> >> I've redirected followups for this post to comp.theory. >> >> Thank you. >> > > Thanks for writing. Sure, I'll limit this. > > Yeah, I've been looking at this, and here's what it seems is the profile, of the resources, about the vector units, on Intel/AMD and ARM. So, first there's that MMX since Pentium is still alive, yet, it's considered sort of aside what are the general purpose registers, if for a sort of "general-auxiliary" use, about the "16 general purpose registers". Then ARM mostly has "32 general purpose registers", with the idea that Intel has 16 (or less) general purpose + registers + 8 old floating-point/MMX SIMD vectors. So, there are basically 16 general purpose registers on each, and 2 of those on ARM. Then, the vector registers basically make for "SSE 4.2" or here for what's SSE3 yet beyond SSE2, about there being vector registers now essentially separate from general registers. So, here the goal is to use the vector registers like large scalars, or at least as arrays of bytes. Well, that's not exactly the goal of the vector/packed/SIMD registers. So, there's a common subset of functionality, and limits within the vector registers, about what can be treated as scalars (with the byte as least-addressable, shift & rotate, and with the logical operations and compare that go straight up and down, in terms of two vector registers their lanes their words their bytes their bits). Basically then there's "double quad-word" or 128 bits, in both the Intel/AMD and ARM, that's about the biggest "scalar" word there, as the data type, for the common subset of instructions abstractly they support. Then, the SSE4.2, has 128-bit vector-registers, that can be operated upon with their DQ for double-quadword variants of instructions, alike scalars, or at least for the byte-wise, if not necessarily the bit-wise, with regards to shift & rotate even multiples of 8 bits. Then AVX with 256-bits, is two of those side-by-side, similarly AVX-512 then, is two of those side-by-side, and ARM SVE, is one or more of those side-by-side, 128-bit double quad-words with "byte-wise" moves like shift & rotate, with regards to using "extract" on ARM to simulate shift & rotate multiples of 8-bits. So, this sort of tiling of the register files, thinking of the registers the memories as a rectangular block of bits, about the register transfer logic moving the bits or computing the bits, basically gives 128-bit 16-long blocks, that can be treated like "byte-addressable scalars". SSE4.2: 1 block (16-many x 128-wide) ARM NEON: 2 blocks (32-many x 128-wide) AVX: 2 blocks (16-many x 256-wide) AVX2: 4 blocks (32-many x 256-wide) AVX-512: 8 blocks (32-many x 512-wide) ARM SVE: 2-20 blocks (32-many x 128-2048-wide) where all the widths are essentially separate units run together in lock-step of "double quad-word type size" byte-addressable "scalars". So, algorithms should be designed to work in 1 block, in the register file, and then scale in these blocks, for vector-wide scalar-word operations (byte-wise). Here then the idea is that the "character machine" basically implements a little scheduler and then making the various findings and matchings in the blocks.
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-07-31 12:55 -0700 |
| Message-ID | <qx6dnbtn6ZAiYPH3nZ2dnZfqn_udnZ2d@giganews.com> |
| In reply to | #143121 |
On 07/30/2026 07:05 AM, Ross Finlayson wrote: > On 07/30/2026 06:49 AM, Ross Finlayson wrote: >> On 07/27/2026 11:45 AM, Ross Finlayson wrote: >>> On 07/27/2026 11:44 AM, Ross Finlayson wrote: >>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote: >>>>> Hello, here I'll post some design notes and a panel discussion with >>>>> some >>>>> chat-bots about making some sense of the "vector-wide scalar word" >>>>> and "character machines", on commodity hardware about ubiquitous >>>>> operations. >>>>> >>>>> >>>>> It's considered at least tangentially relevant to comp.lang.c and >>>>> comp.lang.c++ because for example text is ubiquitous and the targets >>>>> would be low-level, while the higher-level languages would have a >>>>> same sort of patternry, and for example that libc and cstdlib are >>>>> standard, and as with regards to POSIX and Unicode and so on. >>>>> >>>>> Please feel free to excuse or ignore, or comment as freely. >>>>> >>>>> Thanks for reading. >>>>> >>>> >>>> >>>> [ viswath-charmaigne.txt ] >>>> >>>> >>>> >> >> > [ viswath-charmaigne-20260730.txt ] Drift-Find About the finding, then for matching, the idea of "drift-find" is as distinct "anchored-find", about that drift-find is about iterating over offsets and finding matches, without testing each match as anchored-test-match. So, the standard algorithms match byte-wise according to properties/predicates (that at least one predicate matches at least one property) and codepoints/rangepoints (that the byte is within the range, inclusive, of the pair of rangepoints). Then, when matching word-wise, and drifting the input pattern over the input data word, then it's ambiguous simply OR'ing together the standard algorithm SA results. AB pattern AAB data <- ambiguous whether found at offset 0 or 1, or both ABA pattern ABABA data <- ambiguous whether found at offset 0, 1, 2 Then, the idea is to implement an account of the "drift-palindromic" or "keyway comb", that instead of the SA making 0xFF on finding and 0x00 on not finding, that the drifting accumulate with a sparseness matching from the front, and sparseness matching from the back, and that the combined run must have a length matching the pattern length, then that it's an unambiguous match, the result of the finding. forward -> 1011011101111 ... k-many bits for pattern of length k reverse -> 0100100010000 ... k-many bits for pattern of length k Then the idea is that in the drift, as for drift-slip and drift-slide, that the result of the standard algorithm is converted to each of the forward and reverse, those being put on the stack or otherwise collected, then that only when their union is all 1-bits, is it un-ambiguously alike 0xFF. The idea is that the combs are generated, then about whether they confirm the match, or, cancel the match, or about that the findings fiddle the combs, so that only the first byte of matches get indicated as found, and each of the first bytes, as drift is to find all offsets where the pattern matches. Then, the idea of progressive combs breaks the SBC-less, with the idea of calculating all the forward and reverse combs, and to give combs at different offset different progressions of density/sparsity of bits, then to result that only matching combs result all set bits and only where they match. So, then it is BC-less, yet stalls are introduced when storing on the stack the combs each, then that they are worked together what result that only the full matches are found, that S < B < C the cost. Then, the idea might be to first make the naive match, and then make the cancel match, that the arithmetic would work out making no-ops on the matches, and cancels on the mis-matches, since the arithmetic would be indicated by an already ambiguous match, else no arithmetic. So, the idea is to store off pairs of combs for each byte offset, or 2W-many, then to go through the combs and any mis-match results cancelling at that offset. Here that might be alike "optimistic drift", where comb mis-matches are only to make cancels, else matches: canceling the first byte of the match. So, the idea is developing to a) make the ambiguous naive match, then b) make the cancel match, off the first bytes of those. So, the idea is to drift forward, and union together all the findings, then drift backward, and zero the first byte if it's not a match. Then, the drifting case is perhaps much simpler than the drift-palindromic or the comb-fiddling, with the idea that drift-forward makes all matched bytes their characters, then drift-revert invalidates the first _character_ of matches on the way back, then that it results that any matches have their original length, yet, that would possibly invalidate trailing characters of an earlier match, thus getting back into the idea of the drift-palindromic and comb-fiddling. Then, the idea might be to make for canceling the first byte of mismatches, that might be a last byte of an earlier match, about: going back and forth setting the first byte, setting the second byte, and so on, or as with regards to whether the output of the algorithm is as sequence of offsets of first bytes instead of otherwise the SA offset-indicator bit-string. Since the patterns might overlap, then the offset-indicator bit-string itself is ambiguous, about whether to return the first finding, or plurally all the offsets where findings occur. Then, the idea would be to result an offsets tuple, where the offsets range from 0 to W-1, eg 8, 16, 32, 64 for 64, 128, 256, 512 registers, then that those each fit in a byte, for a word of offsets, where the maximum offset thus difference in offsets is W-1, and the maximum count of offsets is W. Then this could be converted to the offset-indicator bit-string, of starts of matches, instead of saturation of matches. char-wise indicator string: bits are set fixed-wise indicator string: starts are set Then, it seems for only marking the first matching character on the match, yet, for the initial/final trailing/leading, then it's wanted to make the bit-string with the plural matches. "Parallel String Matching Philip Pfaffe, Martin Tillmann, Sarah Lutteropp, Bernhard Scheirle, and Kevin Zerr" One idea then is to make counters, and only bytes with counters being the length of the fixed-pattern, are included, about matching any byte in the pattern to any aligned byte in the input, and counting those up what would be the combinations of all the substrings, that all the combinations of the substrings match. Still, not knocking out the first character won't eliminate the starts, yet not each character is a start. Then, the idea of "count of matches", may simply enough make for that differences from 0 indicate overlapping. This then is to drift along, and find the 0xFF matching, increment a counter for that offset, and then when going along, that each increment is a start, and each decrement is an end, then though at multiples of K, is also an end and a start, if no differences. Then, only for fixed-patterns, it seems the idea is to find the starts by checking each offset in the drift, and what results matching, up to that length, gets incremented, or also, that it can just be any positive difference indicates a start, so the pattern can be repeated, then drifted across, and the starts will have increases, and the non-starts won't. "M. O. Külekci: Filter Based Fast Matching of Long Patterns by Using SIMD Instructions" https://www.stringology.org/ "Handbook of Exact String-Matching Algorithms" http://www-igm.univ-mlv.fr/~lecroq/string/ Then, for making drift-diff, is that the pattern can simply be made repeated in the pattern, and it only needs to drift offsets K-1 many, then the counts will have been accumulated, for the diffs to be computed. W/K About building the repeated pattern, there is broadcast or the like, ABC . 012012012012 ... ABCABCABC ... then, the idea of not having a loop, or un-rolling the loop, is basically about that there is binary subdivision, to not explode the number of statement blocks, into block-with-nops, and also to have the shorter statement blocks for the shorter patterns. So, using the standard algorithms SA for matching, then the predicate/rangepoints of the fixed pattern (a fixed-length predicate or fixed-length string or rangepoints), has that drift invokes the standard algorithm, only to compute the counts, then separating the SA the predication, from moving off the result, that the counts are to be collected, then made their diffs. ceil log_2 K -> count drift-shifts Then, for example where K = 1, log_2 1 = 0, the repeated shift makes the match at once. Then, there still needs be checking either "diff" or "even modulo" from the previous match, its count. So, for the fixed pattern alone, then, for the cost of making it repeated in the pattern, then for shifting it K-1 many times, and accumulating the matches, is for having W many entry-points, then the rotation simply occurs K-1 times in the block, un-rolled. Then there's the problem of a) straddling when the pattern straddles the word at B, and b) when the pattern straddles multiple words. The idea is that the prefixes start, and then the remaining pattern gets multi-drifted, which would require enough depth of those rotations, to cover the length of the pattern, or a word, pulling forward the pattern, then also, the pattern, will need to be stored in its entirety or as to that it's loaded from memory in however many words it may straddle. For example, for pattern ABCD, when the input ends AB, then there's an anchored match of CD, then to follow with starting over drifting, where K < W. For the pattern AAAA, when the input ends AAA, then each of A, AA, AAA need anchored matches, or drifting with that "the initial segment pattern is found", ..., about how to treat SHIFT and ROTATE so that basically it can make for the repeated pattern, to start rotated left each of the offsets, about making counts of those. Point being, the findings of the straddlings won't complete until as many words have passed as K fits, or the last word, and, the partial matches from the previous word, carry-in and are to accumulate, that their offsets are in the previous word. About the instruction cache, it makes sense to just have one block, and then just make it so that the arithmetic just results nops, .... SA: star standard algorithms for matching patterns, anchored SA: fixed standard algorithms for anchored/drift fixed strings About the binary indicator-strings, is that 8 words worth of those can fit into a vector register, about GW, the general purpose word, and Gw, in bits, about that there are 64-bits about which to run BSF/FFS on and make to emit offsets. Ideas about signature of reported findings/matches include: 1) a context struct, and functions to return count, to compute the size of the return buffer, then functions to populate the buffer with the offsets, and about character and byte offsets. 2) a fixed-size output buffer, the function accepts the size and the buffer and returns the count of elements in it, which are offsets, returning -1 at EOF (EOI) 3) a fixed-size output buffer, less than pattern/expression max, making capture groups 4) a callback function, called with offset 5) one pass to compute bounds, one pass to fill bounds Example: Deflate algorithm, compression/decompression Compression involves a 32 kiB window, where back-references would be, then the idea that in a block of up to size 64kiB, then the heavy computation is the longest-duplicate detection or "the finding of Huffman codes", as with regards to finding the most and longest duplicates that get the shortest codes, about finding the duplicate, or for long runs or the highly compressible, breaking those down into moduli. So, the idea would be to make it drifting over itself, that the patterns naturally enough start from the front, that there are 32kiB / WB words in the window, eg 2^`5 / 2^7 = 2^8, for 128 bits, 256 words, or that larger vectors would make for larger windows, with just fixing the ratio, then that from the front gets into matching the characters, that each word (8 = 64b, 16 = 128b , 32 = 256b, 64 = 512b, ... bytes) should make its own Huffman codes, then to combine those, making candidates according to those locales, then to make the account for "long" runs by a histogram of modes, and "common" runs as of the combinations of the modes, .... Then, about building histograms, the idea is to make the pattern the rangepoints of itself, i.e., just duplicating the input data, and using that as the pattern, then drifting that along making counts, across the word, then each byte will have how many times it was matched, then to take the max of those, building the histogram from the highest to lowest multplicities (cardinals of the multisets). Then, there's whether those are regular separators, or parts of regular substrings, then about high/low cardinality with regards to principals, modes, majors, minors, and the long tail, then about the ordering-statistics, to build out histograms to make counting arguments about what those are. Looking a bit into the object file organization (PECOFF, ELF) it seems that there are the sections as map to segments with regards to the CALL instructions, about the idea then that the calls will be with regards to the segments, about how big the segments can be, and then about the range of offsets so indicated, or about that many segments, each about PAGE_SIZE size, are indicated, about the locals. nybble 1: alnum punct white coded nybble 2: alnum: alpha digit punct: inner outer joiner affix white: nl space horz vert coded: ctrl utf8 nul nybble 3: alnum/alpha: upper lower alnum/digit: zero whole white/horz: space tab white/vert: nl cr ff vt coded/ctrl: single prefix left right coded/utf8: punct/inner: arith bool cmp res punct/outer: quote paren bracket brace punct/joiner: separator delimiter segment punct/affix: unary ref kleene lang nybble 4: punct/inner/arith: plus minus times slash punct/inner/res: modulo leftshift rightshift punct/inner/bool: and or xor punct/inner/cmp: eq lt gt punct/affix/unary: bang tilde minus punct/affix/ref: dollar asterisk ampersand dot punct/affix/kleene: plus star punct/affix/lang: period question exclamation punct/outer/quote: single double backtick punct/outer/paren: paren-left paren-right brace-left brace-right punct/outer/brack: angle-left angle-right square-left square-right punct/joiner/separator: comma semicolon punct/joiner/delimiter: comma pipe tab punct/joiner/connector: underscore colon slash backslash Here the idea is that breaking out punctuation is about that the usages are overloaded, so that the properties have that the characters have multiple properties, so that then according to the context, then as by the properties are found matched the predicates. Then, the organization is a curated sort of emphasis for common source files their usual syntax, or the common. Then, the idea is that grammars can provide their own property tables, then that here the first byte is always included, to make for UTF-8 and NUL and control characters, and the second byte is "source text" and also "data text". Then, the standard algorithm will be matching one or more bytes, here usually two bytes, that the indicated terminals as they usually are in expressions and grammars, get matched, that they match the mask of the first byte and the second byte. Then, the grammar-provided properties would usually indicate escapes, comments, and additions to the above, and accounts of characters that introduce ambiguity, to be disambiguated. As well, the main tables could be over-ridden, about specific differences from "C-style" languages. https://justine.lol/lex/ So, syntax has the "main" and "source" and then expression/grammar driven. About the logic, there gets involved how to make composable what result the "anchored" or "atomic" (sub-)expressions and terminals. The properties/predicates and codepoints/rangepoints can be combined, where leaving 0's matches none. The matching of the properties/predicates should be inclusive or exclusive, "match all" or "match any", here it's default "match any" (so predicated). The compositions of "yes/no/maybe" and "union/intersect/setminus" are to get figured out, how combinations of predicates are to be combined, basically as of the composition of classes, besides AND, IOR, XOR, NOR. The, the element of compositions is to result the character classes, then as with regards to the character classes having both the predicates/properties and codepoints/rangepoints, the main or default ones, and then union/intersection/setminus of those, and about complement classes. https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Regular_expressions/Character_class https://tc39.es/ecma262/multipage/text-processing.html#table-nonbinary-unicode-properties https://unicode.org/reports/tr18/#General_Category_Property https://unicode.org/reports/tr18/#Compatibility_Properties The Unicode TR18 for regular expressions is very useful and could be considered normative. https://unicode.org/reports/tr18/#Resolving_Character_Ranges_with_Strings About shift/rotate on the vector registers, it seems that there's a problem since there's a limit of 16 bytes for 128 bits (SSE2) , for packed-shift-right-logical-double-qword, PSRLDQ, the xmm register, that there isn't a byte-wise shift, for ymm/zmm registers, as they get split into lanes, ..., and shifting both the double-quadwords would make a void in the middle. Then, the ymm/zmm would have to be treated as separate units, for example piling in the instructions on both sides using the same offsets and computing for alternatives and so on. https://www.felixcloutier.com/x86/ https://mischasan.wordpress.com/2011/04/04/what-is-sse-good-for-2-bit-vector-operations/ https://www.scs.stanford.edu/~zyedidia/arm64/sveindex.html It looks similar with ARM. Then the idea would be to work up to double-quadwords or 128-bits the 16-bytes, as with regards then to making the acts being round-robin'ed to each of the packed double-quadwords, then about updating the anchors the offsets in lock-step. Then it's figured that the acts on the machines, that output the bit-string indicators of the byte offsets about smearing/unsmearing and byte and character offsets, would have a tag of what was found and matched in terms of the expression/grammar, that resulted the indicators, then that it's serialized what makes the matches/productions. Then for ARM NEON it looks like there's no double-quadword shift (128-bits) only each of the packed dwords (32-bit), "SIMD" on NEON. There is a REV64 instruction on ARM as might be about BSWAP, then with the idea though that shift byte-wise is the idea, and NEON instructions are "on each double-word", 32-bits. Then it might make sense just to divide-and-conquer, yet the lock-step item gets involved with having a common view of the input data and a given offset as the current sort of state-of-the-machine. "VEXT can be used to implement a moving window on data from two vectors, useful in FIR filters. For permutation, it can also be used to simulate a byte-wise rotate operation, when using the same vector for both input operands." -- https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/coding-for-neon---part-5-rearranging-vectors So, that then can effect "vector byte-wise right shift", basically loading from the end of the zero vector and the beginning of the vector to be shifted. It's considered a MOV so it leads to stalls. Then in SVE there's EXTQ, which is also organized about 128-bit double-QWORDS. About smearing and byte/character offsets then, those would mostly go to the vector registers as a bit-sequence indicator will indicate starts of characters in the byte-sequence. Yeah, I've been looking at this, and here's what it seems is the profile, of the resources, about the vector units, on Intel/AMD and ARM. So, first there's that MMX since Pentium is still alive, yet, it's considered sort of aside what are the general purpose registers, if for a sort of "general-auxiliary" use, about the "16 general purpose registers". Then ARM mostly has "32 general purpose registers", with the idea that Intel has 16 (or less) general purpose + registers + 8 old floating-point/MMX SIMD vectors. So, there are basically 16 general purpose registers on each, and 2 of those on ARM. Then, the vector registers basically make for "SSE 4.2" or here for what's SSE3 yet beyond SSE2, about there being vector registers now essentially separate from general registers. So, here the goal is to use the vector registers like large scalars, or at least as arrays of bytes. Well, that's not exactly the goal of the vector/packed/SIMD registers. So, there's a common subset of functionality, and limits within the vector registers, about what can be treated as scalars (with the byte as least-addressable, shift & rotate, and with the logical operations and compare that go straight up and down, in terms of two vector registers their lanes their words their bytes their bits). Basically then there's "double quad-word" or 128 bits, in both the Intel/AMD and ARM, that's about the biggest "scalar" word there, as the data type, for the common subset of instructions abstractly they support. Then, the SSE4.2, has 128-bit vector-registers, that can be operated upon with their DQ for double-quadword variants of instructions, alike scalars, or at least for the byte-wise, if not necessarily the bit-wise, with regards to shift & rotate even multiples of 8 bits. Then AVX with 256-bits, is two of those side-by-side, similarly AVX-512 then, is two of those side-by-side, and ARM SVE, is one or more of those side-by-side, 128-bit double quad-words with "byte-wise" moves like shift & rotate, with regards to using "extract" on ARM to simulate shift & rotate multiples of 8-bits. So, this sort of tiling of the register files, thinking of the registers the memories as a rectangular block of bits, about the register transfer logic moving the bits or computing the bits, basically gives 128-bit 16-long blocks, that can be treated like "byte-addressable scalars". SSE4.2: 1 block (16-many x 128-wide) ARM NEON: 2 blocks (32-many x 128-wide) AVX: 2 blocks (16-many x 256-wide) AVX2: 4 blocks (32-many x 256-wide) AVX-512: 8 blocks (32-many x 512-wide) ARM SVE: 2-20 blocks (32-many x 128-2048-wide) where all the widths are essentially separate units run together in lock-step of "double quad-word type size" byte-addressable "scalars". So, algorithms should be designed to work in 1 block, in the register file, and then scale in these blocks, for vector-wide scalar-word operations (byte-wise). Here then the idea is that the "character machine" basically implements a little scheduler and then making the various findings and matchings in the blocks. Then, figuring for making a "scheduler" is after a "plan", figuring that the expressions and grammars have their events of representatives and productions, then as with regards to the operation of "matchings" and "parsings", in the machine, then as with regards to the static machine, "the engine". So, overall, the functional units of the machine and engine are 128b = 16B wide, and 16-registers deep, then as with regards to the notion of scheduling the units as with regards to various and evolving "standard algorithms" SA, and then a model of the 64b = 8B wide, and 8-registers deep, for fallback to core 64-bit general purpose their auxiliary registers, or as for reference and fallback implementations in higher-level languages.
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-07-31 13:05 -0700 |
| Message-ID | <5--dnVBtWaqEnfD3nZ2dnZfqn_dg4p2d@giganews.com> |
| In reply to | #143162 |
On 07/31/2026 12:55 PM, Ross Finlayson wrote: > On 07/30/2026 07:05 AM, Ross Finlayson wrote: >> On 07/30/2026 06:49 AM, Ross Finlayson wrote: >>> On 07/27/2026 11:45 AM, Ross Finlayson wrote: >>>> On 07/27/2026 11:44 AM, Ross Finlayson wrote: >>>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote: >>>>>> Hello, here I'll post some design notes and a panel discussion with >>>>>> some >>>>>> chat-bots about making some sense of the "vector-wide scalar word" >>>>>> and "character machines", on commodity hardware about ubiquitous >>>>>> operations. >>>>>> >>>>>> >>>>>> It's considered at least tangentially relevant to comp.lang.c and >>>>>> comp.lang.c++ because for example text is ubiquitous and the targets >>>>>> would be low-level, while the higher-level languages would have a >>>>>> same sort of patternry, and for example that libc and cstdlib are >>>>>> standard, and as with regards to POSIX and Unicode and so on. >>>>>> >>>>>> Please feel free to excuse or ignore, or comment as freely. >>>>>> >>>>>> Thanks for reading. >>>>>> >>>>> >>>>> >>>>> [ viswath-charmaigne.txt ] >>>>> >>>>> >>>>> >>> >>> >> > [ Excuse, replied to an earlier post before dropping comp.lang.c, comp.lang.c++, please ignore, as follow-ups are to comp.theory. It's appreciated the tolerance or absence thereof. -- ]
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-08-02 10:27 -0700 |
| Message-ID | <Cjqdncy-TPx54PL3nZ2dnZfqn_WdnZ2d@giganews.com> |
| In reply to | #143162 |
On 07/31/2026 05:03 PM, Ross Finlayson wrote:
> On 07/31/2026 12:55 PM, Ross Finlayson wrote:
>> On 07/30/2026 07:05 AM, Ross Finlayson wrote:
>>> On 07/30/2026 06:49 AM, Ross Finlayson wrote:
>>>> On 07/27/2026 11:45 AM, Ross Finlayson wrote:
>>>>> On 07/27/2026 11:44 AM, Ross Finlayson wrote:
>>>>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote:
>>>>>>> Hello, here I'll post some design notes and a panel discussion with
>>>>>>> some
>>>>>>> chat-bots about making some sense of the "vector-wide scalar word"
>>>>>>> and "character machines", on commodity hardware about ubiquitous
>>>>>>> operations.
>>>>>>>
>>>>>>>
>>>>>>> It's considered at least tangentially relevant to comp.lang.c and
>>>>>>> comp.lang.c++ because for example text is ubiquitous and the targets
>>>>>>> would be low-level, while the higher-level languages would have a
>>>>>>> same sort of patternry, and for example that libc and cstdlib are
>>>>>>> standard, and as with regards to POSIX and Unicode and so on.
>>>>>>>
>>>>>>> Please feel free to excuse or ignore, or comment as freely.
>>>>>>>
>>>>>>> Thanks for reading.
>>>>>>>
>>>>>>
>>>>>>
>>>>>> [ viswath-charmaigne.txt ]
>>>>>>
>>>>>>
>>>>>>
>>>>
>>>>
>>>
>>
>
>
[ viswath-charmaigne-20260801.txt ]
Standard Algorithms and vr-blocks
There are two standard-algorithms.
sa-free1
sa-fixed
The vr-block is 16-many 128b v-register vector registers.
The idea is that the finding, byte-wise, starts with the codepoint
itself in vr-1, then their properties.
Then follows the predicate(s), and rangepoint(s), where a predicate
is one v-registers, and a rangepoint is two v-registers.
Then, the idea is that the standard algorithm, result making a
bit-string for sa-free1/sa-star, and a bit-string for sa-fixed/sa-drift,
about when always both are computed, then that the interpretation is
according to the mode, so that it's always the same brief instruction
listing.
vr-1: codepoints (bytes, contiguous as codepoints, text))
vr-2: properties main (text)
vr-3: properties secondary (text)
vr-4: properties tertiary (text) -x
vr-5: predicates main (pattern)
vr-6: predicates secondary (pattern)
vr-7: predicates tertiary (pattern) -x
vr-8: rangepoint-upper (pattern)
vr-9: rangepoint-lower (pattern)
vr-10: complement-predicates (pattern) -x
vr-11: complement-rangepoints (pattern) -x
vr-12: complement-result (pattern)
vr-13: varibyte-only: codepoint varibyte indices (nybble
bytes-encountered, nybble bytes-remaining, text)
vr-14: varibyte-only: rangepoint varibyte indices (..., pattern)
vr-16: result
vr-15: memo/maintenance
The v-registers are either "text" for the input text, "pattern" for the
input pattern, or "maintenance/memo/result" for other state.
So, that's 14/16 of the vector registers in the 16/16 vr-block occupied,
or 12/16 for non-variable-byte encodings. Then, it's figured to toss
the tertiary properties/predicates for 2 vector registers the
intermediate results / temporaries. Then, the complement-predicates and
complement-rangepoints might be tossed, with the idea about
making "composite and complement classes" in the standard algorithm,
vis-a-vis, union/intersection/setminus, and complement, though it's
figured to include complement in the standard algorithm, since it
results a positive property for membership, for character-classes
defined by complement.
vr-constant-zero
vr-constant-ones
vr-temporary-A
vr-temporary-B
Then, the idea is that with the properties are "match-any", with the
idea that according to presence and indicators, is to make findings
(resulting 1-byte indicator bits set) about either and both nybbles,
and either and each properties. Then, "don't care" is indicated by
all 1's in the predicates, or as about the "Filtering and Finding" below.
The first properties are considered structural, since there are
placeholders for NUL, BOM, and UTF-8 in the first byte, and about the UTF-8.
The rangepoints make findings when the codepoint is within the
range, except that rangepoints never match \0 or NUL, where
NUL is indicated by the main or primary property byte.
About rangepoints, it would be required that the UTF-8 or
UTF-16 in the high/low result overall a range matching,
about though that the arithmetic is from the high byte,
vis-a-vis, high/low surrogate pairs in UTF-16, as with regards
to that UTF-8 codepoints match the order of the Unicode codepoints,
and then with regards to the comparison cascading down.
So, the rangepoint is complicated, since in the multi-byte and
vari-byte, when the initial segment, for example the first byte, is
equal, then comparison is indicated by following bytes, about that this
involves synthesizing the operation on the vector registers, or, working
backwards on those whence found, that a first sort of finding, when
UTF-8 or vari-byte, populates the varibyte indices, then that the
indices must match as the numbers are to be in the range, eg, to
multiply the rangepoints by their XOR, resulting zeros, and to then be
excluding zero in the standard algorithm.
The varibyte-only then is upon loading the input data and input pattern,
that for each of the 16 bytes in the input, 2 bits encode 0, 1, 2, 3, of
any detected leading-byte (for UTF-8), so that 16 x 2b = 32b, then the
idea is to lookup each of 4 bytes of those, into a table that thus
encodes the bytes-encountered and bytes-remaining, putting those
together as nybbles under the bytes, then that the smearing can work the
indices, that as part of "initialiation, shift, trim" is "initialize,
shift, smear, trim", or IST for not-vari-byte and ISST for vari-byte.
Thus, the gathering stage (not stall-less) loads main-properties,
then makes detection and gathering of varibyte-indices,
then besides that properties according to the expression context,
in the gathering stage (not stall-less).
Then, the varibyte stage computes the indices, and then the
character classes or properties are assigned across the width
of the variable-length character under its codepoints, so
then that the smearing, among various variable-length codepoints,
when shifting and smearing, shifts the predicates and smears
(unsmears respectively) the predicates.
Smearing then is complicated, since it's to maintain the characters,
of the patterns, as under the characters, of the input, while
computing their associations via the bytes. Then it's figured
that characters of different sizes can't match codepoint-wise,
while, their properties as same for each can match, then about
when squeezing a longer character under a shorter character,
about how to maintain the relation of the character classes
the properties/predicates and the codepoints/rangepoints
byte-wise and char-wise, their findings and matchings.
So, for shifting and smearing, it looks to involve the pattern,
and keeping a copy of the pattern for the original pattern,
and derived lengths, then that each shift is an operation afresh
off that, when before the idea was to shift it byte-wise, moving
the pattern along, when searching across/find-first or across/find-long,
and the drifting findings, vis-a-vis down/find-next and when instead it
would be simpler to match-multiple, about find-first/find-long and
find-next/find-plex.
The idea of smearing here is that according to vertical byte-wise
arithmetic: that the evaluations occur as for matching characters,
in the input data, i.e. to the offsets/extents in the input data.
Then, when shifting a pattern, there are inputs the input text
I of length W and input pattern P of length K. These are organized
byte-wise, where the most-significant-byte of the v-register with
16B, which defines W, is the most-significant bytes of the pattern
P of length K, which is left-aligned in the input patterns (predicates
and rangepoints) in vector-registers, also of length W, then,
the vari-byte shifting/smearing, or about IST and ISST/ISVST procedures,
would probably be occurring on the g-registers as part of maintenance.
Then offsets in I and P are both zero-indexed as byte-indexed,
and, character-indexed, or octets and characters:
I_b
P_b
I_c
P_c
and the idea is that the pattern of character is to have that
the bytes of character c in the pattern P are underneath the bytes of
character c in the input I.
Then, the idea that when a smearing results a squeeze, or a smearing
results a spread, that the vari-byte indices condition
matching/not-matching the rangepoints, while, the predicates match the
properties.
squeeze: when smearing reduces a P_c to the width of I_c
spread: when smearing increments a P_c to the width of I_c
For the non-vari-byte cases, these can be trivial, when 1-byte = 1-char,
and shifts of distance d-many bytes, have that d(b) = d(c). Then, in the
vari-byte, the shifts of distance d involve the partial sums of the
Input, the offsets in bytes, define what must the offsets of the
Pattern, then about also maintaining these when straddling,
the running offsets, when shifting patterns in drifting.
Then, the multi-byte and variable-byte get involved for
the character-set and character-encoding, that these
are parameters, and indicate whether smashing/smearing
is relevant, and the lookups of the properties.
--cs-multibyte = 1, 2, 4
--cs-varibyte = f, t
--character-set-and-encoding
UTF-8 (multibyte 1, varibyte t)
ASCII (multibyte 1, varibyte f)
ISO-8859-*, CP-*
UTF-16 (multibyte 2, varibyte t)
UCS2 (multibyte 2, varibyte f)
UTF-32 (multibyte 4, varibyte f)
--character-set-endianness
BE
LE
These then indicate whether smearing/smashing occurs,
as with regards to the size of property lookup tables,
and with regards to in the algorithm whether the found-bytes
is the same or different than the found-chars, the offset.
The single-byte character-sets are differentiated by what
main-properties they load, then the multi-byte character
sets involve their greater lookup table, the properties
that populate for the input data the inevitable accounts
of covering their alphabets. Then there could be made
accounts of when there's only ASCII data in UCS2 or UTF-32,
for examples, that would probably see results as from
packing the input instead of over-riding/over-loading
the properties and making ignorance of the un-used bytes.
Then, the output is a packed vector of four 16b words:
bytes-found
char-starts
....
that there is 1b for each of the 16b in the bit-string one
for each of the 16B of the input data.
Filtering in Finding
The usual point of filtering is exclusion, then that with
regards to no-filter meaning inclusion, about that the
presence of predicates-1, 2, 3 make to indicate that when
they're absent/zero they're not contributing to AND, and
when they're present then they require AND, of the other
predicates, and rangepoint findings.
It's figured that either/or exclusive of properties/predicates
and codepoints/rangepoints is making the finding.
It's figured that these comprise classes, then as with regards
to that the complement-classes are to be indicated, in which
case to reverse the finding of the codepoints and rangepoints.
Then, a notion of adding a negation mask (vr10-vr12), basically
is to indicate when the intent is to match the complement.
The idea then is to indicate in the usual routine of filtering,
that it results the findings are so conditioned.
"Gaining performance increase requires many modifications in various
different libraries, like ffmpeg, v8, libpng, pixman or libjpeg-turbo."
- https://research.samsung.com/blog/RISC-V-and-Vectorization
It's figured that the properties/predicates make findings
(the arithmetically computed results the indicators, to be
evaluated, of membership in character-classes), with the
binary logic that AND's together the properties and predicates:
// vr-temporary-A = vr-properties-main AND vr-predicates-main
MOV vr-temporary-main, vr-predicates-main;
PAND vr-temporary-main, vr-properties-main;
then that the packed-compare compares to non-zero:
// vr-temporary-A = vr-temporary-A EQUAL vr-constant-zero
PCMPEQB vr-temporary-A, vr-constant-zero;
Then the bytes at under each I_c while have 0xFF for not-found,
and 0x00 for found, to be inverting these for 0x00 for not-found,
and 0xFF for found.
// vr-temporary-A = vr-temporary-A XOR vr-constant-one
PXOR vr-temporary-A, vr-constant-ones
That makes for "match-any", where the properties-byte is two
nybbles each with at most one bit set, and the predicates-byte
is two nybbles with zero or more bits set.
Then, that gets into the predicates otherwise might be match-all,
since, for example, that the first nybble a closed-category indicates
the category of the second nybble a closed-category. Then the idea
is that for a character class to match all of alnum, for example, that
the predicate is alnum/alpha+digit, to union the character-classes,
then that as with regards to alnum/alpha itself, that it would
spuriously or wrongly match alnum/digit, since at least one digit
matches, alnum's. Then, that gets into the idea of having two predicates
for each property, one any-match the other all-match, where yet the
v-registers are already occupied.
Then, an idea is to make for the 16-deep vr-block, an outline of
the 32-deep vrs-block, or "vector-register-serialization", with the
idea that the vr-blocks can be laid in memory by a scheduler and
then processed apiece, with regard to their serialization, vrs-blocks.
Since smearing affects both the predicates and rangepoints, the pattern,
while neither the properties nor codepoints, the text, then gets
involved that what result the values of the predicates and rangepoints,
the IST and ISVST procedures, when find-first/find-long, vis-a-vis
find-next/find-plex, that the predicates/rangepoints are to
unambiguously represent some of the composition of character classes,
among union/intersection/setminus/complement, here about union as
"match-any" and intersection as "match-all" for properties/predicates,
and complement as via an indicator, then as with regards to making that
overall the "complement" register encodes the cases to make the
arithmetic result.
vr-complement <-> vr-filterlogic
bit 1: properties-main any/all
bit 2: properties-secondary any/all
bit 3: properties-tertiary
bit 4: properties- ...
then that it's figured that the combination of main/secondary/... is
always "AND",
bit 5: complement-predicate
bit 6: complement-rangepoints
bit 7: complement-close
(bit 8: character-ignore)
Then, the challenge of that is that there isn't packed byte-wise shift,
where there's packed word-wise shift, though there could be used the
arithmetic alike drift-diff-fixed, to indicate from the synthesized
arithmetic, ..., about that each of the stages in the algorithm, is only
concerned with one of the bits.
Then, this is involved with the interfaces & internals of algorithms &
procedures,
where the algorithms are on the v-registers and the procedures are on
the g-registers,
according to the vr-block the input-text and input-pattern and their
derived properties
and conditions.
The v-registers basically have these operations, dyadic, that place the
result in the destination on x86, as with regards to that thusly being
how register allocation would be realized also for Arm et cetera. These
operations are byte-wise, meaning that they are for the packed
instructions as naturally or specifically on bytes, "built-in",
or to be synthesized from other instructions and procedures when not
otherwise present, "synthesized". These dyadic operations on the
v-registers have 128b operands, source, temporary, and destination.
built-in:
AND ("and", "&")
IOR (inclusive "or", "|")
XOR ("exclusive or", "^")
CMP ("compare")
synthesized:
ADD/ACC ("add", "accumulate", "+")
INC ("increment", "++")
SHR ("shift right", ">> (x8)", shift-right byte-wise)
(Intel has PADDB, ARM does not, idea being a synthesized accumulator
that is for the 16b instruction that on even/odd bytes add 0x0100 or
0x0001 for the high/low 8b byte of the 16b word, a procedure. Intel has
shift-right, PSRLDQ the packed-shift-right-data-logical-double-quadword
= 128 bits, by an immediate from 0-16, while Arm has EXT extract, to
extract the left-side from the constant zero vector's end and the right
side from the shifted operand's front.)
The above-described algorithms are all in those, then with regards to
the accounts of "drif-diff-fixed", is about tallying sums horizontally,
byte-wise, and detecting diffs and in segments of the vari-byte, to
result that then a procedure can convert that in the general-purpose to
a bit-sequence of indicators in 16 bits, or, a v-register of 0xFF and
0x00 bytes, 16 bytes.
Then, the general-purpose or g-registers have more of their own sort of
usual allocation problem for register allocation and layout, usually
with the idea that procedures are proscribed in the assembler/machine
code, then for the interfaces & internals, what result the entry-point
to the code-blocks or cd-blocks (instruction listings) about the memory
cache and the instruction cache, with the idea that cd-blocks of
algorithms are most-cached, and cd-blocks of procedures are
second-cached, then the higher-level operation is after that, with
regards to the "plan" and the "scheduler" as among "procedures".
The "vrr-blocks", then are the account of the map of the register file
itself, since the above built-in/synthesized operations are on the
128b wide, and the 16-deep, while the register files are variously
16-32 deep and multiples of 128b wide. So, the procedures would
involved treating a section of the vrr-blocks, and blocks recursively,
as a vr-block, according to load and store, or copy, according to
coordinates of the vr-blocks within the vrr-block the register file (of
the v-registers).
SSE4: 1 block
NEON: 2 blocks ("vertical", 2x 16-many registers, 128b wide)
AVX/AVX2: 2 blocks ("horizontal", 16-many registers, 2 x 128b wide)
AVX512: 8 blocks ("horizontal x vertical", 2 x 16-many registers, 4 x
128b wide)
SVE: ... (2 x 16-many registers, "S" x 128b wide, 128 ... 2048, 16 x
128b wide)
Then, about the composition of bvr-blocks, basically is the idea that
the results of the standard algorithms, then have for making procedures,
and that the "across" and "down" of the routines, make any sort of
mapping to the vr-blocks as independent and as of "free-lists" of
vr-blocks, instead of organization in the vertical/horizontal about
logical composition of vr-blocks, instead that mostly the vrr-block
procedures involve a "base-block" or "block 0" of the brr-block, where
operations like VEXTRACT128 and VINSERT128 (in AVXV2):
https://www.felixcloutier.com/x86/vinserti128:vinserti32x4:vinserti64x2:vinserti32x8:vinserti64x4
https://www.felixcloutier.com/x86/vextractf128:vextractf32x4:vextractf64x2:vextractf32x8:vextractf64x4
then with regards to "built-in" and "synthesized" operations would need
be figured out, where for example NEON/AVX512/SVE have upper and lower
vr-blocks addressable independently, the vertical, while the horizontal
is aliased into the 128, 256, 512, ..., bit registers, about "INSERT"
and "EXTRACT" procedures, then about what algorithms run on what
vr-blocks, the "standard algorithms", among sa-free1, sa-fixed, about
the anchored and drifting, and so on.
https://www.felixcloutier.com/x86/movdqa:vmovdqa32:vmovdqa64
https://www.felixcloutier.com/x86/vpbroadcastb:vpbroadcastw:vpbroadcastd:vpbroadcastq
Then, that would suggest that the memory organization in the address
space, would be as of the row-vector vis-a-vis the column vector, about
a "m-block", the memory block, about "mr" and "mc", then about VMOVDQA,
that INSERT and EXTRACT involve a temporary register that would be
outside of the vr-block model, in the vvr-block.
Various scheduling approaches suggest themselves, about use-cases of
"search/find", of one pattern, and prefetching and chunking, and about
"recognize/parse" of one expression/grammar, and multi-match
("multi-search"), and of various accounts with standard algorithms and
the SBC-less, and SBC-free, in the algorithms, and accounts of the
BC-less and C-less in the procedures and of the routine, and routines.
algorithm: internal (SBCF-less/SBCF-free, "the algorithm")
procedure: internal interface ("wide-internal")
routines: external interface ("wide-external")
Then, for Stall/Branch/Call-less the implementation, also has introduced
Fault, for SBCF-less, Stall/Branch/Call/Fault-less, where the idea is
that the cost of these conditions is S < B < C < F, then for validating
(not-invalidating) that algorithms are SBCF-free and that SBCF-less is
an ideal to attain for procedures, vis-a-vis correctness, where
error-modeling and bounds-modeling involved in Fault and so on are for
correctness, as paramount, built-up from the bottom-up for performance
(and conformance).
Procedures then suggest themselves.
SCHEDULE (initiated via external)
PLAN
DATA-LOAD-TEXT (memory, via external)
DATA-LOOKUP-PROPERTIES-MAIN (memory/lookup)
CHAR-INDEX-VARIBYTE
DATA-LOOKUP-PROPERTIES-SECONDARY (memory/lookup, internal/external)
DATA-LOAD-PATTERN (memory, via external)
RUN
INITIALIZE-SHIFT-TRIM
INITIALIZE-SHIFT-SMEAR-TRIM
DRIFT-DIFF-FIXED
RECEIVE (internal, result of algorithm)
EVALUATE (external, result of algorithm)
CONTINUE (iterate, recurse)
RETURN (return control)
Then, for accounts of the vr-blocks and their memory representations,
then for aligned loads onto the aliased registers of multiple vr-blocks
as initialized for their plan, is for accounts of "across" and "down"
what's scheduled and planned, according to attributes of the expressions
and what result the evaluations, constructing the blocks
opportunistically in the unbounded or bounded memory, as a result of
compiling the expression's content and resulting the load/lookup-ed
vr-blocks for then the main routine or procedure RUN.
Then, tapping away at this, then it is looking this way,
then for the idea that it's fundamental to text algorithms
and then for things like Internet Text Protocols and so on
their implementation, vis-a-vis standard libraries and the
text algorithms and so on, and string-matching ideas and
the like, about SSE4.2 (or, 16-many registers of one 128b wide
block) or NEON, and then AVX2 (with treating the aliasing in the
virtual-vector-register-block) and then AVX512+ and SVE+
as about same, for things like "Hi-Po I/O: Hippoio Internet Servers"
and also about tapping away at "AATU" and the like.
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-08-06 10:12 -0700 |
| Message-ID | <7ZucnT440LkuXen3nZ2dnZfqnPSdnZ2d@giganews.com> |
| In reply to | #143176 |
On 08/04/2026 04:00 PM, Ross Finlayson wrote:
> On 08/04/2026 03:50 PM, Ross Finlayson wrote:
>> On 08/03/2026 07:42 AM, Ross Finlayson wrote:
>>> On 08/02/2026 10:27 AM, Ross Finlayson wrote:
>>>> On 07/31/2026 05:03 PM, Ross Finlayson wrote:
>>>>> On 07/31/2026 12:55 PM, Ross Finlayson wrote:
>>>>>> On 07/30/2026 07:05 AM, Ross Finlayson wrote:
>>>>>>> On 07/30/2026 06:49 AM, Ross Finlayson wrote:
>>>>>>>> On 07/27/2026 11:45 AM, Ross Finlayson wrote:
>>>>>>>>> On 07/27/2026 11:44 AM, Ross Finlayson wrote:
>>>>>>>>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote:
>>>>>>>>>>> Hello, here I'll post some design notes and a panel discussion
>>>>>>>>>>> with
>>>>>>>>>>> some
>>>>>>>>>>> chat-bots about making some sense of the "vector-wide scalar
>>>>>>>>>>> word"
>>>>>>>>>>> and "character machines", on commodity hardware about ubiquitous
>>>>>>>>>>> operations.
>>>>>>>>>>>
>>>>>>>>>>>
>>>>>>>>>>> It's considered at least tangentially relevant to comp.lang.c
>>>>>>>>>>> and
>>>>>>>>>>> comp.lang.c++ because for example text is ubiquitous and the
>>>>>>>>>>> targets
>>>>>>>>>>> would be low-level, while the higher-level languages would
>>>>>>>>>>> have a
>>>>>>>>>>> same sort of patternry, and for example that libc and cstdlib
>>>>>>>>>>> are
>>>>>>>>>>> standard, and as with regards to POSIX and Unicode and so on.
>>>>>>>>>>>
>>>>>>>>>>> Please feel free to excuse or ignore, or comment as freely.
>>>>>>>>>>>
>>>>>>>>>>> Thanks for reading.
>>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>> [ viswath-charmaigne.txt ]
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>>>
>>>>>>>>
>>>>>>>>
>>>>>>>
>>>>>>
>>>>>
>>>>>
>>>>
>>>
>>>
>>
>
[ viswath-charmaigne-20260805.txt ]
Note: Straddle and Smear
The input-text and input-pattern, on the vr-block for the standard
algorithms, are sub-sequences of the input-text and input-pattern, that
make parameters of the memory access, to the inputs.
Memory access is safe aligned to PAGE_SIZE and in increments of
PAGE_SIZE. A process may access memory allocated to it according to
alignment to PAGE_SIZE and in increments of PAGE_SIZE. This is mostly
universally 4096B, 4 Kibibytes, while though it's a system property
powers of 2 of 512B.
The cacheline or LINE_SIZE is mostly universally 512b = 64B.
The v-register word-size or W is 16B = 128b.
So, the input-text, or T, and input-pattern, or P, has that the
input-text, is the usual contents of memory of text data, while the
pattern P, is what results as that its sub-regions are the contents of
the pattern registers in the vr-block.
The i'th character of T is T[i],
while the i'th stel of W is W[i],
the i'th patel of P is P[i]. In the vari-byte,
these don't necessarily align.
The contents of the v-registers, have that
the properties/codepoint of T[i], and
the predicates/rangepoints of P[i],
these do necessarily align, starting
at the same W[j], ending at the same W[k]
T: abcdef
tcp
tp1
tp2
tvi
P:
pc: pattern condition
pp1: pattern class primary
pp2: pattern class secondary
prl: pattern range low
prh: pattern range high
pvi: pattern vari-byte indices
So, the simple case of alignment and stride for T and P is when
the codepoints and rangepoints are uni-stel, with one stel for
each codepoint in T and one stel for each patel in P.
Then, edge and corner cases abound.
Case: prl and prh and tcp have different lengths
When prl and prh and tcp have the same lengths, then
arithmetic can establish g.t.e. and l.t.e,
len(prl) = len(prh) = len(tcp)
prl <= tcp ?
prh >= tcp ?
on the encoded form of the codepoint, i.e. UTF-8 or UTF-16.
When prl and prh have the same lengths, and tcp has
a different length, then tcp is not in the range,
since one or the other or
len(prl) = len(prh) != len(tcp)
! prl <= tcp
or
! prh >= tcp
When prl and prh have different lengths, and tcp's length
is between theirs, then there are the cases.
len(tcp) > len(prl) implies prl < tcp
len(tcp) < len(prh) implies prh > tcp
then when the lengths are different, makes for that the
CMPTRANS procedure is to determine the lengths of the
codepoints/rangepoints, and for each of high and low,
eliminate the possible cases, then result the CMPTRANS,
which makes comparison stel-wise, transitively.
So, the rangepoint patel may be wider than the codepoint, and still
contain the codepoint, yet if the rangepoint is narrower than the
codepoint, then it does not contain the codepoint.
Since the patel is no wider than the stels of the codepoint, then in
the case of _smearing_, the rangepoints may be compresed in the
patel, when the rangepoint is wider than the stels of the codepoint.
So, smearing or SMEAR, involves
"spreading", len(P[i]) < len(T[i]) implies spreading, and
"stuffing", len(P[i]) > len(T[i]) implies stuffing.
Then, spreading is somewhat simplified, that the P[i]
is simply padded, with its pp/pc1/pc2 duplicated,
while stuffing is complicated, since that the longer
rangepoint or rangepoints don't fit, and would lose information,
thus that the longer rangepoints need be "compressed",
then that information encoded in the stels-encountered/stels-remaining
to indicate that it's "compressed", and the information encoded in
the rangepoint-upper and rangepoint-lower, to establish what
the effective values are, for CMPTRANS.
The speed of CMPTRANS or more important than the
speed of SMEAR, since CMPTRANS is invoked on each (texel, patel)
in the character-class-matching-logic, while SMEAR is only
invoked in ISST, shifting/drifting the pattern. That said,
while CMPTTRANS is invoked in sa-free1, sa-fixed, sa-stars,
sa-drift, and W/S-many times in sa-stars and sa-drift, SMEAR
is invoked W/S-many times in sa-stars and sa-drift, and
invoked on each ISST, when a consumed match advances the offset.
VARIBYTE-DETECT
VARIBYTE-COUNT
The VARIBYTE-DETECT, according to whether the character-set-encoding
is UTF-8 or UTF-16, detects leading/trailing stels or
high-surrogate/low-surrogate stels, then, the VARIBYTE-COUNT operates
over runs of of a bit-sequence of detected char-starts and char-ends,
populating the stels-encountered/stels-remaining, what are the
"vr-varibyte" registers for the texels and the patels.
The stels-encountered/stels-remaining or SESR, is a byte,
and the low nybble is 2b for encountered 0-3 and 2b for
remaining 3-0. It's non-zero for texels/patels for multi-stel
characters/patterns, it's zero for texels/un-smeared patels.
Then, the high bit is set for the case of spreading and stuffing,
about how to compress the patel, so that when it's un-stuffed,
the original value results, or as stuffed, that the evaluation of
CMPTRANS is consistent.
The first idea is to store the value in the low seven bits, or,
excluding control character to have 10xxxxxx for xxxxxx = 64-many
characters, for example the alphanumeric. Then 11xxxxxx would be for
storing in the vr-maintenance of 128b, a compressed lookup-line.
Computing the VARIYBTE first, then, there's a word W subscripted i,
then for that in UTF-8 and UTF-16 the high bits indicate VARIBYTE
characters.
https://datatracker.ietf.org/doc/html/rfc2044
https://en.wikipedia.org/wiki/UTF-8
https://en.wikipedia.org/wiki/UTF-16
UTF-8 leading byte leading bits:
1: 0b
2: 110b
3: 1110b
4: 11110b
5: 111110b
UTF-8 trailling byte leading bits:
X: 10b
The UTF-8 RFC has for five-many bytes, while, the usual idea is
that Unicode only needs 2^21 bits worth for four-many bytes.
high-surrogate: starts with 0xD8, 1101 10b
low-surrogate: starts with 0xDC, 1101 11b
Then, for UTF-8, the idea seems to be to count the leading on-bits,
then decode that, while for UTF-16, seems to be to match for 0xD,
then count the leading on-bits. and decode that.
temp = UTF16 & 0xDC00
if (temp & 0xD000 != 0xD000) -> not multi (SESR = 0, 0)
if (temp & 0xD400 == 0xD499) -> low-surrogate (SESR = 1, 0)
-> high-surrogate (SESR = 0, 1)
Then what I'm looking for is prefix_rank(B) that results the
rank of the prefix code
0 = 0
10... = 1, trailing
110... = 2, leading 2
1110... = 3, leading 3
11110... = 4, leading 4
111110... = 5, leading 5
then decode from that the count. What I figure is to find the first
zero, eg find-first-clear, or find-first-set on the negation.
So, it seems to make for INSERTGV and EXTRACTVG alongside INSERT
and EXTRACT, where the "G" is for g-register and v for v-register
EXTRACTLO
EXTRACTHI
INSERTHI
INSERTLO
Then, for the byte operation "FFZERO" or "UTF8TAG", to extract EXTRACTLO
the hi word of a v-register T or P_up or P_dn, into a g-register,
UTF8TAG that 64-bit register, INSERTLO, while having the VV-block
aliasing implicit.
INSERTHI( vr-texel-varibyte, UTF8TAG(EXTRACTHI(vr-texel)))
INSERTLO( vr-texel-varibyte, UTF8TAG(EXTRACTLO(vr-texel)))
Then what I'd want for FFZERO is arithmetic that first clears all the
bits after the first zero, then 6X shifts off the first bit and adds the
carry onto the result, so the count of bits accumulates. Then,
find-first-clear is probably more direct. Otherwise there's basically
bit-test looping/incrementing to finding the first clear-bit and then
that's the prefix_rank.
Then, to compute the stels-encountered/stels-remaining for UTF8, is
to subtract one from the prefix_rank, that gives stels-remaining,
where stels-encountered is 0, then that those are run out, incrementing
the one and decirementing the other while the other.
Then, writing the expression, nested expression, and that it's
to reverse-unroll to an instruction listing, brings the idea that
the equivalent functional/procedural forms have these implicit :
subscripts,
bounds,
arrangements,
cases,
that are then for the "typed, templating assembler language" the
idea that each input and output type has its array bounds and
widths and constituent words, then that the implication is to
derive the un-rolling of that, then for example where "v-texel"
is in a "vr-block" in a "vvr-block", more implicits for the reference
the accessor, that variables are accessors and functions are
interpolators, that the
accessors, as sources and destinations, or sources and sinks, and
interpolators, that have some specialization for the types x dimensions,
then compose in a way that has a normal form as a code listing of
instructions, ..., so that then those are broken-out (enumerated)
and written-out.
So, to define FILL_VARIBYTE as
INSERTHILO( v-texel-varibyte, UTF8TAG(EXTRACTHILO(v-texel)))
is that it enumerates over hi and lo for EXTRACT and INSERT, and places
them back since they're aligned, skipping over that the UTF8TAG was
applied, which is applied to both hi and lo, each of its 8 bytes.
FILL_TEXEL_VARIBYTE (vr-texel, vr-texel-varibyte) =
EXTRACTHILO(vr-texel) . UTF8TAG . INSERTHILO(vr-texel-varibyte)
FILL_VR_VARIBYTE(vr-src, vr-dst) = EXTRACT(vr-src) . UTF8TAG .
INSERT(vr-dst)
Then, the context of an accessor, and contexts of EXTRACT and INSERT,
have that those are accessors,
FILL_VR_VARIBYTE(vr-src(vr-block-1), vr-dst(vr-block-1)) =
EXTRACT(vr-src) . UTF8TAG . INSERT(vr-dst)
has that accessors are making accessors, then for example, that "." is
both like concatenation, and, like dereference, where it were so that
the compilation off the declarations, generated (made derived) that in a
higher-level language, it results a code-model, for a sort
of "interpolating-interpreter", and a "functional language", that has a
natural form as a procedural language, if not so much vice-versa.
https://github.com/codereport/array-language-comparisons
So, to sort this out a bit, figure that there's an assembler language
with instruction listings, or a procedural language like "BASIC" or
something. Then, the idea is that accounts of loops are left out,
instead for accounts of "interpolations", interpolating from action to
action, with that the loops are implicit, and of fixed dimensions and
un-rolled, then about that the operations: are dyadic usually with
two-many operands src/dst, yet they also have the implicit data-type its
width, and as it contains other data-types as a collection its length,
those being usually enough ratios of powers-of-two,
so that then when concatenating instructions or interpolators
(transformers), that between the poles of them are interpolated the
instructions of the inner body, un-rolled, or for example when two
instructions are defined to be specialized through an outer body, the
implicits of that.
This then would be key for writing the SIMD/SWAR, because the SIMD
functions would be the same as the SWAR, ..., and about how it goes that
then besides between operands of the same size/data-type/dimensions, are
halving/doubling or insert/extract.
Then, why this is relevant to languages like Java or C++ or C, is that
the expressions like EXTRACT . UTF8TAG . INSERT, happen to look just
like field references, for what are structs of "interpolator bodies",
that then at compile-time, a recursive and exhaustive sort of building
up the declarations and definitions, can then result the simple sort of
combinators I suppose, since the types are sort of simplified because
they're only data-types with relative-width and relative-length, and
signed or unsigned integer or floating-point definition in the
instruction sets, then that "re-write rules" or "template
specializations" can be written and found in the "recursive and
exhaustive" sort of compilation, these kinds of things.
LOAD_OR_STORE(src, const, dst) = LOAD(src) . OR(const) . STORE(dst)
Then, this idea of an "accessors and interpolators" language, for
example, it's a sort of language. The idea is that interpolators are
assignable, and then that's at compile-time, and results an instruction
listing, that's fixed.
So, the mode of expression (specification) is to be figured out, with
the idea that where type-theory is usually organized about the
set-theoretic and is-a/has-a (as of collections of relation and
membership), here this is a sort of typing specific to the ordinal
instead of the cardinal as it were, about these relative width and
length quantities, is more of a "sits and fits" instead of a
"is-a/has-a" type outline.
Note: Lookup and fixed-lookup in code & data
Replacing 256B memory lookup with arithmetic instruction.
The idea here is to replace an abstract memory lookup with
arithmetic, basically XOR'ing the value with the index, comparing
to zero, which results 1 only if a match, packed in the arithmetic
so 16 at a time the lookup, and then multiplies the lookup-value
by 0 for no-match and 1 for match, and then accumulating that
into a sum, only one code will match, so it's on the order of 256
instructions for computing the lookup of 16-many, where it's figured
that simply using random access will instead make 16-many lookups,
that also involves extracting/inserting the bytes.
So, the initial main-class lookup can be in code&data instead of memory,
figuring that it's SBC-less. Then for the secondary or dynamic lookup,
that can also be compiled into code&data. About that filling the
instruction-cache instead of the memory-cache,
Note: Straddle and Smear, continued
So, straddle and smear have that their internal operations are
fragments after the fragments that make for "stellate and slice",
about the input work the the input-text and input-pattern, that
it's figured that the input-text in texels (encoded characters) is
from a file or stream and the input-pattern in patels (laid-out
representations of evaluable values of character-classes), has
that the operations of "find" and "match", which don't make
any account of transformations, yet have that "stellate" and
"slice" are implicit, to result the character-offset after the
stel-offset, and then the pattern-offset after its patel-offsets, about
the structure of pattern, that makes it fungible to then apply
the stellate and slice to the input-text and input-pattern apiece,
that then the patels are aligned under the texels in the vector,
so that the vector operations operate on them in parallel.
typedef unsigned char uint8_t;
typedef unsigned short int uint16_t;
typedef unsigned long int uint32_t;
typedef unsigned long long int uint64_t;
// uint128_t, no built-in type
typedef union v_reg {
uint8_t stels1[16];
uint16_t stels2[8];
uint32_t stels4[4];
uint64_t stels8[2];
uint128_t stel
} v_reg;
typedef union g_reg {
uint8_t stels1[16];
uint16_t stels2[8];
uint32_t stels4[4];
uint64_t stel;
} g_reg;
typedef struct vr_block {
v_reg registers[16];
} vr_block;
typedef struct gr_block {
g_reg registers[8];
} gr_blockl
unsigned char BLOCK_COUNT = 2;
typedef struct vvr_block {
vr_block blocks[BLOCK_COUNT];
} vvr_block;
The idea here is to write a C representation of the entire model of the
Viswath then the Charmaigne, since, that way all the algorithms are
to get tested for expressibility and correctness in C-code as
pseudo-code, then about the above, making for the fundamental routines,
that simply have assembler instructions the built-in and synthesized,
the expressible.
unsigned int WV = 16;
unsigned int WG = 8;
unsigned int DVRB = 16;
v_reg* AND(v_reg src, v_reg dst) {
for (int i = 0; i < WV; i++) dst.stels1[i] = src.stels1[i] & dst.stels1[i];
return &dst;
}
Then, the input-text is considered pretty simple, or that it's a region
of memory.
typedef unsigned char* start_of_input;
typedef unsigned char* end_of_input;
type struct input_window {
start_of_input begin;
end_of_input end;
start_of_input aligned_begin;
end_of_input aligned_end;
off_t extent;
off_t extent_aligned;
off_t word_count;
off_t inset_start;
off_t inset_end;
} input_window;
struct input_window _input_window(start_of_input begin, end_of_input end) {
struct input_window _;
_.begin = begin;
_.end = end
_.inset_begin = begin % W;
_.inset_end = W - (end % W);
_.extent = end - begin;
_.extent_aligned = end_aligned - begin_aligned;
_.word_count = _.extent_aligned / W;
_.aligned_begin = begin = _.inset_begin;
_.aligned_end = end + _.aligned_end;
return _;
}
Then, the contents of the v-register, is that it is W-many bytes, where
W is a constant the width of a v-register, and an array V, or sometimes
overloaded an array W, subscripted/index (0, W-1).
off_t input_window_word_count(input_window _) {
return _.word_count;
}
Then, the work is the overall routine, then within that, is a loop
over the words, then the standard algorithm/procedure/maintenance
on those.
v_register load(struct input_window i_w, off_t window_index) {
struct v_register _;
for (int i = 0; i < W; i++) _.stels1[i] = *(i_w.begin_aligned + W *
window_index);
}
off_t word_inset_left(struct input_window i_w, off_t window_index) {
return window_index == 0 ? i_w.inset_start : 0;
}
off_t word_inset_right(struct input_window i_w, off_t window_index) {
return window_index == i_w.word_count - 1 ? i_w.inset_end : 0;
}
unsigned char* begin;
unsigned char* end;
struct input_window input_text = _input_window(begin, end);
for (off_t i = 0; i < input_text.window_count; i++) {
struct v_register v_reg = load(input_text, i);
off_t inset_left = word_inset_left(input_text, i);
off_t inset_right = word_inset_right(input_text, i);
}
Then, the pattern detail is similar, that though is contrived in its
construction, about that while the input_text is just the uninterpreted
octet-sequence betwen begin and end then so aligned, the input_pattern
is a tuple of uninterpreted octet sequences, or "patels".
typedef union nybbled_t {
uint8_t _byte;
struct _nybbles {
unsigned int hi: 4;
unsigned int lo: 4;
}
} nybbled_t;
typedef nybbled_t properties_primary_t;
typedef union nybbled_or_uint_t {
nybbled_t nybbled;
uint8_t _uint;
} nybbled_or_uint_t;
It's figured for these union types that v-registers operate
on bytes/stels, while g-registers operate on bits.
typedef nybbled_or_uint_t properties_secondary_t;
typedef struct patel_t {
nybbled_t predicates_primary;
nybbled_or_uint_t predicates_secondary;
uint8_t rangepoint_upper[4];
uint8_t rangepoint_lower[4];
uint8_t conditions;
} patel_t;
The texel then is similar.
typedef struct texel_t {
uint8_t codepoint[4];
nybbled_t properties_primary;
nybbled_or_uint_t properties_secondary;
} texel_t;
Then both texel's codepoint and patel's rangepoints
have the vari-byte (vari-stel) stels-encountered / stels-remaining.
typedef union sesr_t {
uint8_t _stel;
struct _sesr {
unsigned int _reserved_padding: 4;
unsigned int stels_encountered: 2;
unsigned int stels_remaining; 2;
}
}
The conditions are flags.
typedef struct conditions_t {
unsigned int properties_primary_all_match: 1; // default any-match, 0
unsigned int properties_secondary_all_match: 1; // default any-match, 0
unsigned int properties_secondary_exact_match: 1; // default
any/all-match, default 0
unsigned int complement_predicates: 1; // default 0
unsigned int complement_rangepoints: 1; // default 0
unsigned int complement_result: 1; // default 0
} conditions_t;
typedef struct patel_t {
properties_primary_t predicates_primary;
properties_secondary_t predicates_secondary;
// right-aligned
uint8_t rangepoint_upper[4];
uint8_t rangepoint_lower[4];
off_t len_rangepoint_upper
off_t len_rangepoint_lower;
sesr_t sesr_upper;
sesr_t sesr_lower;
conditions_t conditions;
} patel_t;
The texel then is similar.
typedef struct texel_t {
uint8_t codepoint[4];
off_t len_codepoint;
sesr_t sesr;
properties_primary_t properties_primary;
properties_secondary_t properties_secondary;
} texel_t;
Then, the it's figured that the actual contents of
the codepoints/rangepoints is that the texels/patels
have references to the word V they are in, or the context
of the word and the previous and next word in the input_text
and input_pattern, with regards to extracting the input_text
from memory, and extracting and shifting the input_pattern
from memory.
typedef struct codepoint_ref_t {
v_register v;
off_t index;
off_t offset;
off_t extent;
} codepoint_ref_t;
typedef patel_vector_t {
off_t stel_shift;
};
Then, there's involved the compression or "stuffing" in smearing,
or how to resolve the reference to the codepoint, for the rangepoints
in the patel-vector, then for what's to result the "standard algorithm"
the "character-class-matching-logic", as defined already by patels
under texels, and vectorized findings of character-class-matching-logic.
Then, for the stuffing, is to be figured out how to encode the
rangepoints, which already have involved that the lower-bound and
upper-bound may be of different lengths, since they are encoded
codepoints in the character-set-encoding themselves, about the
implementation of CMPTRANS the procedure, and ISST
(Initialize-Shift-Smear-Trim), where IST (Initialize-Shift-Trim) is
simplified when all the codepoints and rangepoints are single-stel or
uni-stel,
since then that's just about copying the vector, shifting it the
increment plus the inset_left, and trimming it past the inset_right.
Then, the input_text is never modified, merely copied (loaded, inserted
on the v-register in the vr-block in the vvr-block), while the
input_pattern also is never modified, yet the codepoint_ref_t's are to
be describing a sliding/shifting window over the input_pattern, thus
that in the vr-block, the offset of the pattern aligns patels under the
texels, stel-per-stel, that thusly the "standard algorithm" is SBC-free,
and for simple (smooth) data, also that "standard procedure" is under a
small constant.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2026-08-06 12:37 -0700 |
| Message-ID | <1152npc$3pi9$2@dont-email.me> |
| In reply to | #143244 |
On 8/6/2026 10:12 AM, Ross Finlayson wrote:
[...]
Fwiw, my older region allocator works wonders... This is pre C11 :^)
https://pastebin.com/raw/f37a23918
#if ! defined (RALLOC_H)
# define RALLOC_H
# if defined (__cplusplus)
extern "C" {
# endif
/**************************************************************/
#include <stddef.h>
#include <assert.h>
#if defined (_MSC_VER)
/* warning C4116: unnamed type definition in parentheses */
# pragma warning (disable : 4116)
#endif
#if ! defined (NDEBUG)
# include <stdio.h>
# define RALLOC_DBG_PRINTF(mp_exp) printf mp_exp
#else
# define RALLOC_DBG_PRINTF(mp_exp) ((void)0)
#endif
#if ! defined (RALLOC_UINTPTR_TYPE)
# define RALLOC_UINTPTR_TYPE size_t
#endif
typedef RALLOC_UINTPTR_TYPE ralloc_uintptr_type;
typedef char ralloc_static_assert[
sizeof(ralloc_uintptr_type) == sizeof(void*) ? 1 : -1
];
enum ralloc_align_enum {
ALIGN_ENUM
};
struct ralloc_align_struct {
char pad;
double type;
};
union ralloc_align_max {
char char_;
short int short_;
int int_;
long int long_;
float float_;
double double_;
long double long_double_;
void* ptr_;
void* (*fptr_) (void*);
enum ralloc_align_enum enum_;
struct ralloc_align_struct struct_;
size_t size_t_;
ptrdiff_t ptrdiff_t;
};
#define RALLOC_ALIGN_OF(mp_type) \
offsetof( \
struct { \
char pad_RALLOC_ALIGN_OF; \
mp_type type_RALLOC_ALIGN_OF; \
}, \
type_RALLOC_ALIGN_OF \
)
#define RALLOC_ALIGN_MAX RALLOC_ALIGN_OF(union ralloc_align_max)
#define RALLOC_ALIGN_UP(mp_ptr, mp_align) \
((void*)( \
(((ralloc_uintptr_type)(mp_ptr)) + ((mp_align) - 1)) \
& ~(((mp_align) - 1)) \
))
#define RALLOC_ALIGN_ASSERT(mp_ptr, mp_align) \
(((void*)(mp_ptr)) == RALLOC_ALIGN_UP(mp_ptr, mp_align))
struct region {
unsigned char* buffer;
size_t size;
size_t offset;
};
static void
rinit(
struct region* const self,
void* buffer,
size_t size
) {
self->buffer = buffer;
self->size = size;
self->offset = 0;
RALLOC_DBG_PRINTF((
"rinit(%p) {\n"
" buffer = %p\n"
" size = %lu\n"
"}\n\n\n",
(void*)self,
buffer,
(unsigned long int)size
));
}
static void*
rallocex(
struct region* const self,
size_t size,
size_t align
) {
unsigned char* align_buffer;
size_t offset = self->offset;
unsigned char* raw_buffer = self->buffer + offset;
if (! size) {
size = 1;
}
if (! align) {
align = RALLOC_ALIGN_MAX;
}
assert(align == 1 || RALLOC_ALIGN_ASSERT(align, 2));
align_buffer = RALLOC_ALIGN_UP(raw_buffer, align);
assert(RALLOC_ALIGN_ASSERT(align_buffer, align));
size += align_buffer - raw_buffer;
if (offset + size > self->size) {
return NULL;
}
self->offset = offset + size;
RALLOC_DBG_PRINTF((
"rallocex(%p) {\n"
" size = %lu\n"
" alignment = %lu\n"
" origin offset = %lu\n"
" final offset = %lu\n"
" raw_buffer = %p\n"
" align_buffer = %p\n"
" size adjustment = %lu\n"
" final size = %lu\n"
"}\n\n\n",
(void*)self,
(unsigned long int)size - (align_buffer - raw_buffer),
(unsigned long int)align,
(unsigned long int)offset,
(unsigned long int)self->offset,
(void*)raw_buffer,
(void*)align_buffer,
(unsigned long int)(align_buffer - raw_buffer),
(unsigned long int)size
));
return align_buffer;
}
#define ralloc(mp_self, mp_size) \
rallocex((mp_self), (mp_size), RALLOC_ALIGN_MAX)
#define ralloct(mp_self, mp_count, mp_type) \
rallocex( \
(mp_self), \
sizeof(mp_type) * (mp_count),\
RALLOC_ALIGN_OF(mp_type) \
)
static void
rflush(
struct region* const self
) {
self->offset = 0;
RALLOC_DBG_PRINTF((
"rflush(%p) {}\n\n\n",
(void*)self
));
}
#undef RALLOC_DBG_PRINTF
#undef RALLOC_UINTPTR_TYPE
/**************************************************************/
# if defined (__cplusplus)
}
# endif
#endif
[toc] | [prev] | [next] | [standalone]
| From | Ross Finlayson <ross.a.finlayson@gmail.com> |
|---|---|
| Date | 2026-08-31 08:16 -0700 |
| Message-ID | <m6ecnQlIyuFABwj3nZ2dnZfqn_SdnZ2d@giganews.com> |
| In reply to | #143244 |
Well, I've been tapping away at this. It's been figured out what are the common/portable idioms for vector-wide scalar-word operation on "128b-wide v-registers" in "multiples of 128b-wide full-registers" so that the operations are entirely single-instruction/multiple-data and that the profiles of SSE4.2, AVX2, AVX512, NEON, SVE or AMD/ARM entirely support these then that the alias-less/stall-less/branch-less/call-less/fault-less with costs A < S < B < C < F make it so for that these are the portability layer aligned with the commodity units the what result being multiples of 128b (16B-wide) virtual-vector-registers, then about things like UTF-8 and UTF-16 then for the standard sorts of algorithms sa-free1 and sa-fixed in the "character machine". Looking around now I've heard about "StringZilla", which has its own sort of account, that though with regards to character-data it expands it to 32-bits, while this account leaves the data in-place its wire-format network-order encounter-order, with regards to "stels, copels, patels, texels, and graphemels". /* file:vwsw-sse42.S */ /* NB: "128-bit Legacy SSE" */ .include "vwsw-common.S" .include "vwsw-amd.S" .set COUNT_BLOCKS, 1 .set COUNT_BANKS, 1 .set FW, W * COUNT_BLOCKS .set FD, D * COUNT_BANKS .section .rodata .balign FW .include "vwsw-constants.S" .set vr_1, %xmm0 .set vr_2, %xmm1 .set vr_3, %xmm2 .set vr_4, %xmm3 .set vr_5, %xmm4 .set vr_6, %xmm5 .set vr_7, %xmm6 .set vr_8, %xmm7 .set vr_9, %xmm8 .set vr_10, %xmm9 .set vr_11, %xmm10 .set tmpA, %xmm11 .set tmpB, %xmm12 .set tmpC, %xmm13 .set tmpD, %xmm14 .set tmpE, %xmm15 /* zero-out all registers (non-temporaries) */ .macro _zeroall _zero vr_1 _zero vr_2 _zero vr_3 _zero vr_4 _zero vr_5 _zero vr_6 _zero vr_7 _zero vr_8 _zero vr_9 _zero vr_10 _zero vr_11 .endm /* load FWB aligned */ .macro _load dst, mem movdqa \mem, \dst .endm /* store FWB aligned */ .macro _stor mem, src movdqa \src, \mem .endm /* copy from src to dst */ .macro _copy dst, src movdqa \src, \dst .endm /* extract a WB word from src at block/bank to dst */ .macro _extract dst, src, imm_block=0, imm_bank=0 /* no-op in SSE4.2 */ .endm /* insert a WB word from src to dst at block/bank */ .macro _insert dst, src, imm_block=0, imm_bank=0 /* no-op in SSE4.2 */ .endm /* bit-wise-logical AND full-register */ .macro _and dst, src pand \src, \dst .endm /* bit-wise-logical OR full-register */ .macro _ior dst, src por \src, \dst .endm /* bit-wise-logical XOR full-register */ .macro _xor dst, src pxor \src, \dst .endm /* byte-wise compare dst to rhs result 0x00/0xFF in dst */ .macro _cmpeq dst, rhs pcmpeqb \rhs, \dst .endm /* byte-wise compare dst < rhs (unsigned) to result 0x00/0xFF in dst */ /* preserves: rhs */ .macro _cmplt dst, rhs, tmp1 _tsus \rhs, \tmp1 _tsus \dst, \tmp1 _copy \tmp1, \rhs pcmpgtb \rhs, \dst _copy \dst, \rhs _copy \rhs, \tmp1 _tuss \rhs, \tmp1 .endm /* byte-wise compare dst > rhs (unsigned) to result 0x00/0xFF in dst */ /* preserves: rhs */ .macro _cmpgt dst, rhs, tmp1 _tsus \rhs, \tmp1 _tsus \dst, \tmp1 pcmpgtb \dst, \rhs _tuss \rhs, \tmp1 .endm /* bit-wise-logical NOT */ .macro _not dst, tmp1 _ones \tmp1 _xor \dst, \tmp1 .endm /* load constant 0x00 byte-wise into dst */ .macro _zero dst _xor \dst, \dst .endm /* load constant 0xFF byte-wise into dst */ .macro _ones dst _zero \dst _cmpeq \dst, \dst .endm /* load constant 0x80 byte-wise into dst */ .macro _bithi dst _load \dst, constant_bithi(%rip) .endm /* load constant 0x01 byte-wise into dst */ .macro _bitlo dst _load \dst, constant_bitlo(%rip) .endm /* load constant 0xF0 byte-wise into dst */ .macro _nybhi dst _load \dst, constant_nybhi(%rip) .endm /* load constant 0x0F byte-wise into dst */ .macro _nyblo dst _load \dst, constant_nyblo(%rip) .endm /* load constant 0x10 byte-wise into dst */ .macro _nybhibitlo dst _load \dst, constant_nybhibitlo(%rip) .endm /* load constant 0x08 byte-wise into dst */ .macro _nyblobithi dst _load \dst, constant_nyblobithi(%rip) .endm /* shift right byte-wise (effective division) */ .macro _shr dst, imm_byte_offset, tmp1 psrldq \imm_byte_offset, \dst .endm /* shift left byte-wise (effective multiplication) */ .macro _shl dst, imm_byte_offset, tmp1 pslldq \imm_byte_offset, \dst .endm /* shift forward byte-wise */ .macro _shf dst, imm_byte_offset, tmp1 _shl \dst, \imm_byte_offset .endm /* shift backward byte-wise */ .macro _shb dst, imm_byte_offset, tmp1 _shr \dst, \imm_byte_offset .endm /* press each of FWB from src_g byte-wise saturated 0x00/0xFF into dst_v */ .macro _press dst_v, src_g movq \src_g, \dst_v pshufb constant_pressctrl(%rip), \dst_v _and \dst_v, constant_pressmask(%rip) _cmpeq \dst_v, constant_pressmask(%rip) .endm /* pick each of FWB from src_v high-bit as FWb into dst_g */ .macro _pick dst_g, src_v, tmp1, g_tmp pmovmskb \src_v, \dst_g .endm /* translate-signed-to-unsigned byte-wise in dst */ .macro _tsus dst, tmp1 _bithi \tmp1 paddb \tmp1, \dst .endm /* translate-unsigned-to-signed byte-wise in dst */ .macro _tuss dst, tmp1 _bithi \tmp1 paddb \tmp1, \dst .endm /* true-if-false byte-wise translates 0x00 to 0xFF in dst */ .macro _tiff dst, tmp1, tmp2 _zero \tmp1 _copy \tmp2, \dst _cmpeq \tmp2, \tmp1 _ior \dst, \tmp2 .endm /* saturate the high-bit set/clear to 0xFF/0x00 in dst */ .macro _sathi dst, tmp1 _bithi \tmp1 _and \tmp1, \dst _cmpeq \dst, \tmp1 .endm /* saturate the low-bit set/clear to 0xFF/0x00 in dst */ .macro _satlo dst, tmp1 _bitlo \tmp1 _and \tmp1, \dst _cmpeq \dst, \tmp1 .endm /* saturate any set-bit to 0xFF else 0x00 in dst */ .macro _saton dst, tmp1 _zero \tmp1 _cmpeq \dst, \tmp1 _not \dst, \tmp1 .endm .macro _tease1 dst, imm_bit_offset, tmp1 _load \tmp1, constant_tease1(%rip) psrlw (8 - 1 - \imm_bit_offset), \dst _and \dst, tmp1 .endm .macro _tease2 dst, imm_bit_offset, tmp1 _load \tmp1, constant_tease2(%rip) psrlw (8 - 2 - \imm_bit_offset), \dst _and \dst, tmp1 .endm .macro _tease3 dst, imm_bit_offset, tmp1 _load \tmp1, constant_tease3(%rip) psrlw (8 - 3 - \imm_bit_offset), \dst _and \dst, tmp1 .endm .macro _tease4 dst, imm_bit_offset, tmp1 _load \tmp1, constant_tease4(%rip) psrlw (8 - 4 - \imm_bit_offset), \dst _and \dst, tmp1 .endm
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-08-03 20:14 +0200 |
| Subject | Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) |
| Message-ID | <114qlpk$t7q2$1@solani.org> |
| In reply to | #143063 |
Hi, I even don't remember exactly why I landed in comp.theory. A yes, because Rossy Boy, was hooked on SIMD and didn't understand Hack. But the Hack work, rather belongs to my Alma Mater Zurich and my personal heros, Gutnecht and Wirth, who wrote a one pass Modula compiler during some christmas holidays, back then when I was student. Not sure whether the Ljubljana School can do that, when I read this here: Finite Algebraic Effects as dicts and such https://www.philipzucker.com/bdd_term_alg_effects/ I only find gibberish like: - “Data” is somehow less mysterious to me than “computation”. [..] I don’t even know what “computation” really is - In temporal logic, there is a logic CTL which talks about computation trees. - Algerbaic (LoL) effects is almost a complete hackery abuse of the notion of arity and that’s neat. - Then there are 10^10 etc vectors which show up if you discretize 3d/4d space. [..] or reinforcement learning. - Etc.. WTF is this guy smoking? I mean he even doesn't uses math notation, only posts Python code fragments à go go, possibly a Python brain damage. But still less sever than Rossy Boys. Bye Ross Finlayson schrieb: > Hello, here I'll post some design notes and a panel discussion with some > chat-bots about making some sense of the "vector-wide scalar word" > and "character machines", on commodity hardware about ubiquitous > operations. > > > It's considered at least tangentially relevant to comp.lang.c and > comp.lang.c++ because for example text is ubiquitous and the targets > would be low-level, while the higher-level languages would have a > same sort of patternry, and for example that libc and cstdlib are > standard, and as with regards to POSIX and Unicode and so on. > > Please feel free to excuse or ignore, or comment as freely. > > Thanks for reading. >
[toc] | [prev] | [next] | [standalone]
| From | Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> |
|---|---|
| Date | 2026-08-04 02:21 +0800 |
| Subject | Re: Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) |
| Message-ID | <1%4cS.103380$aXr.938@fx18.ams4> |
| In reply to | #143190 |
On 04/08/2026 2:14 AM, Mild Shock wrote: > Hi, > > I even don't remember exactly why I landed > in comp.theory. A yes, because Rossy Boy, > was hooked on SIMD and didn't understand Hack. > > But the Hack work, rather belongs to my > Alma Mater Zurich and my personal heros, Gutnecht > and Wirth, who wrote a one pass Modula That's interesting. Have you read /Software Engineering with Modula-2 and Ada/ (1984) by Richard Wiener and Richard Sincovec? I have it on my shelf, and haven't gotten to read it yet. > > compiler during some christmas holidays, > back then when I was student. Not sure > whether the Ljubljana School can do that, > > when I read this here: > > Finite Algebraic Effects as dicts and such > https://www.philipzucker.com/bdd_term_alg_effects/ > > I only find gibberish like: > - “Data” is somehow less mysterious to me > than “computation”. [..] I don’t even > know what “computation” really is Computation is at the core just a calculation. Humans used to do this, and there's a good documentary about it titled /Hidden Figures/. I assume everyone here has seen it. > > - In temporal logic, there is a logic CTL > which talks about computation trees. > > - Algerbaic (LoL) effects is almost a complete > hackery abuse of the notion of arity > and that’s neat. Are you using /algebraic effects/ when playing League of Legends? I tried it once, but discovered it's a gameplay that doesn't appeal to me. I didn't think to use /algebraic effects/ in it. > > - Then there are 10^10 etc vectors which > show up if you discretize 3d/4d space. [..] > or reinforcement learning. I don't know why you have that many vectors visiting, but please treat them with hospitality according to Zeus' laws. > > - Etc.. > > WTF is this guy smoking? I mean he even > doesn't uses math notation, only posts > Python code fragments à go go, > > possibly a Python brain damage. Or he just works in the ministry of silly walks? Enjoy! -- Johann | email: invalid -> com | http://www.myrkraverk.com/blog/ I'm not from the Internet, I just work there. | via Easynews.com https://bsky.app/profile/myrkraverk.bsky.social
[toc] | [prev] | [next] | [standalone]
| From | Mild Shock <janburse@fastmail.fm> |
|---|---|
| Date | 2026-08-03 20:35 +0200 |
| Subject | pi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School) |
| Message-ID | <114qn2b$t8nm$1@solani.org> |
| In reply to | #143191 |
Hi,
I found that this here:
public final static class RendezVous {
private final Semaphore head = new Semaphore(0);
private final Semaphore tail = new Semaphore(1);
private Object data;
public void put(Object data) throws InterruptedException {
tail.acquire();
this.data = data;
head.release();
}
public Object take() throws InterruptedException {
Object res;
head.acquire();
res = data;
tail.release();
return res;
}
}
Is almost as fast as ArrayBlockingQueue(4),
in a producer worker consumer scenario.
So I considering using the above for the
pi-WAM channels. It would be also closer
to pi-calculus by Robin Milner.
Bye
Johann 'Myrkraverk' Oskarsson schrieb:
> On 04/08/2026 2:14 AM, Mild Shock wrote:
>> Hi,
>>
>> I even don't remember exactly why I landed
>> in comp.theory. A yes, because Rossy Boy,
>> was hooked on SIMD and didn't understand Hack.
>>
>> But the Hack work, rather belongs to my
>> Alma Mater Zurich and my personal heros, Gutnecht
>> and Wirth, who wrote a one pass Modula
>
> That's interesting. Have you read /Software Engineering
> with Modula-2 and Ada/ (1984) by Richard Wiener and Richard
> Sincovec? I have it on my shelf, and haven't gotten to read
> it yet.
>
>>
>> compiler during some christmas holidays,
>> back then when I was student. Not sure
>> whether the Ljubljana School can do that,
>>
>> when I read this here:
>>
>> Finite Algebraic Effects as dicts and such
>> https://www.philipzucker.com/bdd_term_alg_effects/
>>
>> I only find gibberish like:
>> - “Data” is somehow less mysterious to me
>> than “computation”. [..] I don’t even
>> know what “computation” really is
>
> Computation is at the core just a calculation. Humans
> used to do this, and there's a good documentary about it
> titled /Hidden Figures/. I assume everyone here has seen
> it.
>
>>
>> - In temporal logic, there is a logic CTL
>> which talks about computation trees.
>>
>> - Algerbaic (LoL) effects is almost a complete
>> hackery abuse of the notion of arity
>> and that’s neat.
>
> Are you using /algebraic effects/ when playing League
> of Legends? I tried it once, but discovered it's a
> gameplay that doesn't appeal to me. I didn't think to
> use /algebraic effects/ in it.
>
>>
>> - Then there are 10^10 etc vectors which
>> show up if you discretize 3d/4d space. [..]
>> or reinforcement learning.
>
> I don't know why you have that many vectors visiting,
> but please treat them with hospitality according to
> Zeus' laws.
>
>>
>> - Etc..
>>
>> WTF is this guy smoking? I mean he even
>> doesn't uses math notation, only posts
>> Python code fragments à go go,
>>
>> possibly a Python brain damage.
>
> Or he just works in the ministry of silly walks?
>
>
> Enjoy!
[toc] | [prev] | [next] | [standalone]
Page 9 of 10 — ← Prev page 1 … 7 8 [9] 10 Next page →
Back to top | Article view | comp.theory
csiph-web