Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #124392 > unrolled thread

Viswath & Charmaigne (vector-wide scalar-word and character machines)

Started byRoss Finlayson <ross.a.finlayson@gmail.com>
First post2026-07-27 11:43 -0700
Last post2026-08-03 20:40 +0200
Articles 20 on this page of 120 — 7 participants

Back to article view | Back to comp.lang.c++


Contents

  Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 11:43 -0700
    Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-28 02:47 +0800
      Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:07 -0700
        Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Mild Shock <janburse@fastmail.fm> - 2026-07-28 00:25 +0200
          Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 16:09 -0700
            Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 16:33 -0700
          Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:48 -0700
        Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 07:44 -0700
          You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:11 +0200
            Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:25 -0700
              Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 23:49 +0800
                Re: You are still chewing on SIMD. LoL (Was: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:59 -0700
                I don't care about Java, pi-WAM is pi-calculus and WAM (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:05 +0200
                  Underneath pi-WAM is Hack VM, you can goto (Was: I don't care about Java, pi-WAM is pi-calculus and WAM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:09 +0200
                    New addition to π-WAM is π-WAM Assembly (Was: Underneath pi-WAM is Hack VM, you can goto) Mild Shock <janburse@fastmail.fm> - 2026-08-09 19:45 +0200
                  Re: I don't care about Java, pi-WAM is pi-calculus and WAM (Re: You are still chewing on SIMD. LoL) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 09:11 -0700
                    A yellow mustard called Rossy Body (Was: I don't care about Java, pi-WAM is pi-calculus and WAM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:30 +0200
                      Ignoramus or Ignorabimus: I don't care (π-WAM) (Was: A yellow mustard called Rossy Body) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:31 +0200
            Hurry Rossy Boy, the blue bus is waiting (Was: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:32 +0200
              Re: Hurry Rossy Boy, the blue bus is waiting (Was: You are still chewing on SIMD. LoL) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:36 -0700
              Look how they advertized CUDA and logical threads (Was: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:43 +0200
                Forget any arithmetization of product FSA (Was: Look how they advertized CUDA and logical threads) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:47 +0200
                  comp.lang.lisp (was: Re: Forget any arithmetization of product FSA) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 23:53 +0800
                Re: Look how they advertized CUDA and logical threads (Was: Hurry Rossy Boy, the blue bus is waiting) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 08:53 -0700
            There are two versions of Hack VM (Was: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:14 +0200
              Hack VM has also a Prolog spec (Was: There are two versions of Hack VM) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:17 +0200
                A better compiler is planned / What do you target? (Was: Hack VM has also a Prolog spec) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:22 +0200
                  It’s called . . . . enshittification (About the price tag for using a multifile/1) Mild Shock <janburse@fastmail.fm> - 2026-08-14 00:50 +0200
            Budget AI Laptop 2026 versus Cray T3D 1995 (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-08-05 14:24 +0200
      A brain desease of 20 days [Rossy Boy] (Was: Viswath & Charmaigne (vector-wide scalar-word and character machines)) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:38 +0200
        Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:05 +0200
          Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 11:13 -0700
            I don't use Rust, you are crazy [Jump off a bridge, idiot] (Was: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:22 +0200
              Standing on the shoulders of giants (Re: I don't use Rust, you are crazy [Jump off a bridge, idiot]) Mild Shock <janburse@fastmail.fm> - 2026-08-04 03:20 +0200
                You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc.. (Re: Standing on the shoulders of giants) Mild Shock <janburse@fastmail.fm> - 2026-08-04 15:18 +0200
                  How Rossy Boys plagiarism works [Copy Paste Slop] (Re: You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc..) Mild Shock <janburse@fastmail.fm> - 2026-08-04 17:56 +0200
            Postgres is in C! Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 03:46 +0800
              Re: Postgres is in C! Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 13:47 -0700
                Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:00 +0200
                  Re: Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:05 +0800
                    Re: Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) scott@slp53.sl.home (Scott Lurndal) - 2026-07-30 14:46 +0000
                    Re: Please don't extend your cross posting / What does abstract mean? Cóilín Nioclásín Glostéir <thanks-to@Taf.com> - 2026-07-30 16:00 +0000
                Please don't extend your cross posting / What does abstract mean? (Was: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:05 +0200
                Mars, the MIPS emulator in Java (was: Re: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:20 +0800
                  Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 06:59 -0700
                    Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 07:24 -0700
                      Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 09:58 +0800
                    Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 23:19 +0800
                      Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 09:23 -0700
                        Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 09:36 -0700
                          Decorum (was: Re: Mars, the MIPS emulator in Java) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 23:07 +0800
                            Re: Decorum Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-31 23:09 +0800
                            Re: Decorum Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 08:36 -0700
                              Re: Decorum Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 08:44 -0700
                              Re: Decorum Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 10:38 +0800
                        Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 10:15 +0800
                          Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-04 23:31 -0700
                            Re: Mars, the MIPS emulator in Java "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:42 -0700
                              Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 13:24 -0700
                                Re: Mars, the MIPS emulator in Java "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 13:30 -0700
                                Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 13:45 -0700
                            Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-05 20:42 -0700
                              Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 09:55 +0800
                                Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-07 05:13 -0700
                                  Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-07 05:59 -0700
                            Re: Mars, the MIPS emulator in Java Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 09:53 +0800
                              Re: Mars, the MIPS emulator in Java Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-06 20:11 -0700
              Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: Postgres is in C!) Mild Shock <janburse@fastmail.fm> - 2026-07-29 22:49 +0200
                A funny Q16.16 experiment with Hack (Was: Hack ecosystem ignorance paired with paranoia [Nand to Tetris]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:10 +0200
                  Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Re: A funny Q16.16 experiment with Hack) Mild Shock <janburse@fastmail.fm> - 2026-07-30 11:27 +0200
                    Re: Bullshit Authorized by Sarah Connor [EyeProlog Failure] (Re: Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM]) Mild Shock <janburse@fastmail.fm> - 2026-08-12 20:30 +0200
                RCan library(ironpaw) repurpose FFT hardware [Glimps into Ryzen AI 7 350] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-08-01 02:33 +0200
                Can library(ironpaw) repurpose FFT hardware [Glimps into Ryzen AI 7 350] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-08-01 02:34 +0200
              Turbo Vison, again (was: Re: Postgres is in C!) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-30 21:31 +0800
        He uses "FIFO objects", and DMA and Noc [Glimps into Ryzen AI 7 350] (Re: A brain desease of 20 days [Rossy Boy])) Mild Shock <janburse@fastmail.fm> - 2026-08-01 12:15 +0200
          Tablet and phone UBS-C remote debugging (Re: He uses "FIFO objects", and DMA and Noc) Mild Shock <janburse@fastmail.fm> - 2026-08-01 12:18 +0200
            NPUs doing 2d chess comms (Manhattan Distance or L1 Norm) (Re: Tablet and phone UBS-C remote debugging) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:11 +0200
              NACK retransmission might double Manhattan Distance (Re: NPUs doing 2d chess comms) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:23 +0200
    Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-07-27 21:18 +0200
      Re: Rossy Boy is neither Einstein nor Zweistein (Re: Viswath & Charmaigne) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 15:35 -0700
        Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-07-28 11:25 +0200
          Re: Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-28 20:39 -0700
            confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:15 +0200
              Gemini, DeepSeek, OpenAI more clever than rossy boy (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:18 +0200
              Re: confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 17:21 +0800
                In AI Acceleration nobody cares about CivetWeb (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:27 +0200
                  Re: In AI Acceleration nobody cares about CivetWeb (Was: confused rossy boy is confused) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 17:40 +0800
                    Your strictness is your problem , not mine [See WebLLM] (Was: In AI Acceleration nobody cares about CivetWeb) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:46 +0200
                      Graphics Processing with Fortran 77 (was: Re: Your strictness is your problem , not mine) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-07-29 18:10 +0800
                        I am not in C, it is theory and C++ [Hybrid Approaches from KOAN/Fortran-S] (Was: Graphics Processing with Fortran 77) Mild Shock <janburse@fastmail.fm> - 2026-07-29 12:43 +0200
                          Java picky concerning JIT-ing [Luckier with C++/C or FORTRAN compilers?] (Re: I am not in C, it is theory and C++) Mild Shock <janburse@fastmail.fm> - 2026-07-29 12:53 +0200
                  Run with minimum HTTPS and .mjs type (Re: In AI Acceleration nobody cares about CivetWeb) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:48 +0200
                    Lamas in a cradle and Lamas on the edge [Red Pyjama] (Was: Run with minimum HTTPS and .mjs type) Mild Shock <janburse@fastmail.fm> - 2026-07-29 13:05 +0200
                      Synthetic Multilanguage Autoformalization Dataset [Informath project] (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-08-08 09:22 +0200
                        Six Proofs and Generally Inteligent Systems [EyeProlog Pseudo Scientism] (Re: Synthetic Multilanguage Autoformalization Dataset [Informath project]) Mild Shock <janburse@fastmail.fm> - 2026-08-15 15:20 +0200
                          Everybody does eat and sleep [The SK hynix Story] (Re: Six Proofs and Generally Inteligent Systems [EyeProlog Pseudo Scientism]) Mild Shock <janburse@fastmail.fm> - 2026-08-15 18:49 +0200
                Re: confused rossy boy is confused (Was: Clueless about MIMD as usual [Flynn's Taxonomy]) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-29 14:42 -0700
                  Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:37 +0200
                    Re: Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-02 14:40 -0700
                      Re: Even send_color and recv_color can block [Cerebras Waver] (Was: confused rossy boy is confused) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-02 14:43 -0700
                      You don't understand that compute shaders are tasks (Was: Even send_color and recv_color can block [Cerebras Waver]) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:45 +0200
                      You don't understand that compute shaders are tasks (Re: Even send_color and recv_color can block [Cerebras Waver]) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:46 +0200
                        Ignoramus / Ignorabimus Barometer: Almost 1 Month (Re: You don't understand that compute shaders are tasks) Mild Shock <janburse@fastmail.fm> - 2026-08-03 00:09 +0200
                          Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month) Mild Shock <janburse@fastmail.fm> - 2026-08-03 02:08 +0200
                            Re: Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 11:55 -0700
                              You are not correctly thinking (Was: Homework: Game Engine in WebGPU) Mild Shock <janburse@fastmail.fm> - 2026-08-03 21:04 +0200
                                Re: You are not correctly thinking (Was: Homework: Game Engine in WebGPU) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 12:32 -0700
                                  You don't understand the economy of an AI Laptop (Was: You are not correctly thinking) Mild Shock <janburse@fastmail.fm> - 2026-08-03 22:24 +0200
                                    You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) Mild Shock <janburse@fastmail.fm> - 2026-08-03 22:37 +0200
                                      Re: You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop ) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 14:29 -0700
                                        Know nothing and forget what you posted day before (Was: You don't understand producer , workers , consumer) Mild Shock <janburse@fastmail.fm> - 2026-08-03 23:38 +0200
    Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-30 06:49 -0700
      Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-07-30 14:55 -0700
      Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 12:55 -0700
        Re: Viswath & Charmaigne (vector-wide scalar-word and character machines) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 13:05 -0700
    Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:14 +0200
      Re: Ljubljana School versus Zurich School (Was: Viswath & Charmaigne) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 02:21 +0800
        pi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:35 +0200
          A spinlock rewrite will be necessary (Was: pi-WAM uses ADA RendezVous) Mild Shock <janburse@fastmail.fm> - 2026-08-03 21:00 +0200
      Big thanks to Ljubljana School [Searching 0xCAFFEE] (Was: Ljubljana School versus Zurich School) Mild Shock <janburse@fastmail.fm> - 2026-08-03 20:40 +0200

Page 6 of 6 — ← Prev page 1 2 3 4 5 [6]


#124541 — You don't understand that compute shaders are tasks (Was: Even send_color and recv_color can block [Cerebras Waver])

FromMild Shock <janburse@fastmail.fm>
Date2026-08-02 23:45 +0200
SubjectYou don't understand that compute shaders are tasks (Was: Even send_color and recv_color can block [Cerebras Waver])
Message-ID<114odp8$rl67$1@solani.org>
In reply to#124539
Hi,

You are a moron. In WebGPU computer sharers
are tasks not hardware kernels. Forget your
WebGL nonsense cookbooks.

WebGPU is much more elastic.

You are just a moron.

Bye

P.S.: Take this example, I don't have 4096 kernels:

1.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Still it runs, how is this done? The Ryzen has
only around 512 kernels. Newer Ryzen havae 1024
kernels. This is till below 4096 logical threads.

So how is it done?

Chris M. Thomasson schrieb:
> On 8/2/2026 2:37 PM, Mild Shock wrote:
>> Hi,
>>
>> If only the fucking moron Chris M. Thomasson would
>> stop spamming his nonsense, he doesn't listen at
>> all. Problem, he cannot read, he knows nothing.
>>
>> Its very common that compute shaders can block,
>> when they are used for General Purpose computation
>> on GPUs (GPGPU). If only he would pull out his
>>
>> finger from his asshole, and stop thinking in his
>> WebGL legacy code stash nonsense. Even the
>> Cerebras Waver has blocking:
>>
>> "Cerebras Software Language (CSL), send_color
>> and recv_color are parameters passed to tile
>> programs to manage data routing and virtual
>> channels (called colors) across processing
>> elements (PEs) on the wafer
>>
>> Yes, both send and receive operations can block
>> on a Cerebras Processing Element (PE), primarily
>> due to the system's hardware-enforced backpressure
>> mechanism. Because the Cerebras Wafer-Scale Engine
>> (WSE) relies on a fine-grained,
>>
>> dataflow-driven architecture, blocking prevents
>> data loss when hardware resources are
>> fully saturated."
>>
>> Blocking and Unblocking
>> https://sdk.cerebras.ai/computing-with-cerebras#blocking-and-unblocking
> 
> Strive to never make a compute shader wait on something, like an empty 
> condition of a queue, stack.
> 
> 
>> Chris M. Thomasson is an annoyance and an idiot.
>> He is a total waste of time. And represents those
>> people who cannot use their brain.
> 
> I don't think you have coded compute shaders before? If so, cool, but wow.
> 
> [...]

[toc] | [prev] | [next] | [standalone]


#124542 — You don't understand that compute shaders are tasks (Re: Even send_color and recv_color can block [Cerebras Waver])

FromMild Shock <janburse@fastmail.fm>
Date2026-08-02 23:46 +0200
SubjectYou don't understand that compute shaders are tasks (Re: Even send_color and recv_color can block [Cerebras Waver])
Message-ID<114odrg$rl67$2@solani.org>
In reply to#124539
Hi,

You are a moron. In WebGPU computer sharers
are tasks not hardware kernels. Forget your
WebGL nonsense cookbooks.

WebGPU is much more elastic.

You are just a moron.

Bye

P.S.: Take this example, I don't have 4096 kernels:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Still it runs, how is this done? The Ryzen has
only around 512 kernels. Newer Ryzen havae 1024
kernels. This is till below 4096 logical threads.

So how is it done?


Chris M. Thomasson schrieb:
> On 8/2/2026 2:37 PM, Mild Shock wrote:
>> Hi,
>>
>> If only the fucking moron Chris M. Thomasson would
>> stop spamming his nonsense, he doesn't listen at
>> all. Problem, he cannot read, he knows nothing.
>>
>> Its very common that compute shaders can block,
>> when they are used for General Purpose computation
>> on GPUs (GPGPU). If only he would pull out his
>>
>> finger from his asshole, and stop thinking in his
>> WebGL legacy code stash nonsense. Even the
>> Cerebras Waver has blocking:
>>
>> "Cerebras Software Language (CSL), send_color
>> and recv_color are parameters passed to tile
>> programs to manage data routing and virtual
>> channels (called colors) across processing
>> elements (PEs) on the wafer
>>
>> Yes, both send and receive operations can block
>> on a Cerebras Processing Element (PE), primarily
>> due to the system's hardware-enforced backpressure
>> mechanism. Because the Cerebras Wafer-Scale Engine
>> (WSE) relies on a fine-grained,
>>
>> dataflow-driven architecture, blocking prevents
>> data loss when hardware resources are
>> fully saturated."
>>
>> Blocking and Unblocking
>> https://sdk.cerebras.ai/computing-with-cerebras#blocking-and-unblocking
> 
> Strive to never make a compute shader wait on something, like an empty 
> condition of a queue, stack.
> 
> 
>> Chris M. Thomasson is an annoyance and an idiot.
>> He is a total waste of time. And represents those
>> people who cannot use their brain.
> 
> I don't think you have coded compute shaders before? If so, cool, but wow.
> 
> [...]

[toc] | [prev] | [next] | [standalone]


#124544 — Ignoramus / Ignorabimus Barometer: Almost 1 Month (Re: You don't understand that compute shaders are tasks)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 00:09 +0200
SubjectIgnoramus / Ignorabimus Barometer: Almost 1 Month (Re: You don't understand that compute shaders are tasks)
Message-ID<114of69$rlsh$3@solani.org>
In reply to#124542
Hi,

This was archived on Jul 9, 2026:

 > 11.4 Giga Lips with a Budget Laptop
 > https://github.com/Jean-Luc-Picard-2021/gigabudget

Still today on Aug 03, 2026, the usenet
community still struggles with the experiment,
doesn't know the meaning and implications,

especially clueless about 4096 shaders and
modern GPU elasticity. Woa! Thats impressive.
Especially Chris M. Thomasson has a still ongoing

hard time with this little WebGPU experiment.

Bye

Mild Shock schrieb:
> Hi,
> 
> You are a moron. In WebGPU computer sharers
> are tasks not hardware kernels. Forget your
> WebGL nonsense cookbooks.
> 
> WebGPU is much more elastic.
> 
> You are just a moron.
> 
> Bye
> 
> P.S.: Take this example, I don't have 4096 kernels:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Still it runs, how is this done? The Ryzen has
> only around 512 kernels. Newer Ryzen havae 1024
> kernels. This is till below 4096 logical threads.
> 
> So how is it done?
> 
> 
> Chris M. Thomasson schrieb:
>> On 8/2/2026 2:37 PM, Mild Shock wrote:
>>> Hi,
>>>
>>> If only the fucking moron Chris M. Thomasson would
>>> stop spamming his nonsense, he doesn't listen at
>>> all. Problem, he cannot read, he knows nothing.
>>>
>>> Its very common that compute shaders can block,
>>> when they are used for General Purpose computation
>>> on GPUs (GPGPU). If only he would pull out his
>>>
>>> finger from his asshole, and stop thinking in his
>>> WebGL legacy code stash nonsense. Even the
>>> Cerebras Waver has blocking:
>>>
>>> "Cerebras Software Language (CSL), send_color
>>> and recv_color are parameters passed to tile
>>> programs to manage data routing and virtual
>>> channels (called colors) across processing
>>> elements (PEs) on the wafer
>>>
>>> Yes, both send and receive operations can block
>>> on a Cerebras Processing Element (PE), primarily
>>> due to the system's hardware-enforced backpressure
>>> mechanism. Because the Cerebras Wafer-Scale Engine
>>> (WSE) relies on a fine-grained,
>>>
>>> dataflow-driven architecture, blocking prevents
>>> data loss when hardware resources are
>>> fully saturated."
>>>
>>> Blocking and Unblocking
>>> https://sdk.cerebras.ai/computing-with-cerebras#blocking-and-unblocking
>>
>> Strive to never make a compute shader wait on something, like an empty 
>> condition of a queue, stack.
>>
>>
>>> Chris M. Thomasson is an annoyance and an idiot.
>>> He is a total waste of time. And represents those
>>> people who cannot use their brain.
>>
>> I don't think you have coded compute shaders before? If so, cool, but 
>> wow.
>>
>> [...]
> 

[toc] | [prev] | [next] | [standalone]


#124545 — Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 02:08 +0200
SubjectHomework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month)
Message-ID<114om5c$s9h3$3@solani.org>
In reply to#124544
Hi,

Now that the debate with Chris M. Thomasson
has culminated in questions of elasticity,
I suggest this homework:

- Game Engine in WebGPU
   It will support the life cycle of sprites,
   like sprites comming out of nowhere,
   and being destroyed by arms,
   just like in Space invader.

This would be surely a fantastic exercise,
to see what a GPU can do and cannot do,
in respect of life cycle of threads, especially

modern GPUs that sell the CUDA dream.

Have Fun!

Become a nosomatic AI chirurgeon.

Bye

Mild Shock schrieb:
 > Hi,
 >
 > A nosomatic AI chirurgeon is a halfling student
 > of sickness, and a master of the ebb and flow of
 > the energies of life and death of data packets.
 >
 > He is a air bender, water bender and earth bender
 > in one person, using OpenVINO to juggle with
 > CPU, GPU and NPU.
 >
 > Last but not least he can freely switch between
 > symbolic and neural representation of knowledge
 > forms, there is no abyss for him.
 >
 > Bye

Mild Shock schrieb:
> Hi,
> 
> This was archived on Jul 9, 2026:
> 
>  > 11.4 Giga Lips with a Budget Laptop
>  > https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Still today on Aug 03, 2026, the usenet
> community still struggles with the experiment,
> doesn't know the meaning and implications,
> 
> especially clueless about 4096 shaders and
> modern GPU elasticity. Woa! Thats impressive.
> Especially Chris M. Thomasson has a still ongoing
> 
> hard time with this little WebGPU experiment.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> You are a moron. In WebGPU computer sharers
>> are tasks not hardware kernels. Forget your
>> WebGL nonsense cookbooks.
>>
>> WebGPU is much more elastic.
>>
>> You are just a moron.
>>
>> Bye
>>
>> P.S.: Take this example, I don't have 4096 kernels:
>>
>> 11.4 Giga Lips with a Budget Laptop
>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> Still it runs, how is this done? The Ryzen has
>> only around 512 kernels. Newer Ryzen havae 1024
>> kernels. This is till below 4096 logical threads.
>>
>> So how is it done?
>>
>>
>> Chris M. Thomasson schrieb:
>>> On 8/2/2026 2:37 PM, Mild Shock wrote:
>>>> Hi,
>>>>
>>>> If only the fucking moron Chris M. Thomasson would
>>>> stop spamming his nonsense, he doesn't listen at
>>>> all. Problem, he cannot read, he knows nothing.
>>>>
>>>> Its very common that compute shaders can block,
>>>> when they are used for General Purpose computation
>>>> on GPUs (GPGPU). If only he would pull out his
>>>>
>>>> finger from his asshole, and stop thinking in his
>>>> WebGL legacy code stash nonsense. Even the
>>>> Cerebras Waver has blocking:
>>>>
>>>> "Cerebras Software Language (CSL), send_color
>>>> and recv_color are parameters passed to tile
>>>> programs to manage data routing and virtual
>>>> channels (called colors) across processing
>>>> elements (PEs) on the wafer
>>>>
>>>> Yes, both send and receive operations can block
>>>> on a Cerebras Processing Element (PE), primarily
>>>> due to the system's hardware-enforced backpressure
>>>> mechanism. Because the Cerebras Wafer-Scale Engine
>>>> (WSE) relies on a fine-grained,
>>>>
>>>> dataflow-driven architecture, blocking prevents
>>>> data loss when hardware resources are
>>>> fully saturated."
>>>>
>>>> Blocking and Unblocking
>>>> https://sdk.cerebras.ai/computing-with-cerebras#blocking-and-unblocking
>>>
>>> Strive to never make a compute shader wait on something, like an 
>>> empty condition of a queue, stack.
>>>
>>>
>>>> Chris M. Thomasson is an annoyance and an idiot.
>>>> He is a total waste of time. And represents those
>>>> people who cannot use their brain.
>>>
>>> I don't think you have coded compute shaders before? If so, cool, but 
>>> wow.
>>>
>>> [...]
>>
> 

[toc] | [prev] | [next] | [standalone]


#124552 — Re: Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month)

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-08-03 11:55 -0700
SubjectRe: Homework: Game Engine in WebGPU (Re: Ignoramus / Ignorabimus Barometer: Almost 1 Month)
Message-ID<114qo7e$1ip48$1@dont-email.me>
In reply to#124545
On 8/2/2026 5:08 PM, Mild Shock wrote:
[...]
>>>> I don't think you have coded compute shaders before? If so, cool, 
>>>> but wow.

Never mind. You are too hostile. Not worth it. Sorry. Plonk.

[toc] | [prev] | [next] | [standalone]


#124554 — You are not correctly thinking (Was: Homework: Game Engine in WebGPU)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 21:04 +0200
SubjectYou are not correctly thinking (Was: Homework: Game Engine in WebGPU)
Message-ID<114qoo8$t9up$1@solani.org>
In reply to#124552
Hi,

You are not correctly thinking.
I am not using WebGL. I use WebGPU.
Spinning is perfectly fine. I will

soon give proof. Meanwhile enjoy
this use case, so that you understand
the goal of Prolog "inferencing" for

a simple example:

"We try to find 0xCAFFEE in enumerating 4
6-bit digits and the baseline is Dogelog
Player VM in a browser. The CPU backend
with 64 logical threads is already 20
times faster, partly due to its 32-bit
specialization. The GPU backend with
4096 logical threads boosts a further
factor of 7 times."

GPU Backend: Find 0xCAFFEE with π-WAM
https://medium.com/2989/8890efd3503c

If you don't understand the goal, and
the benefits of the goal, all your
thinking will anyways be incorrect.

Bye

Chris M. Thomasson schrieb:
> On 8/2/2026 5:08 PM, Mild Shock wrote:
> [...]
>>>>> ***I don't think** you have coded compute 
>>>>> shaders before? If so, cool, but wow.
> 
> Never mind. You are too hostile. Not worth it. Sorry. Plonk.

[toc] | [prev] | [next] | [standalone]


#124555 — Re: You are not correctly thinking (Was: Homework: Game Engine in WebGPU)

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-08-03 12:32 -0700
SubjectRe: You are not correctly thinking (Was: Homework: Game Engine in WebGPU)
Message-ID<114qqc3$1jea5$1@dont-email.me>
In reply to#124554
On 8/3/2026 12:04 PM, Mild Shock wrote:
> Hi,
> 
> You are not correctly thinking.
> I am not using WebGL. I use WebGPU.
> Spinning is perfectly fine. I will

Wait... Before I totally plonk... Spinning is fine in a compute shader? 
Really? If so my FIFO queue fetch-add-only tweak from Dimity's would 
work fine. Also, Dmitry's CAS based one is good as well. My tweak 
version of his have different tradeoffs... I personally would not want 
to use any of them in a compute shader, never spin and/or wait! Strive 
for it, really hard, first... But, well, does your system have "waiting 
primitives" so you don't have to spin? Also, if you do spin you need 
some sort of backoff, right? Aka PAUSE on x86, etc... Or notice in my 
FIFO one can take the ticket and spin on it later as in a backoff is 
doing other real work.

Akin to my special mutex pattern that can be found here in this group. 
Iirc the thread is entitled:

fun with a mutex...



So, I am using dirextc12 and modern opengl for my compute shaders right 
now. GLSL as my lang. I need to provide some state for them to work 
with. Aka, textures and uniforms.




> 
> soon give proof. Meanwhile enjoy
> this use case, so that you understand
> the goal of Prolog "inferencing" for
> 
> a simple example:
> 
> "We try to find 0xCAFFEE in enumerating 4
> 6-bit digits and the baseline is Dogelog
> Player VM in a browser. The CPU backend
> with 64 logical threads is already 20
> times faster, partly due to its 32-bit
> specialization. The GPU backend with
> 4096 logical threads boosts a further
> factor of 7 times."
> 
> GPU Backend: Find 0xCAFFEE with π-WAM
> https://medium.com/2989/8890efd3503c
> 
> If you don't understand the goal, and
> the benefits of the goal, all your
> thinking will anyways be incorrect.
> 
> Bye
> 
> Chris M. Thomasson schrieb:
>> On 8/2/2026 5:08 PM, Mild Shock wrote:
>> [...]
>>>>>> ***I don't think** you have coded compute shaders before? If so, 
>>>>>> cool, but wow.
>>
>> Never mind. You are too hostile. Not worth it. Sorry. Plonk.
> 

Fun with a mutex:


(read all...)
____________________________________
// A Fun Mutex Pattern? Or, a Nightmare? Humm...
// By: Chris M. Thomasson
//___________________________________________________


#include <iostream>
#include <random>
#include <numeric>
#include <algorithm>
#include <thread>
#include <atomic>
#include <mutex>


#define CT_WORKERS (42)
#define CT_ITERS (996699)
#define CT_BACKOFFS (42)
#define CT_RAND_MAX (20)
#define CT_RAND_THRESHOLD (5)


struct ct_shared
{
     std::mutex m_fun_mutex;
     std::atomic<unsigned long> m_other_work = { 0 };
     int m_test_count0 = 0;

     void
     sanity_check_dump() const
     {
         std::cout << "(ct_shared:" << this << ")->" <<
                      "m_test_count0 = " << m_test_count0 << ", " <<
                      "m_other_work = " << 
m_other_work.load(std::memory_order_relaxed) << "\n";
     }

     bool
     sanity_check_validate() const
     {
         return (m_test_count0 == CT_ITERS * CT_WORKERS);
     }
};



void
ct_worker_entry(
     ct_shared& shared
) {
     //std::cout << "ct_worker_entry" << std::endl; // testing thread 
race for sure...

     // Thread Local...
     std::random_device rnd_seed = { };
     std::mt19937 rnd_gen(rnd_seed());
     std::uniform_int_distribution<unsigned long> rnd_dist(0, CT_RAND_MAX);

     for (unsigned long i = 0; i < CT_ITERS; ++i)
     {
         // Lock logic...
         {
             unsigned long backoff = 0;

             while (! shared.m_fun_mutex.try_lock())
             {
                 unsigned long rnd0 = rnd_dist(rnd_gen);

                 if (rnd0 > CT_RAND_THRESHOLD || backoff > CT_BACKOFFS)
                 {
                     shared.m_fun_mutex.lock();
                     break;
                 }

                 // do other work... :^)
                 shared.m_other_work.fetch_add(1, 
std::memory_order_relaxed);

                 // but not too much work... ;^o
                 ++backoff;
             }
         }

             // Critical Section...
             {
                 shared.m_test_count0 = shared.m_test_count0 + 1;
             }

         // Unlock
         {
             shared.m_fun_mutex.unlock();
         }
     }
}


int main()
{
     // Hello... :^)
     {
         std::cout << "Hello ct_fun_mutex... lol? ;^) ver:(0.0.0)\n";
         std::cout << "By: Chris M. Thomasson\n";
         std::cout << 
"____________________________________________________\n";
         std::cout.flush();
     }

     // Create our fun things... ;^)
     ct_shared shared = { };
     std::thread workers[CT_WORKERS] = { };

     // Lanuch...
     {
         std::cout << "Launching Threads...\n";
         std::cout.flush();

         for (unsigned long i = 0; i < CT_WORKERS; ++i)
         {
             workers[i] = std::thread(ct_worker_entry, std::ref(shared));
         }
     }

     // Join...
     {
         std::cout << "Joining Threads... (computing :^)\n";
         std::cout.flush();
         for (unsigned long i = 0; i < CT_WORKERS; ++i)
         {
             workers[i].join();
         }
     }

     // Sanity Check...
     {
         shared.sanity_check_dump();

         if (! shared.sanity_check_validate())
         {
             std::cout << "\n\n**** Pardon my French, but FUCK!!!!! 
****\n" << std::endl;
         }

         else
         {
             std::cout << "\nWe are Sane!\n\n";
             std::cout << "We completed " <<
                 shared.m_other_work.load(std::memory_order_relaxed) <<
                 " work items while waiting for the mutex..." << std::endl;
         }
     }

     // Fin...
     {
         std::cout << 
"____________________________________________________\n";
         std::cout << "Fin... :^)\n" << std::endl;
     }

     return 0;
}
____________________________________

Any luck? Its fun to see how many work items were completed when the 
mutex was contended...

[toc] | [prev] | [next] | [standalone]


#124556 — You don't understand the economy of an AI Laptop (Was: You are not correctly thinking)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 22:24 +0200
SubjectYou don't understand the economy of an AI Laptop (Was: You are not correctly thinking)
Message-ID<114qtdv$td35$1@solani.org>
In reply to#124555
Hi,

Why do you even open your mouth if you
don't use WebGPU / WGSL? This beyond my
comprehension. OpenGL was phased out by

Apple years ago. It only lives on some
linux boxes. Also you probably don't use
an AI Laptop. Just make a simple calculation,

if you have 512 Kernels, and oversubscribe
4096 logical threads. Then each Kernel runs
4 logical threads. If one of these 4 logical

threads spins, how much performance is lost?
25% of this single kernel. And there are
still 511 Kernels. Spinning is totally fine,

thats why WGSL provides CAS, and not some
waitlists. The kernels are the wait lists itself
doing the following when spinning:

NOP
NOP
NOP
Etc..

Until the a condition is met. You even don't
need backoff, because you cannot pause. The
only pause you can do is a barrier.

But if the condition is not met while the
barrier is met, what will you do?

Bye

Chris M. Thomasson schrieb:
> So, I am using dirextc12 and modern opengl for my 
> compute shaders right  now. GLSL as my lang. I need 
> to provide some state for them to work 
> with. Aka, textures and uniforms.

[toc] | [prev] | [next] | [standalone]


#124557 — You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop )

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 22:37 +0200
SubjectYou don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop )
Message-ID<114qu69$tdm6$1@solani.org>
In reply to#124556
Hi,

I you use atomicAdd() you have the same friction
as if you use Queue put() or take(). There is
no difference. The only difference is unbounded

versus bounded. I tried to explain that like
100-times already. Your comment here:

 > Any luck? Its fun to see how many work
 > items were completed when the mutex was contended...

Says to me you don't understand queues. They
are not mutexes. Because you don't understand
queues, you also don't understand OpenMP

parallelism and patterns such as producer,
workers, consumer. Contention is usually minimal,
the workers just fetch work items from the

producer, and then do some workload. And
then hand the result to the consumer. If
you use atomicAdd() you have the same friction

as if you use Queue put() or take(). There
is no difference. The only difference is unbounded
versus bounded. I tried to explain that

like 100-times already.

Bye

Mild Shock schrieb:
> Hi,
> 
> Why do you even open your mouth if you
> don't use WebGPU / WGSL? This beyond my
> comprehension. OpenGL was phased out by
> 
> Apple years ago. It only lives on some
> linux boxes. Also you probably don't use
> an AI Laptop. Just make a simple calculation,
> 
> if you have 512 Kernels, and oversubscribe
> 4096 logical threads. Then each Kernel runs
> 4 logical threads. If one of these 4 logical
> 
> threads spins, how much performance is lost?
> 25% of this single kernel. And there are
> still 511 Kernels. Spinning is totally fine,
> 
> thats why WGSL provides CAS, and not some
> waitlists. The kernels are the wait lists itself
> doing the following when spinning:
> 
> NOP
> NOP
> NOP
> Etc..
> 
> Until the a condition is met. You even don't
> need backoff, because you cannot pause. The
> only pause you can do is a barrier.
> 
> But if the condition is not met while the
> barrier is met, what will you do?
> 
> Bye
> 
> Chris M. Thomasson schrieb:
>> So, I am using dirextc12 and modern opengl for my compute shaders 
>> right  now. GLSL as my lang. I need to provide some state for them to 
>> work with. Aka, textures and uniforms.

[toc] | [prev] | [next] | [standalone]


#124560 — Re: You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop )

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-08-03 14:29 -0700
SubjectRe: You don't understand producer , workers , consumer (Was: You don't understand the economy of an AI Laptop )
Message-ID<114r181$1lmn5$2@dont-email.me>
In reply to#124557
On 8/3/2026 1:37 PM, Mild Shock wrote:
> Hi,
> 
> I you use atomicAdd() you have the same friction
> as if you use Queue put() or take(). There is
> no difference. The only difference is unbounded
> 
> versus bounded. I tried to explain that like
> 100-times already. Your comment here:
> 
>  > Any luck? Its fun to see how many work
>  > items were completed when the mutex was contended...
> 
> Says to me you don't understand queues. They
> are not mutexes. Because you don't understand
> queues, you also don't understand OpenMP
> 
> parallelism and patterns such as producer,
> workers, consumer. Contention is usually minimal,
> the workers just fetch work items from the
> 
> producer, and then do some workload. And
> then hand the result to the consumer. If
> you use atomicAdd() you have the same friction
> 
> as if you use Queue put() or take(). There
> is no difference. The only difference is unbounded
> versus bounded. I tried to explain that[...]

lol. I forgot to add you to my killfile. Damn it! Anyway, I know all 
about them. Sigh. Peace be with you.

[toc] | [prev] | [next] | [standalone]


#124564 — Know nothing and forget what you posted day before (Was: You don't understand producer , workers , consumer)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 23:38 +0200
SubjectKnow nothing and forget what you posted day before (Was: You don't understand producer , workers , consumer)
Message-ID<114r1no$tfll$4@solani.org>
In reply to#124560
Hi,

Know nothing and forget what you posted
day before. You are the most unfocused
idiotic liar and spammer I have ever met.

Maybe produce some results or shut up!

Bye

Chris M. Thomasson schrieb:
> On 8/3/2026 1:37 PM, Mild Shock wrote:
>> Hi,
>>
>> I you use atomicAdd() you have the same friction
>> as if you use Queue put() or take(). There is
>> no difference. The only difference is unbounded
>>
>> versus bounded. I tried to explain that like
>> 100-times already. Your comment here:
>>
>>  > Any luck? Its fun to see how many work
>>  > items were completed when the mutex was contended...
>>
>> Says to me you don't understand queues. They
>> are not mutexes. Because you don't understand
>> queues, you also don't understand OpenMP
>>
>> parallelism and patterns such as producer,
>> workers, consumer. Contention is usually minimal,
>> the workers just fetch work items from the
>>
>> producer, and then do some workload. And
>> then hand the result to the consumer. If
>> you use atomicAdd() you have the same friction
>>
>> as if you use Queue put() or take(). There
>> is no difference. The only difference is unbounded
>> versus bounded. I tried to explain that[...]
> 
> lol. I forgot to add you to my killfile. Damn it! Anyway, I know all 
> about them. Sigh. Peace be with you.

[toc] | [prev] | [next] | [standalone]


#124464

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-30 06:49 -0700
Message-ID<EAmdnf06E6Hly_b3nZ2dnZfqnPqdnZ2d@giganews.com>
In reply to#124392
On 07/27/2026 11:45 AM, Ross Finlayson wrote:
> On 07/27/2026 11:44 AM, Ross Finlayson wrote:
>> On 07/27/2026 11:43 AM, Ross Finlayson wrote:
>>> Hello, here I'll post some design notes and a panel discussion with some
>>> chat-bots about making some sense of the "vector-wide scalar word"
>>> and "character machines", on commodity hardware about ubiquitous
>>> operations.
>>>
>>>
>>> It's considered at least tangentially relevant to comp.lang.c and
>>> comp.lang.c++ because for example text is ubiquitous and the targets
>>> would be low-level, while the higher-level languages would have a
>>> same sort of patternry, and for example that libc and cstdlib are
>>> standard, and as with regards to POSIX and Unicode and so on.
>>>
>>> Please feel free to excuse or ignore, or comment as freely.
>>>
>>> Thanks for reading.
>>>
>>
>>
>> [ viswath-charmaigne.txt ]
>>
>>
>>


[ viswath-charmaigne-20270727_b.txt ]

About smearing and unsmearing, it's figured to make for
"smear-detection" and "smear-correction", and for the
"unsmear-detection" and
"unsmear-correction", basically that smearing is indicated by variously:

multiple-byte characters
escape characters and translated characters
control-characters with payloads/bodies

with mostly the case being multiple-byte and escape-translations.

The idea of detection and correction is about comprehension and
expression, about what comprehensions, or classifications, occur,
according to what expressions, have as their implicits the contexts.

So, it's figured that it starts with bytes, then, for source text, first
there are the main or base classes, alnum/punct/white/coded, then, for
coded, it's to be established whether those are non-printable control
characters, which mostly are to be avoided or invalidated unless there
are particular comprehensible payloads representing sub-expressions, or
they're UTF-8 codepoints, which is figured to be the default.


ASCII -> UTF-8?
UCS2 -> BE|LE +BOM? -> UTF-16
UCS2 -> UTF-16?

Then, the idea is that first the source-main class is applied, or, about
there being a proto-class that's "coded and non-coded", and for example
about line-breaks or otherwise field-separators and record-separators.


So, it's figured that for "source" languages it's ASCII-centric, so the
base character classes are loaded first, then the smear/unsmear for
UTF-8 or otherwise the multi-byte is ASCII-peripheral, then that UCS-2
got UTF-16 has a similar account with regards to the smashing,

https://www.autoitconsulting.com/site/development/utf-8-utf-16-text-encoding-detection-library/

(An article suggests to detect UCS2/UTF-16 by looking for the
Byte-Order-Marker, then for newlines, then for a preponderance of ASCII
characters.)

https://en.wikipedia.org/wiki/Charset_detection



So, then presuming UTF-8, then gets back to figuring out smearing and
straddling of smearing, about that UTF-8 bytes get smeared and the masks
for their predicates also get smeared, then when they straddle the
codes-themselves, that the context of the character is carried across
the boundary (splitting/stitching).


About the control-characters, then these are for example the "DEC VT" or
"ECMA-48", "ISO 6429", "DEC STD 070", like from "XTerm control
sequences" by Moy, Gildea, and Dickey, mostly to be avoided, yet
variously where anything that's not a "single-character function", is to
be avoided, and that since SPACE, TAB, NL, CR, FF, VT are considered
white-space not coded, has that coded characters make for invalidation,
though there's a simple enough account that the data following
control-characters with parameters in sequences are detectable.

So, coded/ nybbles are first:

alnum/
punct/
white/

coded/ctrl
coded/utf8
coded/nul
coded/bom

Then, a first-pass over the buffer is always starting with context of
the straddle-stitching whether a UTF-8 character or what kind of
control character its sequence is at what state, that what gets derived
for UTF-8 characters as secondary is either a nybble with the
count-total and count-remaining, or, count-encountered and count-remaining.

1
2
3
4


When straddling, it's un-known whether there are remaining bytes,
about basically to have a separate part of the nybble for the straddle

straddling/
split/
stitching/

The idea is that the smear/unsmearing is indicated by the word, for
the properties, then that for the code-point, that's inserted with
the stitching, about that

splitting is only at the end of a word, and
stitching is only at the beginning of a word

for forward search.

So, first the main class is determined, then, conditioned on whether
there exists either a "max-length" or a null character is the
End-of-Input, and conditioned on whether there's a "Start-of-Input" offset,
about offsets and extents, the main class is determined from the
Start-of-Input (usually somewhere in the initial word) and End-of-Input,
then making the lookup of the main class.

Another point of straddle and splitting and stitching is for the fixed
match case, while it's usually figured that the fixed string being
matched fits within a word, arbitrarily it crosses multiple words or is
more than word length, then that when there's an initial-segment match,
to be matching the trailing-segment. So, in splitting UTF-8 codes, it's
known that the code extends, yet not how far, yet in splitting fixed
strings, it's known that the initial-segment matches, not if the
trailing-segment matches.


Then, matching the "fixed" also gets into matching more
widely, about the expressions and grammars. From taking
a look into outlines of Hyperscan and Vectorscan (regex and
multiple-regex matching engines employing vector techniques
from Intel and ARM respectively), there are notions of the
"decomposition" of expressions, then about what's promontory
and matching the "fixed", first fixed-length then fixed-content,
when matching what would be "longest sub-matches", then
to recursively bridge the definite sub-matches.


So, the context of the findings and matchings start to develop,
with the idea that by the presence in the context, that actions
occur, otherwise for nothing or no-ops.

Afore-Input: Start-of-Input, at the beginning of a "walk", and beginning
of a "word"
Afore-Stitch: at the beginning of a word, there's stitching to occur

After-Split: at the end of a word, there's definitely/possibly a splot
After-Input: End-of-Input, at the end of a "walk", and end of a "word".


Here "walk" has the usual notions of "tree-traversals", that instead
here "walk" (or "work") is the notion here of the sequence action,
then for "work". Then "Afore" and "After", or "Before" and "Behind",
make for that they're same-length identifiers and also that they're
in the same lexicographic order.

Before-Stitch
Behind-Split

Afore-Stitch
After-Split

Among-Straddle (Among, Amidst)


So, the context then is for register state and stack contents, that
the indicators of the above as "positive presence" then is to make
for that the adjustments to the offsets and extents and the shifts
is according to those, otherwise no-ops. Then the idea is that a
"working" starts with a given context according to the expression,
then that as various of the "findings" make findings, they push either
context to act on the stack, or no-ops on the stack, then the stack
results being a fixed-size for the working according to the expression,
then the actions are always popping off a fixed amount of actions
and no-ops, with no branching, just computed "presence".



1) work starts
compute any misalignment / Start-of-Input
load word (or bytes-into-word when no-misaligned-loads)

2) word starts

(resolve startings)
(resolve endings)
(resolve stitches)

lookup/load main class
find coded
find splits
find UTF-8
find cntrl

lookup expression/grammar classes
find

(resolve splits)
(resolve straddles, byte-straddles, word-straddles)


The idea is that the predicates (properties/predicates or
code-points/range-points), are to get shifted and trimmed,
or initialized, shifted, and trimmed, so that it results the trimmings
or truncations, then have that the properties/predicates
or code-points/range-points will result matches in what results
of the initialized, shifted, and trimmed.

1) initialize (copy) the predicate/range-points
2) shift to find-start, find-continue
3) trim about the offset, extent
4) find-continue

About code-points/range-points, what's figured is that
it's always inclusive the bounds of the range, then that
the matching of a single code-point is always the matching
of two range-points that happen to be equal, so that matching
either a code-point or a range, is the same operation,
that:
not-less-than-lower && not greater-than-upper
which makes finding of range-points, also works for code-points.


So, the usual idea is that there are the various findings occurring,

find-longest-match:
shift and repeat byte-wise across the word

find-nearest-exit:

find-near:
find-far:


Then, for an expression or expressions, and grammar or grammars,
is the idea of making multi-matches, that the idea is that each of
the possibles make their exercise, and then to result after the word
is worked by each of the sub-expressions, to collate the results, or
to emit the results, then onto the next word.

Basically there is a difference among productions about whether matching
or finding is among "alternatives" or "potentials", with the idea that
matching "alternatives" is vertical while matching "potentials" is
horizontal, that a finding in terms of the NFA/DFA basically enters
either an "arc" or a "transition", that an "arc" is in the "potentials"
to make a "plant" of the "potential plant", vis-a-vis the arcs/plants
and transitions/states.

Then, an alternative has matching the first character, then whether it
introduces a potential, about that the single-character matches then
as for "double-bracket" or "triple-quote", make for that those sorts of
potentials are as according to the bracketed/quoted/escaped
expressions/grammars, to be defining the rules of the machine.



finding potentials then is about this sort of account:

the word is N-many bytes wide

property/predicate: 1 register property, 1 register predicate -> 1
register indicators
codepoint/rangepoint: 1 register codepoint, 2 registers rangepoints -> 1
register indicators

union of findings: 2 registers indicators, 1 register indicators
intersection of findings: 2 registers indicators, 1 register indicators
setminus: ...
complement

The finding then has either a "required" or "optional" next item, when
it's in finding potentials, then across the N-many bytes, the count-down
of the initialization/shift/trim begins, then to be running down the
bytes making each match, while it continues "find-continue", or,
regardless, then that the resulting indicators look for the first
contiguous block of matches.

Then the A/B/other or likely/less-likely/un-likely, is about making the
findings, and automatically composing with making the next findings, or
as that that's in matchings, to adjust the finding as it goes along,
according to that in regular expressions it's a next match, then as with
regards to when there's backtracking and greedy/lazy or among the
greedy/possessive/... regular expressions.


The composition and decomposition of the grammars and expressions, is to
result that after EBNF and regex, the composition and decomposition,
about how to orient the productions and sub-expressions, and their
logic, toward that then alternatives and potentials are arranged their
consequences.

op: + | - | * | / | %
expr: expr op expr

( <-> )

Here the idea is that the balancing of the parentheses and their
relation to the precedence so indicated, is otherwise as according to
left-to-right and right-to-left, about then what induces the potentials
within the balanced parentheses to make expressions, about then the
evalation order of the expressions so indicated, then as with regards to
"concatenation", the most usual operation in strings,

op: /
expr: expr op expr

that when a rule mentions itself it induces a potential, and that when
it has branches that it induces alternatives.

number-initial
number: [non-zero-digit] number

identifier-body: [identifier-body-char] identifier-body
identifier: [identifier-initial] [identifier-body]

keyword: "kw1" | "kw2" | "kw3"

header:
body:
trailer:

sequences "..." introduce sequences (concatenation)
branches "|" introduce alternatives
mentions "<-" introduce potentials
options "[]" introduce options

directionality-left "<" introduces left-balancing, pairing
directionality-right ">" introduces right-balancing, pairing

The directionality or balancing/pairing is indicated when
the left-most and the right-most of the sequence so make
it indicated, the left-most and right-most of a production
of a grammar, or representation/representative of an expression.

op: /
expr: [(] expr op expr [)]

Here the expression has the left-and-right paired, and that
they're only optional mutually, i.e. both or neither, about
a sub-class of optional that's "both-or-neither".


Then, escapes introduce what is a smashing, since the idea
of escapes is that they're symbol-escapes not syntax-escapes,
vis-a-vis quoting, what itself is a syntax-escape, and comments,
what is a syntax-escape, about the escapement, and balancing
and pairing and nested escapes.

So, about the bounds and the offsets, there are the windows
(the coding regions) and the ledges (the ends of the straddles),
then for what goes on the stack of actions, and what is to result
making the stack of findings, is about the organization of

offsets
extents
bounds (offset + extent or offset, offset)

then about the window-bounds and the ledge-bounds,
in terms of those being the word-bounds, and the bounds
of the finding.


union | intersection | complement | setminus

Here complement is usually enough "not", or as
with regards to the entire space of code-points,
about where "not X " is both "universe setminus X"
and "setminus X", about expressions with universes
or "worlds of words". This is that usual accounts of language
are constructively defined as after the alphabet, that here
the alphabet is already "complete" in the sense of the range
of code-points, about then to make for where classes get
defined by ranges or indviduals the range-points, then
in terms of "not" and "complement" and "setminus",
about the logic of union and intersection.


https://wyssmann.com/blog/2019/11/extended-backus-naur-form-ebnf/
https://datatracker.ietf.org/doc/html/rfc2234 (ABNF)


ABNF in RFC2234 introduces ideas of incrementally-defined rules (3.3)
when they are alternatives, here about "composable grammars"
and the ideas of schemas of grammars.

Here there's a fundamental difference between range-points and
alternatives, since range-points are found by code-points while
alternatives would each have their own findings.

Both backtracking and balancing involve state, vis-a-vis,
the "lookahead", the "lookback", and here with regards
to "backstack", and "depthstack", or "pairstack".

The idea of "pairstack" then is each of "backstack"
and "depthstack", about that when crossing words,
while still making a finding, is that the previous words
get pushed on the backstack, then that for balancing
pairs, get pushed on the depthstack, or for example both.


A glossary develops:

register
g-register: a general-purpose register
v-register: a vector register

byte: an octet of bits, interpreted as unsigned integer or bit-flags
nybble: half a byte
word: the v-register word


character-set: a collection of elements of a language
character-encoding: content/layout/format of a character set
character: a member of a character-set
character-class: an attribute of a character or its bytes as properties
or rangepoints

input: a region in memory of contiguous character data, one or more
register words

bit-wise: operating according to index of bits
byte-wise: operating according to index of bytes

offset:
extent:
bounds:

indicators: bit-values 1 yes 0 no

properties: a byte of indicators of a categorical class
predicates: selected interest bits to indicate predicates finding
matching categorical classes
code-points: the byte or bytes that comprise a character
range-points: a lower and upper bound that defines a range of characters
inclusive or individual character

lookup-table: a 256-entry table containing properties for code-points
lookup-line: a linear-lookup cache
lookup-tree: a btree-lookup cache
lookup-file: a backing file for unboundedly many entries

expressions: components and sub-components of regular expressions
representations: examples that match expressions
grammars: rules of composition of expressions
productions: examples that match grammar rules

act: the execution of an instruction of instructions
finding, findings: act, results of making indicators of
properties/predicates or codepoints/rangepoints
matching, matchings: act, results of finding making indicating
representations, productions

made-match
mis-match

working: making findings and matchings over the input
wording: (not a word, working within a word)

straddling: when multi-byte codes cross words
splitting: working either side of a split of a straddling code
stitching: mending both sides of a split of a straddling code

smearing/unsmearing
smashing/unsmashing

backtracking
balancing

backstack
depthstack
pairstack


afore-stitch: cases of straddle, a: start of buffer, before stitch
before-split: cases of straddle, b: end of buffer, before split
after-split: cases of straddle, a: start of buffer, after split
behind-stitch: cases of straddle, b: end of buffer, after stitch



Then, the idea of that it's as a sort of dance (with steps),
or the "rhythm of work" is about the presence of cases
that maintain the context:

work-context
word-context

then about the

initialization
shifting/rotating
trimming

after the

work-offsets
word-offsets

then emitting and maintaining bounds of representatives/productions
of the expressions/grammars.


Then the idea is that for a given offset, the predicates/rangepoints
get popped off the stack, the default algorithm for predicates and
the default algorithm for rangepoints get invoked, or rather, that
a structure makes for defining "relative registers" and having both
the kinds on the same stack, then for example where when there's
potential that the passing predicate gets pushed back on the stack,
or for example that there's made round-robin of all the possible
alternatives on the stack.

Then, making a match results resetting the stack, for example
from the contents of the stack, when making multiple match.

So, in the context, there are predicates and rangepoints, these
are of various sorts.

1) a predicate/range-point is just a duplicated next-char to be spread
and then making finding, the entire word
2) a predicate/range-point is a fixed-length with an extent, to be
making finding

Among the sorts are various cases about whether there's
matching-many (repetitions) or matching-multiple (alternatives),
then for example match-1-alternative or match-all-alternatives
(multi-matching).

Then, next to the predicate/rangepoint or the definition that results
what it is, is about what matches it makes according to its findings,
the matches then being events in the representatives/productions.



Prime Rings and Prime Multisets

As an aside about an example arithmetization, there's the
idea that multisets can be embodied in an integer as primes,
with a catalog of prime numbers to members, then another
idea is about prime rings, finite rings of prime modulus.
The idea is that a given width unsigned integer can maintain
the state of a number of prime rings. For example, Z_5 the
prime ring with five elements, can be represented with 2s,
and then the multiplicity of 2's in the factorization of a number,
is the modulus of the prime ring 0-4.

2^5 = 32

Then, for example with pairs 2, 7 and 3, 5, then an integer
with range >= 7^2 * 5^3 * 3^5 * 2^7 can maintain within
it four prime rings, Z_2 Z_3 Z_5 Z_7 respectively. Then
computing the modulus (or value in the ring 0 to n-1)
is a matter of determining the multiplicity of the given
corresponding factor, while incrementing the ring is a
matter of checking whether b^n-1 is a factor, and dividing
that out to make zero in the ring, else multiplying in b,
to result incrementing in the ring Z_n. It would be usual
enough to instead make for that simply bits and multiples
of bits embody rings, then with just using increment and
modulo on them, then that to store these rings would
take 1-bit for 2, 2-bits for 3, 3-bits for 5 and 7, and so on.

Then, where that might make sense, is when for example
a state transition affects multiple prime rings, that it's a
matter of multiplying in their product to increment both
rings, vis-a-vis setting the relevant bits and adding them
in, then with regards to overflow, either in the adders as
among the bit-packed prime-rings, or in the multipliers
among the prime-backed prime-rings. Prime rings are
useful since when incrementing them each apiece, they
are not zero except when they have common factors of
the counts of increments.


Finders their Ways

So, the finders are basically working across, or down,
across in sequences, and down in alternatives. Then,
there's also that finding is either anchored as prefix-matching,
or drifting as substring-matching.

anchored: prefix-matching (from current offset)
drifting: substring-matching (across offsets)

sequence matching: fixed or likelies
alternative matching: among alternatives

Then, the idea is that the stack of work is the source of
the finders and the matchers, where the finders are the
literals that work in the standard machines, while the matchers
coordinate reaching through arcs to plants, or transitions to states,
that result representatives or productions, then what to do with those.

The standard algorithms are of these kinds:

properties/predicates:
AND the bits to result set bits meaning property = predicate
CMP-to-zero the bits to zero to result 0xFF bytes when all bits are
clear, else 0x00
NOT the bits to result 0xFF when all bits are set

PMOVMSKB the bytes to bits from v-reg to g-reg
BSF the bits to find byte-offsets where property satisfies at least one
predicate

codepoints/rangepoints
CMP-for-gte the lower bound
CMP-for-lte the upper bound
AND the comparisons meaning codepoint between rangepoints
NOT the bits to result 0xFF when all bits are set


PMOVMSKB the bytes to bits from v-reg to g-reg
BSF the bits to find byte-offsets where codepoints between rangepoints

fixed-string sub-string
XOR the bits to result clear bits meaning codepoints match
CMP-to-zero the bits to zero to result 0xFF bytes when all bits are
clear, else 0x00

PMOVMSKB the bytes to bits from v-reg to g-reg
BSF the bits to find byte-offsets where fixed-string equals substring


The predicates make unions, eg, to match either alnum or punct, about
the union of character classes.


Then, the standard algorithm must involve the union, intersection, and
complement/setminus,
about expressions their usual composition. The idea is that these form a
recursive sort of
account, according to implicit and explicit precedence, that result
invoking the standard
algorithms above, to result the bytes to bits from v-reg to g-reg.

These are figured to generally be "yes/no/maybe's" or "sure/yes/no's",
about making
for the the union and intersection of the thing otherwise, that are
pretty simple for
predicates A and B.

union A, B = A || B
intersection A, B = A && B
setminus A \ B = A && !B



So, with regards to the character-set and character-encoding, it's
figured that by default it's Unicode with UTF-8, and that source
texts are overwhelmingly printable ASCII, then that there are also
very usual files that are either UCS2 or UTF-16, or UTF-32. Then, before
the "work" function is along the lines of "detect/inspect", that
otherwise the character-set and character-encoding are assumed
invariants, then that there's as with regards to Internet messages their
declared character-set and character-encoding, and the accounts of
comments and escapes from localedef.


Then, the usual account of each word is mostly clarified, then to get
into the specific semantics of multi-byte characters (characters
generally as both printable and non-printable "characters" then as with
regards to "ligatures" generally and "escapes" generally.

The actions on multi-byte characters mostly are as with regards to
figuring their sparse (or, not completely dense) offsets their first
byte, that first there is the main class its properties, then to be
making the UTF-8 code-points into runs of bytes their characters.


So, the main-class or ascii-class properties are loaded first, instead
of first having a utf-8/non-utf-8 class, since, the distribution of the
content is overwhelmingly printable ASCII (and common control whitespace).


Then, the detection of the coded/ items that are UTF-8 encoding
items follows, with "spotting", and then about the data structures
that indicate the offsets and extents of UTF-8 encoded characters,
to then implement the "smearing", and about escape characters
that result literals, when those are "smashing".

spotting: identifying offsets and extents of UTF-8 characters,
thusly the sparseness/spotting of offsets of characters in the bytes

smearing: extending the sections of predicates according to spotting

Then, for rangepoints gets involved an example, that the ranges are
to be encoded correspondingly into ranges of the UTF-8 encoded
characters. It's figured that contiguous ranges of UTF-8 characters
have contiguous ranges of their encoded bytes.


https://en.wikipedia.org/wiki/Regular_expression
https://en.wikipedia.org/wiki/Parsing_expression_grammar
https://en.wikipedia.org/wiki/Raku_rules
https://en.wikipedia.org/wiki/Recursive_descent_parser
https://en.wikipedia.org/wiki/Thompson%27s_construction


Looking at Thompson's and Glushkov's construction for making
NFA's from expressions, then as with regards to the notion of
minimization after the outer-product or powerset making a DFA,
here is for making what actions are possible, to identify the arcs
and plants, in terms of making of those transitions and states,
about establishing the mutual interpretability of the models
of actions in prefix-matching as usual NFA's/DFA's give, with
regards to prefix- and substring- matching.

It's figured that regular language have forward recognizers,
then as with regards to backtracking and balancing, about
where the recognizer has those, that then gets into limits.

Here the idea of the predictive parser is basically for something
like where Thompson's constructive is said to guarantee that
at most two arcs exit a state, then the idea is that the predicates
can be so combinatorially enumerated, or as what so describes
the matchers, to make consecutive or plural matches in one
"operation", for plural-matches, vis-a-vis multi-matches which
is the idea of having multiple expressions of grammars, about
making plural-predictive predicates and rangepoints, off of
usual constructions of NFA's, that certain predictions are
simpler than others.

Plural Cases

literals: prefix or postfix (suffix)

A usual idea for matching literals is as about the initial-segment
and trailing segment, or, leading segment and final-segment,
where the initial-segment or final-segment is a fixed-string,
while the trailing-segment or leading-segment is variable length,
of a given class, or equivalently, when the class has range-points.
I.e., besides the notion of combining properties/predicates and
code-points/range-points, is to have the fixed-string be the
initial-segment or final-segment, and then the trailing-segment
or leading-segment is a different range in the predicate word,
then that the standard algorithm finds matches for literals
(numeric literals). It's not dissimilar for string literals, about
necessarily enough the escapement, and then also for finding forward
and finding reverse, in the word, and then checking for gaps,
retracting until checking for empty strings, for string or character
literals.

Then the idea is that any of those can be found and matched in
one "run", i.e. a stall-less, branch-less, call-less list of less than
a few or less than a few dozens or less than a few hundreds
instructions, that runs in less than one microsecond.

[toc] | [prev] | [next] | [standalone]


#124479

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2026-07-30 14:55 -0700
Message-ID<114gh8p$25jib$3@kst.eternal-september.org>
In reply to#124464
Ross Finlayson <ross.a.finlayson@gmail.com> writes:
[48 lines deleted]
> RF, good to join the panel. I appreciate the format—direct address and
> genuine exchange rather than parallel monologues.
[4368 lines deleted]

Ross, this is not a "panel".  This is a thread cross-posted to
three newsgroups, comp.theory, comp.lang,c, and comp.lang.c++.

You've just posted more than 4000 lines of text that, as far as I
can tell, have nothing to do with the C or C++ programming languages.

Maybe the discussion is appropriate to comp.theory, which is a
cesspool these days, but in comp.lang.c and comp.lang.c++ we would
very much like to discuss the programming languages that are the
topic of the respective newsgroups without being bombarded with
arrogantly off-topic posts.

I won't try to reason with Johann 'Myrkraverk' Oskarsson, who
seems to enjoy posting to irrelevant newsgroups for some reason,
but perhaps you can do something.  If you're not talking about the
C or C++ programming language, please don't post to comp.lang.c or
comp.lang.c++ -- even if you're posting a followup to a post that
was cross-posted to those groups.  (You'll have to manually edit the
"Newsgroups:" header line.)

I've redirected followups for this post to comp.theory.

Thank you.

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#124516

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-31 12:55 -0700
Message-ID<qx6dnbtn6ZAiYPH3nZ2dnZfqn_udnZ2d@giganews.com>
In reply to#124464
On 07/30/2026 07:05 AM, Ross Finlayson wrote:
> On 07/30/2026 06:49 AM, Ross Finlayson wrote:
>> On 07/27/2026 11:45 AM, Ross Finlayson wrote:
>>> On 07/27/2026 11:44 AM, Ross Finlayson wrote:
>>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote:
>>>>> Hello, here I'll post some design notes and a panel discussion with
>>>>> some
>>>>> chat-bots about making some sense of the "vector-wide scalar word"
>>>>> and "character machines", on commodity hardware about ubiquitous
>>>>> operations.
>>>>>
>>>>>
>>>>> It's considered at least tangentially relevant to comp.lang.c and
>>>>> comp.lang.c++ because for example text is ubiquitous and the targets
>>>>> would be low-level, while the higher-level languages would have a
>>>>> same sort of patternry, and for example that libc and cstdlib are
>>>>> standard, and as with regards to POSIX and Unicode and so on.
>>>>>
>>>>> Please feel free to excuse or ignore, or comment as freely.
>>>>>
>>>>> Thanks for reading.
>>>>>
>>>>
>>>>
>>>> [ viswath-charmaigne.txt ]
>>>>
>>>>
>>>>
>>
>>
>


[ viswath-charmaigne-20260730.txt ]

Drift-Find

About the finding, then for matching, the idea of "drift-find" is as
distinct "anchored-find", about that drift-find is about iterating over
offsets and finding matches, without testing each match as
anchored-test-match.

So, the standard algorithms match byte-wise according to
properties/predicates (that at least one predicate matches at least one 
property) and
codepoints/rangepoints (that the byte is within the range, inclusive, of 
the pair of rangepoints).

Then, when matching word-wise, and drifting the input pattern over the
input data word, then it's ambiguous simply OR'ing together the standard 
algorithm SA
results.

AB pattern
AAB data <- ambiguous whether found at offset 0 or 1, or both

ABA pattern
ABABA data <- ambiguous whether found at offset 0, 1, 2

Then, the idea is to implement an account of the "drift-palindromic" or
"keyway comb", that instead of the SA making 0xFF on finding and 0x00 on 
not finding,
that the drifting accumulate with a sparseness matching from the front, 
and sparseness
matching from the back, and that the combined run must have a length 
matching the
pattern length, then that it's an unambiguous match, the result of the 
finding.

forward -> 1011011101111 ... k-many bits for pattern of length k
reverse -> 0100100010000 ... k-many bits for pattern of length k

Then the idea is that in the drift, as for drift-slip and drift-slide,
that the result of the standard algorithm is converted to each of the 
forward and
reverse, those being put on the stack or otherwise collected, then that 
only when their
union is all 1-bits, is it un-ambiguously alike 0xFF.

The idea is that the combs are generated, then about whether they confirm
the match, or, cancel the match, or about that the findings fiddle the
combs, so that only the first byte of matches get indicated as found, 
and each
of the first bytes, as drift is to find all offsets where the pattern 
matches.

Then, the idea of progressive combs breaks the SBC-less, with the idea of
calculating all the forward and reverse combs, and to give combs at
different offset different progressions of density/sparsity of bits, then to
result that only matching combs result all set bits and only where they 
match. So,
then it is BC-less, yet stalls are introduced when storing on the stack the
combs each, then that they are worked together what result that only
the full matches are found, that S < B < C the cost.

Then, the idea might be to first make the naive match, and then make
the cancel match, that the arithmetic would work out making no-ops
on the matches, and cancels on the mis-matches, since the arithmetic
would be indicated by an already ambiguous match, else no arithmetic.

So, the idea is to store off pairs of combs for each byte offset, or
2W-many, then to go through the combs and any mis-match results 
cancelling at
that offset.

Here that might be alike "optimistic drift", where comb mis-matches are
only to make cancels, else matches: canceling the first byte of the match.

So, the idea is developing to a) make the ambiguous naive match,
then b) make the cancel match, off the first bytes of those.

So, the idea is to drift forward, and union together all the findings,
then drift backward, and zero the first byte if it's not a match.


Then, the drifting case is perhaps much simpler than the drift-palindromic
or the comb-fiddling, with the idea that drift-forward makes all matched
bytes their characters, then drift-revert invalidates the first 
_character_ of
matches on the way back, then that it results that any matches have their
original length, yet, that would possibly invalidate trailing characters 
of an earlier
match, thus getting back into the idea of the drift-palindromic and 
comb-fiddling.


Then, the idea might be to make for canceling the first byte of mismatches,
that might be a last byte of an earlier match, about: going back and
forth setting the first byte, setting the second byte, and so on, or as 
with regards
to whether the output of the algorithm is as sequence of offsets of 
first bytes
instead of otherwise the SA offset-indicator bit-string.

Since the patterns might overlap, then the offset-indicator bit-string
itself is ambiguous, about whether to return the first finding, or 
plurally all
the offsets where findings occur.

Then, the idea would be to result an offsets tuple, where the offsets
range from 0 to W-1, eg 8, 16, 32, 64 for 64, 128, 256, 512 registers,
then that those each fit in a byte, for a word of offsets, where the
maximum offset thus difference in offsets is W-1, and the maximum
count of offsets is W. Then this could be converted to the offset-indicator
bit-string, of starts of matches, instead of saturation of matches.


char-wise indicator string: bits are set
fixed-wise indicator string: starts are set

Then, it seems for only marking the first matching character on the
match, yet, for the initial/final trailing/leading, then it's wanted to
make the bit-string with the plural matches.


"Parallel String Matching
Philip Pfaffe, Martin Tillmann, Sarah Lutteropp, Bernhard Scheirle, and
Kevin Zerr"


One idea then is to make counters, and only bytes with counters being
the length of the fixed-pattern, are included, about matching any byte
in the pattern to any aligned byte in the input, and counting those up
what would be the combinations of all the substrings, that all the
combinations of the substrings match.

Still, not knocking out the first character won't eliminate the starts,
yet not each character is a start.


Then, the idea of "count of matches", may simply enough make
for that differences from 0 indicate overlapping.

This then is to drift along, and find the 0xFF matching, increment
a counter for that offset, and then when going along, that each
increment is a start, and each decrement is an end, then though
at multiples of K, is also an end and a start, if no differences.

Then, only for fixed-patterns, it seems the idea is to find the starts
by checking each offset in the drift, and what results matching,
up to that length, gets incremented, or also, that it can just be
any positive difference indicates a start, so the pattern can be
repeated, then drifted across, and the starts will have increases,
and the non-starts won't.

"M. O. Külekci: Filter Based Fast Matching of Long Patterns by Using
SIMD Instructions"

https://www.stringology.org/

"Handbook of Exact String-Matching Algorithms"
http://www-igm.univ-mlv.fr/~lecroq/string/


Then, for making drift-diff, is that the pattern can simply be made
repeated
in the pattern, and it only needs to drift offsets K-1 many, then the
counts
will have been accumulated, for the diffs to be computed.

W/K

About building the repeated pattern, there is broadcast or the like,

ABC .
012012012012 ...
ABCABCABC ...

then, the idea of not having a loop, or un-rolling the loop, is basically
about that there is binary subdivision, to not explode the number
of statement blocks, into block-with-nops, and also to have the
shorter statement blocks for the shorter patterns.

So, using the standard algorithms SA for matching, then the
predicate/rangepoints of the fixed pattern (a fixed-length predicate or 
fixed-length string or
rangepoints), has that drift invokes the standard algorithm, only to 
compute the
counts, then separating the SA the predication, from moving off the 
result, that the
counts are to be collected, then made their diffs.

ceil log_2 K -> count drift-shifts

Then, for example where K = 1, log_2 1 = 0, the repeated shift makes the
match at once.

Then, there still needs be checking either "diff" or "even modulo" from
the previous match, its count.

So, for the fixed pattern alone, then, for the cost of making it
repeated in the pattern, then for shifting it K-1 many times, and 
accumulating the matches, is
for having W many entry-points, then the rotation simply occurs K-1 
times in the
block, un-rolled.

Then there's the problem of a) straddling when the pattern straddles the
word at B, and b) when the pattern straddles multiple words. The idea is 
that the
prefixes start, and then the remaining pattern gets multi-drifted, which 
would require
enough depth of those rotations, to cover the length of the pattern, or 
a word,
pulling forward the pattern, then also, the pattern, will need to be 
stored in its entirety or as to
that it's loaded from memory in however many words it may straddle.

For example, for pattern ABCD, when the input ends AB, then there's an
anchored match of CD, then to follow with starting over drifting, where 
K < W. For the
pattern AAAA, when the input ends AAA, then each of A, AA, AAA need 
anchored matches,
or drifting with that "the initial segment pattern is found", ..., about 
how to
treat SHIFT and ROTATE so that basically it can make for the repeated 
pattern, to start rotated
left each of the offsets, about making counts of those. Point being, the 
findings of the
straddlings won't complete until as many words have passed as K fits, or 
the last word, and, the
partial matches from the previous word, carry-in and are to accumulate, 
that their offsets
are in the previous word.

About the instruction cache, it makes sense to just have one block, and
then just make it so that the arithmetic just results nops, ....

SA: star
standard algorithms for matching patterns, anchored

SA: fixed
standard algorithms for anchored/drift fixed strings

About the binary indicator-strings, is that 8 words worth of those can
fit into a vector register, about GW, the general purpose word, and Gw, 
in bits, about
that there are 64-bits about which to run BSF/FFS on and make to emit 
offsets.

Ideas about signature of reported findings/matches include:

1) a context struct, and functions to return count,
to compute the size of the return buffer, then
functions to populate the buffer with the offsets,
and about character and byte offsets.

2) a fixed-size output buffer, the function accepts the
size and the buffer and returns the count of elements in it,
which are offsets, returning -1 at EOF (EOI)

3) a fixed-size output buffer, less than pattern/expression max,
making capture groups

4) a callback function, called with offset

5) one pass to compute bounds, one pass to fill bounds


Example: Deflate algorithm, compression/decompression

Compression involves a 32 kiB window, where back-references
would be, then the idea that in a block of up to size 64kiB, then
the heavy computation is the longest-duplicate detection or
"the finding of Huffman codes", as with regards to finding the
most and longest duplicates that get the shortest codes, about
finding the duplicate, or for long runs or the highly compressible,
breaking those down into moduli.

So, the idea would be to make it drifting over itself, that the patterns
naturally enough start from the front,
that there are 32kiB / WB words in the window, eg 2^`5 / 2^7 = 2^8, for
128 bits, 256 words, or that larger vectors would make for larger 
windows, with
just fixing the ratio, then that from the front gets into matching the 
characters, that each word
(8 = 64b, 16 = 128b , 32 = 256b, 64 = 512b, ... bytes) should make its
own Huffman codes, then to combine those, making candidates according to 
those locales,
then to make the account for "long" runs by a histogram of modes, and 
"common" runs
as of the combinations of the modes, ....

Then, about building histograms, the idea is to make the pattern the
rangepoints of itself, i.e., just duplicating the input data, and using 
that as the
pattern, then drifting that along making counts, across the word, then 
each byte will
have how many times it was matched, then to take the max of those, 
building the
histogram from the highest to lowest multplicities (cardinals of the 
multisets).

Then, there's whether those are regular separators, or parts of regular
substrings, then about high/low cardinality with regards to principals, 
modes,
majors, minors, and the long tail, then about the ordering-statistics, 
to build out
histograms to make counting arguments about what those are.



Looking a bit into the object file organization (PECOFF, ELF) it seems that
there are the sections as map to segments with regards to the CALL
instructions, about the idea then that the calls will be with regards to 
the segments,
about how big the segments can be, and then about the range of offsets so
indicated, or about that many segments, each about PAGE_SIZE size, are 
indicated,
about the locals.


nybble 1: alnum punct white coded

nybble 2:

alnum: alpha digit

punct: inner outer joiner affix

white: nl space horz vert

coded: ctrl utf8 nul

nybble 3:

alnum/alpha: upper lower

alnum/digit: zero whole

white/horz: space tab

white/vert: nl cr ff vt

coded/ctrl: single prefix left right

coded/utf8:

punct/inner: arith bool cmp res

punct/outer: quote paren bracket brace

punct/joiner: separator delimiter segment

punct/affix: unary ref kleene lang



nybble 4:

punct/inner/arith: plus minus times slash
punct/inner/res: modulo leftshift rightshift
punct/inner/bool: and or xor
punct/inner/cmp: eq lt gt

punct/affix/unary: bang tilde minus
punct/affix/ref: dollar asterisk ampersand dot
punct/affix/kleene: plus star
punct/affix/lang: period question exclamation

punct/outer/quote: single double backtick

punct/outer/paren: paren-left paren-right brace-left brace-right
punct/outer/brack: angle-left angle-right square-left square-right

punct/joiner/separator: comma semicolon
punct/joiner/delimiter: comma pipe tab
punct/joiner/connector: underscore colon slash backslash

Here the idea is that breaking out punctuation
is about that the usages are overloaded, so that
the properties have that the characters have multiple
properties, so that then according to the context,
then as by the properties are found matched the predicates.

Then, the organization is a curated sort of emphasis for
common source files their usual syntax, or the common.
Then, the idea is that grammars can provide their own
property tables, then that here the first byte is always
included, to make for UTF-8 and NUL and control characters,
and the second byte is "source text" and also "data text".


Then, the standard algorithm will be matching one or more
bytes, here usually two bytes, that the indicated terminals
as they usually are in expressions and grammars, get matched,
that they match the mask of the first byte and the second byte.

Then, the grammar-provided properties would usually
indicate escapes, comments, and additions to the above,
and accounts of characters that introduce ambiguity, to
be disambiguated. As well, the main tables could be
over-ridden, about specific differences from "C-style"
languages.

https://justine.lol/lex/

So, syntax has the "main" and "source" and then expression/grammar driven.

About the logic, there gets involved how to make composable what
result the "anchored" or "atomic" (sub-)expressions and terminals.

The properties/predicates and codepoints/rangepoints can be combined,
where leaving 0's matches none.

The matching of the properties/predicates should be inclusive or
exclusive, "match all" or "match any", here it's default "match any"
(so predicated).

The compositions of "yes/no/maybe" and "union/intersect/setminus"
are to get figured out, how combinations of predicates are to be combined,
basically as of the composition of classes, besides AND, IOR, XOR, NOR.

The, the element of compositions is to result the character classes,
then as with regards to the character classes having both the
predicates/properties and codepoints/rangepoints, the main or default 
ones, and then
union/intersection/setminus of those, and about complement classes.


https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Regular_expressions/Character_class

https://tc39.es/ecma262/multipage/text-processing.html#table-nonbinary-unicode-properties
https://unicode.org/reports/tr18/#General_Category_Property
https://unicode.org/reports/tr18/#Compatibility_Properties

The Unicode TR18 for regular expressions is very useful and could be
considered normative.

https://unicode.org/reports/tr18/#Resolving_Character_Ranges_with_Strings


About shift/rotate on the vector registers, it seems that there's a
problem since there's a limit of 16 bytes for 128 bits (SSE2) , for
packed-shift-right-logical-double-qword, PSRLDQ, the xmm register, that 
there isn't a byte-wise shift, for ymm/zmm
registers, as they get split into lanes, ..., and shifting both the 
double-quadwords would make a
void in the middle. Then, the ymm/zmm would have to be treated as 
separate units, for
example piling in the instructions on both sides using the same offsets 
and computing for
alternatives and so on.

https://www.felixcloutier.com/x86/
https://mischasan.wordpress.com/2011/04/04/what-is-sse-good-for-2-bit-vector-operations/
https://www.scs.stanford.edu/~zyedidia/arm64/sveindex.html

It looks similar with ARM.

Then the idea would be to work up to double-quadwords or 128-bits the
16-bytes, as with regards then to making the acts being round-robin'ed 
to each of the
packed double-quadwords, then about updating the anchors the offsets in 
lock-step.

Then it's figured that the acts on the machines, that output the
bit-string indicators of the byte offsets about smearing/unsmearing and 
byte and character
offsets, would have a tag of what was found and matched in terms of the 
expression/grammar,
that resulted the indicators, then that it's serialized what makes the
matches/productions.

Then for ARM NEON it looks like there's no double-quadword shift (128-bits)
only each of the packed dwords (32-bit), "SIMD" on NEON.

There is a REV64 instruction on ARM as might be about BSWAP, then with
the idea though that shift byte-wise is the idea, and NEON instructions
are "on each double-word", 32-bits.

Then it might make sense just to divide-and-conquer, yet the lock-step
item gets involved with having a common view of the input data and a 
given offset as
the current sort of state-of-the-machine.

"VEXT can be used to implement a moving window on data from two vectors,
useful in FIR filters. For permutation, it can also be used to simulate
a byte-wise
rotate operation, when using the same vector for both input operands."

--
https://developer.arm.com/community/arm-community-blogs/b/architectures-and-processors-blog/posts/coding-for-neon---part-5-rearranging-vectors

So, that then can effect "vector byte-wise right shift", basically loading
from the end of the zero vector and the beginning of the vector to
be shifted.

It's considered a MOV so it leads to stalls. Then in SVE there's EXTQ,
which is also organized about 128-bit double-QWORDS.



About smearing and byte/character offsets then, those would mostly
go to the vector registers as a bit-sequence indicator will indicate
starts of characters in the byte-sequence.




Yeah, I've been looking at this, and here's what it seems
is the profile, of the resources, about the vector units,
on Intel/AMD and ARM.

So, first there's that MMX since Pentium is still alive,
yet, it's considered sort of aside what are the general
purpose registers, if for a sort of "general-auxiliary"
use, about the "16 general purpose registers". Then ARM
mostly has "32 general purpose registers", with the idea
that Intel has 16 (or less) general purpose + registers
+ 8 old floating-point/MMX SIMD vectors.

So, there are basically 16 general purpose registers
on each, and 2 of those on ARM.

Then, the vector registers basically make for "SSE 4.2"
or here for what's SSE3 yet beyond SSE2, about there being
vector registers now essentially separate from general registers.

So, here the goal is to use the vector registers like large
scalars, or at least as arrays of bytes. Well, that's not
exactly the goal of the vector/packed/SIMD registers. So,
there's a common subset of functionality, and limits within
the vector registers, about what can be treated as scalars
(with the byte as least-addressable, shift & rotate, and
with the logical operations and compare that go straight
up and down, in terms of two vector registers their lanes
their words their bytes their bits).

Basically then there's "double quad-word" or 128 bits,
in both the Intel/AMD and ARM, that's about the biggest
"scalar" word there, as the data type, for the common
subset of instructions abstractly they support.

Then, the SSE4.2, has 128-bit vector-registers, that
can be operated upon with their DQ for double-quadword
variants of instructions, alike scalars, or at least
for the byte-wise, if not necessarily the bit-wise,
with regards to shift & rotate even multiples of 8 bits.

Then AVX with 256-bits, is two of those side-by-side,
similarly AVX-512 then, is two of those side-by-side,
and ARM SVE, is one or more of those side-by-side,
128-bit double quad-words with "byte-wise" moves like
shift & rotate, with regards to using "extract" on
ARM to simulate shift & rotate multiples of 8-bits.

So, this sort of tiling of the register files, thinking
of the registers the memories as a rectangular block of
bits, about the register transfer logic moving the bits
or computing the bits, basically gives 128-bit 16-long
blocks, that can be treated like "byte-addressable scalars".

SSE4.2: 1 block (16-many x 128-wide)
ARM NEON: 2 blocks (32-many x 128-wide)
AVX: 2 blocks (16-many x 256-wide)
AVX2: 4 blocks (32-many x 256-wide)
AVX-512: 8 blocks (32-many x 512-wide)
ARM SVE: 2-20 blocks (32-many x 128-2048-wide)

where all the widths are essentially separate units
run together in lock-step of "double quad-word type size"
byte-addressable "scalars".

So, algorithms should be designed to work in 1 block,
in the register file, and then scale in these blocks,
for vector-wide scalar-word operations (byte-wise).

Here then the idea is that the "character machine"
basically implements a little scheduler and then
making the various findings and matchings in the blocks.



Then, figuring for making a "scheduler" is after a "plan",
figuring that the expressions and grammars have their
events of representatives and productions, then as
with regards to the operation of "matchings" and
"parsings", in the machine, then as with regards to
the static machine, "the engine".


So, overall, the functional units of the machine and engine
are 128b = 16B wide, and 16-registers deep, then as with
regards to the notion of scheduling the units as with
regards to various and evolving "standard algorithms" SA,
and then a model of the 64b = 8B wide, and 8-registers deep,
for fallback to core 64-bit general purpose their auxiliary registers,
or as for reference and fallback implementations in higher-level
languages.








[toc] | [prev] | [next] | [standalone]


#124517

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-31 13:05 -0700
Message-ID<5--dnVBtWaqEnfD3nZ2dnZfqn_dg4p2d@giganews.com>
In reply to#124516
On 07/31/2026 12:55 PM, Ross Finlayson wrote:
> On 07/30/2026 07:05 AM, Ross Finlayson wrote:
>> On 07/30/2026 06:49 AM, Ross Finlayson wrote:
>>> On 07/27/2026 11:45 AM, Ross Finlayson wrote:
>>>> On 07/27/2026 11:44 AM, Ross Finlayson wrote:
>>>>> On 07/27/2026 11:43 AM, Ross Finlayson wrote:
>>>>>> Hello, here I'll post some design notes and a panel discussion with
>>>>>> some
>>>>>> chat-bots about making some sense of the "vector-wide scalar word"
>>>>>> and "character machines", on commodity hardware about ubiquitous
>>>>>> operations.
>>>>>>
>>>>>>
>>>>>> It's considered at least tangentially relevant to comp.lang.c and
>>>>>> comp.lang.c++ because for example text is ubiquitous and the targets
>>>>>> would be low-level, while the higher-level languages would have a
>>>>>> same sort of patternry, and for example that libc and cstdlib are
>>>>>> standard, and as with regards to POSIX and Unicode and so on.
>>>>>>
>>>>>> Please feel free to excuse or ignore, or comment as freely.
>>>>>>
>>>>>> Thanks for reading.
>>>>>>
>>>>>
>>>>>
>>>>> [ viswath-charmaigne.txt ]
>>>>>
>>>>>
>>>>>
>>>
>>>
>>
>

[ Excuse, replied to an earlier post before dropping comp.lang.c,
comp.lang.c++, please ignore, as follow-ups are to comp.theory.
It's appreciated the tolerance or absence thereof. -- ]


[toc] | [prev] | [next] | [standalone]


#124548 — Ljubljana School versus Zurich School (Was: Viswath & Charmaigne)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 20:14 +0200
SubjectLjubljana School versus Zurich School (Was: Viswath & Charmaigne)
Message-ID<114qlpk$t7q2$1@solani.org>
In reply to#124392
Hi,

I even don't remember exactly why I landed
in comp.theory. A yes, because Rossy Boy,
was hooked on SIMD and didn't understand Hack.

But the Hack work, rather belongs to my
Alma Mater Zurich and my personal heros, Gutnecht
and Wirth, who wrote a one pass Modula

compiler during some christmas holidays,
back then when I was student. Not sure
whether the Ljubljana School can do that,

when I read this here:

Finite Algebraic Effects as dicts and such
https://www.philipzucker.com/bdd_term_alg_effects/

I only find gibberish like:
- “Data” is somehow less mysterious to me
   than “computation”.  [..] I don’t even
   know what “computation” really is

- In temporal logic, there is a logic CTL
   which talks about computation trees.

- Algerbaic (LoL) effects is almost a complete
   hackery abuse of the notion of arity
   and that’s neat.

- Then there are 10^10 etc vectors which
   show up if you discretize 3d/4d space. [..]
   or reinforcement learning.

- Etc..

WTF is this guy smoking? I mean he even
doesn't uses math notation, only posts
Python code fragments à go go,

possibly a Python brain damage.

But still less sever than Rossy Boys.

Bye

Ross Finlayson schrieb:
> Hello, here I'll post some design notes and a panel discussion with some
> chat-bots about making some sense of the "vector-wide scalar word"
> and "character machines", on commodity hardware about ubiquitous 
> operations.
> 
> 
> It's considered at least tangentially relevant to comp.lang.c and
> comp.lang.c++ because for example text is ubiquitous and the targets
> would be low-level, while the higher-level languages would have a
> same sort of patternry, and for example that libc and cstdlib are
> standard, and as with regards to POSIX and Unicode and so on.
> 
> Please feel free to excuse or ignore, or comment as freely.
> 
> Thanks for reading.
> 

[toc] | [prev] | [next] | [standalone]


#124549 — Re: Ljubljana School versus Zurich School (Was: Viswath & Charmaigne)

FromJohann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid>
Date2026-08-04 02:21 +0800
SubjectRe: Ljubljana School versus Zurich School (Was: Viswath & Charmaigne)
Message-ID<1%4cS.103380$aXr.938@fx18.ams4>
In reply to#124548
On 04/08/2026 2:14 AM, Mild Shock wrote:
> Hi,
> 
> I even don't remember exactly why I landed
> in comp.theory. A yes, because Rossy Boy,
> was hooked on SIMD and didn't understand Hack.
> 
> But the Hack work, rather belongs to my
> Alma Mater Zurich and my personal heros, Gutnecht
> and Wirth, who wrote a one pass Modula

That's interesting.  Have you read /Software Engineering
with Modula-2 and Ada/ (1984) by Richard Wiener and Richard
Sincovec?  I have it on my shelf, and haven't gotten to read
it yet.

> 
> compiler during some christmas holidays,
> back then when I was student. Not sure
> whether the Ljubljana School can do that,
> 
> when I read this here:
> 
> Finite Algebraic Effects as dicts and such
> https://www.philipzucker.com/bdd_term_alg_effects/
> 
> I only find gibberish like:
> - “Data” is somehow less mysterious to me
>    than “computation”.  [..] I don’t even
>    know what “computation” really is

Computation is at the core just a calculation.  Humans
used to do this, and there's a good documentary about it
titled /Hidden Figures/.  I assume everyone here has seen
it.

> 
> - In temporal logic, there is a logic CTL
>    which talks about computation trees.
> 
> - Algerbaic (LoL) effects is almost a complete
>    hackery abuse of the notion of arity
>    and that’s neat.

Are you using /algebraic effects/ when playing League
of Legends?  I tried it once, but discovered it's a
gameplay that doesn't appeal to me.  I didn't think to
use /algebraic effects/ in it.

> 
> - Then there are 10^10 etc vectors which
>    show up if you discretize 3d/4d space. [..]
>    or reinforcement learning.

I don't know why you have that many vectors visiting,
but please treat them with hospitality according to
Zeus' laws.

> 
> - Etc..
> 
> WTF is this guy smoking? I mean he even
> doesn't uses math notation, only posts
> Python code fragments à go go,
> 
> possibly a Python brain damage.

Or he just works in the ministry of silly walks?


Enjoy!
-- 
Johann | email: invalid -> com | http://www.myrkraverk.com/blog/
I'm not from the Internet, I just work there. | via Easynews.com
https://bsky.app/profile/myrkraverk.bsky.social

[toc] | [prev] | [next] | [standalone]


#124550 — pi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 20:35 +0200
Subjectpi-WAM uses ADA RendezVous (Was: Ljubljana School versus Zurich School)
Message-ID<114qn2b$t8nm$1@solani.org>
In reply to#124549
Hi,

I found that this here:

     public final static class RendezVous {
         private final Semaphore head = new Semaphore(0);
         private final Semaphore tail = new Semaphore(1);
         private Object data;

         public void put(Object data) throws InterruptedException {
             tail.acquire();
             this.data = data;
             head.release();
         }

         public Object take() throws InterruptedException {
             Object res;
             head.acquire();
             res = data;
             tail.release();
             return res;
         }
     }

Is almost as fast as ArrayBlockingQueue(4),
in a producer worker consumer scenario.

So I considering using the above for the
pi-WAM channels. It would be also closer

to pi-calculus by Robin Milner.

Bye

Johann 'Myrkraverk' Oskarsson schrieb:
> On 04/08/2026 2:14 AM, Mild Shock wrote:
>> Hi,
>>
>> I even don't remember exactly why I landed
>> in comp.theory. A yes, because Rossy Boy,
>> was hooked on SIMD and didn't understand Hack.
>>
>> But the Hack work, rather belongs to my
>> Alma Mater Zurich and my personal heros, Gutnecht
>> and Wirth, who wrote a one pass Modula
> 
> That's interesting.  Have you read /Software Engineering
> with Modula-2 and Ada/ (1984) by Richard Wiener and Richard
> Sincovec?  I have it on my shelf, and haven't gotten to read
> it yet.
> 
>>
>> compiler during some christmas holidays,
>> back then when I was student. Not sure
>> whether the Ljubljana School can do that,
>>
>> when I read this here:
>>
>> Finite Algebraic Effects as dicts and such
>> https://www.philipzucker.com/bdd_term_alg_effects/
>>
>> I only find gibberish like:
>> - “Data” is somehow less mysterious to me
>>    than “computation”.  [..] I don’t even
>>    know what “computation” really is
> 
> Computation is at the core just a calculation.  Humans
> used to do this, and there's a good documentary about it
> titled /Hidden Figures/.  I assume everyone here has seen
> it.
> 
>>
>> - In temporal logic, there is a logic CTL
>>    which talks about computation trees.
>>
>> - Algerbaic (LoL) effects is almost a complete
>>    hackery abuse of the notion of arity
>>    and that’s neat.
> 
> Are you using /algebraic effects/ when playing League
> of Legends?  I tried it once, but discovered it's a
> gameplay that doesn't appeal to me.  I didn't think to
> use /algebraic effects/ in it.
> 
>>
>> - Then there are 10^10 etc vectors which
>>    show up if you discretize 3d/4d space. [..]
>>    or reinforcement learning.
> 
> I don't know why you have that many vectors visiting,
> but please treat them with hospitality according to
> Zeus' laws.
> 
>>
>> - Etc..
>>
>> WTF is this guy smoking? I mean he even
>> doesn't uses math notation, only posts
>> Python code fragments à go go,
>>
>> possibly a Python brain damage.
> 
> Or he just works in the ministry of silly walks?
> 
> 
> Enjoy!

[toc] | [prev] | [next] | [standalone]


#124553 — A spinlock rewrite will be necessary (Was: pi-WAM uses ADA RendezVous)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 21:00 +0200
SubjectA spinlock rewrite will be necessary (Was: pi-WAM uses ADA RendezVous)
Message-ID<114qogr$t9ob$1@solani.org>
In reply to#124550
Hi,

But because I do a grouping of logical threads
before I go on physical threads, a spinlock
rewrite will be necessary.

I did already a spinlock rewrite, using
a class Spinlock instead of the class Semaphore.
But ultimately I would switch from put() to

an offer() API, that returns a boolean, and
this can be used to skip instructions or otherwise
react in the Hack VM. Same for take() would

need to replace by poll() with repercussions
to Hack VM again. This is much to the dismay
of Chris M. Thomasson, who thinks spinning

is strictly forbidden. But I will sing the song:

    I'm a spinner, I'm a sinner
    I spin on CAS loops for my dinner
    Some call it busy-wait, I call it fate
    When the queue is empty, I just rotate

Bye

Mild Shock schrieb:
> Hi,
> 
> I found that this here:
> 
>      public final static class RendezVous {
>          private final Semaphore head = new Semaphore(0);
>          private final Semaphore tail = new Semaphore(1);
>          private Object data;
> 
>          public void put(Object data) throws InterruptedException {
>              tail.acquire();
>              this.data = data;
>              head.release();
>          }
> 
>          public Object take() throws InterruptedException {
>              Object res;
>              head.acquire();
>              res = data;
>              tail.release();
>              return res;
>          }
>      }
> 
> Is almost as fast as ArrayBlockingQueue(4),
> in a producer worker consumer scenario.
> 
> So I considering using the above for the
> pi-WAM channels. It would be also closer
> 
> to pi-calculus by Robin Milner.
> 
> Bye
> 
> Johann 'Myrkraverk' Oskarsson schrieb:
>> On 04/08/2026 2:14 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> I even don't remember exactly why I landed
>>> in comp.theory. A yes, because Rossy Boy,
>>> was hooked on SIMD and didn't understand Hack.
>>>
>>> But the Hack work, rather belongs to my
>>> Alma Mater Zurich and my personal heros, Gutnecht
>>> and Wirth, who wrote a one pass Modula
>>
>> That's interesting.  Have you read /Software Engineering
>> with Modula-2 and Ada/ (1984) by Richard Wiener and Richard
>> Sincovec?  I have it on my shelf, and haven't gotten to read
>> it yet.
>>
>>>
>>> compiler during some christmas holidays,
>>> back then when I was student. Not sure
>>> whether the Ljubljana School can do that,
>>>
>>> when I read this here:
>>>
>>> Finite Algebraic Effects as dicts and such
>>> https://www.philipzucker.com/bdd_term_alg_effects/
>>>
>>> I only find gibberish like:
>>> - “Data” is somehow less mysterious to me
>>>    than “computation”.  [..] I don’t even
>>>    know what “computation” really is
>>
>> Computation is at the core just a calculation.  Humans
>> used to do this, and there's a good documentary about it
>> titled /Hidden Figures/.  I assume everyone here has seen
>> it.
>>
>>>
>>> - In temporal logic, there is a logic CTL
>>>    which talks about computation trees.
>>>
>>> - Algerbaic (LoL) effects is almost a complete
>>>    hackery abuse of the notion of arity
>>>    and that’s neat.
>>
>> Are you using /algebraic effects/ when playing League
>> of Legends?  I tried it once, but discovered it's a
>> gameplay that doesn't appeal to me.  I didn't think to
>> use /algebraic effects/ in it.
>>
>>>
>>> - Then there are 10^10 etc vectors which
>>>    show up if you discretize 3d/4d space. [..]
>>>    or reinforcement learning.
>>
>> I don't know why you have that many vectors visiting,
>> but please treat them with hospitality according to
>> Zeus' laws.
>>
>>>
>>> - Etc..
>>>
>>> WTF is this guy smoking? I mean he even
>>> doesn't uses math notation, only posts
>>> Python code fragments à go go,
>>>
>>> possibly a Python brain damage.
>>
>> Or he just works in the ministry of silly walks?
>>
>>
>> Enjoy!
> 

[toc] | [prev] | [next] | [standalone]


#124551 — Big thanks to Ljubljana School [Searching 0xCAFFEE] (Was: Ljubljana School versus Zurich School)

FromMild Shock <janburse@fastmail.fm>
Date2026-08-03 20:40 +0200
SubjectBig thanks to Ljubljana School [Searching 0xCAFFEE] (Was: Ljubljana School versus Zurich School)
Message-ID<114qnbd$t8ve$1@solani.org>
In reply to#124548
Hi,

I have nevertheless to thank the Ljubljana
School, especially this blog post:

Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/

Which raised my interest in Hack. Meanwhile
I could produce this toy eample:

"We try to find 0xCAFFEE in enumerating 4
6-bit digits and the baseline is Dogelog
Player VM in a browser. The CPU backend
with 64 logical threads is already 20
times faster, partly due to its 32-bit
specialization. The GPU backend with
4096 logical threads boosts a further
factor of 7 times."

GPU Backend: Find 0xCAFFEE with π-WAM
https://medium.com/2989/8890efd3503c

LoL

Bye

Mild Shock schrieb:
> Hi,
> 
> I even don't remember exactly why I landed
> in comp.theory. A yes, because Rossy Boy,
> was hooked on SIMD and didn't understand Hack.
> 
> But the Hack work, rather belongs to my
> Alma Mater Zurich and my personal heros, Gutnecht
> and Wirth, who wrote a one pass Modula
> 
> compiler during some christmas holidays,
> back then when I was student. Not sure
> whether the Ljubljana School can do that,
> 
> when I read this here:
> 
> Finite Algebraic Effects as dicts and such
> https://www.philipzucker.com/bdd_term_alg_effects/
> 
> I only find gibberish like:
> - “Data” is somehow less mysterious to me
>    than “computation”.  [..] I don’t even
>    know what “computation” really is
> 
> - In temporal logic, there is a logic CTL
>    which talks about computation trees.
> 
> - Algerbaic (LoL) effects is almost a complete
>    hackery abuse of the notion of arity
>    and that’s neat.
> 
> - Then there are 10^10 etc vectors which
>    show up if you discretize 3d/4d space. [..]
>    or reinforcement learning.
> 
> - Etc..
> 
> WTF is this guy smoking? I mean he even
> doesn't uses math notation, only posts
> Python code fragments à go go,
> 
> possibly a Python brain damage.
> 
> But still less sever than Rossy Boys.
> 
> Bye
> 
> Ross Finlayson schrieb:
>> Hello, here I'll post some design notes and a panel discussion with some
>> chat-bots about making some sense of the "vector-wide scalar word"
>> and "character machines", on commodity hardware about ubiquitous 
>> operations.
>>
>>
>> It's considered at least tangentially relevant to comp.lang.c and
>> comp.lang.c++ because for example text is ubiquitous and the targets
>> would be low-level, while the higher-level languages would have a
>> same sort of patternry, and for example that libc and cstdlib are
>> standard, and as with regards to POSIX and Unicode and so on.
>>
>> Please feel free to excuse or ignore, or comment as freely.
>>
>> Thanks for reading.
>>
> 

[toc] | [prev] | [standalone]


Page 6 of 6 — ← Prev page 1 2 3 4 5 [6]

Back to top | Article view | comp.lang.c++


csiph-web