Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > sci.logic > #348207 > unrolled thread

The Wuhan Virus that destroyed Python [ggml Manifesto]

Started byMild Shock <janburse@fastmail.fm>
First post2026-07-22 21:01 +0200
Last post2026-08-09 21:19 +0200
Articles 20 on this page of 98 — 6 participants

Back to article view | Back to sci.logic


Contents

  The Wuhan Virus that destroyed Python [ggml Manifesto] Mild Shock <janburse@fastmail.fm> - 2026-07-22 21:01 +0200
    Deadlock Exorcism: Switch from Push to Pull [A pi-calculus Specification of Prolog] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:25 +0200
      Why do you even need a mpmc queue? [Thunder Kittens] (Re: Deadlock Exorcism: Switch from Push to Pull) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:45 +0200
        Trivial balancing example for (int i=0; i<global_id; i++) (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:55 +0200
          The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:19 +0200
            Enqueue/dequeue need not be fast and can spinn ["fairness" questions] (Re: The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:23 +0200
        And, where did I talk about rockets? [Hint its about xAI's Grok] (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 01:25 +0200
          Why forget something, that was never on my mind (Re: And, where did I talk about rockets? [Hint its about xAI's Grok]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:49 +0200
        Example Mandel Brot rendering [Faster with MIMD] (Was: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:56 +0200
    Potential Python Recovery: Free Threading [3.13 release] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 10:21 +0200
    The things XILINX braught to the AMD table (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 18:48 +0200
      NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-23 19:13 +0200
        Micro penis mother sung arias (Re: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:40 +0200
          Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:27 +0200
            Ignoramus or Ignorabimus: I don't care [(Re: Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:35 +0200
              You are a moron, brainless putin payed (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:01 +0200
                Yeah keep reading my posts, uninspired fool (Re: You are a moron, brainless putin payed) Mild Shock <janburse@fastmail.fm> - 2026-07-24 19:47 +0200
              Out of the blue accusation span 15 days [Empirical USENET study] (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:27 +0200
              A brain desease of 20 days [Rossy Boy] (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:41 +0200
                I didn't use a Ryzen Halo, whats wrong with you? (Re: A brain desease of 20 days [Rossy Boy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:25 +0200
                Ignoramus / Ignorabimus Barometer: Almost 1 Month (Re: A brain desease of 20 days [Rossy Boy]) Mild Shock <janburse@fastmail.fm> - 2026-08-03 00:07 +0200
        Re: NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-28 14:17 +0200
        ASML stocks are plunging, bye bye dutchies (Re: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-28 14:18 +0200
    Little Data Center on Your Palm [AI Laptops for 500 USD] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 17:59 +0200
      2008: 4 Blades + Tesla S1070 versus 2026: 1 AI Laptop (Re: Little Data Center on Your Palm [AI Laptops for 500 USD]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:15 +0200
      Budget AI Laptop 2026 versus Cray T3D 1995 (Was: Little Data Center on Your Palm) Mild Shock <janburse@fastmail.fm> - 2026-08-05 14:21 +0200
        Re: Budget AI Laptop 2026 versus Cray T3D 1995 R Kym Horsell <kym@sdf.org> - 2026-08-05 21:11 +0000
          AI Alarmist with Supercomputer on Yacht [Horsy Boy] (Was: Budget AI Laptop 2026 versus Cray T3D 1995) Mild Shock <janburse@fastmail.fm> - 2026-08-06 13:21 +0200
            Re: AI Alarmist with Supercomputer on Yacht [Horsy Boy] R Kym Horsell <kym@sdf.org> - 2026-08-06 11:56 +0000
    Hurry the blue bus doesnt stop indefinitely (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:37 +0200
      Not SIMD, a MIMD design for NVIDIA Volta (Re: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:58 +0200
        Could take 3-4 months find machine / browser (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-24 21:16 +0200
        The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-26 19:54 +0200
          The turbo capping of AI Laptops (Was: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-26 20:00 +0200
          Re: The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
          Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
            miniTriton CUDA is an alternative to torch variants (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:52 +0200
              Andrej Karpathy original gangster of Budget Laptop (Re: miniTriton CUDA is an alternative to torch variants) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:54 +0200
            The evolution of hardware and GPT-2 training (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 10:57 +0200
              How speed up π-WAM with vector operations (Re: The evolution of hardware and GPT-2 training) Mild Shock <janburse@fastmail.fm> - 2026-07-27 11:10 +0200
                AI accelerator extend from GPU to CPU [Zero Copying] (Re: How speed up π-WAM with vector operations) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:21 +0200
                  The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:22 +0200
                    Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 07:34 -0700
                      Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 17:12 +0200
                      Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kym@sdf.com> - 2026-07-27 15:43 +0000
                        Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kymhorsell@gmail.com> - 2026-07-27 15:46 +0000
                        π-WAM is not adding decimals, it is removing decimals (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:34 +0200
                          In Budget Laptops the TOPS come with low energy footprint (Re: π-WAM is not adding decimals, it is removing decimals) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:45 +0200
      Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Was: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-27 16:56 +0200
        Rossy Boy is neither Einstein nor Zweistein (Was: Potato Computer owner impressed by Ukraine Tech) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:25 +0200
          You are still chewing on SIMD. LoL (Re: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:12 +0200
            Hurry Rossy Boy, the blue bus is waiting (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:53 +0200
              Look how they advertized CUDA and logical threads (Re: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:55 +0200
                Forget any arithmetization of product FSA (Re: Look how they advertized CUDA and logical threads) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:58 +0200
            Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:06 +0200
              I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:25 +0200
                Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: I don't use Rust, you are crazy) Mild Shock <janburse@fastmail.fm> - 2026-07-29 22:52 +0200
                  A funny Q16.16 experiment with Hack (Re: Hack ecosystem ignorance paired with paranoia [Nand to Tetris]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:11 +0200
                    Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Was: A funny Q16.16 experiment with Hack) Mild Shock <janburse@fastmail.fm> - 2026-07-30 11:25 +0200
                  I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:37 +0200
                    For WebGPU I first had SIMD in mind (Re: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:51 +0200
                      Corr.: 4 Months --> 4 Weeks (Re: For WebGPU I first had SIMD in mind) Mild Shock <janburse@fastmail.fm> - 2026-07-30 20:05 +0200
                    MIPS is a big Huffman mess [But Hack could do it] (Re: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:32 +0200
                      Not declarative with PHI (Φ) nodes (Was: MIPS is a big Huffman mess [But Hack could do it]) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:38 +0200
                      Not declarative with PHI (Φ) nodes (Re: MIPS is a big Huffman mess [But Hack could do it]) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:39 +0200
                      Re: MIPS is a big Huffman mess [But Hack could do it] (Re: I wrote Hack VM for π-WAM from scratch) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 09:52 -0700
                        Re: MIPS is a big Huffman mess [But Hack could do it] (Re: I wrote Hack VM for π-WAM from scratch) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 10:00 -0700
                          Re: MIPS is a big Huffman mess [But Hack could do it] (Re: I wrote Hack VM for π-WAM from scratch) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-31 10:07 -0700
                    Quo Vadis: Extend investigations to WebNN (Was: I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java]) Mild Shock <janburse@fastmail.fm> - 2026-07-31 20:43 +0200
                Standing on the shoulders of giants (Re: I don't use Rust, you are crazy [Jump off a bridge, idiot]) Mild Shock <janburse@fastmail.fm> - 2026-08-04 03:19 +0200
                  You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc.. (Re: Standing on the shoulders of giants) Mild Shock <janburse@fastmail.fm> - 2026-08-04 15:19 +0200
                    How Rossy Boys plagiarism works [Copy Paste Slop] (Re: You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc.. ) Mild Shock <janburse@fastmail.fm> - 2026-08-04 17:58 +0200
                      Statistics gave up, no salient truth [Signal Collapse] Re: How Rossy Boys plagiarism works [Copy Paste Slop] (Re: You Thief! Stealing Szemeredi, Aristotle, Leibniz, etc.. ) Mild Shock <janburse@fastmail.fm> - 2026-08-04 18:17 +0200
    Got it. Or are you too stupid? [New Usenet Mantra] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:59 +0200
    Lamas in a cradle and Lamas on the edge [Red Pyjama] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 13:03 +0200
      AI Accelerators and ISO Prolog multi-threading (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:02 +0200
        Actor/Erlang is dead, no Thread and Mailbox conflation [golang channels] (Re: AI Accelerators and ISO Prolog multi-threading) (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:04 +0200
          Can library(ironpaw) repurpose FFT hardware [Glimps into Ryzen AI 7 350] (Re: Actor/Erlang is dead, no Thread and Mailbox conflation ) Mild Shock <janburse@fastmail.fm> - 2026-08-01 02:32 +0200
      Tablet and phone UBS-C remote debugging (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-08-01 12:19 +0200
        NPUs doing 2d chess comms (Manhattan Distance or L1 Norm) (Re: Tablet and phone UBS-C remote debugging) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:12 +0200
          NACK retransmission might double Manhattan Distance (Re: NPUs doing 2d chess comms) Mild Shock <janburse@fastmail.fm> - 2026-08-01 14:24 +0200
            I am using WebGPU, and not WebGL (Re: NACK retransmission might double Manhattan Distance) Mild Shock <janburse@fastmail.fm> - 2026-08-02 00:47 +0200
              Texture inside my compute shader makes no sense (Re: I am using WebGPU, and not WebGL) Mild Shock <janburse@fastmail.fm> - 2026-08-02 02:40 +0200
                Prolog inferencing and not canvasing fancy stuff (Re: Texture inside my compute shader makes no sense) Mild Shock <janburse@fastmail.fm> - 2026-08-02 02:42 +0200
                  It’s called . . . . enshittification (About the price tag for using a multifile/1) Mild Shock <janburse@fastmail.fm> - 2026-08-14 00:49 +0200
        Chris M. Thomasson can ask 100 more questions (Was: Tablet and phone UBS-C remote debugging) Mild Shock <janburse@fastmail.fm> - 2026-08-02 02:46 +0200
          npm install webgpu [Google Dawn] (Re: Chris M. Thomasson can ask 100 more questions) Mild Shock <janburse@fastmail.fm> - 2026-08-02 03:03 +0200
            GPU elasticity was already invented in 2008 with CUDA (Re: npm install webgpu [Google Dawn]) Mild Shock <janburse@fastmail.fm> - 2026-08-03 00:01 +0200
      Synthetic Multilanguage Autoformalization Dataset [Informath project] (Was: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-08-08 09:17 +0200
        Re: Synthetic Multilanguage Autoformalization Dataset [Informath project] (Was: Lamas in a cradle and Lamas on the edge [Red Pyjama]) x3 <x@x.net> - 2026-08-08 11:45 -0700
          Nice try Rossy Boy --> **plonk** (Was: Synthetic Multilanguage Autoformalization Dataset [Informath project]) Mild Shock <janburse@fastmail.fm> - 2026-08-08 23:02 +0200
            Ethernal September Idiots Gone (Was: Nice try Rossy Boy --> **plonk**) Mild Shock <janburse@fastmail.fm> - 2026-08-08 23:15 +0200
    Even send_color and recv_color can block [Cerebras Waver] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-08-02 23:39 +0200
    GPU Elasticity: Collective Communications Libraries (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-08-07 14:35 +0200
      What are Flits and Phits? [Network on a Chip] (Was: GPU Elasticity: Collective Communications Libraries) Mild Shock <janburse@fastmail.fm> - 2026-08-07 18:05 +0200
        The Mac Neo is a Budget Monster [GPU Channels] (Re: What are Flits and Phits? [Network on a Chip]) Mild Shock <janburse@fastmail.fm> - 2026-08-11 16:25 +0200
          The luminaries of duct-tape engineering [Sweeney and Torvald] (Re: The Mac Neo is a Budget Monster [GPU Channels]) Mild Shock <janburse@fastmail.fm> - 2026-08-11 16:54 +0200
    Cristallina: Thank you for the Beam (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-08-09 21:19 +0200

Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →


#348251 — AI accelerator extend from GPU to CPU [Zero Copying] (Re: How speed up π-WAM with vector operations)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 13:21 +0200
SubjectAI accelerator extend from GPU to CPU [Zero Copying] (Re: How speed up π-WAM with vector operations)
Message-ID<1147evl$glbs$2@solani.org>
In reply to#348246
Hi,

The nice thing about AI accelerators, pioneered
maybe by Apple Silicon and their unified memory.
The AMD APU model can be extended so that

vector and matrix operations become uniformly
available for GPU and CPU. With unified memory
already a vector operation such as:

vec_mul_add(X, 2, 3, Y)

Only needs the X and Y address. But I havent
got my head around yet how this is all organized.
Maybe a GPU has still its own GEMM cores,

but you find Apple Silicon C++/C source code,
that taps into vector and matrix operations
by Zero Copying. The Copying is left to the DMA

of the vector or matrix operation. And moderated
by the various caches. Leading to the slogan, that
multiple floating point operations become zero cost:

Some teaching can be found here
https://www.hpc-ch.org/category/topics/course-workshop/

Bye

Mild Shock schrieb:
> Hi,
> 
> One could critisize that my π-WAM doesn't
> utilize GPU to the fullest, since its GPU
> backend prototype only uses scalar operations
> 
> and no vector or matrix operations. And
> modern GPUs thrive on vector and matrix
> operations. Especially matrix operations giving
> 
> a boost of a factor 15x or so. There are
> many papers already showing how Prolog can be
> mapped to matrix operations. Only this research
> 
> is completely ignored by Prolog systems such as
> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
> illustrate what vector operations could do
> 
> for π-WAM, take this compilation of the Prolog
> goal between(0,1023,X), Y is X*2+3:
> 
> int X;
> int Y;
> for (X=0; X < 1024; X++) {
>      Y=X*2+3;
>      [...]
> }
> 
> With vector operations, and vectors of size
> 32 one could do:
> 
> int X1;
> int[] X = new int[32];
> int X3;
> int[] Y = new int[32];
> for (X1 = 0; X1 < 1024 / 32; X1++) {
>      for (int X2 = 0; X2 < 32; X2++)
>         X[X2] = X1*32+X2;
>      vec_mul_add(X, 2, 3, Y);
>      [..]
> }
> 
> Have Fun!
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> While Huggingfaces hired GG in 2026,
>> AK was hired by Anthropic in 2026:
>>
>> Andrej Karpathy (born 23 October 1986[3])
>> is a Slovak-Canadian AI researcher, who
>> co-founded and formerly worked at OpenAI
>> In 2026 he joined Anthropic as part of
>> the pretraining team.
>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>
>> But his nanochat archivement has an
>> interesting time line:
>>
>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>> The best ChatGPT that $100 can buy.
>> https://github.com/karpathy/nanochat
>>
>> But what hardware was the enabler. What is the
>> NVIDIA H100 GPU even. Well the thingy is surely not
>> a Budget Laptop, performance pretty much
>>
>> dependence on data elememt size, the H100 NVL
>> version (*), and when using tensor operations,
>> and not only scalar operations:
>>
>> 8-bit towards 3000 tera flops
>> 16-bit towards 1500 tera flops
>> 32-bit towards 900 tera flops
>>
>> Cool! I guess this experiment would tap into 60
>> tera flops, since it only uses scalar operations so far:
>>
>> 11.4 Giga Lips with a Budget Laptop
>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> You could perform it by migration the web application
>> using WebGPU into a node.js standalone application
>> using the dawn library for GPU access.
>>
>> Bye
>>
>> (*) 
>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
>>
>>
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Whats this "forget" trope of glue sniffing
>>> Rossy Boy with his herpes blisters?
>>>
>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>  > forget Hungarians and Bulgarians.
>>>
>>> Why should I forget Bulgarians,
>>> they are never on my mind. Do you
>>> see me doing ggml stuff?
>>>
>>> I only hypothesized that it is
>>> over for Python as the machine
>>> learning language or AI inferencing
>>>
>>> locally on AI laptops language, and
>>> made the ggml case, so I already forgot
>>> about them. Which might give you a glimps,
>>>
>>> why WebGPU was used for this here:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> Is an interesting choice. Even
>>> github has some Languages statistics,
>>> giving an account what I used:
>>>
>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>
>>> Have Fun!
>>>
>>> Bye
>>>
>>> P.S.: The example below is not p-adics,
>>> you complete imbecil moron. Its just:
>>>
>>> 7-11 cubic Solution by Pritchard & Gries
>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>
>>> Ross Finlayson schrieb:
>>>  > On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>  >> Hi,
>>>  >>
>>>  >> You see it all boils down to find your inner peace
>>>  >> by an immaculate inception of some queue datatype.
>>>  >>
>>>  >> KOAN/Fortran-S was an early 1990s research programming
>>>  >> system for distributed-memory multiprocessors . Developed
>>>  >> at ENS Lyon in the early 1990s . Often listed alongside
>>>  >> other historical parallel programming efforts.
>>>  >>
>>>  >> The Message Passing: The research explicitly
>>>  >> compared the SVM approach against message passing
>>>  >> on the same hardware . The finding was that SVM
>>>  >> could achieve good performance without the low-level
>>>  >>
>>>  >> complexity of managing explicit messages, though
>>>  >> the best results often came from a hybrid approach (sic!)
>>>  >> Here is an interesting baseline, from Java,
>>>  >> a class ElevenSingle that only does:
>>>  >>
>>>  >>      public static void run() {
>>>  >>          for (int A = 1; A < 192; A++) {
>>>  >>              int Y = (771-A)/3;
>>>  >>              for (int B = A; B < Y; B++) {
>>>  >>                  int Z = (771-A-B)/2;
>>>  >>                  for (int C = B; C < Z; C++) {
>>>  >>                      int D = 711-A-B-C;
>>>  >>                      if (A*B*C == 711000000/D &&
>>>  >>                            711000000 % D == 0)
>>>  >>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>  >>                  }
>>>  >>              }
>>>  >>          }
>>>  >>      }
>>>  >>
>>>  >> And then compare it to ElevenMulti, doing some
>>>  >> Work Balancing Scheduler Tetris Game with 8 cores:
>>>  >>
>>>  >> ElevenSingle
>>>  >> A=120, B=125, C=150, D=316
>>>  >> 6.628 ms
>>>  >>
>>>  >> ElevenMulti
>>>  >> A=120, B=125, C=150, D=316
>>>  >> 1.941 ms
>>>  >>
>>>  >> Not great, not terrible!
>>>  >>
>>>  >> Bye
>>>  >
>>>  > Oh, that's just "tricks of p-adic arithmetic".
>>>  >
>>>  > Like other sock-puppet howler trolls, when confronted
>>>  > with its base incredulity, it will descend to its
>>>  > lower levers of the pathos variety.
>>>  >
>>>  > You might be happier learning about Julia trees and
>>>  > raster ops, instead of shilling yet another Ramanujan
>>>  > series without saying how it's made.
>>>  >
>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>  > forget Hungarians and Bulgarians.
>>>  >
>>>  >
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348252 — The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 13:22 +0200
SubjectThe invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<1147f1f$glbs$3@solani.org>
In reply to#348251
Hi,

But the example gives also way to vector
and matrix registers. The int[] X and
int[] Y could be also held in vector

registers. Compilers can also optimize
away int[] Y, and use a inline modification,
in case X isn't used later, then playing

the role of Y:

vec_mul_add(X, 2, 3, X)

Vector and matrix registers in modern GPUs
emerged from distinct architectural milestones:
vector-like register files developed with
early programmable 3D vertex/pixel pipelines

in the late 1990s to early 2000s. While
dedicated multi-dimensional matrix registers
(Tensor Cores/Matrix Cores) were invented by
NVIDIA in 2017, starting with the Tesla

V100 (Volta microarchitecture):

 From Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell

You see the scheduling of tensure core occupation
scheduling in the above article, including memory
and register flow, following the section:

MMA Instruction Overview

It went through a couple of generations, leading
to Tensor Memory (TMEM) and collective operations,
basically realizing the PIM idea:

Processing-in-Memory Tutorials
https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/

Have Fun!

Bye


Mild Shock schrieb:
> Hi,
> 
> The nice thing about AI accelerators, pioneered
> maybe by Apple Silicon and their unified memory.
> The AMD APU model can be extended so that
> 
> vector and matrix operations become uniformly
> available for GPU and CPU. With unified memory
> already a vector operation such as:
> 
> vec_mul_add(X, 2, 3, Y)
> 
> Only needs the X and Y address. But I havent
> got my head around yet how this is all organized.
> Maybe a GPU has still its own GEMM cores,
> 
> but you find Apple Silicon C++/C source code,
> that taps into vector and matrix operations
> by Zero Copying. The Copying is left to the DMA
> 
> of the vector or matrix operation. And moderated
> by the various caches. Leading to the slogan, that
> multiple floating point operations become zero cost:
> 
> Some teaching can be found here
> https://www.hpc-ch.org/category/topics/course-workshop/
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> One could critisize that my π-WAM doesn't
>> utilize GPU to the fullest, since its GPU
>> backend prototype only uses scalar operations
>>
>> and no vector or matrix operations. And
>> modern GPUs thrive on vector and matrix
>> operations. Especially matrix operations giving
>>
>> a boost of a factor 15x or so. There are
>> many papers already showing how Prolog can be
>> mapped to matrix operations. Only this research
>>
>> is completely ignored by Prolog systems such as
>> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
>> illustrate what vector operations could do
>>
>> for π-WAM, take this compilation of the Prolog
>> goal between(0,1023,X), Y is X*2+3:
>>
>> int X;
>> int Y;
>> for (X=0; X < 1024; X++) {
>>      Y=X*2+3;
>>      [...]
>> }
>>
>> With vector operations, and vectors of size
>> 32 one could do:
>>
>> int X1;
>> int[] X = new int[32];
>> int X3;
>> int[] Y = new int[32];
>> for (X1 = 0; X1 < 1024 / 32; X1++) {
>>      for (int X2 = 0; X2 < 32; X2++)
>>         X[X2] = X1*32+X2;
>>      vec_mul_add(X, 2, 3, Y);
>>      [..]
>> }
>>
>> Have Fun!
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> While Huggingfaces hired GG in 2026,
>>> AK was hired by Anthropic in 2026:
>>>
>>> Andrej Karpathy (born 23 October 1986[3])
>>> is a Slovak-Canadian AI researcher, who
>>> co-founded and formerly worked at OpenAI
>>> In 2026 he joined Anthropic as part of
>>> the pretraining team.
>>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>>
>>> But his nanochat archivement has an
>>> interesting time line:
>>>
>>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>>> The best ChatGPT that $100 can buy.
>>> https://github.com/karpathy/nanochat
>>>
>>> But what hardware was the enabler. What is the
>>> NVIDIA H100 GPU even. Well the thingy is surely not
>>> a Budget Laptop, performance pretty much
>>>
>>> dependence on data elememt size, the H100 NVL
>>> version (*), and when using tensor operations,
>>> and not only scalar operations:
>>>
>>> 8-bit towards 3000 tera flops
>>> 16-bit towards 1500 tera flops
>>> 32-bit towards 900 tera flops
>>>
>>> Cool! I guess this experiment would tap into 60
>>> tera flops, since it only uses scalar operations so far:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> You could perform it by migration the web application
>>> using WebGPU into a node.js standalone application
>>> using the dawn library for GPU access.
>>>
>>> Bye
>>>
>>> (*) 
>>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
>>>
>>>
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Whats this "forget" trope of glue sniffing
>>>> Rossy Boy with his herpes blisters?
>>>>
>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>  > forget Hungarians and Bulgarians.
>>>>
>>>> Why should I forget Bulgarians,
>>>> they are never on my mind. Do you
>>>> see me doing ggml stuff?
>>>>
>>>> I only hypothesized that it is
>>>> over for Python as the machine
>>>> learning language or AI inferencing
>>>>
>>>> locally on AI laptops language, and
>>>> made the ggml case, so I already forgot
>>>> about them. Which might give you a glimps,
>>>>
>>>> why WebGPU was used for this here:
>>>>
>>>> 11.4 Giga Lips with a Budget Laptop
>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> Is an interesting choice. Even
>>>> github has some Languages statistics,
>>>> giving an account what I used:
>>>>
>>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>>
>>>> Have Fun!
>>>>
>>>> Bye
>>>>
>>>> P.S.: The example below is not p-adics,
>>>> you complete imbecil moron. Its just:
>>>>
>>>> 7-11 cubic Solution by Pritchard & Gries
>>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>  >> Hi,
>>>>  >>
>>>>  >> You see it all boils down to find your inner peace
>>>>  >> by an immaculate inception of some queue datatype.
>>>>  >>
>>>>  >> KOAN/Fortran-S was an early 1990s research programming
>>>>  >> system for distributed-memory multiprocessors . Developed
>>>>  >> at ENS Lyon in the early 1990s . Often listed alongside
>>>>  >> other historical parallel programming efforts.
>>>>  >>
>>>>  >> The Message Passing: The research explicitly
>>>>  >> compared the SVM approach against message passing
>>>>  >> on the same hardware . The finding was that SVM
>>>>  >> could achieve good performance without the low-level
>>>>  >>
>>>>  >> complexity of managing explicit messages, though
>>>>  >> the best results often came from a hybrid approach (sic!)
>>>>  >> Here is an interesting baseline, from Java,
>>>>  >> a class ElevenSingle that only does:
>>>>  >>
>>>>  >>      public static void run() {
>>>>  >>          for (int A = 1; A < 192; A++) {
>>>>  >>              int Y = (771-A)/3;
>>>>  >>              for (int B = A; B < Y; B++) {
>>>>  >>                  int Z = (771-A-B)/2;
>>>>  >>                  for (int C = B; C < Z; C++) {
>>>>  >>                      int D = 711-A-B-C;
>>>>  >>                      if (A*B*C == 711000000/D &&
>>>>  >>                            711000000 % D == 0)
>>>>  >>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>  >>                  }
>>>>  >>              }
>>>>  >>          }
>>>>  >>      }
>>>>  >>
>>>>  >> And then compare it to ElevenMulti, doing some
>>>>  >> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>  >>
>>>>  >> ElevenSingle
>>>>  >> A=120, B=125, C=150, D=316
>>>>  >> 6.628 ms
>>>>  >>
>>>>  >> ElevenMulti
>>>>  >> A=120, B=125, C=150, D=316
>>>>  >> 1.941 ms
>>>>  >>
>>>>  >> Not great, not terrible!
>>>>  >>
>>>>  >> Bye
>>>>  >
>>>>  > Oh, that's just "tricks of p-adic arithmetic".
>>>>  >
>>>>  > Like other sock-puppet howler trolls, when confronted
>>>>  > with its base incredulity, it will descend to its
>>>>  > lower levers of the pathos variety.
>>>>  >
>>>>  > You might be happier learning about Julia trees and
>>>>  > raster ops, instead of shilling yet another Ramanujan
>>>>  > series without saying how it's made.
>>>>  >
>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>  > forget Hungarians and Bulgarians.
>>>>  >
>>>>  >
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348254 — Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-27 07:34 -0700
SubjectRe: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<M-KdneYpK_8w8fr3nZ2dnZfqnPSdnZ2d@giganews.com>
In reply to#348252
On 07/27/2026 04:22 AM, Mild Shock wrote:
> Hi,
>
> But the example gives also way to vector
> and matrix registers. The int[] X and
> int[] Y could be also held in vector
>
> registers. Compilers can also optimize
> away int[] Y, and use a inline modification,
> in case X isn't used later, then playing
>
> the role of Y:
>
> vec_mul_add(X, 2, 3, X)
>
> Vector and matrix registers in modern GPUs
> emerged from distinct architectural milestones:
> vector-like register files developed with
> early programmable 3D vertex/pixel pipelines
>
> in the late 1990s to early 2000s. While
> dedicated multi-dimensional matrix registers
> (Tensor Cores/Matrix Cores) were invented by
> NVIDIA in 2017, starting with the Tesla
>
> V100 (Volta microarchitecture):
>
>  From Volta To Blackwell
> https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
>
>
> You see the scheduling of tensure core occupation
> scheduling in the above article, including memory
> and register flow, following the section:
>
> MMA Instruction Overview
>
> It went through a couple of generations, leading
> to Tensor Memory (TMEM) and collective operations,
> basically realizing the PIM idea:
>
> Processing-in-Memory Tutorials
> https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/
>
>
> Have Fun!
>
> Bye
>
>
> Mild Shock schrieb:
>> Hi,
>>
>> The nice thing about AI accelerators, pioneered
>> maybe by Apple Silicon and their unified memory.
>> The AMD APU model can be extended so that
>>
>> vector and matrix operations become uniformly
>> available for GPU and CPU. With unified memory
>> already a vector operation such as:
>>
>> vec_mul_add(X, 2, 3, Y)
>>
>> Only needs the X and Y address. But I havent
>> got my head around yet how this is all organized.
>> Maybe a GPU has still its own GEMM cores,
>>
>> but you find Apple Silicon C++/C source code,
>> that taps into vector and matrix operations
>> by Zero Copying. The Copying is left to the DMA
>>
>> of the vector or matrix operation. And moderated
>> by the various caches. Leading to the slogan, that
>> multiple floating point operations become zero cost:
>>
>> Some teaching can be found here
>> https://www.hpc-ch.org/category/topics/course-workshop/
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> One could critisize that my π-WAM doesn't
>>> utilize GPU to the fullest, since its GPU
>>> backend prototype only uses scalar operations
>>>
>>> and no vector or matrix operations. And
>>> modern GPUs thrive on vector and matrix
>>> operations. Especially matrix operations giving
>>>
>>> a boost of a factor 15x or so. There are
>>> many papers already showing how Prolog can be
>>> mapped to matrix operations. Only this research
>>>
>>> is completely ignored by Prolog systems such as
>>> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
>>> illustrate what vector operations could do
>>>
>>> for π-WAM, take this compilation of the Prolog
>>> goal between(0,1023,X), Y is X*2+3:
>>>
>>> int X;
>>> int Y;
>>> for (X=0; X < 1024; X++) {
>>>      Y=X*2+3;
>>>      [...]
>>> }
>>>
>>> With vector operations, and vectors of size
>>> 32 one could do:
>>>
>>> int X1;
>>> int[] X = new int[32];
>>> int X3;
>>> int[] Y = new int[32];
>>> for (X1 = 0; X1 < 1024 / 32; X1++) {
>>>      for (int X2 = 0; X2 < 32; X2++)
>>>         X[X2] = X1*32+X2;
>>>      vec_mul_add(X, 2, 3, Y);
>>>      [..]
>>> }
>>>
>>> Have Fun!
>>>
>>> Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> While Huggingfaces hired GG in 2026,
>>>> AK was hired by Anthropic in 2026:
>>>>
>>>> Andrej Karpathy (born 23 October 1986[3])
>>>> is a Slovak-Canadian AI researcher, who
>>>> co-founded and formerly worked at OpenAI
>>>> In 2026 he joined Anthropic as part of
>>>> the pretraining team.
>>>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>>>
>>>> But his nanochat archivement has an
>>>> interesting time line:
>>>>
>>>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>>>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>>>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>>>> The best ChatGPT that $100 can buy.
>>>> https://github.com/karpathy/nanochat
>>>>
>>>> But what hardware was the enabler. What is the
>>>> NVIDIA H100 GPU even. Well the thingy is surely not
>>>> a Budget Laptop, performance pretty much
>>>>
>>>> dependence on data elememt size, the H100 NVL
>>>> version (*), and when using tensor operations,
>>>> and not only scalar operations:
>>>>
>>>> 8-bit towards 3000 tera flops
>>>> 16-bit towards 1500 tera flops
>>>> 32-bit towards 900 tera flops
>>>>
>>>> Cool! I guess this experiment would tap into 60
>>>> tera flops, since it only uses scalar operations so far:
>>>>
>>>> 11.4 Giga Lips with a Budget Laptop
>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> You could perform it by migration the web application
>>>> using WebGPU into a node.js standalone application
>>>> using the dawn library for GPU access.
>>>>
>>>> Bye
>>>>
>>>> (*)
>>>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet
>>>>
>>>>
>>>>
>>>> Mild Shock schrieb:
>>>>> Hi,
>>>>>
>>>>> Whats this "forget" trope of glue sniffing
>>>>> Rossy Boy with his herpes blisters?
>>>>>
>>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>>  > forget Hungarians and Bulgarians.
>>>>>
>>>>> Why should I forget Bulgarians,
>>>>> they are never on my mind. Do you
>>>>> see me doing ggml stuff?
>>>>>
>>>>> I only hypothesized that it is
>>>>> over for Python as the machine
>>>>> learning language or AI inferencing
>>>>>
>>>>> locally on AI laptops language, and
>>>>> made the ggml case, so I already forgot
>>>>> about them. Which might give you a glimps,
>>>>>
>>>>> why WebGPU was used for this here:
>>>>>
>>>>> 11.4 Giga Lips with a Budget Laptop
>>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>>
>>>>> Is an interesting choice. Even
>>>>> github has some Languages statistics,
>>>>> giving an account what I used:
>>>>>
>>>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>>>
>>>>> Have Fun!
>>>>>
>>>>> Bye
>>>>>
>>>>> P.S.: The example below is not p-adics,
>>>>> you complete imbecil moron. Its just:
>>>>>
>>>>> 7-11 cubic Solution by Pritchard & Gries
>>>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>>>
>>>>> Ross Finlayson schrieb:
>>>>>  > On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>>  >> Hi,
>>>>>  >>
>>>>>  >> You see it all boils down to find your inner peace
>>>>>  >> by an immaculate inception of some queue datatype.
>>>>>  >>
>>>>>  >> KOAN/Fortran-S was an early 1990s research programming
>>>>>  >> system for distributed-memory multiprocessors . Developed
>>>>>  >> at ENS Lyon in the early 1990s . Often listed alongside
>>>>>  >> other historical parallel programming efforts.
>>>>>  >>
>>>>>  >> The Message Passing: The research explicitly
>>>>>  >> compared the SVM approach against message passing
>>>>>  >> on the same hardware . The finding was that SVM
>>>>>  >> could achieve good performance without the low-level
>>>>>  >>
>>>>>  >> complexity of managing explicit messages, though
>>>>>  >> the best results often came from a hybrid approach (sic!)
>>>>>  >> Here is an interesting baseline, from Java,
>>>>>  >> a class ElevenSingle that only does:
>>>>>  >>
>>>>>  >>      public static void run() {
>>>>>  >>          for (int A = 1; A < 192; A++) {
>>>>>  >>              int Y = (771-A)/3;
>>>>>  >>              for (int B = A; B < Y; B++) {
>>>>>  >>                  int Z = (771-A-B)/2;
>>>>>  >>                  for (int C = B; C < Z; C++) {
>>>>>  >>                      int D = 711-A-B-C;
>>>>>  >>                      if (A*B*C == 711000000/D &&
>>>>>  >>                            711000000 % D == 0)
>>>>>  >>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>>  >>                  }
>>>>>  >>              }
>>>>>  >>          }
>>>>>  >>      }
>>>>>  >>
>>>>>  >> And then compare it to ElevenMulti, doing some
>>>>>  >> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>>  >>
>>>>>  >> ElevenSingle
>>>>>  >> A=120, B=125, C=150, D=316
>>>>>  >> 6.628 ms
>>>>>  >>
>>>>>  >> ElevenMulti
>>>>>  >> A=120, B=125, C=150, D=316
>>>>>  >> 1.941 ms
>>>>>  >>
>>>>>  >> Not great, not terrible!
>>>>>  >>
>>>>>  >> Bye
>>>>>  >
>>>>>  > Oh, that's just "tricks of p-adic arithmetic".
>>>>>  >
>>>>>  > Like other sock-puppet howler trolls, when confronted
>>>>>  > with its base incredulity, it will descend to its
>>>>>  > lower levers of the pathos variety.
>>>>>  >
>>>>>  > You might be happier learning about Julia trees and
>>>>>  > raster ops, instead of shilling yet another Ramanujan
>>>>>  > series without saying how it's made.
>>>>>  >
>>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>>  > forget Hungarians and Bulgarians.
>>>>>  >
>>>>>  >
>>>>>
>>>>
>>>
>>
>

That's bullshit, and alike those talking heads that
sniff their way into talking about many-core jumbo-trons,
the super-scalar is as old as the scalar and Cray and examples alike
the Connection Machine what made all the craze of neural nets
is old-wrapped-as-new.

Fabless chips did it already.


Data centers should pay a 10000% excise on electricity,
wherever it comes from, a natural regulator of inverted economies.

And by ten thousand percent I really mean a ten thousand percent.

[toc] | [prev] | [next] | [standalone]


#348257 — Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Was: The invention of vector and matrix registers [NVIDIA Volta])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 17:12 +0200
SubjectMaybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Was: The invention of vector and matrix registers [NVIDIA Volta])
Message-ID<1147sgk$geia$1@solani.org>
In reply to#348254
Hi,

I guess Rossy Boys mother was so disappointed
in the 50's that is son didn't become the
next Einstain, physics was the ultinate idol,

so that Rossy Boy was left rotting in the
basement. But Rossy Boys indoctrination was
not spurious, he now is conditioned on

Einstein. Maybe NVIDIA should have named
its Tesla V100 card NVIDIA Einstein. You would
then see Rossy Boy toe sucking the graphic

card, in his pyjamas in the basement.

Bye

Ross Finlayson schrieb:
> On 07/27/2026 04:22 AM, Mild Shock wrote:
>> Hi,
>>
>> But the example gives also way to vector
>> and matrix registers. The int[] X and
>> int[] Y could be also held in vector
>>
>> registers. Compilers can also optimize
>> away int[] Y, and use a inline modification,
>> in case X isn't used later, then playing
>>
>> the role of Y:
>>
>> vec_mul_add(X, 2, 3, X)
>>
>> Vector and matrix registers in modern GPUs
>> emerged from distinct architectural milestones:
>> vector-like register files developed with
>> early programmable 3D vertex/pixel pipelines
>>
>> in the late 1990s to early 2000s. While
>> dedicated multi-dimensional matrix registers
>> (Tensor Cores/Matrix Cores) were invented by
>> NVIDIA in 2017, starting with the Tesla
>>
>> V100 (Volta microarchitecture):
>>
>>  From Volta To Blackwell
>> https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell 
>>
>>
>>
>> You see the scheduling of tensure core occupation
>> scheduling in the above article, including memory
>> and register flow, following the section:
>>
>> MMA Instruction Overview
>>
>> It went through a couple of generations, leading
>> to Tensor Memory (TMEM) and collective operations,
>> basically realizing the PIM idea:
>>
>> Processing-in-Memory Tutorials
>> https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/ 
>>
>>
>>
>> Have Fun!
>>
>> Bye
> That's bullshit, and alike those talking heads that
> sniff their way into talking about many-core jumbo-trons,
> the super-scalar is as old as the scalar and Cray and examples alike
> the Connection Machine what made all the craze of neural nets
> is old-wrapped-as-new.
> 
> Fabless chips did it already.
> 
> 
> Data centers should pay a 10000% excise on electricity,
> wherever it comes from, a natural regulator of inverted economies.
> 
> And by ten thousand percent I really mean a ten thousand percent.
> 
> 

[toc] | [prev] | [next] | [standalone]


#348260 — Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])

FromR Kym Horsell <kym@sdf.com>
Date2026-07-27 15:43 +0000
SubjectRe: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<1147ub3$2542$1@nnrp.usenet.blueworldhosting.com>
In reply to#348254
In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
...
> Data centers should pay a 10000% excise on electricity,
> wherever it comes from, a natural regulator of inverted economies.
> And by ten thousand percent I really mean a ten thousand percent.

And what would a huge surcharge do?
Almost always end up affecting the less powerful end of society
with increased costs to services the AI industry will be doing
more and more of over time.

I started a little data center (exaflops.com) many years ago.
In those distant days people (in fact one was a prof of computer
science) told me you could never make money running a supercomputer. 
LOL. :)

I've had many years to watch the trends and a far more efficient
way to solve resource problems in this area is to change the
algorithms. There is vast room for improvement, mostly because
of prevailing attitudes.

I used to do competetion data science as a sideline. Companies
would pay almost any price to get an extra decimal place in
the accuracy of their forecasting processes. But typically
they were trying to supercharge a system that should be scrapped
and re-designed from scratch. One area I'm thinking of is
investment. I had a customer one time -- like many times --
ask to improve a system that predicted the future price of
various stocks. The idea (for them) was to have as accurate a
prediction of what some stock would be worth in a week or a month's
time so that some moron could use the information to decide when
to buy or sell the thing.

I tried to argue the efficient thing was to create a system that
takes the human out of the loop altogether. It doesnt provide info
for someone to decide whether or not to follow the advice --
that is just introducing more noise into the loop and probably
cancels any benefit of adding a couple decimal places of precision.
What you *should* do is make a system that is tuned to robustly
maximize the profit from managing a portfolio.

Of course they wouldnt come at that. You can't suggest taking the
managers out of the loop. :)

Another idea relevant to current AI methods might be to curtail
use of typical neural net algorithms. Many of them try to squeeze
the best performance of some NN during the training  phase in
the hope the resulting system will generalize well enough to be useful
on new data. But there's kind-of a law that the harder you train
some system to perform a task well, the less well they can subsuently
perform a more general version of the same thing. It's amusing when
you look at the graphs of NN being trained and then tested that
given a more general problem to solve after being trained to solve
similar problems very very well the poor old NN does worse that it
would have done if it had 0 training in the first place.

It's not like we dont know how to improve this kind of performance.
Try less hard in the training phase or make it "more noisy".
Turns out genetic methods are just the ticket for this.
The training produces less over-fitting and the resulting system
generalizes better than it did before training and more importantly
it takes maybe an order of magnitude crunching to produce a good answer
than the usual over-fit answer.

Anyway. Have to go and feed the cat.

-- 
Nothing in life is to be feared, it is only to be understood. 
Now is the time to understand more, so that we may fear less.
-- M. Curie

[toc] | [prev] | [next] | [standalone]


#348261 — Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])

FromR Kym Horsell <kymhorsell@gmail.com>
Date2026-07-27 15:46 +0000
SubjectRe: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<1147uh4$2542$2@nnrp.usenet.blueworldhosting.com>
In reply to#348260
In comp.lang.prolog R Kym Horsell <kym@sdf.com> wrote:
> generalizes better than it did before training and more importantly
> it takes maybe an order of magnitude crunching to produce a good answer
                                      /\ less
> than the usual over-fit answer.

[toc] | [prev] | [next] | [standalone]


#348263 — π-WAM is not adding decimals, it is removing decimals (Was: The invention of vector and matrix registers [NVIDIA Volta])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:34 +0200
Subjectπ-WAM is not adding decimals, it is removing decimals (Was: The invention of vector and matrix registers [NVIDIA Volta])
Message-ID<11481af$h2se$1@solani.org>
In reply to#348260
Hi,

Come on Horsy Boy, you can do better. I
no where wrote something about curve
fitting and/or increasing the precision of

float point numbers:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

What makes you think LIPS measures precision?
You should know better as a 50% Prologer.

I explictily wrote here what the goal is:

"shave off some of the TOPS to do Prolog inferencing"

What are TOPS? Its a metric for GPUs:

TOPS stands for “Trillions of Operations Per Second.”
https://www.lenovo.com/us/en/glossary/tops-in-computing/

See for yourself what is behind my post:

11.4 Giga Lips with a Budget Laptop
At the end of 2025 we acquired a couple of AI Laptops , that were still 
cheap, since RAM prices had not yet rocketed. The intend was to tap into 
the Copilot+ certified hardware, and shave off some of the TOPS to do 
Prolog inferencing. Amazingly our π-WAM can churn 11.4 GIGA LIPS.

GPUs have evolved form lock-step to independent thread scheduling. This 
made it possible to port the Hack VM variant, that forms the basis for 
our π-WAM, to WebGPU computer shaders. Using NUM_SHADERS = 4096 we could 
produce 11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M.

See also:

Medium Article - 11.4 Giga Lips
https://medium.com/2989/899b0d5c027b

So just get lost with your crazy irrelevant rant.
When I get more LIPS, things run faster, and
I remove digits from the time dimension.

Got it. Or are you too stupid?

Bye

R Kym Horsell schrieb:
> In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
> ...
>> Data centers should pay a 10000% excise on electricity,
>> wherever it comes from, a natural regulator of inverted economies.
>> And by ten thousand percent I really mean a ten thousand percent.
> 
> And what would a huge surcharge do?
> Almost always end up affecting the less powerful end of society
> with increased costs to services the AI industry will be doing
> more and more of over time.
> 
> I started a little data center (exaflops.com) many years ago.
> In those distant days people (in fact one was a prof of computer
> science) told me you could never make money running a supercomputer.
> LOL. :)
> 
> I've had many years to watch the trends and a far more efficient
> way to solve resource problems in this area is to change the
> algorithms. There is vast room for improvement, mostly because
> of prevailing attitudes.
> 
> I used to do competetion data science as a sideline. Companies
> would pay almost any price to get an extra decimal place in
> the accuracy of their forecasting processes. But typically
> they were trying to supercharge a system that should be scrapped
> and re-designed from scratch. One area I'm thinking of is
> investment. I had a customer one time -- like many times --
> ask to improve a system that predicted the future price of
> various stocks. The idea (for them) was to have as accurate a
> prediction of what some stock would be worth in a week or a month's
> time so that some moron could use the information to decide when
> to buy or sell the thing.
> 
> I tried to argue the efficient thing was to create a system that
> takes the human out of the loop altogether. It doesnt provide info
> for someone to decide whether or not to follow the advice --
> that is just introducing more noise into the loop and probably
> cancels any benefit of adding a couple decimal places of precision.
> What you *should* do is make a system that is tuned to robustly
> maximize the profit from managing a portfolio.
> 
> Of course they wouldnt come at that. You can't suggest taking the
> managers out of the loop. :)
> 
> Another idea relevant to current AI methods might be to curtail
> use of typical neural net algorithms. Many of them try to squeeze
> the best performance of some NN during the training  phase in
> the hope the resulting system will generalize well enough to be useful
> on new data. But there's kind-of a law that the harder you train
> some system to perform a task well, the less well they can subsuently
> perform a more general version of the same thing. It's amusing when
> you look at the graphs of NN being trained and then tested that
> given a more general problem to solve after being trained to solve
> similar problems very very well the poor old NN does worse that it
> would have done if it had 0 training in the first place.
> 
> It's not like we dont know how to improve this kind of performance.
> Try less hard in the training phase or make it "more noisy".
> Turns out genetic methods are just the ticket for this.
> The training produces less over-fitting and the resulting system
> generalizes better than it did before training and more importantly
> it takes maybe an order of magnitude crunching to produce a good answer
> than the usual over-fit answer.
> 
> Anyway. Have to go and feed the cat.
> 

[toc] | [prev] | [next] | [standalone]


#348264 — In Budget Laptops the TOPS come with low energy footprint (Re: π-WAM is not adding decimals, it is removing decimals)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:45 +0200
SubjectIn Budget Laptops the TOPS come with low energy footprint (Re: π-WAM is not adding decimals, it is removing decimals)
Message-ID<11481vf$h3a6$2@solani.org>
In reply to#348263
Hi,

Because of the mobile GPU design, the TOPS,
aka “Trillions of Operations Per Second.”
come with not extremly high power consumption.

Especially the presence of vector and matrix
operations can lower the energy consumption,
since they can avoid redundant memory access.

Its quite a difference between discrete graphic
cards and accelerator iGPUs that are directly
on the silicon chip, and have mobile design.

So basically with newer AI Laptops you get more
performence units for less energy units.

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
> Come on Horsy Boy, you can do better. I
> no where wrote something about curve
> fitting and/or increasing the precision of
> 
> float point numbers:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> What makes you think LIPS measures precision?
> You should know better as a 50% Prologer.
> 
> I explictily wrote here what the goal is:
> 
> "shave off some of the TOPS to do Prolog inferencing"
> 
> What are TOPS? Its a metric for GPUs:
> 
> TOPS stands for “Trillions of Operations Per Second.”
> https://www.lenovo.com/us/en/glossary/tops-in-computing/
> 
> See for yourself what is behind my post:
> 
> 11.4 Giga Lips with a Budget Laptop
> At the end of 2025 we acquired a couple of AI Laptops , that were still 
> cheap, since RAM prices had not yet rocketed. The intend was to tap into 
> the Copilot+ certified hardware, and shave off some of the TOPS to do 
> Prolog inferencing. Amazingly our π-WAM can churn 11.4 GIGA LIPS.
> 
> GPUs have evolved form lock-step to independent thread scheduling. This 
> made it possible to port the Hack VM variant, that forms the basis for 
> our π-WAM, to WebGPU computer shaders. Using NUM_SHADERS = 4096 we could 
> produce 11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M.
> 
> See also:
> 
> Medium Article - 11.4 Giga Lips
> https://medium.com/2989/899b0d5c027b
> 
> So just get lost with your crazy irrelevant rant.
> When I get more LIPS, things run faster, and
> I remove digits from the time dimension.
> 
> Got it. Or are you too stupid?
> 
> Bye
> 
> R Kym Horsell schrieb:
>> In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
>> ...
>>> Data centers should pay a 10000% excise on electricity,
>>> wherever it comes from, a natural regulator of inverted economies.
>>> And by ten thousand percent I really mean a ten thousand percent.
>>
>> And what would a huge surcharge do?
>> Almost always end up affecting the less powerful end of society
>> with increased costs to services the AI industry will be doing
>> more and more of over time.
>>
>> I started a little data center (exaflops.com) many years ago.
>> In those distant days people (in fact one was a prof of computer
>> science) told me you could never make money running a supercomputer.
>> LOL. :)
>>
>> I've had many years to watch the trends and a far more efficient
>> way to solve resource problems in this area is to change the
>> algorithms. There is vast room for improvement, mostly because
>> of prevailing attitudes.
>>
>> I used to do competetion data science as a sideline. Companies
>> would pay almost any price to get an extra decimal place in
>> the accuracy of their forecasting processes. But typically
>> they were trying to supercharge a system that should be scrapped
>> and re-designed from scratch. One area I'm thinking of is
>> investment. I had a customer one time -- like many times --
>> ask to improve a system that predicted the future price of
>> various stocks. The idea (for them) was to have as accurate a
>> prediction of what some stock would be worth in a week or a month's
>> time so that some moron could use the information to decide when
>> to buy or sell the thing.
>>
>> I tried to argue the efficient thing was to create a system that
>> takes the human out of the loop altogether. It doesnt provide info
>> for someone to decide whether or not to follow the advice --
>> that is just introducing more noise into the loop and probably
>> cancels any benefit of adding a couple decimal places of precision.
>> What you *should* do is make a system that is tuned to robustly
>> maximize the profit from managing a portfolio.
>>
>> Of course they wouldnt come at that. You can't suggest taking the
>> managers out of the loop. :)
>>
>> Another idea relevant to current AI methods might be to curtail
>> use of typical neural net algorithms. Many of them try to squeeze
>> the best performance of some NN during the training  phase in
>> the hope the resulting system will generalize well enough to be useful
>> on new data. But there's kind-of a law that the harder you train
>> some system to perform a task well, the less well they can subsuently
>> perform a more general version of the same thing. It's amusing when
>> you look at the graphs of NN being trained and then tested that
>> given a more general problem to solve after being trained to solve
>> similar problems very very well the poor old NN does worse that it
>> would have done if it had 0 training in the first place.
>>
>> It's not like we dont know how to improve this kind of performance.
>> Try less hard in the training phase or make it "more noisy".
>> Turns out genetic methods are just the ticket for this.
>> The training produces less over-fitting and the resulting system
>> generalizes better than it did before training and more importantly
>> it takes maybe an order of magnitude crunching to produce a good answer
>> than the usual over-fit answer.
>>
>> Anyway. Have to go and feed the cat.
>>
> 

[toc] | [prev] | [next] | [standalone]


#348255 — Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Was: Hurry the blue bus doesnt stop indefinitely)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 16:56 +0200
SubjectPotato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Was: Hurry the blue bus doesnt stop indefinitely)
Message-ID<1147rj3$gdpk$1@solani.org>
In reply to#348233
Hi,

Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world

country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.

The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:

1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI

Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its

keyboard and some telephathy module ?

Bye

Mild Shock schrieb:
> Hi,
> 
> Ride the snake
> He's old and his skin is cold
> The west is the best
> The west is the best
> Get here and we'll do the rest
> The blue bus is calling us
> The blue bus is calling us
> Driver, where you taking us?
> 
> Apocalypse Now intro: The Doors, The End {1979}
> https://www.youtube.com/watch?v=CIrvSJwwJUE
> 
> Bye
> 
>  > Hi,
>  >
>  > Again I posted everything here:
>  >
>  >> 11.4 Giga Lips with a Budget Laptop
>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>  >
>  > The repo says, same time when I posted
>  > the link first time:
>  >
>  >> This repository was archived by the
>  >> owner on Jul 9, 2026. It is now read-only.
>  >
>  > Now a USENET user, who had already entitled
>  > himself for a couple of irrational accusations
>  >
>  > towards my side, is asking this question:
>  >
>  > Chris M. Thomasson schrieb, Jul 24, 2026
>  >> Show an outline of what you
>  >> need you compute shader to do?
>  >
>  > Bravo, thats a delay of a wooping 15 days.
>  >
>  > Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Remember when first all local AI was Python
>> and PyTorch APIs. And then suddently people started
>> using bare metal C/C++ Code. Here is the story:
>>
>> How it started:
>>
>> GPT-J or GPT-J-6B is an open-source large
>> language model (LLM) developed by EleutherAI
>> in 2021. As the name suggests, it is a
>> generative pre-trained transformer model
>> designed to produce human-like text that
>> continues from a prompt.
>> https://www.eleuther.ai/
>>
>> How it was going [Georgi Gerganov]:
>>
>> So a few days later comes out the LLaMA, I do
>> some calculations and I figure out “Okay, 65
>> billion parameters. You probably need about
>> 40 gigs of RAM, with 4-bit quantization. So
>> this can run on a MacBook. Why not do it?”
>>
>> Why I was able to do it so quickly - basically,
>> for all that I saw it’s pretty much GPT-J architecture
>> with some modifications, like some extra memorization
>> layers. It’s minor changes. Basically, again, the
>> existing code for the GPT-J, I just simply
>> modified it there, it happened pretty quickly.
>> https://changelog.com/podcast/532
>>
>> Georgi Gerganov, Bulgarian, now with Hugging
>> Face, ggml-cann also running on Chinese AI chips.
>> ggml Manifesto https://github.com/ggml-org/ggml
>>
>> Bye
> 

[toc] | [prev] | [next] | [standalone]


#348262 — Rossy Boy is neither Einstein nor Zweistein (Was: Potato Computer owner impressed by Ukraine Tech)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:25 +0200
SubjectRossy Boy is neither Einstein nor Zweistein (Was: Potato Computer owner impressed by Ukraine Tech)
Message-ID<11480pn$h2e6$1@solani.org>
In reply to#348255
Hi,

Rossy Boy is neither Einstein nor Zweistein.
He is not Einstein since Einstein is already dead:

Albert Einstein (1879 - 1955)
https://de.wikipedia.org/wiki/Albert_Einstein

He is also not Zweistein, since he doesn't
understand concepts such as:

- NVIDIA Volta ff. architecture

Also his hands are small, and his breath stinks,
and he lives in the basement of his mother.

Bye

Mild Shock schrieb:
> Hi,
> 
> Slowly I start understanding numbnuts like
> Rossy Boy who don't understand tech, although
> they are from UK and not from a 3rd world
> 
> country, and also I start understanding morons
> like Micro Penis, who are behind a curtain,
> and cannot access a lot of tech.
> 
> The same holds for SWI Prologs newest campaign
> that probably adresses some poor indians that
> have neither 5G nor Macs:
> 
> 1:38:01 The Kyiv keynote disaster
> https://www.youtube.com/watch?v=U8goS6B3BbI
> 
> Woa! Real time download of Scala, Closure,
> etc.. Whats the magic behind that? Some SWI
> point of sale, downloading it via its
> 
> keyboard and some telephathy module ?
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Ride the snake
>> He's old and his skin is cold
>> The west is the best
>> The west is the best
>> Get here and we'll do the rest
>> The blue bus is calling us
>> The blue bus is calling us
>> Driver, where you taking us?
>>
>> Apocalypse Now intro: The Doors, The End {1979}
>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>
>> Bye
>>
>>  > Hi,
>>  >
>>  > Again I posted everything here:
>>  >
>>  >> 11.4 Giga Lips with a Budget Laptop
>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>  >
>>  > The repo says, same time when I posted
>>  > the link first time:
>>  >
>>  >> This repository was archived by the
>>  >> owner on Jul 9, 2026. It is now read-only.
>>  >
>>  > Now a USENET user, who had already entitled
>>  > himself for a couple of irrational accusations
>>  >
>>  > towards my side, is asking this question:
>>  >
>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>  >> Show an outline of what you
>>  >> need you compute shader to do?
>>  >
>>  > Bravo, thats a delay of a wooping 15 days.
>>  >
>>  > Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Remember when first all local AI was Python
>>> and PyTorch APIs. And then suddently people started
>>> using bare metal C/C++ Code. Here is the story:
>>>
>>> How it started:
>>>
>>> GPT-J or GPT-J-6B is an open-source large
>>> language model (LLM) developed by EleutherAI
>>> in 2021. As the name suggests, it is a
>>> generative pre-trained transformer model
>>> designed to produce human-like text that
>>> continues from a prompt.
>>> https://www.eleuther.ai/
>>>
>>> How it was going [Georgi Gerganov]:
>>>
>>> So a few days later comes out the LLaMA, I do
>>> some calculations and I figure out “Okay, 65
>>> billion parameters. You probably need about
>>> 40 gigs of RAM, with 4-bit quantization. So
>>> this can run on a MacBook. Why not do it?”
>>>
>>> Why I was able to do it so quickly - basically,
>>> for all that I saw it’s pretty much GPT-J architecture
>>> with some modifications, like some extra memorization
>>> layers. It’s minor changes. Basically, again, the
>>> existing code for the GPT-J, I just simply
>>> modified it there, it happened pretty quickly.
>>> https://changelog.com/podcast/532
>>>
>>> Georgi Gerganov, Bulgarian, now with Hugging
>>> Face, ggml-cann also running on Chinese AI chips.
>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>
>>> Bye
>>
> 

[toc] | [prev] | [next] | [standalone]


#348278 — You are still chewing on SIMD. LoL (Re: Rossy Boy is neither Einstein nor Zweistein)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:12 +0200
SubjectYou are still chewing on SIMD. LoL (Re: Rossy Boy is neither Einstein nor Zweistein)
Message-ID<114d58q$kg61$2@solani.org>
In reply to#348262
Hi,

You are still chewing on SIMD. LoL

Ross Finlayson schrieb:
 > Then the idea is that any of those can be found and matched in
 > one "run", i.e. a stall-less, branch-less, call-less list of less than
 > a few or less than a few dozens or less than a few hundreds
 > instructions, the results "findings" in data and corresponding
 > "matchings" of expressions, that runs in less than one microsecond.

You cannot make the mental translation that if you have:

Ross Finlayson schrieb:
 > So, the context then is for register state and stack contents, that
 > the indicators of the above as "positive presence" then is to make
 > for that the adjustments to the offsets and extents and the shifts
 > is according to those, otherwise no-ops. Then the idea is that a

As independent logical thread state, that automatically MIMD follows?

Whats the problem to solve then?

Bye

Mild Shock schrieb:
> Hi,
> 
> Rossy Boy is neither Einstein nor Zweistein.
> He is not Einstein since Einstein is already dead:
> 
> Albert Einstein (1879 - 1955)
> https://de.wikipedia.org/wiki/Albert_Einstein
> 
> He is also not Zweistein, since he doesn't
> understand concepts such as:
> 
> - NVIDIA Volta ff. architecture
> 
> Also his hands are small, and his breath stinks,
> and he lives in the basement of his mother.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Slowly I start understanding numbnuts like
>> Rossy Boy who don't understand tech, although
>> they are from UK and not from a 3rd world
>>
>> country, and also I start understanding morons
>> like Micro Penis, who are behind a curtain,
>> and cannot access a lot of tech.
>>
>> The same holds for SWI Prologs newest campaign
>> that probably adresses some poor indians that
>> have neither 5G nor Macs:
>>
>> 1:38:01 The Kyiv keynote disaster
>> https://www.youtube.com/watch?v=U8goS6B3BbI
>>
>> Woa! Real time download of Scala, Closure,
>> etc.. Whats the magic behind that? Some SWI
>> point of sale, downloading it via its
>>
>> keyboard and some telephathy module ?
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Ride the snake
>>> He's old and his skin is cold
>>> The west is the best
>>> The west is the best
>>> Get here and we'll do the rest
>>> The blue bus is calling us
>>> The blue bus is calling us
>>> Driver, where you taking us?
>>>
>>> Apocalypse Now intro: The Doors, The End {1979}
>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>
>>> Bye
>>>
>>>  > Hi,
>>>  >
>>>  > Again I posted everything here:
>>>  >
>>>  >> 11.4 Giga Lips with a Budget Laptop
>>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>  >
>>>  > The repo says, same time when I posted
>>>  > the link first time:
>>>  >
>>>  >> This repository was archived by the
>>>  >> owner on Jul 9, 2026. It is now read-only.
>>>  >
>>>  > Now a USENET user, who had already entitled
>>>  > himself for a couple of irrational accusations
>>>  >
>>>  > towards my side, is asking this question:
>>>  >
>>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>>  >> Show an outline of what you
>>>  >> need you compute shader to do?
>>>  >
>>>  > Bravo, thats a delay of a wooping 15 days.
>>>  >
>>>  > Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Remember when first all local AI was Python
>>>> and PyTorch APIs. And then suddently people started
>>>> using bare metal C/C++ Code. Here is the story:
>>>>
>>>> How it started:
>>>>
>>>> GPT-J or GPT-J-6B is an open-source large
>>>> language model (LLM) developed by EleutherAI
>>>> in 2021. As the name suggests, it is a
>>>> generative pre-trained transformer model
>>>> designed to produce human-like text that
>>>> continues from a prompt.
>>>> https://www.eleuther.ai/
>>>>
>>>> How it was going [Georgi Gerganov]:
>>>>
>>>> So a few days later comes out the LLaMA, I do
>>>> some calculations and I figure out “Okay, 65
>>>> billion parameters. You probably need about
>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>> this can run on a MacBook. Why not do it?”
>>>>
>>>> Why I was able to do it so quickly - basically,
>>>> for all that I saw it’s pretty much GPT-J architecture
>>>> with some modifications, like some extra memorization
>>>> layers. It’s minor changes. Basically, again, the
>>>> existing code for the GPT-J, I just simply
>>>> modified it there, it happened pretty quickly.
>>>> https://changelog.com/podcast/532
>>>>
>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>> Face, ggml-cann also running on Chinese AI chips.
>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>
>>>> Bye
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348279 — Hurry Rossy Boy, the blue bus is waiting (Re: You are still chewing on SIMD. LoL)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:53 +0200
SubjectHurry Rossy Boy, the blue bus is waiting (Re: You are still chewing on SIMD. LoL)
Message-ID<114d7lj$ki33$1@solani.org>
In reply to#348278
Hi,

Hurry Rossy Boy, the blue bus is waiting.
There is a quite a hyperbole from here:

Tesla S1070 in 2008
700 Watts , 1 Terra Flop
SOLVE TOMORROW’S PROBLEMS TODAY
https://www.azken.com/download/Tesla_DS_S1070_EU.pdf

To here:

Blackwell GPU in 2026
575 Watts, 104.8 Terra Flops ( RTX 5090 )
 From Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell

But somehow the S1070 had already Massively-
Parallel, Many-Core Architecture, and forms
of MIMD, since it had 960 / 240 = 4 cores.

960 scalar processor cores (240 per GPU).
But possibly more resticted inside work
groups, than later NVIDIA Volta ff

architecture with independent thread state.

Bye

Disclaimer: The above is only a very rough
RTX 5090 spec. Its doesn't say what value
format and what vector/matrics ops were

used. Also energy consumption may vary.

Mild Shock schrieb:
> Hi,
> 
> You are still chewing on SIMD. LoL
> 
> Ross Finlayson schrieb:
>  > Then the idea is that any of those can be found and matched in
>  > one "run", i.e. a stall-less, branch-less, call-less list of less than
>  > a few or less than a few dozens or less than a few hundreds
>  > instructions, the results "findings" in data and corresponding
>  > "matchings" of expressions, that runs in less than one microsecond.
> 
> You cannot make the mental translation that if you have:
> 
> Ross Finlayson schrieb:
>  > So, the context then is for register state and stack contents, that
>  > the indicators of the above as "positive presence" then is to make
>  > for that the adjustments to the offsets and extents and the shifts
>  > is according to those, otherwise no-ops. Then the idea is that a
> 
> As independent logical thread state, that automatically MIMD follows?
> 
> Whats the problem to solve then?
> 
> Bye

[toc] | [prev] | [next] | [standalone]


#348280 — Look how they advertized CUDA and logical threads (Re: Hurry Rossy Boy, the blue bus is waiting)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:55 +0200
SubjectLook how they advertized CUDA and logical threads (Re: Hurry Rossy Boy, the blue bus is waiting)
Message-ID<114d7ol$ki33$3@solani.org>
In reply to#348279
Hi,

So what does NUM_SHADERS = 4096 shaders mean here?

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Its only the number of logical threads.

CUDA™ TEChNOLOGY UNLOCkS ThE POWER OF TESLA MANY-CORE PROCESSORS
The CUDA C compiler simplifies  many-core programming
by enabling code development in a high-level language
and optimizing code to run on systems without knowledge of
how many cores are in the hardware.

CUDA applications automatically take advantage of more
cores or fewer cores in a system, so they can scale from
entry-level notebook GPUs to high end GPUs in technical
workstations  and further into racks of GPUs in data
centers. This allows developers to

“code once” and deploy on a range of systems, as well as
scale forward in time as future GPUs deliver more
performance per watt and more cores per processor. The benefit
for software users is the opportunity to boost computing
performance simply by adding GPUs or using their

existing GPUs in new ways.
https://www.azken.com/download/Tesla_DS_S1070_EU.pdf

Bye

Mild Shock schrieb:
> Hi,
> 
> Hurry Rossy Boy, the blue bus is waiting.
> There is a quite a hyperbole from here:
> 
> Tesla S1070 in 2008
> 700 Watts , 1 Terra Flop
> SOLVE TOMORROW’S PROBLEMS TODAY
> https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
> 
> To here:
> 
> Blackwell GPU in 2026
> 575 Watts, 104.8 Terra Flops ( RTX 5090 )
>  From Volta To Blackwell
> https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell 
> 
> 
> But somehow the S1070 had already Massively-
> Parallel, Many-Core Architecture, and forms
> of MIMD, since it had 960 / 240 = 4 cores.
> 
> 960 scalar processor cores (240 per GPU).
> But possibly more resticted inside work
> groups, than later NVIDIA Volta ff
> 
> architecture with independent thread state.
> 
> Bye
> 
> Disclaimer: The above is only a very rough
> RTX 5090 spec. Its doesn't say what value
> format and what vector/matrics ops were
> 
> used. Also energy consumption may vary.
> 
> Mild Shock schrieb:
>> Hi,
>>
>> You are still chewing on SIMD. LoL
>>
>> Ross Finlayson schrieb:
>>  > Then the idea is that any of those can be found and matched in
>>  > one "run", i.e. a stall-less, branch-less, call-less list of less than
>>  > a few or less than a few dozens or less than a few hundreds
>>  > instructions, the results "findings" in data and corresponding
>>  > "matchings" of expressions, that runs in less than one microsecond.
>>
>> You cannot make the mental translation that if you have:
>>
>> Ross Finlayson schrieb:
>>  > So, the context then is for register state and stack contents, that
>>  > the indicators of the above as "positive presence" then is to make
>>  > for that the adjustments to the offsets and extents and the shifts
>>  > is according to those, otherwise no-ops. Then the idea is that a
>>
>> As independent logical thread state, that automatically MIMD follows?
>>
>> Whats the problem to solve then?
>>
>> Bye

[toc] | [prev] | [next] | [standalone]


#348281 — Forget any arithmetization of product FSA (Re: Look how they advertized CUDA and logical threads)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:58 +0200
SubjectForget any arithmetization of product FSA (Re: Look how they advertized CUDA and logical threads)
Message-ID<114d7vl$ki33$6@solani.org>
In reply to#348280
Hi,

Because of this parallelism you anyway
need to forget about any arithmetization
of product FSA (finite-state automata).

Just forget it. What modern GPU provide
is a kind of hirarchical viewpoint. You
can have barriers in groups etc..

So you can exercise control over your
mongolian horde of logical threads in
a kind of multilevel schema.

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
> So what does NUM_SHADERS = 4096 shaders mean here?
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Its only the number of logical threads.
> 
> CUDA™ TEChNOLOGY UNLOCkS ThE POWER OF TESLA MANY-CORE PROCESSORS
> The CUDA C compiler simplifies  many-core programming
> by enabling code development in a high-level language
> and optimizing code to run on systems without knowledge of
> how many cores are in the hardware.
> 
> CUDA applications automatically take advantage of more
> cores or fewer cores in a system, so they can scale from
> entry-level notebook GPUs to high end GPUs in technical
> workstations  and further into racks of GPUs in data
> centers. This allows developers to
> 
> “code once” and deploy on a range of systems, as well as
> scale forward in time as future GPUs deliver more
> performance per watt and more cores per processor. The benefit
> for software users is the opportunity to boost computing
> performance simply by adding GPUs or using their
> 
> existing GPUs in new ways.
> https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Hurry Rossy Boy, the blue bus is waiting.
>> There is a quite a hyperbole from here:
>>
>> Tesla S1070 in 2008
>> 700 Watts , 1 Terra Flop
>> SOLVE TOMORROW’S PROBLEMS TODAY
>> https://www.azken.com/download/Tesla_DS_S1070_EU.pdf
>>
>> To here:
>>
>> Blackwell GPU in 2026
>> 575 Watts, 104.8 Terra Flops ( RTX 5090 )
>>  From Volta To Blackwell
>> https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell 
>>
>>
>> But somehow the S1070 had already Massively-
>> Parallel, Many-Core Architecture, and forms
>> of MIMD, since it had 960 / 240 = 4 cores.
>>
>> 960 scalar processor cores (240 per GPU).
>> But possibly more resticted inside work
>> groups, than later NVIDIA Volta ff
>>
>> architecture with independent thread state.
>>
>> Bye
>>
>> Disclaimer: The above is only a very rough
>> RTX 5090 spec. Its doesn't say what value
>> format and what vector/matrics ops were
>>
>> used. Also energy consumption may vary.
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> You are still chewing on SIMD. LoL
>>>
>>> Ross Finlayson schrieb:
>>>  > Then the idea is that any of those can be found and matched in
>>>  > one "run", i.e. a stall-less, branch-less, call-less list of less 
>>> than
>>>  > a few or less than a few dozens or less than a few hundreds
>>>  > instructions, the results "findings" in data and corresponding
>>>  > "matchings" of expressions, that runs in less than one microsecond.
>>>
>>> You cannot make the mental translation that if you have:
>>>
>>> Ross Finlayson schrieb:
>>>  > So, the context then is for register state and stack contents, that
>>>  > the indicators of the above as "positive presence" then is to make
>>>  > for that the adjustments to the offsets and extents and the shifts
>>>  > is according to those, otherwise no-ops. Then the idea is that a
>>>
>>> As independent logical thread state, that automatically MIMD follows?
>>>
>>> Whats the problem to solve then?
>>>
>>> Bye
> 

[toc] | [prev] | [next] | [standalone]


#348283 — Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] (Re: You are still chewing on SIMD. LoL)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 20:06 +0200
SubjectRossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] (Re: You are still chewing on SIMD. LoL)
Message-ID<114dfer$ko2q$3@solani.org>
In reply to#348278
Hi,

Rossy Boys tears could cool a data center,
he thinks there exists no literature about
serial algorithms of parallel stuff, and

he also thinks normal forms lead to optimizing
something. LoL, what a utter bullshit. I did
alreay a serial implementation of a parallel

simulation of my pi-WAM. Just lookup the literature
about pi-calulus. I published it a few days ago,
its part of 2.2.4 released already:

Parallel π-WAM: An Interleaved Synchronous Emulator
https://medium.com/2989/0196089e143a

Whats your point, Rossy Boy? Except you post pretend
nonsense not knowing what you are doing?

Bye

Ross Finlayson schrieb:
 > No, troll, these are serial algorithms their optimized forms.
 >
 > Normal sorts of forms, ....
 >
 >
 > Yeah, everybody already figured out "interpreters" and
 > "programs" and "spawning".
 >
 > Go spawn yourself.
 >


Mild Shock schrieb:
> Hi,
> 
> You are still chewing on SIMD. LoL
> 
> Ross Finlayson schrieb:
>  > Then the idea is that any of those can be found and matched in
>  > one "run", i.e. a stall-less, branch-less, call-less list of less than
>  > a few or less than a few dozens or less than a few hundreds
>  > instructions, the results "findings" in data and corresponding
>  > "matchings" of expressions, that runs in less than one microsecond.
> 
> You cannot make the mental translation that if you have:
> 
> Ross Finlayson schrieb:
>  > So, the context then is for register state and stack contents, that
>  > the indicators of the above as "positive presence" then is to make
>  > for that the adjustments to the offsets and extents and the shifts
>  > is according to those, otherwise no-ops. Then the idea is that a
> 
> As independent logical thread state, that automatically MIMD follows?
> 
> Whats the problem to solve then?
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Rossy Boy is neither Einstein nor Zweistein.
>> He is not Einstein since Einstein is already dead:
>>
>> Albert Einstein (1879 - 1955)
>> https://de.wikipedia.org/wiki/Albert_Einstein
>>
>> He is also not Zweistein, since he doesn't
>> understand concepts such as:
>>
>> - NVIDIA Volta ff. architecture
>>
>> Also his hands are small, and his breath stinks,
>> and he lives in the basement of his mother.
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Slowly I start understanding numbnuts like
>>> Rossy Boy who don't understand tech, although
>>> they are from UK and not from a 3rd world
>>>
>>> country, and also I start understanding morons
>>> like Micro Penis, who are behind a curtain,
>>> and cannot access a lot of tech.
>>>
>>> The same holds for SWI Prologs newest campaign
>>> that probably adresses some poor indians that
>>> have neither 5G nor Macs:
>>>
>>> 1:38:01 The Kyiv keynote disaster
>>> https://www.youtube.com/watch?v=U8goS6B3BbI
>>>
>>> Woa! Real time download of Scala, Closure,
>>> etc.. Whats the magic behind that? Some SWI
>>> point of sale, downloading it via its
>>>
>>> keyboard and some telephathy module ?
>>>
>>> Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Ride the snake
>>>> He's old and his skin is cold
>>>> The west is the best
>>>> The west is the best
>>>> Get here and we'll do the rest
>>>> The blue bus is calling us
>>>> The blue bus is calling us
>>>> Driver, where you taking us?
>>>>
>>>> Apocalypse Now intro: The Doors, The End {1979}
>>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>>
>>>> Bye
>>>>
>>>>  > Hi,
>>>>  >
>>>>  > Again I posted everything here:
>>>>  >
>>>>  >> 11.4 Giga Lips with a Budget Laptop
>>>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>  >
>>>>  > The repo says, same time when I posted
>>>>  > the link first time:
>>>>  >
>>>>  >> This repository was archived by the
>>>>  >> owner on Jul 9, 2026. It is now read-only.
>>>>  >
>>>>  > Now a USENET user, who had already entitled
>>>>  > himself for a couple of irrational accusations
>>>>  >
>>>>  > towards my side, is asking this question:
>>>>  >
>>>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>>>  >> Show an outline of what you
>>>>  >> need you compute shader to do?
>>>>  >
>>>>  > Bravo, thats a delay of a wooping 15 days.
>>>>  >
>>>>  > Bye
>>>>
>>>> Mild Shock schrieb:
>>>>> Hi,
>>>>>
>>>>> Remember when first all local AI was Python
>>>>> and PyTorch APIs. And then suddently people started
>>>>> using bare metal C/C++ Code. Here is the story:
>>>>>
>>>>> How it started:
>>>>>
>>>>> GPT-J or GPT-J-6B is an open-source large
>>>>> language model (LLM) developed by EleutherAI
>>>>> in 2021. As the name suggests, it is a
>>>>> generative pre-trained transformer model
>>>>> designed to produce human-like text that
>>>>> continues from a prompt.
>>>>> https://www.eleuther.ai/
>>>>>
>>>>> How it was going [Georgi Gerganov]:
>>>>>
>>>>> So a few days later comes out the LLaMA, I do
>>>>> some calculations and I figure out “Okay, 65
>>>>> billion parameters. You probably need about
>>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>>> this can run on a MacBook. Why not do it?”
>>>>>
>>>>> Why I was able to do it so quickly - basically,
>>>>> for all that I saw it’s pretty much GPT-J architecture
>>>>> with some modifications, like some extra memorization
>>>>> layers. It’s minor changes. Basically, again, the
>>>>> existing code for the GPT-J, I just simply
>>>>> modified it there, it happened pretty quickly.
>>>>> https://changelog.com/podcast/532
>>>>>
>>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>>> Face, ggml-cann also running on Chinese AI chips.
>>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>>
>>>>> Bye
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348284 — I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 20:25 +0200
SubjectI don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator])
Message-ID<114dgj9$kora$3@solani.org>
In reply to#348283
Hi,

I don't use Rust, you are crazy. First of
all the parallel simulator is 100% written
in Prolog, should also run in ISO Prolog,

enhanced by a library(lists). Second I only
mentioned that WebGPU / WGSL, the language
there has a Rust inspired language.

Its not Rust. Whats wrong with you? Why do
you adress your weariness of life to me.
I am neither thief, nor can I help you

with your frustration, and histeric outbursts.
Maybe just be a man and jump off a bridge, idiot.
Or tame your frustration, usenet is not for

you alone, your stupid asshole.

Bye

Ross Finlayson schrieb:
 > 
https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835
 > I don't much care about Rust.
 >
 > .. gibberish ..
 >
 > Thief.

Mild Shock schrieb:
> Hi,
> 
> Rossy Boys tears could cool a data center,
> he thinks there exists no literature about
> serial algorithms of parallel stuff, and
> 
> he also thinks normal forms lead to optimizing
> something. LoL, what a utter bullshit. I did
> alreay a serial implementation of a parallel
> 
> simulation of my pi-WAM. Just lookup the literature
> about pi-calulus. I published it a few days ago,
> its part of 2.2.4 released already:
> 
> Parallel π-WAM: An Interleaved Synchronous Emulator
> https://medium.com/2989/0196089e143a
> 
> Whats your point, Rossy Boy? Except you post pretend
> nonsense not knowing what you are doing?
> 
> Bye
> 
> Ross Finlayson schrieb:
>  > No, troll, these are serial algorithms their optimized forms.
>  >
>  > Normal sorts of forms, ....
>  >
>  >
>  > Yeah, everybody already figured out "interpreters" and
>  > "programs" and "spawning".
>  >
>  > Go spawn yourself.
>  >
> 
> 
> Mild Shock schrieb:
>> Hi,
>>
>> You are still chewing on SIMD. LoL
>>
>> Ross Finlayson schrieb:
>>  > Then the idea is that any of those can be found and matched in
>>  > one "run", i.e. a stall-less, branch-less, call-less list of less than
>>  > a few or less than a few dozens or less than a few hundreds
>>  > instructions, the results "findings" in data and corresponding
>>  > "matchings" of expressions, that runs in less than one microsecond.
>>
>> You cannot make the mental translation that if you have:
>>
>> Ross Finlayson schrieb:
>>  > So, the context then is for register state and stack contents, that
>>  > the indicators of the above as "positive presence" then is to make
>>  > for that the adjustments to the offsets and extents and the shifts
>>  > is according to those, otherwise no-ops. Then the idea is that a
>>
>> As independent logical thread state, that automatically MIMD follows?
>>
>> Whats the problem to solve then?
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Rossy Boy is neither Einstein nor Zweistein.
>>> He is not Einstein since Einstein is already dead:
>>>
>>> Albert Einstein (1879 - 1955)
>>> https://de.wikipedia.org/wiki/Albert_Einstein
>>>
>>> He is also not Zweistein, since he doesn't
>>> understand concepts such as:
>>>
>>> - NVIDIA Volta ff. architecture
>>>
>>> Also his hands are small, and his breath stinks,
>>> and he lives in the basement of his mother.
>>>
>>> Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Slowly I start understanding numbnuts like
>>>> Rossy Boy who don't understand tech, although
>>>> they are from UK and not from a 3rd world
>>>>
>>>> country, and also I start understanding morons
>>>> like Micro Penis, who are behind a curtain,
>>>> and cannot access a lot of tech.
>>>>
>>>> The same holds for SWI Prologs newest campaign
>>>> that probably adresses some poor indians that
>>>> have neither 5G nor Macs:
>>>>
>>>> 1:38:01 The Kyiv keynote disaster
>>>> https://www.youtube.com/watch?v=U8goS6B3BbI
>>>>
>>>> Woa! Real time download of Scala, Closure,
>>>> etc.. Whats the magic behind that? Some SWI
>>>> point of sale, downloading it via its
>>>>
>>>> keyboard and some telephathy module ?
>>>>
>>>> Bye
>>>>
>>>> Mild Shock schrieb:
>>>>> Hi,
>>>>>
>>>>> Ride the snake
>>>>> He's old and his skin is cold
>>>>> The west is the best
>>>>> The west is the best
>>>>> Get here and we'll do the rest
>>>>> The blue bus is calling us
>>>>> The blue bus is calling us
>>>>> Driver, where you taking us?
>>>>>
>>>>> Apocalypse Now intro: The Doors, The End {1979}
>>>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>>>
>>>>> Bye
>>>>>
>>>>>  > Hi,
>>>>>  >
>>>>>  > Again I posted everything here:
>>>>>  >
>>>>>  >> 11.4 Giga Lips with a Budget Laptop
>>>>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>>  >
>>>>>  > The repo says, same time when I posted
>>>>>  > the link first time:
>>>>>  >
>>>>>  >> This repository was archived by the
>>>>>  >> owner on Jul 9, 2026. It is now read-only.
>>>>>  >
>>>>>  > Now a USENET user, who had already entitled
>>>>>  > himself for a couple of irrational accusations
>>>>>  >
>>>>>  > towards my side, is asking this question:
>>>>>  >
>>>>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>>>>  >> Show an outline of what you
>>>>>  >> need you compute shader to do?
>>>>>  >
>>>>>  > Bravo, thats a delay of a wooping 15 days.
>>>>>  >
>>>>>  > Bye
>>>>>
>>>>> Mild Shock schrieb:
>>>>>> Hi,
>>>>>>
>>>>>> Remember when first all local AI was Python
>>>>>> and PyTorch APIs. And then suddently people started
>>>>>> using bare metal C/C++ Code. Here is the story:
>>>>>>
>>>>>> How it started:
>>>>>>
>>>>>> GPT-J or GPT-J-6B is an open-source large
>>>>>> language model (LLM) developed by EleutherAI
>>>>>> in 2021. As the name suggests, it is a
>>>>>> generative pre-trained transformer model
>>>>>> designed to produce human-like text that
>>>>>> continues from a prompt.
>>>>>> https://www.eleuther.ai/
>>>>>>
>>>>>> How it was going [Georgi Gerganov]:
>>>>>>
>>>>>> So a few days later comes out the LLaMA, I do
>>>>>> some calculations and I figure out “Okay, 65
>>>>>> billion parameters. You probably need about
>>>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>>>> this can run on a MacBook. Why not do it?”
>>>>>>
>>>>>> Why I was able to do it so quickly - basically,
>>>>>> for all that I saw it’s pretty much GPT-J architecture
>>>>>> with some modifications, like some extra memorization
>>>>>> layers. It’s minor changes. Basically, again, the
>>>>>> existing code for the GPT-J, I just simply
>>>>>> modified it there, it happened pretty quickly.
>>>>>> https://changelog.com/podcast/532
>>>>>>
>>>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>>>> Face, ggml-cann also running on Chinese AI chips.
>>>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>>>
>>>>>> Bye
>>>>>
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348285 — Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: I don't use Rust, you are crazy)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 22:52 +0200
SubjectHack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: I don't use Rust, you are crazy)
Message-ID<114dp6a$kf2n$4@solani.org>
In reply to#348284
Hi,

 > Who exactly is the thief? Does this person
 > have stats in the Rogue class in dungeons
 > and dragons?

The conspiracy theory of a stealing of Torso VDBE,
by Rossy Boy, is probably a result of complete
ignorance of the Hack ecosystem.

Hack is a very popular computer science project,
with a couple of subprojects in hardware and
software. It goes also by the name Nand to Tetris,

and is programming language agnositic. You can do
Hack experiments in any programming language, be
it BASIC, ADA or Rust. Nobody cares.

The gist are projects like here, first to
educate yourself about Hack:

https://www.nand2tetris.org/course

And then to use Hack in different contexts:

https://www.nand2tetris.org/copy-of-talks

For didactic purposes, I used Hack for my WebGPU
experiment. I didn't even take a look at Torso
VDBE, why should I? Hack is nicely documented,

has even a book, and fusing the two 16-bit
instruction types A and D, into a single 32-bit
instruction stream, is nowhere patented.

Bye

Mild Shock schrieb:
> Hi,
> 
> I don't use Rust, you are crazy. First of
> all the parallel simulator is 100% written
> in Prolog, should also run in ISO Prolog,
> 
> enhanced by a library(lists). Second I only
> mentioned that WebGPU / WGSL, the language
> there has a Rust inspired language.
> 
> Its not Rust. Whats wrong with you? Why do
> you adress your weariness of life to me.
> I am neither thief, nor can I help you
> 
> with your frustration, and histeric outbursts.
> Maybe just be a man and jump off a bridge, idiot.
> Or tame your frustration, usenet is not for
> 
> you alone, your stupid asshole.
> 
> Bye
> 
> Ross Finlayson schrieb:
>  > 
> https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835 
> 
>  > I don't much care about Rust.
>  >
>  > .. gibberish ..
>  >
>  > Thief.
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Rossy Boys tears could cool a data center,
>> he thinks there exists no literature about
>> serial algorithms of parallel stuff, and
>>
>> he also thinks normal forms lead to optimizing
>> something. LoL, what a utter bullshit. I did
>> alreay a serial implementation of a parallel
>>
>> simulation of my pi-WAM. Just lookup the literature
>> about pi-calulus. I published it a few days ago,
>> its part of 2.2.4 released already:
>>
>> Parallel π-WAM: An Interleaved Synchronous Emulator
>> https://medium.com/2989/0196089e143a
>>
>> Whats your point, Rossy Boy? Except you post pretend
>> nonsense not knowing what you are doing?
>>
>> Bye
>>
>> Ross Finlayson schrieb:
>>  > No, troll, these are serial algorithms their optimized forms.
>>  >
>>  > Normal sorts of forms, ....
>>  >
>>  >
>>  > Yeah, everybody already figured out "interpreters" and
>>  > "programs" and "spawning".
>>  >
>>  > Go spawn yourself.
>>  >
>>
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> You are still chewing on SIMD. LoL
>>>
>>> Ross Finlayson schrieb:
>>>  > Then the idea is that any of those can be found and matched in
>>>  > one "run", i.e. a stall-less, branch-less, call-less list of less 
>>> than
>>>  > a few or less than a few dozens or less than a few hundreds
>>>  > instructions, the results "findings" in data and corresponding
>>>  > "matchings" of expressions, that runs in less than one microsecond.
>>>
>>> You cannot make the mental translation that if you have:
>>>
>>> Ross Finlayson schrieb:
>>>  > So, the context then is for register state and stack contents, that
>>>  > the indicators of the above as "positive presence" then is to make
>>>  > for that the adjustments to the offsets and extents and the shifts
>>>  > is according to those, otherwise no-ops. Then the idea is that a
>>>
>>> As independent logical thread state, that automatically MIMD follows?
>>>
>>> Whats the problem to solve then?
>>>
>>> Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Rossy Boy is neither Einstein nor Zweistein.
>>>> He is not Einstein since Einstein is already dead:
>>>>
>>>> Albert Einstein (1879 - 1955)
>>>> https://de.wikipedia.org/wiki/Albert_Einstein
>>>>
>>>> He is also not Zweistein, since he doesn't
>>>> understand concepts such as:
>>>>
>>>> - NVIDIA Volta ff. architecture
>>>>
>>>> Also his hands are small, and his breath stinks,
>>>> and he lives in the basement of his mother.
>>>>
>>>> Bye
>>>>
>>>> Mild Shock schrieb:
>>>>> Hi,
>>>>>
>>>>> Slowly I start understanding numbnuts like
>>>>> Rossy Boy who don't understand tech, although
>>>>> they are from UK and not from a 3rd world
>>>>>
>>>>> country, and also I start understanding morons
>>>>> like Micro Penis, who are behind a curtain,
>>>>> and cannot access a lot of tech.
>>>>>
>>>>> The same holds for SWI Prologs newest campaign
>>>>> that probably adresses some poor indians that
>>>>> have neither 5G nor Macs:
>>>>>
>>>>> 1:38:01 The Kyiv keynote disaster
>>>>> https://www.youtube.com/watch?v=U8goS6B3BbI
>>>>>
>>>>> Woa! Real time download of Scala, Closure,
>>>>> etc.. Whats the magic behind that? Some SWI
>>>>> point of sale, downloading it via its
>>>>>
>>>>> keyboard and some telephathy module ?
>>>>>
>>>>> Bye
>>>>>
>>>>> Mild Shock schrieb:
>>>>>> Hi,
>>>>>>
>>>>>> Ride the snake
>>>>>> He's old and his skin is cold
>>>>>> The west is the best
>>>>>> The west is the best
>>>>>> Get here and we'll do the rest
>>>>>> The blue bus is calling us
>>>>>> The blue bus is calling us
>>>>>> Driver, where you taking us?
>>>>>>
>>>>>> Apocalypse Now intro: The Doors, The End {1979}
>>>>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>>>>
>>>>>> Bye
>>>>>>
>>>>>>  > Hi,
>>>>>>  >
>>>>>>  > Again I posted everything here:
>>>>>>  >
>>>>>>  >> 11.4 Giga Lips with a Budget Laptop
>>>>>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>>>  >
>>>>>>  > The repo says, same time when I posted
>>>>>>  > the link first time:
>>>>>>  >
>>>>>>  >> This repository was archived by the
>>>>>>  >> owner on Jul 9, 2026. It is now read-only.
>>>>>>  >
>>>>>>  > Now a USENET user, who had already entitled
>>>>>>  > himself for a couple of irrational accusations
>>>>>>  >
>>>>>>  > towards my side, is asking this question:
>>>>>>  >
>>>>>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>>>>>  >> Show an outline of what you
>>>>>>  >> need you compute shader to do?
>>>>>>  >
>>>>>>  > Bravo, thats a delay of a wooping 15 days.
>>>>>>  >
>>>>>>  > Bye
>>>>>>
>>>>>> Mild Shock schrieb:
>>>>>>> Hi,
>>>>>>>
>>>>>>> Remember when first all local AI was Python
>>>>>>> and PyTorch APIs. And then suddently people started
>>>>>>> using bare metal C/C++ Code. Here is the story:
>>>>>>>
>>>>>>> How it started:
>>>>>>>
>>>>>>> GPT-J or GPT-J-6B is an open-source large
>>>>>>> language model (LLM) developed by EleutherAI
>>>>>>> in 2021. As the name suggests, it is a
>>>>>>> generative pre-trained transformer model
>>>>>>> designed to produce human-like text that
>>>>>>> continues from a prompt.
>>>>>>> https://www.eleuther.ai/
>>>>>>>
>>>>>>> How it was going [Georgi Gerganov]:
>>>>>>>
>>>>>>> So a few days later comes out the LLaMA, I do
>>>>>>> some calculations and I figure out “Okay, 65
>>>>>>> billion parameters. You probably need about
>>>>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>>>>> this can run on a MacBook. Why not do it?”
>>>>>>>
>>>>>>> Why I was able to do it so quickly - basically,
>>>>>>> for all that I saw it’s pretty much GPT-J architecture
>>>>>>> with some modifications, like some extra memorization
>>>>>>> layers. It’s minor changes. Basically, again, the
>>>>>>> existing code for the GPT-J, I just simply
>>>>>>> modified it there, it happened pretty quickly.
>>>>>>> https://changelog.com/podcast/532
>>>>>>>
>>>>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>>>>> Face, ggml-cann also running on Chinese AI chips.
>>>>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>>>>
>>>>>>> Bye
>>>>>>
>>>>>
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#348286 — A funny Q16.16 experiment with Hack (Re: Hack ecosystem ignorance paired with paranoia [Nand to Tetris])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 23:11 +0200
SubjectA funny Q16.16 experiment with Hack (Re: Hack ecosystem ignorance paired with paranoia [Nand to Tetris])
Message-ID<114dqan$kfq8$2@solani.org>
In reply to#348285
Hi,

This seems to be a funny Q16.16 experiment.
It shows that an integerish Hack can do
floatish stuff, by using binary fixpoint:

Raytracing on the Hack computer
2021/06/13 - im alex
https://blog.alexqua.ch/posts/from-nand-to-raytracer/

That it uses Rust is arbitrary. Feel free
to do it in C, C++, FORTRAN or Java. I guess
these languages all have basic arithmethic,

right? Maybe not a long jump always?

Bye

Mild Shock schrieb:
> Hi,
> 
>  > Who exactly is the thief? Does this person
>  > have stats in the Rogue class in dungeons
>  > and dragons?
> 
> The conspiracy theory of a stealing of Torso VDBE,
> by Rossy Boy, is probably a result of complete
> ignorance of the Hack ecosystem.
> 
> Hack is a very popular computer science project,
> with a couple of subprojects in hardware and
> software. It goes also by the name Nand to Tetris,
> 
> and is programming language agnositic. You can do
> Hack experiments in any programming language, be
> it BASIC, ADA or Rust. Nobody cares.
> 
> The gist are projects like here, first to
> educate yourself about Hack:
> 
> https://www.nand2tetris.org/course
> 
> And then to use Hack in different contexts:
> 
> https://www.nand2tetris.org/copy-of-talks
> 
> For didactic purposes, I used Hack for my WebGPU
> experiment. I didn't even take a look at Torso
> VDBE, why should I? Hack is nicely documented,
> 
> has even a book, and fusing the two 16-bit
> instruction types A and D, into a single 32-bit
> instruction stream, is nowhere patented.
> 
> Bye

[toc] | [prev] | [next] | [standalone]


#348288 — Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Was: A funny Q16.16 experiment with Hack)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-30 11:25 +0200
SubjectSummer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Was: A funny Q16.16 experiment with Hack)
Message-ID<114f5af$lao5$2@solani.org>
In reply to#348286
Hi,

Woa! Thats a very sad and non fitting statement:

"This was before I was indoctrinated into
ISO Prolog and the ways of monotonic logic
programming. Shen Prolog has many semantic
and syntactic limitations that Scryer Prolog
does not. Also, I now know constraints are a
much better, purer solution to the problems
mode declarations were meant to address"
https://github.com/mthom/scryer-prolog/issues/3410#issuecomment-5030471183

Ok, here is the summer challenge, thats the easy one:

     SQL --> Prolog --> WAN

Here come two variations, slightly mindboggling maybe?

     SQL --> AST --> VDBE

     SQL --> Prolog+Modes --> π-WAM

Bye

BTW: What is VDBE? Some abstract machine, that can
be used to run SQL, following some ideas here:

Database Co-Design With Asynchronous I/O
https://penberg.org/papers/penberg-edgesys24.pdf

Or to run Doom:

Doom on the Turso VDBE
https://github.com/tursodatabase/turso-vdbe-doom-example

What if we would run Doom with π-WAM, on a GPU,
using multiple shaders. We could add some ray tracing.

Mild Shock schrieb:
> Hi,
> 
> This seems to be a funny Q16.16 experiment.
> It shows that an integerish Hack can do
> floatish stuff, by using binary fixpoint:
> 
> Raytracing on the Hack computer
> 2021/06/13 - im alex
> https://blog.alexqua.ch/posts/from-nand-to-raytracer/
> 
> That it uses Rust is arbitrary. Feel free
> to do it in C, C++, FORTRAN or Java. I guess
> these languages all have basic arithmethic,
> 
> right? Maybe not a long jump always?
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>>  > Who exactly is the thief? Does this person
>>  > have stats in the Rogue class in dungeons
>>  > and dragons?
>>
>> The conspiracy theory of a stealing of Torso VDBE,
>> by Rossy Boy, is probably a result of complete
>> ignorance of the Hack ecosystem.
>>
>> Hack is a very popular computer science project,
>> with a couple of subprojects in hardware and
>> software. It goes also by the name Nand to Tetris,
>>
>> and is programming language agnositic. You can do
>> Hack experiments in any programming language, be
>> it BASIC, ADA or Rust. Nobody cares.
>>
>> The gist are projects like here, first to
>> educate yourself about Hack:
>>
>> https://www.nand2tetris.org/course
>>
>> And then to use Hack in different contexts:
>>
>> https://www.nand2tetris.org/copy-of-talks
>>
>> For didactic purposes, I used Hack for my WebGPU
>> experiment. I didn't even take a look at Torso
>> VDBE, why should I? Hack is nicely documented,
>>
>> has even a book, and fusing the two 16-bit
>> instruction types A and D, into a single 32-bit
>> instruction stream, is nowhere patented.
>>
>> Bye

[toc] | [prev] | [next] | [standalone]


#348289 — I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Hack ecosystem ignorance paired with paranoia)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-30 19:37 +0200
SubjectI wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Hack ecosystem ignorance paired with paranoia)
Message-ID<114g24j$m15d$3@solani.org>
In reply to#348285
Hi,

 > Just downloading some other person's code

I didn't do that, I wrote Hack VM for pi-WAM
from scratch, over the last 4 weeks. I came
back from holidays on end of June 2026, and now

we have end of July 2026. But its only possible
because the instruction set is very smal, like
ca. 8 functions and ca. 8 modes and ca. 8 conditions,

so its ca. 8 x 8 x 8 = 512 opcodes, each has an
A parameter and a D parameter simultaneously.
It has currently the following CPU backends:

  - Now supports interleaved synchronous emulation.
  - Now supports warp parallelism via Java platform threads.
  - Now supports warp parallelism via Python system threads.
  - Now supports warp parallelism via JavaScript worker threads.
  - Note: For Python free threads are not yet fully tested.
  - Note: For JavaScript web workers are not yet fully tested.

https://www.dogelog.ch/typtab/doclet/book/14_install/05_notes22/110_224.html

But frankly I came to encounter Hack not from
the usual university curriculum web resources,
but indirectly through a post about a Prolog

emulation of Hack, using constrained horn clauses (CHC):

Verifying Nand2Tetris Assembly
https://www.philipzucker.com/nand2tetris-chc/

The binary encoding is currently that the functions,
modes and conditions eat up a nibble (4-bit), in
total 12-bit, which I use then 10-bit for A parameter

and 10-bit for D parameter. I used AI freemium, Codex
by ChatGPT from within IntelliJ to do some fragment
code translations automatically from Java to JavaScript

or from JavaScript to Python.

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
>  > Who exactly is the thief? Does this person
>  > have stats in the Rogue class in dungeons
>  > and dragons?
> 
> The conspiracy theory of a stealing of Torso VDBE,
> by Rossy Boy, is probably a result of complete
> ignorance of the Hack ecosystem.
> 
> Hack is a very popular computer science project,
> with a couple of subprojects in hardware and
> software. It goes also by the name Nand to Tetris,
> 
> and is programming language agnositic. You can do
> Hack experiments in any programming language, be
> it BASIC, ADA or Rust. Nobody cares.
> 
> The gist are projects like here, first to
> educate yourself about Hack:
> 
> https://www.nand2tetris.org/course
> 
> And then to use Hack in different contexts:
> 
> https://www.nand2tetris.org/copy-of-talks
> 
> For didactic purposes, I used Hack for my WebGPU
> experiment. I didn't even take a look at Torso
> VDBE, why should I? Hack is nicely documented,
> 
> has even a book, and fusing the two 16-bit
> instruction types A and D, into a single 32-bit
> instruction stream, is nowhere patented.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> I don't use Rust, you are crazy. First of
>> all the parallel simulator is 100% written
>> in Prolog, should also run in ISO Prolog,
>>
>> enhanced by a library(lists). Second I only
>> mentioned that WebGPU / WGSL, the language
>> there has a Rust inspired language.
>>
>> Its not Rust. Whats wrong with you? Why do
>> you adress your weariness of life to me.
>> I am neither thief, nor can I help you
>>
>> with your frustration, and histeric outbursts.
>> Maybe just be a man and jump off a bridge, idiot.
>> Or tame your frustration, usenet is not for
>>
>> you alone, your stupid asshole.
>>
>> Bye
>>
>> Ross Finlayson schrieb:
>>  > 
>> https://www.theregister.com/databases/2026/07/29/after-rewriting-sqlite-in-rust-turso-turns-its-sights-on-postgres/5279835 
>>
>>  > I don't much care about Rust.
>>  >
>>  > .. gibberish ..
>>  >
>>  > Thief.
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Rossy Boys tears could cool a data center,
>>> he thinks there exists no literature about
>>> serial algorithms of parallel stuff, and
>>>
>>> he also thinks normal forms lead to optimizing
>>> something. LoL, what a utter bullshit. I did
>>> alreay a serial implementation of a parallel
>>>
>>> simulation of my pi-WAM. Just lookup the literature
>>> about pi-calulus. I published it a few days ago,
>>> its part of 2.2.4 released already:
>>>
>>> Parallel π-WAM: An Interleaved Synchronous Emulator
>>> https://medium.com/2989/0196089e143a
>>>
>>> Whats your point, Rossy Boy? Except you post pretend
>>> nonsense not knowing what you are doing?
>>>
>>> Bye
>>>
>>> Ross Finlayson schrieb:
>>>  > No, troll, these are serial algorithms their optimized forms.
>>>  >
>>>  > Normal sorts of forms, ....
>>>  >
>>>  >
>>>  > Yeah, everybody already figured out "interpreters" and
>>>  > "programs" and "spawning".
>>>  >
>>>  > Go spawn yourself.
>>>  >
>>>
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> You are still chewing on SIMD. LoL
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > Then the idea is that any of those can be found and matched in
>>>>  > one "run", i.e. a stall-less, branch-less, call-less list of less 
>>>> than
>>>>  > a few or less than a few dozens or less than a few hundreds
>>>>  > instructions, the results "findings" in data and corresponding
>>>>  > "matchings" of expressions, that runs in less than one microsecond.
>>>>
>>>> You cannot make the mental translation that if you have:
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > So, the context then is for register state and stack contents, that
>>>>  > the indicators of the above as "positive presence" then is to make
>>>>  > for that the adjustments to the offsets and extents and the shifts
>>>>  > is according to those, otherwise no-ops. Then the idea is that a
>>>>
>>>> As independent logical thread state, that automatically MIMD follows?
>>>>
>>>> Whats the problem to solve then?
>>>>
>>>> Bye
>>>>
>>>> Mild Shock schrieb:
>>>>> Hi,
>>>>>
>>>>> Rossy Boy is neither Einstein nor Zweistein.
>>>>> He is not Einstein since Einstein is already dead:
>>>>>
>>>>> Albert Einstein (1879 - 1955)
>>>>> https://de.wikipedia.org/wiki/Albert_Einstein
>>>>>
>>>>> He is also not Zweistein, since he doesn't
>>>>> understand concepts such as:
>>>>>
>>>>> - NVIDIA Volta ff. architecture
>>>>>
>>>>> Also his hands are small, and his breath stinks,
>>>>> and he lives in the basement of his mother.
>>>>>
>>>>> Bye
>>>>>
>>>>> Mild Shock schrieb:
>>>>>> Hi,
>>>>>>
>>>>>> Slowly I start understanding numbnuts like
>>>>>> Rossy Boy who don't understand tech, although
>>>>>> they are from UK and not from a 3rd world
>>>>>>
>>>>>> country, and also I start understanding morons
>>>>>> like Micro Penis, who are behind a curtain,
>>>>>> and cannot access a lot of tech.
>>>>>>
>>>>>> The same holds for SWI Prologs newest campaign
>>>>>> that probably adresses some poor indians that
>>>>>> have neither 5G nor Macs:
>>>>>>
>>>>>> 1:38:01 The Kyiv keynote disaster
>>>>>> https://www.youtube.com/watch?v=U8goS6B3BbI
>>>>>>
>>>>>> Woa! Real time download of Scala, Closure,
>>>>>> etc.. Whats the magic behind that? Some SWI
>>>>>> point of sale, downloading it via its
>>>>>>
>>>>>> keyboard and some telephathy module ?
>>>>>>
>>>>>> Bye
>>>>>>
>>>>>> Mild Shock schrieb:
>>>>>>> Hi,
>>>>>>>
>>>>>>> Ride the snake
>>>>>>> He's old and his skin is cold
>>>>>>> The west is the best
>>>>>>> The west is the best
>>>>>>> Get here and we'll do the rest
>>>>>>> The blue bus is calling us
>>>>>>> The blue bus is calling us
>>>>>>> Driver, where you taking us?
>>>>>>>
>>>>>>> Apocalypse Now intro: The Doors, The End {1979}
>>>>>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>>>>>
>>>>>>> Bye
>>>>>>>
>>>>>>>  > Hi,
>>>>>>>  >
>>>>>>>  > Again I posted everything here:
>>>>>>>  >
>>>>>>>  >> 11.4 Giga Lips with a Budget Laptop
>>>>>>>  >> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>>>>  >
>>>>>>>  > The repo says, same time when I posted
>>>>>>>  > the link first time:
>>>>>>>  >
>>>>>>>  >> This repository was archived by the
>>>>>>>  >> owner on Jul 9, 2026. It is now read-only.
>>>>>>>  >
>>>>>>>  > Now a USENET user, who had already entitled
>>>>>>>  > himself for a couple of irrational accusations
>>>>>>>  >
>>>>>>>  > towards my side, is asking this question:
>>>>>>>  >
>>>>>>>  > Chris M. Thomasson schrieb, Jul 24, 2026
>>>>>>>  >> Show an outline of what you
>>>>>>>  >> need you compute shader to do?
>>>>>>>  >
>>>>>>>  > Bravo, thats a delay of a wooping 15 days.
>>>>>>>  >
>>>>>>>  > Bye
>>>>>>>
>>>>>>> Mild Shock schrieb:
>>>>>>>> Hi,
>>>>>>>>
>>>>>>>> Remember when first all local AI was Python
>>>>>>>> and PyTorch APIs. And then suddently people started
>>>>>>>> using bare metal C/C++ Code. Here is the story:
>>>>>>>>
>>>>>>>> How it started:
>>>>>>>>
>>>>>>>> GPT-J or GPT-J-6B is an open-source large
>>>>>>>> language model (LLM) developed by EleutherAI
>>>>>>>> in 2021. As the name suggests, it is a
>>>>>>>> generative pre-trained transformer model
>>>>>>>> designed to produce human-like text that
>>>>>>>> continues from a prompt.
>>>>>>>> https://www.eleuther.ai/
>>>>>>>>
>>>>>>>> How it was going [Georgi Gerganov]:
>>>>>>>>
>>>>>>>> So a few days later comes out the LLaMA, I do
>>>>>>>> some calculations and I figure out “Okay, 65
>>>>>>>> billion parameters. You probably need about
>>>>>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>>>>>> this can run on a MacBook. Why not do it?”
>>>>>>>>
>>>>>>>> Why I was able to do it so quickly - basically,
>>>>>>>> for all that I saw it’s pretty much GPT-J architecture
>>>>>>>> with some modifications, like some extra memorization
>>>>>>>> layers. It’s minor changes. Basically, again, the
>>>>>>>> existing code for the GPT-J, I just simply
>>>>>>>> modified it there, it happened pretty quickly.
>>>>>>>> https://changelog.com/podcast/532
>>>>>>>>
>>>>>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>>>>>> Face, ggml-cann also running on Chinese AI chips.
>>>>>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>>>>>
>>>>>>>> Bye
>>>>>>>
>>>>>>
>>>>>
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →

Back to top | Article view | sci.logic


csiph-web