Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > sci.math > #646933 > unrolled thread

The Wuhan Virus that destroyed Python [ggml Manifesto]

Started byMild Shock <janburse@fastmail.fm>
First post2026-07-22 21:00 +0200
Last post2026-07-29 17:05 +0200
Articles 20 on this page of 76 — 7 participants

Back to article view | Back to sci.math


Contents

  The Wuhan Virus that destroyed Python [ggml Manifesto] Mild Shock <janburse@fastmail.fm> - 2026-07-22 21:00 +0200
    Deadlock Exorcism: Switch from Push to Pull [A pi-calculus Specification of Prolog] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:23 +0200
      Why do you even need a mpmc queue? [Thunder Kittens] (Re: Deadlock Exorcism: Switch from Push to Pull) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:43 +0200
        Trivial balancing example for (int i=0; i<global_id; i++) (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:57 +0200
          Enqueue/dequeue need not be fast and can spinn ["fairness" questions] (Was: Trivial balancing example for (int i=0; i<global_id; i++)) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:11 +0200
            The Pixel Phone AI Experiment Song (Re: Enqueue/dequeue need not be fast and can spinn ["fairness" questions] ) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:21 +0200
        Re: Why do you even need a mpmc queue? [Thunder Kittens] (Re: Deadlock Exorcism: Switch from Push to Pull) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-23 08:24 -0700
    Potential Python Recovery: Free Threading [3.13 release] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 10:19 +0200
      Re: Potential Python Recovery: Free Threading [3.13 release] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Ross Valikhanov <kavna@rl.ru> - 2026-07-23 16:01 +0000
    Re: The Wuhan Virus that destroyed Python [ggml Manifesto] Ramon Dubenkov <omd@nnk.ru> - 2026-07-23 13:38 +0000
    The things XILINX braught to the AMD table (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 18:47 +0200
      NIVIDIA evacuated its Chinese market [Tau Scaling] (Was: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-23 19:11 +0200
      NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-23 19:12 +0200
        Re: NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Lane W <cactus_DAC@yahoo.com> - 2026-07-23 11:22 -0600
          Micro penis mother sung arias (Was: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:38 +0200
            Re: Micro penis mother sung arias (Was: NVIDIA evacuated its Chinese market [Tau Scaling]) Lane W <cactus_DAC@yahoo.com> - 2026-07-24 07:15 -0600
              Micro penis brain is in constant hiatus (Was: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:24 +0200
                Re: Micro penis brain is in constant hiatus (Was: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:36 +0200
                Ignoramus or Ignorabimus: I don't care (π-WAM) (Re: Micro penis brain is in constant hiatus) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:38 +0200
                  Re: Ignoramus or Ignorabimus: I don't care (π-WAM) (Re: Micro penis brain is in constant hiatus) Lane W <cactus_DAC@yahoo.com> - 2026-07-24 08:31 -0600
                    You are a moron, brainless putin payed (Was: Ignoramus or Ignorabimus: I don't care (π-WAM)) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:01 +0200
                      Re: You are a moron, brainless putin payed (Was: Ignoramus or Ignorabimus: I don't care (π-WAM)) Lane W <cactus_DAC@yahoo.com> - 2026-07-24 10:27 -0600
                        Yeah keep reading my posts, uninspired fool (Was: You are a moron, brainless putin payed) Mild Shock <janburse@fastmail.fm> - 2026-07-24 19:45 +0200
                          Re: Yeah keep reading my posts, uninspired fool (Was: You are a moron, brainless putin payed) Lane W <cactus_DAC@yahoo.com> - 2026-07-24 12:11 -0600
                            LoL (Was: Yeah keep reading my posts, uninspired fool ) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:12 +0200
                              Re: LoL (Was: Yeah keep reading my posts, uninspired fool ) Lane W <cactus_DAC@yahoo.com> - 2026-07-24 12:53 -0600
                  Out of the blue accusation span 15 days [Empirical USENET study] (Was: Ignoramus or Ignorabimus: I don't care (π-WAM)) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:26 +0200
                  A brain desease of 20 days [Rossy Boy] (Re: Ignoramus or Ignorabimus: I don't care (π-WAM)) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:40 +0200
        ASML stocks are plunging, bye bye dutchies (Was: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-28 14:17 +0200
    Little Data Center on Your Palm [AI Laptops for 500 USD] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto] Mild Shock <janburse@fastmail.fm> - 2026-07-24 17:58 +0200
      2008: 4 Blades + Tesla S1070 versus 2026: 1 AI Laptop (Re: Little Data Center on Your Palm [AI Laptops for 500 USD]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:16 +0200
      Re: Little Data Center on Your Palm [AI Laptops for 500 USD] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto] Bradford Babkoff <ffb@odbb.ru> - 2026-07-24 18:05 +0000
        LoL (Was: Little Data Center on Your Palm [AI Laptops for 500 USD]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:11 +0200
    Hurry the blue bus doesnt stop indefinitely (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:36 +0200
      Not SIMD, a MIMD design for NVIDIA Volta (Re: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:57 +0200
        Could take 3-4 months find machine / browser (Was Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-24 21:15 +0200
        The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-26 19:52 +0200
          The turbo capping of AI Laptops (Re: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-26 20:01 +0200
          Re: The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-26 20:33 -0700
            Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:14 +0200
              miniTriton CUDA is an alternative to torch variants (Was: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:40 +0200
                Andrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:51 +0200
                  Re: Andrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 01:41 -0700
              Re: Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 01:38 -0700
              The evolution of hardware and GPT-2 training (Was: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 10:56 +0200
                How speed up π-WAM with vector operations (Was: The evolution of hardware and GPT-2 training) Mild Shock <janburse@fastmail.fm> - 2026-07-27 11:08 +0200
                  AI accelerator extend from GPU to CPU [Zero Copying] (Was: How speed up π-WAM with vector operations) Mild Shock <janburse@fastmail.fm> - 2026-07-27 11:21 +0200
                    The invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:20 +0200
                      Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Re: The invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 17:14 +0200
                      π-WAM is not adding decimals, it is removing decimals (Re: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:36 +0200
                        In Budget Laptops the TOPS come with low energy footprint (Was: π-WAM is not adding decimals, it is removing decimals) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:44 +0200
          Java picky concerning JIT-ing [Luckier with C++/C or FORTRAN compilers?] (Re: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 12:55 +0200
      Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Re: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-27 16:57 +0200
    Got it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 19:00 +0200
      Re: Got it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Kim Baitchorov <bvhkoc@bmc.ru> - 2026-07-27 22:38 +0000
        Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-07-28 11:29 +0200
          confused rossy boy is confused (Re: Clueless about MIMD as usual [Flynn's Taxonomy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:20 +0200
            Gemini, DeepSeek, OpenAI more clever than rossy boy (Re: confused rossy boy is confused) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:21 +0200
              In AI Acceleration nobody cares about CivetWeb (Re: Gemini, DeepSeek, OpenAI more clever than rossy boy) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:29 +0200
                Run with minimum HTTPS and .mjs type (Re: In AI Acceleration nobody cares about CivetWeb) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:49 +0200
                  Your strictness is your problem , not mine [See WebLLM] (Re: Run with minimum HTTPS and .mjs type) Mild Shock <janburse@fastmail.fm> - 2026-07-29 11:51 +0200
                Re: In AI Acceleration nobody cares about CivetWeb (Re: Gemini, DeepSeek, OpenAI more clever than rossy boy) Lane W <cactus_DAC@yahoo.com> - 2026-07-29 07:06 -0600
                Re: In AI Acceleration nobody cares about CivetWeb (Re: Gemini, DeepSeek, OpenAI more clever than rossy boy) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 07:20 -0700
          You are still chewing on SIMD. LoL (Re: Clueless about MIMD as usual [Flynn's Taxonomy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:14 +0200
            Hurry Rossy Boy, the blue bus is waiting (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:54 +0200
              Look how they advertized CUDA and logical threads (Re: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:56 +0200
                Forget any arithmetization of product FSA (Re: Look how they advertized CUDA and logical threads) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:57 +0200
            Re: You are still chewing on SIMD. LoL (Re: Clueless about MIMD as usual [Flynn's Taxonomy]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 10:48 -0700
              Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] (Was: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:03 +0200
                I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:24 +0200
                  Re: I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 12:27 -0700
                    Re: I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-29 13:40 -0700
                Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: Rossy Boys tears could cool a data center) Mild Shock <janburse@fastmail.fm> - 2026-07-29 22:51 +0200
    Lamas in a cradle and Lamas on the edge [Red Pyjama] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 13:02 +0200
      AI Accelerators and ISO Prolog multi-threading (Was: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:01 +0200
        Actor/Erlang is dead, no Thread and Mailbox conflation [golang channels] (Was: AI Accelerators and ISO Prolog multi-threading) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:05 +0200

Page 3 of 4 — ← Prev page 1 2 [3] 4  Next page →


#647003 — miniTriton CUDA is an alternative to torch variants (Was: Why forget Bulgarians, never on my mind)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 09:40 +0200
SubjectminiTriton CUDA is an alternative to torch variants (Was: Why forget Bulgarians, never on my mind)
Message-ID<114721k$gc17$1@solani.org>
In reply to#647002
Hi,

Some counter PyTorch Python trends are
for example OpenAIs Triton. And the variant
miniTriton CUDA vibe produced by Kimi K3 (sic!):

"We further tested whether Kimi K3 could build
a GPU programming system from scratch. Kimi K3
developed MiniTriton, a compact Triton-like
compiler with its own tile-level IR layer over
MLIR, optimization passes, and a PTX code-
generation pipeline.

Across supported roofline benchmarks, MiniTriton
delivers performance on par with or better than
Triton and torch.compile — beating Triton on
certain workloads. Beyond microbenchmarks,
MiniTriton sustains end-to-end nanoGPT training
with stable convergence, the loss curve

closely tracking the reference with only minor
divergence — validating the full pipeline on a
realistic workload. These results demonstrate
that Kimi K3 can build a coherent end-to-end
compiler — from DSL frontend and IR passes to
PTX codegen and runtime — rather than isolated

kernels; its from-scratch Tensor Core path
already rivals Triton’s extensively optimized stack."

GPU Compiler Development
https://www.kimi.com/blog/kimi-k3

Although many GPU corporate stuff is anonymized,
and some AI papers have lists of 30 authors. Here
nanoGPT is mentioned which is tied to the name

Andrej Karpathy. See also here:

Update Nov 2025 nanoGPT has a new and
improved cousin called nanochat.
https://github.com/karpathy/nanogpt

But as can be seen, he moved on to another project.

Bye

Mild Shock schrieb:
> Hi,
> 
> Whats this "forget" trope of glue sniffing
> Rossy Boy with his herpes blisters?
> 
>  > Bulgarians, that's some real Boris and Natasha crap,
>  > forget Hungarians and Bulgarians.
> 
> Why should I forget Bulgarians,
> they are never on my mind. Do you
> see me doing ggml stuff?
> 
> I only hypothesized that it is
> over for Python as the machine
> learning language or AI inferencing
> 
> locally on AI laptops language, and
> made the ggml case, so I already forgot
> about them. Which might give you a glimps,
> 
> why WebGPU was used for this here:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Is an interesting choice. Even
> github has some Languages statistics,
> giving an account what I used:
> 
> HTML 67.5% JavaScript 23.1% CSS 9.4%
> 
> Have Fun!
> 
> Bye
> 
> P.S.: The example below is not p-adics,
> you complete imbecil moron. Its just:
> 
> 7-11 cubic Solution by Pritchard & Gries
> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
> 
> Ross Finlayson schrieb:
>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> You see it all boils down to find your inner peace
>>> by an immaculate inception of some queue datatype.
>>>
>>> KOAN/Fortran-S was an early 1990s research programming
>>> system for distributed-memory multiprocessors . Developed
>>> at ENS Lyon in the early 1990s . Often listed alongside
>>> other historical parallel programming efforts.
>>>
>>> The Message Passing: The research explicitly
>>> compared the SVM approach against message passing
>>> on the same hardware . The finding was that SVM
>>> could achieve good performance without the low-level
>>>
>>> complexity of managing explicit messages, though
>>> the best results often came from a hybrid approach (sic!)
>>> Here is an interesting baseline, from Java,
>>> a class ElevenSingle that only does:
>>>
>>>      public static void run() {
>>>          for (int A = 1; A < 192; A++) {
>>>              int Y = (771-A)/3;
>>>              for (int B = A; B < Y; B++) {
>>>                  int Z = (771-A-B)/2;
>>>                  for (int C = B; C < Z; C++) {
>>>                      int D = 711-A-B-C;
>>>                      if (A*B*C == 711000000/D &&
>>>                            711000000 % D == 0)
>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>                  }
>>>              }
>>>          }
>>>      }
>>>
>>> And then compare it to ElevenMulti, doing some
>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>
>>> ElevenSingle
>>> A=120, B=125, C=150, D=316
>>> 6.628 ms
>>>
>>> ElevenMulti
>>> A=120, B=125, C=150, D=316
>>> 1.941 ms
>>>
>>> Not great, not terrible!
>>>
>>> Bye
>>
>> Oh, that's just "tricks of p-adic arithmetic".
>>
>> Like other sock-puppet howler trolls, when confronted
>> with its base incredulity, it will descend to its
>> lower levers of the pathos variety.
>>
>> You might be happier learning about Julia trees and
>> raster ops, instead of shilling yet another Ramanujan
>> series without saying how it's made.
>>
>> Bulgarians, that's some real Boris and Natasha crap,
>> forget Hungarians and Bulgarians.
>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647004 — Andrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 09:51 +0200
SubjectAndrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants)
Message-ID<11472mm$gcev$1@solani.org>
In reply to#647003
Hi,

Andrej Karpathy was bascially the original gangster
of doing not only AI inferencing but also AI
learning on a Budget Laptop. The nanoGPT project

states the following:

"I only have a macbook (or other cheap
computer). No worries, we can still train a
GPT but we want to dial things down a notch.
I recommend getting the bleeding edge PyTorch
nightly (select it here when installing) as
it is currently quite likely to make your
code more efficient."
https://github.com/karpathy/nanogpt

But meanwhile he has moved to a higher price
segment. Not sure whether he will climbe
down to a lower price segment again:

For example, you can train your own GPT-2
capability LLM (which cost ~$43,000 to train in
2019) for only $48 (~2 hours of 8XH100 GPU node)
and then talk to it over a simple CLI. On a spot
instance, the total cost can be closer to ~$15.
https://github.com/karpathy/nanochat

Bt he taps into the model to rent GPU which
is available with prices in the range of 1-2 $
per hour. Even in Switzerland one can do that,

for example using the provider Exoscale. Since
he rents a cluster of 8 cards of type H100, this
explains his training price still in the 2 digit range.

Bye

P.S.: I could also do my experiment here with
rented GPU cards, and then draw a comparison
from budget laptop to the rented GPU time market:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

But testing rented GPU is not high priority.

Mild Shock schrieb:
> Hi,
> 
> Some counter PyTorch Python trends are
> for example OpenAIs Triton. And the variant
> miniTriton CUDA vibe produced by Kimi K3 (sic!):
> 
> "We further tested whether Kimi K3 could build
> a GPU programming system from scratch. Kimi K3
> developed MiniTriton, a compact Triton-like
> compiler with its own tile-level IR layer over
> MLIR, optimization passes, and a PTX code-
> generation pipeline.
> 
> Across supported roofline benchmarks, MiniTriton
> delivers performance on par with or better than
> Triton and torch.compile — beating Triton on
> certain workloads. Beyond microbenchmarks,
> MiniTriton sustains end-to-end nanoGPT training
> with stable convergence, the loss curve
> 
> closely tracking the reference with only minor
> divergence — validating the full pipeline on a
> realistic workload. These results demonstrate
> that Kimi K3 can build a coherent end-to-end
> compiler — from DSL frontend and IR passes to
> PTX codegen and runtime — rather than isolated
> 
> kernels; its from-scratch Tensor Core path
> already rivals Triton’s extensively optimized stack."
> 
> GPU Compiler Development
> https://www.kimi.com/blog/kimi-k3
> 
> Although many GPU corporate stuff is anonymized,
> and some AI papers have lists of 30 authors. Here
> nanoGPT is mentioned which is tied to the name
> 
> Andrej Karpathy. See also here:
> 
> Update Nov 2025 nanoGPT has a new and
> improved cousin called nanochat.
> https://github.com/karpathy/nanogpt
> 
> But as can be seen, he moved on to another project.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Whats this "forget" trope of glue sniffing
>> Rossy Boy with his herpes blisters?
>>
>>  > Bulgarians, that's some real Boris and Natasha crap,
>>  > forget Hungarians and Bulgarians.
>>
>> Why should I forget Bulgarians,
>> they are never on my mind. Do you
>> see me doing ggml stuff?
>>
>> I only hypothesized that it is
>> over for Python as the machine
>> learning language or AI inferencing
>>
>> locally on AI laptops language, and
>> made the ggml case, so I already forgot
>> about them. Which might give you a glimps,
>>
>> why WebGPU was used for this here:
>>
>> 11.4 Giga Lips with a Budget Laptop
>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> Is an interesting choice. Even
>> github has some Languages statistics,
>> giving an account what I used:
>>
>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>
>> Have Fun!
>>
>> Bye
>>
>> P.S.: The example below is not p-adics,
>> you complete imbecil moron. Its just:
>>
>> 7-11 cubic Solution by Pritchard & Gries
>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>
>> Ross Finlayson schrieb:
>>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>> Hi,
>>>>
>>>> You see it all boils down to find your inner peace
>>>> by an immaculate inception of some queue datatype.
>>>>
>>>> KOAN/Fortran-S was an early 1990s research programming
>>>> system for distributed-memory multiprocessors . Developed
>>>> at ENS Lyon in the early 1990s . Often listed alongside
>>>> other historical parallel programming efforts.
>>>>
>>>> The Message Passing: The research explicitly
>>>> compared the SVM approach against message passing
>>>> on the same hardware . The finding was that SVM
>>>> could achieve good performance without the low-level
>>>>
>>>> complexity of managing explicit messages, though
>>>> the best results often came from a hybrid approach (sic!)
>>>> Here is an interesting baseline, from Java,
>>>> a class ElevenSingle that only does:
>>>>
>>>>      public static void run() {
>>>>          for (int A = 1; A < 192; A++) {
>>>>              int Y = (771-A)/3;
>>>>              for (int B = A; B < Y; B++) {
>>>>                  int Z = (771-A-B)/2;
>>>>                  for (int C = B; C < Z; C++) {
>>>>                      int D = 711-A-B-C;
>>>>                      if (A*B*C == 711000000/D &&
>>>>                            711000000 % D == 0)
>>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>                  }
>>>>              }
>>>>          }
>>>>      }
>>>>
>>>> And then compare it to ElevenMulti, doing some
>>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>
>>>> ElevenSingle
>>>> A=120, B=125, C=150, D=316
>>>> 6.628 ms
>>>>
>>>> ElevenMulti
>>>> A=120, B=125, C=150, D=316
>>>> 1.941 ms
>>>>
>>>> Not great, not terrible!
>>>>
>>>> Bye
>>>
>>> Oh, that's just "tricks of p-adic arithmetic".
>>>
>>> Like other sock-puppet howler trolls, when confronted
>>> with its base incredulity, it will descend to its
>>> lower levers of the pathos variety.
>>>
>>> You might be happier learning about Julia trees and
>>> raster ops, instead of shilling yet another Ramanujan
>>> series without saying how it's made.
>>>
>>> Bulgarians, that's some real Boris and Natasha crap,
>>> forget Hungarians and Bulgarians.
>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647006 — Re: Andrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants)

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-27 01:41 -0700
SubjectRe: Andrej Karpathy original gangster of Budget Laptop (Was: miniTriton CUDA is an alternative to torch variants)
Message-ID<0jWdnUWmIeh5hPr3nZ2dnZfqnPSdnZ2d@giganews.com>
In reply to#647004
On 07/27/2026 12:51 AM, Mild Shock wrote:
> Hi,
>
> Andrej Karpathy was bascially the original gangster
> of doing not only AI inferencing but also AI
> learning on a Budget Laptop. The nanoGPT project
>
> states the following:
>
> "I only have a macbook (or other cheap
> computer). No worries, we can still train a
> GPT but we want to dial things down a notch.
> I recommend getting the bleeding edge PyTorch
> nightly (select it here when installing) as
> it is currently quite likely to make your
> code more efficient."
> https://github.com/karpathy/nanogpt
>
> But meanwhile he has moved to a higher price
> segment. Not sure whether he will climbe
> down to a lower price segment again:
>
> For example, you can train your own GPT-2
> capability LLM (which cost ~$43,000 to train in
> 2019) for only $48 (~2 hours of 8XH100 GPU node)
> and then talk to it over a simple CLI. On a spot
> instance, the total cost can be closer to ~$15.
> https://github.com/karpathy/nanochat
>
> Bt he taps into the model to rent GPU which
> is available with prices in the range of 1-2 $
> per hour. Even in Switzerland one can do that,
>
> for example using the provider Exoscale. Since
> he rents a cluster of 8 cards of type H100, this
> explains his training price still in the 2 digit range.
>
> Bye
>
> P.S.: I could also do my experiment here with
> rented GPU cards, and then draw a comparison
> from budget laptop to the rented GPU time market:
>
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
>
> But testing rented GPU is not high priority.
>
> Mild Shock schrieb:
>> Hi,
>>
>> Some counter PyTorch Python trends are
>> for example OpenAIs Triton. And the variant
>> miniTriton CUDA vibe produced by Kimi K3 (sic!):
>>
>> "We further tested whether Kimi K3 could build
>> a GPU programming system from scratch. Kimi K3
>> developed MiniTriton, a compact Triton-like
>> compiler with its own tile-level IR layer over
>> MLIR, optimization passes, and a PTX code-
>> generation pipeline.
>>
>> Across supported roofline benchmarks, MiniTriton
>> delivers performance on par with or better than
>> Triton and torch.compile — beating Triton on
>> certain workloads. Beyond microbenchmarks,
>> MiniTriton sustains end-to-end nanoGPT training
>> with stable convergence, the loss curve
>>
>> closely tracking the reference with only minor
>> divergence — validating the full pipeline on a
>> realistic workload. These results demonstrate
>> that Kimi K3 can build a coherent end-to-end
>> compiler — from DSL frontend and IR passes to
>> PTX codegen and runtime — rather than isolated
>>
>> kernels; its from-scratch Tensor Core path
>> already rivals Triton’s extensively optimized stack."
>>
>> GPU Compiler Development
>> https://www.kimi.com/blog/kimi-k3
>>
>> Although many GPU corporate stuff is anonymized,
>> and some AI papers have lists of 30 authors. Here
>> nanoGPT is mentioned which is tied to the name
>>
>> Andrej Karpathy. See also here:
>>
>> Update Nov 2025 nanoGPT has a new and
>> improved cousin called nanochat.
>> https://github.com/karpathy/nanogpt
>>
>> But as can be seen, he moved on to another project.
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Whats this "forget" trope of glue sniffing
>>> Rossy Boy with his herpes blisters?
>>>
>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>  > forget Hungarians and Bulgarians.
>>>
>>> Why should I forget Bulgarians,
>>> they are never on my mind. Do you
>>> see me doing ggml stuff?
>>>
>>> I only hypothesized that it is
>>> over for Python as the machine
>>> learning language or AI inferencing
>>>
>>> locally on AI laptops language, and
>>> made the ggml case, so I already forgot
>>> about them. Which might give you a glimps,
>>>
>>> why WebGPU was used for this here:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> Is an interesting choice. Even
>>> github has some Languages statistics,
>>> giving an account what I used:
>>>
>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>
>>> Have Fun!
>>>
>>> Bye
>>>
>>> P.S.: The example below is not p-adics,
>>> you complete imbecil moron. Its just:
>>>
>>> 7-11 cubic Solution by Pritchard & Gries
>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>
>>> Ross Finlayson schrieb:
>>>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>> Hi,
>>>>>
>>>>> You see it all boils down to find your inner peace
>>>>> by an immaculate inception of some queue datatype.
>>>>>
>>>>> KOAN/Fortran-S was an early 1990s research programming
>>>>> system for distributed-memory multiprocessors . Developed
>>>>> at ENS Lyon in the early 1990s . Often listed alongside
>>>>> other historical parallel programming efforts.
>>>>>
>>>>> The Message Passing: The research explicitly
>>>>> compared the SVM approach against message passing
>>>>> on the same hardware . The finding was that SVM
>>>>> could achieve good performance without the low-level
>>>>>
>>>>> complexity of managing explicit messages, though
>>>>> the best results often came from a hybrid approach (sic!)
>>>>> Here is an interesting baseline, from Java,
>>>>> a class ElevenSingle that only does:
>>>>>
>>>>>      public static void run() {
>>>>>          for (int A = 1; A < 192; A++) {
>>>>>              int Y = (771-A)/3;
>>>>>              for (int B = A; B < Y; B++) {
>>>>>                  int Z = (771-A-B)/2;
>>>>>                  for (int C = B; C < Z; C++) {
>>>>>                      int D = 711-A-B-C;
>>>>>                      if (A*B*C == 711000000/D &&
>>>>>                            711000000 % D == 0)
>>>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>>                  }
>>>>>              }
>>>>>          }
>>>>>      }
>>>>>
>>>>> And then compare it to ElevenMulti, doing some
>>>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>>
>>>>> ElevenSingle
>>>>> A=120, B=125, C=150, D=316
>>>>> 6.628 ms
>>>>>
>>>>> ElevenMulti
>>>>> A=120, B=125, C=150, D=316
>>>>> 1.941 ms
>>>>>
>>>>> Not great, not terrible!
>>>>>
>>>>> Bye
>>>>
>>>> Oh, that's just "tricks of p-adic arithmetic".
>>>>
>>>> Like other sock-puppet howler trolls, when confronted
>>>> with its base incredulity, it will descend to its
>>>> lower levers of the pathos variety.
>>>>
>>>> You might be happier learning about Julia trees and
>>>> raster ops, instead of shilling yet another Ramanujan
>>>> series without saying how it's made.
>>>>
>>>> Bulgarians, that's some real Boris and Natasha crap,
>>>> forget Hungarians and Bulgarians.
>>>>
>>>>
>>>
>>
>

Hopfield + Kohonen and some arithmetic coding,
dirt simple since the '90's, generative programming
and online psychiatrists are around since the '60's,
chat-bots are simple in the scheme of things,
and very widely varied in their implementation,
feedback-directed optimization,
pile arithmetic coding and call it vectors on big data
and say that's a requirement, when really it's just
_bloat_ and a sales-case for _bloat_ and it's bloated
and it _bloats_ thee. Bloater.

[toc] | [prev] | [next] | [standalone]


#647005 — Re: Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta)

FromRoss Finlayson <ross.a.finlayson@gmail.com>
Date2026-07-27 01:38 -0700
SubjectRe: Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S] (Was: Not SIMD, a MIMD design for NVIDIA Volta)
Message-ID<QyKdnQSWua1lhfr3nZ2dnZfqn_SdnZ2d@giganews.com>
In reply to#647002
On 07/27/2026 12:14 AM, Mild Shock wrote:
> Hi,
>
> Whats this "forget" trope of glue sniffing
> Rossy Boy with his herpes blisters?
>
>  > Bulgarians, that's some real Boris and Natasha crap,
>  > forget Hungarians and Bulgarians.
>
> Why should I forget Bulgarians,
> they are never on my mind. Do you
> see me doing ggml stuff?
>
> I only hypothesized that it is
> over for Python as the machine
> learning language or AI inferencing
>
> locally on AI laptops language, and
> made the ggml case, so I already forgot
> about them. Which might give you a glimps,
>
> why WebGPU was used for this here:
>
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
>
> Is an interesting choice. Even
> github has some Languages statistics,
> giving an account what I used:
>
> HTML 67.5% JavaScript 23.1% CSS 9.4%
>
> Have Fun!
>
> Bye
>
> P.S.: The example below is not p-adics,
> you complete imbecil moron. Its just:
>
> 7-11 cubic Solution by Pritchard & Gries
> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>
> Ross Finlayson schrieb:
>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> You see it all boils down to find your inner peace
>>> by an immaculate inception of some queue datatype.
>>>
>>> KOAN/Fortran-S was an early 1990s research programming
>>> system for distributed-memory multiprocessors . Developed
>>> at ENS Lyon in the early 1990s . Often listed alongside
>>> other historical parallel programming efforts.
>>>
>>> The Message Passing: The research explicitly
>>> compared the SVM approach against message passing
>>> on the same hardware . The finding was that SVM
>>> could achieve good performance without the low-level
>>>
>>> complexity of managing explicit messages, though
>>> the best results often came from a hybrid approach (sic!)
>>> Here is an interesting baseline, from Java,
>>> a class ElevenSingle that only does:
>>>
>>>      public static void run() {
>>>          for (int A = 1; A < 192; A++) {
>>>              int Y = (771-A)/3;
>>>              for (int B = A; B < Y; B++) {
>>>                  int Z = (771-A-B)/2;
>>>                  for (int C = B; C < Z; C++) {
>>>                      int D = 711-A-B-C;
>>>                      if (A*B*C == 711000000/D &&
>>>                            711000000 % D == 0)
>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>                  }
>>>              }
>>>          }
>>>      }
>>>
>>> And then compare it to ElevenMulti, doing some
>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>
>>> ElevenSingle
>>> A=120, B=125, C=150, D=316
>>> 6.628 ms
>>>
>>> ElevenMulti
>>> A=120, B=125, C=150, D=316
>>> 1.941 ms
>>>
>>> Not great, not terrible!
>>>
>>> Bye
>>
>> Oh, that's just "tricks of p-adic arithmetic".
>>
>> Like other sock-puppet howler trolls, when confronted
>> with its base incredulity, it will descend to its
>> lower levers of the pathos variety.
>>
>> You might be happier learning about Julia trees and
>> raster ops, instead of shilling yet another Ramanujan
>> series without saying how it's made.
>>
>> Bulgarians, that's some real Boris and Natasha crap,
>> forget Hungarians and Bulgarians.
>>
>>
>

Ah, then the arithmetic is a trick,
and the sock-puppet howler troll drools nonsense.


There are lots of tricks of arithmetic.

Systolic flow machines their ideas are around a long time.

Shut Up

[toc] | [prev] | [next] | [standalone]


#647007 — The evolution of hardware and GPT-2 training (Was: Why forget Bulgarians, never on my mind)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 10:56 +0200
SubjectThe evolution of hardware and GPT-2 training (Was: Why forget Bulgarians, never on my mind)
Message-ID<11476g4$ftjl$1@solani.org>
In reply to#647002
Hi,

While Huggingfaces hired GG in 2026,
AK was hired by Anthropic in 2026:

Andrej Karpathy (born 23 October 1986[3])
is a Slovak-Canadian AI researcher, who
co-founded and formerly worked at OpenAI
In 2026 he joined Anthropic as part of
the pretraining team.
https://en.wikipedia.org/wiki/Andrej_Karpathy

But his nanochat archivement has an
interesting time line:

168 hours , Original OpenAI GPT-2 checkpoint, 2019
3 hours , d24 baseline, slightly overtrained, Jan 29 2026
1 1/2 hour, autoresearch round 2, Mar 14 2026
The best ChatGPT that $100 can buy.
https://github.com/karpathy/nanochat

But what hardware was the enabler. What is the
NVIDIA H100 GPU even. Well the thingy is surely not
a Budget Laptop, performance pretty much

dependence on data elememt size, the H100 NVL
version (*), and when using tensor operations,
and not only scalar operations:

8-bit towards 3000 tera flops
16-bit towards 1500 tera flops
32-bit towards 900 tera flops

Cool! I guess this experiment would tap into 60
tera flops, since it only uses scalar operations so far:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

You could perform it by migration the web application
using WebGPU into a node.js standalone application
using the dawn library for GPU access.

Bye

(*) 
https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet

Mild Shock schrieb:
> Hi,
> 
> Whats this "forget" trope of glue sniffing
> Rossy Boy with his herpes blisters?
> 
>  > Bulgarians, that's some real Boris and Natasha crap,
>  > forget Hungarians and Bulgarians.
> 
> Why should I forget Bulgarians,
> they are never on my mind. Do you
> see me doing ggml stuff?
> 
> I only hypothesized that it is
> over for Python as the machine
> learning language or AI inferencing
> 
> locally on AI laptops language, and
> made the ggml case, so I already forgot
> about them. Which might give you a glimps,
> 
> why WebGPU was used for this here:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Is an interesting choice. Even
> github has some Languages statistics,
> giving an account what I used:
> 
> HTML 67.5% JavaScript 23.1% CSS 9.4%
> 
> Have Fun!
> 
> Bye
> 
> P.S.: The example below is not p-adics,
> you complete imbecil moron. Its just:
> 
> 7-11 cubic Solution by Pritchard & Gries
> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
> 
> Ross Finlayson schrieb:
>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> You see it all boils down to find your inner peace
>>> by an immaculate inception of some queue datatype.
>>>
>>> KOAN/Fortran-S was an early 1990s research programming
>>> system for distributed-memory multiprocessors . Developed
>>> at ENS Lyon in the early 1990s . Often listed alongside
>>> other historical parallel programming efforts.
>>>
>>> The Message Passing: The research explicitly
>>> compared the SVM approach against message passing
>>> on the same hardware . The finding was that SVM
>>> could achieve good performance without the low-level
>>>
>>> complexity of managing explicit messages, though
>>> the best results often came from a hybrid approach (sic!)
>>> Here is an interesting baseline, from Java,
>>> a class ElevenSingle that only does:
>>>
>>>      public static void run() {
>>>          for (int A = 1; A < 192; A++) {
>>>              int Y = (771-A)/3;
>>>              for (int B = A; B < Y; B++) {
>>>                  int Z = (771-A-B)/2;
>>>                  for (int C = B; C < Z; C++) {
>>>                      int D = 711-A-B-C;
>>>                      if (A*B*C == 711000000/D &&
>>>                            711000000 % D == 0)
>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>                  }
>>>              }
>>>          }
>>>      }
>>>
>>> And then compare it to ElevenMulti, doing some
>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>
>>> ElevenSingle
>>> A=120, B=125, C=150, D=316
>>> 6.628 ms
>>>
>>> ElevenMulti
>>> A=120, B=125, C=150, D=316
>>> 1.941 ms
>>>
>>> Not great, not terrible!
>>>
>>> Bye
>>
>> Oh, that's just "tricks of p-adic arithmetic".
>>
>> Like other sock-puppet howler trolls, when confronted
>> with its base incredulity, it will descend to its
>> lower levers of the pathos variety.
>>
>> You might be happier learning about Julia trees and
>> raster ops, instead of shilling yet another Ramanujan
>> series without saying how it's made.
>>
>> Bulgarians, that's some real Boris and Natasha crap,
>> forget Hungarians and Bulgarians.
>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647008 — How speed up π-WAM with vector operations (Was: The evolution of hardware and GPT-2 training)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 11:08 +0200
SubjectHow speed up π-WAM with vector operations (Was: The evolution of hardware and GPT-2 training)
Message-ID<114776f$fu5f$1@solani.org>
In reply to#647007
Hi,

One could critisize that my π-WAM doesn't
utilize GPU to the fullest, since its GPU
backend prototype only uses scalar operations

and no vector or matrix operations. And
modern GPUs thrive on vector and matrix
operations. Especially matrix operations giving

a boost of a factor 15x or so. There are
many papers already showing how Prolog can be
mapped to matrix operations. Only this research

is completely ignored by Prolog systems such as
SICStus, Ciao, SWI, ECLiPSe etc.. But lets
illustrate what vector operations could do

for π-WAM, take this compilation of the Prolog
goal between(0,1023,X), Y is X*2+3:

int X;
int Y;
for (X=0; X < 1024; X++) {
     Y=X*2+3;
     [...]
}

With vector operations, and vectors of size
32 one could do:

int X1;
int[] X = new int[32];
int X3;
int[] Y = new int[32];
for (X1 = 0; X1 < 1024 / 32; X1++) {
     for (int X2 = 0; X2 < 32; X2++)
        X[X2] = X1*32+X2;
     vec_mul_add(X, 2, 3, Y);
     [..]
}

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
> While Huggingfaces hired GG in 2026,
> AK was hired by Anthropic in 2026:
> 
> Andrej Karpathy (born 23 October 1986[3])
> is a Slovak-Canadian AI researcher, who
> co-founded and formerly worked at OpenAI
> In 2026 he joined Anthropic as part of
> the pretraining team.
> https://en.wikipedia.org/wiki/Andrej_Karpathy
> 
> But his nanochat archivement has an
> interesting time line:
> 
> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
> 1 1/2 hour, autoresearch round 2, Mar 14 2026
> The best ChatGPT that $100 can buy.
> https://github.com/karpathy/nanochat
> 
> But what hardware was the enabler. What is the
> NVIDIA H100 GPU even. Well the thingy is surely not
> a Budget Laptop, performance pretty much
> 
> dependence on data elememt size, the H100 NVL
> version (*), and when using tensor operations,
> and not only scalar operations:
> 
> 8-bit towards 3000 tera flops
> 16-bit towards 1500 tera flops
> 32-bit towards 900 tera flops
> 
> Cool! I guess this experiment would tap into 60
> tera flops, since it only uses scalar operations so far:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> You could perform it by migration the web application
> using WebGPU into a node.js standalone application
> using the dawn library for GPU access.
> 
> Bye
> 
> (*) 
> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
> 
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Whats this "forget" trope of glue sniffing
>> Rossy Boy with his herpes blisters?
>>
>>  > Bulgarians, that's some real Boris and Natasha crap,
>>  > forget Hungarians and Bulgarians.
>>
>> Why should I forget Bulgarians,
>> they are never on my mind. Do you
>> see me doing ggml stuff?
>>
>> I only hypothesized that it is
>> over for Python as the machine
>> learning language or AI inferencing
>>
>> locally on AI laptops language, and
>> made the ggml case, so I already forgot
>> about them. Which might give you a glimps,
>>
>> why WebGPU was used for this here:
>>
>> 11.4 Giga Lips with a Budget Laptop
>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> Is an interesting choice. Even
>> github has some Languages statistics,
>> giving an account what I used:
>>
>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>
>> Have Fun!
>>
>> Bye
>>
>> P.S.: The example below is not p-adics,
>> you complete imbecil moron. Its just:
>>
>> 7-11 cubic Solution by Pritchard & Gries
>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>
>> Ross Finlayson schrieb:
>>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>> Hi,
>>>>
>>>> You see it all boils down to find your inner peace
>>>> by an immaculate inception of some queue datatype.
>>>>
>>>> KOAN/Fortran-S was an early 1990s research programming
>>>> system for distributed-memory multiprocessors . Developed
>>>> at ENS Lyon in the early 1990s . Often listed alongside
>>>> other historical parallel programming efforts.
>>>>
>>>> The Message Passing: The research explicitly
>>>> compared the SVM approach against message passing
>>>> on the same hardware . The finding was that SVM
>>>> could achieve good performance without the low-level
>>>>
>>>> complexity of managing explicit messages, though
>>>> the best results often came from a hybrid approach (sic!)
>>>> Here is an interesting baseline, from Java,
>>>> a class ElevenSingle that only does:
>>>>
>>>>      public static void run() {
>>>>          for (int A = 1; A < 192; A++) {
>>>>              int Y = (771-A)/3;
>>>>              for (int B = A; B < Y; B++) {
>>>>                  int Z = (771-A-B)/2;
>>>>                  for (int C = B; C < Z; C++) {
>>>>                      int D = 711-A-B-C;
>>>>                      if (A*B*C == 711000000/D &&
>>>>                            711000000 % D == 0)
>>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>                  }
>>>>              }
>>>>          }
>>>>      }
>>>>
>>>> And then compare it to ElevenMulti, doing some
>>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>
>>>> ElevenSingle
>>>> A=120, B=125, C=150, D=316
>>>> 6.628 ms
>>>>
>>>> ElevenMulti
>>>> A=120, B=125, C=150, D=316
>>>> 1.941 ms
>>>>
>>>> Not great, not terrible!
>>>>
>>>> Bye
>>>
>>> Oh, that's just "tricks of p-adic arithmetic".
>>>
>>> Like other sock-puppet howler trolls, when confronted
>>> with its base incredulity, it will descend to its
>>> lower levers of the pathos variety.
>>>
>>> You might be happier learning about Julia trees and
>>> raster ops, instead of shilling yet another Ramanujan
>>> series without saying how it's made.
>>>
>>> Bulgarians, that's some real Boris and Natasha crap,
>>> forget Hungarians and Bulgarians.
>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647009 — AI accelerator extend from GPU to CPU [Zero Copying] (Was: How speed up π-WAM with vector operations)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 11:21 +0200
SubjectAI accelerator extend from GPU to CPU [Zero Copying] (Was: How speed up π-WAM with vector operations)
Message-ID<11477vd$fup5$1@solani.org>
In reply to#647008
Hi,

The nice thing about AI accelerators, pioneered
maybe by Apple Silicon and their unified memory.
The AMD APU model can be extended so that

vector and matrix operations become uniformly
available for GPU and CPU. With unified memory
already a vector operation such as:

vec_mul_add(X, 2, 3, Y)

Only needs the X and Y address. But I havent
got my head around yet how this is all organized.
Maybe a GPU has still its own GEMM cores,

but you find Apple Silicon C++/C source code,
that taps into vector and matrix operations
by Zero Copying. The Copying is left to the DMA

of the vector or matrix operation. And moderated
by the various caches. Leading to the slogan, that
multiple floating point operations become zero cost:

Some teaching can be found here
https://www.hpc-ch.org/category/topics/course-workshop/

Bye

Mild Shock schrieb:
> Hi,
> 
> One could critisize that my π-WAM doesn't
> utilize GPU to the fullest, since its GPU
> backend prototype only uses scalar operations
> 
> and no vector or matrix operations. And
> modern GPUs thrive on vector and matrix
> operations. Especially matrix operations giving
> 
> a boost of a factor 15x or so. There are
> many papers already showing how Prolog can be
> mapped to matrix operations. Only this research
> 
> is completely ignored by Prolog systems such as
> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
> illustrate what vector operations could do
> 
> for π-WAM, take this compilation of the Prolog
> goal between(0,1023,X), Y is X*2+3:
> 
> int X;
> int Y;
> for (X=0; X < 1024; X++) {
>      Y=X*2+3;
>      [...]
> }
> 
> With vector operations, and vectors of size
> 32 one could do:
> 
> int X1;
> int[] X = new int[32];
> int X3;
> int[] Y = new int[32];
> for (X1 = 0; X1 < 1024 / 32; X1++) {
>      for (int X2 = 0; X2 < 32; X2++)
>         X[X2] = X1*32+X2;
>      vec_mul_add(X, 2, 3, Y);
>      [..]
> }
> 
> Have Fun!
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> While Huggingfaces hired GG in 2026,
>> AK was hired by Anthropic in 2026:
>>
>> Andrej Karpathy (born 23 October 1986[3])
>> is a Slovak-Canadian AI researcher, who
>> co-founded and formerly worked at OpenAI
>> In 2026 he joined Anthropic as part of
>> the pretraining team.
>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>
>> But his nanochat archivement has an
>> interesting time line:
>>
>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>> The best ChatGPT that $100 can buy.
>> https://github.com/karpathy/nanochat
>>
>> But what hardware was the enabler. What is the
>> NVIDIA H100 GPU even. Well the thingy is surely not
>> a Budget Laptop, performance pretty much
>>
>> dependence on data elememt size, the H100 NVL
>> version (*), and when using tensor operations,
>> and not only scalar operations:
>>
>> 8-bit towards 3000 tera flops
>> 16-bit towards 1500 tera flops
>> 32-bit towards 900 tera flops
>>
>> Cool! I guess this experiment would tap into 60
>> tera flops, since it only uses scalar operations so far:
>>
>> 11.4 Giga Lips with a Budget Laptop
>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> You could perform it by migration the web application
>> using WebGPU into a node.js standalone application
>> using the dawn library for GPU access.
>>
>> Bye
>>
>> (*) 
>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
>>
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Whats this "forget" trope of glue sniffing
>>> Rossy Boy with his herpes blisters?
>>>
>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>  > forget Hungarians and Bulgarians.
>>>
>>> Why should I forget Bulgarians,
>>> they are never on my mind. Do you
>>> see me doing ggml stuff?
>>>
>>> I only hypothesized that it is
>>> over for Python as the machine
>>> learning language or AI inferencing
>>>
>>> locally on AI laptops language, and
>>> made the ggml case, so I already forgot
>>> about them. Which might give you a glimps,
>>>
>>> why WebGPU was used for this here:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> Is an interesting choice. Even
>>> github has some Languages statistics,
>>> giving an account what I used:
>>>
>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>
>>> Have Fun!
>>>
>>> Bye
>>>
>>> P.S.: The example below is not p-adics,
>>> you complete imbecil moron. Its just:
>>>
>>> 7-11 cubic Solution by Pritchard & Gries
>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>
>>> Ross Finlayson schrieb:
>>>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>> Hi,
>>>>>
>>>>> You see it all boils down to find your inner peace
>>>>> by an immaculate inception of some queue datatype.
>>>>>
>>>>> KOAN/Fortran-S was an early 1990s research programming
>>>>> system for distributed-memory multiprocessors . Developed
>>>>> at ENS Lyon in the early 1990s . Often listed alongside
>>>>> other historical parallel programming efforts.
>>>>>
>>>>> The Message Passing: The research explicitly
>>>>> compared the SVM approach against message passing
>>>>> on the same hardware . The finding was that SVM
>>>>> could achieve good performance without the low-level
>>>>>
>>>>> complexity of managing explicit messages, though
>>>>> the best results often came from a hybrid approach (sic!)
>>>>> Here is an interesting baseline, from Java,
>>>>> a class ElevenSingle that only does:
>>>>>
>>>>>      public static void run() {
>>>>>          for (int A = 1; A < 192; A++) {
>>>>>              int Y = (771-A)/3;
>>>>>              for (int B = A; B < Y; B++) {
>>>>>                  int Z = (771-A-B)/2;
>>>>>                  for (int C = B; C < Z; C++) {
>>>>>                      int D = 711-A-B-C;
>>>>>                      if (A*B*C == 711000000/D &&
>>>>>                            711000000 % D == 0)
>>>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>>                  }
>>>>>              }
>>>>>          }
>>>>>      }
>>>>>
>>>>> And then compare it to ElevenMulti, doing some
>>>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>>
>>>>> ElevenSingle
>>>>> A=120, B=125, C=150, D=316
>>>>> 6.628 ms
>>>>>
>>>>> ElevenMulti
>>>>> A=120, B=125, C=150, D=316
>>>>> 1.941 ms
>>>>>
>>>>> Not great, not terrible!
>>>>>
>>>>> Bye
>>>>
>>>> Oh, that's just "tricks of p-adic arithmetic".
>>>>
>>>> Like other sock-puppet howler trolls, when confronted
>>>> with its base incredulity, it will descend to its
>>>> lower levers of the pathos variety.
>>>>
>>>> You might be happier learning about Julia trees and
>>>> raster ops, instead of shilling yet another Ramanujan
>>>> series without saying how it's made.
>>>>
>>>> Bulgarians, that's some real Boris and Natasha crap,
>>>> forget Hungarians and Bulgarians.
>>>>
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647013 — The invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 13:20 +0200
SubjectThe invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<1147etb$glbs$1@solani.org>
In reply to#647009
Hi,

But the example gives also way to vector
and matrix registers. The int[] X and
int[] Y could be also held in vector

registers. Compilers can also optimize
away int[] Y, and use a inline modification,
in case X isn't used later, then playing

the role of Y:

vec_mul_add(X, 2, 3, X)

Vector and matrix registers in modern GPUs
emerged from distinct architectural milestones:
vector-like register files developed with
early programmable 3D vertex/pixel pipelines

in the late 1990s to early 2000s. While
dedicated multi-dimensional matrix registers
(Tensor Cores/Matrix Cores) were invented by
NVIDIA in 2017, starting with the Tesla

V100 (Volta microarchitecture):

 From Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell

You see the scheduling of tensure core occupation
scheduling in the above article, including memory
and register flow, following the section:

MMA Instruction Overview

It went through a couple of generations, leading
to Tensor Memory (TMEM) and collective operations,
basically realizing the PIM idea:

Processing-in-Memory Tutorials
https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
> The nice thing about AI accelerators, pioneered
> maybe by Apple Silicon and their unified memory.
> The AMD APU model can be extended so that
> 
> vector and matrix operations become uniformly
> available for GPU and CPU. With unified memory
> already a vector operation such as:
> 
> vec_mul_add(X, 2, 3, Y)
> 
> Only needs the X and Y address. But I havent
> got my head around yet how this is all organized.
> Maybe a GPU has still its own GEMM cores,
> 
> but you find Apple Silicon C++/C source code,
> that taps into vector and matrix operations
> by Zero Copying. The Copying is left to the DMA
> 
> of the vector or matrix operation. And moderated
> by the various caches. Leading to the slogan, that
> multiple floating point operations become zero cost:
> 
> Some teaching can be found here
> https://www.hpc-ch.org/category/topics/course-workshop/
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> One could critisize that my π-WAM doesn't
>> utilize GPU to the fullest, since its GPU
>> backend prototype only uses scalar operations
>>
>> and no vector or matrix operations. And
>> modern GPUs thrive on vector and matrix
>> operations. Especially matrix operations giving
>>
>> a boost of a factor 15x or so. There are
>> many papers already showing how Prolog can be
>> mapped to matrix operations. Only this research
>>
>> is completely ignored by Prolog systems such as
>> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
>> illustrate what vector operations could do
>>
>> for π-WAM, take this compilation of the Prolog
>> goal between(0,1023,X), Y is X*2+3:
>>
>> int X;
>> int Y;
>> for (X=0; X < 1024; X++) {
>>      Y=X*2+3;
>>      [...]
>> }
>>
>> With vector operations, and vectors of size
>> 32 one could do:
>>
>> int X1;
>> int[] X = new int[32];
>> int X3;
>> int[] Y = new int[32];
>> for (X1 = 0; X1 < 1024 / 32; X1++) {
>>      for (int X2 = 0; X2 < 32; X2++)
>>         X[X2] = X1*32+X2;
>>      vec_mul_add(X, 2, 3, Y);
>>      [..]
>> }
>>
>> Have Fun!
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> While Huggingfaces hired GG in 2026,
>>> AK was hired by Anthropic in 2026:
>>>
>>> Andrej Karpathy (born 23 October 1986[3])
>>> is a Slovak-Canadian AI researcher, who
>>> co-founded and formerly worked at OpenAI
>>> In 2026 he joined Anthropic as part of
>>> the pretraining team.
>>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>>
>>> But his nanochat archivement has an
>>> interesting time line:
>>>
>>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>>> The best ChatGPT that $100 can buy.
>>> https://github.com/karpathy/nanochat
>>>
>>> But what hardware was the enabler. What is the
>>> NVIDIA H100 GPU even. Well the thingy is surely not
>>> a Budget Laptop, performance pretty much
>>>
>>> dependence on data elememt size, the H100 NVL
>>> version (*), and when using tensor operations,
>>> and not only scalar operations:
>>>
>>> 8-bit towards 3000 tera flops
>>> 16-bit towards 1500 tera flops
>>> 32-bit towards 900 tera flops
>>>
>>> Cool! I guess this experiment would tap into 60
>>> tera flops, since it only uses scalar operations so far:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> You could perform it by migration the web application
>>> using WebGPU into a node.js standalone application
>>> using the dawn library for GPU access.
>>>
>>> Bye
>>>
>>> (*) 
>>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
>>>
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Whats this "forget" trope of glue sniffing
>>>> Rossy Boy with his herpes blisters?
>>>>
>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>  > forget Hungarians and Bulgarians.
>>>>
>>>> Why should I forget Bulgarians,
>>>> they are never on my mind. Do you
>>>> see me doing ggml stuff?
>>>>
>>>> I only hypothesized that it is
>>>> over for Python as the machine
>>>> learning language or AI inferencing
>>>>
>>>> locally on AI laptops language, and
>>>> made the ggml case, so I already forgot
>>>> about them. Which might give you a glimps,
>>>>
>>>> why WebGPU was used for this here:
>>>>
>>>> 11.4 Giga Lips with a Budget Laptop
>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> Is an interesting choice. Even
>>>> github has some Languages statistics,
>>>> giving an account what I used:
>>>>
>>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>>
>>>> Have Fun!
>>>>
>>>> Bye
>>>>
>>>> P.S.: The example below is not p-adics,
>>>> you complete imbecil moron. Its just:
>>>>
>>>> 7-11 cubic Solution by Pritchard & Gries
>>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>>
>>>> Ross Finlayson schrieb:
>>>>> On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>>> Hi,
>>>>>>
>>>>>> You see it all boils down to find your inner peace
>>>>>> by an immaculate inception of some queue datatype.
>>>>>>
>>>>>> KOAN/Fortran-S was an early 1990s research programming
>>>>>> system for distributed-memory multiprocessors . Developed
>>>>>> at ENS Lyon in the early 1990s . Often listed alongside
>>>>>> other historical parallel programming efforts.
>>>>>>
>>>>>> The Message Passing: The research explicitly
>>>>>> compared the SVM approach against message passing
>>>>>> on the same hardware . The finding was that SVM
>>>>>> could achieve good performance without the low-level
>>>>>>
>>>>>> complexity of managing explicit messages, though
>>>>>> the best results often came from a hybrid approach (sic!)
>>>>>> Here is an interesting baseline, from Java,
>>>>>> a class ElevenSingle that only does:
>>>>>>
>>>>>>      public static void run() {
>>>>>>          for (int A = 1; A < 192; A++) {
>>>>>>              int Y = (771-A)/3;
>>>>>>              for (int B = A; B < Y; B++) {
>>>>>>                  int Z = (771-A-B)/2;
>>>>>>                  for (int C = B; C < Z; C++) {
>>>>>>                      int D = 711-A-B-C;
>>>>>>                      if (A*B*C == 711000000/D &&
>>>>>>                            711000000 % D == 0)
>>>>>>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>>>                  }
>>>>>>              }
>>>>>>          }
>>>>>>      }
>>>>>>
>>>>>> And then compare it to ElevenMulti, doing some
>>>>>> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>>>
>>>>>> ElevenSingle
>>>>>> A=120, B=125, C=150, D=316
>>>>>> 6.628 ms
>>>>>>
>>>>>> ElevenMulti
>>>>>> A=120, B=125, C=150, D=316
>>>>>> 1.941 ms
>>>>>>
>>>>>> Not great, not terrible!
>>>>>>
>>>>>> Bye
>>>>>
>>>>> Oh, that's just "tricks of p-adic arithmetic".
>>>>>
>>>>> Like other sock-puppet howler trolls, when confronted
>>>>> with its base incredulity, it will descend to its
>>>>> lower levers of the pathos variety.
>>>>>
>>>>> You might be happier learning about Julia trees and
>>>>> raster ops, instead of shilling yet another Ramanujan
>>>>> series without saying how it's made.
>>>>>
>>>>> Bulgarians, that's some real Boris and Natasha crap,
>>>>> forget Hungarians and Bulgarians.
>>>>>
>>>>>
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647017 — Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Re: The invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 17:14 +0200
SubjectMaybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Re: The invention of vector and matrix registers [NVIDIA Volta] (Was: AI accelerator extend from GPU to CPU [Zero Copying])
Message-ID<1147skb$geia$2@solani.org>
In reply to#647013
Hi,

I guess Rossy Boys mother was so disappointed
in the 50's that her son didn't become the
next Einstein, physics was the ultimate idol,

so that Rossy Boy was left rotting in the
basement. But Rossy Boys indoctrination was
not spurious, he now is conditioned on

Einstein. Maybe NVIDIA should have named
its Tesla V100 card NVIDIA Einstein. You would
then see Rossy Boy toe sucking the graphic

card, in his pyjamas in the basement.

Bye

Bye Ross Finlayson schrieb:
 > On 07/27/2026 04:22 AM, Mild Shock wrote:
 >> Hi,
 >>
 >> But the example gives also way to vector
 >> and matrix registers. The int[] X and
 >> int[] Y could be also held in vector
 >>
 >> registers. Compilers can also optimize
 >> away int[] Y, and use a inline modification,
 >> in case X isn't used later, then playing
 >>
 >> the role of Y:
 >>
 >> vec_mul_add(X, 2, 3, X)
 >>
 >> Vector and matrix registers in modern GPUs
 >> emerged from distinct architectural milestones:
 >> vector-like register files developed with
 >> early programmable 3D vertex/pixel pipelines
 >>
 >> in the late 1990s to early 2000s. While
 >> dedicated multi-dimensional matrix registers
 >> (Tensor Cores/Matrix Cores) were invented by
 >> NVIDIA in 2017, starting with the Tesla
 >>
 >> V100 (Volta microarchitecture):
 >>
 >>  From Volta To Blackwell
 >> 
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell
 >>
 >>
 >> You see the scheduling of tensure core occupation
 >> scheduling in the above article, including memory
 >> and register flow, following the section:
 >>
 >> MMA Instruction Overview
 >>
 >> It went through a couple of generations, leading
 >> to Tensor Memory (TMEM) and collective operations,
 >> basically realizing the PIM idea:
 >>
 >> Processing-in-Memory Tutorials
 >> 
https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/
 >>
 >>
 >> Have Fun!
 >>
 >> Bye
 > That's bullshit, and alike those talking heads that
 > sniff their way into talking about many-core jumbo-trons,
 > the super-scalar is as old as the scalar and Cray and examples alike
 > the Connection Machine what made all the craze of neural nets
 > is old-wrapped-as-new.
 >
 > Fabless chips did it already.
 >
 >
 > Data centers should pay a 10000% excise on electricity,
 > wherever it comes from, a natural regulator of inverted economies.
 >
 > And by ten thousand percent I really mean a ten thousand percent.

[toc] | [prev] | [next] | [standalone]


#647022 — π-WAM is not adding decimals, it is removing decimals (Re: The invention of vector and matrix registers [NVIDIA Volta])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:36 +0200
Subjectπ-WAM is not adding decimals, it is removing decimals (Re: The invention of vector and matrix registers [NVIDIA Volta])
Message-ID<11481dg$h2se$2@solani.org>
In reply to#647013
Hi,

Come on Horsy Boy, you can do better. I
no where wrote something about curve
fitting and/or increasing the precision of

float point numbers:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

What makes you think LIPS measures precision?
You should know better as a 50% Prologer.

I explictily wrote here what the goal is:

"shave off some of the TOPS to do Prolog inferencing"

What are TOPS? Its a metric for GPUs:

TOPS stands for “Trillions of Operations Per Second.”
https://www.lenovo.com/us/en/glossary/tops-in-computing/

See for yourself what is behind my post:

11.4 Giga Lips with a Budget Laptop
At the end of 2025 we acquired a couple of AI Laptops , that were still 
cheap, since RAM prices had not yet rocketed. The intend was to tap into 
the Copilot+ certified hardware, and shave off some of the TOPS to do 
Prolog inferencing. Amazingly our π-WAM can churn 11.4 GIGA LIPS.

GPUs have evolved form lock-step to independent thread scheduling. This 
made it possible to port the Hack VM variant, that forms the basis for 
our π-WAM, to WebGPU computer shaders. Using NUM_SHADERS = 4096 we could 
produce 11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M.

See also:

Medium Article - 11.4 Giga Lips
https://medium.com/2989/899b0d5c027b

So just get lost with your crazy irrelevant rant.
When I get more LIPS, things run faster, and
I remove digits from the time dimension.

Got it. Or are you too stupid?

Bye

R Kym Horsell schrieb:
 > In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
 > ...
 >> Data centers should pay a 10000% excise on electricity,
 >> wherever it comes from, a natural regulator of inverted economies.
 >> And by ten thousand percent I really mean a ten thousand percent.
 >
 > And what would a huge surcharge do?
 > Almost always end up affecting the less powerful end of society
 > with increased costs to services the AI industry will be doing
 > more and more of over time.
 >
 > I started a little data center (exaflops.com) many years ago.
 > In those distant days people (in fact one was a prof of computer
 > science) told me you could never make money running a supercomputer.
 > LOL.
 >
 > I've had many years to watch the trends and a far more efficient
 > way to solve resource problems in this area is to change the
 > algorithms. There is vast room for improvement, mostly because
 > of prevailing attitudes.
 >
 > I used to do competetion data science as a sideline. Companies
 > would pay almost any price to get an extra decimal place in
 > the accuracy of their forecasting processes. But typically
 > they were trying to supercharge a system that should be scrapped
 > and re-designed from scratch. One area I'm thinking of is
 > investment. I had a customer one time -- like many times --
 > ask to improve a system that predicted the future price of
 > various stocks. The idea (for them) was to have as accurate a
 > prediction of what some stock would be worth in a week or a month's
 > time so that some moron could use the information to decide when
 > to buy or sell the thing.
 >
 > I tried to argue the efficient thing was to create a system that
 > takes the human out of the loop altogether. It doesnt provide info
 > for someone to decide whether or not to follow the advice --
 > that is just introducing more noise into the loop and probably
 > cancels any benefit of adding a couple decimal places of precision.
 > What you *should* do is make a system that is tuned to robustly
 > maximize the profit from managing a portfolio.
 >
 > Of course they wouldnt come at that. You can't suggest taking the
 > managers out of the loop.
 >
 > Another idea relevant to current AI methods might be to curtail
 > use of typical neural net algorithms. Many of them try to squeeze
 > the best performance of some NN during the training  phase in
 > the hope the resulting system will generalize well enough to be useful
 > on new data. But there's kind-of a law that the harder you train
 > some system to perform a task well, the less well they can subsuently
 > perform a more general version of the same thing. It's amusing when
 > you look at the graphs of NN being trained and then tested that
 > given a more general problem to solve after being trained to solve
 > similar problems very very well the poor old NN does worse that it
 > would have done if it had 0 training in the first place.
 >
 > It's not like we dont know how to improve this kind of performance.
 > Try less hard in the training phase or make it "more noisy".
 > Turns out genetic methods are just the ticket for this.
 > The training produces less over-fitting and the resulting system
 > generalizes better than it did before training and more importantly
 > it takes maybe an order of magnitude crunching to produce a good answer
 > than the usual over-fit answer.
 >
 > Anyway. Have to go and feed the cat.
 >

[toc] | [prev] | [next] | [standalone]


#647023 — In Budget Laptops the TOPS come with low energy footprint (Was: π-WAM is not adding decimals, it is removing decimals)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:44 +0200
SubjectIn Budget Laptops the TOPS come with low energy footprint (Was: π-WAM is not adding decimals, it is removing decimals)
Message-ID<11481tk$h3a6$1@solani.org>
In reply to#647022
Hi,

Because of the mobile GPU design, the TOPS,
aka “Trillions of Operations Per Second.”
come with not extremly high power consumption.

Especially the presence of vector and matrix
operations can lower the energy consumption,
since they can avoid redundant memory access.

Its quite a difference between discrete graphic
cards and accelerator iGPUs that are directly
on the silicon chip, and have mobile design.

So basically with newer AI Laptops you get more
performence units for less energy units.

Have Fun!

Bye

Mild Shock schrieb:
> Hi,
> 
> Come on Horsy Boy, you can do better. I
> no where wrote something about curve
> fitting and/or increasing the precision of
> 
> float point numbers:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> What makes you think LIPS measures precision?
> You should know better as a 50% Prologer.
> 
> I explictily wrote here what the goal is:
> 
> "shave off some of the TOPS to do Prolog inferencing"
> 
> What are TOPS? Its a metric for GPUs:
> 
> TOPS stands for “Trillions of Operations Per Second.”
> https://www.lenovo.com/us/en/glossary/tops-in-computing/
> 
> See for yourself what is behind my post:
> 
> 11.4 Giga Lips with a Budget Laptop
> At the end of 2025 we acquired a couple of AI Laptops , that were still 
> cheap, since RAM prices had not yet rocketed. The intend was to tap into 
> the Copilot+ certified hardware, and shave off some of the TOPS to do 
> Prolog inferencing. Amazingly our π-WAM can churn 11.4 GIGA LIPS.
> 
> GPUs have evolved form lock-step to independent thread scheduling. This 
> made it possible to port the Hack VM variant, that forms the basis for 
> our π-WAM, to WebGPU computer shaders. Using NUM_SHADERS = 4096 we could 
> produce 11.4 Giga Lips on a Ryzen AI 7 350 w/ Radeon 860M.
> 
> See also:
> 
> Medium Article - 11.4 Giga Lips
> https://medium.com/2989/899b0d5c027b
> 
> So just get lost with your crazy irrelevant rant.
> When I get more LIPS, things run faster, and
> I remove digits from the time dimension.
> 
> Got it. Or are you too stupid?
> 
> Bye
> 
> R Kym Horsell schrieb:
>  > In comp.lang.prolog Ross Finlayson <ross.a.finlayson@gmail.com> wrote:
>  > ...
>  >> Data centers should pay a 10000% excise on electricity,
>  >> wherever it comes from, a natural regulator of inverted economies.
>  >> And by ten thousand percent I really mean a ten thousand percent.
>  >
>  > And what would a huge surcharge do?
>  > Almost always end up affecting the less powerful end of society
>  > with increased costs to services the AI industry will be doing
>  > more and more of over time.
>  >
>  > I started a little data center (exaflops.com) many years ago.
>  > In those distant days people (in fact one was a prof of computer
>  > science) told me you could never make money running a supercomputer.
>  > LOL.
>  >
>  > I've had many years to watch the trends and a far more efficient
>  > way to solve resource problems in this area is to change the
>  > algorithms. There is vast room for improvement, mostly because
>  > of prevailing attitudes.
>  >
>  > I used to do competetion data science as a sideline. Companies
>  > would pay almost any price to get an extra decimal place in
>  > the accuracy of their forecasting processes. But typically
>  > they were trying to supercharge a system that should be scrapped
>  > and re-designed from scratch. One area I'm thinking of is
>  > investment. I had a customer one time -- like many times --
>  > ask to improve a system that predicted the future price of
>  > various stocks. The idea (for them) was to have as accurate a
>  > prediction of what some stock would be worth in a week or a month's
>  > time so that some moron could use the information to decide when
>  > to buy or sell the thing.
>  >
>  > I tried to argue the efficient thing was to create a system that
>  > takes the human out of the loop altogether. It doesnt provide info
>  > for someone to decide whether or not to follow the advice --
>  > that is just introducing more noise into the loop and probably
>  > cancels any benefit of adding a couple decimal places of precision.
>  > What you *should* do is make a system that is tuned to robustly
>  > maximize the profit from managing a portfolio.
>  >
>  > Of course they wouldnt come at that. You can't suggest taking the
>  > managers out of the loop.
>  >
>  > Another idea relevant to current AI methods might be to curtail
>  > use of typical neural net algorithms. Many of them try to squeeze
>  > the best performance of some NN during the training  phase in
>  > the hope the resulting system will generalize well enough to be useful
>  > on new data. But there's kind-of a law that the harder you train
>  > some system to perform a task well, the less well they can subsuently
>  > perform a more general version of the same thing. It's amusing when
>  > you look at the graphs of NN being trained and then tested that
>  > given a more general problem to solve after being trained to solve
>  > similar problems very very well the poor old NN does worse that it
>  > would have done if it had 0 training in the first place.
>  >
>  > It's not like we dont know how to improve this kind of performance.
>  > Try less hard in the training phase or make it "more noisy".
>  > Turns out genetic methods are just the ticket for this.
>  > The training produces less over-fitting and the resulting system
>  > generalizes better than it did before training and more importantly
>  > it takes maybe an order of magnitude crunching to produce a good answer
>  > than the usual over-fit answer.
>  >
>  > Anyway. Have to go and feed the cat.
>  >
> 

[toc] | [prev] | [next] | [standalone]


#647035 — Java picky concerning JIT-ing [Luckier with C++/C or FORTRAN compilers?] (Re: The Koan of pi-WAM queues [FORTRAN-S])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 12:55 +0200
SubjectJava picky concerning JIT-ing [Luckier with C++/C or FORTRAN compilers?] (Re: The Koan of pi-WAM queues [FORTRAN-S])
Message-ID<114cm6u$jm45$2@solani.org>
In reply to#646999
Hi,

Small correction, the code below should use <=
in the for loops. But Java was extrem picky
concerning JIT-ing of the for loops, refuse

to JIT a <= based loop, so I rewrote a
corrected solution that matches:

7-11 cubic Solution by Pritchard & Gries
https://www.cs.cornell.edu/gries/TechReports/83-574.pdf

Into the following code:

     public static void run() {
         for (int A = 1; A < 193; A++) {
             int Y = (771-A)/3+1;
             for (int B = A; B < Y; B++) {
                 int Z = (771-A-B)/2+1;
                 for (int C = B; C < Z; C++) {
                     int D = 711-A-B-C;
                     if (A *B*C == 711000000/D && 711000000 % D == 0)
                         /* System.out.println("A="+A+", B="+B+", 
C="+C+", D="+D) */ ;
                 }
             }
         }
     }

Bye

Mild Shock schrieb:
> Hi,
> 
> You see it all boils down to find your inner peace
> by an immaculate inception of some queue datatype.
> 
> KOAN/Fortran-S was an early 1990s research programming
> system for distributed-memory multiprocessors . Developed
> at ENS Lyon in the early 1990s . Often listed alongside
> other historical parallel programming efforts.
> 
> The Message Passing: The research explicitly
> compared the SVM approach against message passing
> on the same hardware . The finding was that SVM
> could achieve good performance without the low-level
> 
> complexity of managing explicit messages, though
> the best results often came from a hybrid approach (sic!)
> Here is an interesting baseline, from Java,
> a class ElevenSingle that only does:
> 
>      public static void run() {
>          for (int A = 1; A < 192; A++) {
>              int Y = (771-A)/3;
>              for (int B = A; B < Y; B++) {
>                  int Z = (771-A-B)/2;
>                  for (int C = B; C < Z; C++) {
>                      int D = 711-A-B-C;
>                      if (A*B*C == 711000000/D &&
>                            711000000 % D == 0)
>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>                  }
>              }
>          }
>      }
> 
> And then compare it to ElevenMulti, doing some
> Work Balancing Scheduler Tetris Game with 8 cores:
> 
> ElevenSingle
> A=120, B=125, C=150, D=316
> 6.628 ms
> 
> ElevenMulti
> A=120, B=125, C=150, D=316
> 1.941 ms
> 
> Not great, not terrible!
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Its not tested on some Single Instruction/
>> Multiple Data (SIMD) GPU. It was only tested on
>> AI Laptops with Multiple instruction, Multiple
>>
>> Data (GPU) architecture for the scalar registers
>> per logical thread. As introduced by NVIDIA Volta
>> in around 2017:
>>
>>  > the first product was not announced until May 2017
>>  > https://en.wikipedia.org/wiki/Volta_%28microarchitecture%29
>>
>> Although I wrote the code of Hack VM with SIMD
>> in mind, I never tested it on a pure SIMD GPU,
>> and I never ported boot.mjs or boot2.mjs to
>>
>> WebGL2 / GLSL. I uploaded WebGPU / WGSL. Among the
>> tester I had were these AI Laptops, that could all
>> run WebGPU / WGSL in a browser:
>>
>>  > Intel(R) Core(TM) Ultra 7 258V
>>  > AMD Ryzen AI 7 350 w/ Radeon 860M
>>  > Apple A18 Pro, Darwin Kernel Version 25.5.0
>>  > Snapdragon(R) X - X126100 - Qualcomm(R) Oryon(TM) CPU
>>
>> Some AI Laptops had WebGPU / WGSL still behind
>> a browser flag, since its relatively new on ARM.
>> Also the above AI Laptops have all a iGPU and
>>
>> not a separate GPU card.
>>
>> Bye
>>
>> Mild Shock schrieb:> Hi,
>>  >
>>  >  > Show an outline of what you need you compute shader to do?
>>  >
>>  > Its all on GitHub , for the 100-th time .
>>  > Just RTFM , i.e. study the repo and the
>>  > medim article. Just follow this link:
>>  >
>>  > 11.4 Giga Lips with a Budget Laptop
>>  > https://github.com/Jean-Luc-Picard-2021/gigabudget
>>  >
>>  > Whats wrong with you guys, did the AI boom
>>  > suck out all your braincells. I really have
>>  > no words for being that stupid and slow.
>>  >
>>  > Bye
>>  >
>>  > In particular the repo contains two versions
>>  > of a Hack VM, written in WebGPU / WGSL:
>>  >
>>  > Hack VM: Version 1.0
>>  > 
>> https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/main/course/example63/boot.mjs 
>>
>>  >
>>  >
>>  > Hack VM: Version 2.0
>>  > 
>> https://github.com/Jean-Luc-Picard-2021/gigabudget/blob/main/course/example64/boot2.mjs 
>>
>>  >
>>  >
>>  > Version 1.0 is for a single compute shader
>>  > expriment. And Version 2.o is for a multi
>>  > compute shader experiment.
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Ride the snake
>>> He's old and his skin is cold
>>> The west is the best
>>> The west is the best
>>> Get here and we'll do the rest
>>> The blue bus is calling us
>>> The blue bus is calling us
>>> Driver, where you taking us?
>>>
>>> Apocalypse Now intro: The Doors, The End {1979}
>>> https://www.youtube.com/watch?v=CIrvSJwwJUE
>>>
>>> Bye
>>>
>>>> Hi,
>>>>
>>>> Again I posted everything here:
>>>>
>>>>> 11.4 Giga Lips with a Budget Laptop
>>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> The repo says, same time when I posted
>>>> the link first time:
>>>>
>>>>> This repository was archived by the
>>>>> owner on Jul 9, 2026. It is now read-only.
>>>>
>>>> Now a USENET user, who had already entitled
>>>> himself for a couple of irrational accusations
>>>>
>>>> towards my side, is asking this question:
>>>>
>>>> Chris M. Thomasson schrieb, Jul 24, 2026
>>>>> Show an outline of what you
>>>>> need you compute shader to do?
>>>>
>>>> Bravo, thats a delay of a wooping 15 days.
>>>>
>>>> Bye 
>>>
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Remember when first all local AI was Python
>>>> and PyTorch APIs. And then suddently people strated
>>>> using bare metal C/C++ Code. Here is the story:
>>>>
>>>> How it started:
>>>>
>>>> GPT-J or GPT-J-6B is an open-source large
>>>> language model (LLM) developed by EleutherAI
>>>> in 2021. As the name suggests, it is a
>>>> generative pre-trained transformer model
>>>> designed to produce human-like text that
>>>> continues from a prompt.
>>>> https://www.eleuther.ai/
>>>>
>>>> How it was going [Georgi Gerganov]:
>>>>
>>>> So a few days later comes out the LLaMA, I do
>>>> some calculations and I figure out “Okay, 65
>>>> billion parameters. You probably need about
>>>> 40 gigs of RAM, with 4-bit quantization. So
>>>> this can run on a MacBook. Why not do it?”
>>>>
>>>> Why I was able to do it so quickly - basically,
>>>> for all that I saw it’s pretty much GPT-J architecture
>>>> with some modifications, like some extra memorization
>>>> layers. It’s minor changes. Basically, again, the
>>>> existing code for the GPT-J, I just simply
>>>> modified it there, it happened pretty quickly.
>>>> https://changelog.com/podcast/532
>>>>
>>>> Georgi Gerganov, Bulgarian, now with Hugging
>>>> Face, ggml-cann also running on Chinese AI chips.
>>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>>
>>>> Bye
>>>>
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


#647015 — Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Re: Hurry the blue bus doesnt stop indefinitely)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 16:57 +0200
SubjectPotato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Re: Hurry the blue bus doesnt stop indefinitely)
Message-ID<1147rl7$gdpk$2@solani.org>
In reply to#646987
Hi,

Slowly I start understanding numbnuts like
Rossy Boy who don't understand tech, although
they are from UK and not from a 3rd world

country, and also I start understanding morons
like Micro Penis, who are behind a curtain,
and cannot access a lot of tech.

The same holds for SWI Prologs newest campaign
that probably adresses some poor indians that
have neither 5G nor Macs:

1:38:01 The Kyiv keynote disaster
https://www.youtube.com/watch?v=U8goS6B3BbI

Woa! Real time download of Scala, Closure,
etc.. Whats the magic behind that? Some SWI
point of sale, downloading it via its

keyboard and some telephathy module ?

Bye

Mild Shock schrieb:
> Hi,
> 
> Ride the snake
> He's old and his skin is cold
> The west is the best
> The west is the best
> Get here and we'll do the rest
> The blue bus is calling us
> The blue bus is calling us
> Driver, where you taking us?
> 
> Apocalypse Now intro: The Doors, The End {1979}
> https://www.youtube.com/watch?v=CIrvSJwwJUE
> 
> Bye
> 
>> Hi,
>>
>> Again I posted everything here:
>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>
>> The repo says, same time when I posted
>> the link first time:
>>
>>> This repository was archived by the
>>> owner on Jul 9, 2026. It is now read-only.
>>
>> Now a USENET user, who had already entitled
>> himself for a couple of irrational accusations
>>
>> towards my side, is asking this question:
>>
>> Chris M. Thomasson schrieb, Jul 24, 2026
>>> Show an outline of what you
>>> need you compute shader to do?
>>
>> Bravo, thats a delay of a wooping 15 days.
>>
>> Bye 
> 
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Remember when first all local AI was Python
>> and PyTorch APIs. And then suddently people strated
>> using bare metal C/C++ Code. Here is the story:
>>
>> How it started:
>>
>> GPT-J or GPT-J-6B is an open-source large
>> language model (LLM) developed by EleutherAI
>> in 2021. As the name suggests, it is a
>> generative pre-trained transformer model
>> designed to produce human-like text that
>> continues from a prompt.
>> https://www.eleuther.ai/
>>
>> How it was going [Georgi Gerganov]:
>>
>> So a few days later comes out the LLaMA, I do
>> some calculations and I figure out “Okay, 65
>> billion parameters. You probably need about
>> 40 gigs of RAM, with 4-bit quantization. So
>> this can run on a MacBook. Why not do it?”
>>
>> Why I was able to do it so quickly - basically,
>> for all that I saw it’s pretty much GPT-J architecture
>> with some modifications, like some extra memorization
>> layers. It’s minor changes. Basically, again, the
>> existing code for the GPT-J, I just simply
>> modified it there, it happened pretty quickly.
>> https://changelog.com/podcast/532
>>
>> Georgi Gerganov, Bulgarian, now with Hugging
>> Face, ggml-cann also running on Chinese AI chips.
>> ggml Manifesto https://github.com/ggml-org/ggml
>>
>> Bye
>>
> 

[toc] | [prev] | [next] | [standalone]


#647024 — Got it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 19:00 +0200
SubjectGot it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])
Message-ID<11482s3$h3vk$2@solani.org>
In reply to#646933
Hi,

Ok, guys lets face it. You are a bunch of
morons. When did I do this post:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Yes on Jul 9, 2026, now we have Jul 27, 2026.
Thats a wooping 18 days meanwhile.
And you still don't get the meaning and

implications of the post. Like you even
don't get what "budget" nowdays means in
terms of performance units and energy units?

And what LIPS means, drawn from TOPS,
in terms of applications? Shame on you guys!
You are a bunch of brainless idiots.

Bye

Mild Shock schrieb:
> Hi,
> 
> Remember when first all local AI was Python
> and PyTorch APIs. And then suddently people strated
> using bare metal C/C++ Code. Here is the story:
> 
> How it started:
> 
> GPT-J or GPT-J-6B is an open-source large
> language model (LLM) developed by EleutherAI
> in 2021. As the name suggests, it is a
> generative pre-trained transformer model
> designed to produce human-like text that
> continues from a prompt.
> https://www.eleuther.ai/
> 
> How it was going [Georgi Gerganov]:
> 
> So a few days later comes out the LLaMA, I do
> some calculations and I figure out “Okay, 65
> billion parameters. You probably need about
> 40 gigs of RAM, with 4-bit quantization. So
> this can run on a MacBook. Why not do it?”
> 
> Why I was able to do it so quickly - basically,
> for all that I saw it’s pretty much GPT-J architecture
> with some modifications, like some extra memorization
> layers. It’s minor changes. Basically, again, the
> existing code for the GPT-J, I just simply
> modified it there, it happened pretty quickly.
> https://changelog.com/podcast/532
> 
> Georgi Gerganov, Bulgarian, now with Hugging
> Face, ggml-cann also running on Chinese AI chips.
> ggml Manifesto https://github.com/ggml-org/ggml
> 
> Bye
> 

[toc] | [prev] | [next] | [standalone]


#647025 — Re: Got it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])

FromKim Baitchorov <bvhkoc@bmc.ru>
Date2026-07-27 22:38 +0000
SubjectRe: Got it. Or are you too stupid? [New Usenet Mantra] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])
Message-ID<1148mlj$35a76$1@news.nntp4.net>
In reply to#647024
Mild Shock wrote:

> And what LIPS means, drawn from TOPS,
> in terms of applications? Shame on you guys!
> You are a bunch of brainless idiots.

post the link, i want to buy one, then kiss my ass and fuck off. Git links 
are for fools like you are.

[toc] | [prev] | [next] | [standalone]


#647027 — Clueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-28 11:29 +0200
SubjectClueless about MIMD as usual [Flynn's Taxonomy] (Was: Rossy Boy is neither Einstein nor Zweistein)
Message-ID<1149spc$i9f4$3@solani.org>
In reply to#647025
Hi,

Moron there is no SIMT. As I already wrote:

 > He is also not Zweistein, since he doesn't
 > understand concepts such as:
 >
 > - NVIDIA Volta ff. architecture

But you had the SIMD and MIMD disctinction
alreay in OpenMP (via #pragma omp simd and
#pragma omp parallel):

Flynn's Taxonomy classifies computer architectures
according to how many instruction streams (processes)
and data streams they can process simultaneously,
dividing them into four categories:
SISD, SIMD, MISD, and MIMD.
https://www.geeksforgeeks.org/computer-organization-architecture/computer-architecture-flynns-taxonomy/

Its not so difficult to understand what
the NVIDIA Volta ff. architecture means.

Bye

Ross Finlayson schrieb:
 > They're considered really quite simple,
 > each of those threads is simple, SIMT.

Ross Finlayson schrieb:
 > On 06/25/2021 07:54 PM, Archimedes Plutonium wrote:
 >> On Monday, June 21, 2021 at 12:00:21 PM UTC-5, Graham Cooper wrote:
 >>> On Tuesday, June 22, 2021 at 2:54:40 AM UTC+10, burs...@gmail.com 
wrote:
 >>>> Try yourself:
 >>>> misc.prolog.compound.parenthesis.missing
 >>>>
 >>>> LMAO!
 >>>
 >>> Jan you work too hard. nobody wants theorem provers on prolog
 >>>
 >>> ASIMO tech is going to LISP which will just have a UNIFY routine
 >>>
 >>> but people can LEARN PROLOG if you EFF OFF!
 >>>
 >>>
 >>>
 >>> VOTE NOW! BAN JAN

[toc] | [prev] | [next] | [standalone]


#647030 — confused rossy boy is confused (Re: Clueless about MIMD as usual [Flynn's Taxonomy])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 11:20 +0200
Subjectconfused rossy boy is confused (Re: Clueless about MIMD as usual [Flynn's Taxonomy])
Message-ID<114cglk$k0r4$3@solani.org>
In reply to#647027
Hi,

Confused rossy boy is confused. We are
not building a stupid web server, where
a listener thread spawns service threads,

and to avoid malloc and free, reuses
a pool, or some shitty fork join framework.
The producer and consumer example I posted

elsewhere archived a dataflow without
malloc and free of threads. You are miles
away from what we are doing here.

Bye

>> Ross Finlayson schrieb:
> This is with infinity and continuity,
> SIMT is a worker pool. 

Mild Shock schrieb:
> Hi,
> 
> Moron there is no SIMT. As I already wrote:
> 
>  > He is also not Zweistein, since he doesn't
>  > understand concepts such as:
>  >
>  > - NVIDIA Volta ff. architecture
> 
> But you had the SIMD and MIMD disctinction
> alreay in OpenMP (via #pragma omp simd and
> #pragma omp parallel):
> 
> Flynn's Taxonomy classifies computer architectures
> according to how many instruction streams (processes)
> and data streams they can process simultaneously,
> dividing them into four categories:
> SISD, SIMD, MISD, and MIMD.
> https://www.geeksforgeeks.org/computer-organization-architecture/computer-architecture-flynns-taxonomy/ 
> 
> 
> Its not so difficult to understand what
> the NVIDIA Volta ff. architecture means.
> 
> Bye
> 
> Ross Finlayson schrieb:
>  > They're considered really quite simple,
>  > each of those threads is simple, SIMT.
> 
> Ross Finlayson schrieb:
>  > On 06/25/2021 07:54 PM, Archimedes Plutonium wrote:
>  >> On Monday, June 21, 2021 at 12:00:21 PM UTC-5, Graham Cooper wrote:
>  >>> On Tuesday, June 22, 2021 at 2:54:40 AM UTC+10, burs...@gmail.com 
> wrote:
>  >>>> Try yourself:
>  >>>> misc.prolog.compound.parenthesis.missing
>  >>>>
>  >>>> LMAO!
>  >>>
>  >>> Jan you work too hard. nobody wants theorem provers on prolog
>  >>>
>  >>> ASIMO tech is going to LISP which will just have a UNIFY routine
>  >>>
>  >>> but people can LEARN PROLOG if you EFF OFF!
>  >>>
>  >>>
>  >>>
>  >>> VOTE NOW! BAN JAN

[toc] | [prev] | [next] | [standalone]


#647031 — Gemini, DeepSeek, OpenAI more clever than rossy boy (Re: confused rossy boy is confused)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 11:21 +0200
SubjectGemini, DeepSeek, OpenAI more clever than rossy boy (Re: confused rossy boy is confused)
Message-ID<114cgn6$k0r4$4@solani.org>
In reply to#647030
Hi,

I already posted the candidate MPMC queue
to do these things. But my research is
not yet conclusive:

 > Its actually quite amazing. Gemini, DeepSeek,
 > OpenAI all know Dmitriy V'jukov. I have asked
 > the IntelliJ integrated Freeium AI to generate
 >
 > some code for me, I guess their service uses
 > by default OpenAI (Codex), and had it reviewed
 > by Gemini and DeepSeek. These AIs started lecturing
 >
 > me about lazySet() in Java. But I went with set():
 >
 >     private static boolean enqueue(Queue q, Object data) {
 >         int pos = q.enqueuePos.get();
 >         for (; ; ) {
 >             int index = pos & q.bufferMask;
 >             int seq = q.sequences.get(index);
 >             int dif = seq - pos;
 >             if (dif == 0) {
 >                 if (q.enqueuePos.compareAndSet(pos, pos + 1)) {
 >                     q.data[index] = data;
 >                     q.sequences.set(index, pos + 1);
 >                     return true;
 >                 }
 >                 pos = q.enqueuePos.get();
 >             } else if (dif < 0) {
 >                 return false;
 >             } else {
 >                 pos = q.enqueuePos.get();
 >             }
 >         }
 >     }
 >
 > The above version seems to be more suitable
 > for my purpose, since it allows polling, it
 > basically implements offer(). While the
 >
 > version posted on in the lock free group
 > by Chris M. Thomasson implements a spin wait
 > blocking put() already.

Bye

Mild Shock schrieb:
> Hi,
> 
> Confused rossy boy is confused. We are
> not building a stupid web server, where
> a listener thread spawns service threads,
> 
> and to avoid malloc and free, reuses
> a pool, or some shitty fork join framework.
> The producer and consumer example I posted
> 
> elsewhere archived a dataflow without
> malloc and free of threads. You are miles
> away from what we are doing here.
> 
> Bye
> 
>>> Ross Finlayson schrieb:
>> This is with infinity and continuity,
>> SIMT is a worker pool. 
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Moron there is no SIMT. As I already wrote:
>>
>>  > He is also not Zweistein, since he doesn't
>>  > understand concepts such as:
>>  >
>>  > - NVIDIA Volta ff. architecture
>>
>> But you had the SIMD and MIMD disctinction
>> alreay in OpenMP (via #pragma omp simd and
>> #pragma omp parallel):
>>
>> Flynn's Taxonomy classifies computer architectures
>> according to how many instruction streams (processes)
>> and data streams they can process simultaneously,
>> dividing them into four categories:
>> SISD, SIMD, MISD, and MIMD.
>> https://www.geeksforgeeks.org/computer-organization-architecture/computer-architecture-flynns-taxonomy/ 
>>
>>
>> Its not so difficult to understand what
>> the NVIDIA Volta ff. architecture means.
>>
>> Bye
>>
>> Ross Finlayson schrieb:
>>  > They're considered really quite simple,
>>  > each of those threads is simple, SIMT.
>>
>> Ross Finlayson schrieb:
>>  > On 06/25/2021 07:54 PM, Archimedes Plutonium wrote:
>>  >> On Monday, June 21, 2021 at 12:00:21 PM UTC-5, Graham Cooper wrote:
>>  >>> On Tuesday, June 22, 2021 at 2:54:40 AM UTC+10, burs...@gmail.com 
>> wrote:
>>  >>>> Try yourself:
>>  >>>> misc.prolog.compound.parenthesis.missing
>>  >>>>
>>  >>>> LMAO!
>>  >>>
>>  >>> Jan you work too hard. nobody wants theorem provers on prolog
>>  >>>
>>  >>> ASIMO tech is going to LISP which will just have a UNIFY routine
>>  >>>
>>  >>> but people can LEARN PROLOG if you EFF OFF!
>>  >>>
>>  >>>
>>  >>>
>>  >>> VOTE NOW! BAN JAN
> 

[toc] | [prev] | [next] | [standalone]


#647032 — In AI Acceleration nobody cares about CivetWeb (Re: Gemini, DeepSeek, OpenAI more clever than rossy boy)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 11:29 +0200
SubjectIn AI Acceleration nobody cares about CivetWeb (Re: Gemini, DeepSeek, OpenAI more clever than rossy boy)
Message-ID<114ch5q$k1i9$3@solani.org>
In reply to#647031
Hi,

Nobody cares about CivetWeb a C++/C library,
the rossy boy moron refuses to understand this
simple GPU test, that shows some AI Acceleration:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Bye

Mild Shock schrieb:
> Hi,
> 
> I already posted the candidate MPMC queue
> to do these things. But my research is
> not yet conclusive:
> 
>  > Its actually quite amazing. Gemini, DeepSeek,
>  > OpenAI all know Dmitriy V'jukov. I have asked
>  > the IntelliJ integrated Freeium AI to generate
>  >
>  > some code for me, I guess their service uses
>  > by default OpenAI (Codex), and had it reviewed
>  > by Gemini and DeepSeek. These AIs started lecturing
>  >
>  > me about lazySet() in Java. But I went with set():
>  >
>  >     private static boolean enqueue(Queue q, Object data) {
>  >         int pos = q.enqueuePos.get();
>  >         for (; ; ) {
>  >             int index = pos & q.bufferMask;
>  >             int seq = q.sequences.get(index);
>  >             int dif = seq - pos;
>  >             if (dif == 0) {
>  >                 if (q.enqueuePos.compareAndSet(pos, pos + 1)) {
>  >                     q.data[index] = data;
>  >                     q.sequences.set(index, pos + 1);
>  >                     return true;
>  >                 }
>  >                 pos = q.enqueuePos.get();
>  >             } else if (dif < 0) {
>  >                 return false;
>  >             } else {
>  >                 pos = q.enqueuePos.get();
>  >             }
>  >         }
>  >     }
>  >
>  > The above version seems to be more suitable
>  > for my purpose, since it allows polling, it
>  > basically implements offer(). While the
>  >
>  > version posted on in the lock free group
>  > by Chris M. Thomasson implements a spin wait
>  > blocking put() already.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Confused rossy boy is confused. We are
>> not building a stupid web server, where
>> a listener thread spawns service threads,
>>
>> and to avoid malloc and free, reuses
>> a pool, or some shitty fork join framework.
>> The producer and consumer example I posted
>>
>> elsewhere archived a dataflow without
>> malloc and free of threads. You are miles
>> away from what we are doing here.
>>
>> Bye
>>
>>>> Ross Finlayson schrieb:
>>> This is with infinity and continuity,
>>> SIMT is a worker pool. 
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Moron there is no SIMT. As I already wrote:
>>>
>>>  > He is also not Zweistein, since he doesn't
>>>  > understand concepts such as:
>>>  >
>>>  > - NVIDIA Volta ff. architecture
>>>
>>> But you had the SIMD and MIMD disctinction
>>> alreay in OpenMP (via #pragma omp simd and
>>> #pragma omp parallel):
>>>
>>> Flynn's Taxonomy classifies computer architectures
>>> according to how many instruction streams (processes)
>>> and data streams they can process simultaneously,
>>> dividing them into four categories:
>>> SISD, SIMD, MISD, and MIMD.
>>> https://www.geeksforgeeks.org/computer-organization-architecture/computer-architecture-flynns-taxonomy/ 
>>>
>>>
>>> Its not so difficult to understand what
>>> the NVIDIA Volta ff. architecture means.
>>>
>>> Bye
>>>
>>> Ross Finlayson schrieb:
>>>  > They're considered really quite simple,
>>>  > each of those threads is simple, SIMT.
>>>
>>> Ross Finlayson schrieb:
>>>  > On 06/25/2021 07:54 PM, Archimedes Plutonium wrote:
>>>  >> On Monday, June 21, 2021 at 12:00:21 PM UTC-5, Graham Cooper wrote:
>>>  >>> On Tuesday, June 22, 2021 at 2:54:40 AM UTC+10, 
>>> burs...@gmail.com wrote:
>>>  >>>> Try yourself:
>>>  >>>> misc.prolog.compound.parenthesis.missing
>>>  >>>>
>>>  >>>> LMAO!
>>>  >>>
>>>  >>> Jan you work too hard. nobody wants theorem provers on prolog
>>>  >>>
>>>  >>> ASIMO tech is going to LISP which will just have a UNIFY routine
>>>  >>>
>>>  >>> but people can LEARN PROLOG if you EFF OFF!
>>>  >>>
>>>  >>>
>>>  >>>
>>>  >>> VOTE NOW! BAN JAN
>>
> 

[toc] | [prev] | [next] | [standalone]


#647033 — Run with minimum HTTPS and .mjs type (Re: In AI Acceleration nobody cares about CivetWeb)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 11:49 +0200
SubjectRun with minimum HTTPS and .mjs type (Re: In AI Acceleration nobody cares about CivetWeb)
Message-ID<114cib5$k2b7$3@solani.org>
In reply to#647032
Hi,

Maybe there is a Rossy Boy flux generator
web server with infinity and continuity
HTTPS and .mjs type, aka SIMT halucination.

To run the GPU example that is written in HTML,
JavaScript and WebGPU / WGSL, the minium is
possibly a HTTPS server that can deliver the

right mime type for the .mjs extension. Its
then only a bundle of static pages that does
the demonstration. What worked on my side

is the IntelliJ browse button, which then uses
a small local server on its own, sandboxed to
serving some project files.

But this is only how to launch the test pages.

The Rossy Boy SIMT halucination, could also work, who knows?

Bye

Mild Shock schrieb:
> Hi,
> 
> Nobody cares about CivetWeb a C++/C library,
> the rossy boy moron refuses to understand this
> simple GPU test, that shows some AI Acceleration:
> 
> 11.4 Giga Lips with a Budget Laptop
> https://github.com/Jean-Luc-Picard-2021/gigabudget
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> I already posted the candidate MPMC queue
>> to do these things. But my research is
>> not yet conclusive:
>>
>>  > Its actually quite amazing. Gemini, DeepSeek,
>>  > OpenAI all know Dmitriy V'jukov. I have asked
>>  > the IntelliJ integrated Freeium AI to generate
>>  >
>>  > some code for me, I guess their service uses
>>  > by default OpenAI (Codex), and had it reviewed
>>  > by Gemini and DeepSeek. These AIs started lecturing
>>  >
>>  > me about lazySet() in Java. But I went with set():
>>  >
>>  >     private static boolean enqueue(Queue q, Object data) {
>>  >         int pos = q.enqueuePos.get();
>>  >         for (; ; ) {
>>  >             int index = pos & q.bufferMask;
>>  >             int seq = q.sequences.get(index);
>>  >             int dif = seq - pos;
>>  >             if (dif == 0) {
>>  >                 if (q.enqueuePos.compareAndSet(pos, pos + 1)) {
>>  >                     q.data[index] = data;
>>  >                     q.sequences.set(index, pos + 1);
>>  >                     return true;
>>  >                 }
>>  >                 pos = q.enqueuePos.get();
>>  >             } else if (dif < 0) {
>>  >                 return false;
>>  >             } else {
>>  >                 pos = q.enqueuePos.get();
>>  >             }
>>  >         }
>>  >     }
>>  >
>>  > The above version seems to be more suitable
>>  > for my purpose, since it allows polling, it
>>  > basically implements offer(). While the
>>  >
>>  > version posted on in the lock free group
>>  > by Chris M. Thomasson implements a spin wait
>>  > blocking put() already.
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Confused rossy boy is confused. We are
>>> not building a stupid web server, where
>>> a listener thread spawns service threads,
>>>
>>> and to avoid malloc and free, reuses
>>> a pool, or some shitty fork join framework.
>>> The producer and consumer example I posted
>>>
>>> elsewhere archived a dataflow without
>>> malloc and free of threads. You are miles
>>> away from what we are doing here.
>>>
>>> Bye
>>>
>>>>> Ross Finlayson schrieb:
>>>> This is with infinity and continuity,
>>>> SIMT is a worker pool. 
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Moron there is no SIMT. As I already wrote:
>>>>
>>>>  > He is also not Zweistein, since he doesn't
>>>>  > understand concepts such as:
>>>>  >
>>>>  > - NVIDIA Volta ff. architecture
>>>>
>>>> But you had the SIMD and MIMD disctinction
>>>> alreay in OpenMP (via #pragma omp simd and
>>>> #pragma omp parallel):
>>>>
>>>> Flynn's Taxonomy classifies computer architectures
>>>> according to how many instruction streams (processes)
>>>> and data streams they can process simultaneously,
>>>> dividing them into four categories:
>>>> SISD, SIMD, MISD, and MIMD.
>>>> https://www.geeksforgeeks.org/computer-organization-architecture/computer-architecture-flynns-taxonomy/ 
>>>>
>>>>
>>>> Its not so difficult to understand what
>>>> the NVIDIA Volta ff. architecture means.
>>>>
>>>> Bye
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > They're considered really quite simple,
>>>>  > each of those threads is simple, SIMT.
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > On 06/25/2021 07:54 PM, Archimedes Plutonium wrote:
>>>>  >> On Monday, June 21, 2021 at 12:00:21 PM UTC-5, Graham Cooper wrote:
>>>>  >>> On Tuesday, June 22, 2021 at 2:54:40 AM UTC+10, 
>>>> burs...@gmail.com wrote:
>>>>  >>>> Try yourself:
>>>>  >>>> misc.prolog.compound.parenthesis.missing
>>>>  >>>>
>>>>  >>>> LMAO!
>>>>  >>>
>>>>  >>> Jan you work too hard. nobody wants theorem provers on prolog
>>>>  >>>
>>>>  >>> ASIMO tech is going to LISP which will just have a UNIFY routine
>>>>  >>>
>>>>  >>> but people can LEARN PROLOG if you EFF OFF!
>>>>  >>>
>>>>  >>>
>>>>  >>>
>>>>  >>> VOTE NOW! BAN JAN
>>>
>>
> 

[toc] | [prev] | [next] | [standalone]


Page 3 of 4 — ← Prev page 1 2 [3] 4  Next page →

Back to top | Article view | sci.math


csiph-web