Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > sci.logic > #348252

The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])

From Mild Shock <janburse@fastmail.fm>
Newsgroups sci.logic, comp.lang.prolog, sci.physics
Subject The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying])
Date 2026-07-27 13:22 +0200
Message-ID <1147f1f$glbs$3@solani.org> (permalink)
References (3 earlier) <1145hle$fdqk$2@solani.org> <11470l1$gapi$3@solani.org> <11476i7$ftjl$2@solani.org> <11477a4$fu5f$2@solani.org> <1147evl$glbs$2@solani.org>

Cross-posted to 3 groups.

Show all headers | View raw


Hi,

But the example gives also way to vector
and matrix registers. The int[] X and
int[] Y could be also held in vector

registers. Compilers can also optimize
away int[] Y, and use a inline modification,
in case X isn't used later, then playing

the role of Y:

vec_mul_add(X, 2, 3, X)

Vector and matrix registers in modern GPUs
emerged from distinct architectural milestones:
vector-like register files developed with
early programmable 3D vertex/pixel pipelines

in the late 1990s to early 2000s. While
dedicated multi-dimensional matrix registers
(Tensor Cores/Matrix Cores) were invented by
NVIDIA in 2017, starting with the Tesla

V100 (Volta microarchitecture):

 From Volta To Blackwell
https://newsletter.semianalysis.com/p/nvidia-tensor-core-evolution-from-volta-to-blackwell

You see the scheduling of tensure core occupation
scheduling in the above article, including memory
and register flow, following the section:

MMA Instruction Overview

It went through a couple of generations, leading
to Tensor Memory (TMEM) and collective operations,
basically realizing the PIM idea:

Processing-in-Memory Tutorials
https://www.sigarch.org/processing-in-memory-tutorials-experiences-from-past-two-years-and-thoughts-looking-forward/

Have Fun!

Bye


Mild Shock schrieb:
> Hi,
> 
> The nice thing about AI accelerators, pioneered
> maybe by Apple Silicon and their unified memory.
> The AMD APU model can be extended so that
> 
> vector and matrix operations become uniformly
> available for GPU and CPU. With unified memory
> already a vector operation such as:
> 
> vec_mul_add(X, 2, 3, Y)
> 
> Only needs the X and Y address. But I havent
> got my head around yet how this is all organized.
> Maybe a GPU has still its own GEMM cores,
> 
> but you find Apple Silicon C++/C source code,
> that taps into vector and matrix operations
> by Zero Copying. The Copying is left to the DMA
> 
> of the vector or matrix operation. And moderated
> by the various caches. Leading to the slogan, that
> multiple floating point operations become zero cost:
> 
> Some teaching can be found here
> https://www.hpc-ch.org/category/topics/course-workshop/
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> One could critisize that my π-WAM doesn't
>> utilize GPU to the fullest, since its GPU
>> backend prototype only uses scalar operations
>>
>> and no vector or matrix operations. And
>> modern GPUs thrive on vector and matrix
>> operations. Especially matrix operations giving
>>
>> a boost of a factor 15x or so. There are
>> many papers already showing how Prolog can be
>> mapped to matrix operations. Only this research
>>
>> is completely ignored by Prolog systems such as
>> SICStus, Ciao, SWI, ECLiPSe etc.. But lets
>> illustrate what vector operations could do
>>
>> for π-WAM, take this compilation of the Prolog
>> goal between(0,1023,X), Y is X*2+3:
>>
>> int X;
>> int Y;
>> for (X=0; X < 1024; X++) {
>>      Y=X*2+3;
>>      [...]
>> }
>>
>> With vector operations, and vectors of size
>> 32 one could do:
>>
>> int X1;
>> int[] X = new int[32];
>> int X3;
>> int[] Y = new int[32];
>> for (X1 = 0; X1 < 1024 / 32; X1++) {
>>      for (int X2 = 0; X2 < 32; X2++)
>>         X[X2] = X1*32+X2;
>>      vec_mul_add(X, 2, 3, Y);
>>      [..]
>> }
>>
>> Have Fun!
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> While Huggingfaces hired GG in 2026,
>>> AK was hired by Anthropic in 2026:
>>>
>>> Andrej Karpathy (born 23 October 1986[3])
>>> is a Slovak-Canadian AI researcher, who
>>> co-founded and formerly worked at OpenAI
>>> In 2026 he joined Anthropic as part of
>>> the pretraining team.
>>> https://en.wikipedia.org/wiki/Andrej_Karpathy
>>>
>>> But his nanochat archivement has an
>>> interesting time line:
>>>
>>> 168 hours , Original OpenAI GPT-2 checkpoint, 2019
>>> 3 hours , d24 baseline, slightly overtrained, Jan 29 2026
>>> 1 1/2 hour, autoresearch round 2, Mar 14 2026
>>> The best ChatGPT that $100 can buy.
>>> https://github.com/karpathy/nanochat
>>>
>>> But what hardware was the enabler. What is the
>>> NVIDIA H100 GPU even. Well the thingy is surely not
>>> a Budget Laptop, performance pretty much
>>>
>>> dependence on data elememt size, the H100 NVL
>>> version (*), and when using tensor operations,
>>> and not only scalar operations:
>>>
>>> 8-bit towards 3000 tera flops
>>> 16-bit towards 1500 tera flops
>>> 32-bit towards 900 tera flops
>>>
>>> Cool! I guess this experiment would tap into 60
>>> tera flops, since it only uses scalar operations so far:
>>>
>>> 11.4 Giga Lips with a Budget Laptop
>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>
>>> You could perform it by migration the web application
>>> using WebGPU into a node.js standalone application
>>> using the dawn library for GPU access.
>>>
>>> Bye
>>>
>>> (*) 
>>> https://resources.nvidia.com/en-us-hopper-architecture/nvidia-tensor-core-gpu-datasheet 
>>>
>>>
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>> Whats this "forget" trope of glue sniffing
>>>> Rossy Boy with his herpes blisters?
>>>>
>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>  > forget Hungarians and Bulgarians.
>>>>
>>>> Why should I forget Bulgarians,
>>>> they are never on my mind. Do you
>>>> see me doing ggml stuff?
>>>>
>>>> I only hypothesized that it is
>>>> over for Python as the machine
>>>> learning language or AI inferencing
>>>>
>>>> locally on AI laptops language, and
>>>> made the ggml case, so I already forgot
>>>> about them. Which might give you a glimps,
>>>>
>>>> why WebGPU was used for this here:
>>>>
>>>> 11.4 Giga Lips with a Budget Laptop
>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> Is an interesting choice. Even
>>>> github has some Languages statistics,
>>>> giving an account what I used:
>>>>
>>>> HTML 67.5% JavaScript 23.1% CSS 9.4%
>>>>
>>>> Have Fun!
>>>>
>>>> Bye
>>>>
>>>> P.S.: The example below is not p-adics,
>>>> you complete imbecil moron. Its just:
>>>>
>>>> 7-11 cubic Solution by Pritchard & Gries
>>>> https://www.cs.cornell.edu/gries/TechReports/83-574.pdf
>>>>
>>>> Ross Finlayson schrieb:
>>>>  > On 07/26/2026 10:52 AM, Mild Shock wrote:
>>>>  >> Hi,
>>>>  >>
>>>>  >> You see it all boils down to find your inner peace
>>>>  >> by an immaculate inception of some queue datatype.
>>>>  >>
>>>>  >> KOAN/Fortran-S was an early 1990s research programming
>>>>  >> system for distributed-memory multiprocessors . Developed
>>>>  >> at ENS Lyon in the early 1990s . Often listed alongside
>>>>  >> other historical parallel programming efforts.
>>>>  >>
>>>>  >> The Message Passing: The research explicitly
>>>>  >> compared the SVM approach against message passing
>>>>  >> on the same hardware . The finding was that SVM
>>>>  >> could achieve good performance without the low-level
>>>>  >>
>>>>  >> complexity of managing explicit messages, though
>>>>  >> the best results often came from a hybrid approach (sic!)
>>>>  >> Here is an interesting baseline, from Java,
>>>>  >> a class ElevenSingle that only does:
>>>>  >>
>>>>  >>      public static void run() {
>>>>  >>          for (int A = 1; A < 192; A++) {
>>>>  >>              int Y = (771-A)/3;
>>>>  >>              for (int B = A; B < Y; B++) {
>>>>  >>                  int Z = (771-A-B)/2;
>>>>  >>                  for (int C = B; C < Z; C++) {
>>>>  >>                      int D = 711-A-B-C;
>>>>  >>                      if (A*B*C == 711000000/D &&
>>>>  >>                            711000000 % D == 0)
>>>>  >>      System.out.println("A="+A+", B="+B+", C="+C+", D="+D);
>>>>  >>                  }
>>>>  >>              }
>>>>  >>          }
>>>>  >>      }
>>>>  >>
>>>>  >> And then compare it to ElevenMulti, doing some
>>>>  >> Work Balancing Scheduler Tetris Game with 8 cores:
>>>>  >>
>>>>  >> ElevenSingle
>>>>  >> A=120, B=125, C=150, D=316
>>>>  >> 6.628 ms
>>>>  >>
>>>>  >> ElevenMulti
>>>>  >> A=120, B=125, C=150, D=316
>>>>  >> 1.941 ms
>>>>  >>
>>>>  >> Not great, not terrible!
>>>>  >>
>>>>  >> Bye
>>>>  >
>>>>  > Oh, that's just "tricks of p-adic arithmetic".
>>>>  >
>>>>  > Like other sock-puppet howler trolls, when confronted
>>>>  > with its base incredulity, it will descend to its
>>>>  > lower levers of the pathos variety.
>>>>  >
>>>>  > You might be happier learning about Julia trees and
>>>>  > raster ops, instead of shilling yet another Ramanujan
>>>>  > series without saying how it's made.
>>>>  >
>>>>  > Bulgarians, that's some real Boris and Natasha crap,
>>>>  > forget Hungarians and Bulgarians.
>>>>  >
>>>>  >
>>>>
>>>
>>
> 

Back to sci.logic | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

The Wuhan Virus that destroyed Python [ggml Manifesto] Mild Shock <janburse@fastmail.fm> - 2026-07-22 21:01 +0200
  Deadlock Exorcism: Switch from Push to Pull [A pi-calculus Specification of Prolog] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:25 +0200
    Why do you even need a mpmc queue? [Thunder Kittens] (Re: Deadlock Exorcism: Switch from Push to Pull) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:45 +0200
      Trivial balancing example for (int i=0; i<global_id; i++) (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:55 +0200
        The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:19 +0200
          Enqueue/dequeue need not be fast and can spinn ["fairness" questions] (Re: The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:23 +0200
      And, where did I talk about rockets? [Hint its about xAI's Grok] (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 01:25 +0200
        Why forget something, that was never on my mind (Re: And, where did I talk about rockets? [Hint its about xAI's Grok]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:49 +0200
      Example Mandel Brot rendering [Faster with MIMD] (Was: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:56 +0200
  Potential Python Recovery: Free Threading [3.13 release] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 10:21 +0200
  The things XILINX braught to the AMD table (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 18:48 +0200
    NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-23 19:13 +0200
      Micro penis mother sung arias (Re: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:40 +0200
        Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:27 +0200
          Ignoramus or Ignorabimus: I don't care [(Re: Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:35 +0200
            You are a moron, brainless putin payed (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:01 +0200
              Yeah keep reading my posts, uninspired fool (Re: You are a moron, brainless putin payed) Mild Shock <janburse@fastmail.fm> - 2026-07-24 19:47 +0200
            Out of the blue accusation span 15 days [Empirical USENET study] (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:27 +0200
  Little Data Center on Your Palm [AI Laptops for 500 USD] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 17:59 +0200
    2008: 4 Blades + Tesla S1070 versus 2026: 1 AI Laptop (Re: Little Data Center on Your Palm [AI Laptops for 500 USD]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:15 +0200
  Hurry the blue bus doesnt stop indefinitely (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:37 +0200
    Not SIMD, a MIMD design for NVIDIA Volta (Re: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:58 +0200
      Could take 3-4 months find machine / browser (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-24 21:16 +0200
      The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-26 19:54 +0200
        The turbo capping of AI Laptops (Was: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-26 20:00 +0200
        Re: The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
        Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
          miniTriton CUDA is an alternative to torch variants (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:52 +0200
            Andrej Karpathy original gangster of Budget Laptop (Re: miniTriton CUDA is an alternative to torch variants) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:54 +0200
          The evolution of hardware and GPT-2 training (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 10:57 +0200
            How speed up π-WAM with vector operations (Re: The evolution of hardware and GPT-2 training) Mild Shock <janburse@fastmail.fm> - 2026-07-27 11:10 +0200
              AI accelerator extend from GPU to CPU [Zero Copying] (Re: How speed up π-WAM with vector operations) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:21 +0200
                The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:22 +0200
                Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 07:34 -0700
                Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 17:12 +0200
                Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kym@sdf.com> - 2026-07-27 15:43 +0000
                Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kymhorsell@gmail.com> - 2026-07-27 15:46 +0000
                π-WAM is not adding decimals, it is removing decimals (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:34 +0200
                In Budget Laptops the TOPS come with low energy footprint (Re: π-WAM is not adding decimals, it is removing decimals) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:45 +0200
    Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Was: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-27 16:56 +0200
      Rossy Boy is neither Einstein nor Zweistein (Was: Potato Computer owner impressed by Ukraine Tech) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:25 +0200
  Got it. Or are you too stupid? [New Usenet Mantra] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:59 +0200

csiph-web