Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > sci.physics > #896394 > unrolled thread

The Wuhan Virus that destroyed Python [ggml Manifesto]

Started byMild Shock <janburse@fastmail.fm>
First post2026-07-22 21:01 +0200
Last post2026-07-29 17:04 +0200
Articles 4 on this page of 64 — 4 participants

Back to article view | Back to sci.physics


Contents

  The Wuhan Virus that destroyed Python [ggml Manifesto] Mild Shock <janburse@fastmail.fm> - 2026-07-22 21:01 +0200
    Deadlock Exorcism: Switch from Push to Pull [A pi-calculus Specification of Prolog] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:25 +0200
      Why do you even need a mpmc queue? [Thunder Kittens] (Re: Deadlock Exorcism: Switch from Push to Pull) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:45 +0200
        Trivial balancing example for (int i=0; i<global_id; i++) (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 08:55 +0200
          The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:19 +0200
            Enqueue/dequeue need not be fast and can spinn ["fairness" questions] (Re: The Pixel Phone AI Experiment Song (Enqueue/dequeue need not be fast and can spinn ["fairness" questions]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 09:23 +0200
        And, where did I talk about rockets? [Hint its about xAI's Grok] (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 01:25 +0200
          Why forget something, that was never on my mind (Re: And, where did I talk about rockets? [Hint its about xAI's Grok]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:49 +0200
        Example Mandel Brot rendering [Faster with MIMD] (Was: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-25 09:56 +0200
    Potential Python Recovery: Free Threading [3.13 release] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 10:21 +0200
    The things XILINX braught to the AMD table (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-23 18:48 +0200
      NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-23 19:13 +0200
        Micro penis mother sung arias (Re: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:40 +0200
          Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:27 +0200
            Ignoramus or Ignorabimus: I don't care [(Re: Micro penis brain is in constant hiatus (Re: Micro penis mother sung arias) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:35 +0200
              You are a moron, brainless putin payed (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:01 +0200
                Yeah keep reading my posts, uninspired fool (Re: You are a moron, brainless putin payed) Mild Shock <janburse@fastmail.fm> - 2026-07-24 19:47 +0200
              Out of the blue accusation span 15 days [Empirical USENET study] (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:27 +0200
              A brain desease of 20 days [Rossy Boy] (Re: Ignoramus or Ignorabimus: I don't care) Mild Shock <janburse@fastmail.fm> - 2026-07-29 18:41 +0200
                I didn't use a Ryzen Halo, whats wrong with you? (Re: A brain desease of 20 days [Rossy Boy]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:25 +0200
        Re: NVIDIA evacuated its Chinese market [Tau Scaling] (Re: The things XILINX braught to the AMD table) Mild Shock <janburse@fastmail.fm> - 2026-07-28 14:17 +0200
        ASML stocks are plunging, bye bye dutchies (Re: NVIDIA evacuated its Chinese market [Tau Scaling]) Mild Shock <janburse@fastmail.fm> - 2026-07-28 14:18 +0200
    Little Data Center on Your Palm [AI Laptops for 500 USD] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 17:59 +0200
      2008: 4 Blades + Tesla S1070 versus 2026: 1 AI Laptop (Re: Little Data Center on Your Palm [AI Laptops for 500 USD]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 18:15 +0200
    Hurry the blue bus doesnt stop indefinitely (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:37 +0200
      Not SIMD, a MIMD design for NVIDIA Volta (Re: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-24 20:58 +0200
        Could take 3-4 months find machine / browser (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-24 21:16 +0200
        The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-26 19:54 +0200
          The turbo capping of AI Laptops (Was: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-26 20:00 +0200
          Re: The Koan of pi-WAM queues [FORTRAN-S] (Re: Not SIMD, a MIMD design for NVIDIA Volta) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
          Why forget Bulgarians, never on my mind (Re: The Koan of pi-WAM queues [FORTRAN-S]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:16 +0200
            miniTriton CUDA is an alternative to torch variants (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:52 +0200
              Andrej Karpathy original gangster of Budget Laptop (Re: miniTriton CUDA is an alternative to torch variants) Mild Shock <janburse@fastmail.fm> - 2026-07-27 09:54 +0200
            The evolution of hardware and GPT-2 training (Re: Why forget Bulgarians, never on my mind) Mild Shock <janburse@fastmail.fm> - 2026-07-27 10:57 +0200
              How speed up π-WAM with vector operations (Re: The evolution of hardware and GPT-2 training) Mild Shock <janburse@fastmail.fm> - 2026-07-27 11:10 +0200
                AI accelerator extend from GPU to CPU [Zero Copying] (Re: How speed up π-WAM with vector operations) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:21 +0200
                  The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 13:22 +0200
                    Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-27 07:34 -0700
                      Maybe they should have named it NVIDIA Einstein [Rossy Boy Toe Sucking] (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 17:12 +0200
                      Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kym@sdf.com> - 2026-07-27 15:43 +0000
                        Re: The invention of vector and matrix registers [NVIDIA Volta] (Re: AI accelerator extend from GPU to CPU [Zero Copying]) R Kym Horsell <kymhorsell@gmail.com> - 2026-07-27 15:46 +0000
                        π-WAM is not adding decimals, it is removing decimals (Was: The invention of vector and matrix registers [NVIDIA Volta]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:34 +0200
                          In Budget Laptops the TOPS come with low energy footprint (Re: π-WAM is not adding decimals, it is removing decimals) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:45 +0200
      Potato Computer owner impressed by Ukraine Tech [Rossy Boys Brother?] (Was: Hurry the blue bus doesnt stop indefinitely) Mild Shock <janburse@fastmail.fm> - 2026-07-27 16:56 +0200
        Rossy Boy is neither Einstein nor Zweistein (Was: Potato Computer owner impressed by Ukraine Tech) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:25 +0200
          You are still chewing on SIMD. LoL (Re: Rossy Boy is neither Einstein nor Zweistein) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:12 +0200
            Hurry Rossy Boy, the blue bus is waiting (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:53 +0200
              Look how they advertized CUDA and logical threads (Re: Hurry Rossy Boy, the blue bus is waiting) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:55 +0200
                Forget any arithmetization of product FSA (Re: Look how they advertized CUDA and logical threads) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:58 +0200
            Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator] (Re: You are still chewing on SIMD. LoL) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:06 +0200
              I don't use Rust, you are crazy [Jump off a bridge, idiot] (Re: Rossy Boys tears could cool a data center [pi-WAM Interleaved Synchronized Emulator]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 20:25 +0200
                Hack ecosystem ignorance paired with paranoia [Nand to Tetris] (Re: I don't use Rust, you are crazy) Mild Shock <janburse@fastmail.fm> - 2026-07-29 22:52 +0200
                  A funny Q16.16 experiment with Hack (Re: Hack ecosystem ignorance paired with paranoia [Nand to Tetris]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 23:11 +0200
                    Summer Challenge: libSQL = Prolog+Modes [VDBE versus π-WAM] (Was: A funny Q16.16 experiment with Hack) Mild Shock <janburse@fastmail.fm> - 2026-07-30 11:25 +0200
                  I wrote Hack VM for π-WAM from scratch [4 Months total JavaScript, Python and Java] (Re: Hack ecosystem ignorance paired with paranoia) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:37 +0200
                    For WebGPU I first had SIMD in mind (Re: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 19:51 +0200
                      Corr.: 4 Months --> 4 Weeks (Re: For WebGPU I first had SIMD in mind) Mild Shock <janburse@fastmail.fm> - 2026-07-30 20:05 +0200
                    MIPS is a big Huffman mess [But Hack could do it] (Re: I wrote Hack VM for π-WAM from scratch) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:32 +0200
                      Not declarative with PHI (Φ) nodes (Was: MIPS is a big Huffman mess [But Hack could do it]) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:38 +0200
                      Not declarative with PHI (Φ) nodes (Re: MIPS is a big Huffman mess [But Hack could do it]) Mild Shock <janburse@fastmail.fm> - 2026-07-30 22:39 +0200
    Got it. Or are you too stupid? [New Usenet Mantra] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-27 18:59 +0200
    Lamas in a cradle and Lamas on the edge [Red Pyjama] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 13:03 +0200
      AI Accelerators and ISO Prolog multi-threading (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:02 +0200
        Actor/Erlang is dead, no Thread and Mailbox conflation [golang channels] (Re: AI Accelerators and ISO Prolog multi-threading) (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama]) Mild Shock <janburse@fastmail.fm> - 2026-07-29 17:04 +0200

Page 4 of 4 — ← Prev page 1 2 3 [4]


#896481 — Got it. Or are you too stupid? [New Usenet Mantra] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-27 18:59 +0200
SubjectGot it. Or are you too stupid? [New Usenet Mantra] (Was: The Wuhan Virus that destroyed Python [ggml Manifesto])
Message-ID<11482ou$h3vk$1@solani.org>
In reply to#896394
Hi,

Ok, guys lets face it. You are a bunch of
morons. When did I do this post:

11.4 Giga Lips with a Budget Laptop
https://github.com/Jean-Luc-Picard-2021/gigabudget

Yes on Jul 9, 2026, now we have Jul 27, 2026.
Thats a wooping 18 days meanwhile.
And you still don't get the meaning and

implications of the post. Like you even
don't get what "budget" nowdays means in
terms of performance units and energy units?

And what LIPS means, drawn from TOPS,
in terms of applications? Shame on you guys!
You are a bunch of brainless idiots.

Bye

Mild Shock schrieb:
> Hi,
> 
> Remember when first all local AI was Python
> and PyTorch APIs. And then suddently people started
> using bare metal C/C++ Code. Here is the story:
> 
> How it started:
> 
> GPT-J or GPT-J-6B is an open-source large
> language model (LLM) developed by EleutherAI
> in 2021. As the name suggests, it is a
> generative pre-trained transformer model
> designed to produce human-like text that
> continues from a prompt.
> https://www.eleuther.ai/
> 
> How it was going [Georgi Gerganov]:
> 
> So a few days later comes out the LLaMA, I do
> some calculations and I figure out “Okay, 65
> billion parameters. You probably need about
> 40 gigs of RAM, with 4-bit quantization. So
> this can run on a MacBook. Why not do it?”
> 
> Why I was able to do it so quickly - basically,
> for all that I saw it’s pretty much GPT-J architecture
> with some modifications, like some extra memorization
> layers. It’s minor changes. Basically, again, the
> existing code for the GPT-J, I just simply
> modified it there, it happened pretty quickly.
> https://changelog.com/podcast/532
> 
> Georgi Gerganov, Bulgarian, now with Hugging
> Face, ggml-cann also running on Chinese AI chips.
> ggml Manifesto https://github.com/ggml-org/ggml
> 
> Bye

[toc] | [prev] | [next] | [standalone]


#896505 — Lamas in a cradle and Lamas on the edge [Red Pyjama] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 13:03 +0200
SubjectLamas in a cradle and Lamas on the edge [Red Pyjama] (Re: The Wuhan Virus that destroyed Python [ggml Manifesto])
Message-ID<114cmmv$jmhc$2@solani.org>
In reply to#896394
Hi,

Why does this Lama have a red pyjama.
Oh, its a baby Lama. Its still in the cradle
and needs some training:

RedPajama-Data-v2
https://github.com/togethercomputer/RedPajama-Data

But then Andrej Karpathy recently showed
GPT-2 training on rented GPUs for less
than 100 USD in less then 2 hours.

So where do these grown up Lamas go.
Well Georgi Gerganov prefered C++/C
when he shouted Llama Llama Red Pyjama.

But you also find WebLLM, wrapping the
underlying C++/C GPU interface via the
W3C standard WebGPU / WGSL, with JavaScript:

In-Browser LLM Inference Engine
https://webllm.mlc.ai/

My experience with WebLLM 6 months
ago on an iPad Pro 2024, still a little early
stage performance and robustness.

But hey hardware of AI mobile iGPUs is
still evolving, and AI laptop, AI smartphones
and AI tablets, will soon feature Chinese

hardware such some new Kirin AI in 2027.

Bye

Mild Shock schrieb:
> Hi,
> 
> Remember when first all local AI was Python
> and PyTorch APIs. And then suddently people started
> using bare metal C/C++ Code. Here is the story:
> 
> How it started:
> 
> GPT-J or GPT-J-6B is an open-source large
> language model (LLM) developed by EleutherAI
> in 2021. As the name suggests, it is a
> generative pre-trained transformer model
> designed to produce human-like text that
> continues from a prompt.
> https://www.eleuther.ai/
> 
> How it was going [Georgi Gerganov]:
> 
> So a few days later comes out the LLaMA, I do
> some calculations and I figure out “Okay, 65
> billion parameters. You probably need about
> 40 gigs of RAM, with 4-bit quantization. So
> this can run on a MacBook. Why not do it?”
> 
> Why I was able to do it so quickly - basically,
> for all that I saw it’s pretty much GPT-J architecture
> with some modifications, like some extra memorization
> layers. It’s minor changes. Basically, again, the
> existing code for the GPT-J, I just simply
> modified it there, it happened pretty quickly.
> https://changelog.com/podcast/532
> 
> Georgi Gerganov, Bulgarian, now with Hugging
> Face, ggml-cann also running on Chinese AI chips.
> ggml Manifesto https://github.com/ggml-org/ggml
> 
> Bye

[toc] | [prev] | [next] | [standalone]


#896506 — AI Accelerators and ISO Prolog multi-threading (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:02 +0200
SubjectAI Accelerators and ISO Prolog multi-threading (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama])
Message-ID<114d4lt$kfj1$2@solani.org>
In reply to#896505
Hi,

Usual question:

 > Why implement both pre-emptive threading
AND cooperative tasks/engines?

I had implemented the ISO proposal in formerly Jekejeke
Prolog, you find the ISO proposal here:

ISO/IEC DTR 13211–5:2007
Prolog multi-threading support
https://logtalk.org/plstd/threads.pdf

But the ISO proposal doesn't match modern WebGPU APIs,
where your logical threads can live remotely in a dedicated GPU
in the VRAM there, and where you would have launch

parameters that say: Hey please run 4096 compute
shaders for me, that have independet thread state. Using
cooperative multi-tasking as the orchestrator works well.

Bye

Mild Shock schrieb:
> Hi,
> 
> Why does this Lama have a red pyjama.
> Oh, its a baby Lama. Its still in the cradle
> and needs some training:
> 
> RedPajama-Data-v2
> https://github.com/togethercomputer/RedPajama-Data
> 
> But then Andrej Karpathy recently showed
> GPT-2 training on rented GPUs for less
> than 100 USD in less then 2 hours.
> 
> So where do these grown up Lamas go.
> Well Georgi Gerganov prefered C++/C
> when he shouted Llama Llama Red Pyjama.
> 
> But you also find WebLLM, wrapping the
> underlying C++/C GPU interface via the
> W3C standard WebGPU / WGSL, with JavaScript:
> 
> In-Browser LLM Inference Engine
> https://webllm.mlc.ai/
> 
> My experience with WebLLM 6 months
> ago on an iPad Pro 2024, still a little early
> stage performance and robustness.
> 
> But hey hardware of AI mobile iGPUs is
> still evolving, and AI laptop, AI smartphones
> and AI tablets, will soon feature Chinese
> 
> hardware such some new Kirin AI in 2027.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Remember when first all local AI was Python
>> and PyTorch APIs. And then suddently people started
>> using bare metal C/C++ Code. Here is the story:
>>
>> How it started:
>>
>> GPT-J or GPT-J-6B is an open-source large
>> language model (LLM) developed by EleutherAI
>> in 2021. As the name suggests, it is a
>> generative pre-trained transformer model
>> designed to produce human-like text that
>> continues from a prompt.
>> https://www.eleuther.ai/
>>
>> How it was going [Georgi Gerganov]:
>>
>> So a few days later comes out the LLaMA, I do
>> some calculations and I figure out “Okay, 65
>> billion parameters. You probably need about
>> 40 gigs of RAM, with 4-bit quantization. So
>> this can run on a MacBook. Why not do it?”
>>
>> Why I was able to do it so quickly - basically,
>> for all that I saw it’s pretty much GPT-J architecture
>> with some modifications, like some extra memorization
>> layers. It’s minor changes. Basically, again, the
>> existing code for the GPT-J, I just simply
>> modified it there, it happened pretty quickly.
>> https://changelog.com/podcast/532
>>
>> Georgi Gerganov, Bulgarian, now with Hugging
>> Face, ggml-cann also running on Chinese AI chips.
>> ggml Manifesto https://github.com/ggml-org/ggml
>>
>> Bye
> 

[toc] | [prev] | [next] | [standalone]


#896507 — Actor/Erlang is dead, no Thread and Mailbox conflation [golang channels] (Re: AI Accelerators and ISO Prolog multi-threading) (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-29 17:04 +0200
SubjectActor/Erlang is dead, no Thread and Mailbox conflation [golang channels] (Re: AI Accelerators and ISO Prolog multi-threading) (Re: Lamas in a cradle and Lamas on the edge [Red Pyjama])
Message-ID<114d4p7$kfj1$3@solani.org>
In reply to#896506
Hi,

Mostlikely for high performance computing à la,
the Actor/Erlang model is dead, they might rely
on MPMC (Multiple Producer, Multiple Consumer)

queue entities separate from the threads. The
ISO Prolog multi-threading support had also such
threads. But besides that was also Actor/Erlang

leaning in practice, like SWI, where threads
have some default queues. So an actor is basically
a Thread and Mailbox conflation. While a MPMC queue

is a kind of separate Mailbox, where multiple
"actors" can read from and write from. A kind of
localized Linda Tuple store.

Which Programming language did adopted the
non-Actor pi-calculus model? Right golang
with its channels.

Bye


Mild Shock schrieb:
> Hi,
> 
> Usual question:
> 
>  > Why implement both pre-emptive threading
> AND cooperative tasks/engines?
> 
> I had implemented the ISO proposal in formerly Jekejeke
> Prolog, you find the ISO proposal here:
> 
> ISO/IEC DTR 13211–5:2007
> Prolog multi-threading support
> https://logtalk.org/plstd/threads.pdf
> 
> But the ISO proposal doesn't match modern WebGPU APIs,
> where your logical threads can live remotely in a dedicated GPU
> in the VRAM there, and where you would have launch
> 
> parameters that say: Hey please run 4096 compute
> shaders for me, that have independet thread state. Using
> cooperative multi-tasking as the orchestrator works well.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> Why does this Lama have a red pyjama.
>> Oh, its a baby Lama. Its still in the cradle
>> and needs some training:
>>
>> RedPajama-Data-v2
>> https://github.com/togethercomputer/RedPajama-Data
>>
>> But then Andrej Karpathy recently showed
>> GPT-2 training on rented GPUs for less
>> than 100 USD in less then 2 hours.
>>
>> So where do these grown up Lamas go.
>> Well Georgi Gerganov prefered C++/C
>> when he shouted Llama Llama Red Pyjama.
>>
>> But you also find WebLLM, wrapping the
>> underlying C++/C GPU interface via the
>> W3C standard WebGPU / WGSL, with JavaScript:
>>
>> In-Browser LLM Inference Engine
>> https://webllm.mlc.ai/
>>
>> My experience with WebLLM 6 months
>> ago on an iPad Pro 2024, still a little early
>> stage performance and robustness.
>>
>> But hey hardware of AI mobile iGPUs is
>> still evolving, and AI laptop, AI smartphones
>> and AI tablets, will soon feature Chinese
>>
>> hardware such some new Kirin AI in 2027.
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Remember when first all local AI was Python
>>> and PyTorch APIs. And then suddently people started
>>> using bare metal C/C++ Code. Here is the story:
>>>
>>> How it started:
>>>
>>> GPT-J or GPT-J-6B is an open-source large
>>> language model (LLM) developed by EleutherAI
>>> in 2021. As the name suggests, it is a
>>> generative pre-trained transformer model
>>> designed to produce human-like text that
>>> continues from a prompt.
>>> https://www.eleuther.ai/
>>>
>>> How it was going [Georgi Gerganov]:
>>>
>>> So a few days later comes out the LLaMA, I do
>>> some calculations and I figure out “Okay, 65
>>> billion parameters. You probably need about
>>> 40 gigs of RAM, with 4-bit quantization. So
>>> this can run on a MacBook. Why not do it?”
>>>
>>> Why I was able to do it so quickly - basically,
>>> for all that I saw it’s pretty much GPT-J architecture
>>> with some modifications, like some extra memorization
>>> layers. It’s minor changes. Basically, again, the
>>> existing code for the GPT-J, I just simply
>>> modified it there, it happened pretty quickly.
>>> https://changelog.com/podcast/532
>>>
>>> Georgi Gerganov, Bulgarian, now with Hugging
>>> Face, ggml-cann also running on Chinese AI chips.
>>> ggml Manifesto https://github.com/ggml-org/ggml
>>>
>>> Bye
>>
> 

[toc] | [prev] | [standalone]


Page 4 of 4 — ← Prev page 1 2 3 [4]

Back to top | Article view | sci.physics


csiph-web