Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > sci.physics > #896319 > unrolled thread

I'm a spinner, I'm a sinner [Dmitry Vyukov for pi-WAM] (Re: Paul Tarau versus Mr. Taskmanager, who would win? [A PDP-11 Humunkulus from 1979]

Started byMild Shock <janburse@fastmail.fm>
First post2026-07-19 11:55 +0200
Last post2026-07-24 15:00 +0200
Articles 11 on this page of 51 — 4 participants

Back to article view | Back to sci.physics

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  I'm a spinner, I'm a sinner [Dmitry Vyukov for pi-WAM] (Re: Paul Tarau versus Mr. Taskmanager, who would win? [A PDP-11 Humunkulus from 1979] Mild Shock <janburse@fastmail.fm> - 2026-07-19 11:55 +0200
    Gemini, DeepSeek, OpenAI all know Dmitriy V'jukov (Re: I'm a spinner, I'm a sinner [Dmitry Vyukov for pi-WAM]) Mild Shock <janburse@fastmail.fm> - 2026-07-20 08:33 +0200
      The Cache Identity Crisis by Micro Penis (Re: Gemini, DeepSeek, OpenAI all know Dmitriy V'jukov) Mild Shock <janburse@fastmail.fm> - 2026-07-20 13:54 +0200
        Just RTFM the RDNA 3.5 specs! [GPU Cache Lines] (Re: The Cache Identity Crisis by Micro Penis) Mild Shock <janburse@fastmail.fm> - 2026-07-20 14:10 +0200
          The large memory tax: ECC RAM (Re: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines]) Mild Shock <janburse@fastmail.fm> - 2026-07-20 14:23 +0200
          Re: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines] (Re: The Cache Identity Crisis by Micro Penis) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-20 12:26 -0700
            Re: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines] (Re: The Cache Identity Crisis by Micro Penis) R Kym Horsell <kymhorsell@gmail.com> - 2026-07-20 22:20 +0000
              Friendly Reminder: GPU 10x more performant than CPU (Was: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines]) Mild Shock <janburse@fastmail.fm> - 2026-07-21 00:39 +0200
                Breaking the CUDA edge in AI by WebGPU (Was: Friendly Reminder: GPU 10x more performant than CPU) Mild Shock <janburse@fastmail.fm> - 2026-07-21 00:56 +0200
                  Like WebAssembly before it, WebGPU has "escaped" the browser. (Was: Breaking the CUDA edge in AI by WebGPU) Mild Shock <janburse@fastmail.fm> - 2026-07-21 01:05 +0200
                    If you are paranoid you can use Falco [Agentic AI] (Re: Like WebAssembly before it, WebGPU has "escaped" the browser.) Mild Shock <janburse@fastmail.fm> - 2026-07-21 22:57 +0200
                      What would an EMACS guru say [Windows Recall] (Re: If you are paranoid you can use Falco [Agentic AI]) Mild Shock <janburse@fastmail.fm> - 2026-07-21 23:35 +0200
                        Decide what you critique tiny winy penis (Re: What would an EMACS guru say [Windows Recall]) Mild Shock <janburse@fastmail.fm> - 2026-07-22 08:17 +0200
                          How confused is tiny winy penis? (Was: Decide what you critique tiny winy penis) Mild Shock <janburse@fastmail.fm> - 2026-07-22 08:28 +0200
                            Maybe change your hobby, become a dog owner? (Was: How confused is tiny winy penis?) Mild Shock <janburse@fastmail.fm> - 2026-07-22 09:37 +0200
                          Even dogs know Switzerland != Germany [Syphilis Brain Micro Penis] (Re: Decide what you critique tiny winy penis ) Mild Shock <janburse@fastmail.fm> - 2026-07-22 11:19 +0200
                            node.js has also Worker isolation , headless (Was: Even dogs know Switzerland != Germany [Syphilis Brain Micro Penis]) Mild Shock <janburse@fastmail.fm> - 2026-07-22 11:29 +0200
                              Abyss hobby interpretation vs. professional understanding [Keep up with the Kardashians] (Re: node.js has also Worker isolation , headless) Mild Shock <janburse@fastmail.fm> - 2026-07-22 11:38 +0200
                                Play stupid games, win classic Usenet prizes [CCCP Troll] (Re: Abyss hobby interpretation vs. professional understanding [Keep up with the Kardashians]) Mild Shock <janburse@fastmail.fm> - 2026-07-22 12:16 +0200
                                  Glue Sniffing 5-Year Old Moron [CCCP Troll] (Was: Play stupid games, win classic Usenet prizes [CCCP Troll]) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:29 +0200
                            My Swift Go 16 AI has no IMEI, are you nuts? (Re: Even dogs know Switzerland != Germany [Syphilis Brain Micro Penis]) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:15 +0200
                              Same nickname and email, could post faster [5 year old moron] (Re: My Swift Go 16 AI has no IMEI, are you nuts? ) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:25 +0200
                              I have nothing to hide, you can find me in search.ch (Re: My Swift Go 16 AI has no IMEI, are you nuts?) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:41 +0200
                                Where did I confirm German via .ch, you are more than nuts! (Re: I have nothing to hide, you can find me in search.ch) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:48 +0200
                                  Ask a Ukrainian Neighbour to do Detective [CCCP Troll] (Re: Where did I confirm German via .ch, you are more than nuts! ) Mild Shock <janburse@fastmail.fm> - 2026-07-22 14:51 +0200
            Re: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines] (Re: The Cache Identity Crisis by Micro Penis) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-20 21:36 -0700
              rossy boy is going paranoid (Was: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines]) Mild Shock <janburse@fastmail.fm> - 2026-07-21 09:01 +0200
                Re: rossy boy is going paranoid (Was: Just RTFM the RDNA 3.5 specs! [GPU Cache Lines]) Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-07-21 00:56 -0700
        L1,..,Ln caches are located on the CPU AND on the GPU (Was: The Cache Identity Crisis by Micro Penis( Mild Shock <janburse@fastmail.fm> - 2026-07-20 19:35 +0200
          GPU Cache Hierarchy: Understanding L1, L2, and VRAM (Was: L1,..,Ln caches are located on the CPU AND on the GPU) Mild Shock <janburse@fastmail.fm> - 2026-07-20 19:40 +0200
            Well thats good, co-location, onto the same processor die (Re: GPU Cache Hierarchy: Understanding L1, L2, and VRAM (Was: L1,..,Ln caches are located on the CPU AND on the GPU) Mild Shock <janburse@fastmail.fm> - 2026-07-20 22:02 +0200
              Where is micro penis mental error? (Re: Well thats good, co-location, onto the same processor die (Re: GPU Cache Hierarchy: Understanding L1, L2, and VRAM (Was: L1,..,Ln caches are located on the CPU AND on the GPU) Mild Shock <janburse@fastmail.fm> - 2026-07-20 22:07 +0200
      I didn't find Futex in WebGPU / WGSL (Re: Gemini, DeepSeek, OpenAI all know Dmitriy V'jukov) Mild Shock <janburse@fastmail.fm> - 2026-07-20 23:26 +0200
        There is no imageAtomicAdd in WGSL (Re: I didn't find Futex in WebGPU / WGSL) Mild Shock <janburse@fastmail.fm> - 2026-07-21 00:15 +0200
          OpenGL is dead. Apple said bye bye / Wayland Compositor (Re: There is no imageAtomicAdd in WGSL) Mild Shock <janburse@fastmail.fm> - 2026-07-21 00:26 +0200
          Flogging a Dead Horse, OpenGL is EOL (Re: There is no imageAtomicAdd in WGSL) Mild Shock <janburse@fastmail.fm> - 2026-07-21 01:39 +0200
            imageAtomicAdd trivial, Dmitry Vyukov requires capacity (Re: Flogging a Dead Horse, OpenGL is EOL) Mild Shock <janburse@fastmail.fm> - 2026-07-21 01:40 +0200
              capacity = 2^n for some n / systolic system (Re: imageAtomicAdd trivial, Dmitry Vyukov requires capacity) Mild Shock <janburse@fastmail.fm> - 2026-07-21 01:41 +0200
                Source of the benchmark for DmitryVyukov (Was: capacity = 2^n for some n / systolic system) Mild Shock <janburse@fastmail.fm> - 2026-07-21 01:44 +0200
          Re: There is no imageAtomicAdd in WGSL (Re: I didn't find Futex in WebGPU / WGSL) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-20 16:59 -0700
            I never used OpenGL Version 4.2 and later (Was: There is no imageAtomicAdd in WGSL) Mild Shock <janburse@fastmail.fm> - 2026-07-21 08:47 +0200
              Because of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later) Mild Shock <janburse@fastmail.fm> - 2026-07-21 09:02 +0200
                Why MIMD is interesting for pi-WAM? Mild Shock <janburse@fastmail.fm> - 2026-07-21 09:18 +0200
                Re: Because of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-22 13:34 -0700
                  Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:09 +0200
                    Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms) Mild Shock <janburse@fastmail.fm> - 2026-07-23 00:53 +0200
                    Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-22 18:00 -0700
                      Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-07-22 18:20 -0700
                    Why do you even need a mpmc queue? [Thunder Kittens] (Was: Do you see the loops, in C code and in Java code? /** Looping **/) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:44 +0200
                      What does pi in pi-WAM mean? (Re: Why do you even need a mpmc queue? [Thunder Kittens]) Mild Shock <janburse@fastmail.fm> - 2026-07-24 14:53 +0200
                        OR-parallelism or AND-parallelism? [MapReduce] (Was: What does pi in pi-WAM mean?) Mild Shock <janburse@fastmail.fm> - 2026-07-24 15:00 +0200

Page 3 of 3 — ← Prev page 1 2 [3]


#896361 — I never used OpenGL Version 4.2 and later (Was: There is no imageAtomicAdd in WGSL)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-21 08:47 +0200
SubjectI never used OpenGL Version 4.2 and later (Was: There is no imageAtomicAdd in WGSL)
Message-ID<113n4m3$5oih$1@solani.org>
In reply to#896357
Hi,

I went to holdays in June 2026, had an idea for
a pi-WAM based on a Hack, the later is described here:

Emulating π-WAM in Dogelog Player
https://medium.com/2989/de9cd29c7d37

The Elements of Computing Systems
https://mitpress.mit.edu/9780262539807

In July 2026 I did the CPU and GPU experiments,
moving from emulator to native executor based in
realizing Hack as a concrete virtual machine,

and not as an abstract machine emulated in Prolog.
The GPU experiments were done in WebGPU / WGSL.
So no, I never used OpenGL Version 4.2 and later.

Also the name imageAtomicAdd indicates that it
imageAtomicAdd is rather from a render shader,
while my GPU experiment uses a compute shader.

Es specially I need GPU compute shaders, which
are not executed in lock step, but rather have
indepdendent thread state, also known as MIMD.

"In computing, multiple instruction, multiple
data (MIMD) is a technique employed to
achieve parallelism. "
https://en.wikipedia.org/wiki/Multiple_instruction,_multiple_data

MIMID showed up 2017 with NVIDIA Volta cards.
But is now realized by Intel Arc, Snapdragon Adreno
and AMD RDNA as well.

Bye

Chris M. Thomasson schrieb:
> On 7/20/2026 3:15 PM, Mild Shock wrote:
> [...]
> 
> Have you ever used imageAtomicAdd before?
> 

[toc] | [prev] | [next] | [standalone]


#896363 — Because of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-21 09:02 +0200
SubjectBecause of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later)
Message-ID<113n5j3$5ouq$3@solani.org>
In reply to#896361
Hi,

Because of MIMD you have to reassess algorithms.
A spin loop which could really hurt non-MIMD
GPUs, might less hurt a MIMD GPU.

Basically you have to reassess algorithms. Be
very exact whether your claims relates to
non-MIMD or to MIMD. You can toy around

with WebGL, mainly made for non-MIMD, here:

https://www.shadertoy.com/

and with WegGPU, mainly made for MIMD, here:

https://compute.toys/

The compute toys page, supports two shader
languages, WGSL and Slang. I made my GPU
experiments only with WGSL.

So I don't know Slang either. Things I
don't know in the GPU world are:

- OpenGL 4.2 and later
- Slang https://shader-slang.org/
- CUDA https://de.wikipedia.org/wiki/CUDA

Things I have meanwhile hands on, and which
I plan to integrate into library(edge/brainfog):

- WGSL https://webgpufundamentals.org/

Bye

Mild Shock schrieb:
> Hi,
> 
> I went to holdays in June 2026, had an idea for
> a pi-WAM based on a Hack, the later is described here:
> 
> Emulating π-WAM in Dogelog Player
> https://medium.com/2989/de9cd29c7d37
> 
> The Elements of Computing Systems
> https://mitpress.mit.edu/9780262539807
> 
> In July 2026 I did the CPU and GPU experiments,
> moving from emulator to native executor based in
> realizing Hack as a concrete virtual machine,
> 
> and not as an abstract machine emulated in Prolog.
> The GPU experiments were done in WebGPU / WGSL.
> So no, I never used OpenGL Version 4.2 and later.
> 
> Also the name imageAtomicAdd indicates that it
> imageAtomicAdd is rather from a render shader,
> while my GPU experiment uses a compute shader.
> 
> Es specially I need GPU compute shaders, which
> are not executed in lock step, but rather have
> indepdendent thread state, also known as MIMD.
> 
> "In computing, multiple instruction, multiple
> data (MIMD) is a technique employed to
> achieve parallelism. "
> https://en.wikipedia.org/wiki/Multiple_instruction,_multiple_data
> 
> MIMID showed up 2017 with NVIDIA Volta cards.
> But is now realized by Intel Arc, Snapdragon Adreno
> and AMD RDNA as well.
> 
> Bye
> 
> Chris M. Thomasson schrieb:
>> On 7/20/2026 3:15 PM, Mild Shock wrote:
>> [...]
>>
>> Have you ever used imageAtomicAdd before?
>>
> 

[toc] | [prev] | [next] | [standalone]


#896364 — Why MIMD is interesting for pi-WAM?

FromMild Shock <janburse@fastmail.fm>
Date2026-07-21 09:18 +0200
SubjectWhy MIMD is interesting for pi-WAM?
Message-ID<113n6go$5ple$9@solani.org>
In reply to#896363
Hi,

Since pi-WAM has two ancestors, namely pi for
pi-calculus and WAM for Warren Abstract Machine,
MIMD is especially interesting for pi-WAM .

To realize some pi-calculus fragment for example
Hoare Communicating Sequential Processes (CSP),
I will not use ADA rendez vous, but are planning

"In computer science, communicating sequential
processes (CSP) is a formal language for
describing patterns of interaction in concurrent systems

CSP was first described by Tony Hoare in a 1978 article
CSP has been practically applied in industry as a
tool for specifying and verifying the concurrent

aspects of a variety of different systems,
such as the T9000 Transputer"
https://en.wikipedia.org/wiki/Communicating_sequential_processes

to use queue. Especially bounded MPMC queues, queues
with a finite capacity that allow multiple producers
and multiple consumers. And here MIMD seems to be

brother in spirit. Just think of pi-WAM being a transputer:

"An important purpose of the feature is
to enable reliable use of programming models
such as producer-consumer within a warp"
https://stackoverflow.com/q/70987051

Will see! I do not expect Micro Penis, Ross Finlayson,
Kim Horsel, or Chris M. Thomasson be helpful in
any way. They rather represent the wall of ignorance

or misunderstanding that such a project as pi-WAM can
face, very naturally. So take my posts as Turing Tests,
to see how much brain USENETS morons have, it also

helps me doing my laboratory hygien, by doing
brainwriting. Although recently it gets a little
annoying, since I am meanwhile repeating for the

5-th time what I already wrote weeks ago.

Bye

[toc] | [prev] | [next] | [standalone]


#896397 — Re: Because of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later)

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-07-22 13:34 -0700
SubjectRe: Because of MIMD you have to reassess algorithms (Re: I never used OpenGL Version 4.2 and later)
Message-ID<113r9gn$396f7$4@dont-email.me>
In reply to#896363
On 7/21/2026 12:02 AM, Mild Shock wrote:
> Hi,
> 
> Because of MIMD you have to reassess algorithms.
> A spin loop which could really hurt non-MIMD
> GPUs, might less hurt a MIMD GPU.
[...]

Huh? atomic fetch-add has no looping. compare LOCK CMPXCHG with LOCK 
XADD on x86.

[toc] | [prev] | [next] | [standalone]


#896399 — Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-23 00:09 +0200
SubjectDo you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)
Message-ID<113rf37$846i$1@solani.org>
In reply to#896397
Hi,

CAS and XADD have no looping, they
are atomic operations, that take some
time but basically have some outcome

with some ACID property and a result
value. What loops is the ADT, the Abstract
Data Type that you implement. Respectively

the client that uses the Abstract Data Type.
In your case you added the loop inside the
Abstract Data Type or lower level aggregate

code of a higher level operation:

Chris M. Thomasson wrote:
void producer(double state) {
     uint32_t ver = XADD(&head, 1);
     cell& c = cells[ver & (N - 1)];
     while (LOAD(&c.ver) != ver) backoff(); /** Looping **/
     c.state = state;
     STORE(&c.ver, ver + 1);
}
https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/a6-Di0GZsyEJ

In my case I added the loop during the client
usage of the ADT:

From: Mild Shock <janburse@fastmail.fm>
Subject: Source of the benchmark for DmitryVyukov
Date: Tue, 21 Jul 2026 01:44:21 +0200

     private static void producer(Queue q) {
         for (int i = 0; i < WORK; i++) {
             Integer val = Integer.valueOf(i);
             while (!enqueue(q, val)) ; /** Looping **/
         }
     }

Do you see the two loops, in your C code
and in my Java code? They are marked with a
comment /** Looping **/ .

You see them, don't you? But I don't know
exactly what backoff() does. Sometimes loops
are spurious yield loops, required because

an ADT cannot gurantee that every yield
implies a certain condition. This is for
example already found in the intrinsinc

monitor of Java, the wait(). You might consult
Doug Lea about the matter and how idiomatic
Java code looks like dealing with

spurious yields.

Bye

Chris M. Thomasson schrieb:
> On 7/21/2026 12:02 AM, Mild Shock wrote:
>> Hi,
>>
>> Because of MIMD you have to reassess algorithms.
>> A spin loop which could really hurt non-MIMD
>> GPUs, might less hurt a MIMD GPU.
> [...]
> 
> Huh? atomic fetch-add has no looping. compare LOCK CMPXCHG with LOCK 
> XADD on x86.
> 

[toc] | [prev] | [next] | [standalone]


#896401 — Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-23 00:53 +0200
SubjectRe: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)
Message-ID<113rhkp$85a4$2@solani.org>
In reply to#896399
Hi,

Your father is regreting not using a
contraceptive. Now there is just one more
moron walking earth, and that moron

is you Lane W alias micro penis.

Bye

Lane W schrieb:
 > Mild Shock wrote:
 >> Hi,
 >>
 >> CAS and XADD have no looping, they
 >> are atomic operations, that take some
 >> time but basically have some outcome
 >
 > My father says C++ reminds him of an abortion.

Mild Shock schrieb:
> Hi,
> 
> CAS and XADD have no looping, they
> are atomic operations, that take some
> time but basically have some outcome
> 
> with some ACID property and a result
> value. What loops is the ADT, the Abstract
> Data Type that you implement. Respectively
> 
> the client that uses the Abstract Data Type.
> In your case you added the loop inside the
> Abstract Data Type or lower level aggregate
> 
> code of a higher level operation:
> 
> Chris M. Thomasson wrote:
> void producer(double state) {
>      uint32_t ver = XADD(&head, 1);
>      cell& c = cells[ver & (N - 1)];
>      while (LOAD(&c.ver) != ver) backoff(); /** Looping **/
>      c.state = state;
>      STORE(&c.ver, ver + 1);
> }
> https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/a6-Di0GZsyEJ
> 
> In my case I added the loop during the client
> usage of the ADT:
> 
> From: Mild Shock <janburse@fastmail.fm>
> Subject: Source of the benchmark for DmitryVyukov
> Date: Tue, 21 Jul 2026 01:44:21 +0200
> 
>      private static void producer(Queue q) {
>          for (int i = 0; i < WORK; i++) {
>              Integer val = Integer.valueOf(i);
>              while (!enqueue(q, val)) ; /** Looping **/
>          }
>      }
> 
> Do you see the two loops, in your C code
> and in my Java code? They are marked with a
> comment /** Looping **/ .
> 
> You see them, don't you? But I don't know
> exactly what backoff() does. Sometimes loops
> are spurious yield loops, required because
> 
> an ADT cannot gurantee that every yield
> implies a certain condition. This is for
> example already found in the intrinsinc
> 
> monitor of Java, the wait(). You might consult
> Doug Lea about the matter and how idiomatic
> Java code looks like dealing with
> 
> spurious yields.
> 
> Bye
> 
> Chris M. Thomasson schrieb:
>> On 7/21/2026 12:02 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> Because of MIMD you have to reassess algorithms.
>>> A spin loop which could really hurt non-MIMD
>>> GPUs, might less hurt a MIMD GPU.
>> [...]
>>
>> Huh? atomic fetch-add has no looping. compare LOCK CMPXCHG with LOCK 
>> XADD on x86.
>>
> 

[toc] | [prev] | [next] | [standalone]


#896405 — Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-07-22 18:00 -0700
SubjectRe: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)
Message-ID<113rp41$3e6dt$1@dont-email.me>
In reply to#896399
On 7/22/2026 3:09 PM, Mild Shock wrote:
> Hi,
> 
> CAS and XADD have no looping, they
> are atomic operations, that take some
> time but basically have some outcome
> 
> with some ACID property and a result
> value. What loops is the ADT, the Abstract
> Data Type that you implement. Respectively
> 
> the client that uses the Abstract Data Type.
> In your case you added the loop inside the
> Abstract Data Type or lower level aggregate
> 
> code of a higher level operation:
> 
> Chris M. Thomasson wrote:
> void producer(double state) {
>      uint32_t ver = XADD(&head, 1);
>      cell& c = cells[ver & (N - 1)];
>      while (LOAD(&c.ver) != ver) backoff(); /** Looping **/
>      c.state = state;
>      STORE(&c.ver, ver + 1);
> }
> https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/a6-Di0GZsyEJ

This was meant to be used in either a backoff, or ideally a futex, or 
return the bakery ticket to the user in the need to wait case. Was never 
meant to be used in a GPU. Dmitry CAS version can be used, but its still 
going to spin on any failed CAS. But, they are very different ways to do 
the same thing and they can be mixed and matched. But on the GPU, try to 
avoid CAS or anything that has to wait.

My XADD version beats Dmitry's in some work loads, but the fact it that 
they can be mixed and matched.

_____________
[from me]
 >> > I think it must be possible to combine e.g. CAS-based consumers and
 >> > XADD-based producers. This can be beneficial if you expect that
 >> > producers usually do not blocks (so you care more about fast-path
 >> > performance rather than blocking behavior).

 >> AFAICT, my tweak should be 100% compatible with the
 >> existing algorithm as-is. IMVHO, it could be a fairly
 >> beneficial addition to the existing API. It opens
 >> up a new way to think about handling contention wrt
 >> using the queue as a whole.

[from my friend]
Good point.
I guess single XADD for MPMC if far superior than most algorithms out
there, and is basically the best one can get for a centralized queue.
I see people still implement MS node-based queue with PDR, and that's
2 CAS loops + indirections + memory allocations + PDR acquire/release
overheads.... veeeeeery slow :)
_______________


 From this post:

https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/q9j14HPbO3sJ

You need to be careful wrt just using a multi-threaded lock-free algo in 
a compute shader. Some are just not made for it.

Why do you even need a mpmc queue in your compute shader anyway?



> 
> In my case I added the loop during the client
> usage of the ADT:
> 
> From: Mild Shock <janburse@fastmail.fm>
> Subject: Source of the benchmark for DmitryVyukov
> Date: Tue, 21 Jul 2026 01:44:21 +0200
> 
>      private static void producer(Queue q) {
>          for (int i = 0; i < WORK; i++) {
>              Integer val = Integer.valueOf(i);
>              while (!enqueue(q, val)) ; /** Looping **/
>          }
>      }
> 
> Do you see the two loops, in your C code
> and in my Java code? They are marked with a
> comment /** Looping **/ .
> 
> You see them, don't you? But I don't know
> exactly what backoff() does. Sometimes loops
> are spurious yield loops, required because
> 
> an ADT cannot gurantee that every yield
> implies a certain condition. This is for
> example already found in the intrinsinc
> 
> monitor of Java, the wait(). You might consult
> Doug Lea about the matter and how idiomatic
> Java code looks like dealing with
> 
> spurious yields.
> 
> Bye
> 
> Chris M. Thomasson schrieb:
>> On 7/21/2026 12:02 AM, Mild Shock wrote:
>>> Hi,
>>>
>>> Because of MIMD you have to reassess algorithms.
>>> A spin loop which could really hurt non-MIMD
>>> GPUs, might less hurt a MIMD GPU.
>> [...]
>>
>> Huh? atomic fetch-add has no looping. compare LOCK CMPXCHG with LOCK 
>> XADD on x86.
>>
> 

[toc] | [prev] | [next] | [standalone]


#896406 — Re: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)

From"Chris M. Thomasson" <chris.m.thomasson.1@gmail.com>
Date2026-07-22 18:20 -0700
SubjectRe: Do you see the loops, in C code and in Java code? /** Looping **/ (Was: Because of MIMD you have to reassess algorithms)
Message-ID<113rq9e$3egn4$1@dont-email.me>
In reply to#896405
On 7/22/2026 6:00 PM, Chris M. Thomasson wrote:
> On 7/22/2026 3:09 PM, Mild Shock wrote:
>> Hi,
>>
>> CAS and XADD have no looping, they
>> are atomic operations, that take some
>> time but basically have some outcome
>>
>> with some ACID property and a result
>> value. What loops is the ADT, the Abstract
>> Data Type that you implement. Respectively
>>
>> the client that uses the Abstract Data Type.
>> In your case you added the loop inside the
>> Abstract Data Type or lower level aggregate
>>
>> code of a higher level operation:
>>
>> Chris M. Thomasson wrote:
>> void producer(double state) {
>>      uint32_t ver = XADD(&head, 1);
>>      cell& c = cells[ver & (N - 1)];
>>      while (LOAD(&c.ver) != ver) backoff(); /** Looping **/
>>      c.state = state;
>>      STORE(&c.ver, ver + 1);
>> }
>> https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/a6-Di0GZsyEJ
> 
> This was meant to be used in either a backoff, or ideally a futex, or 
> return the bakery ticket to the user in the need to wait case. Was never 
> meant to be used in a GPU. Dmitry CAS version can be used, but its still 
> going to spin on any failed CAS. But, they are very different ways to do 
> the same thing and they can be mixed and matched. But on the GPU, try to 
> avoid CAS or anything that has to wait.
> 
> My XADD version beats Dmitry's in some work loads, but the fact it that 
> they can be mixed and matched.
> 
> _____________
> [from me]
>  >> > I think it must be possible to combine e.g. CAS-based consumers and
>  >> > XADD-based producers. This can be beneficial if you expect that
>  >> > producers usually do not blocks (so you care more about fast-path
>  >> > performance rather than blocking behavior).
> 
>  >> AFAICT, my tweak should be 100% compatible with the
>  >> existing algorithm as-is. IMVHO, it could be a fairly
>  >> beneficial addition to the existing API. It opens
>  >> up a new way to think about handling contention wrt
>  >> using the queue as a whole.
> 
> [from my friend]
> Good point.
> I guess single XADD for MPMC if far superior than most algorithms out
> there, and is basically the best one can get for a centralized queue.
> I see people still implement MS node-based queue with PDR, and that's
> 2 CAS loops + indirections + memory allocations + PDR acquire/release
> overheads.... veeeeeery slow :)
> _______________
> 
> 
>  From this post:
> 
> https://groups.google.com/g/lock-free/c/acjQ3-89abE/m/q9j14HPbO3sJ
> 
> You need to be careful wrt just using a multi-threaded lock-free algo in 
> a compute shader. Some are just not made for it.
> 
> Why do you even need a mpmc queue in your compute shader anyway?
> 
> 
> 
>>
>> In my case I added the loop during the client
>> usage of the ADT:
>>
>> From: Mild Shock <janburse@fastmail.fm>
>> Subject: Source of the benchmark for DmitryVyukov
>> Date: Tue, 21 Jul 2026 01:44:21 +0200
>>
>>      private static void producer(Queue q) {
>>          for (int i = 0; i < WORK; i++) {
>>              Integer val = Integer.valueOf(i);
>>              while (!enqueue(q, val)) ; /** Looping **/
>>          }
>>      }
>>
>> Do you see the two loops, in your C code
>> and in my Java code? They are marked with a
>> comment /** Looping **/ .
>>
>> You see them, don't you? But I don't know
>> exactly what backoff() does. Sometimes loops
>> are spurious yield loops, required because[...]

They are bakery algorithms on a bounded buffer. My tweak gets away 
without having to use any CAS. If you try to dequeue something and the 
queue is empty, you either have to wait for an enqueue, or do something 
else. My xadd version is pure ticket based, so a a wait condition can 
use futex, backoff, of let the user use that ticker for whatever it 
wants to.

[toc] | [prev] | [next] | [standalone]


#896423 — Why do you even need a mpmc queue? [Thunder Kittens] (Was: Do you see the loops, in C code and in Java code? /** Looping **/)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-24 14:44 +0200
SubjectWhy do you even need a mpmc queue? [Thunder Kittens] (Was: Do you see the loops, in C code and in Java code? /** Looping **/)
Message-ID<113vmnr$bed5$4@solani.org>
In reply to#896399
Hi,

You don't pay attention, right! I am little
bit disappointed that your attention span is
near zero. I already posted:

From: Mild Shock <janburse@fastmail.fm>
Subject: Why do you even need a mpmc queue? [Thunder Kittens]
Date: Thu, 23 Jul 2026 08:43:03 +0200

 > Hi,
 >
 > Because I use WebGPU and not WebGL. And
 > because WebGPU can adresss modern GPU
 > developed with the NVIDIA Volta evolution,
 >
 > which happened in 2017. Namley that compute
 > shaders are not any more subject to the
 > realization restriction of lock step
 >
 > execution, but have independent thread state.
 > And because there is independent thread state
 > there is also independent time spent for a
 >
 > a work item by each logical thread, if the
 > submitted logical thread uses a lot of branching
 > logic or even loops. But the use of branching
 >
 > and loops is encouraged in independent thread
 > state programming of compute shaders. The variables
 > that can drive such logic are the scalar variables:
 >
 > Tour of WGSL - Control Flow
 > https://google.github.io/tour-of-wgsl/control-flow/
 >
 > Then not to waste GPU compute time, by logical
 > threads doing nothing. You will need to
 > introduce some load balancing among multiple
 >
 > logical threads. And MPMC queues are one way to
 > readize load balancing. Compute shaders with
 > producer and consumer entry points are proposed
 >
 > as fundamental architecture by Thunder Kittens:
 >
 > ThunderKittens: Simple, Fast, and Adorable AI Kernels
 > https://arxiv.org/abs/2410.20399
 >
 > They are used by this SpaceX acquisition:
 >
 > Composer 2 Technical Report
 > https://arxiv.org/abs/2603.24477
 >
 > Thunder Kittens uses Hardware support, i.e. tma_expect().
 >
 > Bye
 >
 > Chris M. Thomasson schrieb:
 >>> never meant to be used in a GPU.
 >>> Dmitry CAS version can be used, but
 >>>
 >>> Why do you even need a mpmc queue
 >>> in your compute shader anyway?

Chris M. Thomasson schrieb:
 > I don't think he knows exactly what he is doing...
 > Why does he need a lock/wait-free queue in a
 > compute shader? What is he trying to do?

[toc] | [prev] | [next] | [standalone]


#896424 — What does pi in pi-WAM mean? (Re: Why do you even need a mpmc queue? [Thunder Kittens])

FromMild Shock <janburse@fastmail.fm>
Date2026-07-24 14:53 +0200
SubjectWhat does pi in pi-WAM mean? (Re: Why do you even need a mpmc queue? [Thunder Kittens])
Message-ID<113vn8h$beun$2@solani.org>
In reply to#896423
Hi,

You can also deduce that I need comms,
from pi in pi-WAM, since pi refers to pi-calculus.
There is also a nice paper, that I have already posted:

A pi-calculus Specification of Prolog
https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf

You asked yourself why no atomic and
only comms? So its as simple as 1+1=2.
But usenet people are usually slow as fuck.

Take your time. You could spin loop to ingest
the topic, i.e. try again in 3-4 months, for
example reading some of the paper. Although

I know thats a totally unrealistic request, asking
a troll to do RTFM and study something. They
rather make themselves a total laughing stock,

play stupid games, win usenet prizes.

Bye

Mild Shock schrieb:
> Hi,
> 
> You don't pay attention, right! I am little
> bit disappointed that your attention span is
> near zero. I already posted:
> 
> From: Mild Shock <janburse@fastmail.fm>
> Subject: Why do you even need a mpmc queue? [Thunder Kittens]
> Date: Thu, 23 Jul 2026 08:43:03 +0200
> 
>  > Hi,
>  >
>  > Because I use WebGPU and not WebGL. And
>  > because WebGPU can adresss modern GPU
>  > developed with the NVIDIA Volta evolution,
>  >
>  > which happened in 2017. Namley that compute
>  > shaders are not any more subject to the
>  > realization restriction of lock step
>  >
>  > execution, but have independent thread state.
>  > And because there is independent thread state
>  > there is also independent time spent for a
>  >
>  > a work item by each logical thread, if the
>  > submitted logical thread uses a lot of branching
>  > logic or even loops. But the use of branching
>  >
>  > and loops is encouraged in independent thread
>  > state programming of compute shaders. The variables
>  > that can drive such logic are the scalar variables:
>  >
>  > Tour of WGSL - Control Flow
>  > https://google.github.io/tour-of-wgsl/control-flow/
>  >
>  > Then not to waste GPU compute time, by logical
>  > threads doing nothing. You will need to
>  > introduce some load balancing among multiple
>  >
>  > logical threads. And MPMC queues are one way to
>  > readize load balancing. Compute shaders with
>  > producer and consumer entry points are proposed
>  >
>  > as fundamental architecture by Thunder Kittens:
>  >
>  > ThunderKittens: Simple, Fast, and Adorable AI Kernels
>  > https://arxiv.org/abs/2410.20399
>  >
>  > They are used by this SpaceX acquisition:
>  >
>  > Composer 2 Technical Report
>  > https://arxiv.org/abs/2603.24477
>  >
>  > Thunder Kittens uses Hardware support, i.e. tma_expect().
>  >
>  > Bye
>  >
>  > Chris M. Thomasson schrieb:
>  >>> never meant to be used in a GPU.
>  >>> Dmitry CAS version can be used, but
>  >>>
>  >>> Why do you even need a mpmc queue
>  >>> in your compute shader anyway?
> 
> Chris M. Thomasson schrieb:
>  > I don't think he knows exactly what he is doing...
>  > Why does he need a lock/wait-free queue in a
>  > compute shader? What is he trying to do?
> 

[toc] | [prev] | [next] | [standalone]


#896425 — OR-parallelism or AND-parallelism? [MapReduce] (Was: What does pi in pi-WAM mean?)

FromMild Shock <janburse@fastmail.fm>
Date2026-07-24 15:00 +0200
SubjectOR-parallelism or AND-parallelism? [MapReduce] (Was: What does pi in pi-WAM mean?)
Message-ID<113vnlq$bf89$1@solani.org>
In reply to#896424
Hi,

Via concurrent logic programming, and the
The Fifth Generation Computer Systems
10-year initiative launched in 1982 by Japan:

Fifth Generation Computer Systems (FGCS)
https://en.wikipedia.org/wiki/Fifth_Generation_Computer_Systems

There is a big fundus of material about
relating logic progra execution to parallel
execution. Unfortunetly the topic somehow

lost steam, and has been reduced to Remote
Procedure Call (RPC) of iterators inside the
Web Prolog initiative. But its evident that

this leads to nowhere, on modern CPU and GPU,
since RPC adds an additional comms aka message,
already to only communicate success or failure.

The other comms aka messages approach stems from
implementing parallel query executors for
relational databases and uses messages for

other things. It had already revival in 2008:

Google spotlights data center inner workings
http://news.cnet.com/8301-10784_3-9955184-7.html

Just google MapReduce!

Bye

Mild Shock schrieb:
> Hi,
> 
> You can also deduce that I need comms,
> from pi in pi-WAM, since pi refers to pi-calculus.
> There is also a nice paper, that I have already posted:
> 
> A pi-calculus Specification of Prolog
> https://scispace.com/pdf/a-pi-calculus-specification-of-prolog-3qf2pf04ud.pdf 
> 
> 
> You asked yourself why no atomic and
> only comms? So its as simple as 1+1=2.
> But usenet people are usually slow as fuck.
> 
> Take your time. You could spin loop to ingest
> the topic, i.e. try again in 3-4 months, for
> example reading some of the paper. Although
> 
> I know thats a totally unrealistic request, asking
> a troll to do RTFM and study something. They
> rather make themselves a total laughing stock,
> 
> play stupid games, win usenet prizes.
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> You don't pay attention, right! I am little
>> bit disappointed that your attention span is
>> near zero. I already posted:
>>
>> From: Mild Shock <janburse@fastmail.fm>
>> Subject: Why do you even need a mpmc queue? [Thunder Kittens]
>> Date: Thu, 23 Jul 2026 08:43:03 +0200
>>
>>  > Hi,
>>  >
>>  > Because I use WebGPU and not WebGL. And
>>  > because WebGPU can adresss modern GPU
>>  > developed with the NVIDIA Volta evolution,
>>  >
>>  > which happened in 2017. Namley that compute
>>  > shaders are not any more subject to the
>>  > realization restriction of lock step
>>  >
>>  > execution, but have independent thread state.
>>  > And because there is independent thread state
>>  > there is also independent time spent for a
>>  >
>>  > a work item by each logical thread, if the
>>  > submitted logical thread uses a lot of branching
>>  > logic or even loops. But the use of branching
>>  >
>>  > and loops is encouraged in independent thread
>>  > state programming of compute shaders. The variables
>>  > that can drive such logic are the scalar variables:
>>  >
>>  > Tour of WGSL - Control Flow
>>  > https://google.github.io/tour-of-wgsl/control-flow/
>>  >
>>  > Then not to waste GPU compute time, by logical
>>  > threads doing nothing. You will need to
>>  > introduce some load balancing among multiple
>>  >
>>  > logical threads. And MPMC queues are one way to
>>  > readize load balancing. Compute shaders with
>>  > producer and consumer entry points are proposed
>>  >
>>  > as fundamental architecture by Thunder Kittens:
>>  >
>>  > ThunderKittens: Simple, Fast, and Adorable AI Kernels
>>  > https://arxiv.org/abs/2410.20399
>>  >
>>  > They are used by this SpaceX acquisition:
>>  >
>>  > Composer 2 Technical Report
>>  > https://arxiv.org/abs/2603.24477
>>  >
>>  > Thunder Kittens uses Hardware support, i.e. tma_expect().
>>  >
>>  > Bye
>>  >
>>  > Chris M. Thomasson schrieb:
>>  >>> never meant to be used in a GPU.
>>  >>> Dmitry CAS version can be used, but
>>  >>>
>>  >>> Why do you even need a mpmc queue
>>  >>> in your compute shader anyway?
>>
>> Chris M. Thomasson schrieb:
>>  > I don't think he knows exactly what he is doing...
>>  > Why does he need a lock/wait-free queue in a
>>  > compute shader? What is he trying to do?
>>
> 

[toc] | [prev] | [standalone]


Page 3 of 3 — ← Prev page 1 2 [3]

Back to top | Article view | sci.physics


csiph-web