Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.lisp > #61377

The Mac Neo is a Budget Monster [GPU Channels] (Re: What are Flits and Phits? [Network on a Chip])

From Mild Shock <janburse@fastmail.fm>
Newsgroups comp.lang.c, comp.lang.lisp, comp.theory
Subject The Mac Neo is a Budget Monster [GPU Channels] (Re: What are Flits and Phits? [Network on a Chip])
Date 2026-08-11 16:27 +0200
Message-ID <115fbgm$9vhm$3@solani.org> (permalink)
References (12 earlier) <5T_bS.71191$4Fu9.56234@fx05.ams4> <114qei6$t1uc$1@solani.org> <114qerq$t244$1@solani.org> <1154jip$296j$3@solani.org> <1154vsi$2mto$3@solani.org>

Cross-posted to 3 groups.

Show all headers | View raw


Hi,

Now I implemented some multiple producer
and multiple consumer channel objects for
WebGPU. The only API to integrate it user

facing into pi-WAM is this single predicate:

/**
  * flit(C):
  * The predicate succeeds in C with a new channel. The channel
  * can be used from within GPU backed π-WAM logical threads.
  */

The Mac Neo is a Budget Monster. While the
Ryzen AI Laptop cost around 1300.- CHF.
The Mac Neo was around 600.- CHF with all

extras. Here some performance results,
checking out whether channel objects scale,
when increasing their number to

communicate the same 1 millon packets:

Java performance:

AI Laptop    Single    Double
Ryzen    705.1    337.4
Neo    669.4    239.9

WebGPU performance:

AI Laptop    Single    Double
Ryzen    731.8    392.9
Neo    932.8    483.5

Cool! Java is also pretty cool, their
semaphore library is top notch. I couldn't
replicate the resulst with JavaScript yet,

seems their Atomics.wait() resp. Atomics.waitAsync()
is totally broken, using futex is mutex for
fools somehow. I also found some gremlins

attacking one of the GPUs. The Intel AI Laptop
fails the above experiment. Maybe its a driver
Vulkan versus OpenCL or something problem,

or the Lunar lake architecture is nonsense.

Bye

Mild Shock schrieb:
> Hi,
> 
> Recently there was a paper somebody mentioning
> a flit doing a ACK or NACK, to express
> backpressure inside a Network on a Chip.
> 
> But what is a flit? It seems multiple
> flits can be used to create the message
> passing in one directiob before the
> 
> ACK or NACK in the other direction?
> 
> "The growing need for performance from
> computing systems drove the industry into
> the multi-core and many-core arena. In this
> setup, the execution of a kernel (a program)
> is split across multiple processors and the
> computation happens in parallel
> 
> Flits represent logical units of information,
> while phits represent the physical domain,
> that is, phits represent the number of bits
> that can be transferred in parallel in a
> single cycle. Consider the Cray T3D. It has
> an interconnection network which uses
> 
> flit level message flow control wherein each
> flit is composed of eight 16-bit phits. That
> means its flit size is 128bits and phit size
> is 16bits. Also consider the IBM SP2 switch.
> It also uses the flit level message flow
> control, but its flit size is equal to its
> phit size, which is set to 8 bits."
> https://en.wikipedia.org/wiki/Flit_(computer_networking)#Example
> 
> Well my idea how this is realized in silicon
> is rather foggy, I mean even the Hack project
> from Nand 2 Tetris, does not show some gate level
> schemes for flits and phits.
> 
> Could be an interesting extension. But somehow
> the image of flits and phits inspired my channel
> objects here below. But I am afraid they are fire
> and forget, no ACK and NACK:
> 
> π-WAM Contest: 1 Million Packets with Prolog
> https://medium.com/2989/ec3e91551773
> 
> Its amazing that a max_size(1) buffer
> can beat an unbounded buffer!
> 
> LoL
> 
> Bye
> 
> Mild Shock schrieb:
>> Hi,
>>
>> How it started, NVIDIA being cool:
>>
>> NCCL provides routines such as all-gather,
>> all-reduce, broadcast, reduce, reduce-scatter,
>> and point-to-point send and receive. These
>> routines are optimized to achieve high
>> bandwidth and low latency over PCIe,
>> NVIDIA NVLink™, and other high-speed
>> interconnects within a node and over
>> NVIDIA networking across nodes.
>> https://developer.nvidia.com/nccl
>>
>> How its going, vLLM trying to be cool:
>>
>> [RFC]: Native Weight Syncing APIs
>> However, there are no standardized methods for
>> performing online weight syncing. Open source projects
>> like SkyRL, VeRL, and TRL need to include their
>> own implementations of the weight syncing
>> infrastructure, leading to added complexity
>> for developers seeking to adopt vLLM as their
>> inference server for post-training workloads.
>> https://github.com/vllm-project/vllm/issues/31848
>>
>> How much Workers are enough? I guess it depends
>> on I/O parallelism, CPU Memory parallelism, CPU
>> Processing parallelism, and now also
>>
>> GPU Memory parallelism and GPU Processing
>> parallelism, and last but least you might have
>> a couple DMAs sitting here and there,
>>
>> or even invoking a sort of RDMA. Quite amazing!
>>
>> Bye
>>
>> Mild Shock schrieb:
>>> Hi,
>>>
>>> Well there are two viewpoint, the "client"
>>> of the GPU, which is the CPU, and the "server"
>>> of the GPU, which is the command processor
>>>
>>> queue of the GPU device. So basically as
>>> a CPU client I can write the memory area,
>>> that is later mapped to my GPU code storage.
>>>
>>> And this way have a compiler, even written
>>> in Prolog, that compiles pi-WAM to my Hack VM,
>>> that can then be then deployed to GPU.
>>>
>>> You could also try the same with a Tiny
>>> LISP VM. And a grown up LISP to act as
>>> the compiler. Would be a similar exercise.
>>>
>>> Have Fun!
>>>
>>> Bye
>>>
>>> Mild Shock schrieb:
>>>> Hi,
>>>>
>>>>  > I've also added comp.theory so Mild Shock can comment.
>>>>
>>>>  > that has different
>>>>  > *  sizeof ( void * ), and
>>>>  > *  sizeof ( void (*)( void ) ),
>>>>
>>>> Could indicate a data RAM and code ROM model.
>>>> Which has then the advantage of:
>>>>
>>>> Modern operating systems like Windows 11
>>>> enforce strict Data Execution Prevention (DEP)
>>>> (or NX/XD bit security features) to prevent
>>>> malicious programs from injecting and executing
>>>> code inside data-only memory regions.
>>>> https://root-nation.com/en/soft-en/lifehacks/en-dep-windows-all-about/
>>>>
>>>> I adopted data RAM and code ROM model for
>>>> pi-WAM from Hack, which has the same separation:
>>>>
>>>> Slide 58, Hack Computer
>>>> https://drive.google.com/file/d/1Z_fxYmmRNXTkAzmZ6YMoX9NXZIRVCKiw/view
>>>>
>>>> But my motivation was not Johnny Depp prevention.
>>>> Rather the caching of GPUs. Because WGSL
>>>> allows storage annotations read_write and
>>>>
>>>> read. I use read_write for the data RAM
>>>> of my Hack VM variant, and read for the
>>>> code ROM of my Hack VM variant. You can
>>>>
>>>> see that here, its open source:
>>>>
>>>> @group(0) @binding(0) var<storage, read> code: array<i32>;
>>>> @group(0) @binding(1) var<storage, read_write> state: array<i32>;
>>>>
>>>> 11.4 Giga Lips with a Budget Laptop
>>>> https://github.com/Jean-Luc-Picard-2021/gigabudget
>>>>
>>>> Hope this Helps!
>>>>
>>>> Bye
>>>>
>>>> Johann 'Myrkraverk' Oskarsson schrieb:
>>>>> On 03/08/2026 6:28 PM, David Brown wrote:
>>>>>> On 03/08/2026 11:41, Richard Harnden wrote:
>>>>>>> On 03/08/2026 09:16, David Brown wrote:
>>>>>>
>>>>>>>> On most targets, function pointers are the same size as void* 
>>>>>>>> pointers. But there are exceptions, with some small 
>>>>>>>> microcontrollers and DSPs having different kinds of pointers 
>>>>>>>> with different sizes, depending on the memory space involved.  I 
>>>>>>>> have yet to see a situation where there was any reason for 
>>>>>>>> storing a function address in a "void*" rather than a more 
>>>>>>>> appropriate typedef, such as :
>>>>>>>>
>>>>>>>>      typedef void (*FVoid)(void);
>>>>>>>
>>>>>>> dlsym requires that pointer-to-function is compatible with a void*
>>>>>>>
>>>>>>
>>>>>> As I say, I have yet to see a situation where using void* for 
>>>>>> function pointers was more appropriate than using a function 
>>>>>> pointer type.  If the OS system calls or standard OS libraries 
>>>>>> makes it a requirement that function pointers are converted to or 
>>>>>> from void* for some calls, then of course you need to follow those 
>>>>>> requirements - it's the people who designed the interfaces that 
>>>>>> made questionable design choices.
>>>>>>
>>>>>
>>>>> Nope, you're wrong.  You're dead wrong.  The world isn't built on C,
>>>>> even though here in comp.lang.c we like to pretend it is.
>>>>>
>>>>> Several language environments allow function generation on the fly,
>>>>> these functions need to be garbage collected.  Common Lisp is an
>>>>> example, therefore comp.lang.lisp is added to this discussion.
>>>>>
>>>>> I've also added comp.theory so Mild Shock can comment.
>>>>>
>>>>> You will have to go out of your way to make a computer architecture
>>>>> incompatible with garbage collected and heap allocated binary code,
>>>>> something I've been told SBCL does internally [1] to create an archi-
>>>>> tecture that has different
>>>>>
>>>>> *  sizeof ( void * ), and
>>>>> *  sizeof ( void (*)( void ) ),
>>>>>
>>>>> and when you do that, I'll just claim you're making a /malicious
>>>>> computer architecture/ and refuse to use it.
>>>>>
>>>>>
>>>>> [1] I've not looked at the code, but told the garbage collector can
>>>>> and will at least move the code around, if not collect it.
>>>>
>>>
>>
> 

Back to comp.lang.lisp | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

Malicious Computer Architecture (was: Re: Prioritize Performance over Correctness) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-03 19:23 +0800
  Re: Malicious Computer Architecture (was: Re: Prioritize Performance over Correctness) scott@slp53.sl.home (Scott Lurndal) - 2026-08-03 14:29 +0000
  Johnny Depp prevention [Windows 11 etc...] (Was: Malicious Computer Architecture) Mild Shock <janburse@fastmail.fm> - 2026-08-03 18:10 +0200
    But how can you deploy. when its ROM? (Was: Johnny Depp prevention [Windows 11 etc...]) Mild Shock <janburse@fastmail.fm> - 2026-08-03 18:15 +0200
      GPU Elasticity: Collective Communications Libraries (Was: But how can you deploy. when its ROM?) Mild Shock <janburse@fastmail.fm> - 2026-08-07 14:37 +0200
        What are Flits and Phits? [Network on a Chip] (Re: GPU Elasticity: Collective Communications Libraries) Mild Shock <janburse@fastmail.fm> - 2026-08-07 18:07 +0200
          Cristallina: Thank you for the Beam (Re: What are Flits and Phits? [Network on a Chip]) Mild Shock <janburse@fastmail.fm> - 2026-08-09 21:22 +0200
          The Mac Neo is a Budget Monster [GPU Channels] (Re: What are Flits and Phits? [Network on a Chip]) Mild Shock <janburse@fastmail.fm> - 2026-08-11 16:27 +0200
            The luminaries of duct-tape engineering [Sweeney and Torvald] (Was: The Mac Neo is a Budget Monster [GPU Channels]) Mild Shock <janburse@fastmail.fm> - 2026-08-11 16:51 +0200
  Re: Malicious Computer Architecture Aidan Kehoe <kehoea@parhasard.net> - 2026-08-03 19:51 +0100
    Re: Malicious Computer Architecture Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 00:22 +0800
      Re: Malicious Computer Architecture Richard Harnden <richard.nospam@gmail.invalid> - 2026-08-04 19:10 +0100
      Re: Malicious Computer Architecture Kaz Kylheku <046-301-5902@kylheku.com> - 2026-08-06 21:55 +0000
        Re: Malicious Computer Architecture Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 06:31 +0800
      Re: Malicious Computer Architecture Aidan Kehoe <kehoea@parhasard.net> - 2026-08-06 23:01 +0100
        Re: Malicious Computer Architecture Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-07 06:45 +0800
    Re: Malicious Computer Architecture Niocláisín Cóilín de Ghlostéir <thanks-to@Taf.com> - 2026-08-04 18:13 +0000
      Re: Malicious Computer Architecture Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 03:48 +0800
        Re: Malicious Computer Architecture Niocláisín Cóilín de Ghlostéir <thanks-to@Taf.com> - 2026-08-04 22:01 +0000
          Re: Malicious Computer Architecture Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 06:25 +0800
  A Case for Impurity: The Applied Pi Calculus (Re: Malicious Computer Architecture) Mild Shock <janburse@fastmail.fm> - 2026-08-04 14:58 +0200
    The Sandcastle of Paul Taraus Interactors (Re: A Case for Impurity: The Applied Pi Calculus) Mild Shock <janburse@fastmail.fm> - 2026-08-04 15:10 +0200

csiph-web