Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch > #111515 > unrolled thread

DMA is obsolete

Started byJohn Levine <johnl@taugh.com>
First post2025-04-26 16:19 +0000
Last post2025-05-04 02:10 +0000
Articles 20 on this page of 42 — 14 participants

Back to article view | Back to comp.arch


Contents

  DMA is obsolete John Levine <johnl@taugh.com> - 2025-04-26 16:19 +0000
    Re: DMA is obsolete Lars Poulsen <lars@cleo.beagle-ears.com> - 2025-04-26 16:28 +0000
      Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-04-26 19:28 +0200
      Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 19:35 +0100
        Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:49 +0000
          Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:37 +0000
        Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-04-28 01:20 +0000
    Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-26 17:29 +0000
      Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-26 19:25 +0000
        Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:01 +0000
          Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 16:12 +0000
        Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:02 +0000
        Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 20:13 +0100
          Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:45 +0000
            Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:44 +0000
        Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-01 13:07 +0000
          Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-01 22:03 +0000
            Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 02:15 +0000
              Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-02 05:34 +0000
                Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 15:02 +0000
                  Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-03 06:11 +0000
                    Re: DMA is obsolete Robert Finch <robfi680@gmail.com> - 2025-05-03 06:32 -0400
                    Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 13:33 +0000
                      IP (was: DMA is obsolete) Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 10:50 -0400
                        Re: IP (was: DMA is obsolete) Thomas Koenig <tkoenig@netcologne.de> - 2025-05-03 15:15 +0000
                          Re: IP (was: DMA is obsolete) John Levine <johnl@taugh.com> - 2025-05-03 15:46 +0000
                            Re: IP (was: DMA is obsolete) cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 16:52 +0000
                              Re: IP (was: DMA is obsolete) scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:31 +0000
                              Re: IP Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 23:04 -0400
                                Re: IP cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-04 09:56 +0000
                                  Re: IP Thomas Koenig <tkoenig@netcologne.de> - 2025-05-04 10:17 +0000
                                    Re: IP mitchalsup@aol.com (MitchAlsup1) - 2025-05-04 18:16 +0000
                                      Re: IP Bill Findlay <findlaybill@blueyonder.co.uk> - 2025-05-04 19:37 +0100
                                    Re: IP Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 21:31 +0000
                    Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 06:44 +0000
                  Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:53 +0000
                    Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-03 23:02 +0000
                    Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-21 12:36 +0000
              Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-02 17:40 +0000
                Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-05-03 14:29 +0200
                  ND-10 (was Re: DMA is obsolete) Lars Poulsen <lars@beagle-ears.com> - 2025-05-03 23:30 +0000
                    Re: ND-10 (was Re: DMA is obsolete) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 02:10 +0000

Page 1 of 3  [1] 2 3  Next page →


#111515 — DMA is obsolete

FromJohn Levine <johnl@taugh.com>
Date2025-04-26 16:19 +0000
SubjectDMA is obsolete
Message-ID<vuj131$fnu$1@gal.iecc.com>
Well, not entirely.  This preprint argues that in environments with
lots of cores and where latency is an issue, programmed I/O can outperform
DMA.

Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects

Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe

Conventional wisdom holds that an efficient interface between an OS
running on a CPU and a high-bandwidth I/O device should use Direct
Memory Access (DMA) to offload data transfer, descriptor rings for
buffering and queuing, and interrupts for asynchrony between cores and
device. In this paper we question this wisdom in the light of two
trends: modern and emerging cache-coherent interconnects like CXL3.0,
and workloads, particularly microservices and serverless computing.
Like some others before us, we argue that the assumptions of the
DMA-based model are obsolete, and in many use-cases programmed I/O,
where the CPU explicitly transfers data and control information to and
from a device via loads and stores, delivers a more efficient system.
However, we push this idea much further. We show, in a real hardware
implementation, the gains in latency for fine-grained communication
achievable using an open cache-coherence protocol which exposes cache
transitions to a smart device, and that throughput is competitive with
DMA over modern interconnects. We also demonstrate three use-cases:
fine-grained RPC-style invocation of functions on an accelerator,
offloading of operators in a streaming dataflow engine, and a network
interface targeting serverless functions, comparing our use of
coherence with both traditional DMA-style interaction and a
highly-optimized implementation using memory-mapped programmed I/O
over PCIe.

https://arxiv.org/abs/2409.08141
-- 
Regards,
John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
Please consider the environment before reading this e-mail. https://jl.ly

[toc] | [next] | [standalone]


#111516

FromLars Poulsen <lars@cleo.beagle-ears.com>
Date2025-04-26 16:28 +0000
Message-ID<slrn100q2dv.eisl.lars@cleo.beagle-ears.com>
In reply to#111515
On 2025-04-26, John Levine <johnl@taugh.com> wrote:
> Well, not entirely.  This preprint argues that in environments with
> lots of cores and where latency is an issue, programmed I/O can outperform
> DMA.
>
> Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
>
> Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>
> Conventional wisdom holds that an efficient interface between an OS
> running on a CPU and a high-bandwidth I/O device should use Direct
> Memory Access (DMA) to offload data transfer, descriptor rings for
> buffering and queuing, and interrupts for asynchrony between cores and
> device. In this paper we question this wisdom in the light of two
> trends: modern and emerging cache-coherent interconnects like CXL3.0,
> and workloads, particularly microservices and serverless computing.
> Like some others before us, we argue that the assumptions of the
> DMA-based model are obsolete, and in many use-cases programmed I/O,
> where the CPU explicitly transfers data and control information to and
> from a device via loads and stores, delivers a more efficient system.
> However, we push this idea much further. We show, in a real hardware
> implementation, the gains in latency for fine-grained communication
> achievable using an open cache-coherence protocol which exposes cache
> transitions to a smart device, and that throughput is competitive with
> DMA over modern interconnects. We also demonstrate three use-cases:
> fine-grained RPC-style invocation of functions on an accelerator,
> offloading of operators in a streaming dataflow engine, and a network
> interface targeting serverless functions, comparing our use of
> coherence with both traditional DMA-style interaction and a
> highly-optimized implementation using memory-mapped programmed I/O
> over PCIe.
>
> https://arxiv.org/abs/2409.08141

What is the difference between DMA and message-passing to another core
doing CMOV loop at the ISA level?

DMA means doing that it the micro-engine instead of at the ISA level.
Same difference.

What am I missing?

[toc] | [prev] | [next] | [standalone]


#111517

FromTerje Mathisen <terje.mathisen@tmsw.no>
Date2025-04-26 19:28 +0200
Message-ID<vuj53m$2s0jv$1@dont-email.me>
In reply to#111516
Lars Poulsen wrote:
> On 2025-04-26, John Levine <johnl@taugh.com> wrote:
>> Well, not entirely.  This preprint argues that in environments with
>> lots of cores and where latency is an issue, programmed I/O can outperform
>> DMA.
>>
>> Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
>>
>> Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
[snip]
>>
>> https://arxiv.org/abs/2409.08141
> 
> What is the difference between DMA and message-passing to another core
> doing CMOV loop at the ISA level?
> 
> DMA means doing that it the micro-engine instead of at the ISA level.
> Same difference.
> 
> What am I missing?
> 

I think, in the end it all comes down to power:

If the DMA engine can move n GB of data using less total power than 
having a regular core do it with programmed IO, then the DMA engine wins.

OTOH, I have argued here in c.arch that for most data input streams, a 
regular core is going to look at the data eventually, and in that case 
the same core can do the work and either process it directly (in 
register file sized or smaller blocks)or work as a prefetcher to first 
load up  $L1-sized blocks and then process that chunk.

On the gripping hand, if this is either going out, or you only need to 
look at a small percentage of the incoming cache lines worth of data, 
then the more power-efficient DMA engine can still win.

Terje

-- 
- <Terje.Mathisen at tmsw.no>
"almost all programming can be viewed as an exercise in caching"

[toc] | [prev] | [next] | [standalone]


#111524

FromTheo <theom+news@chiark.greenend.org.uk>
Date2025-04-27 19:35 +0100
Message-ID<Crc*345aA@news.chiark.greenend.org.uk>
In reply to#111516
Lars Poulsen <lars@cleo.beagle-ears.com> wrote:
> What is the difference between DMA and message-passing to another core
> doing CMOV loop at the ISA level?
> 
> DMA means doing that it the micro-engine instead of at the ISA level.
> Same difference.
> 
> What am I missing?

Width and specialisation.

You can absolutely write a DMA engine in software.  One thing that is
troublesome is that the CPU datapath might be a lot narrower than the number
of bits you can move in a single cycle.  eg on FPGA we can't clock logic
anywhere near the DRAM clock so we end up making a very wide memory bus that
runs at a lower clock - 512/1024/2048/...  bits wide.  You can do that in a
regular ISA using vector registers/instructions but it adds complexity you
don't need.

The other is that there's often some degree of marshalling that needs to
happen - reading scatter/gather lists, formatting packets the right way for
PCIe, filling in the right header fields, etc.  It's more efficient to do
that in hardware than it is to spend multiple instructions per packet doing
it.  Meanwhile the DRAM bandwidth is being wasted.

Of course you can customise your ISA with extra instructions for doing the
heavy lifting, but then arguably it's not really a CPU any more, it's a
'programmable DMA engine'.  The line between the two becomes very blurred.

What you might also have is a hard datapath that's orchestrated by a
tightly-coupled microcontroller.  Then you get the ability to have an engine
with good performance while being able to more flexibly program it, without
having to make a strange vector CPU.

It's easy to make 'a' DMA, but you want to push the maximum bandwidth that
the memory/interconnect can achieve and all the tricks are about getting
there.

Theo

[toc] | [prev] | [next] | [standalone]


#111527

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-04-27 20:49 +0000
Message-ID<499b7179ca8b4650a63c444fdc00c2cd@www.novabbs.org>
In reply to#111524
On Sun, 27 Apr 2025 18:35:08 +0000, Theo wrote:

> Lars Poulsen <lars@cleo.beagle-ears.com> wrote:
>> What is the difference between DMA and message-passing to another core
>> doing CMOV loop at the ISA level?
>>
>> DMA means doing that it the micro-engine instead of at the ISA level.
>> Same difference.
>>
>> What am I missing?
>
> Width and specialisation.
>
> You can absolutely write a DMA engine in software.  One thing that is
> troublesome is that the CPU datapath might be a lot narrower than the
> number
> of bits you can move in a single cycle.  eg on FPGA we can't clock logic
> anywhere near the DRAM clock so we end up making a very wide memory bus
> that
> runs at a lower clock - 512/1024/2048/...  bits wide.  You can do that
> in a
> regular ISA using vector registers/instructions but it adds complexity
> you
> don't need.

With anything at 7nm or smaller, the main core interconnect should be
1 cache line wide (512 bits = 64 bytes :: although IBM's choice of 256
byte cache lines might be troublesome for now.)

> The other is that there's often some degree of marshalling that needs to
> happen - reading scatter/gather lists, formatting packets the right way
> for
> PCIe, filling in the right header fields, etc.  It's more efficient to
> do
> that in hardware than it is to spend multiple instructions per packet
> doing
> it.  Meanwhile the DRAM bandwidth is being wasted.

SW is nadda-verrryyy guud at twiddling bits like HW is.

[toc] | [prev] | [next] | [standalone]


#111529

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-27 22:37 +0000
Message-ID<0lyPP.2742753$OrR5.1836911@fx18.iad>
In reply to#111527
mitchalsup@aol.com (MitchAlsup1) writes:
>On Sun, 27 Apr 2025 18:35:08 +0000, Theo wrote:
>
>> Lars Poulsen <lars@cleo.beagle-ears.com> wrote:
>>> What is the difference between DMA and message-passing to another core
>>> doing CMOV loop at the ISA level?
>>>
>>> DMA means doing that it the micro-engine instead of at the ISA level.
>>> Same difference.
>>>
>>> What am I missing?
>>
>> Width and specialisation.
>>
>> You can absolutely write a DMA engine in software.  One thing that is
>> troublesome is that the CPU datapath might be a lot narrower than the
>> number
>> of bits you can move in a single cycle.  eg on FPGA we can't clock logic
>> anywhere near the DRAM clock so we end up making a very wide memory bus
>> that
>> runs at a lower clock - 512/1024/2048/...  bits wide.  You can do that
>> in a
>> regular ISA using vector registers/instructions but it adds complexity
>> you
>> don't need.
>
>With anything at 7nm or smaller, the main core interconnect should be
>1 cache line wide (512 bits = 64 bytes :: although IBM's choice of 256
>byte cache lines might be troublesome for now.)
>
>> The other is that there's often some degree of marshalling that needs to
>> happen - reading scatter/gather lists, formatting packets the right way
>> for
>> PCIe, filling in the right header fields, etc.  It's more efficient to
>> do
>> that in hardware than it is to spend multiple instructions per packet
>> doing
>> it.  Meanwhile the DRAM bandwidth is being wasted.
>
>SW is nadda-verrryyy guud at twiddling bits like HW is.

IME, the use of content addressible memory[*] is a key component required
for efficient packet processing, and that can only be done in hardware.

[*] For RSS and per-flow serialization et alia.

[toc] | [prev] | [next] | [standalone]


#111531

FromLawrence D'Oliveiro <ldo@nz.invalid>
Date2025-04-28 01:20 +0000
Message-ID<vuml44$21v0f$3@dont-email.me>
In reply to#111524
On 27 Apr 2025 19:35:08 +0100 (BST), Theo wrote:

> Of course you can customise your ISA with extra instructions for doing
> the heavy lifting, but then arguably it's not really a CPU any more,
> it's a 'programmable DMA engine'.  The line between the two becomes very
> blurred.

In the old mainframe world, they called that an “I/O channel”. See also 
the RP2040 chip from the Raspberry Pi Foundation.

[toc] | [prev] | [next] | [standalone]


#111518

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-26 17:29 +0000
Message-ID<CJ8PP.2629170$eNx6.1327988@fx14.iad>
In reply to#111515
John Levine <johnl@taugh.com> writes:
>Well, not entirely.  This preprint argues that in environments with
>lots of cores and where latency is an issue, programmed I/O can outperform
>DMA.
>
>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
>
>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>
 <snip abstract>

>
>https://arxiv.org/abs/2409.08141

Interesting article, thanks for posting.

Their conclusion is not at all surprising for the operations they target
in the paper.   PCI express throughput has increased with each
generation, and PCI express latencies have decreased with each
generation.  There are certain workloads enabled by CXL that
benefit from reduced PCIe latency, but those are primarily
aimed at increasing directly accessible memory.

However,  I expect there are still benefits in using DMA for bulk data
transfer, particularly for network packet handling where
throughput is more interesting than PCI MMIO latency.

One concern that arises from the paper are the security
implications of device access to the cache coherency
protocol.   Not an issue for a well-behaved device, but
potentially problematic in a secure environment with
third-party CXL-mem devices.

At 3Leaf systems, we extended the coherency domain over
IB or 10Gbe Ethernet to encompass multiple servers in a
single coherency domain, which both facilitated I/O
and provided a single shared physical address space across
multiple servers (up to 16).     CXL-mem is basically the same
but using PCIe instead of IB.

Granted, that was close to 20 years ago, and switch latencies
were significant (100ns for IB, far more for Ethernet).

CXL-mem is a similar technology with a different transport (we
looked at Infiniband, 10Ge ethernet and "advanced switching"
(a flavor of PCIe)).  Infiniband was the most mature of the
three technologies and switch latencies were signifincantly lower
for IB than the competing transports. 

Today, my CPOE sells a couple of CXL2.0 enabled PCIe devices for
memory expansion; one has 16 high-end ARM V cores.

Quoting from the article (p.2)
 "   As a second example: for throughput-oriented workloads
  DMA has evolved to efficiently transfer data to and from
  main memory without polluting the CPU cache. However, for
  small, fine-grained interactions, it is important that almost all
  the data gets into the right CPU cache as quickly as possible."

Most modern CPU's support "allocate" hints on inbound DMA
that will automatically place the data in the right CPU cache as
quickly as possible.

Decomposing that packet transfer into CPU loads and stores
in a coherent fabric doesn't gain much, and burns more power
on the "device" than a DMA engine.

It's interesting they use one of the processors (designed in 2012) that we
built over a decade ago (and they mispell our company name :-)
in their research computer.   That processor does have
a mechanism to allocate data in cache on inbound DMA[*]; it's
worth noting that the 48 cores on that processor are in-order
cores.  The text comparing it with a modern i7 at 3.6ghz
doesn't note that.

[*] Although I don't recall if that mechanism was documented
    in the public processor technical documentation.

Their description of the Thunder X-1 processor cache is not accurate;
it's not PIPT, it is PIVT (implemented in such a way as to
appear to software as if it were PIPT).  The V drops a cycle
off the load-to-use latency.

It was also the only ARM64 processor chip we built with a cache-coherent
interconnect until the recent CXL based products.

Overall, a very interesting paper.

[toc] | [prev] | [next] | [standalone]


#111519

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-04-26 19:25 +0000
Message-ID<da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>
In reply to#111518
On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:

> John Levine <johnl@taugh.com> writes:
>>Well, not entirely.  This preprint argues that in environments with
>>lots of cores and where latency is an issue, programmed I/O can
>> outperform
>>DMA.
>>
>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent
>> Interconnects
>>
>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>>
>  <snip abstract>
>
>>
>>https://arxiv.org/abs/2409.08141
>
> Interesting article, thanks for posting.
>
> Their conclusion is not at all surprising for the operations they target
> in the paper.   PCI express throughput has increased with each
> generation, and PCI express latencies have decreased with each
> generation.  There are certain workloads enabled by CXL that
> benefit from reduced PCIe latency, but those are primarily
> aimed at increasing directly accessible memory.
>
> However,  I expect there are still benefits in using DMA for bulk data
> transfer, particularly for network packet handling where
> throughput is more interesting than PCI MMIO latency.

I would like to add a though to the concept under discussion::

Does the paper's conclusion hold better or worse if/when the
core ISA contains both LDM/STM and MM instructions. LDM/STM
allow for several sequential registers to move to/from MMI/O
memory in a single interconnect transaction, while MM allows
for up-to page-sized transfers in a single instruction and
only 2 interconnect transactions.

> One concern that arises from the paper are the security
> implications of device access to the cache coherency
> protocol.   Not an issue for a well-behaved device, but
> potentially problematic in a secure environment with
> third-party CXL-mem devices.

Citation please !?!

Also note:: device DMA goes through I/O MMU which adds a
modicum of security-fencing around device DMA accesses
but also adding latency.

> At 3Leaf systems, we extended the coherency domain over
> IB or 10Gbe Ethernet to encompass multiple servers in a
> single coherency domain, which both facilitated I/O
> and provided a single shared physical address space across
> multiple servers (up to 16). CXL-mem is basically the same
> but using PCIe instead of IB.

IB == InfiniBand ?!?

> Granted, that was close to 20 years ago, and switch latencies
> were significant (100ns for IB, far more for Ethernet).
>
> CXL-mem is a similar technology with a different transport (we
> looked at Infiniband, 10Ge ethernet and "advanced switching"
> (a flavor of PCIe)).  Infiniband was the most mature of the
> three technologies and switch latencies were signifincantly lower
> for IB than the competing transports.
>
> Today, my CPOE sells a couple of CXL2.0 enabled PCIe devices for
> memory expansion; one has 16 high-end ARM V cores.
>
> Quoting from the article (p.2)
>  "   As a second example: for throughput-oriented workloads
>   DMA has evolved to efficiently transfer data to and from
>   main memory without polluting the CPU cache. However, for
>   small, fine-grained interactions, it is important that almost all
>   the data gets into the right CPU cache as quickly as possible."
>
> Most modern CPU's support "allocate" hints on inbound DMA
> that will automatically place the data in the right CPU cache as
> quickly as possible.
>
> Decomposing that packet transfer into CPU loads and stores
> in a coherent fabric doesn't gain much, and burns more power
> on the "device" than a DMA engine.

That was my initial thought--core performing lots of LD/ST to
MMI/O is bound to consume more power than device DMA.

Secondarily, using 1-few cores to perform PIO is not going to
have the data land in the cache of the core that will run when
the data has been transferred. The data lands in the cache doing
PIO and not in the one to receive control after I/O is done.
{{It may still be "closer than" memory--but several cache
coherence protocols take longer cache-cache than dram-cache.}}

> It's interesting they use one of the processors (designed in 2012) that
> we
> built over a decade ago (and they mispell our company name :-)
> in their research computer.   That processor does have
> a mechanism to allocate data in cache on inbound DMA[*]; it's
> worth noting that the 48 cores on that processor are in-order
> cores.  The text comparing it with a modern i7 at 3.6ghz
> doesn't note that.
>
> [*] Although I don't recall if that mechanism was documented
>     in the public processor technical documentation.
>
> Their description of the Thunder X-1 processor cache is not accurate;
> it's not PIPT, it is PIVT (implemented in such a way as to
> appear to software as if it were PIPT).  The V drops a cycle
> off the load-to-use latency.

Generally its VIPT virtual index-physical tag with a few bits
of virtual-aliasing to disambiguate P&s.

> It was also the only ARM64 processor chip we built with a cache-coherent
> interconnect until the recent CXL based products.
>
> Overall, a very interesting paper.

Reminds me of trying to sell a micro x86-64 to AMD as a project.
The µ86 is a small x86-64 core made available as IP in Verilog
where it has/runs the same ISA as main GBOoO x86, but is placed
"out in the PCIe" interconnect--performing I/O services topo-
logically adjacent to the device itself. This allows 1ns access
latencies to DCRs and performing OS queueing of DPCs,... without
bothering the GBOoO cores.

AMD didn't buy the arguments.

[toc] | [prev] | [next] | [standalone]


#111521

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-27 14:01 +0000
Message-ID<SMqPP.2347519$FVcd.1912152@fx10.iad>
In reply to#111519
mitchalsup@aol.com (MitchAlsup1) writes:
>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>
>> John Levine <johnl@taugh.com> writes:
>>>Well, not entirely.  This preprint argues that in environments with
>>>lots of cores and where latency is an issue, programmed I/O can
>>> outperform
>>>DMA.
>>>
>>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent
>>> Interconnects
>>>
>>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>>>
>>  <snip abstract>
>>
>>>
>>>https://arxiv.org/abs/2409.08141
>>
>> Interesting article, thanks for posting.
>>
>> Their conclusion is not at all surprising for the operations they target
>> in the paper.   PCI express throughput has increased with each
>> generation, and PCI express latencies have decreased with each
>> generation.  There are certain workloads enabled by CXL that
>> benefit from reduced PCIe latency, but those are primarily
>> aimed at increasing directly accessible memory.
>>
>> However,  I expect there are still benefits in using DMA for bulk data
>> transfer, particularly for network packet handling where
>> throughput is more interesting than PCI MMIO latency.
>
>I would like to add a though to the concept under discussion::
>
>Does the paper's conclusion hold better or worse if/when the
>core ISA contains both LDM/STM and MM instructions. LDM/STM
>allow for several sequential registers to move to/from MMI/O
>memory in a single interconnect transaction, while MM allows
>for up-to page-sized transfers in a single instruction and
>only 2 interconnect transactions.

The article discusses using the intel 64-byte store
instructions, if I recall correctly.     ARM also has
a 64-byte store in the latest versions of the ISA.

>
>> One concern that arises from the paper are the security
>> implications of device access to the cache coherency
>> protocol.   Not an issue for a well-behaved device, but
>> potentially problematic in a secure environment with
>> third-party CXL-mem devices.
>
>Citation please !?!
>
>Also note:: device DMA goes through I/O MMU which adds a
>modicum of security-fencing around device DMA accesses
>but also adding latency.

Clearly any device that participates in the coherency
protocol can snoop addresses - a covert channel.

>
>> At 3Leaf systems, we extended the coherency domain over
>> IB or 10Gbe Ethernet to encompass multiple servers in a
>> single coherency domain, which both facilitated I/O
>> and provided a single shared physical address space across
>> multiple servers (up to 16). CXL-mem is basically the same
>> but using PCIe instead of IB.
>
>IB == InfiniBand ?!?

Yes.

>
 <snip>

>> Most modern CPU's support "allocate" hints on inbound DMA
>> that will automatically place the data in the right CPU cache as
>> quickly as possible.
>>
>> Decomposing that packet transfer into CPU loads and stores
>> in a coherent fabric doesn't gain much, and burns more power
>> on the "device" than a DMA engine.
>
>That was my initial thought--core performing lots of LD/ST to
>MMI/O is bound to consume more power than device DMA.
>
>Secondarily, using 1-few cores to perform PIO is not going to
>have the data land in the cache of the core that will run when
>the data has been transferred. The data lands in the cache doing
>PIO and not in the one to receive control after I/O is done.

The point of the paper is that the PIO is done into a local
cache line (owned/exclusive) by the device.  The entire cache
line eventually makes its way to memory (or a cache associated
with the processor processing the data) as a bulk transfer
rather than individual stores.

>{{It may still be "closer than" memory--but several cache
>coherence protocols take longer cache-cache than dram-cache.}}

That would reduce the benefits, to be sure.

>
>> It's interesting they use one of the processors (designed in 2012) that
>> we
>> built over a decade ago (and they mispell our company name :-)
>> in their research computer.   That processor does have
>> a mechanism to allocate data in cache on inbound DMA[*]; it's
>> worth noting that the 48 cores on that processor are in-order
>> cores.  The text comparing it with a modern i7 at 3.6ghz
>> doesn't note that.
>>
>> [*] Although I don't recall if that mechanism was documented
>>     in the public processor technical documentation.
>>
>> Their description of the Thunder X-1 processor cache is not accurate;
>> it's not PIPT, it is PIVT (implemented in such a way as to
>> appear to software as if it were PIPT).  The V drops a cycle
>> off the load-to-use latency.
>
>Generally its VIPT virtual index-physical tag with a few bits
>of virtual-aliasing to disambiguate P&s.
>
>> It was also the only ARM64 processor chip we built with a cache-coherent
>> interconnect until the recent CXL based products.
>>
>> Overall, a very interesting paper.
>
>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>The µ86 is a small x86-64 core made available as IP in Verilog
>where it has/runs the same ISA as main GBOoO x86, but is placed
>"out in the PCIe" interconnect--performing I/O services topo-
>logically adjacent to the device itself. This allows 1ns access
>latencies to DCRs and performing OS queueing of DPCs,... without
>bothering the GBOoO cores.
>
>AMD didn't buy the arguments.

[toc] | [prev] | [next] | [standalone]


#111523

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-27 16:12 +0000
Message-ID<THsPP.1380824$f81.221818@fx48.iad>
In reply to#111521
scott@slp53.sl.home (Scott Lurndal) writes:
>mitchalsup@aol.com (MitchAlsup1) writes:
>>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>
>>> John Levine <johnl@taugh.com> writes:
>>>>Well, not entirely.  This preprint argues that in environments with
>>>>lots of cores and where latency is an issue, programmed I/O can
>>>> outperform
>>>>DMA.
>>>>
>>>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent
>>>> Interconnects
>>>>
>>>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>>>>
>>>  <snip abstract>
>>>
>>>>
>>>>https://arxiv.org/abs/2409.08141
>>>
>>> Interesting article, thanks for posting.
>>>
>>> Their conclusion is not at all surprising for the operations they target
>>> in the paper.   PCI express throughput has increased with each
>>> generation, and PCI express latencies have decreased with each
>>> generation.  There are certain workloads enabled by CXL that
>>> benefit from reduced PCIe latency, but those are primarily
>>> aimed at increasing directly accessible memory.
>>>
>>> However,  I expect there are still benefits in using DMA for bulk data
>>> transfer, particularly for network packet handling where
>>> throughput is more interesting than PCI MMIO latency.
>>
>>I would like to add a though to the concept under discussion::
>>
>>Does the paper's conclusion hold better or worse if/when the
>>core ISA contains both LDM/STM and MM instructions. LDM/STM
>>allow for several sequential registers to move to/from MMI/O
>>memory in a single interconnect transaction, while MM allows
>>for up-to page-sized transfers in a single instruction and
>>only 2 interconnect transactions.
>
>The article discusses using the intel 64-byte store
>instructions, if I recall correctly.     ARM also has
>a 64-byte store in the latest versions of the ISA.

They also mentioned that vector instructions weren't helpful
in the Thunder X1 case, due to internal 128-bit bus limitations.

[toc] | [prev] | [next] | [standalone]


#111522

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-27 14:02 +0000
Message-ID<wNqPP.2347520$FVcd.356196@fx10.iad>
In reply to#111519
mitchalsup@aol.com (MitchAlsup1) writes:
>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>

>> Their description of the Thunder X-1 processor cache is not accurate;
>> it's not PIPT, it is PIVT (implemented in such a way as to
>> appear to software as if it were PIPT).  The V drops a cycle
>> off the load-to-use latency.
>
>Generally its VIPT virtual index-physical tag with a few bits
>of virtual-aliasing to disambiguate P&s.

Yes, you are correct, I meant VIPT.

[toc] | [prev] | [next] | [standalone]


#111525

FromTheo <theom+news@chiark.greenend.org.uk>
Date2025-04-27 20:13 +0100
Message-ID<Frc*6b6aA@news.chiark.greenend.org.uk>
In reply to#111519
MitchAlsup1 <mitchalsup@aol.com> wrote:
> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
> 
> > However,  I expect there are still benefits in using DMA for bulk data
> > transfer, particularly for network packet handling where
> > throughput is more interesting than PCI MMIO latency.
> 
> I would like to add a though to the concept under discussion::
> 
> Does the paper's conclusion hold better or worse if/when the
> core ISA contains both LDM/STM and MM instructions. LDM/STM
> allow for several sequential registers to move to/from MMI/O
> memory in a single interconnect transaction, while MM allows
> for up-to page-sized transfers in a single instruction and
> only 2 interconnect transactions.

I think this depends on the scale of your core.  For say a NIC <-> CPU,
maybe the CPU has MM instructions, but perhaps the microcontroller on the
NIC doesn't.  That means eg the CPU can push packets to transmit, but the
NIC is not setup to push packets it has received - it has to ask the CPU to
pull which will slow things down.

You can add that feature of course, but then isn't it just becoming a DMA
engine?

ie it's about control path and datapath.  A controller doesn't need a wide
datapath (it isn't doing much compute) but the data transfer does need a
wide datapath.  If you size a CPU for a wide datapath then you end up
paying costs for that (eg wide GP registers when you don't need them).

> > One concern that arises from the paper are the security
> > implications of device access to the cache coherency
> > protocol.   Not an issue for a well-behaved device, but
> > potentially problematic in a secure environment with
> > third-party CXL-mem devices.
> 
> Citation please !?!

CXL's protection model isn't very good:
https://dl.acm.org/doi/pdf/10.1145/3617580
(declaration: I'm a coauthor)

> Also note:: device DMA goes through I/O MMU which adds a
> modicum of security-fencing around device DMA accesses
> but also adding latency.

Indeed, and page-based lookups are both slow (if you miss in the IOTLB) and
have spatial and temporal security issues.

> > Most modern CPU's support "allocate" hints on inbound DMA
> > that will automatically place the data in the right CPU cache as
> > quickly as possible.
> >
> > Decomposing that packet transfer into CPU loads and stores
> > in a coherent fabric doesn't gain much, and burns more power
> > on the "device" than a DMA engine.
> 
> That was my initial thought--core performing lots of LD/ST to
> MMI/O is bound to consume more power than device DMA.
> 
> Secondarily, using 1-few cores to perform PIO is not going to
> have the data land in the cache of the core that will run when
> the data has been transferred. The data lands in the cache doing
> PIO and not in the one to receive control after I/O is done.
> {{It may still be "closer than" memory--but several cache
> coherence protocols take longer cache-cache than dram-cache.}}

I think this is an 'it depends'.  If you're doing RPC type operations, it
takes more work to warm up the DMA than it does to just do PIO.  If you're
an SSD pulling a large file from flash, DMA is more efficient.  If you're
moving network packets, which involve multiple scatter-gathers per packet,
then maybe some heavy lifting is useful for the address handling.

> > It was also the only ARM64 processor chip we built with a cache-coherent
> > interconnect until the recent CXL based products.
> >
> > Overall, a very interesting paper.
> 
> Reminds me of trying to sell a micro x86-64 to AMD as a project.
> The µ86 is a small x86-64 core made available as IP in Verilog
> where it has/runs the same ISA as main GBOoO x86, but is placed
> "out in the PCIe" interconnect--performing I/O services topo-
> logically adjacent to the device itself. This allows 1ns access
> latencies to DCRs and performing OS queueing of DPCs,... without
> bothering the GBOoO cores.
> 
> AMD didn't buy the arguments.

Intel tried that with the Quark line of 'microcontrollers', which appeared
to be a warmed over P54 Pentium (whether it shared microarchitecture or RTL
I'm not sure).  They were too power hungry and unwieldy to be
microcontrollers - they also couldn't run Debian/x86 despite having an MMU
because they were too old for the LOCK CMPXCHG instruction Debian used (P54
didn't need to worry about concurrency, but we do now).

I think at the end of the day there isn't actually a whole lot of benefit to
running the same ISA on your I/O as on your CPU - there tends to be a fairly
hard line between 'drivers' (on the CPU) and 'firmware' (on the device), and
on the firmware side it's easier to throw in a small RISC (eg RISC-V
nowadays) than anything complicated.

Theo

[toc] | [prev] | [next] | [standalone]


#111526

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-04-27 20:45 +0000
Message-ID<3a1fb2f39007b517f06e610f3786e7cd@www.novabbs.org>
In reply to#111525
On Sun, 27 Apr 2025 19:13:47 +0000, Theo wrote:

> MitchAlsup1 <mitchalsup@aol.com> wrote:
>> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>--------------
>>> One concern that arises from the paper are the security
>>> implications of device access to the cache coherency
>>> protocol.   Not an issue for a well-behaved device, but
>>> potentially problematic in a secure environment with
>>> third-party CXL-mem devices.
>>
>> Citation please !?!
>
> CXL's protection model isn't very good:
> https://dl.acm.org/doi/pdf/10.1145/3617580
> (declaration: I'm a coauthor)

Thanks for the URL !!

>> Secondarily, using 1-few cores to perform PIO is not going to
>> have the data land in the cache of the core that will run when
>> the data has been transferred. The data lands in the cache doing
>> PIO and not in the one to receive control after I/O is done.
>> {{It may still be "closer than" memory--but several cache
>> coherence protocols take longer cache-cache than dram-cache.}}
>
> I think this is an 'it depends'.  If you're doing RPC type operations,
> it
> takes more work to warm up the DMA than it does to just do PIO.

Yes, it takes more cycles for a CPU to tell a device to move memory
from here to there than it takes CPU to just move the memory from
here to there. I was, instead, referring to a CPU where it has an
MM (move memory to memory) instruction, where the instruction is
allowed to be sent over the interconnect (say CRM controller) and
have DRC perform the M2M movement locally.

That is:: all major blocks in the system have their own DMA sequencer.

>   If you're
> an SSD pulling a large file from flash, DMA is more efficient.  If
> you're
> moving network packets, which involve multiple scatter-gathers per
> packet,
> then maybe some heavy lifting is useful for the address handling.
>
>>> It was also the only ARM64 processor chip we built with a cache-coherent
>>> interconnect until the recent CXL based products.
>>>
>>> Overall, a very interesting paper.
>>
>> Reminds me of trying to sell a micro x86-64 to AMD as a project.
>> The µ86 is a small x86-64 core made available as IP in Verilog
>> where it has/runs the same ISA as main GBOoO x86, but is placed
>> "out in the PCIe" interconnect--performing I/O services topo-
>> logically adjacent to the device itself. This allows 1ns access
>> latencies to DCRs and performing OS queueing of DPCs,... without
>> bothering the GBOoO cores.
>>
>> AMD didn't buy the arguments.
>
> Intel tried that with the Quark line of 'microcontrollers', which
> appeared
> to be a warmed over P54 Pentium (whether it shared microarchitecture or
> RTL I'm not sure).

In my case, the LBIO core ran exactly the same ISA as the big central
cores. If the remote cores were present, they would field the interrupt
access the device, schedule further cleanup work, and then kick the main
cores in their side.

In order to be viable, the same OS SW has to run whether the LBIO cores
are present or not.

>                 They were too power hungry and unwieldy to be
> microcontrollers - they also couldn't run Debian/x86 despite having an
> MMU because they were too old for the LOCK CMPXCHG instruction Debian
> used (P54 didn't need to worry about concurrency, but we do now).

It is not generally known, but back in ~2006 when Opteron was in full
swing, the HT fabric to the SouthBridge was actually coherent, we just
did not publish the coherence spec and thus devices could not use it.
But it was present, and if someone happened to know the protocol they
could have used the coherent nature of it.

My LBIO core would have made use of that.

And unlike P54, I was being designed for low power operations, and
it was going to use Opteron building blocks to simplify bug-for-bug
compatibilities.

[toc] | [prev] | [next] | [standalone]


#111530

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-04-27 22:44 +0000
Message-ID<WqyPP.2742754$OrR5.1912672@fx18.iad>
In reply to#111526
mitchalsup@aol.com (MitchAlsup1) writes:
>On Sun, 27 Apr 2025 19:13:47 +0000, Theo wrote:
>
>> MitchAlsup1 <mitchalsup@aol.com> wrote:
>>> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:

>> I think this is an 'it depends'.  If you're doing RPC type operations,
>> it
>> takes more work to warm up the DMA than it does to just do PIO.
>
>Yes, it takes more cycles for a CPU to tell a device to move memory
>from here to there than it takes CPU to just move the memory from
>here to there. I was, instead, referring to a CPU where it has an
>MM (move memory to memory) instruction, where the instruction is
>allowed to be sent over the interconnect (say CRM controller) and
>have DRC perform the M2M movement locally.

If the instruction is not asynchronous, then the differences
between a load-store loop and the MM instruction aren't
significant.    If MM is asynchronous, that complicates the
kernel-user interface.


>>                 They were too power hungry and unwieldy to be
>> microcontrollers - they also couldn't run Debian/x86 despite having an
>> MMU because they were too old for the LOCK CMPXCHG instruction Debian
>> used (P54 didn't need to worry about concurrency, but we do now).
>
>It is not generally known, but back in ~2006 when Opteron was in full
>swing, the HT fabric to the SouthBridge was actually coherent, we just
>did not publish the coherence spec and thus devices could not use it.
>But it was present, and if someone happened to know the protocol they
>could have used the coherent nature of it.

We had the coherent HT spec in 2005 and used it to design our
ASIC that extended the coherency domain over infiniband.  Taped
out in 2008.  Connected two Istanbul CPUs (IIRC) to our ASIC via
HT and connected that mainboard to an IB fabric.  Supported 16 such hosts
in a single-system image with fully coherent distributed memory
using a mellanox DDR IB switch.

The AMD CTO at the time was an advisor to our startup.

[toc] | [prev] | [next] | [standalone]


#111543

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-01 13:07 +0000
Message-ID<vuvrlr$att$1@reader1.panix.com>
In reply to#111519
In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>,
MitchAlsup1 <mitchalsup@aol.com> wrote:
>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>[snip]
>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>The µ86 is a small x86-64 core made available as IP in Verilog
>where it has/runs the same ISA as main GBOoO x86, but is placed
>"out in the PCIe" interconnect--performing I/O services topo-
>logically adjacent to the device itself. This allows 1ns access
>latencies to DCRs and performing OS queueing of DPCs,... without
>bothering the GBOoO cores.
>
>AMD didn't buy the arguments.

I can see it either way; I suppose the argument as to whether I
buy it or not comes down to, "in depends".  How much control do
I, as the OS implementer, have over this core?

If it is yet another hidden core embedded somewhere deep in the
SoC complex and I can't easily interact with it from the OS,
then no thanks: we've got enough of those between MP0, MP1, MP5,
etc, etc.

On the other hand, if it's got a "normal" APIC ID, the OS has
control over it like any other LP, and its coherent with the big
cores, then yeah, sign me up: I've been wanting something like
that for a long time now.

Consider a virtualization application.  A problem with, say,
SR-IOV is that very often the hypervisor wants to interpose some
sort of administrative policy between the virtual function and
whatever it actually corresponds to, but get out of the fast
path for most IO.  This implies a kind of offload architecture
where there's some (presumably software) agent dedicated to
handling IO that can be parameterized with such a policy.  A
core very close to the device could handle that swimmingly,
though I'm not sure it would be enough to do it at (say) line
rate for a 400Gbps NIC or Gen5 NVMe device.

...but why x86_64?  It strikes me that as long as the _data_
formats vis the software-visible ABI are the same, it doesn't
need to use the same ISA.  In fact, I can see advantages to not
doing so.

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111544

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-01 22:03 +0000
Message-ID<5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>
In reply to#111543
On Thu, 1 May 2025 13:07:07 +0000, Dan Cross wrote:

> In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>,
> MitchAlsup1 <mitchalsup@aol.com> wrote:
>>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>[snip]
>>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>>The µ86 is a small x86-64 core made available as IP in Verilog
>>where it has/runs the same ISA as main GBOoO x86, but is placed
>>"out in the PCIe" interconnect--performing I/O services topo-
>>logically adjacent to the device itself. This allows 1ns access
>>latencies to DCRs and performing OS queueing of DPCs,... without
>>bothering the GBOoO cores.
>>
>>AMD didn't buy the arguments.
>
> I can see it either way; I suppose the argument as to whether I
> buy it or not comes down to, "in depends".  How much control do
> I, as the OS implementer, have over this core?

Other than it being placed "away" from the centralized cores,
it runs the same ISA as the main cores has longer latency to
coherent memory and shorter latency to device control registers
--which is why it is placed close to the device itself:: latency.
The big fast centralized core is going to get microsecond latency
from MMI/O device whereas ASIC version will have handful of nano-
second latencies. So the 5 GHZ core sees ~1 microsecond while the
little ASIC sees 10 nanoseconds. ...

> If it is yet another hidden core embedded somewhere deep in the
> SoC complex and I can't easily interact with it from the OS,
> then no thanks: we've got enough of those between MP0, MP1, MP5,
> etc, etc.
>
> On the other hand, if it's got a "normal" APIC ID, the OS has
> control over it like any other LP, and its coherent with the big
> cores, then yeah, sign me up: I've been wanting something like
> that for a long time now.

It is just a core that is cheap enough to put in ASICs, that
can offload some I/O burden without you having to do anything
other than setting some bits in some CRs so interrupts are
routed to this core rather than some more centralized core.

> Consider a virtualization application.  A problem with, say,
> SR-IOV is that very often the hypervisor wants to interpose some
> sort of administrative policy between the virtual function and
> whatever it actually corresponds to, but get out of the fast
> path for most IO.  This implies a kind of offload architecture
> where there's some (presumably software) agent dedicated to
> handling IO that can be parameterized with such a policy.  A

Interesting:: Could you cite any literature, here !?!

> core very close to the device could handle that swimmingly,
> though I'm not sure it would be enough to do it at (say) line
> rate for a 400Gbps NIC or Gen5 NVMe device.

I suspect the 400 GHz NIC needs a rather BIG core to handle the
traffic loads.

> ....but why x86_64?  It strikes me that as long as the _data_
> formats vis the software-visible ABI are the same, it doesn't
> need to use the same ISA.  In fact, I can see advantages to not
> doing so.

Having the remote core run the same OS code as every other core
means the OS developers have fewer hoops to jump through. Bug-for
bug compatibility means that clearing of those CRs just leaves
the core out in the periphery idling and bothering no one.

On the other hand, you buy a motherboard with said ASIC core,
and you can boot the MB without putting a big chip in the
socket--but you may have to deal with scant DRAM since the
big centralized chip contains teh memory controller.


> 	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111545

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-02 02:15 +0000
Message-ID<vv19rs$t2d$1@reader1.panix.com>
In reply to#111544
In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>,
MitchAlsup1 <mitchalsup@aol.com> wrote:
>On Thu, 1 May 2025 13:07:07 +0000, Dan Cross wrote:
>> In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>,
>> MitchAlsup1 <mitchalsup@aol.com> wrote:
>>>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>>[snip]
>>>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>>>The µ86 is a small x86-64 core made available as IP in Verilog
>>>where it has/runs the same ISA as main GBOoO x86, but is placed
>>>"out in the PCIe" interconnect--performing I/O services topo-
>>>logically adjacent to the device itself. This allows 1ns access
>>>latencies to DCRs and performing OS queueing of DPCs,... without
>>>bothering the GBOoO cores.
>>>
>>>AMD didn't buy the arguments.
>>
>> I can see it either way; I suppose the argument as to whether I
>> buy it or not comes down to, "in depends".  How much control do
>> I, as the OS implementer, have over this core?
>
>Other than it being placed "away" from the centralized cores,
>it runs the same ISA as the main cores has longer latency to
>coherent memory and shorter latency to device control registers
>--which is why it is placed close to the device itself:: latency.
>The big fast centralized core is going to get microsecond latency
>from MMI/O device whereas ASIC version will have handful of nano-
>second latencies. So the 5 GHZ core sees ~1 microsecond while the
>little ASIC sees 10 nanoseconds. ...

Yes, I get the argument for WHY you'd do it, I just want to make
sure that it's an ordinary core (albeit one that is far away
from the sockets with the main SoC complexes) that I interact
with in the usual manner.  Compare to, say, MP1 or MP0 on AMD
Zen, where it runs its own (proprietary) firmware that I
interact with via an RPC protocol over an AXI bus, if I interact
with it at all: most OEMs just punt and run AGESA (we don't).

>> If it is yet another hidden core embedded somewhere deep in the
>> SoC complex and I can't easily interact with it from the OS,
>> then no thanks: we've got enough of those between MP0, MP1, MP5,
>> etc, etc.
>>
>> On the other hand, if it's got a "normal" APIC ID, the OS has
>> control over it like any other LP, and its coherent with the big
>> cores, then yeah, sign me up: I've been wanting something like
>> that for a long time now.
>
>It is just a core that is cheap enough to put in ASICs, that
>can offload some I/O burden without you having to do anything
>other than setting some bits in some CRs so interrupts are
>routed to this core rather than some more centralized core.

Sounds good.

>> Consider a virtualization application.  A problem with, say,
>> SR-IOV is that very often the hypervisor wants to interpose some
>> sort of administrative policy between the virtual function and
>> whatever it actually corresponds to, but get out of the fast
>> path for most IO.  This implies a kind of offload architecture
>> where there's some (presumably software) agent dedicated to
>> handling IO that can be parameterized with such a policy.  A
>
>Interesting:: Could you cite any literature, here !?!

Sure.  This paper is a bit older, but gets at the main points:
https://www.usenix.org/system/files/conference/nsdi18/nsdi18-firestone.pdf

I don't know if the details are public for similar technologies
from Amazon or Google.

>> core very close to the device could handle that swimmingly,
>> though I'm not sure it would be enough to do it at (say) line
>> rate for a 400Gbps NIC or Gen5 NVMe device.
>
>I suspect the 400 GHz NIC needs a rather BIG core to handle the
>traffic loads.

Indeed.  Part of the challenge for the hyperscalars is in
meeting that demand while not burning too many host resources,
which are the thing they're actually selling their customer in
the first place.  A lot of folks are pushing this off to the NIC
itself, and I've seen at least one team that implemented NVMe in
firmware on a 100Gbps NIC, exposed via SR-IOV, as part of a
disaggregated storage architecture.

Another option is to push this to the switch; things like Intel
Tofino2 were well-position for this, but of course Intel, in its
infinite wisdom and vision, canc'ed Tofino.

>> ....but why x86_64?  It strikes me that as long as the _data_
>> formats vis the software-visible ABI are the same, it doesn't
>> need to use the same ISA.  In fact, I can see advantages to not
>> doing so.
>
>Having the remote core run the same OS code as every other core
>means the OS developers have fewer hoops to jump through. Bug-for
>bug compatibility means that clearing of those CRs just leaves
>the core out in the periphery idling and bothering no one.

Eh...Having to jump through hoops here matters less to me for
this kind of use case than if I'm trying to use those cores for
general-purpose compute.  Having a separate ISA means I cannot
accidentally run a program meant only for the big cores on the
IO service processors.  As long as the OS has total control over
the execution of the core, and it participates in whatever cache
coherency scheme the rest of the system uses, then the ISA just
isn't that important.

>On the other hand, you buy a motherboard with said ASIC core,
>and you can boot the MB without putting a big chip in the
>socket--but you may have to deal with scant DRAM since the
>big centralized chip contains teh memory controller.

A neat hack for bragging rights, but not terribly practical?

Anyway, it's a neat idea.  It's very reminiscent of IBM channel
controllers, in a way.

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111546

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2025-05-02 05:34 +0000
Message-ID<2025May2.073450@mips.complang.tuwien.ac.at>
In reply to#111545
cross@spitfire.i.gajendra.net (Dan Cross) writes:
>In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>,
>MitchAlsup1 <mitchalsup@aol.com> wrote:
>>Other than it being placed "away" from the centralized cores,
>>it runs the same ISA as the main cores has longer latency to
>>coherent memory and shorter latency to device control registers
>>--which is why it is placed close to the device itself:: latency.
>>The big fast centralized core is going to get microsecond latency
>>from MMI/O device whereas ASIC version will have handful of nano-
>>second latencies. So the 5 GHZ core sees ~1 microsecond while the
>>little ASIC sees 10 nanoseconds. ...
>
>Yes, I get the argument for WHY you'd do it, I just want to make
>sure that it's an ordinary core (albeit one that is far away
>from the sockets with the main SoC complexes) that I interact
>with in the usual manner.

Intel has put 2 Crestmont cores (their then-current E-core, not at all
tiny) on the SoC tile (not the compute tile) of Meteor Lake.  The main
idea there seems to be to save power (Meteor Lake is a laptop CPU)
when doing low-load things like playing videos by keeping the compute
tile powered down.

>>> core very close to the device could handle that swimmingly,
>>> though I'm not sure it would be enough to do it at (say) line
>>> rate for a 400Gbps NIC or Gen5 NVMe device.
>>
>>I suspect the 400 GHz NIC needs a rather BIG core to handle the
>>traffic loads.

Looking at
https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a
Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and
write <18GB/s even to the L1 cache).  However, Chester Lam notes: "A53
offers very low cache bandwidth compared to pretty much any other core
we’ve analyzed."  I think, though, that a small in-order core like the
A53, but with enough load and store buffering and enough bandwidth to
I/O and the memory controller should not have a problem shoveling data
from or to a 400Gb/s NIC.  With 128 bits/cycle in each direction one
would need one transfer per cycle in each direction at 3125MHz to
achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop
overhead.  Given that the A53 typically only has 2GHz, supporting 256
bits/cycle of transfer width (for load and store instructions, i.e.,
along the lines of AVX-256) would be more appropriate.

Going for an OoO core (something like AMD's Bobcat or Intel's
Silvermont) would help achieve the bandwidth goals without excessive
fine-tuning of the software.

>>Having the remote core run the same OS code as every other core
>>means the OS developers have fewer hoops to jump through. Bug-for
>>bug compatibility means that clearing of those CRs just leaves
>>the core out in the periphery idling and bothering no one.
>
>Eh...Having to jump through hoops here matters less to me for
>this kind of use case than if I'm trying to use those cores for
>general-purpose compute.

I think it's the same thing as Greenspun's tenth rule: First you find
that a classical DMA engine is too limiting, then you find that an A53
is too limiting, and eventually you find that it would be practical to
run the ISA of the main cores.  In particular, it allows you to use
the toolchain of the main cores for developing them, and you can also
use the facilities of the main cores (e.g., debugging features that
may be absent of the I/O cores) during development.

>Having a separate ISA means I cannot
>accidentally run a program meant only for the big cores on the
>IO service processors.

Marking the binaries that should be able to run on the IO service
processors with some flag, and letting the component of the OS that
assigns processes to cores heed this flag is not rocket science.  You
probably also don't want to run programs for the I/O processors on the
main cores; whether you use a separate flag for indicating that, or
whether one flag indicates both is an interesting question.

>>On the other hand, you buy a motherboard with said ASIC core,
>>and you can boot the MB without putting a big chip in the
>>socket--but you may have to deal with scant DRAM since the
>>big centralized chip contains teh memory controller.
>
>A neat hack for bragging rights, but not terribly practical?

Very practical for updating the firmware of the board to support the
big chip you want to put in the socket (called "BIOS FlashBack" in
connection with AMD big chips).  In a case where we did not have that
feature, and the board did not support the CPU, we had to buy another
CPU to update the firmware
<https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>.  That's
especially relevant for AM4 boards, because the support chips make it
hard to use more than 16MB Flash for firmware, but the firmware for
all supported big chips does not fit into 16MB.  However, as the case
mentioned above shows, it's also relevant for Intel boards.

- anton
-- 
'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
  Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>

[toc] | [prev] | [next] | [standalone]


#111547

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-02 15:02 +0000
Message-ID<vv2mqb$hem$1@reader1.panix.com>
In reply to#111546
In article <2025May2.073450@mips.complang.tuwien.ac.at>,
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>,
>>>[snip]
>>>I suspect the 400 GHz NIC needs a rather BIG core to handle the
>>>traffic loads.
>
>Looking at
>https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a
>Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and
>write <18GB/s even to the L1 cache).  However, Chester Lam notes: "A53
>offers very low cache bandwidth compared to pretty much any other core
>we’ve analyzed."  I think, though, that a small in-order core like the
>A53, but with enough load and store buffering and enough bandwidth to
>I/O and the memory controller should not have a problem shoveling data
>from or to a 400Gb/s NIC.  With 128 bits/cycle in each direction one
>would need one transfer per cycle in each direction at 3125MHz to
>achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop
>overhead.  Given that the A53 typically only has 2GHz, supporting 256
>bits/cycle of transfer width (for load and store instructions, i.e.,
>along the lines of AVX-256) would be more appropriate.
>
>Going for an OoO core (something like AMD's Bobcat or Intel's
>Silvermont) would help achieve the bandwidth goals without excessive
>fine-tuning of the software.
>
>>>Having the remote core run the same OS code as every other core
>>>means the OS developers have fewer hoops to jump through. Bug-for
>>>bug compatibility means that clearing of those CRs just leaves
>>>the core out in the periphery idling and bothering no one.
>>
>>Eh...Having to jump through hoops here matters less to me for
>>this kind of use case than if I'm trying to use those cores for
>>general-purpose compute.
>
>I think it's the same thing as Greenspun's tenth rule: First you find
>that a classical DMA engine is too limiting, then you find that an A53
>is too limiting, and eventually you find that it would be practical to
>run the ISA of the main cores.  In particular, it allows you to use
>the toolchain of the main cores for developing them,

These are issues solveable with the software architecture and
build system for the host OS.   The important characteristic is
that the software coupling makes architectural sense, and that
simply does not require using the same ISA across IPs.

Indeed, consider AMD's Zen CPUs; the PSP/ASP/whatever it's
called these days is an ARM core while the big CPUs are x86.
I'm pretty sure there's an Xtensa DSP in there to do DRAM and
timing and PCIe link training.  Similarly with the ME on Intel.
A BMC might be running on whatever.  We increasingly see ARM
based SBCs that have small RISC-V microcontroller-class cores
embedded in the SoC for exactly this sort of thing.

At work, our service processor (granted, outside of the SoC but
tightly coupled at the board level) is a Cortex-M7, but we wrote
the OS for that, and we control the host OS that runs on x86,
so the SP and big CPUs can be mutually aware.  Our hardware RoT
is a smaller Cortex-M.  We don't have a BMC on our boards;
everything that it does is either done by the SP or built into
the host OS, both of which are measured by the RoT.

The problem is when such service cores are hidden (as they are
in the case of the PSP, SMU, MPIO, and similar components, to
use AMD as the example) and treated like black boxes by
software.  It's really cool that I can configure the IO crossbar
in useful way tailored to specific configurations, but it's much
less cool that I have to do what amounts to an RPC over the SMN
to some totally undocumented entity somewhere in the SoC to do
it.  Bluntly, as an OS person, I do not want random bits of code
running anywhere on my machine that I am not at least aware of
(yes, this includes firmware blobs on devices).

>and you can also
>use the facilities of the main cores (e.g., debugging features that
>may be absent of the I/O cores) during development.

This is interesting, but we've found it more useful going the
other way around.  We do most of our debugging via the SP.
Since The SP is also responsible for system initialization and
holding x86 in reset until we're reading for it to start
running, it's the obvious nexus for debugging the system
holistically.

I must admit that, since we design our own boards, so we have
options here that those buying from the consumer space or
traditional enterprise vendors don't, but that's one of the
considerable value-adds for hardware/software co-design.

>>Having a separate ISA means I cannot
>>accidentally run a program meant only for the big cores on the
>>IO service processors.
>
>Marking the binaries that should be able to run on the IO service
>processors with some flag, and letting the component of the OS that
>assigns processes to cores heed this flag is not rocket science.

I agree, that's easy.  And yet, mistakes will be made, and there
will be tension between wanting to dedicate those CPUs to IO
services and wanting to use them for GP programs: I can easily
imagine a paper where someone modifies a scheduler to move IO
bound programs to those cores.  Using a different ISA obviates
most of that, and provides an (admittedly modest) security benefit.

And if I already have to modify or configure the OS to
accommodate the existence of these things in the first place,
then accommodating an ISA difference really isn't that much
extra work.  The critical observation is that a typical SMP view
of the world no longer makes sense for the system architecture,
and trying to shoehorn that model onto the hardware reality is
just going to cause frustration.  Better to acknowledge that the

>You
>probably also don't want to run programs for the I/O processors on the
>main cores; whether you use a separate flag for indicating that, or
>whether one flag indicates both is an interesting question.
>
>>>On the other hand, you buy a motherboard with said ASIC core,
>>>and you can boot the MB without putting a big chip in the
>>>socket--but you may have to deal with scant DRAM since the
>>>big centralized chip contains teh memory controller.
>>
>>A neat hack for bragging rights, but not terribly practical?
>
>Very practical for updating the firmware of the board to support the
>big chip you want to put in the socket (called "BIOS FlashBack" in
>connection with AMD big chips).

"BIOS", as loaded from the EFS by the ABL on the PSP on EPYC
class chips, is usually stored in a QSPI flash on the main
board (though starting with Turin you _can_ boot via eSPI).
Strictly speaking, you don't _need_ an x86 core to rewrite that.
On our machines, we do that from the SP, but we don't use AGESA
or UEFI: all of the platform enablement stuff done in PEI and
DXE we do directly in the host OS.

Also, on AMD machines, again considering EPYC, it's up to system
software running on x86 to direct either the SMU or MPIO to
configure DXIO and the rest of the fabric before PCIe link
training even begins (releasing PCIe from PERST is done by
either the SMU or MPIO, depending on the specific
microarchitecture).  Where are these cores, again?  If they're
close to the devices, are they in the root complex or on the far
side of a bridge?  Can they even talk to the rest of the board?

Also, since this is x86, there's the issue of starting them and
getting them to run useful software.  Usually on x86 it's the
responsibility of the BSC to start APs (AGESA usually does CCX
initialization and starts all the threads and do APIC ID
assignment and so on, but then directs them to park and wait for
the OS to do the usual INIT/SIPI/SIPI dance); but if the BSC is
absent because the socket is unpopulated, what starts them?  And
what software are they running?  Again, it's not even clear that
they have access to QSPI to boot into e.g. AGESA; if they've got
some little local ROM or flash or something, then how does the
OS get control them?

Perhaps there's some kind of electrical interlock that brings
them up if the socket is empty, but one must answer the question
of what's responsbile before assuming you can use them, and it
seems like the _best_ course of action would be to leave them
in reset (or even powered off) until explicitly enabled by
software, probably via a write to some magic capability in
config space on a bridge (like how one interacts with SMN now).

>In a case where we did not have that
>feature, and the board did not support the CPU, we had to buy another
>CPU to update the firmware
><https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>.  That's
>especially relevant for AM4 boards, because the support chips make it
>hard to use more than 16MB Flash for firmware, but the firmware for
>all supported big chips does not fit into 16MB.  However, as the case
>mentioned above shows, it's also relevant for Intel boards.

You shouldn't need to boot the host operating system to do that,
though I get on most consumer-grade machines you'll do it via
something that interfaces with AGESA or UEFI.  Most server-grade
machines will have a BMC that can do this independently of the
main CPU, and I should be clear that I'm discounting use cases
for consumer grade boards, where I suspect something like this
is less interesting than on server hardware.

As I mentioned, we build our own boards, and this just isn't an
issue for us, but it's not clear to me that these small cores
would even have access to the QSPI to update flash with an new
EFS image (again, to use the AMD example).

	- Dan C.

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | comp.arch


csiph-web