Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.arch > #111515 > unrolled thread
| Started by | John Levine <johnl@taugh.com> |
|---|---|
| First post | 2025-04-26 16:19 +0000 |
| Last post | 2025-05-04 02:10 +0000 |
| Articles | 20 on this page of 42 — 14 participants |
Back to article view | Back to comp.arch
DMA is obsolete John Levine <johnl@taugh.com> - 2025-04-26 16:19 +0000
Re: DMA is obsolete Lars Poulsen <lars@cleo.beagle-ears.com> - 2025-04-26 16:28 +0000
Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-04-26 19:28 +0200
Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 19:35 +0100
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:49 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:37 +0000
Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-04-28 01:20 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-26 17:29 +0000
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-26 19:25 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:01 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 16:12 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:02 +0000
Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 20:13 +0100
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:45 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:44 +0000
Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-01 13:07 +0000
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-01 22:03 +0000
Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 02:15 +0000
Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-02 05:34 +0000
Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 15:02 +0000
Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-03 06:11 +0000
Re: DMA is obsolete Robert Finch <robfi680@gmail.com> - 2025-05-03 06:32 -0400
Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 13:33 +0000
IP (was: DMA is obsolete) Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 10:50 -0400
Re: IP (was: DMA is obsolete) Thomas Koenig <tkoenig@netcologne.de> - 2025-05-03 15:15 +0000
Re: IP (was: DMA is obsolete) John Levine <johnl@taugh.com> - 2025-05-03 15:46 +0000
Re: IP (was: DMA is obsolete) cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 16:52 +0000
Re: IP (was: DMA is obsolete) scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:31 +0000
Re: IP Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 23:04 -0400
Re: IP cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-04 09:56 +0000
Re: IP Thomas Koenig <tkoenig@netcologne.de> - 2025-05-04 10:17 +0000
Re: IP mitchalsup@aol.com (MitchAlsup1) - 2025-05-04 18:16 +0000
Re: IP Bill Findlay <findlaybill@blueyonder.co.uk> - 2025-05-04 19:37 +0100
Re: IP Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 21:31 +0000
Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 06:44 +0000
Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:53 +0000
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-03 23:02 +0000
Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-21 12:36 +0000
Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-02 17:40 +0000
Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-05-03 14:29 +0200
ND-10 (was Re: DMA is obsolete) Lars Poulsen <lars@beagle-ears.com> - 2025-05-03 23:30 +0000
Re: ND-10 (was Re: DMA is obsolete) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 02:10 +0000
Page 1 of 3 [1] 2 3 Next page →
| From | John Levine <johnl@taugh.com> |
|---|---|
| Date | 2025-04-26 16:19 +0000 |
| Subject | DMA is obsolete |
| Message-ID | <vuj131$fnu$1@gal.iecc.com> |
Well, not entirely. This preprint argues that in environments with lots of cores and where latency is an issue, programmed I/O can outperform DMA. Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe Conventional wisdom holds that an efficient interface between an OS running on a CPU and a high-bandwidth I/O device should use Direct Memory Access (DMA) to offload data transfer, descriptor rings for buffering and queuing, and interrupts for asynchrony between cores and device. In this paper we question this wisdom in the light of two trends: modern and emerging cache-coherent interconnects like CXL3.0, and workloads, particularly microservices and serverless computing. Like some others before us, we argue that the assumptions of the DMA-based model are obsolete, and in many use-cases programmed I/O, where the CPU explicitly transfers data and control information to and from a device via loads and stores, delivers a more efficient system. However, we push this idea much further. We show, in a real hardware implementation, the gains in latency for fine-grained communication achievable using an open cache-coherence protocol which exposes cache transitions to a smart device, and that throughput is competitive with DMA over modern interconnects. We also demonstrate three use-cases: fine-grained RPC-style invocation of functions on an accelerator, offloading of operators in a streaming dataflow engine, and a network interface targeting serverless functions, comparing our use of coherence with both traditional DMA-style interaction and a highly-optimized implementation using memory-mapped programmed I/O over PCIe. https://arxiv.org/abs/2409.08141 -- Regards, John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies", Please consider the environment before reading this e-mail. https://jl.ly
[toc] | [next] | [standalone]
| From | Lars Poulsen <lars@cleo.beagle-ears.com> |
|---|---|
| Date | 2025-04-26 16:28 +0000 |
| Message-ID | <slrn100q2dv.eisl.lars@cleo.beagle-ears.com> |
| In reply to | #111515 |
On 2025-04-26, John Levine <johnl@taugh.com> wrote: > Well, not entirely. This preprint argues that in environments with > lots of cores and where latency is an issue, programmed I/O can outperform > DMA. > > Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects > > Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe > > Conventional wisdom holds that an efficient interface between an OS > running on a CPU and a high-bandwidth I/O device should use Direct > Memory Access (DMA) to offload data transfer, descriptor rings for > buffering and queuing, and interrupts for asynchrony between cores and > device. In this paper we question this wisdom in the light of two > trends: modern and emerging cache-coherent interconnects like CXL3.0, > and workloads, particularly microservices and serverless computing. > Like some others before us, we argue that the assumptions of the > DMA-based model are obsolete, and in many use-cases programmed I/O, > where the CPU explicitly transfers data and control information to and > from a device via loads and stores, delivers a more efficient system. > However, we push this idea much further. We show, in a real hardware > implementation, the gains in latency for fine-grained communication > achievable using an open cache-coherence protocol which exposes cache > transitions to a smart device, and that throughput is competitive with > DMA over modern interconnects. We also demonstrate three use-cases: > fine-grained RPC-style invocation of functions on an accelerator, > offloading of operators in a streaming dataflow engine, and a network > interface targeting serverless functions, comparing our use of > coherence with both traditional DMA-style interaction and a > highly-optimized implementation using memory-mapped programmed I/O > over PCIe. > > https://arxiv.org/abs/2409.08141 What is the difference between DMA and message-passing to another core doing CMOV loop at the ISA level? DMA means doing that it the micro-engine instead of at the ISA level. Same difference. What am I missing?
[toc] | [prev] | [next] | [standalone]
| From | Terje Mathisen <terje.mathisen@tmsw.no> |
|---|---|
| Date | 2025-04-26 19:28 +0200 |
| Message-ID | <vuj53m$2s0jv$1@dont-email.me> |
| In reply to | #111516 |
Lars Poulsen wrote: > On 2025-04-26, John Levine <johnl@taugh.com> wrote: >> Well, not entirely. This preprint argues that in environments with >> lots of cores and where latency is an issue, programmed I/O can outperform >> DMA. >> >> Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects >> >> Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe [snip] >> >> https://arxiv.org/abs/2409.08141 > > What is the difference between DMA and message-passing to another core > doing CMOV loop at the ISA level? > > DMA means doing that it the micro-engine instead of at the ISA level. > Same difference. > > What am I missing? > I think, in the end it all comes down to power: If the DMA engine can move n GB of data using less total power than having a regular core do it with programmed IO, then the DMA engine wins. OTOH, I have argued here in c.arch that for most data input streams, a regular core is going to look at the data eventually, and in that case the same core can do the work and either process it directly (in register file sized or smaller blocks)or work as a prefetcher to first load up $L1-sized blocks and then process that chunk. On the gripping hand, if this is either going out, or you only need to look at a small percentage of the incoming cache lines worth of data, then the more power-efficient DMA engine can still win. Terje -- - <Terje.Mathisen at tmsw.no> "almost all programming can be viewed as an exercise in caching"
[toc] | [prev] | [next] | [standalone]
| From | Theo <theom+news@chiark.greenend.org.uk> |
|---|---|
| Date | 2025-04-27 19:35 +0100 |
| Message-ID | <Crc*345aA@news.chiark.greenend.org.uk> |
| In reply to | #111516 |
Lars Poulsen <lars@cleo.beagle-ears.com> wrote: > What is the difference between DMA and message-passing to another core > doing CMOV loop at the ISA level? > > DMA means doing that it the micro-engine instead of at the ISA level. > Same difference. > > What am I missing? Width and specialisation. You can absolutely write a DMA engine in software. One thing that is troublesome is that the CPU datapath might be a lot narrower than the number of bits you can move in a single cycle. eg on FPGA we can't clock logic anywhere near the DRAM clock so we end up making a very wide memory bus that runs at a lower clock - 512/1024/2048/... bits wide. You can do that in a regular ISA using vector registers/instructions but it adds complexity you don't need. The other is that there's often some degree of marshalling that needs to happen - reading scatter/gather lists, formatting packets the right way for PCIe, filling in the right header fields, etc. It's more efficient to do that in hardware than it is to spend multiple instructions per packet doing it. Meanwhile the DRAM bandwidth is being wasted. Of course you can customise your ISA with extra instructions for doing the heavy lifting, but then arguably it's not really a CPU any more, it's a 'programmable DMA engine'. The line between the two becomes very blurred. What you might also have is a hard datapath that's orchestrated by a tightly-coupled microcontroller. Then you get the ability to have an engine with good performance while being able to more flexibly program it, without having to make a strange vector CPU. It's easy to make 'a' DMA, but you want to push the maximum bandwidth that the memory/interconnect can achieve and all the tricks are about getting there. Theo
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-04-27 20:49 +0000 |
| Message-ID | <499b7179ca8b4650a63c444fdc00c2cd@www.novabbs.org> |
| In reply to | #111524 |
On Sun, 27 Apr 2025 18:35:08 +0000, Theo wrote: > Lars Poulsen <lars@cleo.beagle-ears.com> wrote: >> What is the difference between DMA and message-passing to another core >> doing CMOV loop at the ISA level? >> >> DMA means doing that it the micro-engine instead of at the ISA level. >> Same difference. >> >> What am I missing? > > Width and specialisation. > > You can absolutely write a DMA engine in software. One thing that is > troublesome is that the CPU datapath might be a lot narrower than the > number > of bits you can move in a single cycle. eg on FPGA we can't clock logic > anywhere near the DRAM clock so we end up making a very wide memory bus > that > runs at a lower clock - 512/1024/2048/... bits wide. You can do that > in a > regular ISA using vector registers/instructions but it adds complexity > you > don't need. With anything at 7nm or smaller, the main core interconnect should be 1 cache line wide (512 bits = 64 bytes :: although IBM's choice of 256 byte cache lines might be troublesome for now.) > The other is that there's often some degree of marshalling that needs to > happen - reading scatter/gather lists, formatting packets the right way > for > PCIe, filling in the right header fields, etc. It's more efficient to > do > that in hardware than it is to spend multiple instructions per packet > doing > it. Meanwhile the DRAM bandwidth is being wasted. SW is nadda-verrryyy guud at twiddling bits like HW is.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-27 22:37 +0000 |
| Message-ID | <0lyPP.2742753$OrR5.1836911@fx18.iad> |
| In reply to | #111527 |
mitchalsup@aol.com (MitchAlsup1) writes: >On Sun, 27 Apr 2025 18:35:08 +0000, Theo wrote: > >> Lars Poulsen <lars@cleo.beagle-ears.com> wrote: >>> What is the difference between DMA and message-passing to another core >>> doing CMOV loop at the ISA level? >>> >>> DMA means doing that it the micro-engine instead of at the ISA level. >>> Same difference. >>> >>> What am I missing? >> >> Width and specialisation. >> >> You can absolutely write a DMA engine in software. One thing that is >> troublesome is that the CPU datapath might be a lot narrower than the >> number >> of bits you can move in a single cycle. eg on FPGA we can't clock logic >> anywhere near the DRAM clock so we end up making a very wide memory bus >> that >> runs at a lower clock - 512/1024/2048/... bits wide. You can do that >> in a >> regular ISA using vector registers/instructions but it adds complexity >> you >> don't need. > >With anything at 7nm or smaller, the main core interconnect should be >1 cache line wide (512 bits = 64 bytes :: although IBM's choice of 256 >byte cache lines might be troublesome for now.) > >> The other is that there's often some degree of marshalling that needs to >> happen - reading scatter/gather lists, formatting packets the right way >> for >> PCIe, filling in the right header fields, etc. It's more efficient to >> do >> that in hardware than it is to spend multiple instructions per packet >> doing >> it. Meanwhile the DRAM bandwidth is being wasted. > >SW is nadda-verrryyy guud at twiddling bits like HW is. IME, the use of content addressible memory[*] is a key component required for efficient packet processing, and that can only be done in hardware. [*] For RSS and per-flow serialization et alia.
[toc] | [prev] | [next] | [standalone]
| From | Lawrence D'Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2025-04-28 01:20 +0000 |
| Message-ID | <vuml44$21v0f$3@dont-email.me> |
| In reply to | #111524 |
On 27 Apr 2025 19:35:08 +0100 (BST), Theo wrote: > Of course you can customise your ISA with extra instructions for doing > the heavy lifting, but then arguably it's not really a CPU any more, > it's a 'programmable DMA engine'. The line between the two becomes very > blurred. In the old mainframe world, they called that an “I/O channel”. See also the RP2040 chip from the Raspberry Pi Foundation.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-26 17:29 +0000 |
| Message-ID | <CJ8PP.2629170$eNx6.1327988@fx14.iad> |
| In reply to | #111515 |
John Levine <johnl@taugh.com> writes:
>Well, not entirely. This preprint argues that in environments with
>lots of cores and where latency is an issue, programmed I/O can outperform
>DMA.
>
>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent Interconnects
>
>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>
<snip abstract>
>
>https://arxiv.org/abs/2409.08141
Interesting article, thanks for posting.
Their conclusion is not at all surprising for the operations they target
in the paper. PCI express throughput has increased with each
generation, and PCI express latencies have decreased with each
generation. There are certain workloads enabled by CXL that
benefit from reduced PCIe latency, but those are primarily
aimed at increasing directly accessible memory.
However, I expect there are still benefits in using DMA for bulk data
transfer, particularly for network packet handling where
throughput is more interesting than PCI MMIO latency.
One concern that arises from the paper are the security
implications of device access to the cache coherency
protocol. Not an issue for a well-behaved device, but
potentially problematic in a secure environment with
third-party CXL-mem devices.
At 3Leaf systems, we extended the coherency domain over
IB or 10Gbe Ethernet to encompass multiple servers in a
single coherency domain, which both facilitated I/O
and provided a single shared physical address space across
multiple servers (up to 16). CXL-mem is basically the same
but using PCIe instead of IB.
Granted, that was close to 20 years ago, and switch latencies
were significant (100ns for IB, far more for Ethernet).
CXL-mem is a similar technology with a different transport (we
looked at Infiniband, 10Ge ethernet and "advanced switching"
(a flavor of PCIe)). Infiniband was the most mature of the
three technologies and switch latencies were signifincantly lower
for IB than the competing transports.
Today, my CPOE sells a couple of CXL2.0 enabled PCIe devices for
memory expansion; one has 16 high-end ARM V cores.
Quoting from the article (p.2)
" As a second example: for throughput-oriented workloads
DMA has evolved to efficiently transfer data to and from
main memory without polluting the CPU cache. However, for
small, fine-grained interactions, it is important that almost all
the data gets into the right CPU cache as quickly as possible."
Most modern CPU's support "allocate" hints on inbound DMA
that will automatically place the data in the right CPU cache as
quickly as possible.
Decomposing that packet transfer into CPU loads and stores
in a coherent fabric doesn't gain much, and burns more power
on the "device" than a DMA engine.
It's interesting they use one of the processors (designed in 2012) that we
built over a decade ago (and they mispell our company name :-)
in their research computer. That processor does have
a mechanism to allocate data in cache on inbound DMA[*]; it's
worth noting that the 48 cores on that processor are in-order
cores. The text comparing it with a modern i7 at 3.6ghz
doesn't note that.
[*] Although I don't recall if that mechanism was documented
in the public processor technical documentation.
Their description of the Thunder X-1 processor cache is not accurate;
it's not PIPT, it is PIVT (implemented in such a way as to
appear to software as if it were PIPT). The V drops a cycle
off the load-to-use latency.
It was also the only ARM64 processor chip we built with a cache-coherent
interconnect until the recent CXL based products.
Overall, a very interesting paper.
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-04-26 19:25 +0000 |
| Message-ID | <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org> |
| In reply to | #111518 |
On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
> John Levine <johnl@taugh.com> writes:
>>Well, not entirely. This preprint argues that in environments with
>>lots of cores and where latency is an issue, programmed I/O can
>> outperform
>>DMA.
>>
>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent
>> Interconnects
>>
>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>>
> <snip abstract>
>
>>
>>https://arxiv.org/abs/2409.08141
>
> Interesting article, thanks for posting.
>
> Their conclusion is not at all surprising for the operations they target
> in the paper. PCI express throughput has increased with each
> generation, and PCI express latencies have decreased with each
> generation. There are certain workloads enabled by CXL that
> benefit from reduced PCIe latency, but those are primarily
> aimed at increasing directly accessible memory.
>
> However, I expect there are still benefits in using DMA for bulk data
> transfer, particularly for network packet handling where
> throughput is more interesting than PCI MMIO latency.
I would like to add a though to the concept under discussion::
Does the paper's conclusion hold better or worse if/when the
core ISA contains both LDM/STM and MM instructions. LDM/STM
allow for several sequential registers to move to/from MMI/O
memory in a single interconnect transaction, while MM allows
for up-to page-sized transfers in a single instruction and
only 2 interconnect transactions.
> One concern that arises from the paper are the security
> implications of device access to the cache coherency
> protocol. Not an issue for a well-behaved device, but
> potentially problematic in a secure environment with
> third-party CXL-mem devices.
Citation please !?!
Also note:: device DMA goes through I/O MMU which adds a
modicum of security-fencing around device DMA accesses
but also adding latency.
> At 3Leaf systems, we extended the coherency domain over
> IB or 10Gbe Ethernet to encompass multiple servers in a
> single coherency domain, which both facilitated I/O
> and provided a single shared physical address space across
> multiple servers (up to 16). CXL-mem is basically the same
> but using PCIe instead of IB.
IB == InfiniBand ?!?
> Granted, that was close to 20 years ago, and switch latencies
> were significant (100ns for IB, far more for Ethernet).
>
> CXL-mem is a similar technology with a different transport (we
> looked at Infiniband, 10Ge ethernet and "advanced switching"
> (a flavor of PCIe)). Infiniband was the most mature of the
> three technologies and switch latencies were signifincantly lower
> for IB than the competing transports.
>
> Today, my CPOE sells a couple of CXL2.0 enabled PCIe devices for
> memory expansion; one has 16 high-end ARM V cores.
>
> Quoting from the article (p.2)
> " As a second example: for throughput-oriented workloads
> DMA has evolved to efficiently transfer data to and from
> main memory without polluting the CPU cache. However, for
> small, fine-grained interactions, it is important that almost all
> the data gets into the right CPU cache as quickly as possible."
>
> Most modern CPU's support "allocate" hints on inbound DMA
> that will automatically place the data in the right CPU cache as
> quickly as possible.
>
> Decomposing that packet transfer into CPU loads and stores
> in a coherent fabric doesn't gain much, and burns more power
> on the "device" than a DMA engine.
That was my initial thought--core performing lots of LD/ST to
MMI/O is bound to consume more power than device DMA.
Secondarily, using 1-few cores to perform PIO is not going to
have the data land in the cache of the core that will run when
the data has been transferred. The data lands in the cache doing
PIO and not in the one to receive control after I/O is done.
{{It may still be "closer than" memory--but several cache
coherence protocols take longer cache-cache than dram-cache.}}
> It's interesting they use one of the processors (designed in 2012) that
> we
> built over a decade ago (and they mispell our company name :-)
> in their research computer. That processor does have
> a mechanism to allocate data in cache on inbound DMA[*]; it's
> worth noting that the 48 cores on that processor are in-order
> cores. The text comparing it with a modern i7 at 3.6ghz
> doesn't note that.
>
> [*] Although I don't recall if that mechanism was documented
> in the public processor technical documentation.
>
> Their description of the Thunder X-1 processor cache is not accurate;
> it's not PIPT, it is PIVT (implemented in such a way as to
> appear to software as if it were PIPT). The V drops a cycle
> off the load-to-use latency.
Generally its VIPT virtual index-physical tag with a few bits
of virtual-aliasing to disambiguate P&s.
> It was also the only ARM64 processor chip we built with a cache-coherent
> interconnect until the recent CXL based products.
>
> Overall, a very interesting paper.
Reminds me of trying to sell a micro x86-64 to AMD as a project.
The µ86 is a small x86-64 core made available as IP in Verilog
where it has/runs the same ISA as main GBOoO x86, but is placed
"out in the PCIe" interconnect--performing I/O services topo-
logically adjacent to the device itself. This allows 1ns access
latencies to DCRs and performing OS queueing of DPCs,... without
bothering the GBOoO cores.
AMD didn't buy the arguments.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-27 14:01 +0000 |
| Message-ID | <SMqPP.2347519$FVcd.1912152@fx10.iad> |
| In reply to | #111519 |
mitchalsup@aol.com (MitchAlsup1) writes:
>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>
>> John Levine <johnl@taugh.com> writes:
>>>Well, not entirely. This preprint argues that in environments with
>>>lots of cores and where latency is an issue, programmed I/O can
>>> outperform
>>>DMA.
>>>
>>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent
>>> Interconnects
>>>
>>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe
>>>
>> <snip abstract>
>>
>>>
>>>https://arxiv.org/abs/2409.08141
>>
>> Interesting article, thanks for posting.
>>
>> Their conclusion is not at all surprising for the operations they target
>> in the paper. PCI express throughput has increased with each
>> generation, and PCI express latencies have decreased with each
>> generation. There are certain workloads enabled by CXL that
>> benefit from reduced PCIe latency, but those are primarily
>> aimed at increasing directly accessible memory.
>>
>> However, I expect there are still benefits in using DMA for bulk data
>> transfer, particularly for network packet handling where
>> throughput is more interesting than PCI MMIO latency.
>
>I would like to add a though to the concept under discussion::
>
>Does the paper's conclusion hold better or worse if/when the
>core ISA contains both LDM/STM and MM instructions. LDM/STM
>allow for several sequential registers to move to/from MMI/O
>memory in a single interconnect transaction, while MM allows
>for up-to page-sized transfers in a single instruction and
>only 2 interconnect transactions.
The article discusses using the intel 64-byte store
instructions, if I recall correctly. ARM also has
a 64-byte store in the latest versions of the ISA.
>
>> One concern that arises from the paper are the security
>> implications of device access to the cache coherency
>> protocol. Not an issue for a well-behaved device, but
>> potentially problematic in a secure environment with
>> third-party CXL-mem devices.
>
>Citation please !?!
>
>Also note:: device DMA goes through I/O MMU which adds a
>modicum of security-fencing around device DMA accesses
>but also adding latency.
Clearly any device that participates in the coherency
protocol can snoop addresses - a covert channel.
>
>> At 3Leaf systems, we extended the coherency domain over
>> IB or 10Gbe Ethernet to encompass multiple servers in a
>> single coherency domain, which both facilitated I/O
>> and provided a single shared physical address space across
>> multiple servers (up to 16). CXL-mem is basically the same
>> but using PCIe instead of IB.
>
>IB == InfiniBand ?!?
Yes.
>
<snip>
>> Most modern CPU's support "allocate" hints on inbound DMA
>> that will automatically place the data in the right CPU cache as
>> quickly as possible.
>>
>> Decomposing that packet transfer into CPU loads and stores
>> in a coherent fabric doesn't gain much, and burns more power
>> on the "device" than a DMA engine.
>
>That was my initial thought--core performing lots of LD/ST to
>MMI/O is bound to consume more power than device DMA.
>
>Secondarily, using 1-few cores to perform PIO is not going to
>have the data land in the cache of the core that will run when
>the data has been transferred. The data lands in the cache doing
>PIO and not in the one to receive control after I/O is done.
The point of the paper is that the PIO is done into a local
cache line (owned/exclusive) by the device. The entire cache
line eventually makes its way to memory (or a cache associated
with the processor processing the data) as a bulk transfer
rather than individual stores.
>{{It may still be "closer than" memory--but several cache
>coherence protocols take longer cache-cache than dram-cache.}}
That would reduce the benefits, to be sure.
>
>> It's interesting they use one of the processors (designed in 2012) that
>> we
>> built over a decade ago (and they mispell our company name :-)
>> in their research computer. That processor does have
>> a mechanism to allocate data in cache on inbound DMA[*]; it's
>> worth noting that the 48 cores on that processor are in-order
>> cores. The text comparing it with a modern i7 at 3.6ghz
>> doesn't note that.
>>
>> [*] Although I don't recall if that mechanism was documented
>> in the public processor technical documentation.
>>
>> Their description of the Thunder X-1 processor cache is not accurate;
>> it's not PIPT, it is PIVT (implemented in such a way as to
>> appear to software as if it were PIPT). The V drops a cycle
>> off the load-to-use latency.
>
>Generally its VIPT virtual index-physical tag with a few bits
>of virtual-aliasing to disambiguate P&s.
>
>> It was also the only ARM64 processor chip we built with a cache-coherent
>> interconnect until the recent CXL based products.
>>
>> Overall, a very interesting paper.
>
>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>The µ86 is a small x86-64 core made available as IP in Verilog
>where it has/runs the same ISA as main GBOoO x86, but is placed
>"out in the PCIe" interconnect--performing I/O services topo-
>logically adjacent to the device itself. This allows 1ns access
>latencies to DCRs and performing OS queueing of DPCs,... without
>bothering the GBOoO cores.
>
>AMD didn't buy the arguments.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-27 16:12 +0000 |
| Message-ID | <THsPP.1380824$f81.221818@fx48.iad> |
| In reply to | #111521 |
scott@slp53.sl.home (Scott Lurndal) writes: >mitchalsup@aol.com (MitchAlsup1) writes: >>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: >> >>> John Levine <johnl@taugh.com> writes: >>>>Well, not entirely. This preprint argues that in environments with >>>>lots of cores and where latency is an issue, programmed I/O can >>>> outperform >>>>DMA. >>>> >>>>Rethinking Programmed I/O for Fast Devices, Cheap Cores, and Coherent >>>> Interconnects >>>> >>>>Anastasiia Ruzhanskaia, Pengcheng Xu, David Cock, Timothy Roscoe >>>> >>> <snip abstract> >>> >>>> >>>>https://arxiv.org/abs/2409.08141 >>> >>> Interesting article, thanks for posting. >>> >>> Their conclusion is not at all surprising for the operations they target >>> in the paper. PCI express throughput has increased with each >>> generation, and PCI express latencies have decreased with each >>> generation. There are certain workloads enabled by CXL that >>> benefit from reduced PCIe latency, but those are primarily >>> aimed at increasing directly accessible memory. >>> >>> However, I expect there are still benefits in using DMA for bulk data >>> transfer, particularly for network packet handling where >>> throughput is more interesting than PCI MMIO latency. >> >>I would like to add a though to the concept under discussion:: >> >>Does the paper's conclusion hold better or worse if/when the >>core ISA contains both LDM/STM and MM instructions. LDM/STM >>allow for several sequential registers to move to/from MMI/O >>memory in a single interconnect transaction, while MM allows >>for up-to page-sized transfers in a single instruction and >>only 2 interconnect transactions. > >The article discusses using the intel 64-byte store >instructions, if I recall correctly. ARM also has >a 64-byte store in the latest versions of the ISA. They also mentioned that vector instructions weren't helpful in the Thunder X1 case, due to internal 128-bit bus limitations.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-27 14:02 +0000 |
| Message-ID | <wNqPP.2347520$FVcd.356196@fx10.iad> |
| In reply to | #111519 |
mitchalsup@aol.com (MitchAlsup1) writes: >On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: > >> Their description of the Thunder X-1 processor cache is not accurate; >> it's not PIPT, it is PIVT (implemented in such a way as to >> appear to software as if it were PIPT). The V drops a cycle >> off the load-to-use latency. > >Generally its VIPT virtual index-physical tag with a few bits >of virtual-aliasing to disambiguate P&s. Yes, you are correct, I meant VIPT.
[toc] | [prev] | [next] | [standalone]
| From | Theo <theom+news@chiark.greenend.org.uk> |
|---|---|
| Date | 2025-04-27 20:13 +0100 |
| Message-ID | <Frc*6b6aA@news.chiark.greenend.org.uk> |
| In reply to | #111519 |
MitchAlsup1 <mitchalsup@aol.com> wrote:
> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>
> > However, I expect there are still benefits in using DMA for bulk data
> > transfer, particularly for network packet handling where
> > throughput is more interesting than PCI MMIO latency.
>
> I would like to add a though to the concept under discussion::
>
> Does the paper's conclusion hold better or worse if/when the
> core ISA contains both LDM/STM and MM instructions. LDM/STM
> allow for several sequential registers to move to/from MMI/O
> memory in a single interconnect transaction, while MM allows
> for up-to page-sized transfers in a single instruction and
> only 2 interconnect transactions.
I think this depends on the scale of your core. For say a NIC <-> CPU,
maybe the CPU has MM instructions, but perhaps the microcontroller on the
NIC doesn't. That means eg the CPU can push packets to transmit, but the
NIC is not setup to push packets it has received - it has to ask the CPU to
pull which will slow things down.
You can add that feature of course, but then isn't it just becoming a DMA
engine?
ie it's about control path and datapath. A controller doesn't need a wide
datapath (it isn't doing much compute) but the data transfer does need a
wide datapath. If you size a CPU for a wide datapath then you end up
paying costs for that (eg wide GP registers when you don't need them).
> > One concern that arises from the paper are the security
> > implications of device access to the cache coherency
> > protocol. Not an issue for a well-behaved device, but
> > potentially problematic in a secure environment with
> > third-party CXL-mem devices.
>
> Citation please !?!
CXL's protection model isn't very good:
https://dl.acm.org/doi/pdf/10.1145/3617580
(declaration: I'm a coauthor)
> Also note:: device DMA goes through I/O MMU which adds a
> modicum of security-fencing around device DMA accesses
> but also adding latency.
Indeed, and page-based lookups are both slow (if you miss in the IOTLB) and
have spatial and temporal security issues.
> > Most modern CPU's support "allocate" hints on inbound DMA
> > that will automatically place the data in the right CPU cache as
> > quickly as possible.
> >
> > Decomposing that packet transfer into CPU loads and stores
> > in a coherent fabric doesn't gain much, and burns more power
> > on the "device" than a DMA engine.
>
> That was my initial thought--core performing lots of LD/ST to
> MMI/O is bound to consume more power than device DMA.
>
> Secondarily, using 1-few cores to perform PIO is not going to
> have the data land in the cache of the core that will run when
> the data has been transferred. The data lands in the cache doing
> PIO and not in the one to receive control after I/O is done.
> {{It may still be "closer than" memory--but several cache
> coherence protocols take longer cache-cache than dram-cache.}}
I think this is an 'it depends'. If you're doing RPC type operations, it
takes more work to warm up the DMA than it does to just do PIO. If you're
an SSD pulling a large file from flash, DMA is more efficient. If you're
moving network packets, which involve multiple scatter-gathers per packet,
then maybe some heavy lifting is useful for the address handling.
> > It was also the only ARM64 processor chip we built with a cache-coherent
> > interconnect until the recent CXL based products.
> >
> > Overall, a very interesting paper.
>
> Reminds me of trying to sell a micro x86-64 to AMD as a project.
> The µ86 is a small x86-64 core made available as IP in Verilog
> where it has/runs the same ISA as main GBOoO x86, but is placed
> "out in the PCIe" interconnect--performing I/O services topo-
> logically adjacent to the device itself. This allows 1ns access
> latencies to DCRs and performing OS queueing of DPCs,... without
> bothering the GBOoO cores.
>
> AMD didn't buy the arguments.
Intel tried that with the Quark line of 'microcontrollers', which appeared
to be a warmed over P54 Pentium (whether it shared microarchitecture or RTL
I'm not sure). They were too power hungry and unwieldy to be
microcontrollers - they also couldn't run Debian/x86 despite having an MMU
because they were too old for the LOCK CMPXCHG instruction Debian used (P54
didn't need to worry about concurrency, but we do now).
I think at the end of the day there isn't actually a whole lot of benefit to
running the same ISA on your I/O as on your CPU - there tends to be a fairly
hard line between 'drivers' (on the CPU) and 'firmware' (on the device), and
on the firmware side it's easier to throw in a small RISC (eg RISC-V
nowadays) than anything complicated.
Theo
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-04-27 20:45 +0000 |
| Message-ID | <3a1fb2f39007b517f06e610f3786e7cd@www.novabbs.org> |
| In reply to | #111525 |
On Sun, 27 Apr 2025 19:13:47 +0000, Theo wrote:
> MitchAlsup1 <mitchalsup@aol.com> wrote:
>> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>--------------
>>> One concern that arises from the paper are the security
>>> implications of device access to the cache coherency
>>> protocol. Not an issue for a well-behaved device, but
>>> potentially problematic in a secure environment with
>>> third-party CXL-mem devices.
>>
>> Citation please !?!
>
> CXL's protection model isn't very good:
> https://dl.acm.org/doi/pdf/10.1145/3617580
> (declaration: I'm a coauthor)
Thanks for the URL !!
>> Secondarily, using 1-few cores to perform PIO is not going to
>> have the data land in the cache of the core that will run when
>> the data has been transferred. The data lands in the cache doing
>> PIO and not in the one to receive control after I/O is done.
>> {{It may still be "closer than" memory--but several cache
>> coherence protocols take longer cache-cache than dram-cache.}}
>
> I think this is an 'it depends'. If you're doing RPC type operations,
> it
> takes more work to warm up the DMA than it does to just do PIO.
Yes, it takes more cycles for a CPU to tell a device to move memory
from here to there than it takes CPU to just move the memory from
here to there. I was, instead, referring to a CPU where it has an
MM (move memory to memory) instruction, where the instruction is
allowed to be sent over the interconnect (say CRM controller) and
have DRC perform the M2M movement locally.
That is:: all major blocks in the system have their own DMA sequencer.
> If you're
> an SSD pulling a large file from flash, DMA is more efficient. If
> you're
> moving network packets, which involve multiple scatter-gathers per
> packet,
> then maybe some heavy lifting is useful for the address handling.
>
>>> It was also the only ARM64 processor chip we built with a cache-coherent
>>> interconnect until the recent CXL based products.
>>>
>>> Overall, a very interesting paper.
>>
>> Reminds me of trying to sell a micro x86-64 to AMD as a project.
>> The µ86 is a small x86-64 core made available as IP in Verilog
>> where it has/runs the same ISA as main GBOoO x86, but is placed
>> "out in the PCIe" interconnect--performing I/O services topo-
>> logically adjacent to the device itself. This allows 1ns access
>> latencies to DCRs and performing OS queueing of DPCs,... without
>> bothering the GBOoO cores.
>>
>> AMD didn't buy the arguments.
>
> Intel tried that with the Quark line of 'microcontrollers', which
> appeared
> to be a warmed over P54 Pentium (whether it shared microarchitecture or
> RTL I'm not sure).
In my case, the LBIO core ran exactly the same ISA as the big central
cores. If the remote cores were present, they would field the interrupt
access the device, schedule further cleanup work, and then kick the main
cores in their side.
In order to be viable, the same OS SW has to run whether the LBIO cores
are present or not.
> They were too power hungry and unwieldy to be
> microcontrollers - they also couldn't run Debian/x86 despite having an
> MMU because they were too old for the LOCK CMPXCHG instruction Debian
> used (P54 didn't need to worry about concurrency, but we do now).
It is not generally known, but back in ~2006 when Opteron was in full
swing, the HT fabric to the SouthBridge was actually coherent, we just
did not publish the coherence spec and thus devices could not use it.
But it was present, and if someone happened to know the protocol they
could have used the coherent nature of it.
My LBIO core would have made use of that.
And unlike P54, I was being designed for low power operations, and
it was going to use Opteron building blocks to simplify bug-for-bug
compatibilities.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-04-27 22:44 +0000 |
| Message-ID | <WqyPP.2742754$OrR5.1912672@fx18.iad> |
| In reply to | #111526 |
mitchalsup@aol.com (MitchAlsup1) writes: >On Sun, 27 Apr 2025 19:13:47 +0000, Theo wrote: > >> MitchAlsup1 <mitchalsup@aol.com> wrote: >>> On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: >> I think this is an 'it depends'. If you're doing RPC type operations, >> it >> takes more work to warm up the DMA than it does to just do PIO. > >Yes, it takes more cycles for a CPU to tell a device to move memory >from here to there than it takes CPU to just move the memory from >here to there. I was, instead, referring to a CPU where it has an >MM (move memory to memory) instruction, where the instruction is >allowed to be sent over the interconnect (say CRM controller) and >have DRC perform the M2M movement locally. If the instruction is not asynchronous, then the differences between a load-store loop and the MM instruction aren't significant. If MM is asynchronous, that complicates the kernel-user interface. >> They were too power hungry and unwieldy to be >> microcontrollers - they also couldn't run Debian/x86 despite having an >> MMU because they were too old for the LOCK CMPXCHG instruction Debian >> used (P54 didn't need to worry about concurrency, but we do now). > >It is not generally known, but back in ~2006 when Opteron was in full >swing, the HT fabric to the SouthBridge was actually coherent, we just >did not publish the coherence spec and thus devices could not use it. >But it was present, and if someone happened to know the protocol they >could have used the coherent nature of it. We had the coherent HT spec in 2005 and used it to design our ASIC that extended the coherency domain over infiniband. Taped out in 2008. Connected two Istanbul CPUs (IIRC) to our ASIC via HT and connected that mainboard to an IB fabric. Supported 16 such hosts in a single-system image with fully coherent distributed memory using a mellanox DDR IB switch. The AMD CTO at the time was an advisor to our startup.
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2025-05-01 13:07 +0000 |
| Message-ID | <vuvrlr$att$1@reader1.panix.com> |
| In reply to | #111519 |
In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>, MitchAlsup1 <mitchalsup@aol.com> wrote: >On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: >[snip] >Reminds me of trying to sell a micro x86-64 to AMD as a project. >The µ86 is a small x86-64 core made available as IP in Verilog >where it has/runs the same ISA as main GBOoO x86, but is placed >"out in the PCIe" interconnect--performing I/O services topo- >logically adjacent to the device itself. This allows 1ns access >latencies to DCRs and performing OS queueing of DPCs,... without >bothering the GBOoO cores. > >AMD didn't buy the arguments. I can see it either way; I suppose the argument as to whether I buy it or not comes down to, "in depends". How much control do I, as the OS implementer, have over this core? If it is yet another hidden core embedded somewhere deep in the SoC complex and I can't easily interact with it from the OS, then no thanks: we've got enough of those between MP0, MP1, MP5, etc, etc. On the other hand, if it's got a "normal" APIC ID, the OS has control over it like any other LP, and its coherent with the big cores, then yeah, sign me up: I've been wanting something like that for a long time now. Consider a virtualization application. A problem with, say, SR-IOV is that very often the hypervisor wants to interpose some sort of administrative policy between the virtual function and whatever it actually corresponds to, but get out of the fast path for most IO. This implies a kind of offload architecture where there's some (presumably software) agent dedicated to handling IO that can be parameterized with such a policy. A core very close to the device could handle that swimmingly, though I'm not sure it would be enough to do it at (say) line rate for a 400Gbps NIC or Gen5 NVMe device. ...but why x86_64? It strikes me that as long as the _data_ formats vis the software-visible ABI are the same, it doesn't need to use the same ISA. In fact, I can see advantages to not doing so. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-05-01 22:03 +0000 |
| Message-ID | <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org> |
| In reply to | #111543 |
On Thu, 1 May 2025 13:07:07 +0000, Dan Cross wrote: > In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>, > MitchAlsup1 <mitchalsup@aol.com> wrote: >>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: >>[snip] >>Reminds me of trying to sell a micro x86-64 to AMD as a project. >>The µ86 is a small x86-64 core made available as IP in Verilog >>where it has/runs the same ISA as main GBOoO x86, but is placed >>"out in the PCIe" interconnect--performing I/O services topo- >>logically adjacent to the device itself. This allows 1ns access >>latencies to DCRs and performing OS queueing of DPCs,... without >>bothering the GBOoO cores. >> >>AMD didn't buy the arguments. > > I can see it either way; I suppose the argument as to whether I > buy it or not comes down to, "in depends". How much control do > I, as the OS implementer, have over this core? Other than it being placed "away" from the centralized cores, it runs the same ISA as the main cores has longer latency to coherent memory and shorter latency to device control registers --which is why it is placed close to the device itself:: latency. The big fast centralized core is going to get microsecond latency from MMI/O device whereas ASIC version will have handful of nano- second latencies. So the 5 GHZ core sees ~1 microsecond while the little ASIC sees 10 nanoseconds. ... > If it is yet another hidden core embedded somewhere deep in the > SoC complex and I can't easily interact with it from the OS, > then no thanks: we've got enough of those between MP0, MP1, MP5, > etc, etc. > > On the other hand, if it's got a "normal" APIC ID, the OS has > control over it like any other LP, and its coherent with the big > cores, then yeah, sign me up: I've been wanting something like > that for a long time now. It is just a core that is cheap enough to put in ASICs, that can offload some I/O burden without you having to do anything other than setting some bits in some CRs so interrupts are routed to this core rather than some more centralized core. > Consider a virtualization application. A problem with, say, > SR-IOV is that very often the hypervisor wants to interpose some > sort of administrative policy between the virtual function and > whatever it actually corresponds to, but get out of the fast > path for most IO. This implies a kind of offload architecture > where there's some (presumably software) agent dedicated to > handling IO that can be parameterized with such a policy. A Interesting:: Could you cite any literature, here !?! > core very close to the device could handle that swimmingly, > though I'm not sure it would be enough to do it at (say) line > rate for a 400Gbps NIC or Gen5 NVMe device. I suspect the 400 GHz NIC needs a rather BIG core to handle the traffic loads. > ....but why x86_64? It strikes me that as long as the _data_ > formats vis the software-visible ABI are the same, it doesn't > need to use the same ISA. In fact, I can see advantages to not > doing so. Having the remote core run the same OS code as every other core means the OS developers have fewer hoops to jump through. Bug-for bug compatibility means that clearing of those CRs just leaves the core out in the periphery idling and bothering no one. On the other hand, you buy a motherboard with said ASIC core, and you can boot the MB without putting a big chip in the socket--but you may have to deal with scant DRAM since the big centralized chip contains teh memory controller. > - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2025-05-02 02:15 +0000 |
| Message-ID | <vv19rs$t2d$1@reader1.panix.com> |
| In reply to | #111544 |
In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>, MitchAlsup1 <mitchalsup@aol.com> wrote: >On Thu, 1 May 2025 13:07:07 +0000, Dan Cross wrote: >> In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>, >> MitchAlsup1 <mitchalsup@aol.com> wrote: >>>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote: >>>[snip] >>>Reminds me of trying to sell a micro x86-64 to AMD as a project. >>>The µ86 is a small x86-64 core made available as IP in Verilog >>>where it has/runs the same ISA as main GBOoO x86, but is placed >>>"out in the PCIe" interconnect--performing I/O services topo- >>>logically adjacent to the device itself. This allows 1ns access >>>latencies to DCRs and performing OS queueing of DPCs,... without >>>bothering the GBOoO cores. >>> >>>AMD didn't buy the arguments. >> >> I can see it either way; I suppose the argument as to whether I >> buy it or not comes down to, "in depends". How much control do >> I, as the OS implementer, have over this core? > >Other than it being placed "away" from the centralized cores, >it runs the same ISA as the main cores has longer latency to >coherent memory and shorter latency to device control registers >--which is why it is placed close to the device itself:: latency. >The big fast centralized core is going to get microsecond latency >from MMI/O device whereas ASIC version will have handful of nano- >second latencies. So the 5 GHZ core sees ~1 microsecond while the >little ASIC sees 10 nanoseconds. ... Yes, I get the argument for WHY you'd do it, I just want to make sure that it's an ordinary core (albeit one that is far away from the sockets with the main SoC complexes) that I interact with in the usual manner. Compare to, say, MP1 or MP0 on AMD Zen, where it runs its own (proprietary) firmware that I interact with via an RPC protocol over an AXI bus, if I interact with it at all: most OEMs just punt and run AGESA (we don't). >> If it is yet another hidden core embedded somewhere deep in the >> SoC complex and I can't easily interact with it from the OS, >> then no thanks: we've got enough of those between MP0, MP1, MP5, >> etc, etc. >> >> On the other hand, if it's got a "normal" APIC ID, the OS has >> control over it like any other LP, and its coherent with the big >> cores, then yeah, sign me up: I've been wanting something like >> that for a long time now. > >It is just a core that is cheap enough to put in ASICs, that >can offload some I/O burden without you having to do anything >other than setting some bits in some CRs so interrupts are >routed to this core rather than some more centralized core. Sounds good. >> Consider a virtualization application. A problem with, say, >> SR-IOV is that very often the hypervisor wants to interpose some >> sort of administrative policy between the virtual function and >> whatever it actually corresponds to, but get out of the fast >> path for most IO. This implies a kind of offload architecture >> where there's some (presumably software) agent dedicated to >> handling IO that can be parameterized with such a policy. A > >Interesting:: Could you cite any literature, here !?! Sure. This paper is a bit older, but gets at the main points: https://www.usenix.org/system/files/conference/nsdi18/nsdi18-firestone.pdf I don't know if the details are public for similar technologies from Amazon or Google. >> core very close to the device could handle that swimmingly, >> though I'm not sure it would be enough to do it at (say) line >> rate for a 400Gbps NIC or Gen5 NVMe device. > >I suspect the 400 GHz NIC needs a rather BIG core to handle the >traffic loads. Indeed. Part of the challenge for the hyperscalars is in meeting that demand while not burning too many host resources, which are the thing they're actually selling their customer in the first place. A lot of folks are pushing this off to the NIC itself, and I've seen at least one team that implemented NVMe in firmware on a 100Gbps NIC, exposed via SR-IOV, as part of a disaggregated storage architecture. Another option is to push this to the switch; things like Intel Tofino2 were well-position for this, but of course Intel, in its infinite wisdom and vision, canc'ed Tofino. >> ....but why x86_64? It strikes me that as long as the _data_ >> formats vis the software-visible ABI are the same, it doesn't >> need to use the same ISA. In fact, I can see advantages to not >> doing so. > >Having the remote core run the same OS code as every other core >means the OS developers have fewer hoops to jump through. Bug-for >bug compatibility means that clearing of those CRs just leaves >the core out in the periphery idling and bothering no one. Eh...Having to jump through hoops here matters less to me for this kind of use case than if I'm trying to use those cores for general-purpose compute. Having a separate ISA means I cannot accidentally run a program meant only for the big cores on the IO service processors. As long as the OS has total control over the execution of the core, and it participates in whatever cache coherency scheme the rest of the system uses, then the ISA just isn't that important. >On the other hand, you buy a motherboard with said ASIC core, >and you can boot the MB without putting a big chip in the >socket--but you may have to deal with scant DRAM since the >big centralized chip contains teh memory controller. A neat hack for bragging rights, but not terribly practical? Anyway, it's a neat idea. It's very reminiscent of IBM channel controllers, in a way. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2025-05-02 05:34 +0000 |
| Message-ID | <2025May2.073450@mips.complang.tuwien.ac.at> |
| In reply to | #111545 |
cross@spitfire.i.gajendra.net (Dan Cross) writes: >In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>, >MitchAlsup1 <mitchalsup@aol.com> wrote: >>Other than it being placed "away" from the centralized cores, >>it runs the same ISA as the main cores has longer latency to >>coherent memory and shorter latency to device control registers >>--which is why it is placed close to the device itself:: latency. >>The big fast centralized core is going to get microsecond latency >>from MMI/O device whereas ASIC version will have handful of nano- >>second latencies. So the 5 GHZ core sees ~1 microsecond while the >>little ASIC sees 10 nanoseconds. ... > >Yes, I get the argument for WHY you'd do it, I just want to make >sure that it's an ordinary core (albeit one that is far away >from the sockets with the main SoC complexes) that I interact >with in the usual manner. Intel has put 2 Crestmont cores (their then-current E-core, not at all tiny) on the SoC tile (not the compute tile) of Meteor Lake. The main idea there seems to be to save power (Meteor Lake is a laptop CPU) when doing low-load things like playing videos by keeping the compute tile powered down. >>> core very close to the device could handle that swimmingly, >>> though I'm not sure it would be enough to do it at (say) line >>> rate for a 400Gbps NIC or Gen5 NVMe device. >> >>I suspect the 400 GHz NIC needs a rather BIG core to handle the >>traffic loads. Looking at https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and write <18GB/s even to the L1 cache). However, Chester Lam notes: "A53 offers very low cache bandwidth compared to pretty much any other core we’ve analyzed." I think, though, that a small in-order core like the A53, but with enough load and store buffering and enough bandwidth to I/O and the memory controller should not have a problem shoveling data from or to a 400Gb/s NIC. With 128 bits/cycle in each direction one would need one transfer per cycle in each direction at 3125MHz to achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop overhead. Given that the A53 typically only has 2GHz, supporting 256 bits/cycle of transfer width (for load and store instructions, i.e., along the lines of AVX-256) would be more appropriate. Going for an OoO core (something like AMD's Bobcat or Intel's Silvermont) would help achieve the bandwidth goals without excessive fine-tuning of the software. >>Having the remote core run the same OS code as every other core >>means the OS developers have fewer hoops to jump through. Bug-for >>bug compatibility means that clearing of those CRs just leaves >>the core out in the periphery idling and bothering no one. > >Eh...Having to jump through hoops here matters less to me for >this kind of use case than if I'm trying to use those cores for >general-purpose compute. I think it's the same thing as Greenspun's tenth rule: First you find that a classical DMA engine is too limiting, then you find that an A53 is too limiting, and eventually you find that it would be practical to run the ISA of the main cores. In particular, it allows you to use the toolchain of the main cores for developing them, and you can also use the facilities of the main cores (e.g., debugging features that may be absent of the I/O cores) during development. >Having a separate ISA means I cannot >accidentally run a program meant only for the big cores on the >IO service processors. Marking the binaries that should be able to run on the IO service processors with some flag, and letting the component of the OS that assigns processes to cores heed this flag is not rocket science. You probably also don't want to run programs for the I/O processors on the main cores; whether you use a separate flag for indicating that, or whether one flag indicates both is an interesting question. >>On the other hand, you buy a motherboard with said ASIC core, >>and you can boot the MB without putting a big chip in the >>socket--but you may have to deal with scant DRAM since the >>big centralized chip contains teh memory controller. > >A neat hack for bragging rights, but not terribly practical? Very practical for updating the firmware of the board to support the big chip you want to put in the socket (called "BIOS FlashBack" in connection with AMD big chips). In a case where we did not have that feature, and the board did not support the CPU, we had to buy another CPU to update the firmware <https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>. That's especially relevant for AM4 boards, because the support chips make it hard to use more than 16MB Flash for firmware, but the firmware for all supported big chips does not fit into 16MB. However, as the case mentioned above shows, it's also relevant for Intel boards. - anton -- 'Anyone trying for "industrial quality" ISA should avoid undefined behavior.' Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2025-05-02 15:02 +0000 |
| Message-ID | <vv2mqb$hem$1@reader1.panix.com> |
| In reply to | #111546 |
In article <2025May2.073450@mips.complang.tuwien.ac.at>, Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >cross@spitfire.i.gajendra.net (Dan Cross) writes: >>In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>, >>>[snip] >>>I suspect the 400 GHz NIC needs a rather BIG core to handle the >>>traffic loads. > >Looking at >https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a >Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and >write <18GB/s even to the L1 cache). However, Chester Lam notes: "A53 >offers very low cache bandwidth compared to pretty much any other core >we’ve analyzed." I think, though, that a small in-order core like the >A53, but with enough load and store buffering and enough bandwidth to >I/O and the memory controller should not have a problem shoveling data >from or to a 400Gb/s NIC. With 128 bits/cycle in each direction one >would need one transfer per cycle in each direction at 3125MHz to >achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop >overhead. Given that the A53 typically only has 2GHz, supporting 256 >bits/cycle of transfer width (for load and store instructions, i.e., >along the lines of AVX-256) would be more appropriate. > >Going for an OoO core (something like AMD's Bobcat or Intel's >Silvermont) would help achieve the bandwidth goals without excessive >fine-tuning of the software. > >>>Having the remote core run the same OS code as every other core >>>means the OS developers have fewer hoops to jump through. Bug-for >>>bug compatibility means that clearing of those CRs just leaves >>>the core out in the periphery idling and bothering no one. >> >>Eh...Having to jump through hoops here matters less to me for >>this kind of use case than if I'm trying to use those cores for >>general-purpose compute. > >I think it's the same thing as Greenspun's tenth rule: First you find >that a classical DMA engine is too limiting, then you find that an A53 >is too limiting, and eventually you find that it would be practical to >run the ISA of the main cores. In particular, it allows you to use >the toolchain of the main cores for developing them, These are issues solveable with the software architecture and build system for the host OS. The important characteristic is that the software coupling makes architectural sense, and that simply does not require using the same ISA across IPs. Indeed, consider AMD's Zen CPUs; the PSP/ASP/whatever it's called these days is an ARM core while the big CPUs are x86. I'm pretty sure there's an Xtensa DSP in there to do DRAM and timing and PCIe link training. Similarly with the ME on Intel. A BMC might be running on whatever. We increasingly see ARM based SBCs that have small RISC-V microcontroller-class cores embedded in the SoC for exactly this sort of thing. At work, our service processor (granted, outside of the SoC but tightly coupled at the board level) is a Cortex-M7, but we wrote the OS for that, and we control the host OS that runs on x86, so the SP and big CPUs can be mutually aware. Our hardware RoT is a smaller Cortex-M. We don't have a BMC on our boards; everything that it does is either done by the SP or built into the host OS, both of which are measured by the RoT. The problem is when such service cores are hidden (as they are in the case of the PSP, SMU, MPIO, and similar components, to use AMD as the example) and treated like black boxes by software. It's really cool that I can configure the IO crossbar in useful way tailored to specific configurations, but it's much less cool that I have to do what amounts to an RPC over the SMN to some totally undocumented entity somewhere in the SoC to do it. Bluntly, as an OS person, I do not want random bits of code running anywhere on my machine that I am not at least aware of (yes, this includes firmware blobs on devices). >and you can also >use the facilities of the main cores (e.g., debugging features that >may be absent of the I/O cores) during development. This is interesting, but we've found it more useful going the other way around. We do most of our debugging via the SP. Since The SP is also responsible for system initialization and holding x86 in reset until we're reading for it to start running, it's the obvious nexus for debugging the system holistically. I must admit that, since we design our own boards, so we have options here that those buying from the consumer space or traditional enterprise vendors don't, but that's one of the considerable value-adds for hardware/software co-design. >>Having a separate ISA means I cannot >>accidentally run a program meant only for the big cores on the >>IO service processors. > >Marking the binaries that should be able to run on the IO service >processors with some flag, and letting the component of the OS that >assigns processes to cores heed this flag is not rocket science. I agree, that's easy. And yet, mistakes will be made, and there will be tension between wanting to dedicate those CPUs to IO services and wanting to use them for GP programs: I can easily imagine a paper where someone modifies a scheduler to move IO bound programs to those cores. Using a different ISA obviates most of that, and provides an (admittedly modest) security benefit. And if I already have to modify or configure the OS to accommodate the existence of these things in the first place, then accommodating an ISA difference really isn't that much extra work. The critical observation is that a typical SMP view of the world no longer makes sense for the system architecture, and trying to shoehorn that model onto the hardware reality is just going to cause frustration. Better to acknowledge that the >You >probably also don't want to run programs for the I/O processors on the >main cores; whether you use a separate flag for indicating that, or >whether one flag indicates both is an interesting question. > >>>On the other hand, you buy a motherboard with said ASIC core, >>>and you can boot the MB without putting a big chip in the >>>socket--but you may have to deal with scant DRAM since the >>>big centralized chip contains teh memory controller. >> >>A neat hack for bragging rights, but not terribly practical? > >Very practical for updating the firmware of the board to support the >big chip you want to put in the socket (called "BIOS FlashBack" in >connection with AMD big chips). "BIOS", as loaded from the EFS by the ABL on the PSP on EPYC class chips, is usually stored in a QSPI flash on the main board (though starting with Turin you _can_ boot via eSPI). Strictly speaking, you don't _need_ an x86 core to rewrite that. On our machines, we do that from the SP, but we don't use AGESA or UEFI: all of the platform enablement stuff done in PEI and DXE we do directly in the host OS. Also, on AMD machines, again considering EPYC, it's up to system software running on x86 to direct either the SMU or MPIO to configure DXIO and the rest of the fabric before PCIe link training even begins (releasing PCIe from PERST is done by either the SMU or MPIO, depending on the specific microarchitecture). Where are these cores, again? If they're close to the devices, are they in the root complex or on the far side of a bridge? Can they even talk to the rest of the board? Also, since this is x86, there's the issue of starting them and getting them to run useful software. Usually on x86 it's the responsibility of the BSC to start APs (AGESA usually does CCX initialization and starts all the threads and do APIC ID assignment and so on, but then directs them to park and wait for the OS to do the usual INIT/SIPI/SIPI dance); but if the BSC is absent because the socket is unpopulated, what starts them? And what software are they running? Again, it's not even clear that they have access to QSPI to boot into e.g. AGESA; if they've got some little local ROM or flash or something, then how does the OS get control them? Perhaps there's some kind of electrical interlock that brings them up if the socket is empty, but one must answer the question of what's responsbile before assuming you can use them, and it seems like the _best_ course of action would be to leave them in reset (or even powered off) until explicitly enabled by software, probably via a write to some magic capability in config space on a bridge (like how one interacts with SMN now). >In a case where we did not have that >feature, and the board did not support the CPU, we had to buy another >CPU to update the firmware ><https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>. That's >especially relevant for AM4 boards, because the support chips make it >hard to use more than 16MB Flash for firmware, but the firmware for >all supported big chips does not fit into 16MB. However, as the case >mentioned above shows, it's also relevant for Intel boards. You shouldn't need to boot the host operating system to do that, though I get on most consumer-grade machines you'll do it via something that interfaces with AGESA or UEFI. Most server-grade machines will have a BMC that can do this independently of the main CPU, and I should be clear that I'm discounting use cases for consumer grade boards, where I suspect something like this is less interesting than on server hardware. As I mentioned, we build our own boards, and this just isn't an issue for us, but it's not clear to me that these small cores would even have access to the QSPI to update flash with an new EFS image (again, to use the AMD example). - Dan C.
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | comp.arch
csiph-web