Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch > #111515 > unrolled thread

DMA is obsolete

Started byJohn Levine <johnl@taugh.com>
First post2025-04-26 16:19 +0000
Last post2025-05-04 02:10 +0000
Articles 20 on this page of 42 — 14 participants

Back to article view | Back to comp.arch


Contents

  DMA is obsolete John Levine <johnl@taugh.com> - 2025-04-26 16:19 +0000
    Re: DMA is obsolete Lars Poulsen <lars@cleo.beagle-ears.com> - 2025-04-26 16:28 +0000
      Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-04-26 19:28 +0200
      Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 19:35 +0100
        Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:49 +0000
          Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:37 +0000
        Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-04-28 01:20 +0000
    Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-26 17:29 +0000
      Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-26 19:25 +0000
        Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:01 +0000
          Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 16:12 +0000
        Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 14:02 +0000
        Re: DMA is obsolete Theo <theom+news@chiark.greenend.org.uk> - 2025-04-27 20:13 +0100
          Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-04-27 20:45 +0000
            Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-04-27 22:44 +0000
        Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-01 13:07 +0000
          Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-01 22:03 +0000
            Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 02:15 +0000
              Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-02 05:34 +0000
                Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-02 15:02 +0000
                  Re: DMA is obsolete anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-03 06:11 +0000
                    Re: DMA is obsolete Robert Finch <robfi680@gmail.com> - 2025-05-03 06:32 -0400
                    Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 13:33 +0000
                      IP (was: DMA is obsolete) Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 10:50 -0400
                        Re: IP (was: DMA is obsolete) Thomas Koenig <tkoenig@netcologne.de> - 2025-05-03 15:15 +0000
                          Re: IP (was: DMA is obsolete) John Levine <johnl@taugh.com> - 2025-05-03 15:46 +0000
                            Re: IP (was: DMA is obsolete) cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-03 16:52 +0000
                              Re: IP (was: DMA is obsolete) scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:31 +0000
                              Re: IP Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-03 23:04 -0400
                                Re: IP cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-04 09:56 +0000
                                  Re: IP Thomas Koenig <tkoenig@netcologne.de> - 2025-05-04 10:17 +0000
                                    Re: IP mitchalsup@aol.com (MitchAlsup1) - 2025-05-04 18:16 +0000
                                      Re: IP Bill Findlay <findlaybill@blueyonder.co.uk> - 2025-05-04 19:37 +0100
                                    Re: IP Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 21:31 +0000
                    Re: DMA is obsolete Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 06:44 +0000
                  Re: DMA is obsolete scott@slp53.sl.home (Scott Lurndal) - 2025-05-03 21:53 +0000
                    Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-03 23:02 +0000
                    Re: DMA is obsolete cross@spitfire.i.gajendra.net (Dan Cross) - 2025-05-21 12:36 +0000
              Re: DMA is obsolete mitchalsup@aol.com (MitchAlsup1) - 2025-05-02 17:40 +0000
                Re: DMA is obsolete Terje Mathisen <terje.mathisen@tmsw.no> - 2025-05-03 14:29 +0200
                  ND-10 (was Re: DMA is obsolete) Lars Poulsen <lars@beagle-ears.com> - 2025-05-03 23:30 +0000
                    Re: ND-10 (was Re: DMA is obsolete) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-04 02:10 +0000

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#111555

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2025-05-03 06:11 +0000
Message-ID<2025May3.081100@mips.complang.tuwien.ac.at>
In reply to#111547
cross@spitfire.i.gajendra.net (Dan Cross) writes:
>In article <2025May2.073450@mips.complang.tuwien.ac.at>,
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>I think it's the same thing as Greenspun's tenth rule: First you find
>>that a classical DMA engine is too limiting, then you find that an A53
>>is too limiting, and eventually you find that it would be practical to
>>run the ISA of the main cores.  In particular, it allows you to use
>>the toolchain of the main cores for developing them,
>
>These are issues solveable with the software architecture and
>build system for the host OS.

Certainly, one can work around many bad decisions, and in reality one
has to work around some bad decisions, but the issue here is not
whether "the issues are solvable", but which decision leads to better
or worse consequences.

>The important characteristic is
>that the software coupling makes architectural sense, and that
>simply does not require using the same ISA across IPs.

IP?  Internet Protocol?  Software Coupling sounds to me like a concept
from Constantine out of my Software engineering class.  I guess you
did not mean either, but it's unclear what you mean.

In any case, I have made arguments why it would make sense to use the
same ISA as for the OS for programming the cores that replace DMA
engines.  I will discuss your counterarguments below, but the most
important one to me seems to be that these cores would cost more than
with a different ISA.  There is something to that, but when the
application ISA is cheap to implement (e.g., RV64GC), that cost is
small; it may be more an argument for also selecting the
cheap-to-implement ISA for the OS/application cores.

>Indeed, consider AMD's Zen CPUs; the PSP/ASP/whatever it's
>called these days is an ARM core while the big CPUs are x86.
>I'm pretty sure there's an Xtensa DSP in there to do DRAM and
>timing and PCIe link training.

The PSPs are not programmable by the OS or application programmers, so
using the same ISA would not benefit the OS or application
programmers.  By contrast, the idea for the DMA replacement engines is
that they are programmable by the OS and maybe the application
programmers, and that changes whether the same ISA is beneficial.

What is "ASP/whatever"?

>Similarly with the ME on Intel.

Last I read about it, ME uses a core developed by Intel with IA-32 or
AMD64; but in any case, the ME is not programmable by OS or
application programmers, either.

>A BMC might be running on whatever.

Again, a BMC is not programmable by OS or application programmers.

>We increasingly see ARM
>based SBCs that have small RISC-V microcontroller-class cores
>embedded in the SoC for exactly this sort of thing.

That's interesting; it points to RISC-V being cheaper to implement
than ARM.  As for "that sort of thing", they are all not programmable
by OS or application programmers, so see above.

>Our hardware RoT

?

>The problem is when such service cores are hidden (as they are
>in the case of the PSP, SMU, MPIO, and similar components, to
>use AMD as the example) and treated like black boxes by
>software.  It's really cool that I can configure the IO crossbar
>in useful way tailored to specific configurations, but it's much
>less cool that I have to do what amounts to an RPC over the SMN
>to some totally undocumented entity somewhere in the SoC to do
>it.  Bluntly, as an OS person, I do not want random bits of code
>running anywhere on my machine that I am not at least aware of
>(yes, this includes firmware blobs on devices).

Well, one goes with the other.  If you design the hardware for being
programmed by the OS programmers, you use the same ISA for all the
cores that the OS programmers program, whereas if you design the
hardware as programmed by "firmware" programmers, you use a
cheap-to-implement ISA and design the whole thing such that it is
opaque to OS programmers and only offers some certain capabilities to
OS programmers.

And that's not just limited to ISAs.  A very successful example is the
way that flash memory is usually exposed to OSs: as a block device
like a plain old hard disk, and all the idiosyncracies of flash are
hidden in the device behind a flash translation layer that is
implemented by a microcontroller on the device.

What's "SMN"?

>>and you can also
>>use the facilities of the main cores (e.g., debugging features that
>>may be absent of the I/O cores) during development.
>
>This is interesting, but we've found it more useful going the
>other way around.  We do most of our debugging via the SP.
>Since The SP is also responsible for system initialization and
>holding x86 in reset until we're reading for it to start
>running, it's the obvious nexus for debugging the system
>holistically.

Sure, for debugging on the core-dump level that's useful.  I was
thinking about watchpoint and breakpoint registers and performance
counters that one may not want to implement on the DMA-replacement
core, but that is implemented on the OS/application cores.

>>Marking the binaries that should be able to run on the IO service
>>processors with some flag, and letting the component of the OS that
>>assigns processes to cores heed this flag is not rocket science.
>
>I agree, that's easy.  And yet, mistakes will be made, and there
>will be tension between wanting to dedicate those CPUs to IO
>services and wanting to use them for GP programs: I can easily
>imagine a paper where someone modifies a scheduler to move IO
>bound programs to those cores.  Using a different ISA obviates
>most of that, and provides an (admittedly modest) security benefit.

If there really is such tension, that indicates that such cores would
be useful for general-purpose use.  That makes the case for using the
same ISA even stronger.

As for "mistakes will be made", that also goes the other way: With a
separate toolchain for the DMA-replacement ISA, there is lots of
opportunity for mistakes.

As for "security benefit", where is that supposed to come from?  What
attack scenario do you have in mind where that "security benefit"
could materialize?

>And if I already have to modify or configure the OS to
>accommodate the existence of these things in the first place,
>then accommodating an ISA difference really isn't that much
>extra work.  The critical observation is that a typical SMP view
>of the world no longer makes sense for the system architecture,
>and trying to shoehorn that model onto the hardware reality is
>just going to cause frustration.

The shared-memory multiprocessing view of the world is very
successful, while distributed-memory computers are limited to
supercomputing and other areas where hardware cost still dominates
over software cost (i.e., where the software crisis has not happened
yet); as an example of the lack of success of the distributed-memory
paradigm, take the PlayStation 3; programmers found it too hard to
work with, so they did not use the hardware well, and eventually Sony
decided to go for an SMP machine for the PlayStation 4 and 5.

OTOH, one can say that the way many peripherals work on
general-purpose computers is more along the lines of
distributed-memory; but that's probably due to the relative hardware
and software costs for that peripheral.  Sure, the performance
characteristics are non-uniform (NUMA) in many cases, but 1) caches
tend to smooth over that, and 2) most of the code is not
performance-critical, so it just needs to run, which is easier to
achieve with SMP and harder with distributed memory.

Sure, people have argued for advantages of other models for decades,
like you do now, but SMP has usually won.

>>>>On the other hand, you buy a motherboard with said ASIC core,
>>>>and you can boot the MB without putting a big chip in the
>>>>socket--but you may have to deal with scant DRAM since the
>>>>big centralized chip contains teh memory controller.
>>>
>>>A neat hack for bragging rights, but not terribly practical?
>>
>>Very practical for updating the firmware of the board to support the
>>big chip you want to put in the socket (called "BIOS FlashBack" in
>>connection with AMD big chips).
>
>"BIOS", as loaded from the EFS by the ABL on the PSP on EPYC
>class chips, is usually stored in a QSPI flash on the main
>board (though starting with Turin you _can_ boot via eSPI).
>Strictly speaking, you don't _need_ an x86 core to rewrite that.
>On our machines, we do that from the SP, but we don't use AGESA
>or UEFI: all of the platform enablement stuff done in PEI and
>DXE we do directly in the host OS.

EFS?  ABL?  QSPI? eSPI?  PEI?  DXE?

Anyway, what you do in your special setup does not detract from the
fact that being able to flash the firmware without having a working
main core has turned out to be so useful that out of 218 AM5
motherboards offered in Austria <https://geizhals.at/?cat=mbam5>, 203
have that feature.

>Also, on AMD machines, again considering EPYC, it's up to system
>software running on x86 to direct either the SMU or MPIO to
>configure DXIO and the rest of the fabric before PCIe link
>training even begins (releasing PCIe from PERST is done by
>either the SMU or MPIO, depending on the specific
>microarchitecture).  Where are these cores, again?  If they're
>close to the devices, are they in the root complex or on the far
>side of a bridge?  Can they even talk to the rest of the board?

The core that does the flashing obviously is on the board, not on the
CPU package (which may be absent).  I do not know where on the board
it is.  Typically only one USB port can be used for that, so that may
indicate that a special path may be used for that without initializing
all the USB ports and the other hardware that's necessary for that; I
think that some USB ports are directly connected to the CPU package,
so those would not work anyway.

>>In a case where we did not have that
>>feature, and the board did not support the CPU, we had to buy another
>>CPU to update the firmware
>><https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>.  That's
>>especially relevant for AM4 boards, because the support chips make it
>>hard to use more than 16MB Flash for firmware, but the firmware for
>>all supported big chips does not fit into 16MB.  However, as the case
>>mentioned above shows, it's also relevant for Intel boards.
>
>You shouldn't need to boot the host operating system to do that,
>though I get on most consumer-grade machines you'll do it via
>something that interfaces with AGESA or UEFI.

In the bad old days you had to boot into DOS and run a DOS program for
flashing the BIOS.  Or worse, Windows; not very useful if you don't
have Windows installed on the computer (DOS at least could be booted
from a floppy disk).  My last few experiences in that direction were
firmware flashing as a "BIOS" feature, and the flashback feature
(which has it's own problems, because communication with the user is
limited).

>Most server-grade
>machines will have a BMC that can do this independently of the
>main CPU,

And just in another posting you wrote "but not terribly practical?".
The board I mentioned above where we had to buy a separate CPU for
flashing mentioned a BMC on the feature list, but when we looked in
the manual, we found that the BMC is not delivered with the board, but
has to be bought separately.  There was also no mention that one can
use the BMC for flashing the BIOS.

>and I should be clear that I'm discounting use cases
>for consumer grade boards, where I suspect something like this
>is less interesting than on server hardware.

What makes you think so?  And what do you mean with "something like
this"?

1) "BIOS flashback" is a mostly-standard feature in AM5 (i.e.,
consumer-grade) boards.

2) DMA has been a standard feature in various forms on consumer
hardware since the first IBM PC in 1981, and replacing the DMA engines
with cores running a general-purpose ISA accessible to OS designers
will not be limited to servers; if hardware designers and OS
developers put development time into that, there is no reason for
limiting that effort to servers.  The existence of the LPE-Cores on
Meteor Lake (not a server chip) and the in-order ARM cores on various
smartphone SOCs, the existence of P-Cores and E-Cores on Intel
consumer-grade CPUs, while the server versions of these CPUs have the
E-Cores disabled, and the uniformity of cores on the dedicated server
CPUs indicates that non-uniform cores seem to be hard to sell in
server space.

- anton
-- 
'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
  Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>

[toc] | [prev] | [next] | [standalone]


#111556

FromRobert Finch <robfi680@gmail.com>
Date2025-05-03 06:32 -0400
Message-ID<vv4rcd$3bbi0$1@dont-email.me>
In reply to#111555
On 2025-05-03 2:11 a.m., Anton Ertl wrote:
> cross@spitfire.i.gajendra.net (Dan Cross) writes:
>> In article <2025May2.073450@mips.complang.tuwien.ac.at>,
>> Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>> I think it's the same thing as Greenspun's tenth rule: First you find
>>> that a classical DMA engine is too limiting, then you find that an A53
>>> is too limiting, and eventually you find that it would be practical to
>>> run the ISA of the main cores.  In particular, it allows you to use
>>> the toolchain of the main cores for developing them,
>>
>> These are issues solveable with the software architecture and
>> build system for the host OS.
> 
> Certainly, one can work around many bad decisions, and in reality one
> has to work around some bad decisions, but the issue here is not
> whether "the issues are solvable", but which decision leads to better
> or worse consequences.
> 
>> The important characteristic is
>> that the software coupling makes architectural sense, and that
>> simply does not require using the same ISA across IPs.
> 
> IP?  Internet Protocol?  Software Coupling sounds to me like a concept
> from Constantine out of my Software engineering class.  I guess you
> did not mean either, but it's unclear what you mean.
> 
> In any case, I have made arguments why it would make sense to use the
> same ISA as for the OS for programming the cores that replace DMA
> engines.  I will discuss your counterarguments below, but the most
> important one to me seems to be that these cores would cost more than
> with a different ISA.  There is something to that, but when the
> application ISA is cheap to implement (e.g., RV64GC), that cost is
> small; it may be more an argument for also selecting the
> cheap-to-implement ISA for the OS/application cores.
> 
>> Indeed, consider AMD's Zen CPUs; the PSP/ASP/whatever it's
>> called these days is an ARM core while the big CPUs are x86.
>> I'm pretty sure there's an Xtensa DSP in there to do DRAM and
>> timing and PCIe link training.
> 
> The PSPs are not programmable by the OS or application programmers, so
> using the same ISA would not benefit the OS or application
> programmers.  By contrast, the idea for the DMA replacement engines is
> that they are programmable by the OS and maybe the application
> programmers, and that changes whether the same ISA is beneficial.
> 
> What is "ASP/whatever"?
> 
>> Similarly with the ME on Intel.
> 
> Last I read about it, ME uses a core developed by Intel with IA-32 or
> AMD64; but in any case, the ME is not programmable by OS or
> application programmers, either.
> 
>> A BMC might be running on whatever.
> 
> Again, a BMC is not programmable by OS or application programmers.
> 
>> We increasingly see ARM
>> based SBCs that have small RISC-V microcontroller-class cores
>> embedded in the SoC for exactly this sort of thing.
> 
> That's interesting; it points to RISC-V being cheaper to implement
> than ARM.  As for "that sort of thing", they are all not programmable
> by OS or application programmers, so see above.
> 
>> Our hardware RoT
> 
> ?
> 
>> The problem is when such service cores are hidden (as they are
>> in the case of the PSP, SMU, MPIO, and similar components, to
>> use AMD as the example) and treated like black boxes by
>> software.  It's really cool that I can configure the IO crossbar
>> in useful way tailored to specific configurations, but it's much
>> less cool that I have to do what amounts to an RPC over the SMN
>> to some totally undocumented entity somewhere in the SoC to do
>> it.  Bluntly, as an OS person, I do not want random bits of code
>> running anywhere on my machine that I am not at least aware of
>> (yes, this includes firmware blobs on devices).
> 
> Well, one goes with the other.  If you design the hardware for being
> programmed by the OS programmers, you use the same ISA for all the
> cores that the OS programmers program, whereas if you design the
> hardware as programmed by "firmware" programmers, you use a
> cheap-to-implement ISA and design the whole thing such that it is
> opaque to OS programmers and only offers some certain capabilities to
> OS programmers.
> 
> And that's not just limited to ISAs.  A very successful example is the
> way that flash memory is usually exposed to OSs: as a block device
> like a plain old hard disk, and all the idiosyncracies of flash are
> hidden in the device behind a flash translation layer that is
> implemented by a microcontroller on the device.
> 
> What's "SMN"?
> 
>>> and you can also
>>> use the facilities of the main cores (e.g., debugging features that
>>> may be absent of the I/O cores) during development.
>>
>> This is interesting, but we've found it more useful going the
>> other way around.  We do most of our debugging via the SP.
>> Since The SP is also responsible for system initialization and
>> holding x86 in reset until we're reading for it to start
>> running, it's the obvious nexus for debugging the system
>> holistically.
> 
> Sure, for debugging on the core-dump level that's useful.  I was
> thinking about watchpoint and breakpoint registers and performance
> counters that one may not want to implement on the DMA-replacement
> core, but that is implemented on the OS/application cores.
> 
>>> Marking the binaries that should be able to run on the IO service
>>> processors with some flag, and letting the component of the OS that
>>> assigns processes to cores heed this flag is not rocket science.
>>
>> I agree, that's easy.  And yet, mistakes will be made, and there
>> will be tension between wanting to dedicate those CPUs to IO
>> services and wanting to use them for GP programs: I can easily
>> imagine a paper where someone modifies a scheduler to move IO
>> bound programs to those cores.  Using a different ISA obviates
>> most of that, and provides an (admittedly modest) security benefit.
> 
> If there really is such tension, that indicates that such cores would
> be useful for general-purpose use.  That makes the case for using the
> same ISA even stronger.
> 
> As for "mistakes will be made", that also goes the other way: With a
> separate toolchain for the DMA-replacement ISA, there is lots of
> opportunity for mistakes.
> 
> As for "security benefit", where is that supposed to come from?  What
> attack scenario do you have in mind where that "security benefit"
> could materialize?
> 
>> And if I already have to modify or configure the OS to
>> accommodate the existence of these things in the first place,
>> then accommodating an ISA difference really isn't that much
>> extra work.  The critical observation is that a typical SMP view
>> of the world no longer makes sense for the system architecture,
>> and trying to shoehorn that model onto the hardware reality is
>> just going to cause frustration.
> 
> The shared-memory multiprocessing view of the world is very
> successful, while distributed-memory computers are limited to
> supercomputing and other areas where hardware cost still dominates
> over software cost (i.e., where the software crisis has not happened
> yet); as an example of the lack of success of the distributed-memory
> paradigm, take the PlayStation 3; programmers found it too hard to
> work with, so they did not use the hardware well, and eventually Sony
> decided to go for an SMP machine for the PlayStation 4 and 5.
> 
> OTOH, one can say that the way many peripherals work on
> general-purpose computers is more along the lines of
> distributed-memory; but that's probably due to the relative hardware
> and software costs for that peripheral.  Sure, the performance
> characteristics are non-uniform (NUMA) in many cases, but 1) caches
> tend to smooth over that, and 2) most of the code is not
> performance-critical, so it just needs to run, which is easier to
> achieve with SMP and harder with distributed memory.
> 
> Sure, people have argued for advantages of other models for decades,
> like you do now, but SMP has usually won.
> 
>>>>> On the other hand, you buy a motherboard with said ASIC core,
>>>>> and you can boot the MB without putting a big chip in the
>>>>> socket--but you may have to deal with scant DRAM since the
>>>>> big centralized chip contains teh memory controller.
>>>>
>>>> A neat hack for bragging rights, but not terribly practical?
>>>
>>> Very practical for updating the firmware of the board to support the
>>> big chip you want to put in the socket (called "BIOS FlashBack" in
>>> connection with AMD big chips).
>>
>> "BIOS", as loaded from the EFS by the ABL on the PSP on EPYC
>> class chips, is usually stored in a QSPI flash on the main
>> board (though starting with Turin you _can_ boot via eSPI).
>> Strictly speaking, you don't _need_ an x86 core to rewrite that.
>> On our machines, we do that from the SP, but we don't use AGESA
>> or UEFI: all of the platform enablement stuff done in PEI and
>> DXE we do directly in the host OS.
> 
> EFS?  ABL?  QSPI? eSPI?  PEI?  DXE?
> 
> Anyway, what you do in your special setup does not detract from the
> fact that being able to flash the firmware without having a working
> main core has turned out to be so useful that out of 218 AM5
> motherboards offered in Austria <https://geizhals.at/?cat=mbam5>, 203
> have that feature.
> 
>> Also, on AMD machines, again considering EPYC, it's up to system
>> software running on x86 to direct either the SMU or MPIO to
>> configure DXIO and the rest of the fabric before PCIe link
>> training even begins (releasing PCIe from PERST is done by
>> either the SMU or MPIO, depending on the specific
>> microarchitecture).  Where are these cores, again?  If they're
>> close to the devices, are they in the root complex or on the far
>> side of a bridge?  Can they even talk to the rest of the board?
> 
> The core that does the flashing obviously is on the board, not on the
> CPU package (which may be absent).  I do not know where on the board
> it is.  Typically only one USB port can be used for that, so that may
> indicate that a special path may be used for that without initializing
> all the USB ports and the other hardware that's necessary for that; I
> think that some USB ports are directly connected to the CPU package,
> so those would not work anyway.
> 
>>> In a case where we did not have that
>>> feature, and the board did not support the CPU, we had to buy another
>>> CPU to update the firmware
>>> <https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>.  That's
>>> especially relevant for AM4 boards, because the support chips make it
>>> hard to use more than 16MB Flash for firmware, but the firmware for
>>> all supported big chips does not fit into 16MB.  However, as the case
>>> mentioned above shows, it's also relevant for Intel boards.
>>
>> You shouldn't need to boot the host operating system to do that,
>> though I get on most consumer-grade machines you'll do it via
>> something that interfaces with AGESA or UEFI.
> 
> In the bad old days you had to boot into DOS and run a DOS program for
> flashing the BIOS.  Or worse, Windows; not very useful if you don't
> have Windows installed on the computer (DOS at least could be booted
> from a floppy disk).  My last few experiences in that direction were
> firmware flashing as a "BIOS" feature, and the flashback feature
> (which has it's own problems, because communication with the user is
> limited).
> 
>> Most server-grade
>> machines will have a BMC that can do this independently of the
>> main CPU,
> 
> And just in another posting you wrote "but not terribly practical?".
> The board I mentioned above where we had to buy a separate CPU for
> flashing mentioned a BMC on the feature list, but when we looked in
> the manual, we found that the BMC is not delivered with the board, but
> has to be bought separately.  There was also no mention that one can
> use the BMC for flashing the BIOS.
> 
>> and I should be clear that I'm discounting use cases
>> for consumer grade boards, where I suspect something like this
>> is less interesting than on server hardware.
> 
> What makes you think so?  And what do you mean with "something like
> this"?
> 
> 1) "BIOS flashback" is a mostly-standard feature in AM5 (i.e.,
> consumer-grade) boards.
> 
> 2) DMA has been a standard feature in various forms on consumer
> hardware since the first IBM PC in 1981, and replacing the DMA engines
> with cores running a general-purpose ISA accessible to OS designers
> will not be limited to servers; if hardware designers and OS
> developers put development time into that, there is no reason for
> limiting that effort to servers.  The existence of the LPE-Cores on
> Meteor Lake (not a server chip) and the in-order ARM cores on various
> smartphone SOCs, the existence of P-Cores and E-Cores on Intel
> consumer-grade CPUs, while the server versions of these CPUs have the
> E-Cores disabled, and the uniformity of cores on the dedicated server
> CPUs indicates that non-uniform cores seem to be hard to sell in
> server space.
> 
> - anton

My gut tells me that it would be better to have a “flat” design with all 
processors of the same type. It would likely save a lot of debugging 
headaches. But this is from the perspective of a single developer. I 
think it may not be true however, that there would be more debugging 
headaches if the control CPUs were different than the main CPU. The 
“peripheral processors” would likely be cut down versions of the main 
CPU and have their own idiosyncratic bugs. They end up being a bit 
different anyway. How are bugs rated? I am thinking bugs per LOC 
regardless of CPU used. Sure there is a learning curve for a different 
processor, but that curve is likely short for an experienced person or 
long for a newbie.

I have been pondering how to add test facilities to my own CPU core and 
thinking of using a small co-processor. Possibly a stack machine or 
something like the OPC challenge processor.

[toc] | [prev] | [next] | [standalone]


#111558

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-03 13:33 +0000
Message-ID<vv55vr$6hg$1@reader1.panix.com>
In reply to#111555
In article <2025May3.081100@mips.complang.tuwien.ac.at>,
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>In article <2025May2.073450@mips.complang.tuwien.ac.at>,
>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>>I think it's the same thing as Greenspun's tenth rule: First you find
>>>that a classical DMA engine is too limiting, then you find that an A53
>>>is too limiting, and eventually you find that it would be practical to
>>>run the ISA of the main cores.  In particular, it allows you to use
>>>the toolchain of the main cores for developing them,
>>
>>These are issues solveable with the software architecture and
>>build system for the host OS.
>
>Certainly, one can work around many bad decisions, and in reality one
>has to work around some bad decisions, but the issue here is not
>whether "the issues are solvable", but which decision leads to better
>or worse consequences.

I don't know that either would be "better" or "worse" under any
objective criteria.  They would simply be different.

>>The important characteristic is
>>that the software coupling makes architectural sense, and that
>>simply does not require using the same ISA across IPs.
>
>IP?  Internet Protocol?

When we discuss hardware designs at this level, reusable
components that go into the system are often referred to as "IP
cores" or just "IPs".  For example, a UART might be an IP.

Think of them as building blocks that go into, say, a SoC.

>Software Coupling sounds to me like a concept
>from Constantine out of my Software engineering class.

I have no idea who or what that is, but it seems unrelated.

>I guess you
>did not mean either, but it's unclear what you mean.

It's a very common term in this context.
https://en.wikipedia.org/wiki/Semiconductor_intellectual_property_core

>In any case, I have made arguments why it would make sense to use the
>same ISA as for the OS for programming the cores that replace DMA
>engines.  I will discuss your counterarguments below, but the most
>important one to me seems to be that these cores would cost more than
>with a different ISA.  There is something to that, but when the
>application ISA is cheap to implement (e.g., RV64GC), that cost is
>small; it may be more an argument for also selecting the
>cheap-to-implement ISA for the OS/application cores.

Ok.

>>Indeed, consider AMD's Zen CPUs; the PSP/ASP/whatever it's
>>called these days is an ARM core while the big CPUs are x86.
>>I'm pretty sure there's an Xtensa DSP in there to do DRAM and
>>timing and PCIe link training.
>
>The PSPs are not programmable by the OS or application programmers, so
>using the same ISA would not benefit the OS or application
>programmers.

Its firmware ships in BIOS images.  You can, in fact, interact
with it from the OS.  The only thing that keeps it from being
programmable by the OS is signing keys.

>By contrast, the idea for the DMA replacement engines is
>that they are programmable by the OS and maybe the application
>programmers, and that changes whether the same ISA is beneficial.
>
>What is "ASP/whatever"?

The PSP, or "AMD Platform Security Processor", has many names.
AMD says that "PSP" is the "legacy name", and that the new name
is ASP, for "AMD Secure Processor", and that it provides
"runtime security services"; for example, the PSP implements a
TPM in firmware, and exposes a random number generator that x86
can access via the `RDRAND` instruction.

>>Similarly with the ME on Intel.
>
>Last I read about it, ME uses a core developed by Intel with IA-32 or
>AMD64; but in any case, the ME is not programmable by OS or
>application programmers, either.

I was under the impression that it started out as an ARM core,
but I may be mistaken.

In any case, where do you think its firmware comes from?

>>A BMC might be running on whatever.
>
>Again, a BMC is not programmable by OS or application programmers.

The people working on OpenBMC disagree.

>>We increasingly see ARM
>>based SBCs that have small RISC-V microcontroller-class cores
>>embedded in the SoC for exactly this sort of thing.
>
>That's interesting; it points to RISC-V being cheaper to implement
>than ARM.  As for "that sort of thing", they are all not programmable
>by OS or application programmers, so see above.

No, the entire point is to provide an off-load for things that
are real-time.  They are absolutely meant to be "programmable by
OS or application programmers", which is exactly the sort of
scenario that Mitch's proposed cores would be used for.

Is a GPU programmable?  Yes.  Does it use the same ISA as the
general purpose compute core?  No.

>>Our hardware RoT
>
>?

Root of Trust.

>>The problem is when such service cores are hidden (as they are
>>in the case of the PSP, SMU, MPIO, and similar components, to
>>use AMD as the example) and treated like black boxes by
>>software.  It's really cool that I can configure the IO crossbar
>>in useful way tailored to specific configurations, but it's much
>>less cool that I have to do what amounts to an RPC over the SMN
>>to some totally undocumented entity somewhere in the SoC to do
>>it.  Bluntly, as an OS person, I do not want random bits of code
>>running anywhere on my machine that I am not at least aware of
>>(yes, this includes firmware blobs on devices).
>
>Well, one goes with the other.  If you design the hardware for being
>programmed by the OS programmers, you use the same ISA for all the
>cores that the OS programmers program,

That's a categorical statement that is not well supported.  That
may be what is _usually_ done.  It is not what _has_ to be done,
or even what _should_ be done.

You may feel that ths is the way things should be done, but the
arguments you've presented so far are not persuasive.

>whereas if you design the
>hardware as programmed by "firmware" programmers, you use a
>cheap-to-implement ISA and design the whole thing such that it is
>opaque to OS programmers and only offers some certain capabilities to
>OS programmers.

There is little fundamental difference between "firmware" and
the "OS".  I would further argue that this model of walling off
bits of system programmed with "firmware" from the OS a dated
way of thinking about systems that is actively harmful.  See
Roscoe's OSDI'21 keynote, here:

https://www.usenix.org/conference/osdi21/presentation/fri-keynote

Insisting that we use the congealed model we currently use
because that's how it is done is circular reasoning.

>And that's not just limited to ISAs.  A very successful example is the
>way that flash memory is usually exposed to OSs: as a block device
>like a plain old hard disk, and all the idiosyncracies of flash are
>hidden in the device behind a flash translation layer that is
>implemented by a microcontroller on the device.

You're conflating a hardware interface with firmware.

>What's "SMN"?

The "System Management Network."  This is the thing that AMD
uses inside the SoC to talk between the different components
that make up the system (that is, between the different IPs in
the SoC).  SMN is really a network of AXI buses, but it's how
one can, say, read and write registers on various components.

If you look at, for example,
https://www.amd.com/content/dam/amd/en/documents/processor-tech-docs/programmer-references/55803-ppr-family-17h-model-31h-b0-processors.pdf
And you look at the enry for the SMU registers, you'll see
that they have an "aliasSMN" entry in the instance table; those
can be decoded to a 32-bit number.  That is the SMN address of
that register.  For example, `SMU::THM::THM_TCON_CUR_TMP` is the
thermal register maintained by the SMU that encodes the current
temperature (in normalized units that are scaled from e.g.
degrees C, to accommodate different operating temperature ranges
between different physical parts).  Anyway, if one were to
decode the address in the instance table, one would see that
that register is at SMN address 0x0005_9800.  One accesses SMN
via an address/data pair of registers on a special BDF (0/0/0)
in PCI config space.  If you write that address to offset 0x60
for 0/0/0, and then read form offset 0x64 on 0/0/0, you'll get
the contents of that register.  You can use either port IO or
ECAM for such accesses.

Similarly, consider `PCS::DXIO::PCS_GOPX16_PCS_STATUS1`, which
is a register with multiple instances for each XGMI PCS (before
you ask, "PCS" is "Physical Coding Sublayer" and xGMI is the
socket-to-socket [external] Global Memory Interface).  That is,
these are the SerDes (Serializer/Deserializer) for communicating
between sockets.  Anwyway, the SMN address that corresponds to
PCS 21, serdes aggregator 1, is 0x12ff_0050.

>>>and you can also
>>>use the facilities of the main cores (e.g., debugging features that
>>>may be absent of the I/O cores) during development.
>>
>>This is interesting, but we've found it more useful going the
>>other way around.  We do most of our debugging via the SP.
>>Since The SP is also responsible for system initialization and
>>holding x86 in reset until we're reading for it to start
>>running, it's the obvious nexus for debugging the system
>>holistically.
>
>Sure, for debugging on the core-dump level that's useful.  I was
>thinking about watchpoint and breakpoint registers and performance
>counters that one may not want to implement on the DMA-replacement
>core, but that is implemented on the OS/application cores.

I assumed you were talking about remote hardware debugging
interfaces.  You seem to be talking about just running a
debugger or profiler on the IO offload core.  That's a much
simpler use case.

>>>Marking the binaries that should be able to run on the IO service
>>>processors with some flag, and letting the component of the OS that
>>>assigns processes to cores heed this flag is not rocket science.
>>
>>I agree, that's easy.  And yet, mistakes will be made, and there
>>will be tension between wanting to dedicate those CPUs to IO
>>services and wanting to use them for GP programs: I can easily
>>imagine a paper where someone modifies a scheduler to move IO
>>bound programs to those cores.  Using a different ISA obviates
>>most of that, and provides an (admittedly modest) security benefit.
>
>If there really is such tension, that indicates that such cores would
>be useful for general-purpose use.  That makes the case for using the
>same ISA even stronger.

Incorrect.  It makes it weaker: the whole point is to have
coprocessor cores that are dedicated to IO processing that are
not used for GP compute.  As Mitch said, they're already far
away from DRAM; using them for compute is going to suck.  They
are there to offload IO processing from the big cores; don't
make it easier to abuse their existence.

>As for "mistakes will be made", that also goes the other way: With a
>separate toolchain for the DMA-replacement ISA, there is lots of
>opportunity for mistakes.

I meant runtime mistakes.  You can't run x86 code on them if
they're not an x86 core.

>As for "security benefit", where is that supposed to come from?d

You can't run x86 code on them if they're not an x86 core.

>What
>attack scenario do you have in mind where that "security benefit"
>could materialize?

Someone figures out how to exploit a flaw in the OS whereby some
user thread can execute on an IO coprocessor core, and they
figure out you can speculate on IO transactions, allowing them
to exfiltrate data directly from the IO source.

But, if the OS _cannot_ schedule a user process there, because
it's running an entirely different ISA, then that cannot happen.

>>And if I already have to modify or configure the OS to
>>accommodate the existence of these things in the first place,
>>then accommodating an ISA difference really isn't that much
>>extra work.  The critical observation is that a typical SMP view
>>of the world no longer makes sense for the system architecture,
>>and trying to shoehorn that model onto the hardware reality is
>>just going to cause frustration.
>
>The shared-memory multiprocessing view of the world is very
>successful, while distributed-memory computers are limited to
>supercomputing and other areas where hardware cost still dominates
>over software cost (i.e., where the software crisis has not happened
>yet); as an example of the lack of success of the distributed-memory
>paradigm, take the PlayStation 3; programmers found it too hard to
>work with, so they did not use the hardware well, and eventually Sony
>decided to go for an SMP machine for the PlayStation 4 and 5.

The SoCs you are talking about are already, literally,
"distributed memory computers".  See above about the SMN.

>OTOH, one can say that the way many peripherals work on
>general-purpose computers is more along the lines of
>distributed-memory; but that's probably due to the relative hardware
>and software costs for that peripheral. Sure, the performance
>characteristics are non-uniform (NUMA) in many cases, but 1) caches
>tend to smooth over that, and 2) most of the code is not
>performance-critical, so it just needs to run, which is easier to
>achieve with SMP and harder with distributed memory.
>
>Sure, people have argued for advantages of other models for decades,
>like you do now, but SMP has usually won.

Bluntly, you're making a lot of assumptions and drawing
conclusions from those assumptions.

>>>>>On the other hand, you buy a motherboard with said ASIC core,
>>>>>and you can boot the MB without putting a big chip in the
>>>>>socket--but you may have to deal with scant DRAM since the
>>>>>big centralized chip contains teh memory controller.
>>>>
>>>>A neat hack for bragging rights, but not terribly practical?
>>>
>>>Very practical for updating the firmware of the board to support the
>>>big chip you want to put in the socket (called "BIOS FlashBack" in
>>>connection with AMD big chips).
>>
>>"BIOS", as loaded from the EFS by the ABL on the PSP on EPYC
>>class chips, is usually stored in a QSPI flash on the main
>>board (though starting with Turin you _can_ boot via eSPI).
>>Strictly speaking, you don't _need_ an x86 core to rewrite that.
>>On our machines, we do that from the SP, but we don't use AGESA
>>or UEFI: all of the platform enablement stuff done in PEI and
>>DXE we do directly in the host OS.
>
>EFS?  ABL?  QSPI? eSPI?  PEI?  DXE?

Umm, those are the basic components of the "BIOS" and
surrounding stack as implemented on AMD systems with AGESA and
UEFI.  If you are unaware of what these mean, perhaps you should
spend a little bit of time reading up on how the things you are
frankly making a lot of assumptions about actually work.

In this case, I'm happy to explain a bit, but, frankly, your
response makes it painfully obvious that you really need to
do your own homework here.

* EFS: Embedded File System.  This is the filesystem-like format
  that AMD uses for the data stored in flash that is loaded by
  the PSP.
* ABL: AGESA Boot Loader.  This is a software component that
  runs on the PSP that reads and interprets the "BIOS" image
  in the EFS on flash and loads the x86 code that runs from the
  reset vector into DRAM.
* QSPI: Quad SPI.  This is the physical interface used to access
  the flash that holds the EFS.  It is lined out from the socket
  and thus the CPU so that the PSP can access it.  Other things
  can also access it via a series of muxes; for example, on OCP
  boards like Ruby it's accessable across the DC-SCM connector
  to the BMC so that the BMC can update flash.
* eSPI: enhanced Serial Peripheral Interface.  See the Intel
  spec.  Supported in Genoa, and now in Turin, it's possible to
  boot and AMD EPYC CPU over eSPI.  eSPI is lined out from the
  package. 
* PEI: The "Pre-EFI Initialization" phase of UEFI (Unified
  Extensible Firmware Interface -- the "modern" BIOS).  This is
  the phase where most of the platform enablement stuff is done;
  for example, the PCIe buses are initialized and links are
  trained, for example here:
  https://github.com/openSIL/openSIL/blob/main/xUSL/Mpio/Common/MpioInitFlow.c#L508
* DXE: The "Driver Execution Environment" phase of UEFI, where
  individual _devices_ are found an initialized.
  https://uefi.org/specs/PI/1.9/V1_Overview.html

>Anyway, what you do in your special setup does not detract from the
>fact that being able to flash the firmware without having a working
>main core has turned out to be so useful that out of 218 AM5
>motherboards offered in Austria <https://geizhals.at/?cat=mbam5>, 203
>have that feature.

Sure.  It's useful.  You just don't need to have an x86 core to
do it.

>>Also, on AMD machines, again considering EPYC, it's up to system
>>software running on x86 to direct either the SMU or MPIO to
>>configure DXIO and the rest of the fabric before PCIe link
>>training even begins (releasing PCIe from PERST is done by
>>either the SMU or MPIO, depending on the specific
>>microarchitecture).  Where are these cores, again?  If they're
>>close to the devices, are they in the root complex or on the far
>>side of a bridge?  Can they even talk to the rest of the board?
>
>The core that does the flashing obviously is on the board, not on the
>CPU package (which may be absent).  I do not know where on the board
>it is.

I was referring to Mitch's proposed co-processor cores.  The
point was, that if they're on the distant end of an IO bus that
isn't even configured, and not somehow otherwise connected to
the flash part that holds the BIOS, then they're not going to
help you flash the BIOS without the a socket being populated so
that you've got something that can set up that IO bus so that
those cores can connect to anything useful.  You seem to be
assuming that they're just going to start, in the absense of
the main package, but again, that's a big assumption.

>Typically only one USB port can be used for that, so that may
>indicate that a special path may be used for that without initializing
>all the USB ports and the other hardware that's necessary for that; I
>think that some USB ports are directly connected to the CPU package,
>so those would not work anyway.

Like I said, you could have an electromechanical interlock that
lets the IO coprocessors boot independently and talk directly to
the flash mux if the socket is not populated.  The interface by
which you get the flash image is immaterial at that point.  But
it's not at all clear to me that Mitch had anything like that in
mind.

>>>In a case where we did not have that
>>>feature, and the board did not support the CPU, we had to buy another
>>>CPU to update the firmware
>>><https://www.complang.tuwien.ac.at/anton/asus-p10s-c4l.html>.  That's
>>>especially relevant for AM4 boards, because the support chips make it
>>>hard to use more than 16MB Flash for firmware, but the firmware for
>>>all supported big chips does not fit into 16MB.  However, as the case
>>>mentioned above shows, it's also relevant for Intel boards.
>>
>>You shouldn't need to boot the host operating system to do that,
>>though I get on most consumer-grade machines you'll do it via
>>something that interfaces with AGESA or UEFI.
>
>In the bad old days you had to boot into DOS and run a DOS program for
>flashing the BIOS.  Or worse, Windows; not very useful if you don't
>have Windows installed on the computer (DOS at least could be booted
>from a floppy disk).  My last few experiences in that direction were
>firmware flashing as a "BIOS" feature, and the flashback feature
>(which has it's own problems, because communication with the user is
>limited).
>
>>Most server-grade
>>machines will have a BMC that can do this independently of the
>>main CPU,
>
>And just in another posting you wrote "but not terribly practical?".
>The board I mentioned above where we had to buy a separate CPU for
>flashing mentioned a BMC on the feature list, but when we looked in
>the manual, we found that the BMC is not delivered with the board, but
>has to be bought separately.  There was also no mention that one can
>use the BMC for flashing the BIOS.

Sounds like a problem with the vendor.

>>and I should be clear that I'm discounting use cases
>>for consumer grade boards, where I suspect something like this
>>is less interesting than on server hardware.
>
>What makes you think so?  And what do you mean with "something like
>this"?

"Something like this" meaning a dedicated IO coprocessor on the
far side of the root complex for offloading IO handling.

If you can't see why that might have more applications in the
data center than on the desktop, I don't know what to tell you.
Maybe there are consumer use cases I'm not aware of.

>1) "BIOS flashback" is a mostly-standard feature in AM5 (i.e.,
>consumer-grade) boards.

Of course.

>2) DMA has been a standard feature in various forms on consumer
>hardware since the first IBM PC in 1981, and replacing the DMA engines
>with cores running a general-purpose ISA accessible to OS designers
>will not be limited to servers;

I don't think that was the suggestion.

>if hardware designers and OS
>developers put development time into that, there is no reason for
>limiting that effort to servers.  The existence of the LPE-Cores on
>Meteor Lake (not a server chip) and the in-order ARM cores on various
>smartphone SOCs, the existence of P-Cores and E-Cores on Intel
>consumer-grade CPUs, while the server versions of these CPUs have the
>E-Cores disabled, and the uniformity of cores on the dedicated server
>CPUs indicates that non-uniform cores seem to be hard to sell in
>server space.

The systems you just mentioned were designed for minimizing
power consumption, something that's very useful in the consumer
space (e.g., for battery operated applications, like phones and
laptops) and less useful in the data center space.  However,
having dedicated coprocessors to offload things like IO has a
long history in the mainframe world, but that hasn't filtered
down to the server space in part because it's not well-supported
by software.

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111559 — IP (was: DMA is obsolete)

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2025-05-03 10:50 -0400
SubjectIP (was: DMA is obsolete)
Message-ID<jwvcycp3gfc.fsf-monnier+comp.arch@gnu.org>
In reply to#111558
> When we discuss hardware designs at this level, reusable
> components that go into the system are often referred to as "IP
> cores" or just "IPs".  For example, a UART might be an IP.

FWIW, I hate this terminology which comes from "intellectual
property" since it insists on the value of this only as
a bargaining/power tool rather than for what it actually performs.

> Think of them as building blocks that go into, say, a SoC.

Call them blocks, then.


        Stefan

[toc] | [prev] | [next] | [standalone]


#111560 — Re: IP (was: DMA is obsolete)

FromThomas Koenig <tkoenig@netcologne.de>
Date2025-05-03 15:15 +0000
SubjectRe: IP (was: DMA is obsolete)
Message-ID<vv5buh$3ps1q$1@dont-email.me>
In reply to#111559
Stefan Monnier <monnier@iro.umontreal.ca> schrieb:
>> When we discuss hardware designs at this level, reusable
>> components that go into the system are often referred to as "IP
>> cores" or just "IPs".  For example, a UART might be an IP.
>
> FWIW, I hate this terminology which comes from "intellectual
> property" since it insists on the value of this only as
> a bargaining/power tool rather than for what it actually performs.

It is also a bit misleading. Where I come from, "intellectual
property" refers to patents.

[toc] | [prev] | [next] | [standalone]


#111561 — Re: IP (was: DMA is obsolete)

FromJohn Levine <johnl@taugh.com>
Date2025-05-03 15:46 +0000
SubjectRe: IP (was: DMA is obsolete)
Message-ID<vv5dnq$si6$1@gal.iecc.com>
In reply to#111560
According to Thomas Koenig  <tkoenig@netcologne.de>:
>> FWIW, I hate this terminology which comes from "intellectual
>> property" since it insists on the value of this only as
>> a bargaining/power tool rather than for what it actually performs.
>
>It is also a bit misleading. Where I come from, "intellectual
>property" refers to patents.

Where I come from it also means copyright and trademarks.

I agree that if it's a building block or a core, call it that.

R's,
John


-- 
Regards,
John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
Please consider the environment before reading this e-mail. https://jl.ly

[toc] | [prev] | [next] | [standalone]


#111562 — Re: IP (was: DMA is obsolete)

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-03 16:52 +0000
SubjectRe: IP (was: DMA is obsolete)
Message-ID<vv5hl8$6k0$1@reader1.panix.com>
In reply to#111561
In article <vv5dnq$si6$1@gal.iecc.com>, John Levine  <johnl@taugh.com> wrote:
>According to Thomas Koenig  <tkoenig@netcologne.de>:
>>> FWIW, I hate this terminology which comes from "intellectual
>>> property" since it insists on the value of this only as
>>> a bargaining/power tool rather than for what it actually performs.
>>
>>It is also a bit misleading. Where I come from, "intellectual
>>property" refers to patents.
>
>Where I come from it also means copyright and trademarks.
>
>I agree that if it's a building block or a core, call it that.

You don't have to like the terminology, but that's what is used
across the field.  Sorry if it's uncomfortable, and to be honest
I don't care for it much myself, but them's the breaks.  That's
what AMD calls them, so if we're discussing AMD hardware, it
makes sense to use their terminology.

People in construction probably hate that computer people call
things "blocks" that aren't made of concrete.  I'm sure the
networking people don't like it when the hardware people refer
to "IPs" because of the obvious conflict with TCP/IP.  I'm sure
auto mechanics don't like it when mathematicians talk about
"manifolds" that have nothing to do with car engines.

Ambiguities in terminology abound across fields.  But insisting
that someone not use more or less standard terminology because
it conflicts with something in another field is silly.

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111563 — Re: IP (was: DMA is obsolete)

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-05-03 21:31 +0000
SubjectRe: IP (was: DMA is obsolete)
Message-ID<rWvRP.14996$9zYa.633@fx13.iad>
In reply to#111562
cross@spitfire.i.gajendra.net (Dan Cross) writes:
>In article <vv5dnq$si6$1@gal.iecc.com>, John Levine  <johnl@taugh.com> wrote:
>>According to Thomas Koenig  <tkoenig@netcologne.de>:
>>>> FWIW, I hate this terminology which comes from "intellectual
>>>> property" since it insists on the value of this only as
>>>> a bargaining/power tool rather than for what it actually performs.
>>>
>>>It is also a bit misleading. Where I come from, "intellectual
>>>property" refers to patents.
>>
>>Where I come from it also means copyright and trademarks.
>>
>>I agree that if it's a building block or a core, call it that.
>
>You don't have to like the terminology, but that's what is used
>across the field.  Sorry if it's uncomfortable, and to be honest
>I don't care for it much myself, but them's the breaks.  That's
>what AMD calls them, so if we're discussing AMD hardware, it
>makes sense to use their terminology.

We also call them IP blocks.

[toc] | [prev] | [next] | [standalone]


#111568 — Re: IP

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2025-05-03 23:04 -0400
SubjectRe: IP
Message-ID<jwv4iy113qz.fsf-monnier+comp.arch@gnu.org>
In reply to#111562
> You don't have to like the terminology, but that's what is used
> across the field.  Sorry if it's uncomfortable, and to be honest
> I don't care for it much myself, but them's the breaks.  That's
> what AMD calls them, so if we're discussing AMD hardware, it
> makes sense to use their terminology.
>
> People in construction probably hate that computer people call
> things "blocks" that aren't made of concrete.

That comparison doesn't work, the problem with "IP" is not ambiguity,
but that it's politically/ethically charged.  That's why I hate it:
because I disagree with the politics behind it (and hate the fact "they"
managed to make "everyone" use it, without even paying attention to what
it means).


        Stefan

[toc] | [prev] | [next] | [standalone]


#111570 — Re: IP

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-04 09:56 +0000
SubjectRe: IP
Message-ID<vv7djq$mk0$1@reader1.panix.com>
In reply to#111568
In article <jwv4iy113qz.fsf-monnier+comp.arch@gnu.org>,
Stefan Monnier  <monnier@iro.umontreal.ca> wrote:
>> You don't have to like the terminology, but that's what is used
>> across the field.  Sorry if it's uncomfortable, and to be honest
>> I don't care for it much myself, but them's the breaks.  That's
>> what AMD calls them, so if we're discussing AMD hardware, it
>> makes sense to use their terminology.
>>
>> People in construction probably hate that computer people call
>> things "blocks" that aren't made of concrete.
>
>That comparison doesn't work, the problem with "IP" is not ambiguity,
>but that it's politically/ethically charged.  That's why I hate it:
>because I disagree with the politics behind it (and hate the fact "they"
>managed to make "everyone" use it, without even paying attention to what
>it means).

Well, good luck getting the hardware engineers to change
their nomenclature to suit your sensibilities there.  *shrug*

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111571 — Re: IP

FromThomas Koenig <tkoenig@netcologne.de>
Date2025-05-04 10:17 +0000
SubjectRe: IP
Message-ID<vv7er8$1ob9l$1@dont-email.me>
In reply to#111570
Dan Cross <cross@spitfire.i.gajendra.net> schrieb:
> In article <jwv4iy113qz.fsf-monnier+comp.arch@gnu.org>,
> Stefan Monnier  <monnier@iro.umontreal.ca> wrote:
>>> You don't have to like the terminology, but that's what is used
>>> across the field.  Sorry if it's uncomfortable, and to be honest
>>> I don't care for it much myself, but them's the breaks.  That's
>>> what AMD calls them, so if we're discussing AMD hardware, it
>>> makes sense to use their terminology.
>>>
>>> People in construction probably hate that computer people call
>>> things "blocks" that aren't made of concrete.
>>
>>That comparison doesn't work, the problem with "IP" is not ambiguity,
>>but that it's politically/ethically charged.  That's why I hate it:
>>because I disagree with the politics behind it (and hate the fact "they"
>>managed to make "everyone" use it, without even paying attention to what
>>it means).
>
> Well, good luck getting the hardware engineers to change
> their nomenclature to suit your sensibilities there.  *shrug*

Which begs the quesiton - can an IP with an IP be IP-protected?

The main problem is probably the lack of acronym namespace.  This is
relatively harmless in this context, but can cause serious confusion
when discussing, for example, chemicals with abbreviations.
Serious misunderstanding can ensue, for example when "MC" can
mean either Methyl Chloride (Chloromethane) or Methylene Chloride
(Dichloromethane).

[toc] | [prev] | [next] | [standalone]


#111572 — Re: IP

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-04 18:16 +0000
SubjectRe: IP
Message-ID<5aab775cfbbce60d4811dcebcfd69e33@www.novabbs.org>
In reply to#111571
On Sun, 4 May 2025 10:17:12 +0000, Thomas Koenig wrote:

> Dan Cross <cross@spitfire.i.gajendra.net> schrieb:
>> In article <jwv4iy113qz.fsf-monnier+comp.arch@gnu.org>,
>> Stefan Monnier  <monnier@iro.umontreal.ca> wrote:
>>>> You don't have to like the terminology, but that's what is used
>>>> across the field.  Sorry if it's uncomfortable, and to be honest
>>>> I don't care for it much myself, but them's the breaks.  That's
>>>> what AMD calls them, so if we're discussing AMD hardware, it
>>>> makes sense to use their terminology.
>>>>
>>>> People in construction probably hate that computer people call
>>>> things "blocks" that aren't made of concrete.
>>>
>>>That comparison doesn't work, the problem with "IP" is not ambiguity,
>>>but that it's politically/ethically charged.  That's why I hate it:
>>>because I disagree with the politics behind it (and hate the fact "they"
>>>managed to make "everyone" use it, without even paying attention to what
>>>it means).
>>
>> Well, good luck getting the hardware engineers to change
>> their nomenclature to suit your sensibilities there.  *shrug*
>
> Which begs the quesiton - can an IP with an IP be IP-protected?
>
> The main problem is probably the lack of acronym namespace.  This is
> relatively harmless in this context, but can cause serious confusion
> when discussing, for example, chemicals with abbreviations.
> Serious misunderstanding can ensue, for example when "MC" can
> mean either Methyl Chloride (Chloromethane) or Methylene Chloride
> (Dichloromethane).

DEI stood for "Dale Earnhardt Enterprises" for 2 decades before
the bleeding hearts confiscated it.

[toc] | [prev] | [next] | [standalone]


#111573 — Re: IP

FromBill Findlay <findlaybill@blueyonder.co.uk>
Date2025-05-04 19:37 +0100
SubjectRe: IP
Message-ID<0001HW.2DC7EB760003D89230F37538F@news.individual.net>
In reply to#111572
On 4 May 2025, MitchAlsup1 wrote
(in article<5aab775cfbbce60d4811dcebcfd69e33@www.novabbs.org>):

> DEI stood for "Dale Earnhardt Enterprises" for 2 decades before
> the bleeding hearts confiscated it.

It would seem that those with heart bypasses also cannot spell.

-- 
Bill Findlay

[toc] | [prev] | [next] | [standalone]


#111574 — Re: IP

FromLawrence D'Oliveiro <ldo@nz.invalid>
Date2025-05-04 21:31 +0000
SubjectRe: IP
Message-ID<vv8mbg$2qpl1$4@dont-email.me>
In reply to#111571
On Sun, 4 May 2025 10:17:12 -0000 (UTC), Thomas Koenig wrote:

> Serious misunderstanding can ensue, for example when "MC" can mean
> either Methyl Chloride (Chloromethane) or Methylene Chloride
> (Dichloromethane).

No chemist would refer to either of CH₃Cl or CH₂Cl₂ as “MC”.

[toc] | [prev] | [next] | [standalone]


#111569

FromLawrence D'Oliveiro <ldo@nz.invalid>
Date2025-05-04 06:44 +0000
Message-ID<vv72c8$1dc9f$2@dont-email.me>
In reply to#111555
On Sat, 03 May 2025 06:11:00 GMT, Anton Ertl wrote:

> In any case, I have made arguments why it would make sense to use the
> same ISA as for the OS for programming the cores that replace DMA
> engines.  I will discuss your counterarguments below, but the most
> important one to me seems to be that these cores would cost more than
> with a different ISA.

I think efficiency of implementation is still important enough to outweigh 
that. Case in point: the RP2040 chip from the Raspberry Pi Foundation. 
That has an ARM core, combined with a pair of auxiliary processors not a 
million miles removed from the old mainframe idea of “I/O channels”. Those 
auxiliary processors have sufficient oomph to perform feats such as 
emulating the analog video signal from a 1980s-vintage BBC micro, in real 
time.

Newer versions of the chip have a RISC-V core in there somewhere, too.

[toc] | [prev] | [next] | [standalone]


#111564

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-05-03 21:53 +0000
Message-ID<BfwRP.15340$qm51.4765@fx12.iad>
In reply to#111547
cross@spitfire.i.gajendra.net (Dan Cross) writes:
>In article <2025May2.073450@mips.complang.tuwien.ac.at>,
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>>In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>,
>>>>[snip]
>>>>I suspect the 400 GHz NIC needs a rather BIG core to handle the
>>>>traffic loads.
>>
>>Looking at
>>https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a
>>Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and
>>write <18GB/s even to the L1 cache).  However, Chester Lam notes: "A53
>>offers very low cache bandwidth compared to pretty much any other core
>>we’ve analyzed."  I think, though, that a small in-order core like the
>>A53, but with enough load and store buffering and enough bandwidth to
>>I/O and the memory controller should not have a problem shoveling data
>>from or to a 400Gb/s NIC.  With 128 bits/cycle in each direction one
>>would need one transfer per cycle in each direction at 3125MHz to
>>achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop
>>overhead. 

Running any SoC at 3+gHz requires significant effort in the
back-end and to ensure timing closure on the front end (and
affects floorplanning).  All this adds to the cost to build
and manufacture the chips.

It may be more productive to consider widening the internal
buses to be 256 or 512 bits wide.

> Given that the A53 typically only has 2GHz, supporting 256
>>bits/cycle of transfer width (for load and store instructions, i.e.,
>>along the lines of AVX-256) would be more appropriate.

Better to just use custom hardware for data movement and
add accelerators for specific activities (such as crypto).

Back in the late 70's the Burroughs B4900 used 8085 processor
chips in the I/O controllers (and in the maintenance processor).
The 8085 was primarily concerned with data movement and supported
aggregate bandwidth of 8Mbytes/second between each I/O controller
and memory (there could be up to two IOPs, each responsible for
32 channels).

>>>Eh...Having to jump through hoops here matters less to me for
>>>this kind of use case than if I'm trying to use those cores for
>>>general-purpose compute.
>>
>>I think it's the same thing as Greenspun's tenth rule: First you find
>>that a classical DMA engine is too limiting, then you find that an A53
>>is too limiting, and eventually you find that it would be practical to
>>run the ISA of the main cores.  In particular, it allows you to use
>>the toolchain of the main cores for developing them,
>
>These are issues solveable with the software architecture and
>build system for the host OS.   The important characteristic is
>that the software coupling makes architectural sense, and that
>simply does not require using the same ISA across IPs.

I think there are good reasons to have specialized (or low cost,
e.g. riscv) ancilliary cores in a processor package.   Having
been on both sides of the keep them proprietary vs. fully document
them for the OS folks argument, I remain ambivilent.

There are good reasons for both positions.  The same reasons behind
the MP1.5 spec and UEFI apply in many cases - widening the
ecosystem and 'you-fix-it' capabilities.   On the other hand,
there may be trade secrets, or system security implications that
might preclude full disclosure.   Once a capability is documented
in the PC world, it tends to live forever good or bad (ISA anyone?)
which may limit future choices in the product line (or discommode
customers).


>At work, our service processor (granted, outside of the SoC but
>tightly coupled at the board level) is a Cortex-M7, but we wrote
>the OS for that, 

What, not Zephyr?

> and we control the host OS that runs on x86,
>so the SP and big CPUs can be mutually aware.  Our hardware RoT
>is a smaller Cortex-M.  We don't have a BMC on our boards;
>everything that it does is either done by the SP or built into
>the host OS, both of which are measured by the RoT.
>
>The problem is when such service cores are hidden (as they are
>in the case of the PSP, SMU, MPIO, and similar components, to
>use AMD as the example) and treated like black boxes by
>software.  It's really cool that I can configure the IO crossbar
>in useful way tailored to specific configurations, but it's much
>less cool that I have to do what amounts to an RPC over the SMN
>to some totally undocumented entity somewhere in the SoC to do
>it.  Bluntly, as an OS person, I do not want random bits of code
>running anywhere on my machine that I am not at least aware of
>(yes, this includes firmware blobs on devices).

As a hardware (and long-time OS) person (not necessarily in that order),
I sympathize, but, yet, see above.


>And if I already have to modify or configure the OS to
>accommodate the existence of these things in the first place,
>then accommodating an ISA difference really isn't that much
>extra work.  The critical observation is that a typical SMP view
>of the world no longer makes sense for the system architecture,
>and trying to shoehorn that model onto the hardware reality is
>just going to cause frustration.  Better to acknowledge that the
>

Most of the hardware should be standardized through ACPI calls,
allowing the underlying implementation to vary over time.


<big snip>

>Also, on AMD machines, again considering EPYC, it's up to system
>software running on x86 to direct either the SMU or MPIO to
>configure DXIO and the rest of the fabric before PCIe link
>training even begins (releasing PCIe from PERST is done by
>either the SMU or MPIO, depending on the specific
>microarchitecture).  Where are these cores, again?  If they're
>close to the devices, are they in the root complex or on the far
>side of a bridge?  Can they even talk to the rest of the board?

It's not the core that's proprietary in this case, it's the
intimate knowledge of the mainboard that is required for that
operation - consider address routing - once a function BAR
is programmed, the SoC fabric needs to be configured to route
that range of addresses to the correct PCI controller to
be converted to PCIe TLPs to the target function.

That routing can be incredibly complicated, requiring
very substantial and rather tricky configuration
steps in several related IP blocks as well as the
inter-cpu routing fabric/mesh.   Getting it right
is difficult, and there's really no reason for the
OS level to be aware of it (and given it is highly
mainboard/SoC dependent, just complicates the
operating software when not behind  a standard
configuration mechanism like ACPI.)

[toc] | [prev] | [next] | [standalone]


#111565

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-03 23:02 +0000
Message-ID<17ad830022e847950d47b90da1b555b7@www.novabbs.org>
In reply to#111564
On Sat, 3 May 2025 21:53:37 +0000, Scott Lurndal wrote:

> cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>>
>>>Looking at
>>>https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important, a
>>>Cortex-A53 would not be up to it (at 1896MHz it can read <12GB/s and
>>>write <18GB/s even to the L1 cache).  However, Chester Lam notes: "A53
>>>offers very low cache bandwidth compared to pretty much any other core
>>>we’ve analyzed."  I think, though, that a small in-order core like the
>>>A53, but with enough load and store buffering and enough bandwidth to
>>>I/O and the memory controller should not have a problem shoveling data
>>>from or to a 400Gb/s NIC.  With 128 bits/cycle in each direction one
>>>would need one transfer per cycle in each direction at 3125MHz to
>>>achieve 400Gb/s, or maybe 4GHz for a dual-issue core to allow for loop
>>>overhead.
>
> Running any SoC at 3+gHz requires significant effort in the
> back-end and to ensure timing closure on the front end (and
> affects floorplanning).  All this adds to the cost to build
> and manufacture the chips.
>
> It may be more productive to consider widening the internal
> buses to be 256 or 512 bits wide.

At smaller than 7nm there seems to be little reason the main
interconnect is not cache-line-wide or cache-line-wide in two
directions. Your typical GPU will have 1024 wires into and out
of each shader core and several other big blocks.

Many cache-lines are 512-bits wide (except for IBM at 4096-bits
wide).

[toc] | [prev] | [next] | [standalone]


#111710

Fromcross@spitfire.i.gajendra.net (Dan Cross)
Date2025-05-21 12:36 +0000
Message-ID<100khbs$elq$1@reader1.panix.com>
In reply to#111564
[Apologies for the excessively long time to respond; it's been a
very busy few weeks!]

In article <BfwRP.15340$qm51.4765@fx12.iad>,
Scott Lurndal <slp53@pacbell.net> wrote:
>cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>[snip]
>>These are issues solveable with the software architecture and
>>build system for the host OS.   The important characteristic is
>>that the software coupling makes architectural sense, and that
>>simply does not require using the same ISA across IPs.
>
>I think there are good reasons to have specialized (or low cost,
>e.g. riscv) ancilliary cores in a processor package.   Having
>been on both sides of the keep them proprietary vs. fully document
>them for the OS folks argument, I remain ambivilent.
>
>There are good reasons for both positions.  The same reasons behind
>the MP1.5 spec and UEFI apply in many cases - widening the
>ecosystem and 'you-fix-it' capabilities.   On the other hand,
>there may be trade secrets, or system security implications that
>might preclude full disclosure.   Once a capability is documented
>in the PC world, it tends to live forever good or bad (ISA anyone?)
>which may limit future choices in the product line (or discommode
>customers).

The idea that one introduces interfaces to facilitate change
across multiple independent dimensions is a good one.  But the
interfaces that we have are awful.

>>At work, our service processor (granted, outside of the SoC but
>>tightly coupled at the board level) is a Cortex-M7, but we wrote
>>the OS for that, 
>
>What, not Zephyr?

lol nope. https://hubris.oxide.computer

>> and we control the host OS that runs on x86,
>>so the SP and big CPUs can be mutually aware.  Our hardware RoT
>>is a smaller Cortex-M.  We don't have a BMC on our boards;
>>everything that it does is either done by the SP or built into
>>the host OS, both of which are measured by the RoT.
>>
>>The problem is when such service cores are hidden (as they are
>>in the case of the PSP, SMU, MPIO, and similar components, to
>>use AMD as the example) and treated like black boxes by
>>software.  It's really cool that I can configure the IO crossbar
>>in useful way tailored to specific configurations, but it's much
>>less cool that I have to do what amounts to an RPC over the SMN
>>to some totally undocumented entity somewhere in the SoC to do
>>it.  Bluntly, as an OS person, I do not want random bits of code
>>running anywhere on my machine that I am not at least aware of
>>(yes, this includes firmware blobs on devices).
>
>As a hardware (and long-time OS) person (not necessarily in that order),
>I sympathize, but, yet, see above.
>
>>And if I already have to modify or configure the OS to
>>accommodate the existence of these things in the first place,
>>then accommodating an ISA difference really isn't that much
>>extra work.  The critical observation is that a typical SMP view
>>of the world no longer makes sense for the system architecture,
>>and trying to shoehorn that model onto the hardware reality is
>>just going to cause frustration.  Better to acknowledge that the
>
>Most of the hardware should be standardized through ACPI calls,
>allowing the underlying implementation to vary over time.

I strongly disagree.  ACPI is a disaster; it may be the diaster
we know, but it's a disaster nonetheless.  Plus, with its tight
entanglement with UEFI, it forces you into that world.

Taking a step back, the real desideratea here might be some
well-defined interface that creates a seam between the hardware
and the OS, but UEFI+ACPI ain't it.

Granted, we have the massive luxury of developing the hardware
and software together and in concert at my job.  I recognize
that that is not the common case.

><big snip>
>
>>Also, on AMD machines, again considering EPYC, it's up to system
>>software running on x86 to direct either the SMU or MPIO to
>>configure DXIO and the rest of the fabric before PCIe link
>>training even begins (releasing PCIe from PERST is done by
>>either the SMU or MPIO, depending on the specific
>>microarchitecture).  Where are these cores, again?  If they're
>>close to the devices, are they in the root complex or on the far
>>side of a bridge?  Can they even talk to the rest of the board?
>
>It's not the core that's proprietary in this case, it's the
>intimate knowledge of the mainboard that is required for that
>operation - consider address routing - once a function BAR
>is programmed, the SoC fabric needs to be configured to route
>that range of addresses to the correct PCI controller to
>be converted to PCIe TLPs to the target function.

Fortunately, we designed and built the board ourselves.  :-D

But this is an issue that the OS must contend with anyway;
consider the case of PCIe hotplug: a device newly inserted in a
system must be programmed with a useful BAR, which largely means
that the OS must be capable of allocating physical address space
to that device and setting up those routes.  One might argue
that, perhaps, the OS ought to use some sort of existing
interface (like an AML flow or whatever) to do so, but
that doesn't change the fact that ultimately the responsibility
belongs to the OS.

I suppose another argument might be that system firwmare just
sets up all the root bridge ports with a hunk of address space
and routes to those, and then all the OS has to do is plop a
number into a BAR when a new device shows up, but you're already
half way there, and again, the OS has to contend with
understanding the routing that firmware has set up, which in
turn introduces non-trivial complexity.

In any event, we found that setting up the routes isn't that
onerous; we defined some structs that make up effectively a DSL
that lets us specify these in a pretty declarative way;
comparing with e.g. AGESA, or OpenSIL, they do things in a
pretty similar way.  In either case, it's up to the board
manufacturers to provide all of that data.

>That routing can be incredibly complicated, requiring
>very substantial and rather tricky configuration
>steps in several related IP blocks as well as the
>inter-cpu routing fabric/mesh.   Getting it right
>is difficult, and there's really no reason for the
>OS level to be aware of it (and given it is highly
>mainboard/SoC dependent, just complicates the
>operating software when not behind  a standard
>configuration mechanism like ACPI.)

Again, I disagree.  What is the OS there for, if not to control
and configurethe hardware and provide useful abstractions to
application software?  Hiding the actual hardware configuration
from the OS means that the OS is not, in fact, the final entity
in charge of the hardware.  Having a second operating system, in
parallel, gives me all of the same problems I had with BIOSes.

This opens all sorts of issues with respect to safety and
security, provenance of software, fault management, and
ultimately, control.  It's my machine, I want to run it as I
see fit.

	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111552

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-02 17:40 +0000
Message-ID<859cc676b91aa1173e44acf0f2be636b@www.novabbs.org>
In reply to#111545
On Fri, 2 May 2025 2:15:24 +0000, Dan Cross wrote:

> In article <5a77c46910dd2100886ce6fc44c4c460@www.novabbs.org>,
> MitchAlsup1 <mitchalsup@aol.com> wrote:
>>On Thu, 1 May 2025 13:07:07 +0000, Dan Cross wrote:
>>> In article <da5b3dea460370fc1fe8ad2323da9bc4@www.novabbs.org>,
>>> MitchAlsup1 <mitchalsup@aol.com> wrote:
>>>>On Sat, 26 Apr 2025 17:29:06 +0000, Scott Lurndal wrote:
>>>>[snip]
>>>>Reminds me of trying to sell a micro x86-64 to AMD as a project.
>>>>The µ86 is a small x86-64 core made available as IP in Verilog
>>>>where it has/runs the same ISA as main GBOoO x86, but is placed
>>>>"out in the PCIe" interconnect--performing I/O services topo-
>>>>logically adjacent to the device itself. This allows 1ns access
>>>>latencies to DCRs and performing OS queueing of DPCs,... without
>>>>bothering the GBOoO cores.
>>>>
>>>>AMD didn't buy the arguments.
>>>
>>> I can see it either way; I suppose the argument as to whether I
>>> buy it or not comes down to, "in depends".  How much control do
>>> I, as the OS implementer, have over this core?
>>
>>Other than it being placed "away" from the centralized cores,
>>it runs the same ISA as the main cores has longer latency to
>>coherent memory and shorter latency to device control registers
>>--which is why it is placed close to the device itself:: latency.
>>The big fast centralized core is going to get microsecond latency
>>from MMI/O device whereas ASIC version will have handful of nano-
>>second latencies. So the 5 GHZ core sees ~1 microsecond while the
>>little ASIC sees 10 nanoseconds. ...
>
> Yes, I get the argument for WHY you'd do it, I just want to make
> sure that it's an ordinary core (albeit one that is far away
> from the sockets with the main SoC complexes) that I interact
> with in the usual manner.  Compare to, say, MP1 or MP0 on AMD
> Zen, where it runs its own (proprietary) firmware that I
> interact with via an RPC protocol over an AXI bus, if I interact
> with it at all: most OEMs just punt and run AGESA (we don't).
>
>>> If it is yet another hidden core embedded somewhere deep in the
>>> SoC complex and I can't easily interact with it from the OS,
>>> then no thanks: we've got enough of those between MP0, MP1, MP5,
>>> etc, etc.
>>>
>>> On the other hand, if it's got a "normal" APIC ID, the OS has
>>> control over it like any other LP, and its coherent with the big
>>> cores, then yeah, sign me up: I've been wanting something like
>>> that for a long time now.
>>
>>It is just a core that is cheap enough to put in ASICs, that
>>can offload some I/O burden without you having to do anything
>>other than setting some bits in some CRs so interrupts are
>>routed to this core rather than some more centralized core.
>
> Sounds good.
>
>>> Consider a virtualization application.  A problem with, say,
>>> SR-IOV is that very often the hypervisor wants to interpose some
>>> sort of administrative policy between the virtual function and
>>> whatever it actually corresponds to, but get out of the fast
>>> path for most IO.  This implies a kind of offload architecture
>>> where there's some (presumably software) agent dedicated to
>>> handling IO that can be parameterized with such a policy.  A
>>
>>Interesting:: Could you cite any literature, here !?!
>
> Sure.  This paper is a bit older, but gets at the main points:
> https://www.usenix.org/system/files/conference/nsdi18/nsdi18-firestone.pdf
>
> I don't know if the details are public for similar technologies
> from Amazon or Google.
>
>>> core very close to the device could handle that swimmingly,
>>> though I'm not sure it would be enough to do it at (say) line
>>> rate for a 400Gbps NIC or Gen5 NVMe device.
>>
>>I suspect the 400 GHz NIC needs a rather BIG core to handle the
>>traffic loads.
>
> Indeed.  Part of the challenge for the hyperscalars is in
> meeting that demand while not burning too many host resources,
> which are the thing they're actually selling their customer in
> the first place.  A lot of folks are pushing this off to the NIC
> itself, and I've seen at least one team that implemented NVMe in
> firmware on a 100Gbps NIC, exposed via SR-IOV, as part of a
> disaggregated storage architecture.
>
> Another option is to push this to the switch; things like Intel
> Tofino2 were well-position for this, but of course Intel, in its
> infinite wisdom and vision, canc'ed Tofino.
>
>>> ....but why x86_64?  It strikes me that as long as the _data_
>>> formats vis the software-visible ABI are the same, it doesn't
>>> need to use the same ISA.  In fact, I can see advantages to not
>>> doing so.
>>
>>Having the remote core run the same OS code as every other core
>>means the OS developers have fewer hoops to jump through. Bug-for
>>bug compatibility means that clearing of those CRs just leaves
>>the core out in the periphery idling and bothering no one.
>
> Eh...Having to jump through hoops here matters less to me for
> this kind of use case than if I'm trying to use those cores for
> general-purpose compute.  Having a separate ISA means I cannot
> accidentally run a program meant only for the big cores on the
> IO service processors.  As long as the OS has total control over
> the execution of the core, and it participates in whatever cache
> coherency scheme the rest of the system uses, then the ISA just
> isn't that important.
>
>>On the other hand, you buy a motherboard with said ASIC core,
>>and you can boot the MB without putting a big chip in the
>>socket--but you may have to deal with scant DRAM since the
>>big centralized chip contains teh memory controller.
>
> A neat hack for bragging rights, but not terribly practical?
>
> Anyway, it's a neat idea.  It's very reminiscent of IBM channel
> controllers, in a way.

It is more like the Peripheral Processors of CDC 6600 that run
ISA of a CDC 6600 without as much fancy execution in periphery.

> 	- Dan C.

[toc] | [prev] | [next] | [standalone]


#111557

FromTerje Mathisen <terje.mathisen@tmsw.no>
Date2025-05-03 14:29 +0200
Message-ID<vv527f$3hdc3$1@dont-email.me>
In reply to#111552
MitchAlsup1 wrote:
> On Fri, 2 May 2025 2:15:24 +0000, Dan Cross wrote:
>>> On the other hand, you buy a motherboard with said ASIC core,
>>> and you can boot the MB without putting a big chip in the
>>> socket--but you may have to deal with scant DRAM since the
>>> big centralized chip contains teh memory controller.
>>
>> A neat hack for bragging rights, but not terribly practical?
>>
>> Anyway, it's a neat idea.  It's very reminiscent of IBM channel
>> controllers, in a way.
> 
> It is more like the Peripheral Processors of CDC 6600 that run
> ISA of a CDC 6600 without as much fancy execution in periphery.

Similar timeframe: The ND10 minis were popular in process control, CERN 
bought a brace of them.

When they later came out with the larger ND100 and then ND500 machines, 
the latter had a 100 (or 10?) as a front-end IO processor, partially 
required because the original ND10 came with a very early version of 
SINTRAN os which didn't have proper/complete IO support, so customers 
had written machine code to handle it.

The 500 wasn't machine code compatible, so all such IO routines then had 
to run on the front-end processor.

Terje

-- 
- <Terje.Mathisen at tmsw.no>
"almost all programming can be viewed as an exercise in caching"

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.arch


csiph-web