Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch > #111620 > unrolled thread

The Seymour Cray Era of Supercomputers

Started byThomas Koenig <tkoenig@netcologne.de>
First post2025-05-17 20:00 +0000
Last post2025-05-26 19:27 -0700
Articles 20 on this page of 71 — 20 participants

Back to article view | Back to comp.arch


Contents

  The Seymour Cray Era of Supercomputers Thomas Koenig <tkoenig@netcologne.de> - 2025-05-17 20:00 +0000
    Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-17 21:27 +0000
      Re: The Seymour Cray Era of Supercomputers Thomas Koenig <tkoenig@netcologne.de> - 2025-05-18 05:46 +0000
        Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-18 18:23 +0300
          Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-18 22:02 +0000
          Re: The Seymour Cray Era of Supercomputers quadibloc <quadibloc@gmail.com> - 2025-05-19 01:08 +0000
            Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 01:56 +0000
              Re: The Seymour Cray Era of Supercomputers quadibloc <quadibloc@gmail.com> - 2025-05-19 03:12 +0000
                OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-19 06:22 +0000
                  Re: OoO execution (was: The Seymour Cray Era of Supercomputers) John Levine <johnl@taugh.com> - 2025-05-19 17:10 +0000
                    Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-19 17:46 +0000
                      Re: OoO execution ze@zerandconsulting.com (Ze) - 2025-05-19 19:09 +0000
                        Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 00:04 +0000
                          Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 00:30 +0000
                            Re: OoO execution scott@slp53.sl.home (Scott Lurndal) - 2025-05-20 13:52 +0000
                      Re: OoO execution (was: The Seymour Cray Era of Supercomputers) George Neuner <gneuner2@comcast.net> - 2025-05-21 12:52 -0400
                        Re: OoO execution Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-21 13:14 -0400
                        Re: OoO execution moi <findlaybill@blueyonder.co.uk> - 2025-05-21 18:47 +0100
                  Re: OoO execution EricP <ThatWouldBeTelling@thevillage.com> - 2025-05-19 14:33 -0400
                  Re: OoO execution quadibloc <quadibloc@gmail.com> - 2025-05-19 19:08 +0000
                    Re: OoO execution Terje Mathisen <terje.mathisen@tmsw.no> - 2025-05-19 22:04 +0200
                      Re: OoO execution Michael S <already5chosen@yahoo.com> - 2025-05-19 23:27 +0300
                      Re: OoO execution John Savard <quadibloc@invalid.invalid> - 2025-07-16 14:27 +0000
                        Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-07-16 18:27 +0000
                          Re: OoO execution Stefan Monnier <monnier@iro.umontreal.ca> - 2025-07-16 17:45 -0400
                          Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-07-26 02:45 +0000
                            Re: OoO execution John Savard <quadibloc@invalid.invalid> - 2025-07-31 20:38 +0000
                              Re: OoO execution Michael S <already5chosen@yahoo.com> - 2025-08-01 15:02 +0300
                                Re: OoO execution John Levine <johnl@taugh.com> - 2025-08-01 15:44 +0000
                  Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Michael S <already5chosen@yahoo.com> - 2025-05-19 23:41 +0300
                    Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 00:01 +0000
                    Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-20 21:21 +0000
                      Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Michael S <already5chosen@yahoo.com> - 2025-05-30 13:28 +0300
                        Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Al Kossow <aek@bitsavers.org> - 2025-05-30 10:51 -0700
                          Re: OoO execution (was: The Seymour Cray Era of Supercomputers) scott@slp53.sl.home (Scott Lurndal) - 2025-05-30 19:36 +0000
                            Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-31 07:57 +0000
                          Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-30 22:05 +0000
                            Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-30 23:01 +0000
                        Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-31 08:10 +0000
                  Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Thomas Koenig <tkoenig@netcologne.de> - 2025-05-29 19:02 +0000
                    Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-29 20:06 +0000
                      Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-29 22:20 +0000
                      Re: OoO execution David Schultz <david.schultz@earthlink.net> - 2025-05-29 18:36 -0500
                Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 07:50 +0000
              Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 16:55 +0300
                Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 23:58 +0000
                  Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-20 13:45 +0300
                    Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 16:59 +0000
                    Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-20 19:59 +0000
                      Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 22:48 +0000
                      Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-21 11:21 +0300
                        Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-21 08:44 +0000
                        Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-21 16:09 +0000
                          Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-21 20:11 +0300
                            Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-21 20:04 +0000
                          Re: The Seymour Cray Era of Supercomputers Lars Poulsen <lars@cleo.beagle-ears.com> - 2025-05-25 21:08 +0000
                            Re: 360/44, The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-25 22:50 +0000
      Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-18 11:33 +0300
        Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-18 22:01 +0000
          Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 16:35 +0300
            Re: The Seymour Cray Era of Supercomputers Al Kossow <aek@bitsavers.org> - 2025-05-19 09:49 -0700
            Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-19 18:14 +0000
              Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 23:11 +0300
                Re: The Seymour Cray Era of Supercomputers BGB <cr88192@gmail.com> - 2025-05-20 01:36 -0500
              Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 05:40 +0000
                Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 16:58 +0000
                  Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-21 00:28 +0000
              Re: The Seymour Cray Era of Supercomputers "Brian G. Lucas" <bagel99@gmail.com> - 2025-05-26 11:48 -0500
                Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-26 18:01 +0000
                  Re: The Seymour Cray Era of Supercomputers EricP <ThatWouldBeTelling@thevillage.com> - 2025-05-26 16:52 -0400
                    Re: The Seymour Cray Era of Supercomputers Al Kossow <aek@bitsavers.org> - 2025-05-26 19:27 -0700

Page 1 of 4  [1] 2 3 4  Next page →


#111620 — The Seymour Cray Era of Supercomputers

FromThomas Koenig <tkoenig@netcologne.de>
Date2025-05-17 20:00 +0000
SubjectThe Seymour Cray Era of Supercomputers
Message-ID<100apst$hsll$1@dont-email.me>
I just finished the book above, and it was a very interesting read
(especially since I worked with these machines around the era the
authors describe, from ~1960 to 1996, ended).

Well worth reading.  Architectural details are not the main focus
of the book; rather, it deals with Seymour Cray's work, but also
with the technical reasons why some architectures were more
successful than others, and how these machines were adopted
and used at different customers.

ISBN  979-8400713705.  I bought it on Amazon.

[toc] | [next] | [standalone]


#111621

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-17 21:27 +0000
Message-ID<afa210f16ab3d6795c61787ad914e7ba@www.novabbs.org>
In reply to#111620
Did the book relate the story of why CRAY-1 presented a DC-load to
the power supply:: that is, the ECL gates were all of the form where
they would switch 20 ma into either the true or the complement out-
put and thus have no AC energy at the power supply level ??

During the CDC 7600 reign, when performing vector calculations,
(even though CDC 7600 was not a vector machine, it could stream
calculations through its execution window at impressive rates);
Certain data bit-patterns in CDC 7600 would cause more Gnd bounce
and Vdd drop than the gates cols accommodate and the machine would
take a data-dependent hard crash.

So, Cray got rid of the problem by presenting a DC-load to the
power supply.

[toc] | [prev] | [next] | [standalone]


#111624

FromThomas Koenig <tkoenig@netcologne.de>
Date2025-05-18 05:46 +0000
Message-ID<100bs7t$rna2$1@dont-email.me>
In reply to#111621
MitchAlsup1 <mitchalsup@aol.com> schrieb:
> Did the book relate the story of why CRAY-1 presented a DC-load to
> the power supply:: that is, the ECL gates were all of the form where
> they would switch 20 ma into either the true or the complement out-
> put and thus have no AC energy at the power supply level ??

That they didn't mention.  They stressed his decision to build
a machine which had good all-round performance, unlike the
predecessors like the STAR or the Texas Instruments ASC (which I
had never heard or read of).

There was one part on the Cray-I design that I found weird.  After
writing that individual transistors would have been faster than
integrated circuits, but were chosen for density and manufacture,
they wrote

"Concerning memory, as in the case of the CPU, Cray did not
choose the fastest individual components, which would have
been magnetic cores" due to their limitations in size.

What they also describe well is the tradeoff between different
vector lengths.  Also very interesting is the mechanisms of getting
the machines adopted by different industries.

[toc] | [prev] | [next] | [standalone]


#111627

FromMichael S <already5chosen@yahoo.com>
Date2025-05-18 18:23 +0300
Message-ID<20250518182303.00003542@yahoo.com>
In reply to#111624
On Sun, 18 May 2025 05:46:37 -0000 (UTC)
Thomas Koenig <tkoenig@netcologne.de> wrote:

> MitchAlsup1 <mitchalsup@aol.com> schrieb:
> > Did the book relate the story of why CRAY-1 presented a DC-load to
> > the power supply:: that is, the ECL gates were all of the form where
> > they would switch 20 ma into either the true or the complement out-
> > put and thus have no AC energy at the power supply level ??  
> 
> That they didn't mention.  They stressed his decision to build
> a machine which had good all-round performance, unlike the
> predecessors like the STAR or the Texas Instruments ASC (which I
> had never heard or read of).
>

May be, that aspect of CRAY-1 was different from STAR and ASC, but not
different from CDC 6600 or 7600 or from top models of Cyber-170 series.

> There was one part on the Cray-I design that I found weird.  After
> writing that individual transistors would have been faster than
> integrated circuits, but were chosen for density and manufacture,
> they wrote
> 

That part sounds correct. Logic ICs used in Cray-1 were indeed slower
than contemporary individual transistors.

> "Concerning memory, as in the case of the CPU, Cray did not
> choose the fastest individual components, which would have
> been magnetic cores" due to their limitations in size.
> 

That part does not sound right. CRAY-1  main memory was made of SRAM
with 48 ns access time. That was 4-5 times faster than contemporary
core memories. Plus, it didn't suffer from destructive read.
One part that is true is that faster and less dense memory components
were available, but they were SRAM as well.

> What they also describe well is the tradeoff between different
> vector lengths.  Also very interesting is the mechanisms of getting
> the machines adopted by different industries.



[toc] | [prev] | [next] | [standalone]


#111628

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-18 22:02 +0000
Message-ID<1ff4496d7140f7061cd6d218abdee5be@www.novabbs.org>
In reply to#111627
On Sun, 18 May 2025 15:23:03 +0000, Michael S wrote:

> On Sun, 18 May 2025 05:46:37 -0000 (UTC)
> Thomas Koenig <tkoenig@netcologne.de> wrote:
>
>> MitchAlsup1 <mitchalsup@aol.com> schrieb:
>>> Did the book relate the story of why CRAY-1 presented a DC-load to
>>> the power supply:: that is, the ECL gates were all of the form where
>>> they would switch 20 ma into either the true or the complement out-
>>> put and thus have no AC energy at the power supply level ??
>>
>> That they didn't mention.  They stressed his decision to build
>> a machine which had good all-round performance, unlike the
>> predecessors like the STAR or the Texas Instruments ASC (which I
>> had never heard or read of).
>>
>
> May be, that aspect of CRAY-1 was different from STAR and ASC, but not
> different from CDC 6600 or 7600 or from top models of Cyber-170 series.

CRAY-1 was relatively fast running scalar while ASC and STAR were not.

>> There was one part on the Cray-I design that I found weird.  After
>> writing that individual transistors would have been faster than
>> integrated circuits, but were chosen for density and manufacture,
>> they wrote
>>
>
> That part sounds correct. Logic ICs used in Cray-1 were indeed slower
> than contemporary individual transistors.
>
>> "Concerning memory, as in the case of the CPU, Cray did not
>> choose the fastest individual components, which would have
>> been magnetic cores" due to their limitations in size.
>>
>
> That part does not sound right. CRAY-1  main memory was made of SRAM
> with 48 ns access time. That was 4-5 times faster than contemporary
> core memories. Plus, it didn't suffer from destructive read.
> One part that is true is that faster and less dense memory components
> were available, but they were SRAM as well.
>
>> What they also describe well is the tradeoff between different
>> vector lengths.  Also very interesting is the mechanisms of getting
>> the machines adopted by different industries.

[toc] | [prev] | [next] | [standalone]


#111632

Fromquadibloc <quadibloc@gmail.com>
Date2025-05-19 01:08 +0000
Message-ID<76948d869e78f8cb511809bd159008fd@www.novabbs.com>
In reply to#111627
On Sun, 18 May 2025 15:23:03 +0000, Michael S wrote:

> On Sun, 18 May 2025 05:46:37 -0000 (UTC)
> Thomas Koenig <tkoenig@netcologne.de> wrote:
>
>> MitchAlsup1 <mitchalsup@aol.com> schrieb:
>>> Did the book relate the story of why CRAY-1 presented a DC-load to
>>> the power supply:: that is, the ECL gates were all of the form where
>>> they would switch 20 ma into either the true or the complement out-
>>> put and thus have no AC energy at the power supply level ??
>>
>> That they didn't mention.  They stressed his decision to build
>> a machine which had good all-round performance, unlike the
>> predecessors like the STAR or the Texas Instruments ASC (which I
>> had never heard or read of).
>>
>
> May be, that aspect of CRAY-1 was different from STAR and ASC, but not
> different from CDC 6600 or 7600 or from top models of Cyber-170 series.

Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary
computers. They were not vector machines. The STAR and the ASC were
vector machines - and because, unlike the Cray-I, vectors was the only
thing they were good at, Amdahl's Law killed them.

John Savard

[toc] | [prev] | [next] | [standalone]


#111635

FromLawrence D'Oliveiro <ldo@nz.invalid>
Date2025-05-19 01:56 +0000
Message-ID<100e352$1d61i$3@dont-email.me>
In reply to#111632
On Mon, 19 May 2025 01:08:11 +0000, quadibloc wrote:

> Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary
> computers. They were not vector machines.

They were pipelined machines. They were orders of magnitude faster than 
anything from IBM. They pioneered the very concept of a “supercomputer”.

There was nothing “ordinary” about that.

[toc] | [prev] | [next] | [standalone]


#111637

Fromquadibloc <quadibloc@gmail.com>
Date2025-05-19 03:12 +0000
Message-ID<e5fc3f66c40e74c1cf09ba5ed5a53c14@www.novabbs.com>
In reply to#111635
On Mon, 19 May 2025 1:56:50 +0000, Lawrence D'Oliveiro wrote:

> On Mon, 19 May 2025 01:08:11 +0000, quadibloc wrote:
>
>> Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary
>> computers. They were not vector machines.
>
> They were pipelined machines. They were orders of magnitude faster than
> anything from IBM. They pioneered the very concept of a “supercomputer”.
>
> There was nothing “ordinary” about that.

Yes, that is a fair comment. Eventually, IBM caught up with the Control
Data 6600 by perfecting pipelining in the IBM 360/91, and then combining
it with cache in the 360/195. From the Pentium II onwards, that's the
way computers are made nowadays.

I didn't mean to belittle the 6600, but simply to note that it lacked
the additional speedup that you get from having a vector machine.
Whereas the STAR-100 and the ASC had the opposite fault: having vectors
was all those machines had going for them, while their scalar portions,
unlike that of the 6600, were very definitely ordinary - and so Amdahl's
Law bit them, as I noted.

John Savard

[toc] | [prev] | [next] | [standalone]


#111639 — OoO execution (was: The Seymour Cray Era of Supercomputers)

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2025-05-19 06:22 +0000
SubjectOoO execution (was: The Seymour Cray Era of Supercomputers)
Message-ID<2025May19.082242@mips.complang.tuwien.ac.at>
In reply to#111637
quadibloc <quadibloc@gmail.com> writes:
>Eventually, IBM caught up with the Control
>Data 6600 by perfecting pipelining in the IBM 360/91, and then combining
>it with cache in the 360/195. From the Pentium II onwards, that's the
>way computers are made nowadays.

Pipelining and caches are already used on the MIPS R2000 in 1986, and
the 486 in 1989.

You are probably thinking of OoO Execution, where people usually write
as if the Tomasulo algorithm of the 360/91 as implemented the modern
concept of OoO execution.  But the 360/91 only did OoO for FP, did not
support branch prediction, had imprecise exceptions, and the Tomasulo
algorithm was used primarily as a workaround for the dearth of FP
registers in the S/360.

The innovation that made OoO execution generally usable rather than a
publicity stunt like the 360/91 is the reorder buffer (ROB), which allows to
retire the instructions in-order, and to cancel speculatively
"executed" instructions after an exception or branch misprediction.

The Pentium Pro (introduced 1995-11-01), HP PA-8000 (introduced
1995-11-02), and MIPS R10000 (introduced 1996-01) are the first
microprocessors which have full-blown OoO execution.

But as someone pointed out to me, IBM has implemented OoO execution
between the 370/195 and the Pentium Pro: The ES/9000 models 900 and
820 (shipping from September 1991) "were the first models with
out-of-order execution since the System/370-195 of 1973. However
unlike the old S/360-91-derived systems, the models 900 and 820 had
full out-of-order execution for both integer and floating-point units,
with precise exception handling, and a fully superscalar pipeline."
<https://en.wikipedia.org/wiki/IBM_System/390#ES/9000>.  So apparently
they had a ROB, and AFAIK were the first machines to have one.  These
models also had a branch target buffer; the article does not mention
branch prediction proper, but given a ROB and a branch target buffer,
it would be surprising if they did not predict branches.

So who came up with the concept of the ROB?  I recently looked at one
of the HPS papers (Hwu, Patt, Shebanov on a High Performance Substrate
for the VAX from the mid-late 80s) again, and there was no ROB in that
paper.  I did not revisit their later papers whether they had it
there.  So apparently ROBs were not known in the mid-1980s in
academia, and in 1991 there was hardware with a ROB commercially
available, and a few years later it appeared in microprocessors.

I wonder how early and how much IBM talked about their ES/9000 OoO
implementation and features, but that may have inspired the teams at
Intel, HP and SGI; or maybe there was an ealier source that inspired
them all, but only in 1995/1996 the number of transistors on a chip
was enough to do OoO on a microprocessor.

Ironically, in the transition to CMOS (i.e., microprocessors) IBM
mainframe processors regressed back to in-order (and, I think,
single-issue) again (but with higher clock rates), and in the early
2000s they looked pretty outdated to me.  In the meantime they have
re-progressed to OoO again AFAIK.

Back to OoO: it's interesting that Tomasulo and the 360/91 are
mentioned often, but the ROB and its inventor(s?), which are at least
as important for the success of OoO execution, isn't.

- anton
-- 
'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
  Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>

[toc] | [prev] | [next] | [standalone]


#111644 — Re: OoO execution (was: The Seymour Cray Era of Supercomputers)

FromJohn Levine <johnl@taugh.com>
Date2025-05-19 17:10 +0000
SubjectRe: OoO execution (was: The Seymour Cray Era of Supercomputers)
Message-ID<100fomr$n4q$1@gal.iecc.com>
In reply to#111639
It appears that Anton Ertl <anton@mips.complang.tuwien.ac.at> said:
>quadibloc <quadibloc@gmail.com> writes:
>>Eventually, IBM caught up with the Control
>>Data 6600 by perfecting pipelining in the IBM 360/91, and then combining
>>it with cache in the 360/195. From the Pentium II onwards, that's the
>>way computers are made nowadays.
>
>Pipelining and caches are already used on the MIPS R2000 in 1986, and
>the 486 in 1989.
>
>You are probably thinking of OoO Execution, where people usually write
>as if the Tomasulo algorithm of the 360/91 as implemented the modern
>concept of OoO execution.  But the 360/91 only did OoO for FP, did not
>support branch prediction, had imprecise exceptions, and the Tomasulo
>algorithm was used primarily as a workaround for the dearth of FP
>registers in the S/360.

The 360/91 had primitive branch prediction in "loop mode".  It had an
eight doublewprd instruction queue (which it confusingly called a stack.)
If a program did a backward branch of less than eight doublewords, it'd
stop prefetching and execute out of the queue until the program fell or
branched out.  It was occasionally worth tweaking assembly code to get
a loop to start on a doubleword boundary (the CNOP assembler op) so it'd
fit and run in loop mode.

-- 
Regards,
John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies",
Please consider the environment before reading this e-mail. https://jl.ly

[toc] | [prev] | [next] | [standalone]


#111645 — Re: OoO execution (was: The Seymour Cray Era of Supercomputers)

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2025-05-19 17:46 +0000
SubjectRe: OoO execution (was: The Seymour Cray Era of Supercomputers)
Message-ID<2025May19.194645@mips.complang.tuwien.ac.at>
In reply to#111644
John Levine <johnl@taugh.com> writes:
>The 360/91 had primitive branch prediction in "loop mode".  It had an
>eight doublewprd instruction queue (which it confusingly called a stack.)
>If a program did a backward branch of less than eight doublewords, it'd
>stop prefetching and execute out of the queue until the program fell or
>branched out.

The 68010 had a similar feature (with a smaller buffer), but I don't
think one would call it branch prediction.  In any case, I meant
speculative execution based on branch prediction (but did not write it
that way), and the 360/91 did not do speculative execution AFAIK.

- anton
-- 
'Anyone trying for "industrial quality" ISA should avoid undefined behavior.'
  Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>

[toc] | [prev] | [next] | [standalone]


#111648 — Re: OoO execution

Fromze@zerandconsulting.com (Ze)
Date2025-05-19 19:09 +0000
SubjectRe: OoO execution
Message-ID<37e1b146f23e2bebfec119d47fad36e4@www.novabbs.com>
In reply to#111645
Wasn't one of the earliest forms of branch prediction the simple
heuristic of always taking it in one direction and not taking it in the
other direction , I seem to remember that being the case for some of the
early pipelined microprocessors. I believe it was called static branch
prediction compared to the more modern dynamic branch prediction.

Nicholas (Nick) King

--

[toc] | [prev] | [next] | [standalone]


#111662 — Re: OoO execution

FromLawrence D'Oliveiro <ldo@nz.invalid>
Date2025-05-20 00:04 +0000
SubjectRe: OoO execution
Message-ID<100ggtj$1sbnn$4@dont-email.me>
In reply to#111648
On Mon, 19 May 2025 19:09:12 +0000, Ze wrote:

> Wasn't one of the earliest forms of branch prediction the simple
> heuristic of always taking it in one direction and not taking it in the
> other direction , I seem to remember that being the case for some of the
> early pipelined microprocessors. I believe it was called static branch
> prediction compared to the more modern dynamic branch prediction.

The simple heuristic I remember was to assume that backward branches would 
be more likely to be taken than not (on the grounds that they were 
probably loops) while forward ones would more likely not be taken (I guess 
as an excuse for not disturbing the pipeline too much).

[toc] | [prev] | [next] | [standalone]


#111663 — Re: OoO execution

Frommitchalsup@aol.com (MitchAlsup1)
Date2025-05-20 00:30 +0000
SubjectRe: OoO execution
Message-ID<c24984395cc6cb02482559555d959e63@www.novabbs.org>
In reply to#111662
On Tue, 20 May 2025 0:04:03 +0000, Lawrence D'Oliveiro wrote:

> On Mon, 19 May 2025 19:09:12 +0000, Ze wrote:
>
>> Wasn't one of the earliest forms of branch prediction the simple
>> heuristic of always taking it in one direction and not taking it in the
>> other direction , I seem to remember that being the case for some of the
>> early pipelined microprocessors. I believe it was called static branch
>> prediction compared to the more modern dynamic branch prediction.
>
> The simple heuristic I remember was to assume that backward branches
> would
> be more likely to be taken than not (on the grounds that they were
> probably loops) while forward ones would more likely not be taken (I
> guess
> as an excuse for not disturbing the pipeline too much).

CDC 7600 used this scheme. Backwards taken, forwards not-taken.
Was about 70% accurate for essentially zero storage and 1 (or few)
gates.

This scheme might have been limited in scope (backwards into the
instruction stack was predicted taken, farther than stack was
predicted not-taken:: I don't remember exactly.

[toc] | [prev] | [next] | [standalone]


#111671 — Re: OoO execution

Fromscott@slp53.sl.home (Scott Lurndal)
Date2025-05-20 13:52 +0000
SubjectRe: OoO execution
Message-ID<lO%WP.57064$RXsc.6962@fx36.iad>
In reply to#111663
mitchalsup@aol.com (MitchAlsup1) writes:
>On Tue, 20 May 2025 0:04:03 +0000, Lawrence D'Oliveiro wrote:
>
>> On Mon, 19 May 2025 19:09:12 +0000, Ze wrote:
>>
>>> Wasn't one of the earliest forms of branch prediction the simple
>>> heuristic of always taking it in one direction and not taking it in the
>>> other direction , I seem to remember that being the case for some of the
>>> early pipelined microprocessors. I believe it was called static branch
>>> prediction compared to the more modern dynamic branch prediction.
>>
>> The simple heuristic I remember was to assume that backward branches
>> would
>> be more likely to be taken than not (on the grounds that they were
>> probably loops) while forward ones would more likely not be taken (I
>> guess
>> as an excuse for not disturbing the pipeline too much).
>
>CDC 7600 used this scheme. Backwards taken, forwards not-taken.
>Was about 70% accurate for essentially zero storage and 1 (or few)
>gates.

Burroughs B4900 re-wrote the branch opcode on each branch to reflect
the last two taken vs. not-taken choices.  There were four opcodes
for each type of branch - taken/taken, taken/not-taken, not-taken/taken
and not-taken/not-taken.

[toc] | [prev] | [next] | [standalone]


#111718 — Re: OoO execution (was: The Seymour Cray Era of Supercomputers)

FromGeorge Neuner <gneuner2@comcast.net>
Date2025-05-21 12:52 -0400
SubjectRe: OoO execution (was: The Seymour Cray Era of Supercomputers)
Message-ID<a70s2kdrpr4i3437u51ekebln93l397gfr@4ax.com>
In reply to#111645
On Mon, 19 May 2025 17:46:45 GMT, anton@mips.complang.tuwien.ac.at
(Anton Ertl) wrote:

>John Levine <johnl@taugh.com> writes:
>>The 360/91 had primitive branch prediction in "loop mode".  It had an
>>eight doublewprd instruction queue (which it confusingly called a stack.)
>>If a program did a backward branch of less than eight doublewords, it'd
>>stop prefetching and execute out of the queue until the program fell or
>>branched out.
>
>The 68010 had a similar feature (with a smaller buffer), but I don't
>think one would call it branch prediction.  In any case, I meant
>speculative execution based on branch prediction (but did not write it
>that way), and the 360/91 did not do speculative execution AFAIK.
>
>- anton

Most DSPs have some kind of "loop buffer" from which they can execute
without fetching code from memory.

[toc] | [prev] | [next] | [standalone]


#111721 — Re: OoO execution

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2025-05-21 13:14 -0400
SubjectRe: OoO execution
Message-ID<jwv34cxvqdp.fsf-monnier+comp.arch@gnu.org>
In reply to#111718
> Most DSPs have some kind of "loop buffer" from which they can execute
> without fetching code from memory.

And Mitch's My 66000 `VEC` instruction takes the idea a step further.


        Stefan

[toc] | [prev] | [next] | [standalone]


#111722 — Re: OoO execution

Frommoi <findlaybill@blueyonder.co.uk>
Date2025-05-21 18:47 +0100
SubjectRe: OoO execution
Message-ID<m96ht9Fpn7iU1@mid.individual.net>
In reply to#111718
On 21/05/2025 17:52, George Neuner wrote:
> On Mon, 19 May 2025 17:46:45 GMT, anton@mips.complang.tuwien.ac.at
> (Anton Ertl) wrote:
> 
>> John Levine <johnl@taugh.com> writes:
>>> The 360/91 had primitive branch prediction in "loop mode".  It had an
>>> eight doublewprd instruction queue (which it confusingly called a stack.)
>>> If a program did a backward branch of less than eight doublewords, it'd
>>> stop prefetching and execute out of the queue until the program fell or
>>> branched out.
>>
>> The 68010 had a similar feature (with a smaller buffer), but I don't
>> think one would call it branch prediction.  In any case, I meant
>> speculative execution based on branch prediction (but did not write it
>> that way), and the 360/91 did not do speculative execution AFAIK.
>>
>> - anton
> 
> Most DSPs have some kind of "loop buffer" from which they can execute
> without fetching code from memory.

The Ferranti Atlas 2 and the EE KDF9 are both prior art.

-- 
Bill F.

[toc] | [prev] | [next] | [standalone]


#111647 — Re: OoO execution

FromEricP <ThatWouldBeTelling@thevillage.com>
Date2025-05-19 14:33 -0400
SubjectRe: OoO execution
Message-ID<1RKWP.151744$rkV6.62600@fx46.iad>
In reply to#111639
Anton Ertl wrote:
> quadibloc <quadibloc@gmail.com> writes:
>> Eventually, IBM caught up with the Control
>> Data 6600 by perfecting pipelining in the IBM 360/91, and then combining
>> it with cache in the 360/195. From the Pentium II onwards, that's the
>> way computers are made nowadays.
> 
> Pipelining and caches are already used on the MIPS R2000 in 1986, and
> the 486 in 1989.
> 
> You are probably thinking of OoO Execution, where people usually write
> as if the Tomasulo algorithm of the 360/91 as implemented the modern
> concept of OoO execution.  But the 360/91 only did OoO for FP, did not
> support branch prediction, had imprecise exceptions, and the Tomasulo
> algorithm was used primarily as a workaround for the dearth of FP
> registers in the S/360.
> 
> The innovation that made OoO execution generally usable rather than a
> publicity stunt like the 360/91 is the reorder buffer (ROB), which allows to
> retire the instructions in-order, and to cancel speculatively
> "executed" instructions after an exception or branch misprediction.
> 
> The Pentium Pro (introduced 1995-11-01), HP PA-8000 (introduced
> 1995-11-02), and MIPS R10000 (introduced 1996-01) are the first
> microprocessors which have full-blown OoO execution.
> 
> But as someone pointed out to me, IBM has implemented OoO execution
> between the 370/195 and the Pentium Pro: The ES/9000 models 900 and
> 820 (shipping from September 1991) "were the first models with
> out-of-order execution since the System/370-195 of 1973. However
> unlike the old S/360-91-derived systems, the models 900 and 820 had
> full out-of-order execution for both integer and floating-point units,
> with precise exception handling, and a fully superscalar pipeline."
> <https://en.wikipedia.org/wiki/IBM_System/390#ES/9000>.  So apparently
> they had a ROB, and AFAIK were the first machines to have one.  These
> models also had a branch target buffer; the article does not mention
> branch prediction proper, but given a ROB and a branch target buffer,
> it would be surprising if they did not predict branches.
> 
> So who came up with the concept of the ROB?  I recently looked at one
> of the HPS papers (Hwu, Patt, Shebanov on a High Performance Substrate
> for the VAX from the mid-late 80s) again, and there was no ROB in that
> paper.  I did not revisit their later papers whether they had it
> there.  So apparently ROBs were not known in the mid-1980s in
> academia, and in 1991 there was hardware with a ROB commercially
> available, and a few years later it appeared in microprocessors.

There were a number of papers that circled around the various ideas.
"Decoupled Access Execute Computer Architectures" uses queues to link
the hardware modules together.
"Implementing Precise Interrupts in Pipelined Processors" first mentions
the ROB but doesn't have a renamer and limited OoO ability.
HPS has rename, reservation stations, and multiple FU but no ROB.

I don't know in what machine all the pieces came together at once
but it looks like about 1986 they figured out to use multiple pipelines
AND rename AND future file AND a ROB AND reservation stations AND multiple
function units AND forwarding buses.

Decoupled Access Execute Computer Architectures,
James E. Smith, 1982

Instruction Issue Logic in Pipelined Supercomputers
Shlomo Weiss, James E Smith, 1984

Implementing Precise Interrupts in Pipelined Processors,
James E. Smith, A. R. Pleszkun, 1985

HPS - A New Microarchitecture Rationale And Introduction,
Yale N. Patt, Wen-mei Hwu, and Michael Shebanow, 1985

> I wonder how early and how much IBM talked about their ES/9000 OoO
> implementation and features, but that may have inspired the teams at
> Intel, HP and SGI; or maybe there was an ealier source that inspired
> them all, but only in 1995/1996 the number of transistors on a chip
> was enough to do OoO on a microprocessor.
> 
> Ironically, in the transition to CMOS (i.e., microprocessors) IBM
> mainframe processors regressed back to in-order (and, I think,
> single-issue) again (but with higher clock rates), and in the early
> 2000s they looked pretty outdated to me.  In the meantime they have
> re-progressed to OoO again AFAIK.
> 
> Back to OoO: it's interesting that Tomasulo and the 360/91 are
> mentioned often, but the ROB and its inventor(s?), which are at least
> as important for the success of OoO execution, isn't.
> 
> - anton

[toc] | [prev] | [next] | [standalone]


#111649 — Re: OoO execution

Fromquadibloc <quadibloc@gmail.com>
Date2025-05-19 19:08 +0000
SubjectRe: OoO execution
Message-ID<8cbd51d6e62b575d9c2bf6b8cfa684af@www.novabbs.com>
In reply to#111639
On Mon, 19 May 2025 6:22:42 +0000, Anton Ertl wrote:

> You are probably thinking of OoO Execution, where people usually write
> as if the Tomasulo algorithm of the 360/91 as implemented the modern
> concept of OoO execution.  But the 360/91 only did OoO for FP, did not
> support branch prediction, had imprecise exceptions, and the Tomasulo
> algorithm was used primarily as a workaround for the dearth of FP
> registers in the S/360.

Yes, I was thinking of OoO execution, as opposed to other forms of
pipelining - basic pipelining was used in the 7094 II and even the 6502.

The Pentium II (and Pentium Pro) also only used OoO for floating-point,
while the 68050 only used OoO for integers!

It's true the 360, with only four floating-point registers, had a dearth
of them, but since having lots of registers was a way that RISC tried to
avoid the need for OoO, I would not say that this invalidated the use of
OoO on the 360/195.

John Savard

[toc] | [prev] | [next] | [standalone]


Page 1 of 4  [1] 2 3 4  Next page →

Back to top | Article view | comp.arch


csiph-web