Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.arch > #111620 > unrolled thread
| Started by | Thomas Koenig <tkoenig@netcologne.de> |
|---|---|
| First post | 2025-05-17 20:00 +0000 |
| Last post | 2025-05-26 19:27 -0700 |
| Articles | 20 on this page of 71 — 20 participants |
Back to article view | Back to comp.arch
The Seymour Cray Era of Supercomputers Thomas Koenig <tkoenig@netcologne.de> - 2025-05-17 20:00 +0000
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-17 21:27 +0000
Re: The Seymour Cray Era of Supercomputers Thomas Koenig <tkoenig@netcologne.de> - 2025-05-18 05:46 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-18 18:23 +0300
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-18 22:02 +0000
Re: The Seymour Cray Era of Supercomputers quadibloc <quadibloc@gmail.com> - 2025-05-19 01:08 +0000
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 01:56 +0000
Re: The Seymour Cray Era of Supercomputers quadibloc <quadibloc@gmail.com> - 2025-05-19 03:12 +0000
OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-19 06:22 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) John Levine <johnl@taugh.com> - 2025-05-19 17:10 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-19 17:46 +0000
Re: OoO execution ze@zerandconsulting.com (Ze) - 2025-05-19 19:09 +0000
Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 00:04 +0000
Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 00:30 +0000
Re: OoO execution scott@slp53.sl.home (Scott Lurndal) - 2025-05-20 13:52 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) George Neuner <gneuner2@comcast.net> - 2025-05-21 12:52 -0400
Re: OoO execution Stefan Monnier <monnier@iro.umontreal.ca> - 2025-05-21 13:14 -0400
Re: OoO execution moi <findlaybill@blueyonder.co.uk> - 2025-05-21 18:47 +0100
Re: OoO execution EricP <ThatWouldBeTelling@thevillage.com> - 2025-05-19 14:33 -0400
Re: OoO execution quadibloc <quadibloc@gmail.com> - 2025-05-19 19:08 +0000
Re: OoO execution Terje Mathisen <terje.mathisen@tmsw.no> - 2025-05-19 22:04 +0200
Re: OoO execution Michael S <already5chosen@yahoo.com> - 2025-05-19 23:27 +0300
Re: OoO execution John Savard <quadibloc@invalid.invalid> - 2025-07-16 14:27 +0000
Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-07-16 18:27 +0000
Re: OoO execution Stefan Monnier <monnier@iro.umontreal.ca> - 2025-07-16 17:45 -0400
Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-07-26 02:45 +0000
Re: OoO execution John Savard <quadibloc@invalid.invalid> - 2025-07-31 20:38 +0000
Re: OoO execution Michael S <already5chosen@yahoo.com> - 2025-08-01 15:02 +0300
Re: OoO execution John Levine <johnl@taugh.com> - 2025-08-01 15:44 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Michael S <already5chosen@yahoo.com> - 2025-05-19 23:41 +0300
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 00:01 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-20 21:21 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Michael S <already5chosen@yahoo.com> - 2025-05-30 13:28 +0300
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Al Kossow <aek@bitsavers.org> - 2025-05-30 10:51 -0700
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) scott@slp53.sl.home (Scott Lurndal) - 2025-05-30 19:36 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-31 07:57 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-30 22:05 +0000
Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-30 23:01 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2025-05-31 08:10 +0000
Re: OoO execution (was: The Seymour Cray Era of Supercomputers) Thomas Koenig <tkoenig@netcologne.de> - 2025-05-29 19:02 +0000
Re: OoO execution mitchalsup@aol.com (MitchAlsup1) - 2025-05-29 20:06 +0000
Re: OoO execution Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-29 22:20 +0000
Re: OoO execution David Schultz <david.schultz@earthlink.net> - 2025-05-29 18:36 -0500
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 07:50 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 16:55 +0300
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-19 23:58 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-20 13:45 +0300
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 16:59 +0000
Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-20 19:59 +0000
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 22:48 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-21 11:21 +0300
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-21 08:44 +0000
Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-21 16:09 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-21 20:11 +0300
Re: The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-21 20:04 +0000
Re: The Seymour Cray Era of Supercomputers Lars Poulsen <lars@cleo.beagle-ears.com> - 2025-05-25 21:08 +0000
Re: 360/44, The Seymour Cray Era of Supercomputers John Levine <johnl@taugh.com> - 2025-05-25 22:50 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-18 11:33 +0300
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-18 22:01 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 16:35 +0300
Re: The Seymour Cray Era of Supercomputers Al Kossow <aek@bitsavers.org> - 2025-05-19 09:49 -0700
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-19 18:14 +0000
Re: The Seymour Cray Era of Supercomputers Michael S <already5chosen@yahoo.com> - 2025-05-19 23:11 +0300
Re: The Seymour Cray Era of Supercomputers BGB <cr88192@gmail.com> - 2025-05-20 01:36 -0500
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-20 05:40 +0000
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-20 16:58 +0000
Re: The Seymour Cray Era of Supercomputers Lawrence D'Oliveiro <ldo@nz.invalid> - 2025-05-21 00:28 +0000
Re: The Seymour Cray Era of Supercomputers "Brian G. Lucas" <bagel99@gmail.com> - 2025-05-26 11:48 -0500
Re: The Seymour Cray Era of Supercomputers mitchalsup@aol.com (MitchAlsup1) - 2025-05-26 18:01 +0000
Re: The Seymour Cray Era of Supercomputers EricP <ThatWouldBeTelling@thevillage.com> - 2025-05-26 16:52 -0400
Re: The Seymour Cray Era of Supercomputers Al Kossow <aek@bitsavers.org> - 2025-05-26 19:27 -0700
Page 1 of 4 [1] 2 3 4 Next page →
| From | Thomas Koenig <tkoenig@netcologne.de> |
|---|---|
| Date | 2025-05-17 20:00 +0000 |
| Subject | The Seymour Cray Era of Supercomputers |
| Message-ID | <100apst$hsll$1@dont-email.me> |
I just finished the book above, and it was a very interesting read (especially since I worked with these machines around the era the authors describe, from ~1960 to 1996, ended). Well worth reading. Architectural details are not the main focus of the book; rather, it deals with Seymour Cray's work, but also with the technical reasons why some architectures were more successful than others, and how these machines were adopted and used at different customers. ISBN 979-8400713705. I bought it on Amazon.
[toc] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-05-17 21:27 +0000 |
| Message-ID | <afa210f16ab3d6795c61787ad914e7ba@www.novabbs.org> |
| In reply to | #111620 |
Did the book relate the story of why CRAY-1 presented a DC-load to the power supply:: that is, the ECL gates were all of the form where they would switch 20 ma into either the true or the complement out- put and thus have no AC energy at the power supply level ?? During the CDC 7600 reign, when performing vector calculations, (even though CDC 7600 was not a vector machine, it could stream calculations through its execution window at impressive rates); Certain data bit-patterns in CDC 7600 would cause more Gnd bounce and Vdd drop than the gates cols accommodate and the machine would take a data-dependent hard crash. So, Cray got rid of the problem by presenting a DC-load to the power supply.
[toc] | [prev] | [next] | [standalone]
| From | Thomas Koenig <tkoenig@netcologne.de> |
|---|---|
| Date | 2025-05-18 05:46 +0000 |
| Message-ID | <100bs7t$rna2$1@dont-email.me> |
| In reply to | #111621 |
MitchAlsup1 <mitchalsup@aol.com> schrieb: > Did the book relate the story of why CRAY-1 presented a DC-load to > the power supply:: that is, the ECL gates were all of the form where > they would switch 20 ma into either the true or the complement out- > put and thus have no AC energy at the power supply level ?? That they didn't mention. They stressed his decision to build a machine which had good all-round performance, unlike the predecessors like the STAR or the Texas Instruments ASC (which I had never heard or read of). There was one part on the Cray-I design that I found weird. After writing that individual transistors would have been faster than integrated circuits, but were chosen for density and manufacture, they wrote "Concerning memory, as in the case of the CPU, Cray did not choose the fastest individual components, which would have been magnetic cores" due to their limitations in size. What they also describe well is the tradeoff between different vector lengths. Also very interesting is the mechanisms of getting the machines adopted by different industries.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2025-05-18 18:23 +0300 |
| Message-ID | <20250518182303.00003542@yahoo.com> |
| In reply to | #111624 |
On Sun, 18 May 2025 05:46:37 -0000 (UTC) Thomas Koenig <tkoenig@netcologne.de> wrote: > MitchAlsup1 <mitchalsup@aol.com> schrieb: > > Did the book relate the story of why CRAY-1 presented a DC-load to > > the power supply:: that is, the ECL gates were all of the form where > > they would switch 20 ma into either the true or the complement out- > > put and thus have no AC energy at the power supply level ?? > > That they didn't mention. They stressed his decision to build > a machine which had good all-round performance, unlike the > predecessors like the STAR or the Texas Instruments ASC (which I > had never heard or read of). > May be, that aspect of CRAY-1 was different from STAR and ASC, but not different from CDC 6600 or 7600 or from top models of Cyber-170 series. > There was one part on the Cray-I design that I found weird. After > writing that individual transistors would have been faster than > integrated circuits, but were chosen for density and manufacture, > they wrote > That part sounds correct. Logic ICs used in Cray-1 were indeed slower than contemporary individual transistors. > "Concerning memory, as in the case of the CPU, Cray did not > choose the fastest individual components, which would have > been magnetic cores" due to their limitations in size. > That part does not sound right. CRAY-1 main memory was made of SRAM with 48 ns access time. That was 4-5 times faster than contemporary core memories. Plus, it didn't suffer from destructive read. One part that is true is that faster and less dense memory components were available, but they were SRAM as well. > What they also describe well is the tradeoff between different > vector lengths. Also very interesting is the mechanisms of getting > the machines adopted by different industries.
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-05-18 22:02 +0000 |
| Message-ID | <1ff4496d7140f7061cd6d218abdee5be@www.novabbs.org> |
| In reply to | #111627 |
On Sun, 18 May 2025 15:23:03 +0000, Michael S wrote: > On Sun, 18 May 2025 05:46:37 -0000 (UTC) > Thomas Koenig <tkoenig@netcologne.de> wrote: > >> MitchAlsup1 <mitchalsup@aol.com> schrieb: >>> Did the book relate the story of why CRAY-1 presented a DC-load to >>> the power supply:: that is, the ECL gates were all of the form where >>> they would switch 20 ma into either the true or the complement out- >>> put and thus have no AC energy at the power supply level ?? >> >> That they didn't mention. They stressed his decision to build >> a machine which had good all-round performance, unlike the >> predecessors like the STAR or the Texas Instruments ASC (which I >> had never heard or read of). >> > > May be, that aspect of CRAY-1 was different from STAR and ASC, but not > different from CDC 6600 or 7600 or from top models of Cyber-170 series. CRAY-1 was relatively fast running scalar while ASC and STAR were not. >> There was one part on the Cray-I design that I found weird. After >> writing that individual transistors would have been faster than >> integrated circuits, but were chosen for density and manufacture, >> they wrote >> > > That part sounds correct. Logic ICs used in Cray-1 were indeed slower > than contemporary individual transistors. > >> "Concerning memory, as in the case of the CPU, Cray did not >> choose the fastest individual components, which would have >> been magnetic cores" due to their limitations in size. >> > > That part does not sound right. CRAY-1 main memory was made of SRAM > with 48 ns access time. That was 4-5 times faster than contemporary > core memories. Plus, it didn't suffer from destructive read. > One part that is true is that faster and less dense memory components > were available, but they were SRAM as well. > >> What they also describe well is the tradeoff between different >> vector lengths. Also very interesting is the mechanisms of getting >> the machines adopted by different industries.
[toc] | [prev] | [next] | [standalone]
| From | quadibloc <quadibloc@gmail.com> |
|---|---|
| Date | 2025-05-19 01:08 +0000 |
| Message-ID | <76948d869e78f8cb511809bd159008fd@www.novabbs.com> |
| In reply to | #111627 |
On Sun, 18 May 2025 15:23:03 +0000, Michael S wrote: > On Sun, 18 May 2025 05:46:37 -0000 (UTC) > Thomas Koenig <tkoenig@netcologne.de> wrote: > >> MitchAlsup1 <mitchalsup@aol.com> schrieb: >>> Did the book relate the story of why CRAY-1 presented a DC-load to >>> the power supply:: that is, the ECL gates were all of the form where >>> they would switch 20 ma into either the true or the complement out- >>> put and thus have no AC energy at the power supply level ?? >> >> That they didn't mention. They stressed his decision to build >> a machine which had good all-round performance, unlike the >> predecessors like the STAR or the Texas Instruments ASC (which I >> had never heard or read of). >> > > May be, that aspect of CRAY-1 was different from STAR and ASC, but not > different from CDC 6600 or 7600 or from top models of Cyber-170 series. Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary computers. They were not vector machines. The STAR and the ASC were vector machines - and because, unlike the Cray-I, vectors was the only thing they were good at, Amdahl's Law killed them. John Savard
[toc] | [prev] | [next] | [standalone]
| From | Lawrence D'Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2025-05-19 01:56 +0000 |
| Message-ID | <100e352$1d61i$3@dont-email.me> |
| In reply to | #111632 |
On Mon, 19 May 2025 01:08:11 +0000, quadibloc wrote: > Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary > computers. They were not vector machines. They were pipelined machines. They were orders of magnitude faster than anything from IBM. They pioneered the very concept of a “supercomputer”. There was nothing “ordinary” about that.
[toc] | [prev] | [next] | [standalone]
| From | quadibloc <quadibloc@gmail.com> |
|---|---|
| Date | 2025-05-19 03:12 +0000 |
| Message-ID | <e5fc3f66c40e74c1cf09ba5ed5a53c14@www.novabbs.com> |
| In reply to | #111635 |
On Mon, 19 May 2025 1:56:50 +0000, Lawrence D'Oliveiro wrote: > On Mon, 19 May 2025 01:08:11 +0000, quadibloc wrote: > >> Yes, but the CDC 6600 and 7600, while powerful computers, were ordinary >> computers. They were not vector machines. > > They were pipelined machines. They were orders of magnitude faster than > anything from IBM. They pioneered the very concept of a “supercomputer”. > > There was nothing “ordinary” about that. Yes, that is a fair comment. Eventually, IBM caught up with the Control Data 6600 by perfecting pipelining in the IBM 360/91, and then combining it with cache in the 360/195. From the Pentium II onwards, that's the way computers are made nowadays. I didn't mean to belittle the 6600, but simply to note that it lacked the additional speedup that you get from having a vector machine. Whereas the STAR-100 and the ASC had the opposite fault: having vectors was all those machines had going for them, while their scalar portions, unlike that of the 6600, were very definitely ordinary - and so Amdahl's Law bit them, as I noted. John Savard
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2025-05-19 06:22 +0000 |
| Subject | OoO execution (was: The Seymour Cray Era of Supercomputers) |
| Message-ID | <2025May19.082242@mips.complang.tuwien.ac.at> |
| In reply to | #111637 |
quadibloc <quadibloc@gmail.com> writes: >Eventually, IBM caught up with the Control >Data 6600 by perfecting pipelining in the IBM 360/91, and then combining >it with cache in the 360/195. From the Pentium II onwards, that's the >way computers are made nowadays. Pipelining and caches are already used on the MIPS R2000 in 1986, and the 486 in 1989. You are probably thinking of OoO Execution, where people usually write as if the Tomasulo algorithm of the 360/91 as implemented the modern concept of OoO execution. But the 360/91 only did OoO for FP, did not support branch prediction, had imprecise exceptions, and the Tomasulo algorithm was used primarily as a workaround for the dearth of FP registers in the S/360. The innovation that made OoO execution generally usable rather than a publicity stunt like the 360/91 is the reorder buffer (ROB), which allows to retire the instructions in-order, and to cancel speculatively "executed" instructions after an exception or branch misprediction. The Pentium Pro (introduced 1995-11-01), HP PA-8000 (introduced 1995-11-02), and MIPS R10000 (introduced 1996-01) are the first microprocessors which have full-blown OoO execution. But as someone pointed out to me, IBM has implemented OoO execution between the 370/195 and the Pentium Pro: The ES/9000 models 900 and 820 (shipping from September 1991) "were the first models with out-of-order execution since the System/370-195 of 1973. However unlike the old S/360-91-derived systems, the models 900 and 820 had full out-of-order execution for both integer and floating-point units, with precise exception handling, and a fully superscalar pipeline." <https://en.wikipedia.org/wiki/IBM_System/390#ES/9000>. So apparently they had a ROB, and AFAIK were the first machines to have one. These models also had a branch target buffer; the article does not mention branch prediction proper, but given a ROB and a branch target buffer, it would be surprising if they did not predict branches. So who came up with the concept of the ROB? I recently looked at one of the HPS papers (Hwu, Patt, Shebanov on a High Performance Substrate for the VAX from the mid-late 80s) again, and there was no ROB in that paper. I did not revisit their later papers whether they had it there. So apparently ROBs were not known in the mid-1980s in academia, and in 1991 there was hardware with a ROB commercially available, and a few years later it appeared in microprocessors. I wonder how early and how much IBM talked about their ES/9000 OoO implementation and features, but that may have inspired the teams at Intel, HP and SGI; or maybe there was an ealier source that inspired them all, but only in 1995/1996 the number of transistors on a chip was enough to do OoO on a microprocessor. Ironically, in the transition to CMOS (i.e., microprocessors) IBM mainframe processors regressed back to in-order (and, I think, single-issue) again (but with higher clock rates), and in the early 2000s they looked pretty outdated to me. In the meantime they have re-progressed to OoO again AFAIK. Back to OoO: it's interesting that Tomasulo and the 360/91 are mentioned often, but the ROB and its inventor(s?), which are at least as important for the success of OoO execution, isn't. - anton -- 'Anyone trying for "industrial quality" ISA should avoid undefined behavior.' Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
[toc] | [prev] | [next] | [standalone]
| From | John Levine <johnl@taugh.com> |
|---|---|
| Date | 2025-05-19 17:10 +0000 |
| Subject | Re: OoO execution (was: The Seymour Cray Era of Supercomputers) |
| Message-ID | <100fomr$n4q$1@gal.iecc.com> |
| In reply to | #111639 |
It appears that Anton Ertl <anton@mips.complang.tuwien.ac.at> said: >quadibloc <quadibloc@gmail.com> writes: >>Eventually, IBM caught up with the Control >>Data 6600 by perfecting pipelining in the IBM 360/91, and then combining >>it with cache in the 360/195. From the Pentium II onwards, that's the >>way computers are made nowadays. > >Pipelining and caches are already used on the MIPS R2000 in 1986, and >the 486 in 1989. > >You are probably thinking of OoO Execution, where people usually write >as if the Tomasulo algorithm of the 360/91 as implemented the modern >concept of OoO execution. But the 360/91 only did OoO for FP, did not >support branch prediction, had imprecise exceptions, and the Tomasulo >algorithm was used primarily as a workaround for the dearth of FP >registers in the S/360. The 360/91 had primitive branch prediction in "loop mode". It had an eight doublewprd instruction queue (which it confusingly called a stack.) If a program did a backward branch of less than eight doublewords, it'd stop prefetching and execute out of the queue until the program fell or branched out. It was occasionally worth tweaking assembly code to get a loop to start on a doubleword boundary (the CNOP assembler op) so it'd fit and run in loop mode. -- Regards, John Levine, johnl@taugh.com, Primary Perpetrator of "The Internet for Dummies", Please consider the environment before reading this e-mail. https://jl.ly
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2025-05-19 17:46 +0000 |
| Subject | Re: OoO execution (was: The Seymour Cray Era of Supercomputers) |
| Message-ID | <2025May19.194645@mips.complang.tuwien.ac.at> |
| In reply to | #111644 |
John Levine <johnl@taugh.com> writes: >The 360/91 had primitive branch prediction in "loop mode". It had an >eight doublewprd instruction queue (which it confusingly called a stack.) >If a program did a backward branch of less than eight doublewords, it'd >stop prefetching and execute out of the queue until the program fell or >branched out. The 68010 had a similar feature (with a smaller buffer), but I don't think one would call it branch prediction. In any case, I meant speculative execution based on branch prediction (but did not write it that way), and the 360/91 did not do speculative execution AFAIK. - anton -- 'Anyone trying for "industrial quality" ISA should avoid undefined behavior.' Mitch Alsup, <c17fcd89-f024-40e7-a594-88a85ac10d20o@googlegroups.com>
[toc] | [prev] | [next] | [standalone]
| From | ze@zerandconsulting.com (Ze) |
|---|---|
| Date | 2025-05-19 19:09 +0000 |
| Subject | Re: OoO execution |
| Message-ID | <37e1b146f23e2bebfec119d47fad36e4@www.novabbs.com> |
| In reply to | #111645 |
Wasn't one of the earliest forms of branch prediction the simple heuristic of always taking it in one direction and not taking it in the other direction , I seem to remember that being the case for some of the early pipelined microprocessors. I believe it was called static branch prediction compared to the more modern dynamic branch prediction. Nicholas (Nick) King --
[toc] | [prev] | [next] | [standalone]
| From | Lawrence D'Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2025-05-20 00:04 +0000 |
| Subject | Re: OoO execution |
| Message-ID | <100ggtj$1sbnn$4@dont-email.me> |
| In reply to | #111648 |
On Mon, 19 May 2025 19:09:12 +0000, Ze wrote: > Wasn't one of the earliest forms of branch prediction the simple > heuristic of always taking it in one direction and not taking it in the > other direction , I seem to remember that being the case for some of the > early pipelined microprocessors. I believe it was called static branch > prediction compared to the more modern dynamic branch prediction. The simple heuristic I remember was to assume that backward branches would be more likely to be taken than not (on the grounds that they were probably loops) while forward ones would more likely not be taken (I guess as an excuse for not disturbing the pipeline too much).
[toc] | [prev] | [next] | [standalone]
| From | mitchalsup@aol.com (MitchAlsup1) |
|---|---|
| Date | 2025-05-20 00:30 +0000 |
| Subject | Re: OoO execution |
| Message-ID | <c24984395cc6cb02482559555d959e63@www.novabbs.org> |
| In reply to | #111662 |
On Tue, 20 May 2025 0:04:03 +0000, Lawrence D'Oliveiro wrote: > On Mon, 19 May 2025 19:09:12 +0000, Ze wrote: > >> Wasn't one of the earliest forms of branch prediction the simple >> heuristic of always taking it in one direction and not taking it in the >> other direction , I seem to remember that being the case for some of the >> early pipelined microprocessors. I believe it was called static branch >> prediction compared to the more modern dynamic branch prediction. > > The simple heuristic I remember was to assume that backward branches > would > be more likely to be taken than not (on the grounds that they were > probably loops) while forward ones would more likely not be taken (I > guess > as an excuse for not disturbing the pipeline too much). CDC 7600 used this scheme. Backwards taken, forwards not-taken. Was about 70% accurate for essentially zero storage and 1 (or few) gates. This scheme might have been limited in scope (backwards into the instruction stack was predicted taken, farther than stack was predicted not-taken:: I don't remember exactly.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-05-20 13:52 +0000 |
| Subject | Re: OoO execution |
| Message-ID | <lO%WP.57064$RXsc.6962@fx36.iad> |
| In reply to | #111663 |
mitchalsup@aol.com (MitchAlsup1) writes: >On Tue, 20 May 2025 0:04:03 +0000, Lawrence D'Oliveiro wrote: > >> On Mon, 19 May 2025 19:09:12 +0000, Ze wrote: >> >>> Wasn't one of the earliest forms of branch prediction the simple >>> heuristic of always taking it in one direction and not taking it in the >>> other direction , I seem to remember that being the case for some of the >>> early pipelined microprocessors. I believe it was called static branch >>> prediction compared to the more modern dynamic branch prediction. >> >> The simple heuristic I remember was to assume that backward branches >> would >> be more likely to be taken than not (on the grounds that they were >> probably loops) while forward ones would more likely not be taken (I >> guess >> as an excuse for not disturbing the pipeline too much). > >CDC 7600 used this scheme. Backwards taken, forwards not-taken. >Was about 70% accurate for essentially zero storage and 1 (or few) >gates. Burroughs B4900 re-wrote the branch opcode on each branch to reflect the last two taken vs. not-taken choices. There were four opcodes for each type of branch - taken/taken, taken/not-taken, not-taken/taken and not-taken/not-taken.
[toc] | [prev] | [next] | [standalone]
| From | George Neuner <gneuner2@comcast.net> |
|---|---|
| Date | 2025-05-21 12:52 -0400 |
| Subject | Re: OoO execution (was: The Seymour Cray Era of Supercomputers) |
| Message-ID | <a70s2kdrpr4i3437u51ekebln93l397gfr@4ax.com> |
| In reply to | #111645 |
On Mon, 19 May 2025 17:46:45 GMT, anton@mips.complang.tuwien.ac.at (Anton Ertl) wrote: >John Levine <johnl@taugh.com> writes: >>The 360/91 had primitive branch prediction in "loop mode". It had an >>eight doublewprd instruction queue (which it confusingly called a stack.) >>If a program did a backward branch of less than eight doublewords, it'd >>stop prefetching and execute out of the queue until the program fell or >>branched out. > >The 68010 had a similar feature (with a smaller buffer), but I don't >think one would call it branch prediction. In any case, I meant >speculative execution based on branch prediction (but did not write it >that way), and the 360/91 did not do speculative execution AFAIK. > >- anton Most DSPs have some kind of "loop buffer" from which they can execute without fetching code from memory.
[toc] | [prev] | [next] | [standalone]
| From | Stefan Monnier <monnier@iro.umontreal.ca> |
|---|---|
| Date | 2025-05-21 13:14 -0400 |
| Subject | Re: OoO execution |
| Message-ID | <jwv34cxvqdp.fsf-monnier+comp.arch@gnu.org> |
| In reply to | #111718 |
> Most DSPs have some kind of "loop buffer" from which they can execute
> without fetching code from memory.
And Mitch's My 66000 `VEC` instruction takes the idea a step further.
Stefan
[toc] | [prev] | [next] | [standalone]
| From | moi <findlaybill@blueyonder.co.uk> |
|---|---|
| Date | 2025-05-21 18:47 +0100 |
| Subject | Re: OoO execution |
| Message-ID | <m96ht9Fpn7iU1@mid.individual.net> |
| In reply to | #111718 |
On 21/05/2025 17:52, George Neuner wrote: > On Mon, 19 May 2025 17:46:45 GMT, anton@mips.complang.tuwien.ac.at > (Anton Ertl) wrote: > >> John Levine <johnl@taugh.com> writes: >>> The 360/91 had primitive branch prediction in "loop mode". It had an >>> eight doublewprd instruction queue (which it confusingly called a stack.) >>> If a program did a backward branch of less than eight doublewords, it'd >>> stop prefetching and execute out of the queue until the program fell or >>> branched out. >> >> The 68010 had a similar feature (with a smaller buffer), but I don't >> think one would call it branch prediction. In any case, I meant >> speculative execution based on branch prediction (but did not write it >> that way), and the 360/91 did not do speculative execution AFAIK. >> >> - anton > > Most DSPs have some kind of "loop buffer" from which they can execute > without fetching code from memory. The Ferranti Atlas 2 and the EE KDF9 are both prior art. -- Bill F.
[toc] | [prev] | [next] | [standalone]
| From | EricP <ThatWouldBeTelling@thevillage.com> |
|---|---|
| Date | 2025-05-19 14:33 -0400 |
| Subject | Re: OoO execution |
| Message-ID | <1RKWP.151744$rkV6.62600@fx46.iad> |
| In reply to | #111639 |
Anton Ertl wrote: > quadibloc <quadibloc@gmail.com> writes: >> Eventually, IBM caught up with the Control >> Data 6600 by perfecting pipelining in the IBM 360/91, and then combining >> it with cache in the 360/195. From the Pentium II onwards, that's the >> way computers are made nowadays. > > Pipelining and caches are already used on the MIPS R2000 in 1986, and > the 486 in 1989. > > You are probably thinking of OoO Execution, where people usually write > as if the Tomasulo algorithm of the 360/91 as implemented the modern > concept of OoO execution. But the 360/91 only did OoO for FP, did not > support branch prediction, had imprecise exceptions, and the Tomasulo > algorithm was used primarily as a workaround for the dearth of FP > registers in the S/360. > > The innovation that made OoO execution generally usable rather than a > publicity stunt like the 360/91 is the reorder buffer (ROB), which allows to > retire the instructions in-order, and to cancel speculatively > "executed" instructions after an exception or branch misprediction. > > The Pentium Pro (introduced 1995-11-01), HP PA-8000 (introduced > 1995-11-02), and MIPS R10000 (introduced 1996-01) are the first > microprocessors which have full-blown OoO execution. > > But as someone pointed out to me, IBM has implemented OoO execution > between the 370/195 and the Pentium Pro: The ES/9000 models 900 and > 820 (shipping from September 1991) "were the first models with > out-of-order execution since the System/370-195 of 1973. However > unlike the old S/360-91-derived systems, the models 900 and 820 had > full out-of-order execution for both integer and floating-point units, > with precise exception handling, and a fully superscalar pipeline." > <https://en.wikipedia.org/wiki/IBM_System/390#ES/9000>. So apparently > they had a ROB, and AFAIK were the first machines to have one. These > models also had a branch target buffer; the article does not mention > branch prediction proper, but given a ROB and a branch target buffer, > it would be surprising if they did not predict branches. > > So who came up with the concept of the ROB? I recently looked at one > of the HPS papers (Hwu, Patt, Shebanov on a High Performance Substrate > for the VAX from the mid-late 80s) again, and there was no ROB in that > paper. I did not revisit their later papers whether they had it > there. So apparently ROBs were not known in the mid-1980s in > academia, and in 1991 there was hardware with a ROB commercially > available, and a few years later it appeared in microprocessors. There were a number of papers that circled around the various ideas. "Decoupled Access Execute Computer Architectures" uses queues to link the hardware modules together. "Implementing Precise Interrupts in Pipelined Processors" first mentions the ROB but doesn't have a renamer and limited OoO ability. HPS has rename, reservation stations, and multiple FU but no ROB. I don't know in what machine all the pieces came together at once but it looks like about 1986 they figured out to use multiple pipelines AND rename AND future file AND a ROB AND reservation stations AND multiple function units AND forwarding buses. Decoupled Access Execute Computer Architectures, James E. Smith, 1982 Instruction Issue Logic in Pipelined Supercomputers Shlomo Weiss, James E Smith, 1984 Implementing Precise Interrupts in Pipelined Processors, James E. Smith, A. R. Pleszkun, 1985 HPS - A New Microarchitecture Rationale And Introduction, Yale N. Patt, Wen-mei Hwu, and Michael Shebanow, 1985 > I wonder how early and how much IBM talked about their ES/9000 OoO > implementation and features, but that may have inspired the teams at > Intel, HP and SGI; or maybe there was an ealier source that inspired > them all, but only in 1995/1996 the number of transistors on a chip > was enough to do OoO on a microprocessor. > > Ironically, in the transition to CMOS (i.e., microprocessors) IBM > mainframe processors regressed back to in-order (and, I think, > single-issue) again (but with higher clock rates), and in the early > 2000s they looked pretty outdated to me. In the meantime they have > re-progressed to OoO again AFAIK. > > Back to OoO: it's interesting that Tomasulo and the 360/91 are > mentioned often, but the ROB and its inventor(s?), which are at least > as important for the success of OoO execution, isn't. > > - anton
[toc] | [prev] | [next] | [standalone]
| From | quadibloc <quadibloc@gmail.com> |
|---|---|
| Date | 2025-05-19 19:08 +0000 |
| Subject | Re: OoO execution |
| Message-ID | <8cbd51d6e62b575d9c2bf6b8cfa684af@www.novabbs.com> |
| In reply to | #111639 |
On Mon, 19 May 2025 6:22:42 +0000, Anton Ertl wrote: > You are probably thinking of OoO Execution, where people usually write > as if the Tomasulo algorithm of the 360/91 as implemented the modern > concept of OoO execution. But the 360/91 only did OoO for FP, did not > support branch prediction, had imprecise exceptions, and the Tomasulo > algorithm was used primarily as a workaround for the dearth of FP > registers in the S/360. Yes, I was thinking of OoO execution, as opposed to other forms of pipelining - basic pipelining was used in the 7094 II and even the 6502. The Pentium II (and Pentium Pro) also only used OoO for floating-point, while the 68050 only used OoO for integers! It's true the 360, with only four floating-point registers, had a dearth of them, but since having lots of registers was a way that RISC tried to avoid the need for OoO, I would not say that this invalidated the use of OoO on the 360/195. John Savard
[toc] | [prev] | [next] | [standalone]
Page 1 of 4 [1] 2 3 4 Next page →
Back to top | Article view | comp.arch
csiph-web