Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.arch > #2391 > unrolled thread
| Started by | MitchAlsup <MitchAlsup@aol.com> |
|---|---|
| First post | 2011-07-06 14:12 -0700 |
| Last post | 2011-07-19 07:22 +0100 |
| Articles | 20 on this page of 41 — 14 participants |
Back to article view | Back to comp.arch
Re: The Indexed Instruction Problem Solved! MitchAlsup <MitchAlsup@aol.com> - 2011-07-06 14:12 -0700
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-06 15:48 -0700
Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-06 20:36 -0400
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:29 -0700
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:31 -0700
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:38 -0700
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:39 -0700
Re: The Indexed Instruction Problem Solved! timcaffrey@aol.com (Tim McCaffrey) - 2011-07-14 01:36 +0000
Re: The Indexed Instruction Problem Solved! Stephen Fuld <SFuld@alumni.cmu.edu.invalid> - 2011-07-07 08:12 -0700
Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-07 12:24 -0400
Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-08 01:29 +0000
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 18:53 -0700
Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-08 02:17 +0000
Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-09 13:15 -0400
Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-09 18:37 +0000
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:53 -0700
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:55 -0700
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:49 -0700
Re: The Indexed Instruction Problem Solved! Chris Jones <clj@panix.com> - 2011-07-09 16:21 -0400
Re: The Indexed Instruction Problem Solved! Joe Chisolm <jchisolm6@earthlink.net> - 2011-07-07 13:19 -0500
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 12:11 -0700
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 12:13 -0700
Re: The Indexed Instruction Problem Solved! jsavard@excxn.aNOSPAMb.cdn.invalid (John Savard) - 2011-07-07 19:32 +0000
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 04:09 -0700
Re: The Indexed Instruction Problem Solved! Brett Davis <ggtgp@yahoo.com> - 2011-07-12 18:31 -0500
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-13 22:31 -0700
Re: The Indexed Instruction Problem Solved! mac <acolvin@efunct.com> - 2011-07-17 15:31 +0000
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-17 09:50 -0700
Re: The Indexed Instruction Problem Solved! Andrew Reilly <areilly---@bigpond.net.au> - 2011-07-17 23:26 +0000
Re: The Indexed Instruction Problem Solved! Brett Davis <ggtgp@yahoo.com> - 2011-07-17 19:52 -0500
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-17 18:22 -0700
Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-17 21:33 -0700
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 06:34 -0700
Re: The Indexed Instruction Problem Solved! Stephen Fuld <SFuld@alumni.cmu.edu.invalid> - 2011-07-18 09:22 -0700
Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 14:13 +0100
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 07:27 -0700
Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 16:22 +0100
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 09:56 -0700
Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 17:43 +0100
Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 14:58 -0700
Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-19 07:22 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-07 12:11 -0700 |
| Message-ID | <691ec859-feb9-4c6d-8d2c-00c50f72b2c3@5g2000yqb.googlegroups.com> |
| In reply to | #2412 |
On Jul 7, 12:19 pm, Joe Chisolm <jchiso...@earthlink.net> wrote: > Dont forget the Honeywell L66 and follow on DPS8. There is an entire > section in the assembly instruction manual on address modification. There > is 40 something pages on address generation then another 20 something on > plugging that address into the virtual memory addressing option. A pity those aren't up on Bitsavers yet... John Savard
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-07 12:13 -0700 |
| Message-ID | <40a1ca12-5e8a-427f-be2f-00395678d0ca@fq4g2000vbb.googlegroups.com> |
| In reply to | #2412 |
On Jul 7, 12:19 pm, Joe Chisolm <jchiso...@earthlink.net> wrote: > Dont forget the Honeywell L66 and follow on DPS8. There is an entire > section in the assembly instruction manual on address modification. There > is 40 something pages on address generation then another 20 something on > plugging that address into the virtual memory addressing option. ... I see I'm mistaken, and the DPS 8 is on Bitsavers. John Savard
[toc] | [prev] | [next] | [standalone]
| From | jsavard@excxn.aNOSPAMb.cdn.invalid (John Savard) |
|---|---|
| Date | 2011-07-07 19:32 +0000 |
| Message-ID | <4e1609ad.91669@news.aioe.org> |
| In reply to | #2415 |
On Thu, 7 Jul 2011 12:13:05 -0700 (PDT), Quadibloc <jsavard@ecn.ab.ca> wrote, in part: >On Jul 7, 12:19=A0pm, Joe Chisolm <jchiso...@earthlink.net> wrote: > >> Dont forget the Honeywell L66 and follow on DPS8. =A0There is an entire >> section in the assembly instruction manual on address modification. There >> is 40 something pages on address generation then another 20 something on >> plugging that address into the virtual memory addressing option. > >... I see I'm mistaken, and the DPS 8 is on Bitsavers. I think I'll skip trying to copy the indirect then tally mode. I've already included bisequential operation in my example architecture... John Savard http://www.quadibloc.com/index.html
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-07 04:09 -0700 |
| Message-ID | <333003aa-1e8f-4164-a921-47171bb471b3@d1g2000yqm.googlegroups.com> |
| In reply to | #2399 |
On Jul 7, 12:31 am, "Andy \"Krazy\" Glew" <a...@SPAM.comp-arch.net> wrote: > OVERALL OBSERVATION: addressing modes begin the slippery slope to VLIW. Today, though, the technology is at a point where it's doubtful that VLIW will be tried again. VLIW lets instructions explicitly use all the functional units of the computer simultaneously. But a specific program may not _have_ uses for those functional units, no matter how ingenious the compiler. Far better to have a decoupled microarchitecture, in which the processor keeps itself busy handling the differing requirements of multiple threads. The Itanium comes about as close to VLIW as any machine ever will these days - and, while it doesn't fall as far short of VLIW as the Texas Instruments 320C6000, it is still not a real VLIW machine, at least not if one considers a "real VLIW machine" to be something like the Cyberplus. John Savard
[toc] | [prev] | [next] | [standalone]
| From | Brett Davis <ggtgp@yahoo.com> |
|---|---|
| Date | 2011-07-12 18:31 -0500 |
| Message-ID | <ggtgp-55D8EB.18315712072011@netnews.mchsi.com> |
| In reply to | #2391 |
In article <b4ebeaa5-411f-4ee8-8a18-71a49e4e0ed3@glegroupsg2000goo.googlegroups.com>, MitchAlsup <MitchAlsup@aol.com> wrote: > I basically want all address modes to take the same amount of time in the > address generation pipeline. > > The 360 is only burdened with 12-bit positive only offsets, and we saw a > significant benefit to 16-bit +/- offsets in the RISC days; and it was not > until the x86-64 days that I became fully aware of how powerful the x86 > instruction set was when powered by multiple full width instruction decoders > (that give the property of the previous paragraph). > > In a RICS-like instruction set one must fundamentally choose the width of the > offset field (typically 16 but occasionally something strange like 13). > Not so with a x86-like instruction set and offsets (n.e. displacements) > can be 8, 16, 32 bits wide dynamically selected by modes and prefix codes. > But they all retain the "take the same time" property in the address generation stage. So x86 has an advantage in doing a 3 way add, source register plus index register plus offset constant of up to 32 bits. This means three input busses, with the instruction constant bus being cheaper than a register bus? RISC ALUs typically only have two inputs, but I assume RISC also has a constant bus to the ALU? If so an opcode update would be temped to add 3 way adds? At the opposite side of the spectrum, does anyone remember any CPU architectures that only supported direct addressing? Itanium was kinda weak in the address department. I have not looked much at stack processors, are stack architectures a mix with some supporting some address modes, and others just direct addressing?
[toc] | [prev] | [next] | [standalone]
| From | "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> |
|---|---|
| Date | 2011-07-13 22:31 -0700 |
| Message-ID | <4E1E7F1F.4040001@SPAM.comp-arch.net> |
| In reply to | #2475 |
On 7/12/2011 4:31 PM, Brett Davis wrote: > In article > <b4ebeaa5-411f-4ee8-8a18-71a49e4e0ed3@glegroupsg2000goo.googlegroups.com>, > So x86 has an advantage in doing a 3 way add, source register plus index register > plus offset constant of up to 32 bits. > This means three input busses, with the instruction constant bus being cheaper > than a register bus? > RISC ALUs typically only have two inputs, but I assume RISC also has a constant bus > to the ALU? If so an opcode update would be temped to add 3 way adds? > > At the opposite side of the spectrum, does anyone remember any CPU architectures > that only supported direct addressing? Itanium was kinda weak in the address department. AMD 29K.
[toc] | [prev] | [next] | [standalone]
| From | mac <acolvin@efunct.com> |
|---|---|
| Date | 2011-07-17 15:31 +0000 |
| Message-ID | <1670894707332520562.199029acolvin-efunct.com@news.eternal-september.org> |
| In reply to | #2475 |
> At the opposite side of the spectrum, does anyone remember any CPU architectures > that only supported direct addressing? Itanium was kinda weak in the address department. The 70's-era Able <http://sites.google.com/site/macthenaief> as well as some recent VLIW processors designed for embedded systems. A VLIW (and perhaps the Itanium) schedules offset and index arithmetic upstream of the memiry reference. In effect it uses extra integer ALUs as the address generation unit.
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-17 09:50 -0700 |
| Message-ID | <5b3f58f9-d560-4e09-ba6c-490b8edf5ff6@p31g2000vbs.googlegroups.com> |
| In reply to | #2557 |
On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote: > > At the opposite side of the spectrum, does anyone remember any CPU architectures > > that only supported direct addressing? Itanium was kinda weak in the address department. > > The 70's-era Able <http://sites.google.com/site/macthenaief> as well as > some recent VLIW processors designed for embedded systems. A VLIW (and > perhaps the Itanium) schedules offset and index arithmetic upstream of the > memiry reference. In effect it uses extra integer ALUs as the address > generation unit. And that also explains the 29000, since doing it that way makes sense with the RISC paradigm as well. John Savard
[toc] | [prev] | [next] | [standalone]
| From | Andrew Reilly <areilly---@bigpond.net.au> |
|---|---|
| Date | 2011-07-17 23:26 +0000 |
| Message-ID | <98h9dhFhltU2@mid.individual.net> |
| In reply to | #2560 |
On Sun, 17 Jul 2011 09:50:55 -0700, Quadibloc wrote: > On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote: >> > At the opposite side of the spectrum, does anyone remember any CPU >> > architectures that only supported direct addressing? Itanium was >> > kinda weak in the address department. >> >> The 70's-era Able <http://sites.google.com/site/macthenaief> as well as >> some recent VLIW processors designed for embedded systems. A VLIW (and >> perhaps the Itanium) schedules offset and index arithmetic upstream of >> the memiry reference. In effect it uses extra integer ALUs as the >> address generation unit. > > And that also explains the 29000, since doing it that way makes sense > with the RISC paradigm as well. It only makes sense if one cares that maximum flexibility and re-use is made of precious ALU real-estate. On the other hand, address generation is something that probably around half of issued instructions are going to need to do (or have done on their behalf), so having the extra hardware there to do it in parallel with the other ALU ops is surely a win when it comes to instruction space and decode effort and work done per clock? Certainly processors like the 68000 showed us that there are occasions when we want to use more elaborate algorithms to compute our addresses than the simple ones typically encoded in addressing modes, and that having to insert extra move instructions to get values from the "address registers" to the "integer registers" probably isn't a winning formula all the time. There are plenty of other processors that have AGU hardware that operates from normal integer registers (no store port required) without soaking up primary ALU execution slots or blowing out loop schedules, and without all of the extra instruction stream bother of another full-service execution pipeline (VLIW or otherwise). Cheers, -- Andrew
[toc] | [prev] | [next] | [standalone]
| From | Brett Davis <ggtgp@yahoo.com> |
|---|---|
| Date | 2011-07-17 19:52 -0500 |
| Message-ID | <ggtgp-A1357A.19522717072011@netnews.mchsi.com> |
| In reply to | #2568 |
In article <98h9dhFhltU2@mid.individual.net>, Andrew Reilly <areilly---@bigpond.net.au> wrote: > On Sun, 17 Jul 2011 09:50:55 -0700, Quadibloc wrote: > > On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote: > >> The 70's-era Able <http://sites.google.com/site/macthenaief> as well as > >> some recent VLIW processors designed for embedded systems. A VLIW (and > >> perhaps the Itanium) schedules offset and index arithmetic upstream of > >> the memiry reference. In effect it uses extra integer ALUs as the > >> address generation unit. > > > > And that also explains the 29000, since doing it that way makes sense > > with the RISC paradigm as well. > > It only makes sense if one cares that maximum flexibility and re-use is > made of precious ALU real-estate. On the other hand, address generation > is something that probably around half of issued instructions are going > to need to do (or have done on their behalf), so having the extra > hardware there to do it in parallel with the other ALU ops is surely a > win when it comes to instruction space and decode effort and work done > per clock? Certainly processors like the 68000 showed us that there are > occasions when we want to use more elaborate algorithms to compute our > addresses than the simple ones typically encoded in addressing modes, and > that having to insert extra move instructions to get values from the > "address registers" to the "integer registers" probably isn't a winning > formula all the time. There are plenty of other processors that have AGU > hardware that operates from normal integer registers (no store port > required) without soaking up primary ALU execution slots or blowing out > loop schedules, and without all of the extra instruction stream bother of > another full-service execution pipeline (VLIW or otherwise). Itanium seems to only support direct and post increment loads and stores, the post increment can be a constant or a register. Benefits listed are reduced code size, and that a separate increment would consume another I or M ALU unit. (Itanium volume 1, part 2, section 3.5.4) So an instruction like direct loads that does not use both ALU source operand paths is being wasteful. x86 seems to gain a small advantage from having address instructions with 3 sources, with a constant from the instruction being the third source. I assume the x86 instruction crack and recombine takes advantage of this feature for other instructions besides loads/stores? A post-RISC architecture might want to expose this as a general feature. Small constants are used for far more than just addressing. This may be one of the reasons POWER does instruction crack and recombine, to make optimizations like this to the instruction stream. Brett
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-17 18:22 -0700 |
| Message-ID | <b689f420-676b-4230-a1a6-31336baf0104@e7g2000vbw.googlegroups.com> |
| In reply to | #2568 |
On Jul 17, 5:26 pm, Andrew Reilly <areilly...@bigpond.net.au> wrote: > It only makes sense if one cares that maximum flexibility and re-use is > made of precious ALU real-estate. On the other hand, address generation > is something that probably around half of issued instructions are going > to need to do (or have done on their behalf), so having the extra > hardware there to do it in parallel with the other ALU ops is surely a > win when it comes to instruction space and decode effort and work done > per clock? I was going to say that in a pipelined architecture, one wouldn't be doing the address arithmetic at the same time as the arithmetic called for by the instruction itself... but, of course, one could be doing the arithmetic for _some other instruction_ at that time, and so a pipelined conventional CISC machine would indeed strongly benefit from separate hardware for address generation. But if you have a RISC architecture with only simple addressing, the address arithmetic is simply the principal arithmetic of an early instruction. So the pipeline runs smoothly without extra hardware - and indeed, even if one could tell that an instruction was forming an address, there would be no reason to use different hardware because the arithmetic would still be taking place during the execution cycle of that instruction. So the problem isn't implementing the architecture the wrong way. The problem is that the architecture itself forces you to fetch, decode, and execute more instructions to do anything. And yet, RISC is not a "bad idea" either; while cache memory counts for something too, doing stuff in registers instead of RAM whenever possible does help speed execution. John Savard
[toc] | [prev] | [next] | [standalone]
| From | "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> |
|---|---|
| Date | 2011-07-17 21:33 -0700 |
| Message-ID | <4E23B78E.8060204@SPAM.comp-arch.net> |
| In reply to | #2571 |
On 7/17/2011 6:22 PM, Quadibloc wrote: > On Jul 17, 5:26 pm, Andrew Reilly<areilly...@bigpond.net.au> wrote: > >> It only makes sense if one cares that maximum flexibility and re-use is >> made of precious ALU real-estate. On the other hand, address generation >> is something that probably around half of issued instructions are going >> to need to do (or have done on their behalf), so having the extra >> hardware there to do it in parallel with the other ALU ops is surely a >> win when it comes to instruction space and decode effort and work done >> per clock? > > I was going to say that in a pipelined architecture, one wouldn't be > doing the address arithmetic at the same time as the arithmetic called > for by the instruction itself... but, of course, one could be doing > the arithmetic for _some other instruction_ at that time, and so a > pipelined conventional CISC machine would indeed strongly benefit from > separate hardware for address generation. > > But if you have a RISC architecture with only simple addressing, the > address arithmetic is simply the principal arithmetic of an early > instruction. So the pipeline runs smoothly without extra hardware - > and indeed, even if one could tell that an instruction was forming an > address, there would be no reason to use different hardware because > the arithmetic would still be taking place during the execution cycle > of that instruction. > > So the problem isn't implementing the architecture the wrong way. The > problem is that the architecture itself forces you to fetch, decode, > and execute more instructions to do anything. And yet, RISC is not a > "bad idea" either; while cache memory counts for something too, doing > stuff in registers instead of RAM whenever possible does help speed > execution. > > John Savard The problem I always had with RISC addressing modes was stores: STORE M[ basereg+indexreg*scale+offset ] := store-data-reg has three register inputs. Whereas LOAD dest-reg := M[ basereg+indexreg*scale+offset ] has only 2 inputs, and one output. Let alone the problem of fitting a reasonable offset into a 32 bit instruction format with 5 5 bit register numbers and a scale. Forget the offset, and you still have the problem of 3 inputs. You either give in and allow 3 inputs for stores, or you do the ugly and let store have a weaker addressing mode than load.
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-18 06:34 -0700 |
| Message-ID | <c086f8cf-82d2-44d0-b1c7-3792f4017561@g9g2000yqb.googlegroups.com> |
| In reply to | #2573 |
On Jul 17, 10:33 pm, "Andy \"Krazy\" Glew" <a...@SPAM.comp-arch.net> wrote: > You either give in and allow 3 inputs for stores, > or you do the ugly and let store have a weaker addressing mode than load. This, of course, is true for non-RISC architectures as well. Load, Add, Subtract, And, Or, Xor, and so on would have a memory address as an input, a register as the output destination for the final result, and, if applicable, a base and index register as inputs. So the instruction formats just ignore the fact that in a load, the memory location is an input, while the result register is an output... and in a store, it's the other way around. So the problem _isn't_ finding bits for all the fields in an instruction. Three, four, or five bits specify a register if it's an input or an output equally well. So I take it the problem is a hardware design one. It could be something to do with how the instructions are pipelined. Or it could be that registers have to be selected via a multiplexer... so you need three multiplexers to handle a store, but two for a load... since there's only one "memory data register", containing what just got fetched from the bus. I would have thought, though, that in a RISC or non-RISC design, *that* kind of consideration was but a trifle with today's technology - even if people might have agonized over it in the days of discrete transistor computers. Of course, few of those _had_ multiple registers (some did; after all, the Pegasus, not the 360, invented registers - and the PDP-6, the PDP-10, the Univac 1105, and the Sigma were all discrete transistor machines with register banks - although none of those normally used base-displacement addressing). John Savard
[toc] | [prev] | [next] | [standalone]
| From | Stephen Fuld <SFuld@alumni.cmu.edu.invalid> |
|---|---|
| Date | 2011-07-18 09:22 -0700 |
| Message-ID | <j01mjb$ktl$1@dont-email.me> |
| In reply to | #2575 |
On 7/18/2011 6:34 AM, Quadibloc wrote: > On Jul 17, 10:33 pm, "Andy \"Krazy\" Glew"<a...@SPAM.comp-arch.net> > wrote: > >> You either give in and allow 3 inputs for stores, >> or you do the ugly and let store have a weaker addressing mode than load. > > This, of course, is true for non-RISC architectures as well. Load, > Add, Subtract, And, Or, Xor, and so on would have a memory address as > an input, a register as the output destination for the final result, > and, if applicable, a base and index register as inputs. > > So the instruction formats just ignore the fact that in a load, the > memory location is an input, while the result register is an output... > and in a store, it's the other way around. > > So the problem _isn't_ finding bits for all the fields in an > instruction. Three, four, or five bits specify a register if it's an > input or an output equally well. > > So I take it the problem is a hardware design one. > > It could be something to do with how the instructions are pipelined. I take Andy's comment to be that it is a register port issue. Most instructions require two register read ports and one write port, whereas a store instruction required three register read ports. This requires an extra port on the register file if you want to execute the store in a single cycle. But I agree with Nick that if having this extra port is an issue, then just make stores take an extra instruction when the complex modes are needed. Nick mentioned one argument why this would be acceptable, the relative infrequency of stores compared to loads. I would add two others. One is that many of the stores won't require the extra instruction as they don't need the complex addressing mode. The second is that even when it is required, sometimes it has already been computed from an earlier instruction and can be reused if it has been saved. So, while the extra instruction would be required sometimes, it is less often than one might think and requiring it occasionally isn't that bad. That being said, of course not having to require it at all, that is by having the extra read port is certainly better and, at the ISA level, more elegant. :-) -- - Stephen Fuld (e-mail address disguised to prevent spam)
[toc] | [prev] | [next] | [standalone]
| From | nmm1@cam.ac.uk |
|---|---|
| Date | 2011-07-18 14:13 +0100 |
| Message-ID | <j01bik$qcc$1@gosset.csi.cam.ac.uk> |
| In reply to | #2573 |
In article <4E23B78E.8060204@SPAM.comp-arch.net>, Andy \"Krazy\" Glew <andy@SPAM.comp-arch.net> wrote: > >The problem I always had with RISC addressing modes was stores: > >STORE M[ basereg+indexreg*scale+offset ] := store-data-reg > >has three register inputs. > >Whereas > >LOAD dest-reg := M[ basereg+indexreg*scale+offset ] > >has only 2 inputs, and one output. > >Let alone the problem of fitting a reasonable offset into a 32 bit >instruction format with 5 5 bit register numbers and a scale. > >Forget the offset, and you still have the problem of 3 inputs. > >You either give in and allow 3 inputs for stores, >or you do the ugly and let store have a weaker addressing mode than load. Well, you may call it ugly, but it is good engineering! Forget the issue you are talking about - the way that languages are designed means that stores are relatively simple compared with loads. That isn't just a fluke of history, because there are very good analysis reasons to keep stores simple. Some designs allow only 'local' stores but allow 'remote' loads - which becomes very relevant if you allow accesses to have an affinity register for SMP support. This is another case where I am saying that you hardware people should not provide the software people with everything you can, because they will only use the extra power to shoot themselves in their feet :-) Regards, Nick Maclaren.
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-18 07:27 -0700 |
| Message-ID | <392ea6d0-6980-4f64-9a32-e42c9ac272f2@h17g2000yqn.googlegroups.com> |
| In reply to | #2576 |
On Jul 18, 7:13 am, n...@cam.ac.uk wrote: > Well, you may call it ugly, but it is good engineering! Forget the > issue you are talking about - the way that languages are designed > means that stores are relatively simple compared with loads. That > isn't just a fluke of history, because there are very good analysis > reasons to keep stores simple. Some designs allow only 'local' > stores but allow 'remote' loads - which becomes very relevant if > you allow accesses to have an affinity register for SMP support. > > This is another case where I am saying that you hardware people > should not provide the software people with everything you can, > because they will only use the extra power to shoot themselves in > their feet :-) I find this comment quite strange. I would find it very painful if FORTRAN didn't allow subscripted variables on the left-hand side of the equals sign in an assignment statement. That would make the algorithms required for some problems much more complicated and/or slower. And if you don't go to that length, then the only result is that compiling the assignment statement now just generates one instruction to compute the address, and another one to store the value. I don't see what the point o that is; it doesn't prevent the programmer from doing anything, it just makes some things more expensive, so instead of preventing the programmer from shooting himself in the foot, it just gives him an additional way to do so. John Savard
[toc] | [prev] | [next] | [standalone]
| From | nmm1@cam.ac.uk |
|---|---|
| Date | 2011-07-18 16:22 +0100 |
| Message-ID | <j01j3n$r33$1@gosset.csi.cam.ac.uk> |
| In reply to | #2577 |
In article <392ea6d0-6980-4f64-9a32-e42c9ac272f2@h17g2000yqn.googlegroups.com>,
Quadibloc <jsavard@ecn.ab.ca> wrote:
>
>> Well, you may call it ugly, but it is good engineering! Forget the
>> issue you are talking about - the way that languages are designed
>> means that stores are relatively simple compared with loads. That
>> isn't just a fluke of history, because there are very good analysis
>> reasons to keep stores simple. Some designs allow only 'local'
>> stores but allow 'remote' loads - which becomes very relevant if
>> you allow accesses to have an affinity register for SMP support.
>>
>> This is another case where I am saying that you hardware people
>> should not provide the software people with everything you can,
>> because they will only use the extra power to shoot themselves in
>> their feet :-)
>
>I find this comment quite strange.
>
>I would find it very painful if FORTRAN didn't allow subscripted
>variables on the left-hand side of the equals sign in an assignment
>statement.
You have seriously misunderstood, both my point and the Fortran
standard!
In an assignment like <lvalue expr> = <rvalue expr>, almost all
languages specify that the left and right hand expressions must
be evaluated in toto before the assignment is performed. In sane
languages, those expressions do not include stores, though I am
not damning the paradigm <lvalue expr 1> = <lvalue expr 2> =
<lvalue expr 3> = <rvalue expr>, which can be defined cleanly.
So the actual code is required by the standard to be, roughly:
<reg 1> = <lvalue expr> (as an address)
<reg 2> = <rvalue expr>
(location pointed to by) <reg 1> = <reg 2>
Regards,
Nick Maclaren.
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-18 09:56 -0700 |
| Message-ID | <7f0bce2d-8c3d-4b7a-bda8-688d63d69497@df3g2000vbb.googlegroups.com> |
| In reply to | #2578 |
On Jul 18, 9:22 am, n...@cam.ac.uk wrote: > So the actual code is required by the standard to be, roughly: > > <reg 1> = <lvalue expr> (as an address) > <reg 2> = <rvalue expr> > (location pointed to by) <reg 1> = <reg 2> If a compiler that generates <reg 1> = subscript of lvalue, shifted left as required for type <reg 2> = <rvalue expr> store 2,array_address_constant(base reg,reg 1) is in violation of the FORTRAN standard, despite producing indistinguishable results, I'm sure that compiler writers happily ignore that fact. But I am hoping that I am just seriously misunderstanding your post once again. John Savard
[toc] | [prev] | [next] | [standalone]
| From | nmm1@cam.ac.uk |
|---|---|
| Date | 2011-07-18 17:43 +0100 |
| Message-ID | <j01nqp$rrs$1@gosset.csi.cam.ac.uk> |
| In reply to | #2581 |
In article <7f0bce2d-8c3d-4b7a-bda8-688d63d69497@df3g2000vbb.googlegroups.com>, Quadibloc <jsavard@ecn.ab.ca> wrote: > >> So the actual code is required by the standard to be, roughly: >> >> <reg 1> = <lvalue expr> (as an address) >> <reg 2> = <rvalue expr> >> (location pointed to by) <reg 1> = <reg 2> > >If a compiler that generates > > <reg 1> = subscript of lvalue, shifted left as required for type > <reg 2> = <rvalue expr> > store 2,array_address_constant(base reg,reg 1) > >is in violation of the FORTRAN standard, despite producing >indistinguishable results, I'm sure that compiler writers happily >ignore that fact. But I am hoping that I am just seriously >misunderstanding your post once again. No, it isn't, but that wasn't either of my points. The lesser of them was that, because of the required semantics, the complexity of stores generated by the compiler is less than that of loads (as well as there being a lot fewer of them). The greater point is that a few (fairly common) simples cases like that can be optimised, but the general cases (especially for CISC architectures) can't be. For example, it is NOT possible to use an indirection by memory store operation if there is any possibility that either the base address or index might be aliased to the location being stored into. And that can be legal, even in Fortran, and usually has to be assumed in C in the absence of global aliasing analysis. Even with RISC ISAs, that means that compilers sometimes have to cache copies of locations where there is no apparent need to, and failing to do so used to be a fairly common cause of foul code generation bugs. It tended to occur where there was a shortage of registers, the compiler needed to store a complex object, and it was more efficient to recalculate an address than use an extra register or scratch memory location. That didn't always work! There is a related aspect, which particularly affects parallelism and exception handling (INCLUDING at the hardware level). A series of reads can be retried ad lib in part and in any order (subject only to target register independence, which is statically analysable), but a series of stores is much trickier. Indeed, once one has indirection by memory store operations, they cannot be retried at all, in general. As a result, it is generally much cleaner to calculate the target address as a single location if there is any doubt about what might be going on. Regards, Nick Maclaren.
[toc] | [prev] | [next] | [standalone]
| From | Quadibloc <jsavard@ecn.ab.ca> |
|---|---|
| Date | 2011-07-18 14:58 -0700 |
| Message-ID | <25e315c7-78a5-41cb-902e-b326d69a8ed4@e7g2000vbw.googlegroups.com> |
| In reply to | #2583 |
On Jul 18, 10:43 am, n...@cam.ac.uk wrote: > As a result, it is generally much cleaner to calculate the target > address as a single location if there is any doubt about what might > be going on. Now I understand what you were intending to say. But the way I understood what Andy Glew said, he was talking about omitting capabilities from the architecture that would have been used in the optimized simple case. So I still don't agree that these FORTRAN considerations make it reasonable to drop indexed addressing as a feature of store instructions. Although I suppose it could be claimed, as you appear to be doing, that doing so would force compiler writers to do things the "safe" way more often. John Savard
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.arch
csiph-web