Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch > #2391 > unrolled thread

Re: The Indexed Instruction Problem Solved!

Started byMitchAlsup <MitchAlsup@aol.com>
First post2011-07-06 14:12 -0700
Last post2011-07-19 07:22 +0100
Articles 20 on this page of 41 — 14 participants

Back to article view | Back to comp.arch


Contents

  Re: The Indexed Instruction Problem Solved! MitchAlsup <MitchAlsup@aol.com> - 2011-07-06 14:12 -0700
    Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-06 15:48 -0700
    Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-06 20:36 -0400
    Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:29 -0700
      Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:31 -0700
        Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:38 -0700
          Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-06 23:39 -0700
            Re: The Indexed Instruction Problem Solved! timcaffrey@aol.com (Tim McCaffrey) - 2011-07-14 01:36 +0000
          Re: The Indexed Instruction Problem Solved! Stephen Fuld <SFuld@alumni.cmu.edu.invalid> - 2011-07-07 08:12 -0700
          Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-07 12:24 -0400
            Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-08 01:29 +0000
              Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 18:53 -0700
                Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-08 02:17 +0000
              Re: The Indexed Instruction Problem Solved! EricP <ThatWouldBeTelling@thevillage.com> - 2011-07-09 13:15 -0400
                Re: The Indexed Instruction Problem Solved! John Levine <johnl@iecc.com> - 2011-07-09 18:37 +0000
                  Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:53 -0700
                    Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:55 -0700
                Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-09 11:49 -0700
                Re: The Indexed Instruction Problem Solved! Chris Jones <clj@panix.com> - 2011-07-09 16:21 -0400
          Re: The Indexed Instruction Problem Solved! Joe Chisolm <jchisolm6@earthlink.net> - 2011-07-07 13:19 -0500
            Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 12:11 -0700
            Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 12:13 -0700
              Re: The Indexed Instruction Problem Solved! jsavard@excxn.aNOSPAMb.cdn.invalid (John Savard) - 2011-07-07 19:32 +0000
        Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-07 04:09 -0700
    Re: The Indexed Instruction Problem Solved! Brett Davis <ggtgp@yahoo.com> - 2011-07-12 18:31 -0500
      Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-13 22:31 -0700
      Re: The Indexed Instruction Problem Solved! mac <acolvin@efunct.com> - 2011-07-17 15:31 +0000
        Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-17 09:50 -0700
          Re: The Indexed Instruction Problem Solved! Andrew Reilly <areilly---@bigpond.net.au> - 2011-07-17 23:26 +0000
            Re: The Indexed Instruction Problem Solved! Brett Davis <ggtgp@yahoo.com> - 2011-07-17 19:52 -0500
            Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-17 18:22 -0700
              Re: The Indexed Instruction Problem Solved! "Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net> - 2011-07-17 21:33 -0700
                Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 06:34 -0700
                  Re: The Indexed Instruction Problem Solved! Stephen Fuld <SFuld@alumni.cmu.edu.invalid> - 2011-07-18 09:22 -0700
                Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 14:13 +0100
                  Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 07:27 -0700
                    Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 16:22 +0100
                      Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 09:56 -0700
                        Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-18 17:43 +0100
                          Re: The Indexed Instruction Problem Solved! Quadibloc <jsavard@ecn.ab.ca> - 2011-07-18 14:58 -0700
                            Re: The Indexed Instruction Problem Solved! nmm1@cam.ac.uk - 2011-07-19 07:22 +0100

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#2414

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-07 12:11 -0700
Message-ID<691ec859-feb9-4c6d-8d2c-00c50f72b2c3@5g2000yqb.googlegroups.com>
In reply to#2412
On Jul 7, 12:19 pm, Joe Chisolm <jchiso...@earthlink.net> wrote:

> Dont forget the Honeywell L66 and follow on DPS8.  There is an entire
> section in the assembly instruction manual on address modification. There
> is 40 something pages on address generation then another 20 something on
> plugging that address into the virtual memory addressing option.

A pity those aren't up on Bitsavers yet...

John Savard

[toc] | [prev] | [next] | [standalone]


#2415

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-07 12:13 -0700
Message-ID<40a1ca12-5e8a-427f-be2f-00395678d0ca@fq4g2000vbb.googlegroups.com>
In reply to#2412
On Jul 7, 12:19 pm, Joe Chisolm <jchiso...@earthlink.net> wrote:

> Dont forget the Honeywell L66 and follow on DPS8.  There is an entire
> section in the assembly instruction manual on address modification. There
> is 40 something pages on address generation then another 20 something on
> plugging that address into the virtual memory addressing option.

... I see I'm mistaken, and the DPS 8 is on Bitsavers.

John Savard

[toc] | [prev] | [next] | [standalone]


#2416

Fromjsavard@excxn.aNOSPAMb.cdn.invalid (John Savard)
Date2011-07-07 19:32 +0000
Message-ID<4e1609ad.91669@news.aioe.org>
In reply to#2415
On Thu, 7 Jul 2011 12:13:05 -0700 (PDT), Quadibloc <jsavard@ecn.ab.ca>
wrote, in part:

>On Jul 7, 12:19=A0pm, Joe Chisolm <jchiso...@earthlink.net> wrote:
>
>> Dont forget the Honeywell L66 and follow on DPS8. =A0There is an entire
>> section in the assembly instruction manual on address modification. There
>> is 40 something pages on address generation then another 20 something on
>> plugging that address into the virtual memory addressing option.
>
>... I see I'm mistaken, and the DPS 8 is on Bitsavers.

I think I'll skip trying to copy the indirect then tally mode.

I've already included bisequential operation in my example
architecture...

John Savard
http://www.quadibloc.com/index.html

[toc] | [prev] | [next] | [standalone]


#2405

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-07 04:09 -0700
Message-ID<333003aa-1e8f-4164-a921-47171bb471b3@d1g2000yqm.googlegroups.com>
In reply to#2399
On Jul 7, 12:31 am, "Andy \"Krazy\" Glew" <a...@SPAM.comp-arch.net>
wrote:

> OVERALL OBSERVATION:  addressing modes begin the slippery slope to VLIW.

Today, though, the technology is at a point where it's doubtful that
VLIW will be tried again.

VLIW lets instructions explicitly use all the functional units of the
computer simultaneously. But a specific program may not _have_ uses
for those functional units, no matter how ingenious the compiler.

Far better to have a decoupled microarchitecture, in which the
processor keeps itself busy handling the differing requirements of
multiple threads.

The Itanium comes about as close to VLIW as any machine ever will
these days - and, while it doesn't fall as far short of VLIW as the
Texas Instruments 320C6000, it is still not a real VLIW machine, at
least not if one considers a "real VLIW machine" to be something like
the Cyberplus.

John Savard

[toc] | [prev] | [next] | [standalone]


#2475

FromBrett Davis <ggtgp@yahoo.com>
Date2011-07-12 18:31 -0500
Message-ID<ggtgp-55D8EB.18315712072011@netnews.mchsi.com>
In reply to#2391
In article 
<b4ebeaa5-411f-4ee8-8a18-71a49e4e0ed3@glegroupsg2000goo.googlegroups.com>,
 MitchAlsup <MitchAlsup@aol.com> wrote:
> I basically want all address modes to take the same amount of time in the 
> address generation pipeline.
> 
> The 360 is only burdened with 12-bit positive only offsets, and we saw a 
> significant benefit to 16-bit +/- offsets in the RISC days; and it was not 
> until the x86-64 days that I became fully aware of how powerful the x86 
> instruction set was when powered by multiple full width instruction decoders 
> (that give the property of the previous paragraph).
> 
> In a RICS-like instruction set one must fundamentally choose the width of the 
> offset field (typically 16 but occasionally something strange like 13). 
> Not so with a x86-like instruction set and offsets (n.e. displacements) 
> can be 8, 16, 32 bits wide dynamically selected by modes and prefix codes. 
> But they all retain the "take the same time" property in the address generation stage.

So x86 has an advantage in doing a 3 way add, source register plus index register
plus offset constant of up to 32 bits.
This means three input busses, with the instruction constant bus being cheaper
than a register bus?
RISC ALUs typically only have two inputs, but I assume RISC also has a constant bus
to the ALU? If so an opcode update would be temped to add 3 way adds?

At the opposite side of the spectrum, does anyone remember any CPU architectures
that only supported direct addressing? Itanium was kinda weak in the address department.

I have not looked much at stack processors, are stack architectures a mix with some
supporting some address modes, and others just direct addressing?

[toc] | [prev] | [next] | [standalone]


#2507

From"Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net>
Date2011-07-13 22:31 -0700
Message-ID<4E1E7F1F.4040001@SPAM.comp-arch.net>
In reply to#2475
On 7/12/2011 4:31 PM, Brett Davis wrote:
> In article
> <b4ebeaa5-411f-4ee8-8a18-71a49e4e0ed3@glegroupsg2000goo.googlegroups.com>,

> So x86 has an advantage in doing a 3 way add, source register plus index register
> plus offset constant of up to 32 bits.
> This means three input busses, with the instruction constant bus being cheaper
> than a register bus?
> RISC ALUs typically only have two inputs, but I assume RISC also has a constant bus
> to the ALU? If so an opcode update would be temped to add 3 way adds?
>
> At the opposite side of the spectrum, does anyone remember any CPU architectures
> that only supported direct addressing? Itanium was kinda weak in the address department.

AMD 29K.

[toc] | [prev] | [next] | [standalone]


#2557

Frommac <acolvin@efunct.com>
Date2011-07-17 15:31 +0000
Message-ID<1670894707332520562.199029acolvin-efunct.com@news.eternal-september.org>
In reply to#2475
> At the opposite side of the spectrum, does anyone remember any CPU architectures
> that only supported direct addressing? Itanium was kinda weak in the address department.

The 70's-era Able <http://sites.google.com/site/macthenaief> as well as
some recent VLIW processors designed for embedded systems. A VLIW (and
perhaps the Itanium) schedules offset and index arithmetic upstream of the
memiry reference. In effect it uses extra integer ALUs as the address
generation unit.

[toc] | [prev] | [next] | [standalone]


#2560

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-17 09:50 -0700
Message-ID<5b3f58f9-d560-4e09-ba6c-490b8edf5ff6@p31g2000vbs.googlegroups.com>
In reply to#2557
On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote:
> > At the opposite side of the spectrum, does anyone remember any CPU architectures
> > that only supported direct addressing? Itanium was kinda weak in the address department.
>
> The 70's-era Able <http://sites.google.com/site/macthenaief> as well as
> some recent VLIW processors designed for embedded systems. A VLIW (and
> perhaps the Itanium) schedules offset and index arithmetic upstream of the
> memiry reference. In effect it uses extra integer ALUs as the address
> generation unit.

And that also explains the 29000, since doing it that way makes sense
with the RISC paradigm as well.

John Savard

[toc] | [prev] | [next] | [standalone]


#2568

FromAndrew Reilly <areilly---@bigpond.net.au>
Date2011-07-17 23:26 +0000
Message-ID<98h9dhFhltU2@mid.individual.net>
In reply to#2560
On Sun, 17 Jul 2011 09:50:55 -0700, Quadibloc wrote:

> On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote:
>> > At the opposite side of the spectrum, does anyone remember any CPU
>> > architectures that only supported direct addressing? Itanium was
>> > kinda weak in the address department.
>>
>> The 70's-era Able <http://sites.google.com/site/macthenaief> as well as
>> some recent VLIW processors designed for embedded systems. A VLIW (and
>> perhaps the Itanium) schedules offset and index arithmetic upstream of
>> the memiry reference. In effect it uses extra integer ALUs as the
>> address generation unit.
> 
> And that also explains the 29000, since doing it that way makes sense
> with the RISC paradigm as well.

It only makes sense if one cares that maximum flexibility and re-use is 
made of precious ALU real-estate.  On the other hand, address generation 
is something that probably around half of issued instructions are going 
to need to do (or have done on their behalf), so having the extra 
hardware there to do it in parallel with the other ALU ops is surely a 
win when it comes to instruction space and decode effort and work done 
per clock?  Certainly processors like the 68000 showed us that there are 
occasions when we want to use more elaborate algorithms to compute our 
addresses than the simple ones typically encoded in addressing modes, and 
that having to insert extra move instructions to get values from the 
"address registers" to the "integer registers" probably isn't a winning 
formula all the time.  There are plenty of other processors that have AGU 
hardware that operates from normal integer registers (no store port 
required) without soaking up primary ALU execution slots or blowing out 
loop schedules, and without all of the extra instruction stream bother of 
another full-service execution pipeline (VLIW or otherwise).

Cheers,

-- 
Andrew

[toc] | [prev] | [next] | [standalone]


#2569

FromBrett Davis <ggtgp@yahoo.com>
Date2011-07-17 19:52 -0500
Message-ID<ggtgp-A1357A.19522717072011@netnews.mchsi.com>
In reply to#2568
In article <98h9dhFhltU2@mid.individual.net>,
 Andrew Reilly <areilly---@bigpond.net.au> wrote:

> On Sun, 17 Jul 2011 09:50:55 -0700, Quadibloc wrote:
> > On Jul 17, 9:31 am, mac <acol...@efunct.com> wrote:
> >> The 70's-era Able <http://sites.google.com/site/macthenaief> as well as
> >> some recent VLIW processors designed for embedded systems. A VLIW (and
> >> perhaps the Itanium) schedules offset and index arithmetic upstream of
> >> the memiry reference. In effect it uses extra integer ALUs as the
> >> address generation unit.
> > 
> > And that also explains the 29000, since doing it that way makes sense
> > with the RISC paradigm as well.
> 
> It only makes sense if one cares that maximum flexibility and re-use is 
> made of precious ALU real-estate.  On the other hand, address generation 
> is something that probably around half of issued instructions are going 
> to need to do (or have done on their behalf), so having the extra 
> hardware there to do it in parallel with the other ALU ops is surely a 
> win when it comes to instruction space and decode effort and work done 
> per clock?  Certainly processors like the 68000 showed us that there are 
> occasions when we want to use more elaborate algorithms to compute our 
> addresses than the simple ones typically encoded in addressing modes, and 
> that having to insert extra move instructions to get values from the 
> "address registers" to the "integer registers" probably isn't a winning 
> formula all the time.  There are plenty of other processors that have AGU 
> hardware that operates from normal integer registers (no store port 
> required) without soaking up primary ALU execution slots or blowing out 
> loop schedules, and without all of the extra instruction stream bother of 
> another full-service execution pipeline (VLIW or otherwise).

Itanium seems to only support direct and post increment loads and stores,
the post increment can be a constant or a register.
Benefits listed are reduced code size, and that a separate increment would
consume another I or M ALU unit.
(Itanium volume 1, part 2, section 3.5.4)

So an instruction like direct loads that does not use both ALU source 
operand paths is being wasteful.

x86 seems to gain a small advantage from having address instructions
with 3 sources, with a constant from the instruction being the third source.

I assume the x86 instruction crack and recombine takes advantage of this
feature for other instructions besides loads/stores?

A post-RISC architecture might want to expose this as a general feature.
Small constants are used for far more than just addressing.

This may be one of the reasons POWER does instruction crack and recombine,
to make optimizations like this to the instruction stream.

Brett

[toc] | [prev] | [next] | [standalone]


#2571

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-17 18:22 -0700
Message-ID<b689f420-676b-4230-a1a6-31336baf0104@e7g2000vbw.googlegroups.com>
In reply to#2568
On Jul 17, 5:26 pm, Andrew Reilly <areilly...@bigpond.net.au> wrote:

> It only makes sense if one cares that maximum flexibility and re-use is
> made of precious ALU real-estate.  On the other hand, address generation
> is something that probably around half of issued instructions are going
> to need to do (or have done on their behalf), so having the extra
> hardware there to do it in parallel with the other ALU ops is surely a
> win when it comes to instruction space and decode effort and work done
> per clock?

I was going to say that in a pipelined architecture, one wouldn't be
doing the address arithmetic at the same time as the arithmetic called
for by the instruction itself... but, of course, one could be doing
the arithmetic for _some other instruction_ at that time, and so a
pipelined conventional CISC machine would indeed strongly benefit from
separate hardware for address generation.

But if you have a RISC architecture with only simple addressing, the
address arithmetic is simply the principal arithmetic of an early
instruction. So the pipeline runs smoothly without extra hardware -
and indeed, even if one could tell that an instruction was forming an
address, there would be no reason to use different hardware because
the arithmetic would still be taking place during the execution cycle
of that instruction.

So the problem isn't implementing the architecture the wrong way. The
problem is that the architecture itself forces you to fetch, decode,
and execute more instructions to do anything. And yet, RISC is not a
"bad idea" either; while cache memory counts for something too, doing
stuff in registers instead of RAM whenever possible does help speed
execution.

John Savard

[toc] | [prev] | [next] | [standalone]


#2573

From"Andy \"Krazy\" Glew" <andy@SPAM.comp-arch.net>
Date2011-07-17 21:33 -0700
Message-ID<4E23B78E.8060204@SPAM.comp-arch.net>
In reply to#2571
On 7/17/2011 6:22 PM, Quadibloc wrote:
> On Jul 17, 5:26 pm, Andrew Reilly<areilly...@bigpond.net.au>  wrote:
>
>> It only makes sense if one cares that maximum flexibility and re-use is
>> made of precious ALU real-estate.  On the other hand, address generation
>> is something that probably around half of issued instructions are going
>> to need to do (or have done on their behalf), so having the extra
>> hardware there to do it in parallel with the other ALU ops is surely a
>> win when it comes to instruction space and decode effort and work done
>> per clock?
>
> I was going to say that in a pipelined architecture, one wouldn't be
> doing the address arithmetic at the same time as the arithmetic called
> for by the instruction itself... but, of course, one could be doing
> the arithmetic for _some other instruction_ at that time, and so a
> pipelined conventional CISC machine would indeed strongly benefit from
> separate hardware for address generation.
>
> But if you have a RISC architecture with only simple addressing, the
> address arithmetic is simply the principal arithmetic of an early
> instruction. So the pipeline runs smoothly without extra hardware -
> and indeed, even if one could tell that an instruction was forming an
> address, there would be no reason to use different hardware because
> the arithmetic would still be taking place during the execution cycle
> of that instruction.
>
> So the problem isn't implementing the architecture the wrong way. The
> problem is that the architecture itself forces you to fetch, decode,
> and execute more instructions to do anything. And yet, RISC is not a
> "bad idea" either; while cache memory counts for something too, doing
> stuff in registers instead of RAM whenever possible does help speed
> execution.
>
> John Savard


The problem I always had with RISC addressing modes was stores:

STORE M[ basereg+indexreg*scale+offset ] := store-data-reg

has three register inputs.

Whereas

LOAD dest-reg := M[ basereg+indexreg*scale+offset ]

has only 2 inputs, and one output.

Let alone the problem of fitting a reasonable offset into a 32 bit 
instruction format with 5 5 bit register numbers and a scale.

Forget the offset, and you still have the problem of 3 inputs.

You either give in and allow 3 inputs for stores,
or you do the ugly and let store have a weaker addressing mode than load.

[toc] | [prev] | [next] | [standalone]


#2575

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-18 06:34 -0700
Message-ID<c086f8cf-82d2-44d0-b1c7-3792f4017561@g9g2000yqb.googlegroups.com>
In reply to#2573
On Jul 17, 10:33 pm, "Andy \"Krazy\" Glew" <a...@SPAM.comp-arch.net>
wrote:

> You either give in and allow 3 inputs for stores,
> or you do the ugly and let store have a weaker addressing mode than load.

This, of course, is true for non-RISC architectures as well. Load,
Add, Subtract, And, Or, Xor, and so on would have a memory address as
an input, a register as the output destination for the final result,
and, if applicable, a base and index register as inputs.

So the instruction formats just ignore the fact that in a load, the
memory location is an input, while the result register is an output...
and in a store, it's the other way around.

So the problem _isn't_ finding bits for all the fields in an
instruction. Three, four, or five bits specify a register if it's an
input or an output equally well.

So I take it the problem is a hardware design one.

It could be something to do with how the instructions are pipelined.

Or it could be that registers have to be selected via a multiplexer...
so you need three multiplexers to handle a store, but two for a
load... since there's only one "memory data register", containing what
just got fetched from the bus.

I would have thought, though, that in a RISC or non-RISC design,
*that* kind of consideration was but a trifle with today's technology
- even if people might have agonized over it in the days of discrete
transistor computers. Of course, few of those _had_ multiple registers
(some did; after all, the Pegasus, not the 360, invented registers -
and the PDP-6, the PDP-10, the Univac 1105, and the Sigma were all
discrete transistor machines with register banks - although none of
those normally used base-displacement addressing).

John Savard

[toc] | [prev] | [next] | [standalone]


#2580

FromStephen Fuld <SFuld@alumni.cmu.edu.invalid>
Date2011-07-18 09:22 -0700
Message-ID<j01mjb$ktl$1@dont-email.me>
In reply to#2575
On 7/18/2011 6:34 AM, Quadibloc wrote:
> On Jul 17, 10:33 pm, "Andy \"Krazy\" Glew"<a...@SPAM.comp-arch.net>
> wrote:
>
>> You either give in and allow 3 inputs for stores,
>> or you do the ugly and let store have a weaker addressing mode than load.
>
> This, of course, is true for non-RISC architectures as well. Load,
> Add, Subtract, And, Or, Xor, and so on would have a memory address as
> an input, a register as the output destination for the final result,
> and, if applicable, a base and index register as inputs.
>
> So the instruction formats just ignore the fact that in a load, the
> memory location is an input, while the result register is an output...
> and in a store, it's the other way around.
>
> So the problem _isn't_ finding bits for all the fields in an
> instruction. Three, four, or five bits specify a register if it's an
> input or an output equally well.
>
> So I take it the problem is a hardware design one.
>
> It could be something to do with how the instructions are pipelined.

I take Andy's comment to be that it is a register port issue.  Most 
instructions require two register read ports and one write port, whereas 
a store instruction required three register read ports.  This requires 
an extra port on the register file if you want to execute the store in a 
single cycle.

But I agree with Nick that if having this extra port is an issue, then 
just make stores take an extra instruction when the complex modes are 
needed.  Nick mentioned one argument why this would be acceptable, the 
relative infrequency of stores compared to loads.  I would add two 
others.  One is that many of the stores won't require the extra 
instruction as they don't need the complex addressing mode.  The second 
is that even when it is required, sometimes it has already been computed 
from an earlier instruction and can be reused if it has been saved.  So, 
while the extra instruction would be required sometimes, it is less 
often than one might think and requiring it occasionally isn't that bad.

That being said, of course not having to require it at all, that is by 
having the extra read port is certainly better and, at the ISA level, 
more elegant.  :-)



-- 
  - Stephen Fuld
(e-mail address disguised to prevent spam)

[toc] | [prev] | [next] | [standalone]


#2576

Fromnmm1@cam.ac.uk
Date2011-07-18 14:13 +0100
Message-ID<j01bik$qcc$1@gosset.csi.cam.ac.uk>
In reply to#2573
In article <4E23B78E.8060204@SPAM.comp-arch.net>,
Andy \"Krazy\" Glew <andy@SPAM.comp-arch.net> wrote:
>
>The problem I always had with RISC addressing modes was stores:
>
>STORE M[ basereg+indexreg*scale+offset ] := store-data-reg
>
>has three register inputs.
>
>Whereas
>
>LOAD dest-reg := M[ basereg+indexreg*scale+offset ]
>
>has only 2 inputs, and one output.
>
>Let alone the problem of fitting a reasonable offset into a 32 bit 
>instruction format with 5 5 bit register numbers and a scale.
>
>Forget the offset, and you still have the problem of 3 inputs.
>
>You either give in and allow 3 inputs for stores,
>or you do the ugly and let store have a weaker addressing mode than load.

Well, you may call it ugly, but it is good engineering!  Forget the
issue you are talking about - the way that languages are designed
means that stores are relatively simple compared with loads.  That
isn't just a fluke of history, because there are very good analysis
reasons to keep stores simple.  Some designs allow only 'local'
stores but allow 'remote' loads - which becomes very relevant if
you allow accesses to have an affinity register for SMP support.

This is another case where I am saying that you hardware people
should not provide the software people with everything you can,
because they will only use the extra power to shoot themselves in
their feet :-)


Regards,
Nick Maclaren.

[toc] | [prev] | [next] | [standalone]


#2577

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-18 07:27 -0700
Message-ID<392ea6d0-6980-4f64-9a32-e42c9ac272f2@h17g2000yqn.googlegroups.com>
In reply to#2576
On Jul 18, 7:13 am, n...@cam.ac.uk wrote:

> Well, you may call it ugly, but it is good engineering!  Forget the
> issue you are talking about - the way that languages are designed
> means that stores are relatively simple compared with loads.  That
> isn't just a fluke of history, because there are very good analysis
> reasons to keep stores simple.  Some designs allow only 'local'
> stores but allow 'remote' loads - which becomes very relevant if
> you allow accesses to have an affinity register for SMP support.
>
> This is another case where I am saying that you hardware people
> should not provide the software people with everything you can,
> because they will only use the extra power to shoot themselves in
> their feet :-)

I find this comment quite strange.

I would find it very painful if FORTRAN didn't allow subscripted
variables on the left-hand side of the equals sign in an assignment
statement.

That would make the algorithms required for some problems much more
complicated and/or slower.

And if you don't go to that length, then the only result is that
compiling the assignment statement now just generates one instruction
to compute the address, and another one to store the value. I don't
see what the point o that is; it doesn't prevent the programmer from
doing anything, it just makes some things more expensive, so instead
of preventing the programmer from shooting himself in the foot, it
just gives him an additional way to do so.

John Savard

[toc] | [prev] | [next] | [standalone]


#2578

Fromnmm1@cam.ac.uk
Date2011-07-18 16:22 +0100
Message-ID<j01j3n$r33$1@gosset.csi.cam.ac.uk>
In reply to#2577
In article <392ea6d0-6980-4f64-9a32-e42c9ac272f2@h17g2000yqn.googlegroups.com>,
Quadibloc  <jsavard@ecn.ab.ca> wrote:
>
>> Well, you may call it ugly, but it is good engineering! Forget the
>> issue you are talking about - the way that languages are designed
>> means that stores are relatively simple compared with loads. That
>> isn't just a fluke of history, because there are very good analysis
>> reasons to keep stores simple. Some designs allow only 'local'
>> stores but allow 'remote' loads - which becomes very relevant if
>> you allow accesses to have an affinity register for SMP support.
>>
>> This is another case where I am saying that you hardware people
>> should not provide the software people with everything you can,
>> because they will only use the extra power to shoot themselves in
>> their feet :-)
>
>I find this comment quite strange.
>
>I would find it very painful if FORTRAN didn't allow subscripted
>variables on the left-hand side of the equals sign in an assignment
>statement.

You have seriously misunderstood, both my point and the Fortran
standard!

In an assignment like <lvalue expr> = <rvalue expr>, almost all
languages specify that the left and right hand expressions must
be evaluated in toto before the assignment is performed.  In sane
languages, those expressions do not include stores, though I am
not damning the paradigm <lvalue expr 1> = <lvalue expr 2> =
<lvalue expr 3> = <rvalue expr>, which can be defined cleanly.

So the actual code is required by the standard to be, roughly:

    <reg 1> = <lvalue expr> (as an address)
    <reg 2> = <rvalue expr>
    (location pointed to by) <reg 1> = <reg 2>


Regards,
Nick Maclaren.

[toc] | [prev] | [next] | [standalone]


#2581

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-18 09:56 -0700
Message-ID<7f0bce2d-8c3d-4b7a-bda8-688d63d69497@df3g2000vbb.googlegroups.com>
In reply to#2578
On Jul 18, 9:22 am, n...@cam.ac.uk wrote:

> So the actual code is required by the standard to be, roughly:
>
>     <reg 1> = <lvalue expr> (as an address)
>     <reg 2> = <rvalue expr>
>     (location pointed to by) <reg 1> = <reg 2>

If a compiler that generates

 <reg 1> = subscript of lvalue, shifted left as required for type
 <reg 2> = <rvalue expr>
 store 2,array_address_constant(base reg,reg 1)

is in violation of the FORTRAN standard, despite producing
indistinguishable results, I'm sure that compiler writers happily
ignore that fact. But I am hoping that I am just seriously
misunderstanding your post once again.

John Savard

[toc] | [prev] | [next] | [standalone]


#2583

Fromnmm1@cam.ac.uk
Date2011-07-18 17:43 +0100
Message-ID<j01nqp$rrs$1@gosset.csi.cam.ac.uk>
In reply to#2581
In article <7f0bce2d-8c3d-4b7a-bda8-688d63d69497@df3g2000vbb.googlegroups.com>,
Quadibloc  <jsavard@ecn.ab.ca> wrote:
>
>> So the actual code is required by the standard to be, roughly:
>>
>>     <reg 1> = <lvalue expr> (as an address)
>>     <reg 2> = <rvalue expr>
>>     (location pointed to by) <reg 1> = <reg 2>
>
>If a compiler that generates
>
> <reg 1> = subscript of lvalue, shifted left as required for type
> <reg 2> = <rvalue expr>
> store 2,array_address_constant(base reg,reg 1)
>
>is in violation of the FORTRAN standard, despite producing
>indistinguishable results, I'm sure that compiler writers happily
>ignore that fact. But I am hoping that I am just seriously
>misunderstanding your post once again.

No, it isn't, but that wasn't either of my points.  The lesser of
them was that, because of the required semantics, the complexity
of stores generated by the compiler is less than that of loads
(as well as there being a lot fewer of them).

The greater point is that a few (fairly common) simples cases like
that can be optimised, but the general cases (especially for CISC
architectures) can't be.  For example, it is NOT possible to use
an indirection by memory store operation if there is any possibility
that either the base address or index might be aliased to the
location being stored into.  And that can be legal, even in Fortran,
and usually has to be assumed in C in the absence of global aliasing
analysis.

Even with RISC ISAs, that means that compilers sometimes have to
cache copies of locations where there is no apparent need to, and
failing to do so used to be a fairly common cause of foul code
generation bugs.  It tended to occur where there was a shortage
of registers, the compiler needed to store a complex object, and
it was more efficient to recalculate an address than use an extra
register or scratch memory location.  That didn't always work!

There is a related aspect, which particularly affects parallelism
and exception handling (INCLUDING at the hardware level).  A series
of reads can be retried ad lib in part and in any order (subject
only to target register independence, which is statically analysable),
but a series of stores is much trickier.  Indeed, once one has
indirection by memory store operations, they cannot be retried at
all, in general.

As a result, it is generally much cleaner to calculate the target
address as a single location if there is any doubt about what might
be going on.


Regards,
Nick Maclaren.

[toc] | [prev] | [next] | [standalone]


#2584

FromQuadibloc <jsavard@ecn.ab.ca>
Date2011-07-18 14:58 -0700
Message-ID<25e315c7-78a5-41cb-902e-b326d69a8ed4@e7g2000vbw.googlegroups.com>
In reply to#2583
On Jul 18, 10:43 am, n...@cam.ac.uk wrote:

> As a result, it is generally much cleaner to calculate the target
> address as a single location if there is any doubt about what might
> be going on.

Now I understand what you were intending to say. But the way I
understood what Andy Glew said, he was talking about omitting
capabilities from the architecture that would have been used in the
optimized simple case.

So I still don't agree that these FORTRAN considerations make it
reasonable to drop indexed addressing as a feature of store
instructions. Although I suppose it could be claimed, as you appear to
be doing, that doing so would force compiler writers to do things the
"safe" way more often.

John Savard

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.arch


csiph-web