Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compilers > #794 > unrolled thread

Green Compiler ?

Started byAbid <abidmuslim@gmail.com>
First post2012-12-20 02:00 -0800
Last post2012-12-28 10:42 -0500
Articles 20 on this page of 24 — 12 participants

Back to article view | Back to comp.compilers


Contents

  Green Compiler ? Abid <abidmuslim@gmail.com> - 2012-12-20 02:00 -0800
    Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-23 08:18 +0000
      Re: Green Compiler ? Peter Dassow <z80eu@arcor.de> - 2012-12-26 19:31 +0100
        Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-28 03:09 +0000
        Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-28 08:35 +0100
          Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-28 16:14 +0000
          Re: Green Compiler ? Peter Dassow <z80eu@arcor.de> - 2012-12-29 09:35 +0100
            Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-30 08:14 +0100
              Re: Green Compiler ? George Neuner <gneuner2@comcast.net> - 2012-12-31 01:24 -0500
                Re: Green Compiler ? "Jonathan Thornburg" <jthorn@astro.indiana.edu> - 2013-01-02 04:09 +0000
                  Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 18:29 +0000
                Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 05:11 +0000
                Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2013-01-02 07:29 +0100
                  Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 19:52 +0000
                    Re: Green Compiler ? "Charles Richmond" <numerist@aquaporin4.com> - 2013-01-04 08:59 -0600
    Re: Green Compiler ? "Nils M Holm" <nmh@t3x.org> - 2012-12-23 10:01 +0100
      Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-24 05:16 +0100
      Re: Green Compiler ? anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-27 13:36 +0000
        Re: Green Compiler ? "Nils M Holm" <nmh@t3x.org> - 2012-12-28 09:11 +0100
          Re: Green Compiler ? anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-28 16:57 +0000
          Re: Green Compiler ? George Neuner <gneuner2@comcast.net> - 2012-12-28 12:57 -0500
            Re: Green Compiler ? "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-12-30 09:22 +0100
      Re: Green Compiler ? Joshua Cranmer <Pidgeot18@verizon.invalid> - 2012-12-27 21:59 -0600
    Re: Green Compiler ? Walter Banks <walter@bytecraft.com> - 2012-12-28 10:42 -0500

Page 1 of 2  [1] 2  Next page →


#794 — Green Compiler ?

FromAbid <abidmuslim@gmail.com>
Date2012-12-20 02:00 -0800
SubjectGreen Compiler ?
Message-ID<12-12-010@comp.compilers>
Hi:

It seems that the Power Wall is becoming a major  issue, especially for High
Performance Computing. Current compilers work has two dimensional model, i.e.,
all the optimization phases are targeted either towards reducing execution
time or code size. My question is: Do we need to change this model and make it
three dimensional by adding power axis in the search space? If yes, then we
have to revisit all the phases and adjust them or come up with  new cost
models for these phases.

I will be grateful if some one can point me to the latest efforts/projects
with respect to power efficient compilation for high performance computing.

Thanks

[toc] | [next] | [standalone]


#796

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2012-12-23 08:18 +0000
Message-ID<12-12-012@comp.compilers>
In reply to#794
Abid <abidmuslim@gmail.com> wrote:

> It seems that the Power Wall is becoming a major  issue, especially for High
> Performance Computing. Current compilers work has two dimensional model, i.e.,
> all the optimization phases are targeted either towards reducing execution
> time or code size. My question is: Do we need to change this model and make it
> three dimensional by adding power axis in the search space? If yes, then we
> have to revisit all the phases and adjust them or come up with  new cost
> models for these phases.

It should be energy, not power. You can always reduce the power by
doing the calculation slower. Energy is power times time (more
generally, the integral of E(t)dt).

-- glen

[toc] | [prev] | [next] | [standalone]


#806

FromPeter Dassow <z80eu@arcor.de>
Date2012-12-26 19:31 +0100
Message-ID<12-12-022@comp.compilers>
In reply to#796
On 23.12.2012 09:18, glen herrmannsfeldt wrote:
> Abid <abidmuslim@gmail.com> wrote:
>
>> It seems that the Power Wall is becoming a major  issue, especially for High
>> Performance Computing. Current compilers work has two dimensional model, i.e.,
>> all the optimization phases are targeted either towards reducing execution
>> time or code size. My question is: Do we need to change this model and make it
>> three dimensional by adding power axis in the search space? If yes, then we
>> have to revisit all the phases and adjust them or come up with  new cost
>> models for these phases.
>
> It should be energy, not power. You can always reduce the power by
> doing the calculation slower. Energy is power times time (more
> generally, the integral of E(t)dt).

Efficiency is NOT equal to power consumption.  You can do calculations
(e.g. divide something) in more than one way, means by using the
"fastest" operations, or by using the "best" algorithm, means using
the most efficient opcodes (that is NOT using the fastest operations,
e.g. shifting/rotating bits), e.g. using a DIVIDE opcode instead which
can probably cost more clock cycles.  That does not mean just to make
it slow, you should still choose the most efficient way (means doing
something with fewer opcodes, regardless of how many clock cycles are
used for per opcode).  But I'm not sure this kind of optimization
would save a significant percentage compared to an clock cycle
optimized binary.

Regards
  Peter

[toc] | [prev] | [next] | [standalone]


#808

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2012-12-28 03:09 +0000
Message-ID<12-12-024@comp.compilers>
In reply to#806
Peter Dassow <z80eu@arcor.de> wrote:
>> Abid <abidmuslim@gmail.com> wrote:

>>> It seems that the Power Wall is becoming a major  issue, especially for High
>>> Performance Computing.

(snip, I wrote)
>> It should be energy, not power.

> Efficiency is NOT equal to power consumption.

Yes. As I noted, energy not power.

There is a fundamental tradeoff between speed and power.

It is pretty easy to see for CMOS, (well, until recently) where most
of the current is used to charge/discharge capacitors.  More current,
more power at constant supply voltage, charges the capacitors faster.
The energy stored in a capacitor is C*V*V/2, and that energy is
disippated when a gate is turned on or off.

For bipolar logic, such as TTL, it is the charge in the depletion
region that must be removed. Again more current (at fixed Vcc) will do
it faster. There are three families of TTL that differ only in their
supply current and speed.

> You can do calculations (e.g. divide something) in more
> than one way, means by using the "fastest" operations, or by
> using the "best" algorithm, means using the most efficient
> opcodes (that is NOT using the fastest operations,
> e.g. shifting/rotating bits), e.g. using a DIVIDE opcode
> instead which can probably cost more clock cycles.

Until recently, CMOS had almost zero quiescent current.  The current
was almost directly related to gates switching, and so power directly
related to the rate of gate switching.

More recently, at the smallest feature size and gate oxide thickness,
quantum tunneling through the gate oxide has become a significant
current (power) drain.

Even more, the timing characteristics of transistors are changing:

http://www.technion.ac.il/~sbeer/publications/p3.pdf


> That does not mean just to make it slow, you should still
> choose the most efficient way (means doing something with fewer
> opcodes, regardless of how many clock cycles are used for per
> opcode).  But I'm not sure this kind of optimization
> would save a significant percentage compared to an clock cycle
> optimized binary.

ECL and for the most part TTL, run at constant power. Faster
computation (everything else constant) means less energy used.

As above, for CMOS until recently, energy is related to the gates
switching. Processors designs have improved in not using (usually not
even clocking) parts that aren't needed. Turn off the floating point
unit when it isn't needed.

A divide instruction might take the same amount of time as a shift,
but consume much more energy (power * time).

-- glen

[toc] | [prev] | [next] | [standalone]


#812

FromHans-Peter Diettrich <DrDiettrich1@aol.com>
Date2012-12-28 08:35 +0100
Message-ID<12-12-028@comp.compilers>
In reply to#806
Peter Dassow schrieb:

> But I'm not sure this kind of optimization
> would save a significant percentage compared to an clock cycle
> optimized binary.

I doubt that a clock cycle count is related to power consumption.
On modern CPUs, with much parallel processing, parts of the chip may be
inactive while others are operating. E.g. when two instructions execute
in parallel, they will consume about the same amount of power as when
executed sequentially.

Please note that most CMOS processor power consumption results from
switching (stray) capacities, and only a small percentage for leak
currents. E.g. a register or gate consumes such power whenever a bit is
changed, and almost nothing when it has reached an stable state.

DoDi

[toc] | [prev] | [next] | [standalone]


#815

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2012-12-28 16:14 +0000
Message-ID<12-12-031@comp.compilers>
In reply to#812
Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote:

(snip)
> Please note that most CMOS processor power consumption results from
> switching (stray) capacities, and only a small percentage for leak
> currents. E.g. a register or gate consumes such power whenever a bit is
> changed, and almost nothing when it has reached an stable state.

Used to be true. Now, quantum tunneling through the oxide is a
very significant quiescent current for CMOS.

-- glen

[toc] | [prev] | [next] | [standalone]


#818

FromPeter Dassow <z80eu@arcor.de>
Date2012-12-29 09:35 +0100
Message-ID<12-12-034@comp.compilers>
In reply to#812
On 28.12.2012 08:35, Hans-Peter Diettrich wrote:
>
> Please note that most CMOS processor power consumption results from
> switching (stray) capacities, and only a small percentage for leak
> currents. E.g. a register or gate consumes such power whenever a bit is
> changed, and almost nothing when it has reached an stable state.
>
So using extensively registers instead of "conventional" memory (e.g.
DDR-RAM, memory outside a CPU) will save energy (if equal functionality
is given) ?

Regards
  Peter

[toc] | [prev] | [next] | [standalone]


#821

FromHans-Peter Diettrich <DrDiettrich1@aol.com>
Date2012-12-30 08:14 +0100
Message-ID<12-12-037@comp.compilers>
In reply to#818
Peter Dassow schrieb:
> On 28.12.2012 08:35, Hans-Peter Diettrich wrote:
>>
>> Please note that most CMOS processor power consumption results from
>> switching (stray) capacities, and only a small percentage for leak
>> currents. E.g. a register or gate consumes such power whenever a bit is
>> changed, and almost nothing when it has reached an stable state.
>>
> So using extensively registers instead of "conventional" memory (e.g.
> DDR-RAM, memory outside a CPU) will save energy (if equal functionality
> is given) ?

I don't see a relationship here, except that external memory is slow
and [in x86] a couple of caches and address translations are involved
in reading from RAM. But registers are a very scarce resource, so that
frequent loading from memory is hardly avoidable. Register usage is
already optimized by every (good) compiler, reducing runtime and power
consuption at the same time. This was a very narrow bottleneck of the
IA-32 architecture, until the 64 bit (AMD) model increased the number
of registers considerably (16 addressable, of 128 [256?] shadow
registers).  The CPU computation circuitry and usage is the same,
regardless of the operand sources.

DoDi

[toc] | [prev] | [next] | [standalone]


#825

FromGeorge Neuner <gneuner2@comcast.net>
Date2012-12-31 01:24 -0500
Message-ID<13-01-002@comp.compilers>
In reply to#821
On Sun, 30 Dec 2012 08:14:26 +0100, Hans-Peter Diettrich
<DrDiettrich1@aol.com> wrote:

>Peter Dassow schrieb:
>> On 28.12.2012 08:35, Hans-Peter Diettrich wrote:
>>>
>>> Please note that most CMOS processor power consumption results from
>>> switching (stray) capacities, and only a small percentage for leak
>>> currents. E.g. a register or gate consumes such power whenever a bit is
>>> changed, and almost nothing when it has reached an stable state.

Smaller transistors have more leakage.

>> So using extensively registers instead of "conventional" memory (e.g.
>> DDR-RAM, memory outside a CPU) will save energy (if equal functionality
>> is given) ?
>
>I don't see a relationship here, except that external memory is slow
>and [in x86] a couple of caches and address translations are involved
>in reading from RAM.

But associative cache and external memory accesses both are power
intensive.

>But registers are a very scarce resource, so that frequent loading from
>memory is hardly avoidable.

My opinions are colored by experience with DSPs, but I have long
thought that it would be helpful to have a few K-words of non-cache
scratchpad memory very close (1..2 cycles) to the CPU.

George

[toc] | [prev] | [next] | [standalone]


#826

From"Jonathan Thornburg" <jthorn@astro.indiana.edu>
Date2013-01-02 04:09 +0000
Message-ID<13-01-003@comp.compilers>
In reply to#825
George Neuner <gneuner2@comcast.net> wrote:
>>But registers are a very scarce resource, so that frequent loading from
>>memory is hardly avoidable.
>
> My opinions are colored by experience with DSPs, but I have long
> thought that it would be helpful to have a few K-words of non-cache
> scratchpad memory very close (1..2 cycles) to the CPU.

The Cray-2 and Cray-3 had this (I think they called it "local memory").
Alas, compilers had a lot of trouble making good use of it.  As John
McCalpin put it in 1998 (at the time he was at SGI):

# Local memories work fine if an experience performance person gets
# to rewrite the code.  It gets a lot more difficult if you have
# to educate a compiler to do the right thing....

I don't know if compilers have improved significantly in this area.

And on the hardware side, you'd have to figure out what to do on
context switch...
--
-- "Jonathan Thornburg [remove -animal to reply]" <jthorn@astro.indiana-zebra.edu>
   Dept of Astronomy, Indiana University, Bloomington, Indiana, USA
   "C++ is to programming as sex is to reproduction. Better ways might
    technically exist but they're not nearly as much fun." -- Nikolai Irgens
[It sounds like if we keep working, we might reinvent register windows. -John]

[toc] | [prev] | [next] | [standalone]


#831

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2013-01-02 18:29 +0000
Message-ID<13-01-008@comp.compilers>
In reply to#826
Jonathan Thornburg <jthorn@astro.indiana.edu> wrote:
> George Neuner <gneuner2@comcast.net> wrote:

(snip)
>> My opinions are colored by experience with DSPs, but I have long
>> thought that it would be helpful to have a few K-words of non-cache
>> scratchpad memory very close (1..2 cycles) to the CPU.

> The Cray-2 and Cray-3 had this (I think they called it "local memory").
> Alas, compilers had a lot of trouble making good use of it.  As John
> McCalpin put it in 1998 (at the time he was at SGI):

I previously suggested the Cray-1 vector registers, maybe not quite
the same. Still, the being able to move around multiple words in one
operation, especially to get them through a pipelined ALU, seems
a good one.

> # Local memories work fine if an experience performance person gets
> # to rewrite the code.  It gets a lot more difficult if you have
> # to educate a compiler to do the right thing....

I do remember, I believe from the Cray-1 days, suggestions of having
the linker allocate registers. That is, similar to the way memory
relocation is done at link time, vector registers would be assigned
to minimize (as well as can be done statically) register spill.

Now, that was before Fortran had dynamic allocation. (And Fortran
was the popular language for Cray programming.)

> I don't know if compilers have improved significantly in this area.

Compiler technology (and size) has change much over the years.

One ongoing question is how well compilers are doing with Fortran
array expressions. It would seem easier for compilers to generate
code using vector registers for array operations.

> And on the hardware side, you'd have to figure out what to do on
> context switch...

Seems to me that the idea behind batch programming (as contrasted
to time-sharing) is keeping context switch low. With a little luck
you don't have to save the vector registers (or local store) for
every interrupt.

(snip, our moderator wrote)
> [It sounds like if we keep working, we might reinvent register windows. -John]

Is SPARC dead now? I was never sure how well the SPARC register windows,
with in and out registers, actually worked for real code.
That is, if they made the best use of chip resources.

I presume they help more for procedure call register saving than for
contect switch, though.

-- glen
[SPARC is still around, mostly seems to be Sun selling servers to run parent
Oracle's software.  Register windows are awful for context switch, since you
have to dump and restore the whole register stack on each switch if you don't
have some hack like multiple stacks for active processes.  It is also my
impression that SPARC has register windows because it was designed to run code
from PCC, a compiler that did very naive register allocation.  The IBM 801
project was at the same time, and its chip had an ordinary register file because
they found that the compiler could generate code to do the saves and restores
at least as well as the implicit ones that windows do. -John]

[toc] | [prev] | [next] | [standalone]


#827

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2013-01-02 05:11 +0000
Message-ID<13-01-004@comp.compilers>
In reply to#825
George Neuner <gneuner2@comcast.net> wrote:

(snip)
> My opinions are colored by experience with DSPs, but I have long
> thought that it would be helpful to have a few K-words of non-cache
> scratchpad memory very close (1..2 cycles) to the CPU.

You mean like the vector registers of the Cray-1?

It seems that vector processors are not out of fashion, though not so
obvious that they couldn't come back.

They work especially well with the pipelined ALU that the Cray
series used.

-- glen
[It is my impression that vector processors went out of fashion when
they figured out how to build instruction pipelines deep enough to
keep execution units busy with scalar code.  The context switch
problems were and are severe, unless you want to manage a pool
of scratchpads for active processes, and spill as needed, sort of
like the way x86 processors do floating point register management.
-John]

[toc] | [prev] | [next] | [standalone]


#828

FromHans-Peter Diettrich <DrDiettrich1@aol.com>
Date2013-01-02 07:29 +0100
Message-ID<13-01-005@comp.compilers>
In reply to#825
George Neuner schrieb:
> On Sun, 30 Dec 2012 08:14:26 +0100, Hans-Peter Diettrich
> <DrDiettrich1@aol.com> wrote:
>
>> Peter Dassow schrieb:
>>> On 28.12.2012 08:35, Hans-Peter Diettrich wrote:
>>>> Please note that most CMOS processor power consumption results from
>>>> switching (stray) capacities, and only a small percentage for leak
>>>> currents. E.g. a register or gate consumes such power whenever a bit is
>>>> changed, and almost nothing when it has reached an stable state.
>
> Smaller transistors have more leakage.

ACK (tunneling effect).

>>> So using extensively registers instead of "conventional" memory (e.g.
>>> DDR-RAM, memory outside a CPU) will save energy (if equal functionality
>>> is given) ?
>> I don't see a relationship here, except that external memory is slow
>> and [in x86] a couple of caches and address translations are involved
>> in reading from RAM.
>
> But associative cache and external memory accesses both are power
> intensive.

Right, but since every reference to external memory costs *time* in
the first place, *every* compiler already optimizes for best register
usage.  There is nothing that can be done *additionally* in a "green"
compiler.  C already has a "register" keyword, as a compiler hint to
hold a local variable in an register.

The actual use of registers depends on the control flow taken
*actually*. When a subroutine is optimized for using all available
registers, it has to save and restore the registers on entry/exit.
When it actually does nothing, due to given conditions, the time and
energy used for pushing/popping the registers is only wasted.

For that reason some (Texas Instruments?) processors implemented a
register stack, decades ago, with its stack pointer adjusted according
to the number of registers used in a subroutine. This stack could be
moved into the CPU nowadays, eliminating the need for saving registers
in external memory. But this optimization reaches a hard limit on
deeply nested calls, with every subroutine using a high number of
registers. In external memory the register-stack size is adjustable to
program needs, just like ordinary stack size is, but a CPU resident
register stack has a fixed depth. Eventually the register stack still
could be kept in RAM, with an dedicated cache equivalent to the L1/L2
caches. Then the caches would automatically push/pop register contents
depending on their actual *use*, not by fixed push/pop *instruction
sequences*. OTOH we already have nested caches, so that the effect of
an additional register cache is questionable. (see Wikipedia "CPU
cache")

The x86 architecture uses another approach (register renaming), with a
high number of shadow registers (compared to only 16 addressable
registers). I'm not sure, though, how a compiler should generate code
for best use of that model...


>> But registers are a very scarce resource, so that frequent loading from
>> memory is hardly avoidable.
>
> My opinions are colored by experience with DSPs, but I have long
> thought that it would be helpful to have a few K-words of non-cache
> scratchpad memory very close (1..2 cycles) to the CPU.

ACK, but see above considerations on the use of such memory, with its
*size* limited by the architecture, and *usage* depending on subroutine
needs and control flow (subroutine nesting, branches taken...).

DoDi
[There were stacks with the top few registers kept in fast memory in
the Burroughs machines in the 1960s.  It was easy to generate code for
them, but since register coloring was invented in the 1970s, modern
code scheduling for normal registers is much more effective.  -John]

[toc] | [prev] | [next] | [standalone]


#832

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2013-01-02 19:52 +0000
Message-ID<13-01-009@comp.compilers>
In reply to#828
Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote:

(snip, someone wrote)
>> Smaller transistors have more leakage.

> ACK (tunneling effect).

(snip)

> The actual use of registers depends on the control flow taken
> *actually*. When a subroutine is optimized for using all available
> registers, it has to save and restore the registers on entry/exit.
> When it actually does nothing, due to given conditions, the time and
> energy used for pushing/popping the registers is only wasted.

Some use caller saves, some callee saves. In the latter case, the
subroutine knows which registers it will change, and only needs to
save those. It is usual to do that save on entry and restore just
before return, but with a little extra logic the save might be delayed
a little bit.

> For that reason some (Texas Instruments?) processors implemented a
> register stack, decades ago, with its stack pointer adjusted according
> to the number of registers used in a subroutine. This stack could be
> moved into the CPU nowadays, eliminating the need for saving registers
> in external memory.

If I remember the TMS9900, the registers were in memory, so that
pointer just changed where in memory they were.

> But this optimization reaches a hard limit on
> deeply nested calls, with every subroutine using a high number of
> registers. In external memory the register-stack size is adjustable to
> program needs, just like ordinary stack size is, but a CPU resident
> register stack has a fixed depth. Eventually the register stack still
> could be kept in RAM, with an dedicated cache equivalent to the L1/L2
> caches.

The original idea behind the 8087 register stack was that it would be
interrupt driven, spill on overflow, restore on underflow.  The
problem was that the logic wasn't tested before the chip was built,
and then it was too late.  Also, it has only 8 entries.

> Then the caches would automatically push/pop register contents
> depending on their actual *use*, not by fixed push/pop *instruction
> sequences*. OTOH we already have nested caches, so that the effect of
> an additional register cache is questionable. (see Wikipedia "CPU
> cache")

A larger register stack, with working automatic spill/restore,
and a fast enough interrupt handler, might work.

> The x86 architecture uses another approach (register renaming), with a
> high number of shadow registers (compared to only 16 addressable
> registers). I'm not sure, though, how a compiler should generate code
> for best use of that model...

I first knew about register renaming for the floating point unit
of the IBM 360/91. It was a popular example machine for many generations
of books on pipelined processors. With only four floating point
registers, renaming was pretty important for out-of-order execution
for S/360. That was especially true since the 360/91 was designed
to be able to run code not specifically optimized for it.

RISC processors tend to expect compilers to order instructions
appropriately, and also usually have plenty of registers.

(snip)
> [There were stacks with the top few registers kept in fast memory in
> the Burroughs machines in the 1960s.  It was easy to generate code for
> them, but since register coloring was invented in the 1970s, modern
> code scheduling for normal registers is much more effective.  -John]

Hmm. Not that I completely understand it, but it seems to me that
normal general registers are generally used over shorter distances.
Larger ones, such as vector registers, take longer to save and restore,
and might also need to store values for a longer time. It could help
much not to have to save/restore for every interrupt, for example.

-- glen

[toc] | [prev] | [next] | [standalone]


#838

From"Charles Richmond" <numerist@aquaporin4.com>
Date2013-01-04 08:59 -0600
Message-ID<13-01-015@comp.compilers>
In reply to#832
"glen herrmannsfeldt" <gah@ugcs.caltech.edu> wrote in message
> Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote:
>
>     [snip...]                     [snip...]                     [snip...]
>
>> For that reason some (Texas Instruments?) processors implemented a
>> register stack, decades ago, with its stack pointer adjusted according
>> to the number of registers used in a subroutine. This stack could be
>> moved into the CPU nowadays, eliminating the need for saving registers
>> in external memory.
>
> If I remember the TMS9900, the registers were in memory, so that
> pointer just changed where in memory they were.

The TMS9900 did allocate the register "file" in memory, with a pointer
in the CPU to indicate the position of the register file.  Some
versions of the chip had some on-board RAM that could be used for
register allocation.  This RAM was faster to access and would speed
along execution.

--

numerist at aquaporin4 dot com

[toc] | [prev] | [next] | [standalone]


#797

From"Nils M Holm" <nmh@t3x.org>
Date2012-12-23 10:01 +0100
Message-ID<12-12-013@comp.compilers>
In reply to#794
Abid <abidmuslim@gmail.com> wrote:
> Do we need to change this model and make it
> three dimensional by adding power axis in the search space?

In principle, I would say that higher execution speed equals more
grenn-ness. Less time spent dissipating heat means less energy
consumed.

So the question actually is the same as ever: how to we make code
run fast?

- How do we squeeze more meaning into fewer/faster instructions?
- How do we make code so small that it fits in caches?
- How do we organize code execution in such a way that cache
  stalls are minimized? (Locality)

--
Nils M Holm  < n m h @ t 3 x . o r g >  www.t3x.org

[toc] | [prev] | [next] | [standalone]


#801

FromHans-Peter Diettrich <DrDiettrich1@aol.com>
Date2012-12-24 05:16 +0100
Message-ID<12-12-017@comp.compilers>
In reply to#797
Nils M Holm schrieb:
> Abid <abidmuslim@gmail.com> wrote:
>> Do we need to change this model and make it
>> three dimensional by adding power axis in the search space?
>
> In principle, I would say that higher execution speed equals more
> grenn-ness. Less time spent dissipating heat means less energy
> consumed.

Not really. You mean more efficient code, I suppose?

> So the question actually is the same as ever: how to we make code
> run fast?
>
> - How do we squeeze more meaning into fewer/faster instructions?

Okay, but the available machine instructions are limited. In detail on
RISC architectures.

> - How do we make code so small that it fits in caches?

It's not only the code that has to fit into the caches. Instruction
reordering and branch prediction are one thing, but the placement of the
data (variables) is quite another story.

> - How do we organize code execution in such a way that cache
>   stalls are minimized? (Locality)

Such optimization requires compiler hints, so that the compiler can not
only arrange the code, but also the data. This requires knowledge about
the *frequency* of subroutine use, branches taken, and variables used in
these critical pathes.

But who should provide such hints? IMO only a profiling tool can provide
that information, when a program is run with typical (real life) data.

DoDi

[toc] | [prev] | [next] | [standalone]


#807

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2012-12-27 13:36 +0000
Message-ID<12-12-023@comp.compilers>
In reply to#797
"Nils M Holm" <nmh@t3x.org> writes:
>In principle, I would say that higher execution speed equals more
>grenn-ness. Less time spent dissipating heat means less energy
>consumed.

If a computer runs only a given set of batch-style programs a given
set of times, yes.  But that's not how computers are used.  E.g.,
there are games that consume all the resources they are given.  If the
game code runs faster, it just means that the game has better FPS, not
that the computer consumes less energy.  There are also anti-virus
programs etc., and even for batch-style programs, if they run faster,
one often runs them more often.  So it's not so easy.

- anton
--
M. Anton Ertl
anton@mips.complang.tuwien.ac.at
http://www.complang.tuwien.ac.at/anton/

[toc] | [prev] | [next] | [standalone]


#811

From"Nils M Holm" <nmh@t3x.org>
Date2012-12-28 09:11 +0100
Message-ID<12-12-027@comp.compilers>
In reply to#807
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
> "Nils M Holm" <nmh@t3x.org> writes:
> >In principle, I would say that higher execution speed equals more
> >green-ness. Less time spent dissipating heat means less energy
> >consumed.
>
> If a computer runs only a given set of batch-style programs a given
> set of times, yes.  But that's not how computers are used.  E.g.,
> there are games that consume all the resources they are given.  If the
> game code runs faster, it just means that the game has better FPS, not
> that the computer consumes less energy.  There are also anti-virus
> programs etc., and even for batch-style programs, if they run faster,
> one often runs them more often.  So it's not so easy.

So you are basically arguing that people always try to get the maximum
out of their resources and therefore faster code is worse. Indeed, if
this should be the case, you are right.

Your first example was games. Why does a game *have to* use the
maximum number of frames per second? When a given number of frames per
second is sufficient, in the sense that additional frames will not
make the video output any smoother, why increase the frame rate
further? Let the program sleep between frames. In this case fast code
can sleep between frames, slow code can't.

Same for virus scanner: why run them more often?

So I would argue that the issue you raise is a psychological one rather
than a technical one. When people always want to get the maximum out
of their resources, then there is no way to reduce power consumption,
except, maybe, by creating *slower* processors.

Creating slower compiler output would not help, because it will just
burn more CPU cycles to achieve less. Faster code is better, because
it achieves the same in less time. E.g.: a slow interpreter can run
the same algorithm as a fast, compiled executable. Both achieve the
same on the same CPU, but the interpreted code takes longer and hence
wastes energy. So faster code saves energy.

Of course, this point assumes that people want a fixed outcome at the
least cost, and I see that this is often not the case. It seems to
get more and more interesting though, with all those "green" processors,
and other "green" equipment around.

Bottom line: we can create the tools to save energy, but if people do
not want to save energy, nothing will help.

--
Nils M Holm  < n m h @ t 3 x . o r g >  www.t3x.org

[toc] | [prev] | [next] | [standalone]


#816

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2012-12-28 16:57 +0000
Message-ID<12-12-032@comp.compilers>
In reply to#811
"Nils M Holm" <nmh@t3x.org> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> "Nils M Holm" <nmh@t3x.org> writes:
>> >In principle, I would say that higher execution speed equals more
>> >green-ness. Less time spent dissipating heat means less energy
>> >consumed.
...
>So you are basically arguing that people always try to get the maximum
>out of their resources and therefore faster code is worse.

Not worse, but the claims you made above are a bit too simple-minded.
Fast code still has the advantage of being fast; but it does not
necessarily save energy.

>Your first example was games. Why does a game *have to* use the
>maximum number of frames per second? When a given number of frames per
>second is sufficient, in the sense that additional frames will not
>make the video output any smoother, why increase the frame rate
>further? Let the program sleep between frames. In this case fast code
>can sleep between frames, slow code can't.

But do game designers work that way?  If the game engine is so fast
that it can sleep between frames, they put in more features in the
game that make use of the CPU; or the reverse, they don't take as many
features out in the slimming phase when they try to make the game run
smooth enough.  So the faster code buys a more featureful game, but it
does not save energy.

Or at the level of the end user (player), they tune the graphics
details such that they get the maximum details at an acceptable frame
rate on their machine.  So a faster program means more details, not
less energy consumption.

>Same for virus scanner: why run them more often?

More often?  AFAIK, these days the virus scanners are running
constantly in the background and consume one or two cores.  I don't
know if that has a technical reason or is just done to give the users
a feeling of better protection, but that does not matter for the
energy consumption.

>So I would argue that the issue you raise is a psychological one rather
>than a technical one. When people always want to get the maximum out
>of their resources, then there is no way to reduce power consumption,
>except, maybe, by creating *slower* processors.

Or at least processors and machines with lower power consumption under
load (yes, they tend to be slower, the converse is not necessarily
true).  Yes, you can see it as a psychological issue, but that does
not make it less real.  Idle power consumption is also an issue, for
some workloads.

- anton
--
M. Anton Ertl
anton@mips.complang.tuwien.ac.at
http://www.complang.tuwien.ac.at/anton/

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | comp.compilers


csiph-web