Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.compilers > #794 > unrolled thread
| Started by | Abid <abidmuslim@gmail.com> |
|---|---|
| First post | 2012-12-20 02:00 -0800 |
| Last post | 2012-12-28 10:42 -0500 |
| Articles | 20 on this page of 24 — 12 participants |
Back to article view | Back to comp.compilers
Green Compiler ? Abid <abidmuslim@gmail.com> - 2012-12-20 02:00 -0800
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-23 08:18 +0000
Re: Green Compiler ? Peter Dassow <z80eu@arcor.de> - 2012-12-26 19:31 +0100
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-28 03:09 +0000
Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-28 08:35 +0100
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2012-12-28 16:14 +0000
Re: Green Compiler ? Peter Dassow <z80eu@arcor.de> - 2012-12-29 09:35 +0100
Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-30 08:14 +0100
Re: Green Compiler ? George Neuner <gneuner2@comcast.net> - 2012-12-31 01:24 -0500
Re: Green Compiler ? "Jonathan Thornburg" <jthorn@astro.indiana.edu> - 2013-01-02 04:09 +0000
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 18:29 +0000
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 05:11 +0000
Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2013-01-02 07:29 +0100
Re: Green Compiler ? glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-01-02 19:52 +0000
Re: Green Compiler ? "Charles Richmond" <numerist@aquaporin4.com> - 2013-01-04 08:59 -0600
Re: Green Compiler ? "Nils M Holm" <nmh@t3x.org> - 2012-12-23 10:01 +0100
Re: Green Compiler ? Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-12-24 05:16 +0100
Re: Green Compiler ? anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-27 13:36 +0000
Re: Green Compiler ? "Nils M Holm" <nmh@t3x.org> - 2012-12-28 09:11 +0100
Re: Green Compiler ? anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-28 16:57 +0000
Re: Green Compiler ? George Neuner <gneuner2@comcast.net> - 2012-12-28 12:57 -0500
Re: Green Compiler ? "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-12-30 09:22 +0100
Re: Green Compiler ? Joshua Cranmer <Pidgeot18@verizon.invalid> - 2012-12-27 21:59 -0600
Re: Green Compiler ? Walter Banks <walter@bytecraft.com> - 2012-12-28 10:42 -0500
Page 1 of 2 [1] 2 Next page →
| From | Abid <abidmuslim@gmail.com> |
|---|---|
| Date | 2012-12-20 02:00 -0800 |
| Subject | Green Compiler ? |
| Message-ID | <12-12-010@comp.compilers> |
Hi: It seems that the Power Wall is becoming a major issue, especially for High Performance Computing. Current compilers work has two dimensional model, i.e., all the optimization phases are targeted either towards reducing execution time or code size. My question is: Do we need to change this model and make it three dimensional by adding power axis in the search space? If yes, then we have to revisit all the phases and adjust them or come up with new cost models for these phases. I will be grateful if some one can point me to the latest efforts/projects with respect to power efficient compilation for high performance computing. Thanks
[toc] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2012-12-23 08:18 +0000 |
| Message-ID | <12-12-012@comp.compilers> |
| In reply to | #794 |
Abid <abidmuslim@gmail.com> wrote: > It seems that the Power Wall is becoming a major issue, especially for High > Performance Computing. Current compilers work has two dimensional model, i.e., > all the optimization phases are targeted either towards reducing execution > time or code size. My question is: Do we need to change this model and make it > three dimensional by adding power axis in the search space? If yes, then we > have to revisit all the phases and adjust them or come up with new cost > models for these phases. It should be energy, not power. You can always reduce the power by doing the calculation slower. Energy is power times time (more generally, the integral of E(t)dt). -- glen
[toc] | [prev] | [next] | [standalone]
| From | Peter Dassow <z80eu@arcor.de> |
|---|---|
| Date | 2012-12-26 19:31 +0100 |
| Message-ID | <12-12-022@comp.compilers> |
| In reply to | #796 |
On 23.12.2012 09:18, glen herrmannsfeldt wrote: > Abid <abidmuslim@gmail.com> wrote: > >> It seems that the Power Wall is becoming a major issue, especially for High >> Performance Computing. Current compilers work has two dimensional model, i.e., >> all the optimization phases are targeted either towards reducing execution >> time or code size. My question is: Do we need to change this model and make it >> three dimensional by adding power axis in the search space? If yes, then we >> have to revisit all the phases and adjust them or come up with new cost >> models for these phases. > > It should be energy, not power. You can always reduce the power by > doing the calculation slower. Energy is power times time (more > generally, the integral of E(t)dt). Efficiency is NOT equal to power consumption. You can do calculations (e.g. divide something) in more than one way, means by using the "fastest" operations, or by using the "best" algorithm, means using the most efficient opcodes (that is NOT using the fastest operations, e.g. shifting/rotating bits), e.g. using a DIVIDE opcode instead which can probably cost more clock cycles. That does not mean just to make it slow, you should still choose the most efficient way (means doing something with fewer opcodes, regardless of how many clock cycles are used for per opcode). But I'm not sure this kind of optimization would save a significant percentage compared to an clock cycle optimized binary. Regards Peter
[toc] | [prev] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2012-12-28 03:09 +0000 |
| Message-ID | <12-12-024@comp.compilers> |
| In reply to | #806 |
Peter Dassow <z80eu@arcor.de> wrote: >> Abid <abidmuslim@gmail.com> wrote: >>> It seems that the Power Wall is becoming a major issue, especially for High >>> Performance Computing. (snip, I wrote) >> It should be energy, not power. > Efficiency is NOT equal to power consumption. Yes. As I noted, energy not power. There is a fundamental tradeoff between speed and power. It is pretty easy to see for CMOS, (well, until recently) where most of the current is used to charge/discharge capacitors. More current, more power at constant supply voltage, charges the capacitors faster. The energy stored in a capacitor is C*V*V/2, and that energy is disippated when a gate is turned on or off. For bipolar logic, such as TTL, it is the charge in the depletion region that must be removed. Again more current (at fixed Vcc) will do it faster. There are three families of TTL that differ only in their supply current and speed. > You can do calculations (e.g. divide something) in more > than one way, means by using the "fastest" operations, or by > using the "best" algorithm, means using the most efficient > opcodes (that is NOT using the fastest operations, > e.g. shifting/rotating bits), e.g. using a DIVIDE opcode > instead which can probably cost more clock cycles. Until recently, CMOS had almost zero quiescent current. The current was almost directly related to gates switching, and so power directly related to the rate of gate switching. More recently, at the smallest feature size and gate oxide thickness, quantum tunneling through the gate oxide has become a significant current (power) drain. Even more, the timing characteristics of transistors are changing: http://www.technion.ac.il/~sbeer/publications/p3.pdf > That does not mean just to make it slow, you should still > choose the most efficient way (means doing something with fewer > opcodes, regardless of how many clock cycles are used for per > opcode). But I'm not sure this kind of optimization > would save a significant percentage compared to an clock cycle > optimized binary. ECL and for the most part TTL, run at constant power. Faster computation (everything else constant) means less energy used. As above, for CMOS until recently, energy is related to the gates switching. Processors designs have improved in not using (usually not even clocking) parts that aren't needed. Turn off the floating point unit when it isn't needed. A divide instruction might take the same amount of time as a shift, but consume much more energy (power * time). -- glen
[toc] | [prev] | [next] | [standalone]
| From | Hans-Peter Diettrich <DrDiettrich1@aol.com> |
|---|---|
| Date | 2012-12-28 08:35 +0100 |
| Message-ID | <12-12-028@comp.compilers> |
| In reply to | #806 |
Peter Dassow schrieb: > But I'm not sure this kind of optimization > would save a significant percentage compared to an clock cycle > optimized binary. I doubt that a clock cycle count is related to power consumption. On modern CPUs, with much parallel processing, parts of the chip may be inactive while others are operating. E.g. when two instructions execute in parallel, they will consume about the same amount of power as when executed sequentially. Please note that most CMOS processor power consumption results from switching (stray) capacities, and only a small percentage for leak currents. E.g. a register or gate consumes such power whenever a bit is changed, and almost nothing when it has reached an stable state. DoDi
[toc] | [prev] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2012-12-28 16:14 +0000 |
| Message-ID | <12-12-031@comp.compilers> |
| In reply to | #812 |
Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote: (snip) > Please note that most CMOS processor power consumption results from > switching (stray) capacities, and only a small percentage for leak > currents. E.g. a register or gate consumes such power whenever a bit is > changed, and almost nothing when it has reached an stable state. Used to be true. Now, quantum tunneling through the oxide is a very significant quiescent current for CMOS. -- glen
[toc] | [prev] | [next] | [standalone]
| From | Peter Dassow <z80eu@arcor.de> |
|---|---|
| Date | 2012-12-29 09:35 +0100 |
| Message-ID | <12-12-034@comp.compilers> |
| In reply to | #812 |
On 28.12.2012 08:35, Hans-Peter Diettrich wrote: > > Please note that most CMOS processor power consumption results from > switching (stray) capacities, and only a small percentage for leak > currents. E.g. a register or gate consumes such power whenever a bit is > changed, and almost nothing when it has reached an stable state. > So using extensively registers instead of "conventional" memory (e.g. DDR-RAM, memory outside a CPU) will save energy (if equal functionality is given) ? Regards Peter
[toc] | [prev] | [next] | [standalone]
| From | Hans-Peter Diettrich <DrDiettrich1@aol.com> |
|---|---|
| Date | 2012-12-30 08:14 +0100 |
| Message-ID | <12-12-037@comp.compilers> |
| In reply to | #818 |
Peter Dassow schrieb: > On 28.12.2012 08:35, Hans-Peter Diettrich wrote: >> >> Please note that most CMOS processor power consumption results from >> switching (stray) capacities, and only a small percentage for leak >> currents. E.g. a register or gate consumes such power whenever a bit is >> changed, and almost nothing when it has reached an stable state. >> > So using extensively registers instead of "conventional" memory (e.g. > DDR-RAM, memory outside a CPU) will save energy (if equal functionality > is given) ? I don't see a relationship here, except that external memory is slow and [in x86] a couple of caches and address translations are involved in reading from RAM. But registers are a very scarce resource, so that frequent loading from memory is hardly avoidable. Register usage is already optimized by every (good) compiler, reducing runtime and power consuption at the same time. This was a very narrow bottleneck of the IA-32 architecture, until the 64 bit (AMD) model increased the number of registers considerably (16 addressable, of 128 [256?] shadow registers). The CPU computation circuitry and usage is the same, regardless of the operand sources. DoDi
[toc] | [prev] | [next] | [standalone]
| From | George Neuner <gneuner2@comcast.net> |
|---|---|
| Date | 2012-12-31 01:24 -0500 |
| Message-ID | <13-01-002@comp.compilers> |
| In reply to | #821 |
On Sun, 30 Dec 2012 08:14:26 +0100, Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote: >Peter Dassow schrieb: >> On 28.12.2012 08:35, Hans-Peter Diettrich wrote: >>> >>> Please note that most CMOS processor power consumption results from >>> switching (stray) capacities, and only a small percentage for leak >>> currents. E.g. a register or gate consumes such power whenever a bit is >>> changed, and almost nothing when it has reached an stable state. Smaller transistors have more leakage. >> So using extensively registers instead of "conventional" memory (e.g. >> DDR-RAM, memory outside a CPU) will save energy (if equal functionality >> is given) ? > >I don't see a relationship here, except that external memory is slow >and [in x86] a couple of caches and address translations are involved >in reading from RAM. But associative cache and external memory accesses both are power intensive. >But registers are a very scarce resource, so that frequent loading from >memory is hardly avoidable. My opinions are colored by experience with DSPs, but I have long thought that it would be helpful to have a few K-words of non-cache scratchpad memory very close (1..2 cycles) to the CPU. George
[toc] | [prev] | [next] | [standalone]
| From | "Jonathan Thornburg" <jthorn@astro.indiana.edu> |
|---|---|
| Date | 2013-01-02 04:09 +0000 |
| Message-ID | <13-01-003@comp.compilers> |
| In reply to | #825 |
George Neuner <gneuner2@comcast.net> wrote:
>>But registers are a very scarce resource, so that frequent loading from
>>memory is hardly avoidable.
>
> My opinions are colored by experience with DSPs, but I have long
> thought that it would be helpful to have a few K-words of non-cache
> scratchpad memory very close (1..2 cycles) to the CPU.
The Cray-2 and Cray-3 had this (I think they called it "local memory").
Alas, compilers had a lot of trouble making good use of it. As John
McCalpin put it in 1998 (at the time he was at SGI):
# Local memories work fine if an experience performance person gets
# to rewrite the code. It gets a lot more difficult if you have
# to educate a compiler to do the right thing....
I don't know if compilers have improved significantly in this area.
And on the hardware side, you'd have to figure out what to do on
context switch...
--
-- "Jonathan Thornburg [remove -animal to reply]" <jthorn@astro.indiana-zebra.edu>
Dept of Astronomy, Indiana University, Bloomington, Indiana, USA
"C++ is to programming as sex is to reproduction. Better ways might
technically exist but they're not nearly as much fun." -- Nikolai Irgens
[It sounds like if we keep working, we might reinvent register windows. -John]
[toc] | [prev] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2013-01-02 18:29 +0000 |
| Message-ID | <13-01-008@comp.compilers> |
| In reply to | #826 |
Jonathan Thornburg <jthorn@astro.indiana.edu> wrote: > George Neuner <gneuner2@comcast.net> wrote: (snip) >> My opinions are colored by experience with DSPs, but I have long >> thought that it would be helpful to have a few K-words of non-cache >> scratchpad memory very close (1..2 cycles) to the CPU. > The Cray-2 and Cray-3 had this (I think they called it "local memory"). > Alas, compilers had a lot of trouble making good use of it. As John > McCalpin put it in 1998 (at the time he was at SGI): I previously suggested the Cray-1 vector registers, maybe not quite the same. Still, the being able to move around multiple words in one operation, especially to get them through a pipelined ALU, seems a good one. > # Local memories work fine if an experience performance person gets > # to rewrite the code. It gets a lot more difficult if you have > # to educate a compiler to do the right thing.... I do remember, I believe from the Cray-1 days, suggestions of having the linker allocate registers. That is, similar to the way memory relocation is done at link time, vector registers would be assigned to minimize (as well as can be done statically) register spill. Now, that was before Fortran had dynamic allocation. (And Fortran was the popular language for Cray programming.) > I don't know if compilers have improved significantly in this area. Compiler technology (and size) has change much over the years. One ongoing question is how well compilers are doing with Fortran array expressions. It would seem easier for compilers to generate code using vector registers for array operations. > And on the hardware side, you'd have to figure out what to do on > context switch... Seems to me that the idea behind batch programming (as contrasted to time-sharing) is keeping context switch low. With a little luck you don't have to save the vector registers (or local store) for every interrupt. (snip, our moderator wrote) > [It sounds like if we keep working, we might reinvent register windows. -John] Is SPARC dead now? I was never sure how well the SPARC register windows, with in and out registers, actually worked for real code. That is, if they made the best use of chip resources. I presume they help more for procedure call register saving than for contect switch, though. -- glen [SPARC is still around, mostly seems to be Sun selling servers to run parent Oracle's software. Register windows are awful for context switch, since you have to dump and restore the whole register stack on each switch if you don't have some hack like multiple stacks for active processes. It is also my impression that SPARC has register windows because it was designed to run code from PCC, a compiler that did very naive register allocation. The IBM 801 project was at the same time, and its chip had an ordinary register file because they found that the compiler could generate code to do the saves and restores at least as well as the implicit ones that windows do. -John]
[toc] | [prev] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2013-01-02 05:11 +0000 |
| Message-ID | <13-01-004@comp.compilers> |
| In reply to | #825 |
George Neuner <gneuner2@comcast.net> wrote: (snip) > My opinions are colored by experience with DSPs, but I have long > thought that it would be helpful to have a few K-words of non-cache > scratchpad memory very close (1..2 cycles) to the CPU. You mean like the vector registers of the Cray-1? It seems that vector processors are not out of fashion, though not so obvious that they couldn't come back. They work especially well with the pipelined ALU that the Cray series used. -- glen [It is my impression that vector processors went out of fashion when they figured out how to build instruction pipelines deep enough to keep execution units busy with scalar code. The context switch problems were and are severe, unless you want to manage a pool of scratchpads for active processes, and spill as needed, sort of like the way x86 processors do floating point register management. -John]
[toc] | [prev] | [next] | [standalone]
| From | Hans-Peter Diettrich <DrDiettrich1@aol.com> |
|---|---|
| Date | 2013-01-02 07:29 +0100 |
| Message-ID | <13-01-005@comp.compilers> |
| In reply to | #825 |
George Neuner schrieb: > On Sun, 30 Dec 2012 08:14:26 +0100, Hans-Peter Diettrich > <DrDiettrich1@aol.com> wrote: > >> Peter Dassow schrieb: >>> On 28.12.2012 08:35, Hans-Peter Diettrich wrote: >>>> Please note that most CMOS processor power consumption results from >>>> switching (stray) capacities, and only a small percentage for leak >>>> currents. E.g. a register or gate consumes such power whenever a bit is >>>> changed, and almost nothing when it has reached an stable state. > > Smaller transistors have more leakage. ACK (tunneling effect). >>> So using extensively registers instead of "conventional" memory (e.g. >>> DDR-RAM, memory outside a CPU) will save energy (if equal functionality >>> is given) ? >> I don't see a relationship here, except that external memory is slow >> and [in x86] a couple of caches and address translations are involved >> in reading from RAM. > > But associative cache and external memory accesses both are power > intensive. Right, but since every reference to external memory costs *time* in the first place, *every* compiler already optimizes for best register usage. There is nothing that can be done *additionally* in a "green" compiler. C already has a "register" keyword, as a compiler hint to hold a local variable in an register. The actual use of registers depends on the control flow taken *actually*. When a subroutine is optimized for using all available registers, it has to save and restore the registers on entry/exit. When it actually does nothing, due to given conditions, the time and energy used for pushing/popping the registers is only wasted. For that reason some (Texas Instruments?) processors implemented a register stack, decades ago, with its stack pointer adjusted according to the number of registers used in a subroutine. This stack could be moved into the CPU nowadays, eliminating the need for saving registers in external memory. But this optimization reaches a hard limit on deeply nested calls, with every subroutine using a high number of registers. In external memory the register-stack size is adjustable to program needs, just like ordinary stack size is, but a CPU resident register stack has a fixed depth. Eventually the register stack still could be kept in RAM, with an dedicated cache equivalent to the L1/L2 caches. Then the caches would automatically push/pop register contents depending on their actual *use*, not by fixed push/pop *instruction sequences*. OTOH we already have nested caches, so that the effect of an additional register cache is questionable. (see Wikipedia "CPU cache") The x86 architecture uses another approach (register renaming), with a high number of shadow registers (compared to only 16 addressable registers). I'm not sure, though, how a compiler should generate code for best use of that model... >> But registers are a very scarce resource, so that frequent loading from >> memory is hardly avoidable. > > My opinions are colored by experience with DSPs, but I have long > thought that it would be helpful to have a few K-words of non-cache > scratchpad memory very close (1..2 cycles) to the CPU. ACK, but see above considerations on the use of such memory, with its *size* limited by the architecture, and *usage* depending on subroutine needs and control flow (subroutine nesting, branches taken...). DoDi [There were stacks with the top few registers kept in fast memory in the Burroughs machines in the 1960s. It was easy to generate code for them, but since register coloring was invented in the 1970s, modern code scheduling for normal registers is much more effective. -John]
[toc] | [prev] | [next] | [standalone]
| From | glen herrmannsfeldt <gah@ugcs.caltech.edu> |
|---|---|
| Date | 2013-01-02 19:52 +0000 |
| Message-ID | <13-01-009@comp.compilers> |
| In reply to | #828 |
Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote: (snip, someone wrote) >> Smaller transistors have more leakage. > ACK (tunneling effect). (snip) > The actual use of registers depends on the control flow taken > *actually*. When a subroutine is optimized for using all available > registers, it has to save and restore the registers on entry/exit. > When it actually does nothing, due to given conditions, the time and > energy used for pushing/popping the registers is only wasted. Some use caller saves, some callee saves. In the latter case, the subroutine knows which registers it will change, and only needs to save those. It is usual to do that save on entry and restore just before return, but with a little extra logic the save might be delayed a little bit. > For that reason some (Texas Instruments?) processors implemented a > register stack, decades ago, with its stack pointer adjusted according > to the number of registers used in a subroutine. This stack could be > moved into the CPU nowadays, eliminating the need for saving registers > in external memory. If I remember the TMS9900, the registers were in memory, so that pointer just changed where in memory they were. > But this optimization reaches a hard limit on > deeply nested calls, with every subroutine using a high number of > registers. In external memory the register-stack size is adjustable to > program needs, just like ordinary stack size is, but a CPU resident > register stack has a fixed depth. Eventually the register stack still > could be kept in RAM, with an dedicated cache equivalent to the L1/L2 > caches. The original idea behind the 8087 register stack was that it would be interrupt driven, spill on overflow, restore on underflow. The problem was that the logic wasn't tested before the chip was built, and then it was too late. Also, it has only 8 entries. > Then the caches would automatically push/pop register contents > depending on their actual *use*, not by fixed push/pop *instruction > sequences*. OTOH we already have nested caches, so that the effect of > an additional register cache is questionable. (see Wikipedia "CPU > cache") A larger register stack, with working automatic spill/restore, and a fast enough interrupt handler, might work. > The x86 architecture uses another approach (register renaming), with a > high number of shadow registers (compared to only 16 addressable > registers). I'm not sure, though, how a compiler should generate code > for best use of that model... I first knew about register renaming for the floating point unit of the IBM 360/91. It was a popular example machine for many generations of books on pipelined processors. With only four floating point registers, renaming was pretty important for out-of-order execution for S/360. That was especially true since the 360/91 was designed to be able to run code not specifically optimized for it. RISC processors tend to expect compilers to order instructions appropriately, and also usually have plenty of registers. (snip) > [There were stacks with the top few registers kept in fast memory in > the Burroughs machines in the 1960s. It was easy to generate code for > them, but since register coloring was invented in the 1970s, modern > code scheduling for normal registers is much more effective. -John] Hmm. Not that I completely understand it, but it seems to me that normal general registers are generally used over shorter distances. Larger ones, such as vector registers, take longer to save and restore, and might also need to store values for a longer time. It could help much not to have to save/restore for every interrupt, for example. -- glen
[toc] | [prev] | [next] | [standalone]
| From | "Charles Richmond" <numerist@aquaporin4.com> |
|---|---|
| Date | 2013-01-04 08:59 -0600 |
| Message-ID | <13-01-015@comp.compilers> |
| In reply to | #832 |
"glen herrmannsfeldt" <gah@ugcs.caltech.edu> wrote in message > Hans-Peter Diettrich <DrDiettrich1@aol.com> wrote: > > [snip...] [snip...] [snip...] > >> For that reason some (Texas Instruments?) processors implemented a >> register stack, decades ago, with its stack pointer adjusted according >> to the number of registers used in a subroutine. This stack could be >> moved into the CPU nowadays, eliminating the need for saving registers >> in external memory. > > If I remember the TMS9900, the registers were in memory, so that > pointer just changed where in memory they were. The TMS9900 did allocate the register "file" in memory, with a pointer in the CPU to indicate the position of the register file. Some versions of the chip had some on-board RAM that could be used for register allocation. This RAM was faster to access and would speed along execution. -- numerist at aquaporin4 dot com
[toc] | [prev] | [next] | [standalone]
| From | "Nils M Holm" <nmh@t3x.org> |
|---|---|
| Date | 2012-12-23 10:01 +0100 |
| Message-ID | <12-12-013@comp.compilers> |
| In reply to | #794 |
Abid <abidmuslim@gmail.com> wrote: > Do we need to change this model and make it > three dimensional by adding power axis in the search space? In principle, I would say that higher execution speed equals more grenn-ness. Less time spent dissipating heat means less energy consumed. So the question actually is the same as ever: how to we make code run fast? - How do we squeeze more meaning into fewer/faster instructions? - How do we make code so small that it fits in caches? - How do we organize code execution in such a way that cache stalls are minimized? (Locality) -- Nils M Holm < n m h @ t 3 x . o r g > www.t3x.org
[toc] | [prev] | [next] | [standalone]
| From | Hans-Peter Diettrich <DrDiettrich1@aol.com> |
|---|---|
| Date | 2012-12-24 05:16 +0100 |
| Message-ID | <12-12-017@comp.compilers> |
| In reply to | #797 |
Nils M Holm schrieb: > Abid <abidmuslim@gmail.com> wrote: >> Do we need to change this model and make it >> three dimensional by adding power axis in the search space? > > In principle, I would say that higher execution speed equals more > grenn-ness. Less time spent dissipating heat means less energy > consumed. Not really. You mean more efficient code, I suppose? > So the question actually is the same as ever: how to we make code > run fast? > > - How do we squeeze more meaning into fewer/faster instructions? Okay, but the available machine instructions are limited. In detail on RISC architectures. > - How do we make code so small that it fits in caches? It's not only the code that has to fit into the caches. Instruction reordering and branch prediction are one thing, but the placement of the data (variables) is quite another story. > - How do we organize code execution in such a way that cache > stalls are minimized? (Locality) Such optimization requires compiler hints, so that the compiler can not only arrange the code, but also the data. This requires knowledge about the *frequency* of subroutine use, branches taken, and variables used in these critical pathes. But who should provide such hints? IMO only a profiling tool can provide that information, when a program is run with typical (real life) data. DoDi
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-12-27 13:36 +0000 |
| Message-ID | <12-12-023@comp.compilers> |
| In reply to | #797 |
"Nils M Holm" <nmh@t3x.org> writes: >In principle, I would say that higher execution speed equals more >grenn-ness. Less time spent dissipating heat means less energy >consumed. If a computer runs only a given set of batch-style programs a given set of times, yes. But that's not how computers are used. E.g., there are games that consume all the resources they are given. If the game code runs faster, it just means that the game has better FPS, not that the computer consumes less energy. There are also anti-virus programs etc., and even for batch-style programs, if they run faster, one often runs them more often. So it's not so easy. - anton -- M. Anton Ertl anton@mips.complang.tuwien.ac.at http://www.complang.tuwien.ac.at/anton/
[toc] | [prev] | [next] | [standalone]
| From | "Nils M Holm" <nmh@t3x.org> |
|---|---|
| Date | 2012-12-28 09:11 +0100 |
| Message-ID | <12-12-027@comp.compilers> |
| In reply to | #807 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > "Nils M Holm" <nmh@t3x.org> writes: > >In principle, I would say that higher execution speed equals more > >green-ness. Less time spent dissipating heat means less energy > >consumed. > > If a computer runs only a given set of batch-style programs a given > set of times, yes. But that's not how computers are used. E.g., > there are games that consume all the resources they are given. If the > game code runs faster, it just means that the game has better FPS, not > that the computer consumes less energy. There are also anti-virus > programs etc., and even for batch-style programs, if they run faster, > one often runs them more often. So it's not so easy. So you are basically arguing that people always try to get the maximum out of their resources and therefore faster code is worse. Indeed, if this should be the case, you are right. Your first example was games. Why does a game *have to* use the maximum number of frames per second? When a given number of frames per second is sufficient, in the sense that additional frames will not make the video output any smoother, why increase the frame rate further? Let the program sleep between frames. In this case fast code can sleep between frames, slow code can't. Same for virus scanner: why run them more often? So I would argue that the issue you raise is a psychological one rather than a technical one. When people always want to get the maximum out of their resources, then there is no way to reduce power consumption, except, maybe, by creating *slower* processors. Creating slower compiler output would not help, because it will just burn more CPU cycles to achieve less. Faster code is better, because it achieves the same in less time. E.g.: a slow interpreter can run the same algorithm as a fast, compiled executable. Both achieve the same on the same CPU, but the interpreted code takes longer and hence wastes energy. So faster code saves energy. Of course, this point assumes that people want a fixed outcome at the least cost, and I see that this is often not the case. It seems to get more and more interesting though, with all those "green" processors, and other "green" equipment around. Bottom line: we can create the tools to save energy, but if people do not want to save energy, nothing will help. -- Nils M Holm < n m h @ t 3 x . o r g > www.t3x.org
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-12-28 16:57 +0000 |
| Message-ID | <12-12-032@comp.compilers> |
| In reply to | #811 |
"Nils M Holm" <nmh@t3x.org> writes: >Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >> "Nils M Holm" <nmh@t3x.org> writes: >> >In principle, I would say that higher execution speed equals more >> >green-ness. Less time spent dissipating heat means less energy >> >consumed. ... >So you are basically arguing that people always try to get the maximum >out of their resources and therefore faster code is worse. Not worse, but the claims you made above are a bit too simple-minded. Fast code still has the advantage of being fast; but it does not necessarily save energy. >Your first example was games. Why does a game *have to* use the >maximum number of frames per second? When a given number of frames per >second is sufficient, in the sense that additional frames will not >make the video output any smoother, why increase the frame rate >further? Let the program sleep between frames. In this case fast code >can sleep between frames, slow code can't. But do game designers work that way? If the game engine is so fast that it can sleep between frames, they put in more features in the game that make use of the CPU; or the reverse, they don't take as many features out in the slimming phase when they try to make the game run smooth enough. So the faster code buys a more featureful game, but it does not save energy. Or at the level of the end user (player), they tune the graphics details such that they get the maximum details at an acceptable frame rate on their machine. So a faster program means more details, not less energy consumption. >Same for virus scanner: why run them more often? More often? AFAIK, these days the virus scanners are running constantly in the background and consume one or two cores. I don't know if that has a technical reason or is just done to give the users a feeling of better protection, but that does not matter for the energy consumption. >So I would argue that the issue you raise is a psychological one rather >than a technical one. When people always want to get the maximum out >of their resources, then there is no way to reduce power consumption, >except, maybe, by creating *slower* processors. Or at least processors and machines with lower power consumption under load (yes, they tend to be slower, the converse is not necessarily true). Yes, you can see it as a psychological issue, but that does not make it less real. Idle power consumption is also an issue, for some workloads. - anton -- M. Anton Ertl anton@mips.complang.tuwien.ac.at http://www.complang.tuwien.ac.at/anton/
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | comp.compilers
csiph-web