Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #17603 > unrolled thread
| Started by | Mark Wills <forthfreak@gmail.com> |
|---|---|
| First post | 2012-11-27 08:01 -0800 |
| Last post | 2012-11-28 14:21 +0000 |
| Articles | 20 on this page of 112 — 16 participants |
Back to article view | Back to comp.lang.forth
DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 08:01 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-11-27 08:55 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 09:01 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-27 17:42 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 23:50 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 04:41 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 02:48 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 05:27 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 03:51 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 21:56 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-29 01:26 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-29 22:34 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-30 01:42 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 13:18 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-01 01:42 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-03 15:25 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-03 16:18 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-03 21:20 -0500
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-03 17:29 -1000
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 00:12 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-05 11:35 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-04 20:17 -0800
Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-05 08:31 +0200
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:48 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:53 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-05 12:13 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 15:26 -0800
Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-06 06:32 +0200
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 01:07 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 04:23 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:49 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:42 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:30 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-06 09:45 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 21:41 +0000
Re: DTC "A. K." <akk@nospam.org> - 2012-12-06 23:15 +0100
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-08 01:27 +0000
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 11:19 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 03:55 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:44 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:05 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:35 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:20 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-09 01:01 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:10 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:57 +0100
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 19:46 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:23 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:37 +0100
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 14:41 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:15 -0800
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 17:07 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:16 -0800
Re: DTC Brad Eckert <hwfwguy@gmail.com> - 2012-12-07 09:00 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:21 -0800
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 08:01 -1000
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:18 -0800
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 13:48 -1000
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-06 19:03 -0500
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-07 02:59 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 13:13 +0000
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-05 20:39 -0500
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:46 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:47 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:36 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 04:03 -0800
Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:30 -0800
Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-11 23:52 -0500
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-12 10:59 -0800
Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-31 02:43 -0500
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2013-01-02 00:44 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-28 13:54 +0000
Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:16 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-28 06:58 -0500
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 04:50 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 07:08 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 06:02 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:23 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:18 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 15:00 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 09:18 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 16:36 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 11:02 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 17:13 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:03 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 18:12 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-29 14:30 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-29 18:05 +0100
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-29 11:19 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 03:14 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 14:12 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 16:40 +0000
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-01 15:34 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-01 21:23 +0100
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 16:28 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 18:44 +0100
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-12-03 04:50 -0600
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 16:48 +0100
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 15:59 +0000
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-30 15:34 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:36 -0600
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 12:47 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-28 11:29 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:15 -0500
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-11-29 08:52 -1000
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 22:25 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:12 -0500
Re: DTC humptydumpty <ouatubi@gmail.com> - 2012-11-28 02:04 -0800
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:21 +0000
Page 5 of 6 — ← Prev page 1 2 3 4 [5] 6 Next page →
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-28 15:00 +0000 |
| Message-ID | <2012Nov28.160018@mips.complang.tuwien.ac.at> |
| In reply to | #17636 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>Mark Wills <forthfreak@gmail.com> wrote:
>>>
>>>> In a DTC, the CFA is a CALL or a BRANCH/JMP etc to the
>>>> 'handler' (DOCOL et al).
>>>
>>>No: in some cases it's the actual code. Consider a CONSTANT, for
>>>example.
>>
>> Why constants? I would expect that on most DTC systems there is a
>> "JMP/CALL DOCON" at the start of a constant.
>
>It's possible, but a pretty crappy implementation. Why would you do
>that
Because it's closest in many respects to the traditional ITC
implementation. Because it allows defining
: COMPILE, , ;
> when there's such an obvious speedup?
What obvious speedup? How do you get it? How big is it? If it's so
obvious, big, and has no disadvantages, why is it not used in ITC?
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 09:18 -0600 |
| Message-ID | <EPedneRds5eqtivNnZ2dnUVZ8rGdnZ2d@supernews.com> |
| In reply to | #17638 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>>>Mark Wills <forthfreak@gmail.com> wrote: >>>> >>>>> In a DTC, the CFA is a CALL or a BRANCH/JMP etc to the >>>>> 'handler' (DOCOL et al). >>>> >>>>No: in some cases it's the actual code. Consider a CONSTANT, for >>>>example. >>> >>> Why constants? I would expect that on most DTC systems there is a >>> "JMP/CALL DOCON" at the start of a constant. >> >>It's possible, but a pretty crappy implementation. Why would you do >>that > > Because it's closest in many respects to the traditional ITC > implementation. I assume that if you want a traditional ITC implementation you'll use ITC. > Because it allows defining > > : COMPILE, , ; It allows that anyway, assuming that ' returns the address of the start of the code. > What obvious speedup? So that 1 CONSTANT 1 generates mov r0, #1 push r0 next or perhaps mov r0, #1 jmp push or something similar. > How do you get it? How big is it? That depends on the architecture. > If it's so obvious, big, and has no disadvantages, why is it not > used in ITC? ITC can't do it as easily. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-28 16:36 +0000 |
| Message-ID | <2012Nov28.173609@mips.complang.tuwien.ac.at> |
| In reply to | #17639 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> What obvious speedup?
>
>So that
>
> 1 CONSTANT 1
>
>generates
>
> mov r0, #1
> push r0
> next
>
>or perhaps
>
> mov r0, #1
> jmp push
>
>or something similar.
I did not think of that. So much for "obvious".
>> How do you get it? How big is it?
>
>That depends on the architecture.
On IA-32 and AMD-64, you would get a big slowdown for many programs
unless you separate the generated native code from the threaded code
(which requires additional work that has not made it into a number of
native code systems yet, 16 years after the issue became known).
>> If it's so obvious, big, and has no disadvantages, why is it not
>> used in ITC?
>
>ITC can't do it as easily.
I don't think that adding a code field makes a significant difference.
1 constant 1 generates
dw *+cell
mov r0, #1
push r0
next
or perhaps
dw *+cell
mov r0, #1
jmp push
or something similar.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 11:02 -0600 |
| Message-ID | <3_idnYfdWuwR3ivNnZ2dnUVZ8tadnZ2d@supernews.com> |
| In reply to | #17640 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >>> What obvious speedup? >> >>So that >> >> 1 CONSTANT 1 >> >>generates >> >> mov r0, #1 >> push r0 >> next >> >>or perhaps >> >> mov r0, #1 >> jmp push >> >>or something similar. > > I did not think of that. So much for "obvious". > >>> How do you get it? How big is it? >> >>That depends on the architecture. > > On IA-32 and AMD-64, you would get a big slowdown for many programs > unless you separate the generated native code from the threaded code > (which requires additional work that has not made it into a number of > native code systems yet, 16 years after the issue became known). Err, why would mov r0, #1 jmp push be slower than call docon ... the constant 1 ? The latter mixes code and data in the same memory area, the former doesn't. Besides, last time I looked the only penalty was for a write to the same cache line as the code; that's irrelevant here. >>> If it's so obvious, big, and has no disadvantages, why is it not >>> used in ITC? >> >>ITC can't do it as easily. > > I don't think that adding a code field makes a significant difference. > > 1 constant 1 generates > > dw *+cell > mov r0, #1 > push r0 > next > > or perhaps > > dw *+cell > mov r0, #1 > jmp push > > or something similar. I suppose it depends on how you count "significant". In the simple DTC case of mov r0, #1 jmp push you've got (depending on architecture) low or zero code expansion with some performance benefit. It's a tradeoff, like everything else in interpreter design. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-28 17:13 +0000 |
| Message-ID | <2012Nov28.181355@mips.complang.tuwien.ac.at> |
| In reply to | #17641 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> On IA-32 and AMD-64, you would get a big slowdown for many programs
>> unless you separate the generated native code from the threaded code
>> (which requires additional work that has not made it into a number of
>> native code systems yet, 16 years after the issue became known).
>
>Err, why would
>
> mov r0, #1
> jmp push
>
>be slower than
>
> call docon
> ... the constant 1
>
>? The latter mixes code and data in the same memory area, the former
>doesn't.
Good point. Yes, I found that ITC-like DTC is slower than ITC,
because of this issue.
>Besides, last time I looked the only penalty was for a write
>to the same cache line as the code; that's irrelevant here.
I guess that myopic view is what causes the persistence of this
problem. Now consider:
variable foo
1 constant bar
variable boing
: flip bar foo ! bar boing ! ;
flip
Many native code Forth systems have bet on written data not being in
the same cache line as code, and lost; and actually "not being in the
same cache line" is not enough, thanks to prefetching.
>>>ITC can't do it as easily.
>>
>> I don't think that adding a code field makes a significant difference.
...
>I suppose it depends on how you count "significant". In the simple
>DTC case of
>
> mov r0, #1
> jmp push
>
>you've got (depending on architecture) low or zero code expansion with
>some performance benefit.
It depends more on what you mean with "easy". For me it has to do
with programmer effort, not memory usage.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 12:03 -0600 |
| Message-ID | <25KdnaK6jvpvzCvNnZ2dnUVZ8radnZ2d@supernews.com> |
| In reply to | #17642 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > >>Besides, last time I looked the only penalty was for a write >>to the same cache line as the code; that's irrelevant here. > > I guess that myopic view is what causes the persistence of this > problem. In what way is this view "myopic"? It is true. > Now consider: > > variable foo > 1 constant bar > variable boing > : flip bar foo ! bar boing ! ; > flip BOING is a variable. The problem is that BOING's data shares a cache line with some code. That's irrelevant here: we're talking about constants. Their data, which does not change, may be freely mixed with code. > Many native code Forth systems have bet on written data not being in > the same cache line as code, and lost; and actually "not being in > the same cache line" is not enough, thanks to prefetching. If a processor prefetches you have to consider its accesses beyond those explicitly written, but the issue is still that of code and writable data in the same cache line. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-28 18:12 +0000 |
| Message-ID | <2012Nov28.191254@mips.complang.tuwien.ac.at> |
| In reply to | #17643 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>>
>>>Besides, last time I looked the only penalty was for a write
>>>to the same cache line as the code; that's irrelevant here.
>>
>> I guess that myopic view is what causes the persistence of this
>> problem.
>
>In what way is this view "myopic"? It is true.
>
>> Now consider:
>>
>> variable foo
>> 1 constant bar
>> variable boing
>> : flip bar foo ! bar boing ! ;
>> flip
>
>BOING is a variable. The problem is that BOING's data shares a cache
>line with some code. That's irrelevant here: we're talking about
>constants. Their data, which does not change, may be freely mixed
>with code.
It's myopic, because it looks only at the constant, as if it existed
in isolation. So I show an example with some surroundings, and you
insist that the surroundings are irrelevant. They are not. They are
the reason why bigForth is 5 times slower then Gforth on cd16sim and
about 3 times for brew.
>If a processor prefetches you have to consider its accesses beyond
>those explicitly written, but the issue is still that of code and
>writable data in the same cache line.
No, the code and the data can be in different cache lines, and you can
still have cache consistency overhead. Also, on older processors (P5,
K6) even read-only data caused this problem.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 12:32 -0600 |
| Message-ID | <DMOdnf90lcNYxSvNnZ2dnUVZ8rqdnZ2d@supernews.com> |
| In reply to | #17646 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >>> Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>>>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >>> >>>>Besides, last time I looked the only penalty was for a write >>>>to the same cache line as the code; that's irrelevant here. >>> >>> I guess that myopic view is what causes the persistence of this >>> problem. >> >>In what way is this view "myopic"? It is true. >> >>> Now consider: >>> >>> variable foo >>> 1 constant bar >>> variable boing >>> : flip bar foo ! bar boing ! ; >>> flip >> >>BOING is a variable. The problem is that BOING's data shares a cache >>line with some code. That's irrelevant here: we're talking about >>constants. Their data, which does not change, may be freely mixed >>with code. > > It's myopic, because it looks only at the constant, as if it existed > in isolation. So I show an example with some surroundings, and you > insist that the surroundings are irrelevant. They are not. They > are the reason why bigForth is 5 times slower then Gforth on cd16sim > and about 3 times for brew. In the case above it does not matter whether the constant is implemented with a separate data field or not: there will be a slowdown if a variable is written in the same cache line as the code. What does matter is where you put the data for the variable, not the data for the constant. >>If a processor prefetches you have to consider its accesses beyond >>those explicitly written, but the issue is still that of code and >>writable data in the same cache line. > > No, the code and the data can be in different cache lines, and you can > still have cache consistency overhead. You'll have to explain that a bit more; I don't understand the point you're making. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-29 14:30 +0000 |
| Message-ID | <2012Nov29.153018@mips.complang.tuwien.ac.at> |
| In reply to | #17647 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
[...]
>In the case above it does not matter whether the constant is
>implemented with a separate data field or not: there will be a
>slowdown if a variable is written in the same cache line as the code.
>What does matter is where you put the data for the variable, not the
>data for the constant.
Certainly. Traditionally Forth systems have just one space where they
put things, and that tradition persists even into native code systems,
even to this day (although some have been fixed). So a DTC system
will likely have just one space, too, and will run into this problem.
You are right though, that ITC-emulating DTC systems will have this
problem anyway. Primitive-centric systems to a lesser degree (only
for EXECUTE etc.), and hybrid direct/indirect threaded systems only
for CODE words.
>> No, the code and the data can be in different cache lines, and you can
>> still have cache consistency overhead.
>
>You'll have to explain that a bit more; I don't understand the point
>you're making.
If there is a cache line containing code followed by a cache line
containing data, on (at least) some CPUs the code prefetcher will
prefetch the data line and incur cache consistency overhead. See
<http://www.complang.tuwien.ac.at/misc/pentium-switch/> for data.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-11-29 18:05 +0100 |
| Message-ID | <2700090.aQ2KJ0Axev@sunwukong.fritz.box> |
| In reply to | #17647 |
Andrew Haley wrote: > You'll have to explain that a bit more; I don't understand the point > you're making. Let's give some typical code you can find in many applications; in this case we use it as micro-benchmark, it's a variable and the corresponding modifier function side by side: Variable foo# : foo+ 1 foo# +! ; : foos 0 ?DO foo+ LOOP ; Test 1 is with foo#'s slot being in a different cache line as foo+ on a Core2: !time #1000000 foos .time 0,005371 sec ok ' foo+ . 1003350C ok foo# . 100334F8 ok Test 2 is with foo#'s slot and foo+ in the same cache line, same Core2: !time #1000000 foos .time 0,185575 sec ok ' foo+ . 1003351C ok foo# . 10033508 ok Both tests done with bigForth. Factor 37. This is really significant. Go to a different CPU, now instead of Core2, we chose an AMD Zacate, to see how that is affected by the problem: Test 1: 0,008208 sec ok Test 2: 0,319169 sec ok Factor 39. Not that different. Same test with vfxForth on the same Zacate, for timing I use LocalExtern: gettimeofday int gettimeofday ( int * , int * ); 2variable tz 2variable td : @time td tz gettimeofday drop td 2@ 1000000 um* rot 0 d+ ; Test 1: 0.007708s Test 2: 0.319617s Factor 41.5 (the factor is higher, because VFX creates code that is a bit faster than bigForth's code). Test 2 in VFX is more difficult to achieve than in bigForth. First of all, Stephen *does* now use a different memory pool for variables and buffers, so if you declare foo# as VARIABLE, it will get allocated from somewhere else (no problem). If you however declare foo# with CREATE FOO# 0 , it will be close to the executed code, but the tokenizer pads enough between FOO# and the actual increment instruction that it doesn't affect Zacate. The lesson lerned is that: * You should separate code and data * VFX does a good job now on variables, but on CREATEd words, it is accidential * bigForth doesn't, and as it compiles dense code without additional information, it quickly exposes the problem * If you use such a system, declare your variables first, add maybe some padding, and then write your code If I was going to write another code generator, I'd very likely use Gforth's approach: EXECUTE has an indirection (this is pretty cheap on current CPUs), and the dictionary contains variables, headers, and pointer to the actual code, but not the code itself. It might also be possible to change CREATE so that it uses a different area for code and data, so that having a direct EXECUTE is possible. Separated headers might give some benefits, but with a hash table, you don't pollute the cache much. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | Alex McDonald <blog@rivadpm.com> |
|---|---|
| Date | 2012-11-29 11:19 -0800 |
| Message-ID | <1d14e804-7e68-460d-8eb8-a4193562b56c@h5g2000vbb.googlegroups.com> |
| In reply to | #17697 |
On Nov 29, 5:05 pm, Bernd Paysan <bernd.pay...@gmx.de> wrote:
> Andrew Haley wrote:
> > You'll have to explain that a bit more; I don't understand the point
> > you're making.
>
> Let's give some typical code you can find in many applications; in this
> case we use it as micro-benchmark, it's a variable and the corresponding
> modifier function side by side:
>
> Variable foo#
> : foo+ 1 foo# +! ;
> : foos 0 ?DO foo+ LOOP ;
>
> Test 1 is with foo#'s slot being in a different cache line as foo+ on a
> Core2:
>
> !time #1000000 foos .time 0,005371 sec ok
> ' foo+ . 1003350C ok
> foo# . 100334F8 ok
>
> Test 2 is with foo#'s slot and foo+ in the same cache line, same Core2:
>
> !time #1000000 foos .time 0,185575 sec ok
> ' foo+ . 1003351C ok
> foo# . 10033508 ok
>
> Both tests done with bigForth. Factor 37. This is really significant.
>
> Go to a different CPU, now instead of Core2, we chose an AMD Zacate, to
> see how that is affected by the problem:
>
> Test 1: 0,008208 sec ok
> Test 2: 0,319169 sec ok
>
> Factor 39. Not that different.
>
> Same test with vfxForth on the same Zacate, for timing I use
>
> LocalExtern: gettimeofday int gettimeofday ( int * , int * );
> 2variable tz 2variable td
> : @time td tz gettimeofday drop td 2@ 1000000 um* rot 0 d+ ;
>
> Test 1: 0.007708s
> Test 2: 0.319617s
>
> Factor 41.5 (the factor is higher, because VFX creates code that is a
> bit faster than bigForth's code).
>
> Test 2 in VFX is more difficult to achieve than in bigForth. First of
> all, Stephen *does* now use a different memory pool for variables and
> buffers, so if you declare foo# as VARIABLE, it will get allocated from
> somewhere else (no problem). If you however declare foo# with CREATE
> FOO# 0 , it will be close to the executed code, but the tokenizer pads
> enough between FOO# and the actual increment instruction that it doesn't
> affect Zacate.
>
> The lesson lerned is that:
>
> * You should separate code and data
> * VFX does a good job now on variables, but on CREATEd words, it is
> accidential
> * bigForth doesn't, and as it compiles dense code without additional
> information, it quickly exposes the problem
> * If you use such a system, declare your variables first, add maybe some
> padding, and then write your code
>
> If I was going to write another code generator, I'd very likely use
> Gforth's approach: EXECUTE has an indirection (this is pretty cheap on
> current CPUs), and the dictionary contains variables, headers, and
> pointer to the actual code, but not the code itself. It might also be
> possible to change CREATE so that it uses a different area for code and
> data, so that having a direct EXECUTE is possible. Separated headers
> might give some benefits, but with a hash table, you don't pollute the
> cache much.
>
> --
> Bernd Paysan
> "If you want it done right, you have to do it yourself"http://bernd-paysan.de/
The design decision for my modified Win32Forth was to have 3 areas;
code, dictionaries and data. Although the dictionaries and data could
be in one area, the advantages of having an identifiable range of
addresses for each outweighs the very slight (for the implementer) to
non-existent (for the user) inconvenience.
create x 0 , ok
see x
create x ( addr $805164 ) ( 0 -- 1 )
\ ' (comp-cons) compiles; code=$41A71A len=10 type=13
\ defined in (console)
( $0 ) mov ecx $805164 \ B964518000
( $5 ) jmp ' dovar \ E9E068FEFF
( end ) ok
The data address is an immediate. CREATE VARIABLE VALUE and their
children can all use >BODY as the code is identical for each; the JMP
is to DOVAR (sets top of stack from ecx) or DOVAL (fetches tos from
[ecx]). CONSTANTs have no data area, and the value is the constant
directly.
22 constant y ok
see y
22 constant y ( 0 -- 1 )
\ ' (comp-cons) compiles; code=$41A728 len=10 type=11
\ defined in (console)
( $0 ) mov ecx $16 \ B916000000
( $5 ) jmp ' dovar \ E9D268FEFF
( end ) ok
Compiled code eliminates all but references to the data area; the
interpret code shown is not used for these primitives.
: foo x y ; ok
see foo
: foo ( ? -- ? )
\ std call compiles; code=$41A75C len=18 type=1
\ defined in (console)
( $0 ) mov dword { $-4 ebp } eax \ 8945FC
( $3 ) mov eax $16 \ B816000000
( $8 ) mov dword { $-8 ebp } $805164 \ C745F864518000
( $F ) sub ebp $8 \ 83ED08
( $12 ) ret \ C3 ( end ) ok
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-30 03:14 -0600 |
| Message-ID | <DKKdnRsQzsib5CXNnZ2dnUVZ8gednZ2d@supernews.com> |
| In reply to | #17697 |
Bernd Paysan <bernd.paysan@gmx.de> wrote: > Andrew Haley wrote: >> You'll have to explain that a bit more; I don't understand the point >> you're making. > > Let's give some typical code you can find in many applications; in this > case we use it as micro-benchmark, it's a variable and the corresponding > modifier function side by side: Umm no, that's not what I was asking. I was asking why Anton said no when as far as I could tell he was agreeing with me. But never mind. > The lesson lerned is that: > > * You should separate code and data > * VFX does a good job now on variables, but on CREATEd words, it is > accidential > * bigForth doesn't, and as it compiles dense code without additional > information, it quickly exposes the problem > * If you use such a system, declare your variables first, add maybe some > padding, and then write your code > > If I was going to write another code generator, I'd very likely use > Gforth's approach: EXECUTE has an indirection (this is pretty cheap > on current CPUs), and the dictionary contains variables, headers, > and pointer to the actual code, but not the code itself. It might > also be possible to change CREATE so that it uses a different area > for code and data, so that having a direct EXECUTE is possible. I don't understand the point of this. Surely you just put everything write-once (which includes the code and the dictionary entries) in one area and everything variable (HERE and ALLOT) in another. EXECUTE doesn't need an indirection because ' can return a code address; only >BODY needs the indirection. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-30 14:12 +0000 |
| Message-ID | <2012Nov30.151239@mips.complang.tuwien.ac.at> |
| In reply to | #17754 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>I don't understand the point of this. Surely you just put everything
>write-once (which includes the code and the dictionary entries) in one
>area and everything variable (HERE and ALLOT) in another. EXECUTE
>doesn't need an indirection because ' can return a code address;
>only >BODY needs the indirection.
Surely? Gforth puts stuff that the CPU puts into the I-cache in one
area, and stuff that the CPU puts into the D-cache in a different
area:
Advantages of the Gforth approach:
+ Better utilization of I-Cache and D-cache
+ Also avoids the cache consistency issues on processors that
invalidate I-cache lines on data reads (Pentium, K6).
+ No indirection on >BODY
Disadvantage:
- Indirection on EXECUTE.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-30 10:32 -0600 |
| Message-ID | <8cudneSQd9MoQiXNnZ2dnUVZ8nWdnZ2d@supernews.com> |
| In reply to | #17769 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>I don't understand the point of this. Surely you just put everything >>write-once (which includes the code and the dictionary entries) in one >>area and everything variable (HERE and ALLOT) in another. EXECUTE >>doesn't need an indirection because ' can return a code address; >>only >BODY needs the indirection. > > Surely? Gforth puts stuff that the CPU puts into the I-cache in one > area, and stuff that the CPU puts into the D-cache in a different > area: > > Advantages of the Gforth approach: > > + Better utilization of I-Cache and D-cache That's interesting. I suppose Forth words are so extremely small that you can get more than one of them in a single cache line, so it's worth packing them together without intervening stuff. And you might get good enough locality for this to be useful, with a following wind. Is it really worth it? > + Also avoids the cache consistency issues on processors that > invalidate I-cache lines on data reads (Pentium, K6). Uh, yeah. Strictly for retrocomputing fans. > + No indirection on >BODY > > Disadvantage: > > - Indirection on EXECUTE. But surely you can still have ' returning a code address even if you have separated code from dictionary entries. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-11-30 16:40 +0000 |
| Message-ID | <2012Nov30.174037@mips.complang.tuwien.ac.at> |
| In reply to | #17772 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> + Also avoids the cache consistency issues on processors that
>> invalidate I-cache lines on data reads (Pentium, K6).
>
>Uh, yeah. Strictly for retrocomputing fans.
Supposedly Atoms and Larrabee use a core derived from the Pentium. I
would not rely on them not having this issue without doing
measurements first.
>> Disadvantage:
>>
>> - Indirection on EXECUTE.
>
>But surely you can still have ' returning a code address even if
>you have separated code from dictionary entries.
Yes, one can do this, but them a piece of code needs to be generated
for every word (variables, does> children, etc.), just for the sake of
EXECUTE and DEFER. Whether you want to do that is a tradeoff. In
Gforth we do the indirection.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-12-01 15:34 +0000 |
| Message-ID | <2012Dec1.163416@mips.complang.tuwien.ac.at> |
| In reply to | #17772 |
Andrew Haley <andrew29@littlepinkcloud.invalid> writes:
>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
>> + Better utilization of I-Cache and D-cache
>
>That's interesting. I suppose Forth words are so extremely small that
>you can get more than one of them in a single cache line, so it's
>worth packing them together without intervening stuff. And you might
>get good enough locality for this to be useful, with a following wind.
The way I think about this is not on a per-word basis. Instead,
consider the working set of the application: Does it fit in the
I-Cache and in the D-Cache, or conversely, how big would they have to
be to fit? If you mix code with read-only data, the working-set will
grow (both in the I-cache and in the D-cache), and if the caches are
smaller than the working set, you will have more misses in these
caches. Actually, even if the caches are big enough for the working
set, you will have more misses (compulsory misses), although the
absolute number of misses will be small.
>Is it really worth it?
Given that it was easier to implement that based on the earlier Gforth
implementation (which had no separate headers and no native code
(apart from the ITC-emulating DTC trampolines)) than to implement your
approach, yes, this advantage is definitely worth the negative cost.
How small or big the advantage is depends on the application and can
only be determined by implementing both variants in the same Forth
system and measuring the result.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-12-01 21:23 +0100 |
| Message-ID | <5599042.RWDIGzUSXz@sunwukong.fritz.box> |
| In reply to | #17792 |
Anton Ertl wrote: > The way I think about this is not on a per-word basis. Instead, > consider the working set of the application: Does it fit in the > I-Cache and in the D-Cache, or conversely, how big would they have to > be to fit? If you mix code with read-only data, the working-set will > grow (both in the I-cache and in the D-cache), and if the caches are > smaller than the working set, you will have more misses in these > caches. Actually, even if the caches are big enough for the working > set, you will have more misses (compulsory misses), although the > absolute number of misses will be small. Using a separate region to allocate variables as VFX now does is probably further reducing D-cache footprint: To access variables, you don't actually need all its header and stuff - this part is only needed for compilation, and should not pollute the cache during execution. Gforth EC does provide such a capability, for flash systems, where you need to separate write-once and write-many parts. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-12-03 16:28 +0000 |
| Message-ID | <2012Dec3.172809@mips.complang.tuwien.ac.at> |
| In reply to | #17798 |
Bernd Paysan <bernd.paysan@gmx.de> writes:
>Using a separate region to allocate variables as VFX now does is
>probably further reducing D-cache footprint: To access variables, you
>don't actually need all its header and stuff - this part is only needed
>for compilation, and should not pollute the cache during execution.
Yes, there are ways to reduce the D-cache footprint by separating
various kinds of data, and this one is probably a good one. But
separating the code from the data is a very obvious one, because the
code does not even live in the same L1 cache as the data.
>Gforth EC does provide such a capability, for flash systems, where you
>need to separate write-once and write-many parts.
That is a different kind of separation. It may be beneficial with
write-back caches, but if it separates data that is used at the same
time and would be close to each other without that separation, it's
not a clear win. Separating the headers is a win, because they are
normally not used during program execution.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-12-03 18:44 +0100 |
| Message-ID | <7877533.SGPmmKFI4j@sunwukong.fritz.box> |
| In reply to | #17823 |
Anton Ertl wrote: >>Gforth EC does provide such a capability, for flash systems, where you >>need to separate write-once and write-many parts. > > That is a different kind of separation. It may be beneficial with > write-back caches, but if it separates data that is used at the same > time and would be close to each other without that separation, it's > not a clear win. Separating the headers is a win, because they are > normally not used during program execution. That's a side effect of this property: A variable consists of a header (only used during compilation), a code field (only used for compilation and EXECUTE, since Gforth is primitive centric), and a body, where the actual value sits, and this is used during execution. That body lives in RAM, the rest lives in flash. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-12-03 04:50 -0600 |
| Message-ID | <lMadnRQJgbWaGSHNnZ2dnUVZ8oKdnZ2d@supernews.com> |
| In reply to | #17792 |
Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: > Andrew Haley <andrew29@littlepinkcloud.invalid> writes: >>Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote: >>> + Better utilization of I-Cache and D-cache >> >>That's interesting. I suppose Forth words are so extremely small >>that you can get more than one of them in a single cache line, so >>it's worth packing them together without intervening stuff. And you >>might get good enough locality for this to be useful, with a >>following wind. > > The way I think about this is not on a per-word basis. Instead, > consider the working set of the application: Does it fit in the > I-Cache and in the D-Cache, or conversely, how big would they have > to be to fit? If you mix code with read-only data, the working-set > will grow (both in the I-cache and in the D-cache), and if the > caches are smaller than the working set, you will have more misses > in these caches. It's the cache line granularity that really matters, though: you won't gain anything from separating headers and words unless more words in your working set share the same cache lines. A word that's only 20 bytes long may occupy a line of its own regardless of whether its headers are in the same space. Andrew.
[toc] | [prev] | [next] | [standalone]
Page 5 of 6 — ← Prev page 1 2 3 4 [5] 6 Next page →
Back to top | Article view | comp.lang.forth
csiph-web