Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #17603 > unrolled thread

DTC

Started byMark Wills <forthfreak@gmail.com>
First post2012-11-27 08:01 -0800
Last post2012-11-28 14:21 +0000
Articles 20 on this page of 112 — 16 participants

Back to article view | Back to comp.lang.forth


Contents

  DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 08:01 -0800
    Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-11-27 08:55 -0800
      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 09:01 -0800
    Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-27 17:42 -0800
      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 23:50 -0800
        Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 04:41 -0600
          Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 02:48 -0800
            Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 05:27 -0600
              Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 03:51 -0800
                Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 21:56 -0800
                  Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-29 01:26 -0800
                    Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-29 22:34 -0800
                      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-30 01:42 -0800
                        Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 13:18 -0800
                          Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-01 01:42 -0800
                            Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-03 15:25 -0800
                              Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-03 16:18 -0800
                              Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-03 21:20 -0500
                                Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-03 17:29 -1000
                                Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 00:12 -0800
                                  Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-05 11:35 -0800
                                Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-04 20:17 -0800
                                  Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-05 08:31 +0200
                                    Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:48 -0800
                                      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:53 -0800
                                      Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-05 12:13 -0800
                                        Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 15:26 -0800
                                          Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-06 06:32 +0200
                                          Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 01:07 -0800
                                            Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 04:23 -0800
                                              Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:49 +0100
                                                Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:42 -0800
                                                  Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:30 -0800
                                                  Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-06 09:45 -0800
                                                    Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 21:41 +0000
                                                      Re: DTC "A. K." <akk@nospam.org> - 2012-12-06 23:15 +0100
                                                        Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-08 01:27 +0000
                                                          Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 11:19 +0100
                                                            Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 03:55 -0800
                                                              Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:44 +0100
                                                                Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:05 -0800
                                                                  Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:35 +0100
                                                                    Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:20 -0800
                                                                      Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-09 01:01 +0100
                                                                Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:10 -0800
                                                                  Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:57 +0100
                                                                    Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 19:46 +0100
                                                                      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:23 -0800
                                                            Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:37 +0100
                                                              Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 14:41 +0100
                                                              Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:15 -0800
                                                                Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 17:07 +0100
                                                      Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:16 -0800
                                                      Re: DTC Brad Eckert <hwfwguy@gmail.com> - 2012-12-07 09:00 -0800
                                                    Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:21 -0800
                                                  Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 08:01 -1000
                                                    Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:18 -0800
                                                      Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 13:48 -1000
                                              Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-06 19:03 -0500
                                                Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-07 02:59 -0800
                                          Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 13:13 +0000
                                        Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-05 20:39 -0500
                                      Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:46 +0100
                                        Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:47 -0800
                                          Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:36 -0800
                                  Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 04:03 -0800
                  Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:30 -0800
                  Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-11 23:52 -0500
                    Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-12 10:59 -0800
                      Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-31 02:43 -0500
                        Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2013-01-02 00:44 -0800
              Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-28 13:54 +0000
                Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:16 -0800
      Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-28 06:58 -0500
        Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 04:50 -0800
          Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 07:08 -0600
            Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 06:02 -0800
              Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:23 -0600
            Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:18 +0000
              Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:32 -0600
                Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 15:00 +0000
                  Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 09:18 -0600
                    Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 16:36 +0000
                      Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 11:02 -0600
                        Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 17:13 +0000
                          Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:03 -0600
                            Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 18:12 +0000
                              Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:32 -0600
                                Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-29 14:30 +0000
                                Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-29 18:05 +0100
                                  Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-29 11:19 -0800
                                  Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 03:14 -0600
                                    Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 14:12 +0000
                                      Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:32 -0600
                                        Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 16:40 +0000
                                        Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-01 15:34 +0000
                                          Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-01 21:23 +0100
                                            Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 16:28 +0000
                                              Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 18:44 +0100
                                          Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-12-03 04:50 -0600
                                            Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 16:48 +0100
                                            Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 15:59 +0000
                                    Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-30 15:34 +0000
                                      Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:36 -0600
                                        Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 12:47 -0800
                          Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-28 11:29 -0800
          Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:15 -0500
            Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-11-29 08:52 -1000
        Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 22:25 -0800
          Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:12 -0500
    Re: DTC humptydumpty <ouatubi@gmail.com> - 2012-11-28 02:04 -0800
    Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:21 +0000

Page 1 of 6  [1] 2 3 4 5 6  Next page →


#17603 — DTC

FromMark Wills <forthfreak@gmail.com>
Date2012-11-27 08:01 -0800
SubjectDTC
Message-ID<ad599bc5-3e28-4fe0-a2d5-dc9b9dfb2255@eo2g2000vbb.googlegroups.com>
Very nice write-up here about the Forth virtual machine, complete with
links to our very own Anton Ertl's pages:

http://www.wordiq.com/definition/Forth_virtual_machine

One thing that struck me, it mentions that direct threaded is not as
"flexible" as ITC. I was wondering in what respect DTC is less
flexible? Anyone have any opinions/comments?

For context, here is the description of DTC from the above link:

"Direct threading: The addresses in the code are actually the address
of machine language. This is a compromise between speed and space. The
indirect data pointer is lost, at some loss in the language's
flexibility, and this may need to be corrected by a type tag in the
data areas, with an auxiliary table. Some Forth systems have produced
direct-threaded code. On many machines direct-threading is faster than
subroutine threading (see reference below)."

The reference in the above paragraph points to a paper by Anton on
threading benchmarks.

[toc] | [next] | [standalone]


#17606

FromPaul Rubin <no.email@nospam.invalid>
Date2012-11-27 08:55 -0800
Message-ID<7xy5hnastn.fsf@ruckus.brouhaha.com>
In reply to#17603
Mark Wills <forthfreak@gmail.com> writes:
> Very nice write-up here about the Forth virtual machine, complete with
> links to our very own Anton Ertl's pages:
> http://www.wordiq.com/definition/Forth_virtual_machine

Note: That is a mirror of the wikipedia article you can reach with the
same title.  (I won't attempt answering the ITC vs DTC question as I'd
probably get something wrong, but I'm sure others here can explain it).

[toc] | [prev] | [next] | [standalone]


#17608

FromMark Wills <forthfreak@gmail.com>
Date2012-11-27 09:01 -0800
Message-ID<c2ab11a4-aa3d-443e-b5cf-45db969256d8@b12g2000vbg.googlegroups.com>
In reply to#17606
On Nov 27, 4:55 pm, Paul Rubin <no.em...@nospam.invalid> wrote:
> Mark Wills <forthfr...@gmail.com> writes:
> > Very nice write-up here about the Forth virtual machine, complete with
> > links to our very own Anton Ertl's pages:
> >http://www.wordiq.com/definition/Forth_virtual_machine
>
> Note: That is a mirror of the wikipedia article you can reach with the
> same title.  (I won't attempt answering the ITC vs DTC question as I'd
> probably get something wrong, but I'm sure others here can explain it).

Ah! I didn't realise, thanks. Guess I should have checked. Thanks for
pointing it out! :-)

[toc] | [prev] | [next] | [standalone]


#17617

FromHugh Aguilar <hughaguilar96@yahoo.com>
Date2012-11-27 17:42 -0800
Message-ID<0782b510-4476-423b-b520-fa0fb1a730c6@r10g2000pbd.googlegroups.com>
In reply to#17603
On Nov 27, 9:01 am, Mark Wills <forthfr...@gmail.com> wrote:
> Very nice write-up here about the Forth virtual machine, complete with
> links to our very own Anton Ertl's pages:
>
> http://www.wordiq.com/definition/Forth_virtual_machine
>
> One thing that struck me, it mentions that direct threaded is not as
> "flexible" as ITC. I was wondering in what respect DTC is less
> flexible? Anyone have any opinions/comments?

With ITC it is possible to change how a word is interpreted by
changing the pointer at the cfa. With DTC, by comparison, you don't
have a pointer to the code that interprets the word, but rather you
have the code itself pasted in there. It is a major hassle to patch
this code to change how the word is interpreted. This is all academic
anyway --- I've never heard of anybody doing this. I think that it
would be done to provide a debug-interpretation in which the threaded
code is single-stepped through --- that is the only purpose I can
think of, but I haven't done it. I did write a single-step debugger
for my 65c02 cross-compiler, but it was subroutine-threaded --- if I
wanted to debug, I would recompile with the debug option turned on,
which would cause a BRK instruction to get compiled between every
chunk of code (representing a Forth word in the source-code). My
compiler would keep track of where all of these BRK instructions were,
so that when single-stepping through the program it would display the
correct block of source-code with a smiley-face showing where in the
block we were. It would put the user (me) into query-interpret, so I
could examine the 65c02 as necessary. Also, basic information such as
the parameter and return stacks, and some watch variables, was
continually displayed underneath the display of the source-code block.
This was all written in 16-bit UR/Forth, and the target was an Apple-
IIe computer, and they communicated with an RS-232 serial cable. I
wrote that back in maybe 1989. My application program was a symbolic
math program that would do calculus --- I got as far as determining
the derivative of a function, and reducing the equation to simplest
terms, but never got as far as symbolic integration of functions,
which is much more difficult.

I don't mess with debuggers nowadays --- Paul Rubin may find this hard
to believe, but it is not because I don't know how to write a
debugger, but it is because I find testing functions at the console
immediately after writing them to be more efficient.

BTW: The article listed "return threading" under the heading: "Less
often used are." This most likely is a reference to what I recently
learned over on clax and which those guys called "stack threading."
This is what I'm doing in HostForth. The processor return-stack
pointer (rsp on the 64-bit x86) is used as the Forth IP register. This
only works on big processors that don't use the application program's
return stack for their interrupts --- it won't work on micro-
controllers because an interrupt would overwrite the threaded Forth
code that is executing at the time that the interrupt occurs. I'm just
using it in HostForth because it is convenient and reasonably fast.
NEXT is just a single RET instruction. By comparison, in subroutine-
threading, NEXT is a CALL and a RET, so stack threading is faster for
executing colon words containing mostly primitives. For executing
colon words, stack threading requires DOCOLON code pasted in front of
the threaded code, whereas subroutine-threading still just uses a CALL
and a RET, so subroutine-threading is faster for executing colon words
containing mostly other colon words. Also, with subroutine-threading
you get to compile simple primitives as inline machine-code, which
speeds things up a lot. Mostly what kills the speed in any threaded
system is that branch prediction doesn't work, and so iteration comes
out slow -- all threaded schemes are slow because of this --- but I
don't care with HostForth because the only program that will ever be
written in HostForth is the cross-compiler TargForth, which is not
speed critical as it is all compile-time. TargForth will generate
subroutine-threaded code for the micro-controllers, as that code is
speed critical.

[toc] | [prev] | [next] | [standalone]


#17619

FromMark Wills <forthfreak@gmail.com>
Date2012-11-27 23:50 -0800
Message-ID<2e19f99e-2e8d-4de1-976a-3dd8099c14a4@f17g2000vbz.googlegroups.com>
In reply to#17617
On Nov 28, 1:42 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> On Nov 27, 9:01 am, Mark Wills <forthfr...@gmail.com> wrote:
>
> > Very nice write-up here about the Forth virtual machine, complete with
> > links to our very own Anton Ertl's pages:
>
> >http://www.wordiq.com/definition/Forth_virtual_machine
>
> > One thing that struck me, it mentions that direct threaded is not as
> > "flexible" as ITC. I was wondering in what respect DTC is less
> > flexible? Anyone have any opinions/comments?
>
> With ITC it is possible to change how a word is interpreted by
> changing the pointer at the cfa. With DTC, by comparison, you don't
> have a pointer to the code that interprets the word, but rather you
> have the code itself pasted in there.

I don't think that's correct, Hugh. Unless I'm mistaken, you're
thinking of native compiled code.

With DTC, a definition is still a 'thread' of addresses, but they are
the addresses of code, rather than the addresses of addresses of code;
a single cell references exactly one definition, same as ITC.

Maybe the article is mistaken. But I'm trying to think of what the
disadvantages of DTC are. I presume there *are* disadvantages,
otherwise ITC would not have evolved to be the defacto that it was
during the 70's and 80's.

The Rodriguez article sums it up quite nicely:

http://www.bradrodriguez.com/papers/moving1.htm

In the article, Rodriguez states that DTC can result in larger code
size:

"This costs space: every high-level definition in a Z80 Forth (for
example) is now one byte longer, since a 2-byte address has been
replaced by a 3-byte call. But this is not universally true. A 32-bit
68000 Forth may replace a 4-byte address with a 4-byte BSR
instruction, for no net loss. And on the Zilog Super8, which has
machine instructions for DTC Forth, the 2-byte address is replaced by
a 1-byte ENTER instruction, making a DTC Forth smaller on the Super8!"

But there is no mention of a loss of flexibility, which is what the
Wikipedia article states. I can't see any reason for a lack of
flexibility myself.

I guess the original article is simply erroneous, in stating that DTC
is less flexible than DTC.

[toc] | [prev] | [next] | [standalone]


#17621

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2012-11-28 04:41 -0600
Message-ID<Y7-dna8SH_nwdyjNnZ2dnUVZ7sWdnZ2d@supernews.com>
In reply to#17619
Mark Wills <forthfreak@gmail.com> wrote:
> On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>
>> With ITC it is possible to change how a word is interpreted by
>> changing the pointer at the cfa. With DTC, by comparison, you don't
>> have a pointer to the code that interprets the word, but rather you
>> have the code itself pasted in there.
> 
> I don't think that's correct, Hugh.

I'm sure it is.

Andrew.

[toc] | [prev] | [next] | [standalone]


#17622

FromMark Wills <forthfreak@gmail.com>
Date2012-11-28 02:48 -0800
Message-ID<41b8dd36-4a90-4a2a-b4e6-a738eac630f2@f17g2000vbz.googlegroups.com>
In reply to#17621
On Nov 28, 10:41 am, Andrew Haley <andre...@littlepinkcloud.invalid>
wrote:
> Mark Wills <forthfr...@gmail.com> wrote:
> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>
> >> With ITC it is possible to change how a word is interpreted by
> >> changing the pointer at the cfa. With DTC, by comparison, you don't
> >> have a pointer to the code that interprets the word, but rather you
> >> have the code itself pasted in there.
>
> > I don't think that's correct, Hugh.
>
> I'm sure it is.
>
> Andrew.

Eh?

[toc] | [prev] | [next] | [standalone]


#17624

FromAndrew Haley <andrew29@littlepinkcloud.invalid>
Date2012-11-28 05:27 -0600
Message-ID<25-dnfvDKpqEaCjNnZ2dnUVZ8mGdnZ2d@supernews.com>
In reply to#17622
Mark Wills <forthfreak@gmail.com> wrote:
> On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid>
> wrote:
>> Mark Wills <forthfr...@gmail.com> wrote:
>> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>>
>> >> With ITC it is possible to change how a word is interpreted by
>> >> changing the pointer at the cfa. With DTC, by comparison, you don't
>> >> have a pointer to the code that interprets the word, but rather you
>> >> have the code itself pasted in there.
>>
>> > I don't think that's correct, Hugh.
>>
>> I'm sure it is.
> 
> Eh?

I don't understand the problem you're having with my reply.  With ITC
it is possible to change how a word is interpreted by changing the
pointer at the cfa.  This is simply true, there is no doubt about it,
and your comment is incorrect.

Andrew.

[toc] | [prev] | [next] | [standalone]


#17626

FromMark Wills <forthfreak@gmail.com>
Date2012-11-28 03:51 -0800
Message-ID<6ee17c18-62a6-4233-bfe8-ae53343817fb@g6g2000vbk.googlegroups.com>
In reply to#17624
On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid>
wrote:
> Mark Wills <forthfr...@gmail.com> wrote:
> > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid>
> > wrote:
> >> Mark Wills <forthfr...@gmail.com> wrote:
> >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>
> >> >> With ITC it is possible to change how a word is interpreted by
> >> >> changing the pointer at the cfa. With DTC, by comparison, you don't
> >> >> have a pointer to the code that interprets the word, but rather you
> >> >> have the code itself pasted in there.
>
> >> > I don't think that's correct, Hugh.
>
> >> I'm sure it is.
>
> > Eh?
>
> I don't understand the problem you're having with my reply.  With ITC
> it is possible to change how a word is interpreted by changing the
> pointer at the cfa.  This is simply true, there is no doubt about it,
> and your comment is incorrect.
>
> Andrew.- Hide quoted text -
>
> - Show quoted text -

Oh. Okay. Yes, I see. I should have read Hugh's reply more closely
before posting.

So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR
whatever/etc. In DTC it's slightly more complicated, since the CFA
field would contain executable code. In my particular processor of
choice, the CFA field would be two cells wide. Yes, I can see that
patching it to change how it is interpreted could be a pain. Though I
suppose a dedicated helper word(s) could be provided to facilitate it.
Like Hugh says, it would be quite a rare occurence.

Okay. I get it. Sorry for the confusion.

Mark

[toc] | [prev] | [next] | [standalone]


#17658

FromHugh Aguilar <hughaguilar96@yahoo.com>
Date2012-11-28 21:56 -0800
Message-ID<3d8d574a-8ffd-4602-b176-36d6773a0d07@uc4g2000pbc.googlegroups.com>
In reply to#17626
On Nov 28, 4:51 am, Mark Wills <forthfr...@gmail.com> wrote:
> On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid>
> wrote:
>
> > Mark Wills <forthfr...@gmail.com> wrote:
> > > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid>
> > > wrote:
> > >> Mark Wills <forthfr...@gmail.com> wrote:
> > >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>
> > >> >> With ITC it is possible to change how a word is interpreted by
> > >> >> changing the pointer at the cfa. With DTC, by comparison, you don't
> > >> >> have a pointer to the code that interprets the word, but rather you
> > >> >> have the code itself pasted in there.
>
> > >> > I don't think that's correct, Hugh.
>
> > >> I'm sure it is.
>
> > > Eh?
>
> > I don't understand the problem you're having with my reply.  With ITC
> > it is possible to change how a word is interpreted by changing the
> > pointer at the cfa.  This is simply true, there is no doubt about it,
> > and your comment is incorrect.
>
> > Andrew.
>
> Oh. Okay. Yes, I see. I should have read Hugh's reply more closely
> before posting.
>
> So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR
> whatever/etc. In DTC it's slightly more complicated, since the CFA
> field would contain executable code. In my particular processor of
> choice, the CFA field would be two cells wide. Yes, I can see that
> patching it to change how it is interpreted could be a pain. Though I
> suppose a dedicated helper word(s) could be provided to facilitate it.
> Like Hugh says, it would be quite a rare occurence.
>
> Okay. I get it. Sorry for the confusion.
>
> Mark

I think that you have got it, but I'm not sure, so I'll go over it
again:

1.) In ITC what you have in front of the threaded code of the colon
word (or the body of the whatever), is a single pointer to the
interpreter for that kind of word (DOCOLON for colon words, etc.). It
is easy to store a different pointer in that slot, so the word will be
interpreted differently (DOCOLON-WITH-SINGLE-STEPPING for colon words,
for example).

2.) In DTC what you have in front of the threaded code of the colon
word (or the body of the whatever), is the actual machine-code of the
interpreter for that kind of word (DOCOLON for colon words, etc.). It
is difficult to patch this code, because it is actual code, rather
than a pointer to some code.

3.) There is a kind of hybrid between ITC and DTC. This is DTC in the
sense that we have machine-code in front of each word. However, this
machine-code always consists of a single CALL instruction to DOCOLON
etc.. When a CALL is executed, it puts the address just after itself
on the processor return-stack. Normally this is for RET to use to go
back. Here is the clever part though: this is the address of the body
of the Forth word. DOCOLON can load this address into the IP and begin
interpreting. In this case, you get DTC which is faster than ITC, but
you also get an easy way to change how a word is interpreted (just
store a new pointer into the operand of the CALL instruction).

The PDP-11 had an interesting feature. The JSR (its term for CALL)
would store the address after itself into a register, and it would
first push that register onto the return-stack. Effectively, the top
value of the return-stack was held in a register. But that register
could be your IP! You have DTC code and just do a JSR to DOCOLON (#3
above), and DOCOLON automatically gets the address of the threaded
code loaded into the IP. I figured this out way back in 1985 when I
was taking a class in assembly-language at the city college, which was
PDP-11. This works so well, that I had to suppose that the designers
of the PDP-11 were Forth programmers, or at least, were trying to
support DTC threaded code. I've never seen this feature on any other
processor. Even in 1985 though, the PDP-11 was obsolete --- the city
college was still teaching it just because they had all the textbooks,
but the professor cheerfully admitted that the PDP-11 was obsolete and
we would never use what we learned in the real world. I've always
thought that the PDP-11 was pretty cool though --- I wish somebody
would come out with a micro-controller that runs PDP-11 code and RT11
and all that --- maybe on an FPGA.

BTW: There is a discussion of threading over on comp.lang.asm.x86:
https://groups.google.com/group/comp.lang.asm.x86/browse_thread/thread/971adcb57df96272

Mark: Since your TI Forth system is ITC, why don't you take a stab at
writing a single-step source-level debugger? As I mentioned, I wrote
one for my 65c02 system. It is not as difficult as you might suppose.
I did it with screen-file source-code. It can be done with seq-file
source-code though, I would suppose. I don't think that a debugger is
all that useful, but writing one is pretty interesting --- and your
users will be impressed. :-)

Have fun!  Hugh

[toc] | [prev] | [next] | [standalone]


#17676

FromMark Wills <forthfreak@gmail.com>
Date2012-11-29 01:26 -0800
Message-ID<12dc647c-7250-46a4-ab66-2c6929f2b98a@n5g2000vbk.googlegroups.com>
In reply to#17658
On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> On Nov 28, 4:51 am, Mark Wills <forthfr...@gmail.com> wrote:
>
>
>
>
>
> > On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid>
> > wrote:
>
> > > Mark Wills <forthfr...@gmail.com> wrote:
> > > > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid>
> > > > wrote:
> > > >> Mark Wills <forthfr...@gmail.com> wrote:
> > > >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
>
> > > >> >> With ITC it is possible to change how a word is interpreted by
> > > >> >> changing the pointer at the cfa. With DTC, by comparison, you don't
> > > >> >> have a pointer to the code that interprets the word, but rather you
> > > >> >> have the code itself pasted in there.
>
> > > >> > I don't think that's correct, Hugh.
>
> > > >> I'm sure it is.
>
> > > > Eh?
>
> > > I don't understand the problem you're having with my reply.  With ITC
> > > it is possible to change how a word is interpreted by changing the
> > > pointer at the cfa.  This is simply true, there is no doubt about it,
> > > and your comment is incorrect.
>
> > > Andrew.
>
> > Oh. Okay. Yes, I see. I should have read Hugh's reply more closely
> > before posting.
>
> > So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR
> > whatever/etc. In DTC it's slightly more complicated, since the CFA
> > field would contain executable code. In my particular processor of
> > choice, the CFA field would be two cells wide. Yes, I can see that
> > patching it to change how it is interpreted could be a pain. Though I
> > suppose a dedicated helper word(s) could be provided to facilitate it.
> > Like Hugh says, it would be quite a rare occurence.
>
> > Okay. I get it. Sorry for the confusion.
>
> > Mark
>
> I think that you have got it, but I'm not sure, so I'll go over it
> again:
>
> 1.) In ITC what you have in front of the threaded code of the colon
> word (or the body of the whatever), is a single pointer to the
> interpreter for that kind of word (DOCOLON for colon words, etc.). It
> is easy to store a different pointer in that slot, so the word will be
> interpreted differently (DOCOLON-WITH-SINGLE-STEPPING for colon words,
> for example).
>
> 2.) In DTC what you have in front of the threaded code of the colon
> word (or the body of the whatever), is the actual machine-code of the
> interpreter for that kind of word (DOCOLON for colon words, etc.). It
> is difficult to patch this code, because it is actual code, rather
> than a pointer to some code.
>
> 3.) There is a kind of hybrid between ITC and DTC. This is DTC in the
> sense that we have machine-code in front of each word. However, this
> machine-code always consists of a single CALL instruction to DOCOLON
> etc.. When a CALL is executed, it puts the address just after itself
> on the processor return-stack. Normally this is for RET to use to go
> back. Here is the clever part though: this is the address of the body
> of the Forth word. DOCOLON can load this address into the IP and begin
> interpreting. In this case, you get DTC which is faster than ITC, but
> you also get an easy way to change how a word is interpreted (just
> store a new pointer into the operand of the CALL instruction).
>
> The PDP-11 had an interesting feature. The JSR (its term for CALL)
> would store the address after itself into a register, and it would
> first push that register onto the return-stack. Effectively, the top
> value of the return-stack was held in a register. But that register
> could be your IP! You have DTC code and just do a JSR to DOCOLON (#3
> above), and DOCOLON automatically gets the address of the threaded
> code loaded into the IP. I figured this out way back in 1985 when I
> was taking a class in assembly-language at the city college, which was
> PDP-11. This works so well, that I had to suppose that the designers
> of the PDP-11 were Forth programmers, or at least, were trying to
> support DTC threaded code. I've never seen this feature on any other
> processor. Even in 1985 though, the PDP-11 was obsolete --- the city
> college was still teaching it just because they had all the textbooks,
> but the professor cheerfully admitted that the PDP-11 was obsolete and
> we would never use what we learned in the real world. I've always
> thought that the PDP-11 was pretty cool though --- I wish somebody
> would come out with a micro-controller that runs PDP-11 code and RT11
> and all that --- maybe on an FPGA.
>
> BTW: There is a discussion of threading over on comp.lang.asm.x86:https://groups.google.com/group/comp.lang.asm.x86/browse_thread/threa...
>
> Mark: Since your TI Forth system is ITC, why don't you take a stab at
> writing a single-step source-level debugger? As I mentioned, I wrote
> one for my 65c02 system. It is not as difficult as you might suppose.
> I did it with screen-file source-code. It can be done with seq-file
> source-code though, I would suppose. I don't think that a debugger is
> all that useful, but writing one is pretty interesting --- and your
> users will be impressed. :-)
>
> Have fun!  Hugh- Hide quoted text -
>
> - Show quoted text -

Hi Hugh,

Thanks for the clarification.

I'm going to take a serious look at DTC because in my case, it's low
hanging fruit in terms of a 'cheap' way to gain a performance boost.
It's already very fast for what it is, running on a 3 mHZ 16-bit chip
with a multiplexed 8 bit data bus (the fastest Forth ever produced for
that machine).

I figure I can effectively lose DOCOL and EXIT altogether. What I mean
is, normally DOCOL and EXIT are written as subroutines that each colon
definition calls, as described by you above. Well, a branch
instruction on the 9900 is a 4 byte instruction; 2 bytes for the
instruction, and two bytes for the address (it's only a 2 byte
instruction if jumping via a register, but I digress).

The code for DOCOL would be:

DECT RP ; create entry on return stack (DECT=decrement by two)
MOV IP,*RP ; move instruction pointer value to return stack

That's also four bytes (2x 2byte instructions). So no point jumping to
a subroutine; doing so slows things down!

EXIT becomes

MOV *RP+,IP ; pop return address into IP
B *IP  ; go run that code

Also 4 bytes. So, again this gets inlined.

So there's a speed advantage right there. In addition, the single
level of indirection in DTC will increase performance (of the
"interpreter") by around 50% I think.

I'll sit down with a pencil and paper (my favourite method of working
stuff out) and the instruction set cycle counts at the weekend and
work out the saving of the overhead of DOCOL and EXIT. I think it'll
be a no-brainer though.

Also, primitives will not need a NEXT subroutine. Currently, in my
ITC, all primitives call NEXT at the end which moves along the thread
and causes the next word in the thread to be executed. Of course, the
branch to NEXT is again 4 bytes. But in DTC, I think it'll be two
bytes:

B *IP+ \ branch the next address in the thread

So each primitive executes the next word in the thread. There is no
"NEXT" - it's a single instruction. Again, huge payoff.

The more I think about, the more I convince myself!

As I say, DTC is relatively low hanging fruit; I don't if there would
be major complications with converting things like DOES> over. I'll
take a look at that.

The next step after that would be to have the compiler generate
machine code (subroutine threaded) but, on the TMS9900, which has no
stacks at all, there is I think no point in this. STC would possibly
be slower than DTC code.

Interesting stuff.

I'm not man enough to look at things like peephole optimisation on
native-code generating compilers. That's above my paygrade for now ;-)
I'm very interested in it though. And I wouldn't spend the time for
that on a 30 year old hobby. I'd move to an ARM project board and
study it in the context of an ARM based system.

But that's for another day. When I win the lottery, then I'll get onto
it!

[toc] | [prev] | [next] | [standalone]


#17750

FromHugh Aguilar <hughaguilar96@yahoo.com>
Date2012-11-29 22:34 -0800
Message-ID<a260b4eb-0a49-4576-b541-2c74c19a7927@v9g2000pbi.googlegroups.com>
In reply to#17676
On Nov 29, 2:26 am, Mark Wills <forthfr...@gmail.com> wrote:
> On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> > Mark: Since your TI Forth system is ITC, why don't you take a stab at
> > writing a single-step source-level debugger? As I mentioned, I wrote
> > one for my 65c02 system. It is not as difficult as you might suppose.
> > I did it with screen-file source-code. It can be done with seq-file
> > source-code though, I would suppose. I don't think that a debugger is
> > all that useful, but writing one is pretty interesting --- and your
> > users will be impressed. :-)

This suggestion was a bad idea. I wasn't thinking straight when I said
that.

I was able to write a source-level debugger because my 65c02 Forth was
a cross-compiler and it was running on an MS-DOS machine (it was
written in UR/Forth). The compiler needs to generate a large data
structure containing the addresses of every word compiled in every
definition. When the single-stepper is running, it stops at every one
of these (a BRK instruction in my system, although in an ITC system
DOCOLON will stop on every word). This address is looked up and the
corresponding source-code is displayed. The host computer has to have
a lot of memory for that gigantic data-structure, and it has to be
pretty fast. You don't want to try this on a 1980s vintage TI99/4A
computer --- you don't have the memory or the speed to do this --- it
taxed the limits of the 80386 computer that I was using as a host.

When I wrote that suggestion, I had forgotten that you don't have a
cross-compiler, but have an on-board Forth.

> Hi Hugh,
>
> Thanks for the clarification.
>
> I'm going to take a serious look at DTC because in my case, it's low
> hanging fruit in terms of a 'cheap' way to gain a performance boost.
> It's already very fast for what it is, running on a 3 mHZ 16-bit chip
> with a multiplexed 8 bit data bus (the fastest Forth ever produced for
> that machine).

For your old 16-bit computer, DTC should help to speed up the system.
It will also make the programs slightly larger. Instead of a pointer
in front of each colon word, you have a chunk of code. It is true that
NEXT should be smaller, so every primitive will be slightly smaller,
but this won't reduce the size of the system very much --- overall,
more memory will be needed.

A better way to boost the speed, is with some optimization. In many
cases, there are pairs of producers and consumers. For example, LIT is
a producer because it produces some data for the parameter stack, and
+ is a consumer because it consumes some data from the parameter
stack. These pairs are inefficient because the producer pushes data
onto the stack, and the consumer immediately pops that data off the
stack. The solution is to combine them into a single word. For
example, write a primitive LIT_+ that combines what LIT and + do. It
would hold the literal value in a register rather than push it onto
the stack and then pop it off again.

Even with ITC, it is possible to optimize pairs like this. Make your
compiler smart enough to remember what the last word it compiled was.
When it is ready to compile the next word, it checks what the last
word was and, if they are an optimizable pair, it compiles the combo
instead. For example, if the last thing you did was LIT, when you are
about to compile + your compiler will instead back up and get rid of
the LIT and replace it with LIT_+. This kind of peephole-optimization
not only makes your program faster, but smaller as well.

You are right though, that DTC is low-hanging fruit, and much easier
to implement. Peephole-optimization is somewhat more difficult, but
not unreasonably difficult. You can do the peephole-optimization on a
piece-meal basis. Start with + and make it smart enough to combine
with all the likely producers, then do ! and +! and so forth --- you
don't have to do everything at once, just doing + should boost the
speed significantly, and you can go from there.

BTW: I'm switching from DTC to ITC on my system. This is because I
realize (from reading this thread!), that with DTC the DOCOLON code is
scattered all around and won't be in the code cache, whereas with ITC
the whole VM should be in the code cache. That is only an issue on big
processors such as the modern x86 --- it is not an issue on the
TI99/4A. Also, listening to Paul Rubin promote single-step debuggers
has made me feel inclined to provide one for my system, and that is
easier with ITC than DTC --- I don't really like to use a debugger,
but other people do, and it isn't difficult to implement, so I might
as well go ahead and provide one.

[toc] | [prev] | [next] | [standalone]


#17756

FromMark Wills <forthfreak@gmail.com>
Date2012-11-30 01:42 -0800
Message-ID<8746568b-28b4-4466-bea6-5f5a7849126b@r4g2000vbi.googlegroups.com>
In reply to#17750
On Nov 30, 6:34 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> On Nov 29, 2:26 am, Mark Wills <forthfr...@gmail.com> wrote:
>
> > On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> > > Mark: Since your TI Forth system is ITC, why don't you take a stab at
> > > writing a single-step source-level debugger? As I mentioned, I wrote
> > > one for my 65c02 system. It is not as difficult as you might suppose.
> > > I did it with screen-file source-code. It can be done with seq-file
> > > source-code though, I would suppose. I don't think that a debugger is
> > > all that useful, but writing one is pretty interesting --- and your
> > > users will be impressed. :-)
>
> This suggestion was a bad idea. I wasn't thinking straight when I said
> that.
>

Oh. Well. Now you've gone and thrown the gauntlet down, haven't
you?! ;-)

> I was able to write a source-level debugger because my 65c02 Forth was
> a cross-compiler and it was running on an MS-DOS machine (it was
> written in UR/Forth). The compiler needs to generate a large data
> structure containing the addresses of every word compiled in every
> definition. When the single-stepper is running, it stops at every one
> of these (a BRK instruction in my system, although in an ITC system
> DOCOLON will stop on every word). This address is looked up and the
> corresponding source-code is displayed. The host computer has to have
> a lot of memory for that gigantic data-structure, and it has to be
> pretty fast. You don't want to try this on a 1980s vintage TI99/4A
> computer --- you don't have the memory or the speed to do this --- it
> taxed the limits of the 80386 computer that I was using as a host.
>

Well, I have already written a simple debugger. It's not a single
stepper, though. It works like this:

You load the debugger and it modifies : and ; such that each
subsequently defined colon definition makes a call into a word (can't
remember what it's called) that displays the name of the executing
word, and the data-stack.

As it goes, the depth of the return stack is measured, and it uses
this to produce indentations on the on-screen display - thus one can
see how one's program nests and un-nests.

A new word is defined, BREAK, which stops the program, returning to
the command line, and giving a full return stack dump. You can scatter
BREAKs around your code at point where you think it may be going awry.

For example:

: test swap break ;
: harry 3 test ;
: dick 2 harry ;
: tom 1 dick ;
tom

And you'd get output that looked like this:

>tom (1) 1
  >dick (2) 1 2
    >harry (3) 1 2 3
      >test (3) 1 2 3
BREAK in test in harry in dick in tom

Without a break, if you just let the program run, you'd get:

>tom (1) 1
  >dick (2) 1 2
    >harry (3) 1 2 3
      >test (3) 1 2 3
      <test (3) 1 3 2
    <harry (3) 1 3 2
  <dick (3) 1 3 2
<tom (3) 1 3 2

It's fairly simple to extend the above into a single step debugger. An
on-screen display showing the definition currently being executed with
a cursor pointing to the current word is less trivial, but it
possible. Again, I have a starting point, in that I already have SEE
for my system. So I already have code to de-compile a word. So it's
possible, and doesn't require a large list/table in memory, it's just
a different technique. In fact it would be an interesting excercise!

There's just one little itsy bitsy problem: Since I wrote my TRACER
program, I've used it once. And that was to demo it to someone else.
And the only reason I was showing it was to show them that facilities
such as a tracer/debugger can be written in Forth itself and the Forth
environment simply augmented with the functionality (they were
suitably impressed). I don't think I've touched it since. I just debug
at the command line.

In fact, I rarely use SEE. I only use SEE if I'm debugging a compiling
word. The last time I used SEE was a couple of weeks ago when I was
implementing your MACRO: idea (duly implemented as a loadable
extension and working beautifully - thank you for the inspiration!).
SEE has limitations (at least on my system) because some subroutines
in my system are headerless (don't have dictionary entries) so they
display as a ? when de-compiled. It's no problem to me, since I know
what's going on. But a newbie would wonder what's going on.
Unfortunately I don't have the ROM space available to allow headers
for everything. For example, DOES> compiles a DODOES, but DODOES is
headerless. This would be a problem in a single stepping debugger,
because it would not be possible to display the names for headerless
words. This is a limitation of my system due to memory constraints. I
only have 16K. My Forth system is implemented as a plug in cartridge:

http://turboforth.net/about_turboforth.html

>
> For your old 16-bit computer, DTC should help to speed up the system.
> It will also make the programs slightly larger. Instead of a pointer
> in front of each colon word, you have a chunk of code. It is true that
> NEXT should be smaller, so every primitive will be slightly smaller,
> but this won't reduce the size of the system very much --- overall,
> more memory will be needed.
>
I had a look at this yesterday and got myself tied up in knots. I
couldn't work out how to bootstrap the thing; to get it started. How
does the 'interpreter' for a high-level definition execute the words
in the thread. I couldn't figure it out in my lunch break and had to
junk what I had done. Need more time to concentrate. I was missing
something very fundamental. I was using the : SQUARE DUP * ; as my
target but didn't get anywhere.

> A better way to boost the speed, is with some optimization. In many
> cases, there are pairs of producers and consumers. For example, LIT is
> a producer because it produces some data for the parameter stack, and
> + is a consumer because it consumes some data from the parameter
> stack. These pairs are inefficient because the producer pushes data
> onto the stack, and the consumer immediately pops that data off the
> stack. The solution is to combine them into a single word. For
> example, write a primitive LIT_+ that combines what LIT and + do. It
> would hold the literal value in a register rather than push it onto
> the stack and then pop it off again.

This is an excellent suggestion. Perhaps an easier way (at least, in
terms of performing optimisations) is to make every word in the
dictionary immediate. Then, every word can 'look ahead' and see what
is about to be compiled and intervene accordingly. It would be very
difficult to produce a standard Forth with such a system though! I
wonder if anyone has previously experimented with such a technique?

>
> Even with ITC, it is possible to optimize pairs like this. Make your
> compiler smart enough to remember what the last word it compiled was.
> When it is ready to compile the next word, it checks what the last
> word was and, if they are an optimizable pair, it compiles the combo
> instead. For example, if the last thing you did was LIT, when you are
> about to compile + your compiler will instead back up and get rid of
> the LIT and replace it with LIT_+. This kind of peephole-optimization
> not only makes your program faster, but smaller as well.
>
I'll add your peephole suggestion to my "things to look at in the next
version" list. The next version (V2.0) is the version that I tell
myself I'm *not* going to write, but I know I 99.9% probably will.
It's like a bloody drug. It's the classic symptom of wanting to start
with a clean sheet, to implement all the 'lessons learned' that you
spent blood, sweat and tears learning on the first implementation.
There are many aspects of V1.x that have been re-written a couple of
times as I learned (from other Forthers, some here on this list) or
simply discovered (as part of the Forth awakening procees) better way
to do things.

[ and to the nay-sayers: I *do* write Forth code too. Not just a
compiler. But my Forth coding is for fun. I'm still learning. I write
stuff like this:

http://turboforth.net/tutorials/darkstar.html ]

I also want to spend some time looking at Smalltalk though (a project
for 2013) - not writing a smalltalk system, just learning the
language. It's the OOP equivalent of Forth. It's beautiful (though
very slow, I believe). It looks very interesting indeed to me. Despite
being pure OO it shares the idea of terseness and brevity and total
simplicity that Forth has.

Things for 2013:
* VFX (I want to do some simple SCADA stuff using serial and IP comms)
* Smalltalk
* TurboForth V2.0 (maybe - it'll be a part-time when-feel-like-it
thing)

> You are right though, that DTC is low-hanging fruit, and much easier
> to implement. Peephole-optimization is somewhat more difficult, but
> not unreasonably difficult. You can do the peephole-optimization on a
> piece-meal basis. Start with + and make it smart enough to combine
> with all the likely producers, then do ! and +! and so forth --- you
> don't have to do everything at once, just doing + should boost the
> speed significantly, and you can go from there.
>
Yeah. You've got me thinking now! I need a lot more memory to do this
though than I currently have. Still, I plan to make the next version a
64K EPROM but I can go up to 128K in an eprom if I need to. That's 16
8K pages which is a PITA, but doable.

> BTW: I'm switching from DTC to ITC on my system. This is because I
> realize (from reading this thread!), that with DTC the DOCOLON code is
> scattered all around and won't be in the code cache, whereas with ITC
> the whole VM should be in the code cache.

Well, if your high-level definitions make a CALL to DOCOLON then
there's no reason why DOCOLON would not be in the cache. The expense
of the call might be less than the delay induced by a cache miss.

However, I'd urge you to take a step back and a deep breath. I was
reading your post on the x86 group where you are discussing caching
etc. However, I have to point out that if the Forth you are intending
to produce is primarily for embedded systems then the chances of the
embedded system running on an x86 processor are quite low. It's much
more likely to be an ARM variant. In other words, don't allow key
descisions about the architecture of your system to be guided by
relatively un-important architectural constraints of a particular
processor family.

I'd urge you to get out a notebook and pencil. Sit down somewhere
quiet and write a list of key things that you want the system to do.
Design goals. Then put the list away. Reflect on it for a couple of
days and go back and make changes. Iterate. Eventually your thoughts/
ideas/requirements will coalesce. There's your plan/design goals. When
you've got it done, pin it up on the wall above your computer. Let it
be your guide as you develop the *project*, and when you feel a knee-
jerk step-change coming on, consult the plan again! Don't be swayed.
Stick to the plan. Have faith in the design decisions you made
earlier, even if you've had a bright idea.

I failed to make a plan/design and ended up with many many more
iterations/builds/bugs/teeth-knashing/wailing than I should have. It's
okay for me, because it's a hobby system and is given away for free.
Your aspirations are somewhat higher though, wanting a good system for
embedded targets. So I'd urge due consideration and diligence!

Just my two cents, FWIW!

Mark

[toc] | [prev] | [next] | [standalone]


#17783

FromHugh Aguilar <hughaguilar96@yahoo.com>
Date2012-11-30 13:18 -0800
Message-ID<379d76ca-9bee-48b2-a6d9-da39d74594e5@i2g2000pbi.googlegroups.com>
In reply to#17756
On Nov 30, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote:
> On Nov 30, 6:34 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> A new word is defined, BREAK, which stops the program, returning to
> the command line, and giving a full return stack dump. You can scatter
> BREAKs around your code at point where you think it may be going awry.

That is a good technique. I call it QI because it consists of QUERY
INTERPRET.

> There's just one little itsy bitsy problem: Since I wrote my TRACER
> program, I've used it once. And that was to demo it to someone else.
> And the only reason I was showing it was to show them that facilities
> such as a tracer/debugger can be written in Forth itself and the Forth
> environment simply augmented with the functionality (they were
> suitably impressed). I don't think I've touched it since. I just debug
> at the command line.

I agree --- it is easiest to just debug at the command line.

The only Forth single-step source-level debugger that I've ever used
was the one that I wrote myself for my 65c02 cross-compiler --- that
was a long time ago!

> I had a look at this yesterday and got myself tied up in knots. I
> couldn't work out how to bootstrap the thing; to get it started. How
> does the 'interpreter' for a high-level definition execute the words
> in the thread. I couldn't figure it out in my lunch break and had to
> junk what I had done. Need more time to concentrate. I was missing
> something very fundamental. I was using the : SQUARE DUP * ; as my
> target but didn't get anywhere.

ITC is more complicated than DTC because there is an extra level of
indirection. If you can't figure out how DTC works, but you've got ITC
already working, then I'm guessing that you ported your ITC over from
somewhere without really understanding it.

Keep at it --- if you still can't figure it out, contact me by email
and I will show you some code in x86 or whatever assembly-language you
want (not TI9900 though, as I don't know that one).

> > A better way to boost the speed, is with some optimization. In many
> > cases, there are pairs of producers and consumers. For example, LIT is
> > a producer because it produces some data for the parameter stack, and
> > + is a consumer because it consumes some data from the parameter
> > stack. These pairs are inefficient because the producer pushes data
> > onto the stack, and the consumer immediately pops that data off the
> > stack. The solution is to combine them into a single word. For
> > example, write a primitive LIT_+ that combines what LIT and + do. It
> > would hold the literal value in a register rather than push it onto
> > the stack and then pop it off again.
>
> This is an excellent suggestion. Perhaps an easier way (at least, in
> terms of performing optimisations) is to make every word in the
> dictionary immediate. Then, every word can 'look ahead' and see what
> is about to be compiled and intervene accordingly. It would be very
> difficult to produce a standard Forth with such a system though! I
> wonder if anyone has previously experimented with such a technique?

I don't recommend doing that. The way that I'm doing it in my own
system, is to smarten up what COMPILE, does. This doesn't just compile
the xt that it is given, but instead it puts the xt into a queue. When
it does this, it looks to see what xt is already in the queue, and
combines them if possible. On my system, the queue is only a single
item in length, so it is actually a variable not a queue --- I call it
LIMBO --- because the xt in there is in limbo, in the sense that it
has been compiled, but hasn't yet really been compiled.

> I also want to spend some time looking at Smalltalk though (a project
> for 2013) - not writing a smalltalk system, just learning the
> language. It's the OOP equivalent of Forth. It's beautiful (though
> very slow, I believe). It looks very interesting indeed to me. Despite
> being pure OO it shares the idea of terseness and brevity and total
> simplicity that Forth has.

Smalltalk is where the idea of dynamic-OOP originated. For the most
part though, CLOS really is where dynamic-OOP got going (there are
dynamic-OOP systems for Scheme too).

I would recommend learning Scheme or Lisp instead of Smalltalk, as
these still have an active community, which I don't think Smalltalk
does. I'm learning Scheme --- if you learn it too, we could bounce
ideas back and forth by email.

Learning Lisp has always been on my bucket list, but now I'm finally
doing it. :-)

> > BTW: I'm switching from DTC to ITC on my system. This is because I
> > realize (from reading this thread!), that with DTC the DOCOLON code is
> > scattered all around and won't be in the code cache, whereas with ITC
> > the whole VM should be in the code cache.
>
> Well, if your high-level definitions make a CALL to DOCOLON then
> there's no reason why DOCOLON would not be in the cache. The expense
> of the call might be less than the delay induced by a cache miss.

Yes, but the CALL isn't in the code cache with DTC. With ITC however,
you don't have code that calls your interpreter, but each word has a
pointer in the front (at the cfa) that points to the interpreter.

> However, I'd urge you to take a step back and a deep breath. I was
> reading your post on the x86 group where you are discussing caching
> etc. However, I have to point out that if the Forth you are intending
> to produce is primarily for embedded systems then the chances of the
> embedded system running on an x86 processor are quite low. It's much
> more likely to be an ARM variant. In other words, don't allow key
> descisions about the architecture of your system to be guided by
> relatively un-important architectural constraints of a particular
> processor family.

I'm writing two Forth systems. HostForth runs on the host computer
(the x86). TargForth is written in HostForth and it generates the
micro-controller code. I'm must working on HostForth right now.

What I was primarily trying to figure out on comp.lang.asm.x86 is how
to support overlays. I think I've got that figured out now though.
That is pretty important! It involves a fundamental design feature. If
I hadn't thought about this early on, but had left for later, I would
have been in trouble when later on when I discovered that supporting
overlays would be impossible without a complete rewrite. Fundamental
stuff like this really has to be figured out as early as possible.

I'm always able to learn something from those discussions on clax ---
most of those guys are really knowledgeable! --- unlike clf, where it
is mostly just baloney with b.s. frosting.

[toc] | [prev] | [next] | [standalone]


#17788

FromMark Wills <forthfreak@gmail.com>
Date2012-12-01 01:42 -0800
Message-ID<48bb9ca6-351c-471b-81be-9b69733bb33f@bx4g2000vbb.googlegroups.com>
In reply to#17783
On Nov 30, 9:18 pm, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> > I had a look at this yesterday and got myself tied up in knots. I
> > couldn't work out how to bootstrap the thing; to get it started. How
> > does the 'interpreter' for a high-level definition execute the words
> > in the thread. I couldn't figure it out in my lunch break and had to
> > junk what I had done. Need more time to concentrate. I was missing
> > something very fundamental. I was using the : SQUARE DUP * ; as my
> > target but didn't get anywhere.
>
> ITC is more complicated than DTC because there is an extra level of
> indirection. If you can't figure out how DTC works, but you've got ITC
> already working, then I'm guessing that you ported your ITC over from
> somewhere without really understanding it.
>
> Keep at it --- if you still can't figure it out, contact me by email
> and I will show you some code in x86 or whatever assembly-language you
> want (not TI9900 though, as I don't know that one).
>
Thanks. I got it working. It was much simpler than I thought.

It seems it doesn't make much difference on the 9900; it saves a
single MOV assembly instruction in NEXT (one less level of
indirection).

So, NEXT is two assembly instructions instead of three. That takes the
same space as a TMS9900 BRANCH instruction, so, at the end of a
primitive, rather than branching to NEXT, it can just be in-lined.
However, the Kernal runs in 8-bit memory, and, doing the math, it
looks like its faster to put next in 16-bit (0 wait state) memory, and
have primitives branch to NEXT.

It all evens out to pretty much the same. For sure, *not* worth coding
a new system in DTC on the 9900. It's so marginal that it's just not
worth it.

Maybe I'll take a look at a system that generates native machine code.
That opens up all sorts of possibilities for optimisation, such as in-
lining. Words can have a bit reserved in their dictionary entry that
determines is a word is to be in-lined or not. If not, a branch/call
is compiled, otherwise, the code is pasted into the current
definition.

It would be a nice learning exercise to learn about native code
compilers. Maybe I could incrementally add optimisations as I learn
the techniques. The peephole that you mentioned (I got a list of about
25-30 words that could be optimised using the combining literal
technique that you described). I was also reading about constant
folding on wikipedia. It's extremely clever, though I can't currently
see how it is actually implemented.

Any optimising compiler that I wrote for the 4A would have to be
simple, low-hanging fruit optimisations. It's not worth the effort to
write a compiler that compiles to intermediate language that is then
optimised. Simple optimisations would be the order of the day.

How far into HostForth are you?

[toc] | [prev] | [next] | [standalone]


#17840

FromHugh Aguilar <hughaguilar96@yahoo.com>
Date2012-12-03 15:25 -0800
Message-ID<9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com>
In reply to#17788
On Dec 1, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote:
> Thanks. I got it working. It was much simpler than I thought.
>
> It seems it doesn't make much difference on the 9900; it saves a
> single MOV assembly instruction in NEXT (one less level of
> indirection).
>
> So, NEXT is two assembly instructions instead of three. That takes the
> same space as a TMS9900 BRANCH instruction, so, at the end of a
> primitive, rather than branching to NEXT, it can just be in-lined.
> However, the Kernal runs in 8-bit memory, and, doing the math, it
> looks like its faster to put next in 16-bit (0 wait state) memory, and
> have primitives branch to NEXT.
>
> It all evens out to pretty much the same. For sure, *not* worth coding
> a new system in DTC on the 9900. It's so marginal that it's just not
> worth it.

I'm glad you figured out DTC. Most things are simple after you figure
them out! :-)

You are right that there is often not much difference between ITC and
DTC in speed. There was more difference on primitive processors that
lacked addressing-modes and/or didn't have enough registers. That was
a different world --- nowadays, it is all about cache efficiency.

> Maybe I'll take a look at a system that generates native machine code.
> That opens up all sorts of possibilities for optimisation, such as in-
> lining. Words can have a bit reserved in their dictionary entry that
> determines is a word is to be in-lined or not. If not, a branch/call
> is compiled, otherwise, the code is pasted into the current
> definition.

Inlining isn't all that good of a technique on modern processors. Your
code becomes bloated, which causes it to thrash the cache (Hey! That
rhymed! I'm a poet!).

Also, on the modern x86, CALL and RET are very efficient. Doing a CALL
to a function, and it doing a RET back again, is almost as fast as
inlining that function --- and it saves a lot of memory.

SwiftForth inlines small functions, but it just pastes them in. You
see code that pushes a datum onto the stack from a particular
register, and then immediately pops the datum back into that same
register. That isn't optimization! You are better off to just leave
those functions as functions, so you reduce your bloat.

> It would be a nice learning exercise to learn about native code
> compilers. Maybe I could incrementally add optimisations as I learn
> the techniques. The peephole that you mentioned (I got a list of about
> 25-30 words that could be optimised using the combining literal
> technique that you described). I was also reading about constant
> folding on wikipedia. It's extremely clever, though I can't currently
> see how it is actually implemented.
>
> Any optimising compiler that I wrote for the 4A would have to be
> simple, low-hanging fruit optimisations. It's not worth the effort to
> write a compiler that compiles to intermediate language that is then
> optimised. Simple optimisations would be the order of the day.

Well, you could stick with ITC and make your peephole-optimizer
combine word pairs. For example, OVER + would get compiled as a single
function: OVER_+ .

This works quite well. Also, it is not processor dependent. If you get
this to work on your TI9900, you can later port it over directly to
your ARM Forth. You will have to write all of the functions, such as
OVER and + and OVER_+ in ARM assembly, but this is easy. The
complicated part, of recognizing the word pairs and combining them,
will be exactly the same no matter what processor is underneath the
hood.

I wouldn't recommend writing an optimizer for generating machine-code.
That requires a lot of knowledge of the processor under the hood. Very
little that you do on the TI9900 would port over to the ARM, as they
are quite different. Stick with ITC though, and you are largely
processor independent.

Are you planning on jumping to the ARM or to the MSP430 in the future?
You know, you have to abandon that TI99/4A someday! What if you drop
it and it breaks? You can't go to WalMart and buy another one...

> How far into HostForth are you?

Not too far.

I had HLA code for generating optimized machine-code, but that was
getting complicated and I got bogged down. Then I decided to rewrite
in traditional assembly language and generate ITC code. I will only
generate optimized machine-code for the micro-controllers, where it
matters.

I liked HLA, but it is limited to 32-bit x86, and I want to be cutting-
edge for once in my life --- so I'm going with 64-bit x86 instead.

I'm taking it slow on Straight Forth because I have a lot to learn. I
don't know much about low-level stuff, such as caches. I also don't
know much about high-level stuff, such as closures. There seems to be
only a narrow window of mid-level stuff that I know about. lol

It is always worthwhile to learn new ideas, and to better oneself!

[toc] | [prev] | [next] | [standalone]


#17844

FromAlex McDonald <blog@rivadpm.com>
Date2012-12-03 16:18 -0800
Message-ID<50c1958c-3c4a-43aa-8520-a5f3e020bc33@n8g2000vbb.googlegroups.com>
In reply to#17840
On Dec 3, 11:25 pm, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> On Dec 1, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote:


> > Maybe I'll take a look at a system that generates native machine code.
> > That opens up all sorts of possibilities for optimisation, such as in-
> > lining. Words can have a bit reserved in their dictionary entry that
> > determines is a word is to be in-lined or not. If not, a branch/call
> > is compiled, otherwise, the code is pasted into the current
> > definition.
>
> Inlining isn't all that good of a technique on modern processors. Your
> code becomes bloated, which causes it to thrash the cache (Hey! That
> rhymed! I'm a poet!).
>
> Also, on the modern x86, CALL and RET are very efficient. Doing a CALL
> to a function, and it doing a RET back again, is almost as fast as
> inlining that function --- and it saves a lot of memory.

No it doesn't.

>
> SwiftForth inlines small functions, but it just pastes them in. You
> see code that pushes a datum onto the stack from a particular
> register, and then immediately pops the datum back into that same
> register. That isn't optimization! You are better off to just leave
> those functions as functions, so you reduce your bloat.
>

STC Experimental 32bit: 0.06.05 Build: 363

With simple inlining of words <10 bytes long

Test time including overhead               ms     times     ns (each)
Eratosthenes sieve 1899 Primes            120     8190000      14
Fibonacci recursion ( 35 -> 9227465 )      65     9227430       7
Hoare's quick sort (reverse order)        129     2000000      64
Generate random numbers (1024 kb array)    92     262144      350
LZ77 Comp. (400 kb Random Data Mem>Mem)   134     1
Dhrystone (integer)                        81     500000      162
6172839 Dh
rystones/sec
Total:                                    647     1
APP mem: 113,865, CODE mem: 15,273, SYS mem: 5,488 Total: 134,626

Without inlining

Test time including overhead               ms     times     ns (each)
Eratosthenes sieve 1899 Primes            126     8190000      15
Fibonacci recursion ( 35 -> 9227465 )      66     9227430       7
Hoare's quick sort (reverse order)        252     2000000     126
Generate random numbers (1024 kb array)   104     262144      396
LZ77 Comp. (400 kb Random Data Mem>Mem)   155     1
Dhrystone (integer)                        88     500000      176
5681818 Dh
rystones/sec
Total:                                    800     1
APP mem: 113,865, CODE mem: 14,799, SYS mem: 5,488 Total: 134,152

A overall speed decrease of 25% for 470 bytes extra code.

[toc] | [prev] | [next] | [standalone]


#17845

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-12-03 21:20 -0500
Message-ID<k9jmcb$ilv$1@speranza.aioe.org>
In reply to#17840
"Hugh Aguilar" <hughaguilar96@yahoo.com> wrote in message
news:9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com...
...

> Also, on the modern x86, CALL and RET are very efficient. Doing
> a CALL to a function, and it doing a RET back again, is almost
> as fast as inlining that function --- [...]

That may be almost true now.  But, I don't know how true it is
that they're fast now.  I haven't checked the x86 manuals in three
to five years.  Even if they're fast now, CALL and RET still have
some overhead that's not present with inlining.  Historically,
CALL and RET being fast on x86 wasn't true.

The problem with RET on modern x86 - according to those on
c.l.a.x. - is that it must be _matched_ with a CALL or it causes a
slowdown of the processor.  E.g., the RET location was pushed
onto the stack via PUSH instead of by a CALL.


Rod Pemberton


[toc] | [prev] | [next] | [standalone]


#17846

From"Elizabeth D. Rather" <erather@forth.com>
Date2012-12-03 17:29 -1000
Message-ID<0ZKdnRs5KJ-48yDNnZ2dnUVZ_rCdnZ2d@supernews.com>
In reply to#17845
On 12/3/12 4:20 PM, Rod Pemberton wrote:
...
>> Also, on the modern x86, CALL and RET are very efficient. Doing
>> a CALL to a function, and it doing a RET back again, is almost
>> as fast as inlining that function --- [...]
>
> That may be almost true now.  But, I don't know how true it is
> that they're fast now.  I haven't checked the x86 manuals in three
> to five years.  Even if they're fast now, CALL and RET still have
> some overhead that's not present with inlining.  Historically,
> CALL and RET being fast on x86 wasn't true.

Whether it's worth inlining or not really depends on the length of the 
code being inlined. If it's just a few instructions, the ratio of the 
CALL/RET to the code is such that the inlining pays off. For a longer 
sequence, it does not. It also depends on whether the CALL is set up by 
C, which adds overhead for calling sequences that is missing in Forth 
written in Forth/assembler.

Cheers,
Elizabeth

-- 
==================================================
Elizabeth D. Rather   (US & Canada)   800-55-FORTH
FORTH Inc.                         +1 310.999.6784
5959 West Century Blvd. Suite 700
Los Angeles, CA 90045
http://www.forth.com

"Forth-based products and Services for real-time
applications since 1973."
==================================================

[toc] | [prev] | [next] | [standalone]


#17847

FromMark Wills <forthfreak@gmail.com>
Date2012-12-04 00:12 -0800
Message-ID<9e119760-3e3a-4bcb-a7b2-f540695bdbe5@n8g2000vbb.googlegroups.com>
In reply to#17845
On Dec 4, 2:20 am, "Rod Pemberton" <do_not_h...@notemailnotz.cnm>
wrote:
> "Hugh Aguilar" <hughaguila...@yahoo.com> wrote in message
>
> news:9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com...
> ...
>
> > Also, on the modern x86, CALL and RET are very efficient. Doing
> > a CALL to a function, and it doing a RET back again, is almost
> > as fast as inlining that function --- [...]
>
> That may be almost true now.  But, I don't know how true it is
> that they're fast now.  I haven't checked the x86 manuals in three
> to five years.  Even if they're fast now, CALL and RET still have
> some overhead that's not present with inlining.  Historically,
> CALL and RET being fast on x86 wasn't true.
>
> The problem with RET on modern x86 - according to those on
> c.l.a.x. - is that it must be _matched_ with a CALL or it causes a
> slowdown of the processor.  E.g., the RET location was pushed
> onto the stack via PUSH instead of by a CALL.
>
> Rod Pemberton

I would imagine (I have no experience) that with CALL/RET you also run
the risk of the routine that you are CALLing not being in the cache,
which adds a further performance penalty. At least if the the code is
inlined (even with inefficiencies such as pushing to the data stack
and immediately popping again) there is a good chance it's running
from cache.

My knowledge of cache's is very 1990's though; maybe they are a lot
cleverer these days.

What is the difference between a level 1 and a level 2 cache? Is the
level 2 cache a cache for the level 1 cache? So there are two caches
between the CPU and external memory? Is that how it works?

Do caches run 'metrics' on subroutines like "Hmmm... This subroutine
here seems to be called a lot more often than these others. I'll keep
it in my cache where I can access it quickly" or are they simply dumb,
where, if a section of memory is called for, and it's not in the
cache, it reads the memory, plus n bytes into the cache?

Guess I should read up on caches!

I don't have any such complications in my hobby system! It's nice and
simple. The only potential complication is the TMS9995 because it has
instruction prefetch. This means some self-modifying code can trip you
up, but only if you are modifying the instruction immediately in
front. A simple NOP between the instructions fixes that.

[toc] | [prev] | [next] | [standalone]


Page 1 of 6  [1] 2 3 4 5 6  Next page →

Back to top | Article view | comp.lang.forth


csiph-web