Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #17603 > unrolled thread
| Started by | Mark Wills <forthfreak@gmail.com> |
|---|---|
| First post | 2012-11-27 08:01 -0800 |
| Last post | 2012-11-28 14:21 +0000 |
| Articles | 20 on this page of 112 — 16 participants |
Back to article view | Back to comp.lang.forth
DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 08:01 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-11-27 08:55 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 09:01 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-27 17:42 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-27 23:50 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 04:41 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 02:48 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 05:27 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 03:51 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 21:56 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-29 01:26 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-29 22:34 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-30 01:42 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 13:18 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-01 01:42 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-03 15:25 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-03 16:18 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-03 21:20 -0500
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-03 17:29 -1000
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 00:12 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-05 11:35 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-04 20:17 -0800
Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-05 08:31 +0200
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:48 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-04 23:53 -0800
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-05 12:13 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 15:26 -0800
Re: DTC Ron Aaron <rambamist@gmail.com> - 2012-12-06 06:32 +0200
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 01:07 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 04:23 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:49 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:42 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:30 -0800
Re: DTC Paul Rubin <no.email@nospam.invalid> - 2012-12-06 09:45 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 21:41 +0000
Re: DTC "A. K." <akk@nospam.org> - 2012-12-06 23:15 +0100
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-08 01:27 +0000
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 11:19 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 03:55 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:44 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:05 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:35 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:20 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-09 01:01 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:10 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 18:57 +0100
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 19:46 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 11:23 -0800
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-08 13:37 +0100
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 14:41 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-08 06:15 -0800
Re: DTC "A. K." <akk@nospam.org> - 2012-12-08 17:07 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:16 -0800
Re: DTC Brad Eckert <hwfwguy@gmail.com> - 2012-12-07 09:00 -0800
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:21 -0800
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 08:01 -1000
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 14:18 -0800
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-12-06 13:48 -1000
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-06 19:03 -0500
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-07 02:59 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-12-06 13:13 +0000
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-12-05 20:39 -0500
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-06 15:46 +0100
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-12-06 07:47 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-06 08:36 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-12-05 04:03 -0800
Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:30 -0800
Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-11 23:52 -0500
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-12-12 10:59 -0800
Re: DTC David Thompson <dave.thompson2@verizon.net> - 2012-12-31 02:43 -0500
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2013-01-02 00:44 -0800
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-28 13:54 +0000
Re: DTC "Clyde W. Phillips Jr." <cwpjr02@gmail.com> - 2012-12-06 20:16 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-28 06:58 -0500
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 04:50 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 07:08 -0600
Re: DTC Mark Wills <forthfreak@gmail.com> - 2012-11-28 06:02 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:23 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:18 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 08:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 15:00 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 09:18 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 16:36 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 11:02 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 17:13 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:03 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 18:12 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-28 12:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-29 14:30 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-29 18:05 +0100
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-29 11:19 -0800
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 03:14 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 14:12 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:32 -0600
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-30 16:40 +0000
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-01 15:34 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-01 21:23 +0100
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 16:28 +0000
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 18:44 +0100
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-12-03 04:50 -0600
Re: DTC Bernd Paysan <bernd.paysan@gmx.de> - 2012-12-03 16:48 +0100
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-12-03 15:59 +0000
Re: DTC albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-30 15:34 +0000
Re: DTC Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-30 10:36 -0600
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-30 12:47 -0800
Re: DTC Alex McDonald <blog@rivadpm.com> - 2012-11-28 11:29 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:15 -0500
Re: DTC "Elizabeth D. Rather" <erather@forth.com> - 2012-11-29 08:52 -1000
Re: DTC Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-11-28 22:25 -0800
Re: DTC "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-29 04:12 -0500
Re: DTC humptydumpty <ouatubi@gmail.com> - 2012-11-28 02:04 -0800
Re: DTC anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-28 14:21 +0000
Page 1 of 6 [1] 2 3 4 5 6 Next page →
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-27 08:01 -0800 |
| Subject | DTC |
| Message-ID | <ad599bc5-3e28-4fe0-a2d5-dc9b9dfb2255@eo2g2000vbb.googlegroups.com> |
Very nice write-up here about the Forth virtual machine, complete with links to our very own Anton Ertl's pages: http://www.wordiq.com/definition/Forth_virtual_machine One thing that struck me, it mentions that direct threaded is not as "flexible" as ITC. I was wondering in what respect DTC is less flexible? Anyone have any opinions/comments? For context, here is the description of DTC from the above link: "Direct threading: The addresses in the code are actually the address of machine language. This is a compromise between speed and space. The indirect data pointer is lost, at some loss in the language's flexibility, and this may need to be corrected by a type tag in the data areas, with an auxiliary table. Some Forth systems have produced direct-threaded code. On many machines direct-threading is faster than subroutine threading (see reference below)." The reference in the above paragraph points to a paper by Anton on threading benchmarks.
[toc] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-11-27 08:55 -0800 |
| Message-ID | <7xy5hnastn.fsf@ruckus.brouhaha.com> |
| In reply to | #17603 |
Mark Wills <forthfreak@gmail.com> writes: > Very nice write-up here about the Forth virtual machine, complete with > links to our very own Anton Ertl's pages: > http://www.wordiq.com/definition/Forth_virtual_machine Note: That is a mirror of the wikipedia article you can reach with the same title. (I won't attempt answering the ITC vs DTC question as I'd probably get something wrong, but I'm sure others here can explain it).
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-27 09:01 -0800 |
| Message-ID | <c2ab11a4-aa3d-443e-b5cf-45db969256d8@b12g2000vbg.googlegroups.com> |
| In reply to | #17606 |
On Nov 27, 4:55 pm, Paul Rubin <no.em...@nospam.invalid> wrote: > Mark Wills <forthfr...@gmail.com> writes: > > Very nice write-up here about the Forth virtual machine, complete with > > links to our very own Anton Ertl's pages: > >http://www.wordiq.com/definition/Forth_virtual_machine > > Note: That is a mirror of the wikipedia article you can reach with the > same title. (I won't attempt answering the ITC vs DTC question as I'd > probably get something wrong, but I'm sure others here can explain it). Ah! I didn't realise, thanks. Guess I should have checked. Thanks for pointing it out! :-)
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-11-27 17:42 -0800 |
| Message-ID | <0782b510-4476-423b-b520-fa0fb1a730c6@r10g2000pbd.googlegroups.com> |
| In reply to | #17603 |
On Nov 27, 9:01 am, Mark Wills <forthfr...@gmail.com> wrote: > Very nice write-up here about the Forth virtual machine, complete with > links to our very own Anton Ertl's pages: > > http://www.wordiq.com/definition/Forth_virtual_machine > > One thing that struck me, it mentions that direct threaded is not as > "flexible" as ITC. I was wondering in what respect DTC is less > flexible? Anyone have any opinions/comments? With ITC it is possible to change how a word is interpreted by changing the pointer at the cfa. With DTC, by comparison, you don't have a pointer to the code that interprets the word, but rather you have the code itself pasted in there. It is a major hassle to patch this code to change how the word is interpreted. This is all academic anyway --- I've never heard of anybody doing this. I think that it would be done to provide a debug-interpretation in which the threaded code is single-stepped through --- that is the only purpose I can think of, but I haven't done it. I did write a single-step debugger for my 65c02 cross-compiler, but it was subroutine-threaded --- if I wanted to debug, I would recompile with the debug option turned on, which would cause a BRK instruction to get compiled between every chunk of code (representing a Forth word in the source-code). My compiler would keep track of where all of these BRK instructions were, so that when single-stepping through the program it would display the correct block of source-code with a smiley-face showing where in the block we were. It would put the user (me) into query-interpret, so I could examine the 65c02 as necessary. Also, basic information such as the parameter and return stacks, and some watch variables, was continually displayed underneath the display of the source-code block. This was all written in 16-bit UR/Forth, and the target was an Apple- IIe computer, and they communicated with an RS-232 serial cable. I wrote that back in maybe 1989. My application program was a symbolic math program that would do calculus --- I got as far as determining the derivative of a function, and reducing the equation to simplest terms, but never got as far as symbolic integration of functions, which is much more difficult. I don't mess with debuggers nowadays --- Paul Rubin may find this hard to believe, but it is not because I don't know how to write a debugger, but it is because I find testing functions at the console immediately after writing them to be more efficient. BTW: The article listed "return threading" under the heading: "Less often used are." This most likely is a reference to what I recently learned over on clax and which those guys called "stack threading." This is what I'm doing in HostForth. The processor return-stack pointer (rsp on the 64-bit x86) is used as the Forth IP register. This only works on big processors that don't use the application program's return stack for their interrupts --- it won't work on micro- controllers because an interrupt would overwrite the threaded Forth code that is executing at the time that the interrupt occurs. I'm just using it in HostForth because it is convenient and reasonably fast. NEXT is just a single RET instruction. By comparison, in subroutine- threading, NEXT is a CALL and a RET, so stack threading is faster for executing colon words containing mostly primitives. For executing colon words, stack threading requires DOCOLON code pasted in front of the threaded code, whereas subroutine-threading still just uses a CALL and a RET, so subroutine-threading is faster for executing colon words containing mostly other colon words. Also, with subroutine-threading you get to compile simple primitives as inline machine-code, which speeds things up a lot. Mostly what kills the speed in any threaded system is that branch prediction doesn't work, and so iteration comes out slow -- all threaded schemes are slow because of this --- but I don't care with HostForth because the only program that will ever be written in HostForth is the cross-compiler TargForth, which is not speed critical as it is all compile-time. TargForth will generate subroutine-threaded code for the micro-controllers, as that code is speed critical.
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-27 23:50 -0800 |
| Message-ID | <2e19f99e-2e8d-4de1-976a-3dd8099c14a4@f17g2000vbz.googlegroups.com> |
| In reply to | #17617 |
On Nov 28, 1:42 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > On Nov 27, 9:01 am, Mark Wills <forthfr...@gmail.com> wrote: > > > Very nice write-up here about the Forth virtual machine, complete with > > links to our very own Anton Ertl's pages: > > >http://www.wordiq.com/definition/Forth_virtual_machine > > > One thing that struck me, it mentions that direct threaded is not as > > "flexible" as ITC. I was wondering in what respect DTC is less > > flexible? Anyone have any opinions/comments? > > With ITC it is possible to change how a word is interpreted by > changing the pointer at the cfa. With DTC, by comparison, you don't > have a pointer to the code that interprets the word, but rather you > have the code itself pasted in there. I don't think that's correct, Hugh. Unless I'm mistaken, you're thinking of native compiled code. With DTC, a definition is still a 'thread' of addresses, but they are the addresses of code, rather than the addresses of addresses of code; a single cell references exactly one definition, same as ITC. Maybe the article is mistaken. But I'm trying to think of what the disadvantages of DTC are. I presume there *are* disadvantages, otherwise ITC would not have evolved to be the defacto that it was during the 70's and 80's. The Rodriguez article sums it up quite nicely: http://www.bradrodriguez.com/papers/moving1.htm In the article, Rodriguez states that DTC can result in larger code size: "This costs space: every high-level definition in a Z80 Forth (for example) is now one byte longer, since a 2-byte address has been replaced by a 3-byte call. But this is not universally true. A 32-bit 68000 Forth may replace a 4-byte address with a 4-byte BSR instruction, for no net loss. And on the Zilog Super8, which has machine instructions for DTC Forth, the 2-byte address is replaced by a 1-byte ENTER instruction, making a DTC Forth smaller on the Super8!" But there is no mention of a loss of flexibility, which is what the Wikipedia article states. I can't see any reason for a lack of flexibility myself. I guess the original article is simply erroneous, in stating that DTC is less flexible than DTC.
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 04:41 -0600 |
| Message-ID | <Y7-dna8SH_nwdyjNnZ2dnUVZ7sWdnZ2d@supernews.com> |
| In reply to | #17619 |
Mark Wills <forthfreak@gmail.com> wrote: > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > >> With ITC it is possible to change how a word is interpreted by >> changing the pointer at the cfa. With DTC, by comparison, you don't >> have a pointer to the code that interprets the word, but rather you >> have the code itself pasted in there. > > I don't think that's correct, Hugh. I'm sure it is. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-28 02:48 -0800 |
| Message-ID | <41b8dd36-4a90-4a2a-b4e6-a738eac630f2@f17g2000vbz.googlegroups.com> |
| In reply to | #17621 |
On Nov 28, 10:41 am, Andrew Haley <andre...@littlepinkcloud.invalid> wrote: > Mark Wills <forthfr...@gmail.com> wrote: > > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > >> With ITC it is possible to change how a word is interpreted by > >> changing the pointer at the cfa. With DTC, by comparison, you don't > >> have a pointer to the code that interprets the word, but rather you > >> have the code itself pasted in there. > > > I don't think that's correct, Hugh. > > I'm sure it is. > > Andrew. Eh?
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-11-28 05:27 -0600 |
| Message-ID | <25-dnfvDKpqEaCjNnZ2dnUVZ8mGdnZ2d@supernews.com> |
| In reply to | #17622 |
Mark Wills <forthfreak@gmail.com> wrote: > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid> > wrote: >> Mark Wills <forthfr...@gmail.com> wrote: >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: >> >> >> With ITC it is possible to change how a word is interpreted by >> >> changing the pointer at the cfa. With DTC, by comparison, you don't >> >> have a pointer to the code that interprets the word, but rather you >> >> have the code itself pasted in there. >> >> > I don't think that's correct, Hugh. >> >> I'm sure it is. > > Eh? I don't understand the problem you're having with my reply. With ITC it is possible to change how a word is interpreted by changing the pointer at the cfa. This is simply true, there is no doubt about it, and your comment is incorrect. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-28 03:51 -0800 |
| Message-ID | <6ee17c18-62a6-4233-bfe8-ae53343817fb@g6g2000vbk.googlegroups.com> |
| In reply to | #17624 |
On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid> wrote: > Mark Wills <forthfr...@gmail.com> wrote: > > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid> > > wrote: > >> Mark Wills <forthfr...@gmail.com> wrote: > >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > >> >> With ITC it is possible to change how a word is interpreted by > >> >> changing the pointer at the cfa. With DTC, by comparison, you don't > >> >> have a pointer to the code that interprets the word, but rather you > >> >> have the code itself pasted in there. > > >> > I don't think that's correct, Hugh. > > >> I'm sure it is. > > > Eh? > > I don't understand the problem you're having with my reply. With ITC > it is possible to change how a word is interpreted by changing the > pointer at the cfa. This is simply true, there is no doubt about it, > and your comment is incorrect. > > Andrew.- Hide quoted text - > > - Show quoted text - Oh. Okay. Yes, I see. I should have read Hugh's reply more closely before posting. So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR whatever/etc. In DTC it's slightly more complicated, since the CFA field would contain executable code. In my particular processor of choice, the CFA field would be two cells wide. Yes, I can see that patching it to change how it is interpreted could be a pain. Though I suppose a dedicated helper word(s) could be provided to facilitate it. Like Hugh says, it would be quite a rare occurence. Okay. I get it. Sorry for the confusion. Mark
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-11-28 21:56 -0800 |
| Message-ID | <3d8d574a-8ffd-4602-b176-36d6773a0d07@uc4g2000pbc.googlegroups.com> |
| In reply to | #17626 |
On Nov 28, 4:51 am, Mark Wills <forthfr...@gmail.com> wrote: > On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid> > wrote: > > > Mark Wills <forthfr...@gmail.com> wrote: > > > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid> > > > wrote: > > >> Mark Wills <forthfr...@gmail.com> wrote: > > >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > > >> >> With ITC it is possible to change how a word is interpreted by > > >> >> changing the pointer at the cfa. With DTC, by comparison, you don't > > >> >> have a pointer to the code that interprets the word, but rather you > > >> >> have the code itself pasted in there. > > > >> > I don't think that's correct, Hugh. > > > >> I'm sure it is. > > > > Eh? > > > I don't understand the problem you're having with my reply. With ITC > > it is possible to change how a word is interpreted by changing the > > pointer at the cfa. This is simply true, there is no doubt about it, > > and your comment is incorrect. > > > Andrew. > > Oh. Okay. Yes, I see. I should have read Hugh's reply more closely > before posting. > > So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR > whatever/etc. In DTC it's slightly more complicated, since the CFA > field would contain executable code. In my particular processor of > choice, the CFA field would be two cells wide. Yes, I can see that > patching it to change how it is interpreted could be a pain. Though I > suppose a dedicated helper word(s) could be provided to facilitate it. > Like Hugh says, it would be quite a rare occurence. > > Okay. I get it. Sorry for the confusion. > > Mark I think that you have got it, but I'm not sure, so I'll go over it again: 1.) In ITC what you have in front of the threaded code of the colon word (or the body of the whatever), is a single pointer to the interpreter for that kind of word (DOCOLON for colon words, etc.). It is easy to store a different pointer in that slot, so the word will be interpreted differently (DOCOLON-WITH-SINGLE-STEPPING for colon words, for example). 2.) In DTC what you have in front of the threaded code of the colon word (or the body of the whatever), is the actual machine-code of the interpreter for that kind of word (DOCOLON for colon words, etc.). It is difficult to patch this code, because it is actual code, rather than a pointer to some code. 3.) There is a kind of hybrid between ITC and DTC. This is DTC in the sense that we have machine-code in front of each word. However, this machine-code always consists of a single CALL instruction to DOCOLON etc.. When a CALL is executed, it puts the address just after itself on the processor return-stack. Normally this is for RET to use to go back. Here is the clever part though: this is the address of the body of the Forth word. DOCOLON can load this address into the IP and begin interpreting. In this case, you get DTC which is faster than ITC, but you also get an easy way to change how a word is interpreted (just store a new pointer into the operand of the CALL instruction). The PDP-11 had an interesting feature. The JSR (its term for CALL) would store the address after itself into a register, and it would first push that register onto the return-stack. Effectively, the top value of the return-stack was held in a register. But that register could be your IP! You have DTC code and just do a JSR to DOCOLON (#3 above), and DOCOLON automatically gets the address of the threaded code loaded into the IP. I figured this out way back in 1985 when I was taking a class in assembly-language at the city college, which was PDP-11. This works so well, that I had to suppose that the designers of the PDP-11 were Forth programmers, or at least, were trying to support DTC threaded code. I've never seen this feature on any other processor. Even in 1985 though, the PDP-11 was obsolete --- the city college was still teaching it just because they had all the textbooks, but the professor cheerfully admitted that the PDP-11 was obsolete and we would never use what we learned in the real world. I've always thought that the PDP-11 was pretty cool though --- I wish somebody would come out with a micro-controller that runs PDP-11 code and RT11 and all that --- maybe on an FPGA. BTW: There is a discussion of threading over on comp.lang.asm.x86: https://groups.google.com/group/comp.lang.asm.x86/browse_thread/thread/971adcb57df96272 Mark: Since your TI Forth system is ITC, why don't you take a stab at writing a single-step source-level debugger? As I mentioned, I wrote one for my 65c02 system. It is not as difficult as you might suppose. I did it with screen-file source-code. It can be done with seq-file source-code though, I would suppose. I don't think that a debugger is all that useful, but writing one is pretty interesting --- and your users will be impressed. :-) Have fun! Hugh
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-29 01:26 -0800 |
| Message-ID | <12dc647c-7250-46a4-ab66-2c6929f2b98a@n5g2000vbk.googlegroups.com> |
| In reply to | #17658 |
On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > On Nov 28, 4:51 am, Mark Wills <forthfr...@gmail.com> wrote: > > > > > > > On Nov 28, 11:27 am, Andrew Haley <andre...@littlepinkcloud.invalid> > > wrote: > > > > Mark Wills <forthfr...@gmail.com> wrote: > > > > On Nov 28, 10:41?am, Andrew Haley <andre...@littlepinkcloud.invalid> > > > > wrote: > > > >> Mark Wills <forthfr...@gmail.com> wrote: > > > >> > On Nov 28, 1:42?am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > > > >> >> With ITC it is possible to change how a word is interpreted by > > > >> >> changing the pointer at the cfa. With DTC, by comparison, you don't > > > >> >> have a pointer to the code that interprets the word, but rather you > > > >> >> have the code itself pasted in there. > > > > >> > I don't think that's correct, Hugh. > > > > >> I'm sure it is. > > > > > Eh? > > > > I don't understand the problem you're having with my reply. With ITC > > > it is possible to change how a word is interpreted by changing the > > > pointer at the cfa. This is simply true, there is no doubt about it, > > > and your comment is incorrect. > > > > Andrew. > > > Oh. Okay. Yes, I see. I should have read Hugh's reply more closely > > before posting. > > > So, in the CFA of an ITC system there is a pointer to DOCOL, DOVAR > > whatever/etc. In DTC it's slightly more complicated, since the CFA > > field would contain executable code. In my particular processor of > > choice, the CFA field would be two cells wide. Yes, I can see that > > patching it to change how it is interpreted could be a pain. Though I > > suppose a dedicated helper word(s) could be provided to facilitate it. > > Like Hugh says, it would be quite a rare occurence. > > > Okay. I get it. Sorry for the confusion. > > > Mark > > I think that you have got it, but I'm not sure, so I'll go over it > again: > > 1.) In ITC what you have in front of the threaded code of the colon > word (or the body of the whatever), is a single pointer to the > interpreter for that kind of word (DOCOLON for colon words, etc.). It > is easy to store a different pointer in that slot, so the word will be > interpreted differently (DOCOLON-WITH-SINGLE-STEPPING for colon words, > for example). > > 2.) In DTC what you have in front of the threaded code of the colon > word (or the body of the whatever), is the actual machine-code of the > interpreter for that kind of word (DOCOLON for colon words, etc.). It > is difficult to patch this code, because it is actual code, rather > than a pointer to some code. > > 3.) There is a kind of hybrid between ITC and DTC. This is DTC in the > sense that we have machine-code in front of each word. However, this > machine-code always consists of a single CALL instruction to DOCOLON > etc.. When a CALL is executed, it puts the address just after itself > on the processor return-stack. Normally this is for RET to use to go > back. Here is the clever part though: this is the address of the body > of the Forth word. DOCOLON can load this address into the IP and begin > interpreting. In this case, you get DTC which is faster than ITC, but > you also get an easy way to change how a word is interpreted (just > store a new pointer into the operand of the CALL instruction). > > The PDP-11 had an interesting feature. The JSR (its term for CALL) > would store the address after itself into a register, and it would > first push that register onto the return-stack. Effectively, the top > value of the return-stack was held in a register. But that register > could be your IP! You have DTC code and just do a JSR to DOCOLON (#3 > above), and DOCOLON automatically gets the address of the threaded > code loaded into the IP. I figured this out way back in 1985 when I > was taking a class in assembly-language at the city college, which was > PDP-11. This works so well, that I had to suppose that the designers > of the PDP-11 were Forth programmers, or at least, were trying to > support DTC threaded code. I've never seen this feature on any other > processor. Even in 1985 though, the PDP-11 was obsolete --- the city > college was still teaching it just because they had all the textbooks, > but the professor cheerfully admitted that the PDP-11 was obsolete and > we would never use what we learned in the real world. I've always > thought that the PDP-11 was pretty cool though --- I wish somebody > would come out with a micro-controller that runs PDP-11 code and RT11 > and all that --- maybe on an FPGA. > > BTW: There is a discussion of threading over on comp.lang.asm.x86:https://groups.google.com/group/comp.lang.asm.x86/browse_thread/threa... > > Mark: Since your TI Forth system is ITC, why don't you take a stab at > writing a single-step source-level debugger? As I mentioned, I wrote > one for my 65c02 system. It is not as difficult as you might suppose. > I did it with screen-file source-code. It can be done with seq-file > source-code though, I would suppose. I don't think that a debugger is > all that useful, but writing one is pretty interesting --- and your > users will be impressed. :-) > > Have fun! Hugh- Hide quoted text - > > - Show quoted text - Hi Hugh, Thanks for the clarification. I'm going to take a serious look at DTC because in my case, it's low hanging fruit in terms of a 'cheap' way to gain a performance boost. It's already very fast for what it is, running on a 3 mHZ 16-bit chip with a multiplexed 8 bit data bus (the fastest Forth ever produced for that machine). I figure I can effectively lose DOCOL and EXIT altogether. What I mean is, normally DOCOL and EXIT are written as subroutines that each colon definition calls, as described by you above. Well, a branch instruction on the 9900 is a 4 byte instruction; 2 bytes for the instruction, and two bytes for the address (it's only a 2 byte instruction if jumping via a register, but I digress). The code for DOCOL would be: DECT RP ; create entry on return stack (DECT=decrement by two) MOV IP,*RP ; move instruction pointer value to return stack That's also four bytes (2x 2byte instructions). So no point jumping to a subroutine; doing so slows things down! EXIT becomes MOV *RP+,IP ; pop return address into IP B *IP ; go run that code Also 4 bytes. So, again this gets inlined. So there's a speed advantage right there. In addition, the single level of indirection in DTC will increase performance (of the "interpreter") by around 50% I think. I'll sit down with a pencil and paper (my favourite method of working stuff out) and the instruction set cycle counts at the weekend and work out the saving of the overhead of DOCOL and EXIT. I think it'll be a no-brainer though. Also, primitives will not need a NEXT subroutine. Currently, in my ITC, all primitives call NEXT at the end which moves along the thread and causes the next word in the thread to be executed. Of course, the branch to NEXT is again 4 bytes. But in DTC, I think it'll be two bytes: B *IP+ \ branch the next address in the thread So each primitive executes the next word in the thread. There is no "NEXT" - it's a single instruction. Again, huge payoff. The more I think about, the more I convince myself! As I say, DTC is relatively low hanging fruit; I don't if there would be major complications with converting things like DOES> over. I'll take a look at that. The next step after that would be to have the compiler generate machine code (subroutine threaded) but, on the TMS9900, which has no stacks at all, there is I think no point in this. STC would possibly be slower than DTC code. Interesting stuff. I'm not man enough to look at things like peephole optimisation on native-code generating compilers. That's above my paygrade for now ;-) I'm very interested in it though. And I wouldn't spend the time for that on a 30 year old hobby. I'd move to an ARM project board and study it in the context of an ARM based system. But that's for another day. When I win the lottery, then I'll get onto it!
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-11-29 22:34 -0800 |
| Message-ID | <a260b4eb-0a49-4576-b541-2c74c19a7927@v9g2000pbi.googlegroups.com> |
| In reply to | #17676 |
On Nov 29, 2:26 am, Mark Wills <forthfr...@gmail.com> wrote: > On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > Mark: Since your TI Forth system is ITC, why don't you take a stab at > > writing a single-step source-level debugger? As I mentioned, I wrote > > one for my 65c02 system. It is not as difficult as you might suppose. > > I did it with screen-file source-code. It can be done with seq-file > > source-code though, I would suppose. I don't think that a debugger is > > all that useful, but writing one is pretty interesting --- and your > > users will be impressed. :-) This suggestion was a bad idea. I wasn't thinking straight when I said that. I was able to write a source-level debugger because my 65c02 Forth was a cross-compiler and it was running on an MS-DOS machine (it was written in UR/Forth). The compiler needs to generate a large data structure containing the addresses of every word compiled in every definition. When the single-stepper is running, it stops at every one of these (a BRK instruction in my system, although in an ITC system DOCOLON will stop on every word). This address is looked up and the corresponding source-code is displayed. The host computer has to have a lot of memory for that gigantic data-structure, and it has to be pretty fast. You don't want to try this on a 1980s vintage TI99/4A computer --- you don't have the memory or the speed to do this --- it taxed the limits of the 80386 computer that I was using as a host. When I wrote that suggestion, I had forgotten that you don't have a cross-compiler, but have an on-board Forth. > Hi Hugh, > > Thanks for the clarification. > > I'm going to take a serious look at DTC because in my case, it's low > hanging fruit in terms of a 'cheap' way to gain a performance boost. > It's already very fast for what it is, running on a 3 mHZ 16-bit chip > with a multiplexed 8 bit data bus (the fastest Forth ever produced for > that machine). For your old 16-bit computer, DTC should help to speed up the system. It will also make the programs slightly larger. Instead of a pointer in front of each colon word, you have a chunk of code. It is true that NEXT should be smaller, so every primitive will be slightly smaller, but this won't reduce the size of the system very much --- overall, more memory will be needed. A better way to boost the speed, is with some optimization. In many cases, there are pairs of producers and consumers. For example, LIT is a producer because it produces some data for the parameter stack, and + is a consumer because it consumes some data from the parameter stack. These pairs are inefficient because the producer pushes data onto the stack, and the consumer immediately pops that data off the stack. The solution is to combine them into a single word. For example, write a primitive LIT_+ that combines what LIT and + do. It would hold the literal value in a register rather than push it onto the stack and then pop it off again. Even with ITC, it is possible to optimize pairs like this. Make your compiler smart enough to remember what the last word it compiled was. When it is ready to compile the next word, it checks what the last word was and, if they are an optimizable pair, it compiles the combo instead. For example, if the last thing you did was LIT, when you are about to compile + your compiler will instead back up and get rid of the LIT and replace it with LIT_+. This kind of peephole-optimization not only makes your program faster, but smaller as well. You are right though, that DTC is low-hanging fruit, and much easier to implement. Peephole-optimization is somewhat more difficult, but not unreasonably difficult. You can do the peephole-optimization on a piece-meal basis. Start with + and make it smart enough to combine with all the likely producers, then do ! and +! and so forth --- you don't have to do everything at once, just doing + should boost the speed significantly, and you can go from there. BTW: I'm switching from DTC to ITC on my system. This is because I realize (from reading this thread!), that with DTC the DOCOLON code is scattered all around and won't be in the code cache, whereas with ITC the whole VM should be in the code cache. That is only an issue on big processors such as the modern x86 --- it is not an issue on the TI99/4A. Also, listening to Paul Rubin promote single-step debuggers has made me feel inclined to provide one for my system, and that is easier with ITC than DTC --- I don't really like to use a debugger, but other people do, and it isn't difficult to implement, so I might as well go ahead and provide one.
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-11-30 01:42 -0800 |
| Message-ID | <8746568b-28b4-4466-bea6-5f5a7849126b@r4g2000vbi.googlegroups.com> |
| In reply to | #17750 |
On Nov 30, 6:34 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> On Nov 29, 2:26 am, Mark Wills <forthfr...@gmail.com> wrote:
>
> > On Nov 29, 5:56 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote:
> > > Mark: Since your TI Forth system is ITC, why don't you take a stab at
> > > writing a single-step source-level debugger? As I mentioned, I wrote
> > > one for my 65c02 system. It is not as difficult as you might suppose.
> > > I did it with screen-file source-code. It can be done with seq-file
> > > source-code though, I would suppose. I don't think that a debugger is
> > > all that useful, but writing one is pretty interesting --- and your
> > > users will be impressed. :-)
>
> This suggestion was a bad idea. I wasn't thinking straight when I said
> that.
>
Oh. Well. Now you've gone and thrown the gauntlet down, haven't
you?! ;-)
> I was able to write a source-level debugger because my 65c02 Forth was
> a cross-compiler and it was running on an MS-DOS machine (it was
> written in UR/Forth). The compiler needs to generate a large data
> structure containing the addresses of every word compiled in every
> definition. When the single-stepper is running, it stops at every one
> of these (a BRK instruction in my system, although in an ITC system
> DOCOLON will stop on every word). This address is looked up and the
> corresponding source-code is displayed. The host computer has to have
> a lot of memory for that gigantic data-structure, and it has to be
> pretty fast. You don't want to try this on a 1980s vintage TI99/4A
> computer --- you don't have the memory or the speed to do this --- it
> taxed the limits of the 80386 computer that I was using as a host.
>
Well, I have already written a simple debugger. It's not a single
stepper, though. It works like this:
You load the debugger and it modifies : and ; such that each
subsequently defined colon definition makes a call into a word (can't
remember what it's called) that displays the name of the executing
word, and the data-stack.
As it goes, the depth of the return stack is measured, and it uses
this to produce indentations on the on-screen display - thus one can
see how one's program nests and un-nests.
A new word is defined, BREAK, which stops the program, returning to
the command line, and giving a full return stack dump. You can scatter
BREAKs around your code at point where you think it may be going awry.
For example:
: test swap break ;
: harry 3 test ;
: dick 2 harry ;
: tom 1 dick ;
tom
And you'd get output that looked like this:
>tom (1) 1
>dick (2) 1 2
>harry (3) 1 2 3
>test (3) 1 2 3
BREAK in test in harry in dick in tom
Without a break, if you just let the program run, you'd get:
>tom (1) 1
>dick (2) 1 2
>harry (3) 1 2 3
>test (3) 1 2 3
<test (3) 1 3 2
<harry (3) 1 3 2
<dick (3) 1 3 2
<tom (3) 1 3 2
It's fairly simple to extend the above into a single step debugger. An
on-screen display showing the definition currently being executed with
a cursor pointing to the current word is less trivial, but it
possible. Again, I have a starting point, in that I already have SEE
for my system. So I already have code to de-compile a word. So it's
possible, and doesn't require a large list/table in memory, it's just
a different technique. In fact it would be an interesting excercise!
There's just one little itsy bitsy problem: Since I wrote my TRACER
program, I've used it once. And that was to demo it to someone else.
And the only reason I was showing it was to show them that facilities
such as a tracer/debugger can be written in Forth itself and the Forth
environment simply augmented with the functionality (they were
suitably impressed). I don't think I've touched it since. I just debug
at the command line.
In fact, I rarely use SEE. I only use SEE if I'm debugging a compiling
word. The last time I used SEE was a couple of weeks ago when I was
implementing your MACRO: idea (duly implemented as a loadable
extension and working beautifully - thank you for the inspiration!).
SEE has limitations (at least on my system) because some subroutines
in my system are headerless (don't have dictionary entries) so they
display as a ? when de-compiled. It's no problem to me, since I know
what's going on. But a newbie would wonder what's going on.
Unfortunately I don't have the ROM space available to allow headers
for everything. For example, DOES> compiles a DODOES, but DODOES is
headerless. This would be a problem in a single stepping debugger,
because it would not be possible to display the names for headerless
words. This is a limitation of my system due to memory constraints. I
only have 16K. My Forth system is implemented as a plug in cartridge:
http://turboforth.net/about_turboforth.html
>
> For your old 16-bit computer, DTC should help to speed up the system.
> It will also make the programs slightly larger. Instead of a pointer
> in front of each colon word, you have a chunk of code. It is true that
> NEXT should be smaller, so every primitive will be slightly smaller,
> but this won't reduce the size of the system very much --- overall,
> more memory will be needed.
>
I had a look at this yesterday and got myself tied up in knots. I
couldn't work out how to bootstrap the thing; to get it started. How
does the 'interpreter' for a high-level definition execute the words
in the thread. I couldn't figure it out in my lunch break and had to
junk what I had done. Need more time to concentrate. I was missing
something very fundamental. I was using the : SQUARE DUP * ; as my
target but didn't get anywhere.
> A better way to boost the speed, is with some optimization. In many
> cases, there are pairs of producers and consumers. For example, LIT is
> a producer because it produces some data for the parameter stack, and
> + is a consumer because it consumes some data from the parameter
> stack. These pairs are inefficient because the producer pushes data
> onto the stack, and the consumer immediately pops that data off the
> stack. The solution is to combine them into a single word. For
> example, write a primitive LIT_+ that combines what LIT and + do. It
> would hold the literal value in a register rather than push it onto
> the stack and then pop it off again.
This is an excellent suggestion. Perhaps an easier way (at least, in
terms of performing optimisations) is to make every word in the
dictionary immediate. Then, every word can 'look ahead' and see what
is about to be compiled and intervene accordingly. It would be very
difficult to produce a standard Forth with such a system though! I
wonder if anyone has previously experimented with such a technique?
>
> Even with ITC, it is possible to optimize pairs like this. Make your
> compiler smart enough to remember what the last word it compiled was.
> When it is ready to compile the next word, it checks what the last
> word was and, if they are an optimizable pair, it compiles the combo
> instead. For example, if the last thing you did was LIT, when you are
> about to compile + your compiler will instead back up and get rid of
> the LIT and replace it with LIT_+. This kind of peephole-optimization
> not only makes your program faster, but smaller as well.
>
I'll add your peephole suggestion to my "things to look at in the next
version" list. The next version (V2.0) is the version that I tell
myself I'm *not* going to write, but I know I 99.9% probably will.
It's like a bloody drug. It's the classic symptom of wanting to start
with a clean sheet, to implement all the 'lessons learned' that you
spent blood, sweat and tears learning on the first implementation.
There are many aspects of V1.x that have been re-written a couple of
times as I learned (from other Forthers, some here on this list) or
simply discovered (as part of the Forth awakening procees) better way
to do things.
[ and to the nay-sayers: I *do* write Forth code too. Not just a
compiler. But my Forth coding is for fun. I'm still learning. I write
stuff like this:
http://turboforth.net/tutorials/darkstar.html ]
I also want to spend some time looking at Smalltalk though (a project
for 2013) - not writing a smalltalk system, just learning the
language. It's the OOP equivalent of Forth. It's beautiful (though
very slow, I believe). It looks very interesting indeed to me. Despite
being pure OO it shares the idea of terseness and brevity and total
simplicity that Forth has.
Things for 2013:
* VFX (I want to do some simple SCADA stuff using serial and IP comms)
* Smalltalk
* TurboForth V2.0 (maybe - it'll be a part-time when-feel-like-it
thing)
> You are right though, that DTC is low-hanging fruit, and much easier
> to implement. Peephole-optimization is somewhat more difficult, but
> not unreasonably difficult. You can do the peephole-optimization on a
> piece-meal basis. Start with + and make it smart enough to combine
> with all the likely producers, then do ! and +! and so forth --- you
> don't have to do everything at once, just doing + should boost the
> speed significantly, and you can go from there.
>
Yeah. You've got me thinking now! I need a lot more memory to do this
though than I currently have. Still, I plan to make the next version a
64K EPROM but I can go up to 128K in an eprom if I need to. That's 16
8K pages which is a PITA, but doable.
> BTW: I'm switching from DTC to ITC on my system. This is because I
> realize (from reading this thread!), that with DTC the DOCOLON code is
> scattered all around and won't be in the code cache, whereas with ITC
> the whole VM should be in the code cache.
Well, if your high-level definitions make a CALL to DOCOLON then
there's no reason why DOCOLON would not be in the cache. The expense
of the call might be less than the delay induced by a cache miss.
However, I'd urge you to take a step back and a deep breath. I was
reading your post on the x86 group where you are discussing caching
etc. However, I have to point out that if the Forth you are intending
to produce is primarily for embedded systems then the chances of the
embedded system running on an x86 processor are quite low. It's much
more likely to be an ARM variant. In other words, don't allow key
descisions about the architecture of your system to be guided by
relatively un-important architectural constraints of a particular
processor family.
I'd urge you to get out a notebook and pencil. Sit down somewhere
quiet and write a list of key things that you want the system to do.
Design goals. Then put the list away. Reflect on it for a couple of
days and go back and make changes. Iterate. Eventually your thoughts/
ideas/requirements will coalesce. There's your plan/design goals. When
you've got it done, pin it up on the wall above your computer. Let it
be your guide as you develop the *project*, and when you feel a knee-
jerk step-change coming on, consult the plan again! Don't be swayed.
Stick to the plan. Have faith in the design decisions you made
earlier, even if you've had a bright idea.
I failed to make a plan/design and ended up with many many more
iterations/builds/bugs/teeth-knashing/wailing than I should have. It's
okay for me, because it's a hobby system and is given away for free.
Your aspirations are somewhat higher though, wanting a good system for
embedded targets. So I'd urge due consideration and diligence!
Just my two cents, FWIW!
Mark
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-11-30 13:18 -0800 |
| Message-ID | <379d76ca-9bee-48b2-a6d9-da39d74594e5@i2g2000pbi.googlegroups.com> |
| In reply to | #17756 |
On Nov 30, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote: > On Nov 30, 6:34 am, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > A new word is defined, BREAK, which stops the program, returning to > the command line, and giving a full return stack dump. You can scatter > BREAKs around your code at point where you think it may be going awry. That is a good technique. I call it QI because it consists of QUERY INTERPRET. > There's just one little itsy bitsy problem: Since I wrote my TRACER > program, I've used it once. And that was to demo it to someone else. > And the only reason I was showing it was to show them that facilities > such as a tracer/debugger can be written in Forth itself and the Forth > environment simply augmented with the functionality (they were > suitably impressed). I don't think I've touched it since. I just debug > at the command line. I agree --- it is easiest to just debug at the command line. The only Forth single-step source-level debugger that I've ever used was the one that I wrote myself for my 65c02 cross-compiler --- that was a long time ago! > I had a look at this yesterday and got myself tied up in knots. I > couldn't work out how to bootstrap the thing; to get it started. How > does the 'interpreter' for a high-level definition execute the words > in the thread. I couldn't figure it out in my lunch break and had to > junk what I had done. Need more time to concentrate. I was missing > something very fundamental. I was using the : SQUARE DUP * ; as my > target but didn't get anywhere. ITC is more complicated than DTC because there is an extra level of indirection. If you can't figure out how DTC works, but you've got ITC already working, then I'm guessing that you ported your ITC over from somewhere without really understanding it. Keep at it --- if you still can't figure it out, contact me by email and I will show you some code in x86 or whatever assembly-language you want (not TI9900 though, as I don't know that one). > > A better way to boost the speed, is with some optimization. In many > > cases, there are pairs of producers and consumers. For example, LIT is > > a producer because it produces some data for the parameter stack, and > > + is a consumer because it consumes some data from the parameter > > stack. These pairs are inefficient because the producer pushes data > > onto the stack, and the consumer immediately pops that data off the > > stack. The solution is to combine them into a single word. For > > example, write a primitive LIT_+ that combines what LIT and + do. It > > would hold the literal value in a register rather than push it onto > > the stack and then pop it off again. > > This is an excellent suggestion. Perhaps an easier way (at least, in > terms of performing optimisations) is to make every word in the > dictionary immediate. Then, every word can 'look ahead' and see what > is about to be compiled and intervene accordingly. It would be very > difficult to produce a standard Forth with such a system though! I > wonder if anyone has previously experimented with such a technique? I don't recommend doing that. The way that I'm doing it in my own system, is to smarten up what COMPILE, does. This doesn't just compile the xt that it is given, but instead it puts the xt into a queue. When it does this, it looks to see what xt is already in the queue, and combines them if possible. On my system, the queue is only a single item in length, so it is actually a variable not a queue --- I call it LIMBO --- because the xt in there is in limbo, in the sense that it has been compiled, but hasn't yet really been compiled. > I also want to spend some time looking at Smalltalk though (a project > for 2013) - not writing a smalltalk system, just learning the > language. It's the OOP equivalent of Forth. It's beautiful (though > very slow, I believe). It looks very interesting indeed to me. Despite > being pure OO it shares the idea of terseness and brevity and total > simplicity that Forth has. Smalltalk is where the idea of dynamic-OOP originated. For the most part though, CLOS really is where dynamic-OOP got going (there are dynamic-OOP systems for Scheme too). I would recommend learning Scheme or Lisp instead of Smalltalk, as these still have an active community, which I don't think Smalltalk does. I'm learning Scheme --- if you learn it too, we could bounce ideas back and forth by email. Learning Lisp has always been on my bucket list, but now I'm finally doing it. :-) > > BTW: I'm switching from DTC to ITC on my system. This is because I > > realize (from reading this thread!), that with DTC the DOCOLON code is > > scattered all around and won't be in the code cache, whereas with ITC > > the whole VM should be in the code cache. > > Well, if your high-level definitions make a CALL to DOCOLON then > there's no reason why DOCOLON would not be in the cache. The expense > of the call might be less than the delay induced by a cache miss. Yes, but the CALL isn't in the code cache with DTC. With ITC however, you don't have code that calls your interpreter, but each word has a pointer in the front (at the cfa) that points to the interpreter. > However, I'd urge you to take a step back and a deep breath. I was > reading your post on the x86 group where you are discussing caching > etc. However, I have to point out that if the Forth you are intending > to produce is primarily for embedded systems then the chances of the > embedded system running on an x86 processor are quite low. It's much > more likely to be an ARM variant. In other words, don't allow key > descisions about the architecture of your system to be guided by > relatively un-important architectural constraints of a particular > processor family. I'm writing two Forth systems. HostForth runs on the host computer (the x86). TargForth is written in HostForth and it generates the micro-controller code. I'm must working on HostForth right now. What I was primarily trying to figure out on comp.lang.asm.x86 is how to support overlays. I think I've got that figured out now though. That is pretty important! It involves a fundamental design feature. If I hadn't thought about this early on, but had left for later, I would have been in trouble when later on when I discovered that supporting overlays would be impossible without a complete rewrite. Fundamental stuff like this really has to be figured out as early as possible. I'm always able to learn something from those discussions on clax --- most of those guys are really knowledgeable! --- unlike clf, where it is mostly just baloney with b.s. frosting.
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-12-01 01:42 -0800 |
| Message-ID | <48bb9ca6-351c-471b-81be-9b69733bb33f@bx4g2000vbb.googlegroups.com> |
| In reply to | #17783 |
On Nov 30, 9:18 pm, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > > I had a look at this yesterday and got myself tied up in knots. I > > couldn't work out how to bootstrap the thing; to get it started. How > > does the 'interpreter' for a high-level definition execute the words > > in the thread. I couldn't figure it out in my lunch break and had to > > junk what I had done. Need more time to concentrate. I was missing > > something very fundamental. I was using the : SQUARE DUP * ; as my > > target but didn't get anywhere. > > ITC is more complicated than DTC because there is an extra level of > indirection. If you can't figure out how DTC works, but you've got ITC > already working, then I'm guessing that you ported your ITC over from > somewhere without really understanding it. > > Keep at it --- if you still can't figure it out, contact me by email > and I will show you some code in x86 or whatever assembly-language you > want (not TI9900 though, as I don't know that one). > Thanks. I got it working. It was much simpler than I thought. It seems it doesn't make much difference on the 9900; it saves a single MOV assembly instruction in NEXT (one less level of indirection). So, NEXT is two assembly instructions instead of three. That takes the same space as a TMS9900 BRANCH instruction, so, at the end of a primitive, rather than branching to NEXT, it can just be in-lined. However, the Kernal runs in 8-bit memory, and, doing the math, it looks like its faster to put next in 16-bit (0 wait state) memory, and have primitives branch to NEXT. It all evens out to pretty much the same. For sure, *not* worth coding a new system in DTC on the 9900. It's so marginal that it's just not worth it. Maybe I'll take a look at a system that generates native machine code. That opens up all sorts of possibilities for optimisation, such as in- lining. Words can have a bit reserved in their dictionary entry that determines is a word is to be in-lined or not. If not, a branch/call is compiled, otherwise, the code is pasted into the current definition. It would be a nice learning exercise to learn about native code compilers. Maybe I could incrementally add optimisations as I learn the techniques. The peephole that you mentioned (I got a list of about 25-30 words that could be optimised using the combining literal technique that you described). I was also reading about constant folding on wikipedia. It's extremely clever, though I can't currently see how it is actually implemented. Any optimising compiler that I wrote for the 4A would have to be simple, low-hanging fruit optimisations. It's not worth the effort to write a compiler that compiles to intermediate language that is then optimised. Simple optimisations would be the order of the day. How far into HostForth are you?
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-12-03 15:25 -0800 |
| Message-ID | <9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com> |
| In reply to | #17788 |
On Dec 1, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote: > Thanks. I got it working. It was much simpler than I thought. > > It seems it doesn't make much difference on the 9900; it saves a > single MOV assembly instruction in NEXT (one less level of > indirection). > > So, NEXT is two assembly instructions instead of three. That takes the > same space as a TMS9900 BRANCH instruction, so, at the end of a > primitive, rather than branching to NEXT, it can just be in-lined. > However, the Kernal runs in 8-bit memory, and, doing the math, it > looks like its faster to put next in 16-bit (0 wait state) memory, and > have primitives branch to NEXT. > > It all evens out to pretty much the same. For sure, *not* worth coding > a new system in DTC on the 9900. It's so marginal that it's just not > worth it. I'm glad you figured out DTC. Most things are simple after you figure them out! :-) You are right that there is often not much difference between ITC and DTC in speed. There was more difference on primitive processors that lacked addressing-modes and/or didn't have enough registers. That was a different world --- nowadays, it is all about cache efficiency. > Maybe I'll take a look at a system that generates native machine code. > That opens up all sorts of possibilities for optimisation, such as in- > lining. Words can have a bit reserved in their dictionary entry that > determines is a word is to be in-lined or not. If not, a branch/call > is compiled, otherwise, the code is pasted into the current > definition. Inlining isn't all that good of a technique on modern processors. Your code becomes bloated, which causes it to thrash the cache (Hey! That rhymed! I'm a poet!). Also, on the modern x86, CALL and RET are very efficient. Doing a CALL to a function, and it doing a RET back again, is almost as fast as inlining that function --- and it saves a lot of memory. SwiftForth inlines small functions, but it just pastes them in. You see code that pushes a datum onto the stack from a particular register, and then immediately pops the datum back into that same register. That isn't optimization! You are better off to just leave those functions as functions, so you reduce your bloat. > It would be a nice learning exercise to learn about native code > compilers. Maybe I could incrementally add optimisations as I learn > the techniques. The peephole that you mentioned (I got a list of about > 25-30 words that could be optimised using the combining literal > technique that you described). I was also reading about constant > folding on wikipedia. It's extremely clever, though I can't currently > see how it is actually implemented. > > Any optimising compiler that I wrote for the 4A would have to be > simple, low-hanging fruit optimisations. It's not worth the effort to > write a compiler that compiles to intermediate language that is then > optimised. Simple optimisations would be the order of the day. Well, you could stick with ITC and make your peephole-optimizer combine word pairs. For example, OVER + would get compiled as a single function: OVER_+ . This works quite well. Also, it is not processor dependent. If you get this to work on your TI9900, you can later port it over directly to your ARM Forth. You will have to write all of the functions, such as OVER and + and OVER_+ in ARM assembly, but this is easy. The complicated part, of recognizing the word pairs and combining them, will be exactly the same no matter what processor is underneath the hood. I wouldn't recommend writing an optimizer for generating machine-code. That requires a lot of knowledge of the processor under the hood. Very little that you do on the TI9900 would port over to the ARM, as they are quite different. Stick with ITC though, and you are largely processor independent. Are you planning on jumping to the ARM or to the MSP430 in the future? You know, you have to abandon that TI99/4A someday! What if you drop it and it breaks? You can't go to WalMart and buy another one... > How far into HostForth are you? Not too far. I had HLA code for generating optimized machine-code, but that was getting complicated and I got bogged down. Then I decided to rewrite in traditional assembly language and generate ITC code. I will only generate optimized machine-code for the micro-controllers, where it matters. I liked HLA, but it is limited to 32-bit x86, and I want to be cutting- edge for once in my life --- so I'm going with 64-bit x86 instead. I'm taking it slow on Straight Forth because I have a lot to learn. I don't know much about low-level stuff, such as caches. I also don't know much about high-level stuff, such as closures. There seems to be only a narrow window of mid-level stuff that I know about. lol It is always worthwhile to learn new ideas, and to better oneself!
[toc] | [prev] | [next] | [standalone]
| From | Alex McDonald <blog@rivadpm.com> |
|---|---|
| Date | 2012-12-03 16:18 -0800 |
| Message-ID | <50c1958c-3c4a-43aa-8520-a5f3e020bc33@n8g2000vbb.googlegroups.com> |
| In reply to | #17840 |
On Dec 3, 11:25 pm, Hugh Aguilar <hughaguila...@yahoo.com> wrote: > On Dec 1, 2:42 am, Mark Wills <forthfr...@gmail.com> wrote: > > Maybe I'll take a look at a system that generates native machine code. > > That opens up all sorts of possibilities for optimisation, such as in- > > lining. Words can have a bit reserved in their dictionary entry that > > determines is a word is to be in-lined or not. If not, a branch/call > > is compiled, otherwise, the code is pasted into the current > > definition. > > Inlining isn't all that good of a technique on modern processors. Your > code becomes bloated, which causes it to thrash the cache (Hey! That > rhymed! I'm a poet!). > > Also, on the modern x86, CALL and RET are very efficient. Doing a CALL > to a function, and it doing a RET back again, is almost as fast as > inlining that function --- and it saves a lot of memory. No it doesn't. > > SwiftForth inlines small functions, but it just pastes them in. You > see code that pushes a datum onto the stack from a particular > register, and then immediately pops the datum back into that same > register. That isn't optimization! You are better off to just leave > those functions as functions, so you reduce your bloat. > STC Experimental 32bit: 0.06.05 Build: 363 With simple inlining of words <10 bytes long Test time including overhead ms times ns (each) Eratosthenes sieve 1899 Primes 120 8190000 14 Fibonacci recursion ( 35 -> 9227465 ) 65 9227430 7 Hoare's quick sort (reverse order) 129 2000000 64 Generate random numbers (1024 kb array) 92 262144 350 LZ77 Comp. (400 kb Random Data Mem>Mem) 134 1 Dhrystone (integer) 81 500000 162 6172839 Dh rystones/sec Total: 647 1 APP mem: 113,865, CODE mem: 15,273, SYS mem: 5,488 Total: 134,626 Without inlining Test time including overhead ms times ns (each) Eratosthenes sieve 1899 Primes 126 8190000 15 Fibonacci recursion ( 35 -> 9227465 ) 66 9227430 7 Hoare's quick sort (reverse order) 252 2000000 126 Generate random numbers (1024 kb array) 104 262144 396 LZ77 Comp. (400 kb Random Data Mem>Mem) 155 1 Dhrystone (integer) 88 500000 176 5681818 Dh rystones/sec Total: 800 1 APP mem: 113,865, CODE mem: 14,799, SYS mem: 5,488 Total: 134,152 A overall speed decrease of 25% for 470 bytes extra code.
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-12-03 21:20 -0500 |
| Message-ID | <k9jmcb$ilv$1@speranza.aioe.org> |
| In reply to | #17840 |
"Hugh Aguilar" <hughaguilar96@yahoo.com> wrote in message news:9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com... ... > Also, on the modern x86, CALL and RET are very efficient. Doing > a CALL to a function, and it doing a RET back again, is almost > as fast as inlining that function --- [...] That may be almost true now. But, I don't know how true it is that they're fast now. I haven't checked the x86 manuals in three to five years. Even if they're fast now, CALL and RET still have some overhead that's not present with inlining. Historically, CALL and RET being fast on x86 wasn't true. The problem with RET on modern x86 - according to those on c.l.a.x. - is that it must be _matched_ with a CALL or it causes a slowdown of the processor. E.g., the RET location was pushed onto the stack via PUSH instead of by a CALL. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | "Elizabeth D. Rather" <erather@forth.com> |
|---|---|
| Date | 2012-12-03 17:29 -1000 |
| Message-ID | <0ZKdnRs5KJ-48yDNnZ2dnUVZ_rCdnZ2d@supernews.com> |
| In reply to | #17845 |
On 12/3/12 4:20 PM, Rod Pemberton wrote: ... >> Also, on the modern x86, CALL and RET are very efficient. Doing >> a CALL to a function, and it doing a RET back again, is almost >> as fast as inlining that function --- [...] > > That may be almost true now. But, I don't know how true it is > that they're fast now. I haven't checked the x86 manuals in three > to five years. Even if they're fast now, CALL and RET still have > some overhead that's not present with inlining. Historically, > CALL and RET being fast on x86 wasn't true. Whether it's worth inlining or not really depends on the length of the code being inlined. If it's just a few instructions, the ratio of the CALL/RET to the code is such that the inlining pays off. For a longer sequence, it does not. It also depends on whether the CALL is set up by C, which adds overhead for calling sequences that is missing in Forth written in Forth/assembler. Cheers, Elizabeth -- ================================================== Elizabeth D. Rather (US & Canada) 800-55-FORTH FORTH Inc. +1 310.999.6784 5959 West Century Blvd. Suite 700 Los Angeles, CA 90045 http://www.forth.com "Forth-based products and Services for real-time applications since 1973." ==================================================
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-12-04 00:12 -0800 |
| Message-ID | <9e119760-3e3a-4bcb-a7b2-f540695bdbe5@n8g2000vbb.googlegroups.com> |
| In reply to | #17845 |
On Dec 4, 2:20 am, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> wrote: > "Hugh Aguilar" <hughaguila...@yahoo.com> wrote in message > > news:9e679e24-97e8-4efb-868f-a5d0aae6fda2@r10g2000pbd.googlegroups.com... > ... > > > Also, on the modern x86, CALL and RET are very efficient. Doing > > a CALL to a function, and it doing a RET back again, is almost > > as fast as inlining that function --- [...] > > That may be almost true now. But, I don't know how true it is > that they're fast now. I haven't checked the x86 manuals in three > to five years. Even if they're fast now, CALL and RET still have > some overhead that's not present with inlining. Historically, > CALL and RET being fast on x86 wasn't true. > > The problem with RET on modern x86 - according to those on > c.l.a.x. - is that it must be _matched_ with a CALL or it causes a > slowdown of the processor. E.g., the RET location was pushed > onto the stack via PUSH instead of by a CALL. > > Rod Pemberton I would imagine (I have no experience) that with CALL/RET you also run the risk of the routine that you are CALLing not being in the cache, which adds a further performance penalty. At least if the the code is inlined (even with inefficiencies such as pushing to the data stack and immediately popping again) there is a good chance it's running from cache. My knowledge of cache's is very 1990's though; maybe they are a lot cleverer these days. What is the difference between a level 1 and a level 2 cache? Is the level 2 cache a cache for the level 1 cache? So there are two caches between the CPU and external memory? Is that how it works? Do caches run 'metrics' on subroutines like "Hmmm... This subroutine here seems to be called a lot more often than these others. I'll keep it in my cache where I can access it quickly" or are they simply dumb, where, if a section of memory is called for, and it's not in the cache, it reads the memory, plus n bytes into the cache? Guess I should read up on caches! I don't have any such complications in my hobby system! It's nice and simple. The only potential complication is the TMS9995 because it has instruction prefetch. This means some self-modifying code can trip you up, but only if you are modifying the instruction immediately in front. A simple NOP between the instructions fixes that.
[toc] | [prev] | [next] | [standalone]
Page 1 of 6 [1] 2 3 4 5 6 Next page →
Back to top | Article view | comp.lang.forth
csiph-web