Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #16579 > unrolled thread
| Started by | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| First post | 2012-10-22 08:53 -0700 |
| Last post | 2012-10-23 10:52 +0000 |
| Articles | 20 on this page of 172 — 22 participants |
Back to article view | Back to comp.lang.forth
RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000
Page 8 of 9 — ← Prev page 1 2 3 4 5 6 7 [8] 9 Next page →
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-25 00:50 -0400 |
| Message-ID | <k6ag6q$m45$1@speranza.aioe.org> |
| In reply to | #16666 |
"rickman" <gnuarm@gmail.com> wrote in message news:k69gf5$hic$1@dont-email.me... > On 10/24/2012 8:47 AM, Rod Pemberton wrote: ... > > For temporary reasons, most of my Forth stack operators, like DUP SWAP > > etc, aren't currently low-level Forth "primitives". They're implemented > > in high-level Forth using an even lower set of actual stack > > "primitives", which I'll call sub-operators for this thread. These > > sub-operators could be considered to be "stack assembly instructions" > > for my Forth. Currently, only>R and R> are actually "primitives" > > coded in C. Eventually, DUP DROP SWAP OVER will be actual > > low-level Forth "primitives", as they once were. > > Coded in 'C'? Perhaps I missed something. I was talking about a Forth > like CPU. Are you referring to a Forth compiler on a PC? > Yes. You're talking about a Forth like CPU. I'm talking about my Forth interpreter in C. I'm attempting to discuss similarities which are useful. > > So, a word like OVER can be coded in many ways. I have 22 different > > definitions just for OVER in a list, and one can construct many more. > > E.g., > > > > : OVER>R DUP R> SWAP ; > > : OVER 1 PICK ; > > : OVER SWAP TUCK ; > > : OVER SWAP DUP -ROT ; > > : OVER NUP SWAP ; > > etc. > > Yes it can, but why? Over is a very simple construct for either a Forth > like CPU or a Forth compiler. If you implement it as a single instruction, that's great. But, you're not going to fit all Forth words into your instruction set. So, some words are going to be made up of multiple Forth words comprised of multiple instructions. This was just an example of how I went about optimizing such a sequence. Also, some Forth stack words don't fit most CPU's well. For example, some words shift a portion of the data on the stack, e.g., ROT TUCK. Stack shifts are not well suited to most processors. Stacks can push or pop or use other instructions to direct reads and writes, but shifts require a copy routine. > > Which definition you choose affects how fast OVER executes and depends > > on what operations you have available and on how fast each of those > > operations are. > > > > E.g., if>R>R DUP SWAP and -ROT are all very fast machine instructions > > for your processor, then ">R DUP R> SWAP" and "SWAP DUP -ROT" > > should be fast sequences. And, one sequence will be faster than the > > other, depending on how fast each instruction is. But, if -ROT is not a > > machine instruction, e.g., perhaps coded as "ROT ROT" or "SWAP>R > > SWAP R>", or -ROT is a very slow machine instruction, then "SWAP > > DUP -ROT" is more expensive than ">R DUP R> SWAP". > > I don't follow that a four instruction sequence should be "fast" in any > real sense. I don't follow at all why you would have -ROT as a > primitive and not OVER. > You're going to be implementing instructions for your processor. Which instructions you implement and which you don't implement affects how each Forth word is coded. That affects their speed. > > The instructions also need to be balanced acrossed many Forth words. > > Originally, I had the 2xxx series words defined in terms of the simpler > > stack operators, like SWAP DUP etc. When I converted to sub-operators, > > the number of operations per definition dropped dramatically for some > > of the 2xxx definitions. Since the sub-operators are primitives, some > > of the 2xxx definitions became faster. However, after converting all of > > the simpler stack operators to sub-operators, a few of the non-primitive > > simple stack operators ended up with more operations per definition. > > So some words became much faster, while others became slightly slower. > > And, the former primitives became real slow, but they'll be converted > > back eventually. > > Not sure why you are looking at this. If you want to optimize the > implementation, why not look at what is used most often and start by > optimizing those operators? That is why I went to Koopman's book. Not > much return on optimizing things that aren't used so much. There is no quarantee that the instruction frequencies will be the most used in the code that will be run. It merely demonstrates which was the most used in the code that was sampled. I.e., let's say I used Koopman's dynamic instruction frequencies to design a processor. TUCK isn't in the list. So, I didn't put any effort into optimizing TUCK. TUCK is really slow and long sequence for my imaginary processor. Later, I fall in "love" with TUCK. So, I begin to use it everywhere in my imaginary Forth code. Will my code be fast? No. The frequencies are only useful on the sampled code and similar code. If you're not using a code optimizer, then you (or your future users) are going to have to selectively preference the optimized instructions over the unoptimized. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-25 18:58 -0400 |
| Message-ID | <k6cg75$o2d$1@dont-email.me> |
| In reply to | #16684 |
On 10/25/2012 12:50 AM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k69gf5$hic$1@dont-email.me... >> On 10/24/2012 8:47 AM, Rod Pemberton wrote: > ... > >>> For temporary reasons, most of my Forth stack operators, like DUP SWAP >>> etc, aren't currently low-level Forth "primitives". They're implemented >>> in high-level Forth using an even lower set of actual stack >>> "primitives", which I'll call sub-operators for this thread. These >>> sub-operators could be considered to be "stack assembly instructions" >>> for my Forth. Currently, only>R and R> are actually "primitives" >>> coded in C. Eventually, DUP DROP SWAP OVER will be actual >>> low-level Forth "primitives", as they once were. >> >> Coded in 'C'? Perhaps I missed something. I was talking about a Forth >> like CPU. Are you referring to a Forth compiler on a PC? >> > > Yes. You're talking about a Forth like CPU. I'm talking about my Forth > interpreter in C. I'm attempting to discuss similarities which are useful. Ok, now that I get the context, what is the point exactly? Why do you only code the lesser number of primitives when you will eventually code more words as primitives. I'm not following what point you are trying to make. Is this an exercise where you hope to learn how best to optimize primitives? >>> So, a word like OVER can be coded in many ways. I have 22 different >>> definitions just for OVER in a list, and one can construct many more. >>> E.g., >>> >>> : OVER>R DUP R> SWAP ; >>> : OVER 1 PICK ; >>> : OVER SWAP TUCK ; >>> : OVER SWAP DUP -ROT ; >>> : OVER NUP SWAP ; >>> etc. >> >> Yes it can, but why? Over is a very simple construct for either a Forth >> like CPU or a Forth compiler. > > If you implement it as a single instruction, that's great. But, you're not > going to fit all Forth words into your instruction set. So, some words are > going to be made up of multiple Forth words comprised of multiple > instructions. This was just an example of how I went about optimizing such > a sequence. I'm still not following your point. OVER is a very logical word to implement as a primitive. It is used often enough and requires little or nothing in the way of extra hardware for a CPU design and should be a single instruction on most existing CPUs. > Also, some Forth stack words don't fit most CPU's well. For example, some > words shift a portion of the data on the stack, e.g., ROT TUCK. Stack > shifts are not well suited to most processors. Stacks can push or pop or > use other instructions to direct reads and writes, but shifts require a copy > routine. Absolutely. Many Forth words won't map well to either hardware or to any given software implementation. >>> Which definition you choose affects how fast OVER executes and depends >>> on what operations you have available and on how fast each of those >>> operations are. >>> >>> E.g., if>R>R DUP SWAP and -ROT are all very fast machine instructions >>> for your processor, then ">R DUP R> SWAP" and "SWAP DUP -ROT" >>> should be fast sequences. And, one sequence will be faster than the >>> other, depending on how fast each instruction is. But, if -ROT is not a >>> machine instruction, e.g., perhaps coded as "ROT ROT" or "SWAP>R >>> SWAP R>", or -ROT is a very slow machine instruction, then "SWAP >>> DUP -ROT" is more expensive than ">R DUP R> SWAP". >> >> I don't follow that a four instruction sequence should be "fast" in any >> real sense. I don't follow at all why you would have -ROT as a >> primitive and not OVER. >> > > You're going to be implementing instructions for your processor. Which > instructions you implement and which you don't implement affects how each > Forth word is coded. That affects their speed. Ok. For instructions I picked functions that were - 1) Basic enough to be implemented in one clock cycle 2) Were good primitives in a Forth sense 3) Were high on the list of use frequency according to Koopman's book not necessarily in that priority. I did not have a one to one mapping of instructions to existing Forth words, but many instructions were existing Forth words or very similar. Some exceptions were adding the top elements of the return stack and adding the top of data stack to the return stack. These are used to control looping, although they may change if I return to working on this design. >>> The instructions also need to be balanced acrossed many Forth words. >>> Originally, I had the 2xxx series words defined in terms of the simpler >>> stack operators, like SWAP DUP etc. When I converted to sub-operators, >>> the number of operations per definition dropped dramatically for some >>> of the 2xxx definitions. Since the sub-operators are primitives, some >>> of the 2xxx definitions became faster. However, after converting all of >>> the simpler stack operators to sub-operators, a few of the non-primitive >>> simple stack operators ended up with more operations per definition. >>> So some words became much faster, while others became slightly slower. >>> And, the former primitives became real slow, but they'll be converted >>> back eventually. >> >> Not sure why you are looking at this. If you want to optimize the >> implementation, why not look at what is used most often and start by >> optimizing those operators? That is why I went to Koopman's book. Not >> much return on optimizing things that aren't used so much. > > There is no quarantee that the instruction frequencies will be the most used > in the code that will be run. It merely demonstrates which was the most > used in the code that was sampled. Then you have no basis for picking any words for optimization. No matter what words you optimize, if the application doesn't use those instructions very much the optimization won't be effective for that app. That is why I rely on data like Koopman's book, it draws on a range of apps. > I.e., let's say I used Koopman's dynamic instruction frequencies to design a > processor. TUCK isn't in the list. So, I didn't put any effort into > optimizing TUCK. TUCK is really slow and long sequence for my imaginary > processor. Later, I fall in "love" with TUCK. So, I begin to use it > everywhere in my imaginary Forth code. Will my code be fast? No. > > The frequencies are only useful on the sampled code and similar code. If > you're not using a code optimizer, then you (or your future users) are going > to have to selectively preference the optimized instructions over the > unoptimized. Ok, I am all ears. If you have a better idea for how to pick primitive instructions to implement on a dual stack CPU design, I would love to hear about it. The thing above about balancing across multiple words has the same flaw of only applying to some applications as any other method of selection. Rick
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-25 16:07 -0700 |
| Message-ID | <7xy5iucfqq.fsf@ruckus.brouhaha.com> |
| In reply to | #16714 |
rickman <gnuarm@gmail.com> writes: > Ok, I am all ears. If you have a better idea for how to pick > primitive instructions to implement on a dual stack CPU design, I > would love to hear about it. In the J1, I wonder what it would have cost (in CLB's or speed) to add some ways for user code to access the stack slots in array- or register-like fashion. The stacks are implemented as arrays in the Verilog code already. That could save some typical stack juggling and make life easier for compilers. Maybe they could be mapped into memory rather than finding opcode space to do stuff with them.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-25 19:20 -0400 |
| Message-ID | <k6chg9$4fj$1@dont-email.me> |
| In reply to | #16715 |
On 10/25/2012 7:07 PM, Paul Rubin wrote: > rickman<gnuarm@gmail.com> writes: >> Ok, I am all ears. If you have a better idea for how to pick >> primitive instructions to implement on a dual stack CPU design, I >> would love to hear about it. > > In the J1, I wonder what it would have cost (in CLB's or speed) to add > some ways for user code to access the stack slots in array- or > register-like fashion. The stacks are implemented as arrays in the > Verilog code already. That could save some typical stack juggling and > make life easier for compilers. Maybe they could be mapped into memory > rather than finding opcode space to do stuff with them. Without looking at the code for J1 (something that is on my list of things to do) I would say it is not a huge effort. It requires a source of address and a mux to bring that address into the stack RAM along with the control hardware of course. This is not a lot of hardware, but where does the address come from? I don't understand your last sentence at all. BTW, how it is coded in Verilog has nothing to do with it. This isn't software. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-26 20:57 -0400 |
| Message-ID | <k6fbal$dqf$1@speranza.aioe.org> |
| In reply to | #16714 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6cg75$o2d$1@dont-email.me... > On 10/25/2012 12:50 AM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > > news:k69gf5$hic$1@dont-email.me... > >> On 10/24/2012 8:47 AM, Rod Pemberton wrote: ... > >>> For temporary reasons, most of my Forth stack operators, like DUP SWAP > >>> etc, aren't currently low-level Forth "primitives". They're > >>> implemented in high-level Forth using an even lower set of actual > >>> stack "primitives", which I'll call sub-operators for this thread. > >>> These sub-operators could be considered to be "stack assembly > >>> instructions" for my Forth. Currently, only>R and R> are actually > >>> "primitives" coded in C. Eventually, DUP DROP SWAP OVER will > >>> be actual low-level Forth "primitives", as they once were. > >> > >> Coded in 'C'? Perhaps I missed something. I was talking about a Forth > >> like CPU. Are you referring to a Forth compiler on a PC? > >> > > > > Yes. You're talking about a Forth like CPU. I'm talking about my Forth > > interpreter in C. I'm attempting to discuss similarities which are > > useful. > > Ok, now that I get the context, what is the point exactly? Why do you > only code the lesser number of primitives when you will eventually code > more words as primitives. I'm in the process of converting much of my code from precompiled ITC Forth in C (similar to a DTC Forth in assembly) to Forth as ASCII text. The goal was to 1) make it as independent of C as possible and/or reveal hidden C dependencies, and 2) have only primitives, variables, and other low-level words in C. The reality is slightly different. What's left are the primitives, variables, other low-level words in C, and a few critical high-level Forth words in C, as well a program (re)written entirely in terms of primitives to implement the outer interpreter. > I'm not following what point you are trying > to make. Whatever operations you chose to implement for your processor affects how quickly various Forth operations will be. > Is this an exercise where you hope to learn how best to > optimize primitives? No. It's mostly a result of the transformation process. Minimizing primitives allowed additional code to be moved out of the "C space" to the pure "Forth space". > > If you implement it as a single instruction, that's great. But, you're > > not going to fit all Forth words into your instruction set. So, some > > words are going to be made up of multiple Forth words comprised > > of multiple instructions. This was just an example of how I went > > about optimizing such a sequence. > > I'm still not following your point. OVER is a very logical word to > implement as a primitive. It is used often enough and requires little > or nothing in the way of extra hardware for a CPU design and should be a > single instruction on most existing CPUs. Pick a Forth word that you're not going to implement as an instruction. Implement it using multiple instructions. Count how many. Now, if you change your set of instructions, is it implemented in fewer or more instructions? You want a set of instructions which results in fewer, especially for the Forth words you think will be used the most. > > You're going to be implementing instructions for your processor. Which > > instructions you implement and which you don't implement affects how > > each Forth word is coded. That affects their speed. > > Ok. For instructions I picked functions that were - > > 1) Basic enough to be implemented in one clock cycle > 2) Were good primitives in a Forth sense > 3) Were high on the list of use frequency according to Koopman's book > > not necessarily in that priority. > > I did not have a one to one mapping of instructions to existing Forth > words, but many instructions were existing Forth words or very similar. > Some exceptions were adding the top elements of the return stack and > adding the top of data stack to the return stack. These are used to > control looping, although they may change if I return to working on this > design. Ok. > >> [...] If you want to optimize the > >> implementation, why not look at what is used most often and start by > >> optimizing those operators? That is why I went to Koopman's book. Not > >> much return on optimizing things that aren't used so much. > > > > There is no quarantee that the instruction frequencies will be the most > > used in the code that will be run. It merely demonstrates which was > > the most used in the code that was sampled. > > Then you have no basis for picking any words for optimization. No > matter what words you optimize, if the application doesn't use those > instructions very much the optimization won't be effective for that app. True. > That is why I rely on data like Koopman's book, it draws on a range of > apps. Yes, it's representative of the range of apps that were sampled. As long as the code to be executed uses instructions similarly, you'll be fine. > [...] > The thing above about balancing across multiple words > has the same flaw of only applying to some applications as any other > method of selection. If you attempt to balance all instructions so each takes the same amount of time, that results in a RISC like design. Yes? If you implement some instructions for speed and disregard optimization for others, that results in a CISC like design. Yes? Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-27 15:43 -0400 |
| Message-ID | <k6hdga$4je$1@dont-email.me> |
| In reply to | #16757 |
On 10/26/2012 8:57 PM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6cg75$o2d$1@dont-email.me... >> >> Ok, now that I get the context, what is the point exactly? Why do you >> only code the lesser number of primitives when you will eventually code >> more words as primitives. > > I'm in the process of converting much of my code from precompiled ITC Forth > in C (similar to a DTC Forth in assembly) to Forth as ASCII text. The goal > was to 1) make it as independent of C as possible and/or reveal hidden C > dependencies, and 2) have only primitives, variables, and other low-level > words in C. The reality is slightly different. What's left are the > primitives, variables, other low-level words in C, and a few critical > high-level Forth words in C, as well a program (re)written entirely in terms > of primitives to implement the outer interpreter. Maybe you can explain the purpose of wanting Forth in ASCII text? Wouldn't that be the source? Or are you talking about F-code in text? What would be the purpose of doing this? >> I'm not following what point you are trying >> to make. > > Whatever operations you chose to implement for your processor affects how > quickly various Forth operations will be. > >> Is this an exercise where you hope to learn how best to >> optimize primitives? > > No. It's mostly a result of the transformation process. Minimizing > primitives allowed additional code to be moved out of the "C space" to the > pure "Forth space". "Allowed"??? By definition that is minimizing primitives. But why? Typically more primitives are better in most metrics other than possibly size. The main reason is that you need to draw a line somewhere otherwise you end up with all assembly. Too many words is just extra work without much benefit. Minimizing primitives makes the code slower, possibly a bit more compact, but otherwise to what advantage? >> I'm still not following your point. OVER is a very logical word to >> implement as a primitive. It is used often enough and requires little >> or nothing in the way of extra hardware for a CPU design and should be a >> single instruction on most existing CPUs. > > Pick a Forth word that you're not going to implement as an instruction. > Implement it using multiple instructions. Count how many. Now, if you > change your set of instructions, is it implemented in fewer or more > instructions? You want a set of instructions which results in fewer, > especially for the Forth words you think will be used the most. That is the purpose of using surveys like Koopman's. But I have been thinking lately that this may not be the best way to optimize instruction sets. Working with just *my* code would not be good for anything other than the code I've written in the past. > If you attempt to balance all instructions so each takes the same amount > of time, that results in a RISC like design. Yes? > > If you implement some instructions for speed and disregard optimization for > others, that results in a CISC like design. Yes? Does it? One of my design rules is that all instructions take one clock cycle, period. Adding states to control multi-cycle instructions adds to the instruction decode impacts the goal of minimizing the LUT usage. Another goal is minimize RAM usage which is important in space constrained apps. Rather than optimizing the word definitions in Forth equally, I expect to optimize the functions I will implement using the assembly language. These functions will be the ones that use most of the CPU time. Koopman's Forth word frequency is a starting point, but I will tweek the instructions according to what I use in the future. I took a quick look at 16 bit VLIW instructions this morning. I could code each of the three execution units separately and use 11 of the 16 bits leaving 5 bits for literal data in some instances. I could also combine the address unit and the fetch unit since they are closely related and use 4 bits each for the data and address units. This leaves 8 bits for literal data or possibly other controls when literals are not used. If I pursue this I may end up with a new set of possible instructions since address unit and data unit can work independently and in parallel. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-28 05:01 -0400 |
| Message-ID | <k6is21$7j0$1@speranza.aioe.org> |
| In reply to | #16786 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6hdga$4je$1@dont-email.me... > On 10/26/2012 8:57 PM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > > news:k6cg75$o2d$1@dont-email.me... ... > >> Ok, now that I get the context, what is the point exactly? Why do you > >> only code the lesser number of primitives when you will eventually code > >> more words as primitives. > > > > I'm in the process of converting much of my code from precompiled ITC > > Forth in C (similar to a DTC Forth in assembly) to Forth as ASCII text. > > The goal was to 1) make it as independent of C as possible and/or reveal > > hidden C dependencies, and 2) have only primitives, variables, and other > > low-level words in C. The reality is slightly different. What's left > > are the primitives, variables, other low-level words in C, and a few > > critical high-level Forth words in C, as well a program (re)written > > entirely in terms of primitives to implement the outer interpreter. > > Maybe you can explain the purpose of wanting Forth in ASCII text? Didn't I just explain that? > Wouldn't that be the source? Yes. It's Forth source. Previously, the definitions were in C, as ITC address-lists, just like Forth coded in assembly. Now, they're mostly in ASCII text as high-level Forth definitions. > Or are you talking about F-code in text? No. > What would be the purpose of doing this? See 1) and 2) above. > >> Is this an exercise where you hope to learn how best to > >> optimize primitives? > > > > No. It's mostly a result of the transformation process. Minimizing > > primitives allowed additional code to be moved out of the "C space" to > > the pure "Forth space". > > "Allowed"??? By definition that is minimizing primitives. Yes, allowed. I.e., even some of the typical "primitives" can be written in terms of other more primitive "primitives". A typical example might be C@ or BRANCH. For C@, which is typically a primitive, you could construct it using @ and AND. The same is true for BRANCH. It's usually a primitive, but can be easily constructed in Forth for an interpreted Forth (exact definition depends on implementation specific details): \ direct addresses : BRANCH R> @ >R ; \ relative offset with pointer adjustment : BRANCH R> DUP @ + CELL+ >R ; > But why? I thought I stated that. It's a result of eliminating the C code and dependencies, and converting it to Forth-in-Forth. > Typically more primitives are better in most metrics other > than possibly size. Yes, that comes later. Once I've got as much Forth-in-Forth as is possible, and as little Forth-in-C code as is possible, then I can begin (re)implementing more primitives and other higher-level words as primitives that will improve speed. > The main reason is that you need to draw a line somewhere > otherwise you end up with all assembly. ... all C. > Too many words is just extra work without much benefit. Exactly. You're clearly not going to implement all Forth words as instructions on your processor either. One of the things I noticed from the instruction frequencies for Forth, is that even the handful of the most used Forth words were used like only 10% of the time or so ... None of them were used 25% or 40% or 70% of the time. > Minimizing primitives makes the code slower, > possibly a bit more compact, but otherwise to what advantage? To allow the transformation from C to Forth occur. Each word that isn't a primitive can be implemented in high-level Forth. The fewer primitives there are the more code there is that than can be converted. At some point, I will reach the absolute minimum that I can comprehend. At that point, the code is the least dependent on C that I can make it. Then, I can easily reorganize or rewrite my Forth definitions. After that, I can rework them for improved speed, e.g., by adding or converting words to primitives. > > If you attempt to balance all instructions so each takes the same amount > > of time, that results in a RISC like design. Yes? > > > > If you implement some instructions for speed and disregard optimization > > for others, that results in a CISC like design. Yes? > > Does it? One of my design rules is that all instructions take one clock > cycle, period. If an XOR to memory has six steps and an AND to register has four steps? You mentioned MISC. Are breaking AND etc down into RISC-like operations that become your instructions? E.g., move-register-to-memory becomes an instruction? Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-28 15:23 -0400 |
| Message-ID | <k6k0nn$n8o$1@dont-email.me> |
| In reply to | #16801 |
On 10/28/2012 5:01 AM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6hdga$4je$1@dont-email.me... >> On 10/26/2012 8:57 PM, Rod Pemberton wrote: >>> "rickman"<gnuarm@gmail.com> wrote in message >>> news:k6cg75$o2d$1@dont-email.me... > ... > >>>> Ok, now that I get the context, what is the point exactly? Why do you >>>> only code the lesser number of primitives when you will eventually code >>>> more words as primitives. >>> >>> I'm in the process of converting much of my code from precompiled ITC >>> Forth in C (similar to a DTC Forth in assembly) to Forth as ASCII text. >>> The goal was to 1) make it as independent of C as possible and/or reveal >>> hidden C dependencies, and 2) have only primitives, variables, and other >>> low-level words in C. The reality is slightly different. What's left >>> are the primitives, variables, other low-level words in C, and a few >>> critical high-level Forth words in C, as well a program (re)written >>> entirely in terms of primitives to implement the outer interpreter. >> >> Maybe you can explain the purpose of wanting Forth in ASCII text? > > Didn't I just explain that? > >> Wouldn't that be the source? > > Yes. It's Forth source. > > Previously, the definitions were in C, as ITC address-lists, just like Forth > coded in assembly. Now, they're mostly in ASCII text as high-level Forth > definitions. > >> Or are you talking about F-code in text? > > No. > >> What would be the purpose of doing this? > > See 1) and 2) above. No, you didn't explain why you need to use source. You gave goals, not rationals for what you did. If you don't want to discuss this so I can understand, just say so. I'll stop asking. >>>> Is this an exercise where you hope to learn how best to >>>> optimize primitives? >>> >>> No. It's mostly a result of the transformation process. Minimizing >>> primitives allowed additional code to be moved out of the "C space" to >>> the pure "Forth space". >> >> "Allowed"??? By definition that is minimizing primitives. > > Yes, allowed. I.e., even some of the typical "primitives" can be written in > terms of other more primitive "primitives". A typical example might be C@ > or BRANCH. For C@, which is typically a primitive, you could construct it > using @ and AND. The same is true for BRANCH. It's usually a primitive, > but can be easily constructed in Forth for an interpreted Forth (exact > definition depends on implementation specific details): > > \ direct addresses > : BRANCH R> @>R ; > > \ relative offset with pointer adjustment > : BRANCH R> DUP @ + CELL+>R ; > >> But why? > > I thought I stated that. It's a result of eliminating the C code and > dependencies, and converting it to Forth-in-Forth. > >> Typically more primitives are better in most metrics other >> than possibly size. > > Yes, that comes later. Once I've got as much Forth-in-Forth as is possible, > and as little Forth-in-C code as is possible, then I can begin > (re)implementing > more primitives and other higher-level words as primitives that will improve > speed. > >> The main reason is that you need to draw a line somewhere >> otherwise you end up with all assembly. > > ... all C. > >> Too many words is just extra work without much benefit. > > Exactly. You're clearly not going to implement all Forth words as > instructions on your processor either. One of the things I noticed from the > instruction frequencies for Forth, is that even the handful of the most used > Forth words were used like only 10% of the time or so ... None of them > were used 25% or 40% or 70% of the time. Isn't that rather a DUH! You can only have one word with >50% frequency. You can only have three words with >30% frequency, etc. Are you expecting to have lots of words that are used 90% of the time? I really don't follow. >> Minimizing primitives makes the code slower, >> possibly a bit more compact, but otherwise to what advantage? > > To allow the transformation from C to Forth occur. Each word that isn't a > primitive can be implemented in high-level Forth. The fewer primitives > there are the more code there is that than can be converted. At some point, > I will reach the absolute minimum that I can comprehend. At that point, the > code is the least dependent on C that I can make it. Then, I can easily > reorganize or rewrite my Forth definitions. After that, I can rework them > for improved speed, e.g., by adding or converting words to primitives. This sounds very complex. >>> If you attempt to balance all instructions so each takes the same amount >>> of time, that results in a RISC like design. Yes? >>> >>> If you implement some instructions for speed and disregard optimization >>> for others, that results in a CISC like design. Yes? >> >> Does it? One of my design rules is that all instructions take one clock >> cycle, period. > > If an XOR to memory has six steps and an AND to register has four steps? > > You mentioned MISC. Are breaking AND etc down into RISC-like operations > that become your instructions? E.g., move-register-to-memory becomes an > instruction? RISC? I don't use RISC concepts, that would be too complex. MISC - M for Minimal. There are NO registers, there are TWO STACKS, FULL STOP! Instructions are things like, move stack to memory, xor popping the top two items on the data stack and pushing the result (still one operation), branch/call/return. These may not be the same as Forth words, but they can be used to make Forth words. There is no XOR to memory as an instruction. You can learn a lot about MISC by learning the instruction set of the F18A. Chuck uses two address registers which is different from my design. The instructions are similar in scale. They can all be done in one operation and don't require multiple steps. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-29 19:32 -0400 |
| Message-ID | <k6n3ep$lf2$1@speranza.aioe.org> |
| In reply to | #16816 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6k0nn$n8o$1@dont-email.me... > On 10/28/2012 5:01 AM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > > news:k6hdga$4je$1@dont-email.me... > >> On 10/26/2012 8:57 PM, Rod Pemberton wrote: > >>> "rickman"<gnuarm@gmail.com> wrote in message > >>> news:k6cg75$o2d$1@dont-email.me... ... > No, you didn't explain why you need to use source. It's just an intermediate stage. It may remain in part, or may not. I'd like to keep much as source. It's easy to change. But, I'm clearly going to get rid of words that aren't doing much except slow the interpreter down. If I keep the source and achieve my transformation goals, then I will only need to compile a very small C program to implement or execute Forth. Perhaps, maybe, it'll even be suitable for "sandboxing". > You gave goals, not rationals for what you did. I don't understand what you mean by that ... Isn't eliminating dependence on C a rationale too? > >> Too many words is just extra work without much benefit. > > > > Exactly. You're clearly not going to implement all Forth words as > > instructions on your processor either. One of the things I noticed from > > the instruction frequencies for Forth, is that even the handful of the > > most used Forth words were used like only 10% of the time or so ... > > None of them were used 25% or 40% or 70% of the time. > > Isn't that rather a DUH! You can only have one word with >50% > frequency. You can only have three words with >30% frequency, etc. Are > you expecting to have lots of words that are used 90% of the time? I > really don't follow. It's hard to determine where to concentrate your efforts if everything is of minimal use. If the language had a small set of more frequently used operations instead of a large set of minimally used operations, I'd think it would be better suited to implementation on a RISC style microprocessor, perhaps MISC too. Instead, people must analyze Forth words and potential instructions, figure out how to rewrite them, or break them up into other operations, etc until they think they've found the "magic" set of operations or sub-operations which should allow them to very effectively implement Forth. Obviously, even Charles Moore is _still_ working on this ... You're working on this in your way. I'm working this in my way. Alex McDonald in his. Bernd Paysan in his. Stephen Pelc in his. Marcel Hendrix in his. Anton Ertl in his. Andrew Haley in his. Hugh Aguilar in his. Etc. At what point does a highly effective Forth computation model become available? ... The apparent difficulty in finding an effective model implies to me that maybe the Forth language is not such a good fit to the hardware. Yes, I understand that a very simple computational model for Forth ITC interpreters works. But, this is the era of Forth compilers ... > RISC? I don't use RISC concepts, that would be too complex. MISC - M > for Minimal. There are NO registers, there are TWO STACKS, FULL STOP! > Instructions are things like, move stack to memory, xor popping the top > two items on the data stack and pushing the result (still one > operation), branch/call/return. These may not be the same as Forth > words, but they can be used to make Forth words. There is no XOR to > memory as an instruction. Ok. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-30 19:00 -0400 |
| Message-ID | <k6pm77$huv$1@dont-email.me> |
| In reply to | #16832 |
On 10/29/2012 7:32 PM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6k0nn$n8o$1@dont-email.me... >> On 10/28/2012 5:01 AM, Rod Pemberton wrote: >>> "rickman"<gnuarm@gmail.com> wrote in message >>> news:k6hdga$4je$1@dont-email.me... >>>> On 10/26/2012 8:57 PM, Rod Pemberton wrote: >>>>> "rickman"<gnuarm@gmail.com> wrote in message >>>>> news:k6cg75$o2d$1@dont-email.me... > ... > >> No, you didn't explain why you need to use source. > > It's just an intermediate stage. It may remain in part, or may not. I'd > like to keep much as source. It's easy to change. But, I'm clearly going > to get rid of words that aren't doing much except slow the interpreter down. > If I keep the source and achieve my transformation goals, then I will only > need to compile a very small C program to implement or execute Forth. > Perhaps, maybe, it'll even be suitable for "sandboxing". > >> You gave goals, not rationals for what you did. > > I don't understand what you mean by that ... > > Isn't eliminating dependence on C a rationale too? It isn't an explanation of the original question. Why did the chicken cross the road, to get to the other side... There are many ways to "eliminating dependence on C", the question is why did you choose to use Forth source in your design rather than compile the Forth? >>>> Too many words is just extra work without much benefit. >>> >>> Exactly. You're clearly not going to implement all Forth words as >>> instructions on your processor either. One of the things I noticed from >>> the instruction frequencies for Forth, is that even the handful of the >>> most used Forth words were used like only 10% of the time or so ... >>> None of them were used 25% or 40% or 70% of the time. >> >> Isn't that rather a DUH! You can only have one word with>50% >> frequency. You can only have three words with>30% frequency, etc. Are >> you expecting to have lots of words that are used 90% of the time? I >> really don't follow. > > It's hard to determine where to concentrate your efforts if everything is of > minimal use. If the language had a small set of more frequently used > operations instead of a large set of minimally used operations, I'd think it > would be better suited to implementation on a RISC style microprocessor, > perhaps MISC too. Instead, people must analyze Forth words and potential > instructions, figure out how to rewrite them, or break them up into other > operations, etc until they think they've found the "magic" set of operations > or sub-operations which should allow them to very effectively implement > Forth. Obviously, even Charles Moore is _still_ working on this ... You're > working on this in your way. I'm working this in my way. Alex McDonald in > his. Bernd Paysan in his. Stephen Pelc in his. Marcel Hendrix in his. > Anton Ertl in his. Andrew Haley in his. Hugh Aguilar in his. Etc. At > what point does a highly effective Forth computation model become available? > ... The apparent difficulty in finding an effective model implies to me > that maybe the Forth language is not such a good fit to the hardware. Yes, > I understand that a very simple computational model for Forth ITC > interpreters works. But, this is the era of Forth compilers ... > >> RISC? I don't use RISC concepts, that would be too complex. MISC - M >> for Minimal. There are NO registers, there are TWO STACKS, FULL STOP! >> Instructions are things like, move stack to memory, xor popping the top >> two items on the data stack and pushing the result (still one >> operation), branch/call/return. These may not be the same as Forth >> words, but they can be used to make Forth words. There is no XOR to >> memory as an instruction. > > Ok. Ok. I spent a little time today looking at using a 16 bit VLIW instruction set today against two snippets of code I wrote just to test instruction efficiency. One is a memory to memory copy and the other is a memory fill which I think I saw Chuck use for his CPU once. Surprisingly enough, I got zero utility of the parallel operation capability in the first example. I also got virtually no improvement in the number of instructions which means it used TWICE the code space if you don't count the literal setup code which was the same! I'm going to look at some more code, but moving data around in memory is a big one for the apps I am designing this for. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-31 01:23 -0400 |
| Message-ID | <k6qccg$704$1@speranza.aioe.org> |
| In reply to | #16848 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6pm77$huv$1@dont-email.me... [snip, re-org] > I spent a little time today looking at using a 16 bit VLIW instruction > set today against two snippets of code I wrote just to test instruction > efficiency. One is a memory to memory copy and the other is a memory > fill which I think I saw Chuck use for his CPU once. > > Surprisingly enough, I got zero utility of the parallel operation > capability in the first example. I also got virtually no improvement in > the number of instructions which means it used TWICE the code space if > you don't count the literal setup code which was the same! > > I'm going to look at some more code, but moving data around in > memory is a big one for the apps I am designing this for. The old Amiga's had a few coprocessors. I still find them to be novel some 26 or so years later. http://en.wikipedia.org/wiki/Original_Amiga_chipset Does your design include some type of DMA? Or, a "blitter"? IIRC, you said your MISC design had no registers only two stacks for Forth. So, what stack arguments are you using to specify a memory move? e.g., from, to, data size, length? from, end, to? Or, is this a part of the instruction? > [...] the question is why did you choose to use > Forth source in your design rather than compile the Forth? Compiled Forth is what I started with. The system was developed for many years using just that. It's useable. Compiled high-level Forth in C is very close to Forth in assembly. However, after the ability to parse Forth source was implemented, I've found it's also a bit more difficult to work with compiled Forth. It has to be compiled. Forth source doesn't. Generally, the basic Forth definitions for words are the same, or sometimes trivially different between the two forms. E.g., a LIT equivalent is used with numbers in the C code for compiled Forth, but numbers are used directly in the Forth source. A LIT is compiled into definitions by the text/outer interpreter when calling NUMBER or somesuch Forth word(s) ... Some code also can't be made to work exactly the same due to C issues, but they are generally very close, e.g., except branches, DOES> etc. Those differences are sometimes noticeable in the compiled Forth, but generally not easy to detect casually. I want a true Forth system, not a Forth clone, or near Forth environment. Early on, there was almost no Forth. I backfilled needed operations using C code and lots of it. Eventually, I reworked it to be mostly compiled Forth, but still with quite some C. After the Forth parsing words were implemented, I could migrate most compiled high-level Forth definitions to Forth source. Typically, they are the exact same Forth word definitions either way. Sometimes there are trivial differences, e.g., like LIT. With a few exceptions, the compiled definitions are the same binary-wise too. They basically had to be. I was migrating them from compiled Forth to source Forth one at a time. Currently, the time for parsing of the Forth dictionary is insignificant and unnoticeable. So, I don't "see" why compiled Forth would be preferred. Some of the stuff done over the years was very complicated. It was really only needed to progress the program to whatever next stage of development I decided I was working towards. E.g., at one point, I needed to be able to access variables, e.g., dictionary pointer (DP), in both C and Forth. Most were simple to do, but synchronising DP was very trying. Eventually, I eliminated the need for the DP on the C side. But, that required a major rewrite to eliminate the non-Forth C code that was accessing DP. Converting the non-Forth C code to Forth was just a total pain. I knew what I was doing on the C side, but had no idea how to implement the equivalent code in Forth. I ran into numerous other "bottlenecks" along the way too. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-31 14:56 -0400 |
| Message-ID | <k6rs9g$72r$1@dont-email.me> |
| In reply to | #16853 |
On 10/31/2012 1:23 AM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6pm77$huv$1@dont-email.me... > > [snip, re-org] > >> I spent a little time today looking at using a 16 bit VLIW instruction >> set today against two snippets of code I wrote just to test instruction >> efficiency. One is a memory to memory copy and the other is a memory >> fill which I think I saw Chuck use for his CPU once. >> >> Surprisingly enough, I got zero utility of the parallel operation >> capability in the first example. I also got virtually no improvement in >> the number of instructions which means it used TWICE the code space if >> you don't count the literal setup code which was the same! >> >> I'm going to look at some more code, but moving data around in >> memory is a big one for the apps I am designing this for. > > The old Amiga's had a few coprocessors. I still find them to be novel some > 26 or so years later. > > http://en.wikipedia.org/wiki/Original_Amiga_chipset > > Does your design include some type of DMA? Or, a "blitter"? Anything can be added to the FPGA. If you need hard DMA, add it. If you need a bit-blit, add it. It doesn't need to be part of the CPU. > IIRC, you said your MISC design had no registers only two stacks for Forth. > So, what stack arguments are you using to specify a memory move? e.g., > from, to, data size, length? from, end, to? Or, is this a part of the > instruction? The "return" stack is really an address stack and generates the addresses for memory ops. Even with an auto increment, doing a memory copy requires manipulating two addresses with stack ops to rearrange the addresses. I looked at the VLIW instructions today and I realized that there was parallism being wasted and came up with a three instruction loop that could copy memory and count down an index. Remember, that one of my goals is to keep the logic in the design as simple as possible. So all instructions have to use a fairly simple data path without messy, special muxes. I realized that the Fetch+ instruction (with autoincrement) could also do the swap required to get the right address on the top of the return stack and the Store+ will swap them back to get ready for the next Fetch+. The count on the data stack can be decremented as the loop counter in parallel with the Store+. The conditional jump is the third instruction. LITL AddrTo-1 LITL AddrFrm LITL Limit CALL MOVE . . . MOVE: RFRM, PREFETCH Loop1: FTCHP, RSWP STORP, RSWP, DCR, PREFETCH JNZ Loop1 RDRP RDRP DROP, RET The first four instructions are literal loads of the addresses and a move of the count to the data stack. Notice the loop is only three VLIW instructions. Then cleanup is three more instructions. The code size is slightly smaller or the same as with an 8/9 bit implementation. DROP is combined with a RET to end the subroutine. The prefetch is just a command to use the value being loaded to the top of the return stack to read data memory, it has nothing to do with instruction prefetch. The memory in FPGAs is registered and so the read has to happen at the start of the cycle. I may be able to change this by clocking the memory on the opposite phase of the clock, but that will require some consideration. >> [...] the question is why did you choose to use >> Forth source in your design rather than compile the Forth? > > Compiled Forth is what I started with. The system was developed for many > years using just that. It's useable. Compiled high-level Forth in C is > very close to Forth in assembly. However, after the ability to parse Forth > source was implemented, I've found it's also a bit more difficult to work > with compiled Forth. It has to be compiled. Forth source doesn't. > > Generally, the basic Forth definitions for words are the same, or > sometimes trivially different between the two forms. E.g., a LIT > equivalent is used with numbers in the C code for compiled Forth, but > numbers are used directly in the Forth source. A LIT is compiled into > definitions by the text/outer interpreter when calling NUMBER or somesuch > Forth word(s) ... Some code also can't be made to work exactly the same due > to C issues, but they are generally very close, e.g., except branches, DOES> > etc. Those differences are sometimes noticeable in the compiled Forth, but > generally not easy to detect casually. > > I want a true Forth system, not a Forth clone, or near Forth environment. > > Early on, there was almost no Forth. I backfilled needed operations using C > code and lots of it. Eventually, I reworked it to be mostly compiled Forth, > but still with quite some C. After the Forth parsing words were > implemented, I could migrate most compiled high-level Forth definitions to > Forth source. Typically, they are the exact same Forth word definitions > either way. Sometimes there are trivial differences, e.g., like LIT. With > a few exceptions, the compiled definitions are the same binary-wise too. > They basically had to be. I was migrating them from compiled Forth to > source Forth one at a time. > > Currently, the time for parsing of the Forth dictionary is insignificant and > unnoticeable. So, I don't "see" why compiled Forth would be preferred. > > Some of the stuff done over the years was very complicated. It was really > only needed to progress the program to whatever next stage of development I > decided I was working towards. E.g., at one point, I needed to be able to > access variables, e.g., dictionary pointer (DP), in both C and Forth. Most > were simple to do, but synchronising DP was very trying. Eventually, I > eliminated the need for the DP on the C side. But, that required a major > rewrite to eliminate the non-Forth C code that was accessing DP. Converting > the non-Forth C code to Forth was just a total pain. I knew what I was > doing on the C side, but had no idea how to implement the equivalent code > in Forth. I ran into numerous other "bottlenecks" along the way too. > > > Rod Pemberton > > I don't have those issues because technically I'm not writing Forth. I'm using a dual stack CPU. I will use Forth to write the compiler for the CPU and will likely program it in something like Forth, but that is secondary. Rick
[toc] | [prev] | [next] | [standalone]
| From | visualforth@rocketmail.com |
|---|---|
| Date | 2012-10-22 19:44 -0700 |
| Message-ID | <4cdc8277-bbae-4a01-a56b-92c103b1df00@googlegroups.com> |
| In reply to | #16579 |
On Monday, October 22, 2012 11:53:24 AM UTC-4, Brad Eckert wrote: > Hi All, I've been thinking about Novix style processors like the RTX2000. There are many Forth sequences that can be compacted into one instruction, so with a good optimizer the chip can execute several Forth (source) primitives in one machine cycle. I suspect though that such optimization opportunities are the exception rather than the rule. Does anyone here have a feel for the correspondence between Forth source primitives and generated code? This kind of architecture is pretty good if you can get a wide instruction from code space every clock. You have calls and you have everything else, where everything else is a kind of compact VLIW. Implementation can be simple, as shown by James Bowman's J1. The trick is to make the right use of the available bits. The RTX2000 uses two important tricks to make it run fast: First of all, one bit switches between two types of commands: general commands and subroutine calls. This bit could be bit zero. If bit zero is zero, the other bits form the address of the subroutine to be called. That's obvious, because you don't need odd addresses for subroutines. But the RTX2000 uses bit 15 and shifts the address to fit. Second, the aforementioned trick with the return from subroutine bit. These both accomplish speed first hand. The remaining bits should be used to decode all Forth primitives which are needed. I am sure you don't need a frequency of calling, because the Forth OS itself needs all of them. If there are enough bits, it will be possible to use these to run several primitives in one clock cycle - RISCs normally use only one clock cycle per command, the RTX2000 method allows up to four commands run per one clock cycle - it's some kind of Super-RISC. These commands running in one clock cycle should generate the next level of primitives. And of course there is the possibility to use some bits to switch between different kinds of decoding. The more bits, the more possibilities there are, and more commands can be made to run in one clock cycle, not necessarily in parallel.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-23 17:47 -0400 |
| Message-ID | <k6739i$b3k$1@dont-email.me> |
| In reply to | #16608 |
On 10/22/2012 10:44 PM, visualforth@rocketmail.com wrote: > On Monday, October 22, 2012 11:53:24 AM UTC-4, Brad Eckert wrote: >> Hi All, I've been thinking about Novix style processors like the RTX2000. There are many Forth sequences that can be compacted into one instruction, so with a good optimizer the chip can execute several Forth (source) primitives in one machine cycle. I suspect though that such optimization opportunities are the exception rather than the rule. Does anyone here have a feel for the correspondence between Forth source primitives and generated code? This kind of architecture is pretty good if you can get a wide instruction from code space every clock. You have calls and you have everything else, where everything else is a kind of compact VLIW. Implementation can be simple, as shown by James Bowman's J1. > > The trick is to make the right use of the available bits. > > The RTX2000 uses two important tricks to make it run fast: > First of all, one bit switches between two types of commands: general commands and subroutine calls. This bit could be bit zero. If bit zero is zero, the other bits form the address of the subroutine to be called. That's obvious, because you don't need odd addresses for subroutines. But the RTX2000 uses bit 15 and shifts the address to fit. > Second, the aforementioned trick with the return from subroutine bit. > > These both accomplish speed first hand. > > The remaining bits should be used to decode all Forth primitives which are needed. I am sure you don't need a frequency of calling, because the Forth OS itself needs all of them. With a 16 bit or larger instruction, you have enough bits to "encode" each execution unit in a dual stack machine separately. Then you don't need a separate bit for the return operation. It can't be used in parallel with any other Instruction Unit operation or a Return Stack operation since both of these are used to do a return. It can only be used in parallel with a purely Data Stack operation. The next time I look at a design on an FPGA I will take a look at a 16 or 18 bit instruction word that encodes the execution units operations separately. Part of the problem is that 16 is too many! What to do with the remainder? > If there are enough bits, it will be possible to use these to run several primitives in one clock cycle - RISCs normally use only one clock cycle per command, the RTX2000 method allows up to four commands run per one clock cycle - it's some kind of Super-RISC. These commands running in one clock cycle should generate the next level of primitives. And of course there is the possibility to use some bits to switch between different kinds of decoding.. The more bits, the more possibilities there are, and more commands can be made to run in one clock cycle, not necessarily in parallel. Primitives can only be run together if they don't conflict. That's why I want to look at which primitives occur together in the code. Rick
[toc] | [prev] | [next] | [standalone]
| From | visualforth@rocketmail.com |
|---|---|
| Date | 2012-10-23 16:00 -0700 |
| Message-ID | <e2535cbb-1b8d-4178-b5a3-f9085d627490@googlegroups.com> |
| In reply to | #16629 |
On Tuesday, October 23, 2012 5:47:32 PM UTC-4, rickman wrote: > > On 10/22/2012 10:44 PM, visualforth.com wrote: > > The RTX2000 uses two important tricks to make it run fast: > > Second, the aforementioned trick with the return from subroutine bit. > With a 16 bit or larger instruction, you have enough bits to "encode" each > execution unit in a dual stack machine separately. Then you don't need a > separate bit for the return operation. It can't be used in parallel with any > other Instruction Unit operation or a Return Stack operation since both of > these are used to do a return. It can only be used in parallel with a purely > Data Stack operation. If you say so ....
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-23 21:07 -0400 |
| Message-ID | <k67f09$9r4$2@dont-email.me> |
| In reply to | #16631 |
On 10/23/2012 7:00 PM, visualforth@rocketmail.com wrote: > On Tuesday, October 23, 2012 5:47:32 PM UTC-4, rickman wrote: >>> On 10/22/2012 10:44 PM, visualforth.com wrote: >>> The RTX2000 uses two important tricks to make it run fast: >>> Second, the aforementioned trick with the return from subroutine bit. > >> With a 16 bit or larger instruction, you have enough bits to "encode" each >> execution unit in a dual stack machine separately. Then you don't need a >> separate bit for the return operation. It can't be used in parallel with any >> other Instruction Unit operation or a Return Stack operation since both of >> these are used to do a return. It can only be used in parallel with a purely >> Data Stack operation. > > If you say so .... I don't understand. Are you agreeing or passively saying you don't agree? Rick
[toc] | [prev] | [next] | [standalone]
| From | visualforth@rocketmail.com |
|---|---|
| Date | 2012-10-23 18:44 -0700 |
| Message-ID | <3e649d75-d4e1-4e4f-8c3f-e4f629308692@googlegroups.com> |
| In reply to | #16638 |
On Tuesday, October 23, 2012 9:07:21 PM UTC-4, rickman wrote: > On 10/23/2012 7:00 PM, visualforth.com wrote: > On Tuesday, October 23, 2012 5:47:32 PM UTC-4, rickman wrote: >>> On 10/22/2012 10:44 PM, visualforth.com wrote: >>> The RTX2000 uses two important tricks to make it run fast: >>> Second, the aforementioned trick with the return from subroutine bit. > >> With a 16 bit or larger instruction, you have enough bits to "encode" each >> execution unit in a dual stack machine separately. Then you don't need a >> separate bit for the return operation. It can't be used in parallel with any >> other Instruction Unit operation or a Return Stack operation since both of >> these are used to do a return. It can only be used in parallel with a purely >> Data Stack operation. > > If you say so .... > I don't understand. Are you agreeing or passively saying you don't agree? > Rick I have to confess: communication is not easy. I wrote "If you say so ..." to address that you have your own opinion in this matter. My point of view is - I forgot to mention - that a special return bit closes a definition, and instead of having a separate return instruction which needs one clock cycle more, the next address is the start address of the following word. When the RTX2000 was introduced, IIRC any other microprocessor had a special return instruction, which needed a full memory access and normally a lot of clock cycles. That's why back then the main stream teaching was: "avoid subroutines". The structure of Forth is based on "subroutines" - to accelerate that process, a special functionality has been needed: the subroutine bit. Now, fifteen years later, this may be taken for granted.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-24 16:03 -0400 |
| Message-ID | <k69hj8$phb$1@dont-email.me> |
| In reply to | #16639 |
On 10/23/2012 9:44 PM, visualforth@rocketmail.com wrote: > On Tuesday, October 23, 2012 9:07:21 PM UTC-4, rickman wrote: >> On 10/23/2012 7:00 PM, visualforth.com wrote:> On Tuesday, October 23, 2012 5:47:32 PM UTC-4, rickman wrote:>>> On 10/22/2012 10:44 PM, visualforth.com wrote:>>> The RTX2000 uses two important tricks to make it run fast:>>> Second, the aforementioned trick with the return from subroutine bit.> >> With a 16 bit or larger instruction, you have enough bits to "encode" each>> execution unit in a dual stack machine separately. Then you don't need a>> separate bit for the return operation. It can't be used in parallel with any>> other Instruction Unit operation or a Return Stack operation since both of>> these are used to do a return. It can only be used in parallel with a purely>> Data Stack operation.> > If you say so .... > >> I don't understand. Are you agreeing or passively saying you don't agree? >> Rick > > I have to confess: communication is not easy. > I wrote "If you say so ..." to address that you have your own opinion in this matter. > > My point of view is - I forgot to mention - that a special return bit closes a definition, and instead of having a separate return instruction which needs one clock cycle more, the next address is the start address of the following word. When the RTX2000 was introduced, IIRC any other microprocessor had a special return instruction, which needed a full memory access and normally a lot of clock cycles. That's why back then the main stream teaching was: "avoid subroutines". The structure of Forth is based on "subroutines" - to accelerate that process, a special functionality has been needed: the subroutine bit. > Now, fifteen years later, this may be taken for granted. This is all intended to be for discussion. It ended up being a bit long. Please don't take it as anything other than friendly discussion and please feel free to disagree with me. Yeah. You can talk about subroutines being slow or bad or whatever. But they get used. I'm not sure what your point is about this. I don't ever remember being taught to "avoid subroutines". In fact, I was taught just the opposite in a structured programming class. They talked about the advantages of subroutines, like information hiding, decision hiding, module reuse, etc... Forth carries that to the extreme for sure, so optimizing the subroutine linkage is useful. But oddly enough, if you want to optimize returns, the return "bit" doesn't work as well as you might think for a number of reasons. As I said before, only about half the instructions can be combined withe the return function. In Forth code it can be hard to use this bit. If your last word before the ; is an ELSE, what instruction do you combine the return with? If the last word before the ; is another definition and not a primitive you can't combine a return with a call... or can you? Both of these situations allow optimization in a different way. The last definition call of a word definition can be changed from a call to a jump, this is called "tail recursion", IIRC. The ELSE before a ; just means the tail recursion is performed for both cases of the IF. Words like LOOP or R> can't be combined with the return because they are using the return stack. Don't forget there is a cost to that return "bit". It is a bit in the instruction word that might be put to other, more productive use. The Univac 1108 had a cascaded indirect bit that has never been used in any machine since because all of the address bits were better used as... address bits. The return bit may be good in some ways, but can't be used in many cases. Also, don't overrate the savings of the return "bit". It saves an instruction access, but in my machine it is only a single clock cycle. It is only in standard CPUs that a subroutine uses multiple clock cycles. Rick
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-24 13:15 -0700 |
| Message-ID | <7xip9zmxsl.fsf@ruckus.brouhaha.com> |
| In reply to | #16667 |
rickman <gnuarm@gmail.com> writes: > Don't forget there is a cost to that return "bit". It is a bit in the > instruction word that might be put to other, more productive use. I could imagine an instruction set packed into a bit stream using Huffman codes. It would use more decoding hardware but the programs would be smaller. It's possible that "return" would end up as a single bit, if it was much more common than any other instruction. More likely it would be 2 or 3 bits.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-24 16:25 -0400 |
| Message-ID | <k69iro$2om$1@dont-email.me> |
| In reply to | #16671 |
On 10/24/2012 4:15 PM, Paul Rubin wrote: > rickman<gnuarm@gmail.com> writes: >> Don't forget there is a cost to that return "bit". It is a bit in the >> instruction word that might be put to other, more productive use. > > I could imagine an instruction set packed into a bit stream using > Huffman codes. It would use more decoding hardware but the programs > would be smaller. It's possible that "return" would end up as a single > bit, if it was much more common than any other instruction. More likely > it would be 2 or 3 bits. That is essentially what I did with my instruction encoding. I had to work with a fixed instruction size, so I used the extra bits for data (address mostly). In an 8 bit instruction a leading 0 meant the rest of the instruction was literal data. A leading 1 followed by the next three bits said it was some form of call (rel or abs) or jump (several different conditions) with four bits of immediate address to be appended to any literal previously loaded. The remaining 16 instructions were direct opcodes of which return was one, a full 8 bits. Take a look at Forth code sometime and consider how often tail recursion can be used and you will see that the "return" instruction is not so often used. In addition the remaining cases often can't be combined with the return bit. At least that is what I found when I tried it. Rick
[toc] | [prev] | [next] | [standalone]
Page 8 of 9 — ← Prev page 1 2 3 4 5 6 7 [8] 9 Next page →
Back to top | Article view | comp.lang.forth
csiph-web