Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #16579 > unrolled thread
| Started by | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| First post | 2012-10-22 08:53 -0700 |
| Last post | 2012-10-23 10:52 +0000 |
| Articles | 12 on this page of 172 — 22 participants |
Back to article view | Back to comp.lang.forth
RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000
Page 9 of 9 — ← Prev page 1 2 3 4 5 6 7 8 [9]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-10-26 03:15 +0200 |
| Message-ID | <8891969.HGoWPOBEZb@sunwukong.fritz.box> |
| In reply to | #16667 |
rickman wrote: > If the last word before the ; is another > definition and not a primitive you can't combine a return with a > call... or can you? Of course, that's called "tail call elimination" and is a standard technique in compilers on normal CPUs. You turn the call into a jump. It is a bit more complicated in languages like C with their stack frames, but well, a C compiler is a bit more complicated than a Forth compiler. In Forth, it's really just converting the call into a jump. Maybe you can't do it for every call, because the jumps might not cover the entire address space, but if you can, it is a useful optimization. It is about the only optimization I do on the b16 - which doesn't have a return bit, but since combining calls+returns happens quite frequently in Forth, I grab this low-hanging fruit. The RTX2000 had a "horizontal" instruction encoding. This term comes from microcode engines, and on "horizontal" engines, you could encode the operations of all available resources (ALUs, incrementers, register files, etc.) in one wide horizontal line of microcode, while a "vertical" encoding had a selector for the resource, and one instruction for the selected resource. The horizontal microcode executed all in parallel, if useful or not; the code is less dense. The vertical microcode execute all in sequence, the performance then is less good. So even if a combination is not always useful or even doable, a horizontal style will make it possible, because *when* it is doable, it provides a speed advantage. The RTX2000 allowed to combine more than one Forth operation into one instruction, IIRC up to three (with the return being the third). In conventional processors, this sort of instruction combination was touted "superscalar", and Intel's first superscalar processor was the Pentium. Though, as Forth instructions don't do much work, you could argue that mov [offset+reg],reg is already equivalent to three Forth instructions (lit + @ or lit + !). As usual in engineering, these different styles have different tradeoffs. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| Date | 2012-10-24 09:55 -0700 |
| Message-ID | <d916285b-b2e5-46d2-8022-26d0e061dac7@googlegroups.com> |
| In reply to | #16629 |
On Tuesday, October 23, 2012 2:47:32 PM UTC-7, rickman wrote: > > The next time I look at a design on an FPGA I will take a look at a 16 > or 18 bit instruction word that encodes the execution units operations > separately. Part of the problem is that 16 is too many! What to do > with the remainder? > Look at the J1, which weighs in at 200 lines of Verilog. There isn't much instruction decoding, it's mostly small bit fields controlling different muxes.
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-24 10:04 -0700 |
| Message-ID | <7x7gqf94xy.fsf@ruckus.brouhaha.com> |
| In reply to | #16654 |
Brad Eckert <hwfwguy@gmail.com> writes: > Look at the J1, which weighs in at 200 lines of Verilog. There isn't > much instruction decoding, it's mostly small bit fields controlling > different muxes. I wonder how the code density for these tiny Forth cores compares with the typical C target.
[toc] | [prev] | [next] | [standalone]
| From | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| Date | 2012-10-24 12:00 -0700 |
| Message-ID | <90458090-bc72-42cb-82e6-a6988b213994@googlegroups.com> |
| In reply to | #16655 |
On Wednesday, October 24, 2012 10:04:58 AM UTC-7, Paul Rubin wrote: > > I wonder how the code density for these tiny Forth cores compares with > the typical C target. The J1 was used because the Microblaze CPU was too big, too slow, and the C application (for an IP camera) was too big to fit in available memory. The Forth version was 62% smaller: 6349 vs 16380 bytes. See http://excamera.com/sphinx/fpga-j1.html
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-24 23:24 -0700 |
| Message-ID | <7xvcdzm5ml.fsf@ruckus.brouhaha.com> |
| In reply to | #16663 |
Brad Eckert <hwfwguy@gmail.com> writes: > The J1 was used because the Microblaze CPU was too big, too slow, and > the C application (for an IP camera) was too big to fit in available > memory. The Forth version was 62% smaller: 6349 vs 16380 bytes. See > http://excamera.com/sphinx/fpga-j1.html The J1 is really cool and I remember looking at it before. I guess it uses RAM blocks for the stacks? It's very clever, the instruction words are almost like microcode. The hardware folks here would know better than me, but I wonder if the J1 architecture is tuned for FPGA implementation in a way that would lose some advantage (compared to something like b16) in silicon. I notice the paper doesn't say how many CLB's are consumed. I see all the opcodes are used, but a couple of them could be synthesized from others, so maybe it's possible to repurpose them for new instructions. Or maybe some unused combinations of opcodes and flags could be hijacked, if checking for them doesn't cause delays. It's interesting that 22% of instructions in his code are literals. I wonder if most of them are either 0 or 1. I can understand a lot of the Verilog by reading it. I'd sure like to be able to write it too.
[toc] | [prev] | [next] | [standalone]
| From | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| Date | 2012-10-25 09:26 -0700 |
| Message-ID | <cd9d49f8-0d25-4e6d-a83d-7e7a4fafa502@googlegroups.com> |
| In reply to | #16689 |
On Wednesday, October 24, 2012 11:24:05 PM UTC-7, Paul Rubin wrote: > The J1 is really cool and I remember looking at it before. I guess it > uses RAM blocks for the stacks? It's very clever, the instruction words > are almost like microcode. The hardware folks here would know better > than me, but I wonder if the J1 architecture is tuned for FPGA > implementation in a way that would lose some advantage (compared to > something like b16) in silicon. > Stacks use LUT RAM, which is fine for small RAMs. It is geared more toward FPGA implementation than the B16, but that's more implementation than ISA. Forth Day 2010 was an interesting year. Three different CPU cores were presented. The J1, a minimal control processor; the SC20, a more full featured application processor; and Ting's P32, something in between. Minimal processors are a good fit for Forth. Since the PAUSE chain can be made so fast, interrupts are often not needed. Not having interrupts simplifies the CPU design and saves you from having to do a bunch of verification to see where an asynchronous interrupt could break the system. Thinking about the different kinds of stack machines led me to start this thread. The Novix style ISA is good and compact. The compiler needs to be smarter than that of a MISC to take advantage of it.
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-25 10:39 -0700 |
| Message-ID | <7xr4ome9it.fsf@ruckus.brouhaha.com> |
| In reply to | #16699 |
Brad Eckert <hwfwguy@gmail.com> writes: > Stacks use LUT RAM, which is fine for small RAMs. It is geared more > toward FPGA implementation than the B16, but that's more > implementation than ISA. Ah, thanks, I see that now, about the LUT ram. I wonder how much hardware it uses compared to the b16. The amount of verilog code is very small, but some of the things it does look expensive: 1) the stacks are way bigger than the b16's (in addition to being in LUT's) 2) There is a "barrel shifter" (T = N << (T&0x0F) and T = N >> (T&0x0F)). I wonder how often code actually uses it. 3) I'm not sure but I saw something to the effect that on every cycle, it computes all the possible instruction results (T+N, T&N, memory read, etc.) in parallel, so the result is ready by the time the instruction fetch completes. The instruction then just selects which result to use. This sounds power hungry? The paper mentions that putting the stacks in LUT complicates PICK and ROLL, and I see in nuc.fs that they are done as unrolled loops. I wonder what obstacles there might have been to allowing random access to the stacks. That might have made the J1 a more attractive target for conventional compilers, without complicating the Verilog too much.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-25 19:35 -0400 |
| Message-ID | <k6cicb$8ol$1@dont-email.me> |
| In reply to | #16700 |
On 10/25/2012 1:39 PM, Paul Rubin wrote: > Brad Eckert<hwfwguy@gmail.com> writes: >> Stacks use LUT RAM, which is fine for small RAMs. It is geared more >> toward FPGA implementation than the B16, but that's more >> implementation than ISA. > > Ah, thanks, I see that now, about the LUT ram. I wonder how much > hardware it uses compared to the b16. The amount of verilog code is > very small, but some of the things it does look expensive: It may use LUT RAM on a Xilinx or Lattice part, but many FPGAs don't *have* LUT RAM in which case it is done in block RAM. LUT RAM has limitations and it is not unusual to use block RAM to more fully exploit the advantages, like full dual read ports. > 1) the stacks are way bigger than the b16's (in addition to being in LUT's) If they are "way bigger", they probably aren't in LUT RAM. Most devices with LUT RAM are only 16 deep. Otherwise you need to use multiplexors to combined multiple LUTs for each bit. > 2) There is a "barrel shifter" (T = N<< (T&0x0F) and T = N>> > (T&0x0F)). I wonder how often code actually uses it. Good question. Barrel shifters are used in floating point calculations but are often cheap in FPGAs because you can implement them in multipliers if you can tolerate the clock delay as the multipliers are clocked. > 3) I'm not sure but I saw something to the effect that on every cycle, > it computes all the possible instruction results (T+N, T&N, memory read, > etc.) in parallel, so the result is ready by the time the instruction > fetch completes. The instruction then just selects which result to use. > This sounds power hungry? I doubt this is really any more power hungry than most other designs. Every register that changes dissipates power in all of the logic it feeds regardless of whether you use that path or not. The only question is does the J1 have multiple logic paths where otherwise a single path with multiple functions would be used? An subtractor is just an adder with an inverted input and a one on the carry. Or do you use an adder and a subtractor followed by a mux? BTW, the mux to select the function outputs can be as much logic as the functions themselves and extra power as well. > The paper mentions that putting the stacks in LUT complicates PICK and > ROLL, and I see in nuc.fs that they are done as unrolled loops. I > wonder what obstacles there might have been to allowing random access to > the stacks. That might have made the J1 a more attractive target for > conventional compilers, without complicating the Verilog too much. I don't know of any reason why LUT RAM is any different from block RAM in doing PICK or ROLL. I believe ROLL requires you to copy sequential locations in the stack and PICK just copies one element to the top of stack. If the stack is in RAM, regardless of which type of RAM is used, this is all the same. It requires some extra hardware to mux an address which is not the top of stack, but is not precluded and is the same for both types. Rick
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-10-26 18:26 +0200 |
| Message-ID | <22523497.Rbg4U2RY2y@sunwukong.fritz.box> |
| In reply to | #16700 |
Paul Rubin wrote: > 1) the stacks are way bigger than the b16's (in addition to being in > LUT's) On the FPGA implementation, I tried to make it possible to use block RAM for the stacks - so you could make the stacks larger without sacrifying much resources. > 2) There is a "barrel shifter" (T = N << (T&0x0F) and T = N >> > (T&0x0F)). I wonder how often code actually uses it. I doubt it's much code that uses a barrel shifter. > 3) I'm not sure but I saw something to the effect that on every cycle, > it computes all the possible instruction results (T+N, T&N, memory > read, etc.) in parallel, so the result is ready by the time the > instruction > fetch completes. The instruction then just selects which result to > use. This sounds power hungry? Yes, this is potentially more power hungry than what the b16 does. The b16 has a compact ALU, which needs to be set up at the start of the cycle to produce the correct result (it does +, &, |, xor all with the same logic element, the full adder, and needs multiplexers setup to feed the carry through for +). I also only access the RAM when it is necessary. Some FPGA tools have problems converting the b16 case statement into a multiplexer; for these it might help to do that "by hand". > The paper mentions that putting the stacks in LUT complicates PICK and > ROLL, and I see in nuc.fs that they are done as unrolled loops. I > wonder what obstacles there might have been to allowing random access > to > the stacks. That might have made the J1 a more attractive target for > conventional compilers, without complicating the Verilog too much. For a conventional compiler, I would stronlgy recommend to simply add the workspace concept from the transputer. E.g. expand the b16 instructions from 5 to 6 bits, add a workspace pointer register, and have 16 W+x@ and 16 W+x! instructions. If you want to help OOP programming languages, have two registers, one for the current object, one for the workspace, and access only 8 locations in each. However, for these small embedded targets, this extra instructions are not really necessary. The programs there usually have a completely static call tree, so you can allocate the locals statically; as the Keil compiler does for an 8051. The easiest allocation algorithm starts with leaf procedures: Their locals go to the smallest addresses (easiest to address; leaf procedures do most of the work). Then proceed to the next layer in the call graph, and allocate the locals there, starting with the smallest available address, too (addresses reserved by leaf calls are not to be used). And so on, until you get to main(). A function pointer call has to be treated as if all the functions which pointer has been taken are called there. Recursive functions will face a performance hit and need a separate dynamic stack area (same as with Keil), unless the compiler can optimize all locals away. Calling conventions should be Forth-compatible, i.e. you push parameters on the stack when calling a function. The function then might place the parameters in locals. We had a student doing some experiments with a relatively simple C compiler that used a stack based intermediate language, and the conclusion of the student (who did this as summer practica) was that you can target C to it, but you get considerably slower and larger code than with Forth. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-24 16:10 -0400 |
| Message-ID | <k69hve$sgk$1@dont-email.me> |
| In reply to | #16655 |
On 10/24/2012 1:04 PM, Paul Rubin wrote: > Brad Eckert<hwfwguy@gmail.com> writes: >> Look at the J1, which weighs in at 200 lines of Verilog. There isn't >> much instruction decoding, it's mostly small bit fields controlling >> different muxes. > > I wonder how the code density for these tiny Forth cores compares with > the typical C target. I compared my code to a RISC instruction set once, don't recall which one. They had 16 bit instructions and the code did a simple memory copy loop. The RAM sizes were the same. Mine had twice as many instructions but the RISC instructions were twice as large. If the RISC were large enough it might do each instruction in one clock cycle. But my entire CPU would probably be smaller than the register file on the RISC. Remember that this CPU design is optimized for implementing Forth and I was comparing assembly to assembly. Who knows how it would do with C vs. Forth on appropriate CPUs. I do believe that MISC is the way to go for many applications. But CISC/RISC is firmly entrenched. Rick
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-24 13:21 -0700 |
| Message-ID | <7xehknmxia.fsf@ruckus.brouhaha.com> |
| In reply to | #16669 |
rickman <gnuarm@gmail.com> writes: > I do believe that MISC is the way to go for many applications. But > CISC/RISC is firmly entrenched. Looking at the F18 die photo, most of the F18 area is memory arrays and the stacks. The ALU is maybe 10% of it. If they went to 6-bit instructions instead of 5-bit, they could add a bunch more primitives, more registers, a few built-in constants, etc. Perhaps the ALU would become 2x or 3x larger, but there might well be significant improvements in code size and performance to make up for it.
[toc] | [prev] | [next] | [standalone]
| From | stephenXXX@mpeforth.com (Stephen Pelc) |
|---|---|
| Date | 2012-10-23 10:52 +0000 |
| Message-ID | <50867352.85104223@192.168.0.50> |
| In reply to | #16579 |
On Mon, 22 Oct 2012 08:53:22 -0700 (PDT), Brad Eckert <hwfwguy@gmail.com> wrote: >I've been thinking about Novix style processors like the RTX2000. There are= > many Forth sequences that can be compacted into one instruction, so with a= > good optimizer the chip can execute several Forth (source) primitives in o= >ne machine cycle. I suspect though that such optimization opportunities are= > the exception rather than the rule. See: http://www.complang.tuwien.ac.at/anton/euroforth/ef04/pelc-bailey04.pdf This machine has been run in an FPGA. http://www-users.cs.york.ac.uk/~chrisb/main-pages/fpga/fpga-research-interests >Does anyone here have a feel for the correspondence between Forth source pr= >imitives and generated code? Many address generating phrases such as "dup 8 + @" can be collapsed to single instructions. However, the result is not a Forth purist's CPU. Stephen -- Stephen Pelc, stephenXXX@mpeforth.com MicroProcessor Engineering Ltd - More Real, Less Time 133 Hill Lane, Southampton SO15 5AF, England tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691 web: http://www.mpeforth.com - free VFX Forth downloads
[toc] | [prev] | [standalone]
Page 9 of 9 — ← Prev page 1 2 3 4 5 6 7 8 [9]
Back to top | Article view | comp.lang.forth
csiph-web