Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #16579 > unrolled thread
| Started by | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| First post | 2012-10-22 08:53 -0700 |
| Last post | 2012-10-23 10:52 +0000 |
| Articles | 20 on this page of 172 — 22 participants |
Back to article view | Back to comp.lang.forth
RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000
Page 1 of 9 [1] 2 3 4 5 6 7 8 9 Next page →
| From | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| Date | 2012-10-22 08:53 -0700 |
| Subject | RTX2000 optimization |
| Message-ID | <e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com> |
Hi All, I've been thinking about Novix style processors like the RTX2000. There are many Forth sequences that can be compacted into one instruction, so with a good optimizer the chip can execute several Forth (source) primitives in one machine cycle. I suspect though that such optimization opportunities are the exception rather than the rule. Does anyone here have a feel for the correspondence between Forth source primitives and generated code? This kind of architecture is pretty good if you can get a wide instruction from code space every clock. You have calls and you have everything else, where everything else is a kind of compact VLIW. Implementation can be simple, as shown by James Bowman's J1.
[toc] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-10-22 11:21 -0500 |
| Message-ID | <vOqdnTp57Pbl7xjNnZ2dnUVZ8vOdnZ2d@supernews.com> |
| In reply to | #16579 |
Brad Eckert <hwfwguy@gmail.com> wrote: > Hi All, > > I've been thinking about Novix style processors like the > RTX2000. There are many Forth sequences that can be compacted into > one instruction, so with a good optimizer the chip can execute > several Forth (source) primitives in one machine cycle. I suspect > though that such optimization opportunities are the exception rather > than the rule. It happens a lot bacause many phrases are things like OVER + and R> DROP . Also, ; pairs with just about everything. Novix had more of these phrases than RTX2000 because of the way its encoding was done. Andrew.
[toc] | [prev] | [next] | [standalone]
| From | visualforth@rocketmail.com |
|---|---|
| Date | 2012-10-22 09:56 -0700 |
| Message-ID | <3b14c2ce-6e19-4df7-abc2-cf169dcf9665@googlegroups.com> |
| In reply to | #16579 |
On Monday, October 22, 2012 11:53:24 AM UTC-4, Brad Eckert wrote: > Hi All, I've been thinking about Novix style processors like the RTX2000. There are many Forth sequences that can be compacted into one instruction, so with a good optimizer the chip can execute several Forth (source) primitives in one machine cycle. I suspect though that such optimization opportunities are the exception rather than the rule. Does anyone here have a feel for the correspondence between Forth source primitives and generated code? This kind of architecture is pretty good if you can get a wide instruction from code space every clock. You have calls and you have everything else, where everything else is a kind of compact VLIW. Implementation can be simple, as shown by James Bowman's J1. Despite allowing several Forth commands inside one RTX2000 16 bit word a great advantage is the ability to set a special Return from Subroutine bit - that is a real speed accelerator - so a whole executable program may need only one RTX2000 16 bit word! Imagine what additional possibilities you can have with 32 bit words!
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <forthfreak@gmail.com> |
|---|---|
| Date | 2012-10-22 12:10 -0700 |
| Message-ID | <07562d07-693e-48ec-b996-610c4eba1cc9@k21g2000vbj.googlegroups.com> |
| In reply to | #16584 |
On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: > Despite allowing several Forth commands inside one RTX2000 16 bit word a great advantage is the ability to set a special Return from Subroutine bit - that is a real speed accelerator - so a whole executable program may need only one RTX2000 16 bit word! Imagine what additional possibilities you can have with 32 bit words! Hmmm... I'm not following you. What is the function/significance of this special bit? Where is it set? In the call instruction or the return instruction?
[toc] | [prev] | [next] | [standalone]
| From | visualforth@rocketmail.com |
|---|---|
| Date | 2012-10-22 12:55 -0700 |
| Message-ID | <3919a089-b570-4665-8e29-cc20199181b7@googlegroups.com> |
| In reply to | #16585 |
On Monday, October 22, 2012 3:10:29 PM UTC-4, M.R.W Wills wrote: > On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: > Despite allowing several Forth commands inside one RTX2000 16 bit word a great advantage is the ability to set a special Return from Subroutine bit - that is a real speed accelerator - so a whole executable program may need only one RTX2000 16 bit word! Imagine what additional possibilities you can have with 32 bit words! Hmmm... I'm not following you. What is the function/significance of this special bit? Where is it set? In the call instruction or the return instruction? There is no special return instruction. Return is achieved by one bit only: For all instructions except Subroutine Calls or Branch instructions, bit 5 of the instruction code represents the Subroutine Return Bit. If this bit is set to 1, a Return is performed whereby the return address is popped from the Return Stack. Source: Intersil HS-RTX2010RH Data Sheet March 2000 File Number 3961.3, p. 28 http://www.intersil.com/content/dam/Intersil/documents/fn39/fn3961.pdf Subroutine Return Bit, ibid., p. 31, HARRIS RTX2000 ADVANCE INFORMATION, May 1988, p. 16
[toc] | [prev] | [next] | [standalone]
| From | Coos Haak <chforth@hccnet.nl> |
|---|---|
| Date | 2012-10-22 22:02 +0200 |
| Message-ID | <zj8a1eiw6l23$.1bnvz91x2j9mn$.dlg@40tude.net> |
| In reply to | #16585 |
Op Mon, 22 Oct 2012 12:10:28 -0700 (PDT) schreef Mark Wills: > On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: >> Despite allowing several Forth commands inside one RTX2000 16 bit word a great advantage is the ability to set a special Return from Subroutine bit - that is a real speed accelerator - so a whole executable program may need only one RTX2000 16 bit word! Imagine what additional possibilities you can have with 32 bit words! > > Hmmm... I'm not following you. What is the function/significance of > this special bit? Where is it set? In the call instruction or the > return instruction? There is no return instruction, but any instruction may have its return bit set. : SQR DUP + ; may be one instruction word existing of DUP, plus and the return bit. -- Coos CHForth, 16 bit DOS applications http://home.hccnet.nl/j.j.haak/forth.html
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-22 16:50 -0400 |
| Message-ID | <k64bj1$m4n$1@dont-email.me> |
| In reply to | #16585 |
On 10/22/2012 3:10 PM, Mark Wills wrote: > On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: >> Despite allowing several Forth commands inside one RTX2000 16 bit word a great advantage is the ability to set a special Return from Subroutine bit - that is a real speed accelerator - so a whole executable program may need only one RTX2000 16 bit word! Imagine what additional possibilities you can have with 32 bit words! > > Hmmm... I'm not following you. What is the function/significance of > this special bit? Where is it set? In the call instruction or the > return instruction? <this post turned out to be a bit longer than I expected...> Stack machines are inherently pretty simple because of the limited nature of registers and the limited instruction set typically implemented. I did one about ten years ago and found it really consists of three execution units, all of which can operate in parallel. 1) The Instruction Unit contains the instruction memory with the address generation (instruction fetch) as well as the decode (although the decode could also be spread among the units). 2) The Address Unit contains the return stack with stack pointer and the various logic that manipulates elements on the stack, such as auto-increment of addresses and loop counter functions. The address for main memory is provided by this unit in my machine. 3) The Data Unit contains the data stack with stack pointer, the ALU and any other special logic for operations on the data stack. I also included the memory interface in this unit even though the address comes from the Address Unit. Each of these three units is present in any dual stack CPU design. Any given instruction may involve any combination of the three units or may leave some idle. If instructions leave any unit idle it can be combined with another instruction that uses just that unit in a compatible way. For example, as others have indicated, any instruction that is not using the Address Unit, e.g. ADD, SUB, etc. can execute an instruction that only uses the Address Unit, e.g. RET, LOOP, etc. Both instructions have to be using the Instruction Unit in a compatible way which is typically not hard since most instructions are doing the equivalent of NEXT (or NOP in other terms). The only issue with such combining of instructions is the instruction encoding. I encoded to minimize the amount of program space needed which resulted in a minimum width instruction optimized per Koopman's data for instruction frequency. I looked at separating the opcodes for each unit in essence making the MISC equivalent of a VLIW processor, if you can have such a thing... lol I couldn't quite squeeze the instruction into a 9 bit word which is a memory width commonly available in FPGAs. I am now looking at using an FPGA that only supports 16 bit wide memory and don't like the idea of multiplexing the instructions, that just adds a level of logic to the timing path. A wider instruction might just minimize the decode logic and provide for more instruction parallelism at the same time. Another way of paralleling instructions is to just pick unused opcodes and implement the most common parallel instructions. This won't improve timing of the design and may worsen it a bit, but will provide for fewer instructions in a program. The obvious one, return as a separate bit, in parallel with everything, can only be used with about half the instructions in my CPU design. The return stack is used in a lot of them. But if the bit is free... Info that would be VERY useful to me is frequency of use of instructions in combination. This could be pulled from existing code by measuring how often instructions are found adjacent to each other. This may not be a perfect measure, but it would be a great start. Can anyone generate a metric on this similar to Koopman's data on instruction use? Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-22 21:52 -0400 |
| Message-ID | <k64t1j$nlq$1@speranza.aioe.org> |
| In reply to | #16595 |
"rickman" <gnuarm@gmail.com> wrote in message news:k64bj1$m4n$1@dont-email.me... > On 10/22/2012 3:10 PM, Mark Wills wrote: > > On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: > >> Despite allowing several Forth commands inside one RTX2000 16 bit > >> word a great advantage is the ability to set a special Return from > >> Subroutine bit - that is a real speed accelerator - so a whole > >> executable program may need only one RTX2000 16 bit word! > >> Imagine what additional possibilities you can have with 32 bit words! > > > > Hmmm... I'm not following you. What is the function/significance of > > this special bit? Where is it set? In the call instruction or the > > return instruction? > > <this post turned out to be a bit longer than I expected...> > I'm versed in (old) microprocessor design. So, that's easily taken care of: [SNIP] > The only issue with such combining of instructions is the instruction > encoding. I encoded to minimize the amount of program space needed > which resulted in a minimum width instruction optimized per Koopman's > data for instruction frequency. I looked at separating the opcodes for > each unit in essence making the MISC equivalent of a VLIW processor, if > you can have such a thing... lol I couldn't quite squeeze the > instruction into a 9 bit word which is a memory width commonly available > in FPGAs. [...] Forth's built using "primitives" generally need 30 to 40 or so, i.e., 5-bits. Why do you need 9-bits (512)? I'd guess that you're encoding things other than just the instruction, e.g., control-bits, offsets, modes, etc. Koopman's and Ertl's instruction frequency data is basically the same. Using their data is a good choice though. (This is also posted earlier in the thread:) The question for both you (and Brad) is if you create new, faster, more powerful, multiple operation instructions, how do you ensure they are used? Without an optimizer, it's likely the instruction will have a low instruction frequency. I.e., a person is unlikely to use it. In which case, there is no point in using or implementing it. > The obvious one, return as a separate bit, in parallel with everything, > can only be used with about half the instructions in my CPU design. The > return stack is used in a lot of them. But if the bit is free... Years ago, there was a processor that used a few bits, like two, for conditional execution of each instruction. I don't recall what it was, or if it was a Forth processor. It might've been a bit-slice design... > Info that would be VERY useful to me is frequency of use of instructions > in combination. This could be pulled from existing code by measuring > how often instructions are found adjacent to each other. This may not > be a perfect measure, but it would be a great start. Can anyone > generate a metric on this similar to Koopman's data on instruction use? Anton Ertl also has instruction frequency data. I don't recall if he showed combinations or not. I know he or someone created the concept of "super operators" for Forth, which I think is what you're asking about. If he doesn't respond, I'll attempt to locate for you what I previously found. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-10-22 22:03 -0700 |
| Message-ID | <02194379-ee88-4ea5-a65c-ca78c88545e0@tr7g2000pbc.googlegroups.com> |
| In reply to | #16606 |
On Oct 22, 6:48 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> wrote: > Years ago, there was a processor that used a few bits, like two, for > conditional execution of each instruction. I don't recall what it was, or > if it was a Forth processor. It might've been a bit-slice design... Isn't that the way that the ARM works?
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-10-23 03:19 -0500 |
| Message-ID | <0LednT1JJZiAzhvNnZ2dnUVZ8iOdnZ2d@supernews.com> |
| In reply to | #16610 |
Hugh Aguilar <hughaguilar96@yahoo.com> wrote: > On Oct 22, 6:48?pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> > wrote: >> Years ago, there was a processor that used a few bits, like two, for >> conditional execution of each instruction. ?I don't recall what it was, or >> if it was a Forth processor. ?It might've been a bit-slice design... > > Isn't that the way that the ARM works? It was, but it's mostly been dropped in ARM 64 because "Benchmarking shows that modern branch predictors work well enough that predicated execution of instructions does not offer sufficient benefit to justify its significant use of opcode space, and its implementation cost in advanced implementations." Andrew.
[toc] | [prev] | [next] | [standalone]
| From | vandys@vsta.org |
|---|---|
| Date | 2012-10-23 17:54 +0000 |
| Message-ID | <aeo3v5Fi14oU1@mid.individual.net> |
| In reply to | #16614 |
Andrew Haley <andrew29@littlepinkcloud.invalid> wrote: > Hugh Aguilar <hughaguilar96@yahoo.com> wrote: >>> Years ago, there was a processor that used a few bits, like two, for >>> conditional execution of each instruction. >> Isn't that the way that the ARM works? > It was, but it's mostly been dropped in ARM 64 because "Benchmarking > shows that modern branch predictors work well enough that predicated > execution of instructions does not offer sufficient benefit to justify > its significant use of opcode space, and its implementation cost in > advanced implementations." FWIW, the Propeller CPU has conditional execution too. I did a fair amount of hand-coding of assembly for it, and came away not liking its instruction set very much at all. MIPS is my all-time favorite, but I'd even take 32-bit x86 over Propeller. -- Andy Valencia Home page: http://www.vsta.org/andy/ To contact me: http://www.vsta.org/contact/andy.html
[toc] | [prev] | [next] | [standalone]
| From | Hugh Aguilar <hughaguilar96@yahoo.com> |
|---|---|
| Date | 2012-10-24 08:16 -0700 |
| Message-ID | <ddf7bf76-6fbe-430b-a11b-ce75ef7945ea@c20g2000yqe.googlegroups.com> |
| In reply to | #16623 |
On Oct 23, 10:54 am, van...@vsta.org wrote: > Andrew Haley <andre...@littlepinkcloud.invalid> wrote: > > Hugh Aguilar <hughaguila...@yahoo.com> wrote: > >>> Years ago, there was a processor that used a few bits, like two, for > >>> conditional execution of each instruction. > >> Isn't that the way that the ARM works? > > It was, but it's mostly been dropped in ARM 64 because "Benchmarking > > shows that modern branch predictors work well enough that predicated > > execution of instructions does not offer sufficient benefit to justify > > its significant use of opcode space, and its implementation cost in > > advanced implementations." > > FWIW, the Propeller CPU has conditional execution too. I did a fair amount > of hand-coding of assembly for it, and came away not liking its instruction > set very much at all. MIPS is my all-time favorite, but I'd even take 32-bit > x86 over Propeller. > > -- > Andy Valencia > Home page:http://www.vsta.org/andy/ > To contact me:http://www.vsta.org/contact/andy.html I looked at the Propeller advertisements, and the seemed pretty cool to have 8 processors on a board working together. I've been intending to get one. What was the problem with the instruction set? Your comment about the 32-bit x86 seems to imply that the problem with the propeller is register starvation, as that is certainly the problem with the 32-bit x86. The MIPS has a lot of registers and an orthogonal instruction set. It seems to be not very popular compared to the ARM though. I recently bought a Microstik-II from MicroChip. It has both the PIC24 and the PIC32. I am programming the PIC24 which will be the first target for Straight Forth. The PIC32 (which is a MIPS) seems like an obvious second target. I've programmed the PIC16 and PIC24 in the past, and I'm pretty impressed with MicroChip development software and their technical support. What is it that you like about the MIPS? Are you programming the PIC32? Isn't it true that the MIPS was always an academic exercise in the past for VHDL, but the PIC32 is the only commercial product featuring it?
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-23 17:19 -0400 |
| Message-ID | <k671kb$1uc$1@dont-email.me> |
| In reply to | #16606 |
On 10/22/2012 9:52 PM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k64bj1$m4n$1@dont-email.me... >> On 10/22/2012 3:10 PM, Mark Wills wrote: >>> On Oct 22, 5:56 pm, visualfo...@rocketmail.com wrote: > >>>> Despite allowing several Forth commands inside one RTX2000 16 bit >>>> word a great advantage is the ability to set a special Return from >>>> Subroutine bit - that is a real speed accelerator - so a whole >>>> executable program may need only one RTX2000 16 bit word! >>>> Imagine what additional possibilities you can have with 32 bit words! >>> >>> Hmmm... I'm not following you. What is the function/significance of >>> this special bit? Where is it set? In the call instruction or the >>> return instruction? >> >> <this post turned out to be a bit longer than I expected...> >> > > I'm versed in (old) microprocessor design. So, that's easily taken care of: > [SNIP] > >> The only issue with such combining of instructions is the instruction >> encoding. I encoded to minimize the amount of program space needed >> which resulted in a minimum width instruction optimized per Koopman's >> data for instruction frequency. I looked at separating the opcodes for >> each unit in essence making the MISC equivalent of a VLIW processor, if >> you can have such a thing... lol I couldn't quite squeeze the >> instruction into a 9 bit word which is a memory width commonly available >> in FPGAs. [...] > > Forth's built using "primitives" generally need 30 to 40 or so, i.e., > 5-bits. Why do you need 9-bits (512)? I'd guess that you're encoding > things other than just the instruction, e.g., control-bits, offsets, modes, > etc. I was constrained by the memories available in FPGAs. Many allow somewhat flexible word widths of 1, 2, 4, 8, 9, 16 and 18 bits, some even 32 and 36 bits. Multiplexing multiple instructions in one word has a downside in that it requires extra levels of logic in the instruction decode path which affects *all* instructions. I wanted to avoid that. So I started with a 4 bit instruction and found that rather limiting, mainly in the impact on performance since most code as around twice as long as it could be with larger instructions. Literals (both data and address) were especially problematic. Looking at Koopman's data it was clear that anything which could optimize the address fields of calls and other instructions, including immediate data would be a boon. So I tried 8 bit words and used a variable bit with instruction with the remaining bits as immediate data. This was combined with a data extension scheme similar to that used by the Transputer. They used 4 bit instructions with 4 bits of immediate data which would be shifted into larger words. Since this would be the most commonly used instruction I gave it one bit with 7 bit immediate data. The first invocation of a literal instruction pushes the top of return stack with the 7 bit data, sign extended. Each subsequent invocation of the literal instruction shifts 7 more bits into the top of return stack. Calls and Jumps have a four field which is combined with the top of return stack if a literal has been pushed, or just sign extended if not. There remains some 16 opcodes for the various instructions for manipulating data. The 8 bit machine was used in one design. I considered a 9 bit version to fully utilize the block RAM in most FPGAs. The immediate data fields were extended by one bit which I think was significant for jumps and calls, 5 bits vs. 4). In the case of general opcodes the extra bit could be used to provide 32 instructions rather than 16, but I didn't feel this gave much benefit and complicated the instruction decode which was already more complex than I preferred. Another alternative was to use the extra bit to flag a combined Return instruction. I found it could only be used with about half the opcodes I was using because of conflicts. > Koopman's and Ertl's instruction frequency data is basically the same. > Using their data is a good choice though. It is a LOT better than no data at all which is what I have otherwise. > (This is also posted earlier in the thread:) > The question for both you (and Brad) is if you create new, faster, more > powerful, multiple operation instructions, how do you ensure they are used? > Without an optimizer, it's likely the instruction will have a low > instruction frequency. I.e., a person is unlikely to use it. In which > case, there is no point in using or implementing it. Who is this "person"? My design was for me and if I generated enough code to analyze statistically, I would pick the instructions to combine from analyzing my code. It's not like I am selling this design for others to use... not that I wouldn't mind sharing, but you have to read my mind for much of the details. One person asked and I gave him my block diagram with labeled control points and my opcode cheat sheet. He couldn't make heads or tails out of it... lol >> The obvious one, return as a separate bit, in parallel with everything, >> can only be used with about half the instructions in my CPU design. The >> return stack is used in a lot of them. But if the bit is free... > > Years ago, there was a processor that used a few bits, like two, for > conditional execution of each instruction. I don't recall what it was, or > if it was a Forth processor. It might've been a bit-slice design... I've heard of that as well as other "unique" features. An ancient Univac machine had a bit in the address field that flagged indirect. The address fetched had the same bit... they had to add a indirect counter to get out of the infinite loops that could happen. >> Info that would be VERY useful to me is frequency of use of instructions >> in combination. This could be pulled from existing code by measuring >> how often instructions are found adjacent to each other. This may not >> be a perfect measure, but it would be a great start. Can anyone >> generate a metric on this similar to Koopman's data on instruction use? > > Anton Ertl also has instruction frequency data. I don't recall if he showed > combinations or not. I know he or someone created the concept of "super > operators" for Forth, which I think is what you're asking about. If he > doesn't respond, I'll attempt to locate for you what I previously found. That would be greatly interesting. I'm surprised I didn't notice this before. I wish there was some market for a machine like this, but then others would have done this before me. I know Bernd would have been all over this years ago if the market existed as well as others. It seems that if the CPU isn't pipelined and blazing fast, it isn't interesting to most FPGA users. They prefer very high performance CPUs that access MBs of external memory and use 1000's of LUTs. My design has an extensible word size, but will likely never have a C compiler for it. That reminds me of the ZPU. It is a stack machine designed to run C. It was also designed to be as tiny as possible in the minimal configuration with other versions running faster but using more resources. The ported the gcc tools for it. I think the minimal version is slightly smaller than my design, but very slow, maybe 10x... or would that be 10/? Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-24 06:56 -0400 |
| Message-ID | <k68h98$mor$1@speranza.aioe.org> |
| In reply to | #16625 |
"rickman" <gnuarm@gmail.com> wrote in message news:k671kb$1uc$1@dont-email.me... > On 10/22/2012 9:52 PM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > > news:k64bj1$m4n$1@dont-email.me... ... > >> Info that would be VERY useful to me is frequency of use of > >> instructions in combination. This could be pulled from existing > >> code by measuring how often instructions are found adjacent to > >> each other. This may not be a perfect measure, but it would be a > >> great start. Can anyone generate a metric on this similar to > >> Koopman's data on instruction use? > > > > Anton Ertl also has instruction frequency data. I don't recall if he > > showed combinations or not. I know he or someone created the concept > > of "superoperators" for Forth, which I think is what you're asking > > about. If he doesn't respond, I'll attempt to locate for you what I > > previously found. > > That would be greatly interesting. Anton Ertl provided links to instruction frequencies. It lists combinations of instructions also. The pdf he posted uses the term "superinstruction". It also mentions another paper by Proebstring using the term "superoperators". Todd Proebstring "Optimizing an ANSI C Interpreter with Superoperators" http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.94.4382&rep=rep1&type=pdf > It seems that if the CPU isn't pipelined and blazing fast, it isn't > interesting to most FPGA users. Without pipelining, I doubt it's of interest to anyone. The 6502 proved how effective pipelining in a microprocessor is in increasing performance. > That reminds me of the ZPU. It is a stack machine designed to run C. > The name seemed familiar. But, I didn't recall it. It seems you mentioned this circa 2008 in c.l.f. in a thread in which I responded. Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | anton@mips.complang.tuwien.ac.at (Anton Ertl) |
|---|---|
| Date | 2012-10-24 12:26 +0000 |
| Message-ID | <2012Oct24.142629@mips.complang.tuwien.ac.at> |
| In reply to | #16644 |
"Rod Pemberton" <do_not_have@notemailnotz.cnm> writes:
>The pdf he posted uses the term "superinstruction". It also mentions
>another paper by Proebstring using the term "superoperators".
>
>Todd Proebstring "Optimizing an ANSI C Interpreter with Superoperators"
>http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.94.4382&rep=rep1&type=pdf
Superinstructions combine adjacent VM instructions in a sequential
representation of a program. Superoperators combine adjacent
operators in a tree representation. E.g., consider a C statement like
a[i+3] = b[j]
with a tree representation (shown as S-expression) like
(c!
(+
(var a)
(+
(var i)
(const 3)))
(c@
(+
(var b)
(var j))))
and s sequential representation like
var_b var_j + c@ var_a var_i 3 + + c!
Then you can combine a tree pattern like (c! * (c@ *)) into a
superoperator (say c@!, resulting in the following tree:
(c@!
(+
(var a)
(+
(var i)
(const 3)))
(+
(var b)
(var j)))
You cannot do that with superinstructions.
And you can combine the sequence
c@ var_*
into a superinstruction c@var_*, resulting in the following sequence:
var_b var_j + c@var_a var_i 3 + + c!
You cannot do that with superoperators.
Helmut Eller compared superoperators with superinstructions in his
diploma thesis
<http://www.complang.tuwien.ac.at/Diplomarbeiten/eller05.ps.gz> and
found superinstructions to be slightly better.
- anton
--
M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard: http://www.forth200x.org/forth200x.html
EuroForth 2012: http://www.euroforth.org/ef12/
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-24 15:25 -0400 |
| Message-ID | <k69fbv$a6n$1@dont-email.me> |
| In reply to | #16644 |
On 10/24/2012 6:56 AM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message >> It seems that if the CPU isn't pipelined and blazing fast, it isn't >> interesting to most FPGA users. > > Without pipelining, I doubt it's of interest to anyone. The 6502 proved how > effective pipelining in a microprocessor is in increasing performance. I don't follow that. There are lots of MCUs that aren't pipelined. The need for speed is a spectrum. Often there are other things that are more important requirements, like size. How many LUTs is the Leon processor, 5000? Fast, but large. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-25 00:30 -0400 |
| Message-ID | <k6af1l$k0c$1@speranza.aioe.org> |
| In reply to | #16664 |
"rickman" <gnuarm@gmail.com> wrote in message news:k69fbv$a6n$1@dont-email.me... > On 10/24/2012 6:56 AM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > >> It seems that if the CPU isn't pipelined and blazing fast, it isn't > >> interesting to most FPGA users. > > > > Without pipelining, I doubt it's of interest to anyone. The 6502 > > proved how effective pipelining in a microprocessor is in increasing > > performance. > > I don't follow that. There are lots of MCUs that aren't pipelined. The > need for speed is a spectrum. Often there are other things that are > more important requirements, like size. How many LUTs is the Leon > processor, 5000? Fast, but large. > Are we talking about the same thing? http://en.wikipedia.org/wiki/Instruction_pipeline http://en.wikipedia.org/wiki/Pipelining AFAIK, the 6502 was the first microprocessor with pipelining. And, the Cray CDC 6600 had the first central processing unit with it. If we're talking about the same thing and a microprocessor design isn't using pipelining, how do you get any acceptable level of performance? Do you just use faster logic and clock? Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | Paul Rubin <no.email@nospam.invalid> |
|---|---|
| Date | 2012-10-24 22:10 -0700 |
| Message-ID | <7xbofrt9v0.fsf@ruckus.brouhaha.com> |
| In reply to | #16682 |
"Rod Pemberton" <do_not_have@notemailnotz.cnm> writes: > If we're talking about the same thing and a microprocessor design isn't > using pipelining, how do you get any acceptable level of performance? Do > you just use faster logic and clock? You just get less performance. A non-pipelined processor typically uses 3-4 cycles for basic instructions that a pipelined one does in 1 cycle (throughput). That may still be enough for your application. If it's not, pick a different processor, usually at a cost in die area and power consumption.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-25 16:43 -0400 |
| Message-ID | <k6c89i$aqq$1@dont-email.me> |
| In reply to | #16686 |
On 10/25/2012 1:10 AM, Paul Rubin wrote: > "Rod Pemberton"<do_not_have@notemailnotz.cnm> writes: >> If we're talking about the same thing and a microprocessor design isn't >> using pipelining, how do you get any acceptable level of performance? Do >> you just use faster logic and clock? > > You just get less performance. A non-pipelined processor typically uses > 3-4 cycles for basic instructions that a pipelined one does in 1 cycle > (throughput). That may still be enough for your application. If it's > not, pick a different processor, usually at a cost in die area and power > consumption. You are thinking of CISC or RISC processors where the instructions are more complex. As is true for many MISC designs, every instruction on my processor is one clock cycle. Pipelining might help to speed up the performance, but there become complex interactions between adjacent instructions and interrupt handling becomes more complex. I'd like to get back to working on this CPU design. I wasn't able to optimize it in ways I originally set out to do. If I return to it I may be able to do that. Rick
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-10-26 17:55 +0200 |
| Message-ID | <3503068.uZqhCW9t74@sunwukong.fritz.box> |
| In reply to | #16710 |
rickman wrote: > On 10/25/2012 1:10 AM, Paul Rubin wrote: >> "Rod Pemberton"<do_not_have@notemailnotz.cnm> writes: >>> If we're talking about the same thing and a microprocessor design >>> isn't >>> using pipelining, how do you get any acceptable level of >>> performance? Do you just use faster logic and clock? >> >> You just get less performance. A non-pipelined processor typically >> uses 3-4 cycles for basic instructions that a pipelined one does in 1 >> cycle >> (throughput). That may still be enough for your application. If >> it's not, pick a different processor, usually at a cost in die area >> and power consumption. > > You are thinking of CISC or RISC processors where the instructions are > more complex. As is true for many MISC designs, every instruction on > my > processor is one clock cycle. Pipelining might help to speed up the > performance, but there become complex interactions between adjacent > instructions and interrupt handling becomes more complex. The only pipelining that helps performance on a MISC is the instruction prefetch - and that's something MISCs often do. A typical simple RISC processor pipeline has four stages: instruction fetch, register read, ALU operation, register write; a CISC or a RISC with compressed instructions like ARM Thumb needs an additional decode stage. Since a stack based processor doesn't need register read and write (TOS and NOS are directly in the ALU path, the access to the stack to push/pull NOS is in parallel), no further pipeline stage is necessary. The concept of packing multiple sequential instructions into one word helps to reduce the conflicts in such a small pipeline. If your instruction fetch takes one cycle, you only need to fetch the next instruction in the last slot of an instruction. You either could reduce the number of possible instructions there to those which don't conflict with the instruction fetch (e.g. only ALU and stack operations, no load/store operations, no branches), or you delay the prefetch if there is a conflict, or you reduce the number of possible instructions in the first slot - that's what I do on the b16. The first slot can only do a call or a nop, and those two choices are selected right when the instruction has been fetched (do or not do a call). With a complete prefetch, I could do 3 instructions in 3 cycles, without this prefetch, I do 3 instructions in 4 cycles, except if the first instruction is a call - that's single cycle again. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
Page 1 of 9 [1] 2 3 4 5 6 7 8 9 Next page →
Back to top | Article view | comp.lang.forth
csiph-web