Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.forth > #16579 > unrolled thread
| Started by | Brad Eckert <hwfwguy@gmail.com> |
|---|---|
| First post | 2012-10-22 08:53 -0700 |
| Last post | 2012-10-23 10:52 +0000 |
| Articles | 20 on this page of 172 — 22 participants |
Back to article view | Back to comp.lang.forth
RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000
Page 2 of 9 — ← Prev page 1 [2] 3 4 5 6 7 8 9 Next page →
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-26 19:44 -0400 |
| Message-ID | <k6f789$rn9$1@dont-email.me> |
| In reply to | #16737 |
On 10/26/2012 11:55 AM, Bernd Paysan wrote: > rickman wrote: > >> On 10/25/2012 1:10 AM, Paul Rubin wrote: >>> "Rod Pemberton"<do_not_have@notemailnotz.cnm> writes: >>>> If we're talking about the same thing and a microprocessor design >>>> isn't >>>> using pipelining, how do you get any acceptable level of >>>> performance? Do you just use faster logic and clock? >>> >>> You just get less performance. A non-pipelined processor typically >>> uses 3-4 cycles for basic instructions that a pipelined one does in 1 >>> cycle >>> (throughput). That may still be enough for your application. If >>> it's not, pick a different processor, usually at a cost in die area >>> and power consumption. >> >> You are thinking of CISC or RISC processors where the instructions are >> more complex. As is true for many MISC designs, every instruction on >> my >> processor is one clock cycle. Pipelining might help to speed up the >> performance, but there become complex interactions between adjacent >> instructions and interrupt handling becomes more complex. > > The only pipelining that helps performance on a MISC is the instruction > prefetch - and that's something MISCs often do. A typical simple RISC > processor pipeline has four stages: instruction fetch, register read, > ALU operation, register write; a CISC or a RISC with compressed > instructions like ARM Thumb needs an additional decode stage. Since a > stack based processor doesn't need register read and write (TOS and NOS > are directly in the ALU path, the access to the stack to push/pull NOS > is in parallel), no further pipeline stage is necessary. I won't argue this point much, but any operation that uses more than one level of gates (or LUTs in an FPGA) can be pipelined. Pipelining doesn't need to have anything to do with the functional units of a processor. The instruction decode can easily be pipelined by adding registers. The fetch has the issue of the memory access typically being slower than much of the rest of the logic, but if that is buffered into a register the address generation can be pipelined. I believe the Pentium 4 had something like 15 stages to its pipeline which illustrates the problem. The Pentium 4 was only faster in some cases because of the cost of a pipeline flush. Not to mention the huge transistor count. But then the problem they were starting to have to what to do with the transistors that were available with each new process generation! A MISC can be pipelined by adding registers to the ALU, the decode and the stack address generation as well as the instruction address generation. But it adds lots of complexity to the logic. > The concept of packing multiple sequential instructions into one word > helps to reduce the conflicts in such a small pipeline. If your > instruction fetch takes one cycle, you only need to fetch the next > instruction in the last slot of an instruction. You either could reduce > the number of possible instructions there to those which don't conflict > with the instruction fetch (e.g. only ALU and stack operations, no > load/store operations, no branches), or you delay the prefetch if there > is a conflict, or you reduce the number of possible instructions in the > first slot - that's what I do on the b16. The first slot can only do a > call or a nop, and those two choices are selected right when the > instruction has been fetched (do or not do a call). With a complete > prefetch, I could do 3 instructions in 3 cycles, without this prefetch, > I do 3 instructions in 4 cycles, except if the first instruction is a > call - that's single cycle again. > All of these are design details that depend on a lot more than just the things you mention. Requiring a branch instruction to be put off to the next word just so that next instruction word can be fetched efficiently when it might be followed by another instruction fetch is of dubious value. This all depends on many design trade offs. I don't see the value of multiplexing instructions in words in an FPGA CPU which is what I design. In an IC the requirements and trade offs are entirely different. Rick
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-10-27 02:17 +0200 |
| Message-ID | <1843381.MoqF7MJDlU@sunwukong.fritz.box> |
| In reply to | #16750 |
rickman wrote: > I believe the Pentium 4 had something like 15 stages to its pipeline > which illustrates the problem. The Pentium 4 was only faster in some > cases because of the cost of a pipeline flush. Not to mention the > huge > transistor count. But then the problem they were starting to have to > what to do with the transistors that were available with each new > process generation! > > A MISC can be pipelined by adding registers to the ALU, the decode and > the stack address generation as well as the instruction address > generation. But it adds lots of complexity to the logic. The Pentium 4 (first generation) had a double-pumped ALU. I.e. instead of splitting up the ALU into two cycles, it completed the work in one half cycle, and two operations with dependency in the ALU path could be issued in one cycle. The other 15 cycles were spent on decoding and scheduling instructions from the trace cache - they were already pre- decoded. The ALU is not the bottleneck of any CPU. You can possibly pipeline it, but at the cost of dependent instructions, which are fairly frequent (that's why the Pentium 4 team considered this approach - it speeds up dependent instructions). In any case, feel free to choose tradeoffs as you like. For the tasks I've used the b16 with, performance didn't matter. Power consumption and code size did matter. Power consumption means something like "rush to finish", but that's considering the most simple structure you can use - if you can manage it in 30% less time by burning twice the current, it's no good. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-26 23:01 -0400 |
| Message-ID | <k6hcat$tc1$3@dont-email.me> |
| In reply to | #16756 |
On 10/26/2012 8:17 PM, Bernd Paysan wrote: > rickman wrote: >> I believe the Pentium 4 had something like 15 stages to its pipeline >> which illustrates the problem. The Pentium 4 was only faster in some >> cases because of the cost of a pipeline flush. Not to mention the >> huge >> transistor count. But then the problem they were starting to have to >> what to do with the transistors that were available with each new >> process generation! >> >> A MISC can be pipelined by adding registers to the ALU, the decode and >> the stack address generation as well as the instruction address >> generation. But it adds lots of complexity to the logic. > > The Pentium 4 (first generation) had a double-pumped ALU. I.e. instead > of splitting up the ALU into two cycles, it completed the work in one > half cycle, and two operations with dependency in the ALU path could be > issued in one cycle. The other 15 cycles were spent on decoding and > scheduling instructions from the trace cache - they were already pre- > decoded. > > The ALU is not the bottleneck of any CPU. You can possibly pipeline it, > but at the cost of dependent instructions, which are fairly frequent > (that's why the Pentium 4 team considered this approach - it speeds up > dependent instructions). > > In any case, feel free to choose tradeoffs as you like. For the tasks > I've used the b16 with, performance didn't matter. Power consumption > and code size did matter. Power consumption means something like "rush > to finish", but that's considering the most simple structure you can use > - if you can manage it in 30% less time by burning twice the current, > it's no good. > I don't think you understand me. I am not saying pipelining is an important thing to use. It all depends on your requirements. I am simply responding to your statement that "The only pipelining that helps performance on a MISC is the instruction prefetch". I simply don't agree with that because what holds back your MISC design depends on your MISC design. BTW, my CPU design doesn't have a prefetch. It just fetches each instruction in parallel with the current instruction execution. Because the instruction fetch unit is in parallel with the other units there is no problem with the instruction needing to be "prefetched". It would only be if the design were pipelined that the fetch would be a "prefetch". I don't intend to even try to pipeline my CPU design mainly because I don't have the time or interest. I prefer to try to optimize design tradeoffs which don't complicate the design, but should actually simplify it. In other words, optimizing by simplifying... much like what Chuck prefers. But for now I'll keep the SWAP instruction and will call an XOR an XOR. I guess stuff like IF ELSE THEN and OR instead of XOR is what you get when you work just to please yourself! Still, you have to recognize the brilliance of the man. Rick
[toc] | [prev] | [next] | [standalone]
| From | Bernd Paysan <bernd.paysan@gmx.de> |
|---|---|
| Date | 2012-10-27 22:18 +0200 |
| Message-ID | <1762397.1u9z86ECTl@sunwukong.fritz.box> |
| In reply to | #16783 |
rickman wrote: > I don't think you understand me. Partly, probably because you insist on calling your design non-pipelined and your parallel instruction fetch unit not a "prefetch", even though it is. My opinion is that the only pipelining that helps a MISC is the instruction prefetch. You can pipeline other parts, *but it doesn't help*. I.e. it doesn't provide better performance, because you will be waiting for dependent instructions. The only low-hanging fruit you have is to make the instruction fetch unit work in parallel with the other units. The ALU and the multiplexers in a MISC are all small and fast, and the first limit you get is the memory timing limit. > I am not saying pipelining is an > important thing to use. It all depends on your requirements. I am > simply responding to your statement that "The only pipelining that > helps > performance on a MISC is the instruction prefetch". I simply don't > agree with that because what holds back your MISC design depends on > your MISC design. BTW, my CPU design doesn't have a prefetch. It just > fetches each instruction in parallel with the current instruction > execution. This is the definition of a prefetch. You fetch the next instruction before you complete with the current instruction. > Because the instruction fetch unit is in parallel with the > other units there is no problem with the instruction needing to be > "prefetched". It would only be if the design were pipelined that the > fetch would be a "prefetch". As you have just explained, your design is pipelined, and you do a prefetch. A non-pipelined execution fetches current instruction and executes it, and when done would start to fetch the next instruction. Nothing parallel. > I don't intend to even try to pipeline my CPU design mainly because I > don't have the time or interest. I prefer to try to optimize design > tradeoffs which don't complicate the design, but should actually > simplify it. In other words, optimizing by simplifying... much like > what Chuck prefers. But for now I'll keep the SWAP instruction and > will call an XOR an XOR. Calling an XOR and XOR is really wise, and SWAP is a useful instruction. I don't know why Chuck wants to deliberately confuse people with his OR=XOR and -=INVERT. If he wants to be compact, he could use ~ for invert, & for AND, | for OR, and ^ for XOR. Nobody would complain. > I guess stuff like IF ELSE THEN and OR instead of XOR is what you get > when you work just to please yourself! Still, you have to recognize > the brilliance of the man. HP renamed IF ELSE THEN to IF condition THEN a ELSE b ENDIF. IF is really just syntactic sugar. This is a lot less confusing for people who don't know much about Forth, because it looks exactly like Algol. But in fact, it's just Forth in disguise. If we follow the advice above, to use C's short symbols, we could have ? : ; for IF ELSE THEN, and ; could be a general terminator of controll structures - only if it sees colon-sys it would terminate the definition as such. -- Bernd Paysan "If you want it done right, you have to do it yourself" http://bernd-paysan.de/
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-27 19:19 -0400 |
| Message-ID | <k6hq5t$s0f$1@dont-email.me> |
| In reply to | #16789 |
On 10/27/2012 4:18 PM, Bernd Paysan wrote: > rickman wrote: >> I don't think you understand me. > > Partly, probably because you insist on calling your design non-pipelined > and your parallel instruction fetch unit not a "prefetch", even though > it is. Ok, I would like to understand why you think it is a prefetch. Each instruction fetches the next sequential instruction unless the instruction is a branch, call or return in which case the appropriate instruction is fetched. No instruction is fetched ahead of time, no instruction is ever tossed away. The instruction to be executed next is always fetched. I think the issue is where I put the register in the fetch loop. Instead of putting the register between the address generation and memory, the register is at the output of the memory because the block RAMs have a register there and I don't get a choice. So I moved the register in the loop from the input of the memory to the output. This doesn't change the way the single cycle machine works, so I don't call it a prefetch. > My opinion is that the only pipelining that helps a MISC is the > instruction prefetch. You can pipeline other parts, *but it doesn't > help*. I.e. it doesn't provide better performance, because you will be > waiting for dependent instructions. The only low-hanging fruit you have > is to make the instruction fetch unit work in parallel with the other > units. The ALU and the multiplexers in a MISC are all small and fast, > and the first limit you get is the memory timing limit. Pipelining speeds up the cycle time of an operation. If the instruction fetch is not the limiting function in the cycle time of the machine then pipelining the function that is limiting the cycle time will provide a faster cycle time, although the overall performance may or may not improve. >> I am not saying pipelining is an >> important thing to use. It all depends on your requirements. I am >> simply responding to your statement that "The only pipelining that >> helps >> performance on a MISC is the instruction prefetch". I simply don't >> agree with that because what holds back your MISC design depends on >> your MISC design. BTW, my CPU design doesn't have a prefetch. It just >> fetches each instruction in parallel with the current instruction >> execution. > > This is the definition of a prefetch. You fetch the next instruction > before you complete with the current instruction. Not really before, at the same time. The instruction fetch does not depend on the current instruction unless it is a "fetch" type instruction. The conditionals check flags in FFs from the previous cycles. >> Because the instruction fetch unit is in parallel with the >> other units there is no problem with the instruction needing to be >> "prefetched". It would only be if the design were pipelined that the >> fetch would be a "prefetch". > > As you have just explained, your design is pipelined, and you do a > prefetch. A non-pipelined execution fetches current instruction and > executes it, and when done would start to fetch the next instruction. > Nothing parallel. I must have mis-stated something. The design is NOT pipelined. Everything happens in one clock cycle and is done. A pipelined operation is split over multiple cycles. My design simply moves the register in the cycle from the start of the fetch to the start of the decode. This doesn't change the way the cycle works. >> I don't intend to even try to pipeline my CPU design mainly because I >> don't have the time or interest. I prefer to try to optimize design >> tradeoffs which don't complicate the design, but should actually >> simplify it. In other words, optimizing by simplifying... much like >> what Chuck prefers. But for now I'll keep the SWAP instruction and >> will call an XOR an XOR. > > Calling an XOR and XOR is really wise, and SWAP is a useful instruction. > I don't know why Chuck wants to deliberately confuse people with his > OR=XOR and -=INVERT. If he wants to be compact, he could use ~ for > invert,& for AND, | for OR, and ^ for XOR. Nobody would complain. Well, who came up with IF ELSE THEN? That is pretty well accepted, but it is the same thing. I can live with those opcode names, -, OR, etc. I won't use them in my design, at least in part because I want my opcodes to be all text, ADD, NOP, INVERT... I just wish it was all documented better... >> I guess stuff like IF ELSE THEN and OR instead of XOR is what you get >> when you work just to please yourself! Still, you have to recognize >> the brilliance of the man. > > HP renamed IF ELSE THEN to IF condition THEN a ELSE b ENDIF. IF is > really just syntactic sugar. This is a lot less confusing for people > who don't know much about Forth, because it looks exactly like Algol. > But in fact, it's just Forth in disguise. I've never liked the term, "syntactic sugar" because it has no defined meaning. > If we follow the advice above, to use C's short symbols, we could have ? > : ; for IF ELSE THEN, and ; could be a general terminator of controll > structures - only if it sees colon-sys it would terminate the definition > as such. Ok. I don't mind typing longer words. I use VHDL when I design FPGAs, so obviously I like to type. ;7) Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-28 04:21 -0400 |
| Message-ID | <k6ipnf$28h$1@speranza.aioe.org> |
| In reply to | #16793 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6hq5t$s0f$1@dont-email.me... ... > The design is NOT pipelined. Ok. > Everything happens in one clock cycle and is done. Are there still multiple steps or operations per clock? (Yes?) E.g., an ADD instruction may have four operations: read of memory into an ALU register, read of uP (microprocessor) register to ALU register, adds ALU registers, stores ALU result register result back into uP register. > A pipelined operation is split over multiple cycles. If Intel designed an internal PLL (phase-locked loop) into their uP's, they could use a slower external clock with a faster internal clock and then claim that their CISC processor is non-pipelined ... So, yours is different how? ;-) I.e., the point being pipelining is not just splitting an operation over multiple cycles, but also executing another instruction at the same time. > My design simply moves the register in the cycle from the > start of the fetch to the start of the decode. This doesn't > change the way the cycle works. So, you should've told BP you had a single register instruction cache or a 1-element FIFO buffer... :-) Lol! Joking aside, your design is synchronous, not asynchronous, because of the FPGA, yes? Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-28 14:36 -0400 |
| Message-ID | <k6ju0o$5ib$1@dont-email.me> |
| In reply to | #16800 |
On 10/28/2012 4:21 AM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6hq5t$s0f$1@dont-email.me... > ... > >> The design is NOT pipelined. > > Ok. > >> Everything happens in one clock cycle and is done. > > Are there still multiple steps or operations per clock? (Yes?) I'm not sure what you are asking. There are any number of things going on in parallel. For example, when a stack is pushed, the pointer is incremented, the memory is written, the next instruction is fetched and the PC being used is registered. But none of this is pipelined, it all happens in one clock. > E.g., an ADD instruction may have four operations: read of memory into an > ALU register, read of uP (microprocessor) register to ALU register, adds ALU > registers, stores ALU result register result back into uP register. You are talking about a RISC or CISC machine. Instructions in a MISC are typically selected to run in one clock cycle and not take multiple steps. e.g. R> is just a pop of the return stack, a push on the data stack and a fetch of the next instruction... one clock cycle and all in parallel. >> A pipelined operation is split over multiple cycles. > > If Intel designed an internal PLL (phase-locked loop) into their uP's, they > could use a slower external clock with a faster internal clock and then > claim that their CISC processor is non-pipelined ... Really, so their processor won't take N times the instruction time to recover from a branch? > So, yours is different how? ;-) > > I.e., the point being pipelining is not just splitting an operation over > multiple cycles, but also executing another instruction at the same time. Yes, by inserting registers in the MIDDLE of a function so that it takes multiple clocks to execute, but each stage can be processing a separate instruction. >> My design simply moves the register in the cycle from the >> start of the fetch to the start of the decode. This doesn't >> change the way the cycle works. > > So, you should've told BP you had a single register instruction cache > or a 1-element FIFO buffer... :-) Lol! > > Joking aside, your design is synchronous, not asynchronous, because of the > FPGA, yes? Yes, it is synchronous. There are only three ways to create an async design, one is a custom chip, e.g. the GA144, two is an async FPGA, e.g. Achronix which pretty much only supports the large players, or three wire your own stuff from MSI chips. I won't be doing any of these unless someone with deep pockets comes along asking me to. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-29 19:06 -0400 |
| Message-ID | <k6n1ub$i0b$1@speranza.aioe.org> |
| In reply to | #16813 |
"rickman" <gnuarm@gmail.com> wrote in message news:k6ju0o$5ib$1@dont-email.me... > On 10/28/2012 4:21 AM, Rod Pemberton wrote: > > "rickman"<gnuarm@gmail.com> wrote in message > > news:k6hq5t$s0f$1@dont-email.me... ... > >> Everything happens in one clock cycle and is done. > > > > Are there still multiple steps or operations per clock? (Yes?) > > I'm not sure what you are asking. There are any number of things going > on in parallel. For example, when a stack is pushed, the pointer is > incremented, the memory is written, the next instruction is fetched and > the PC being used is registered. But none of this is pipelined, it all > happens in one clock. > I was asking if the operations are occuring sequentially per one clock - one after another. But, you just stated that they're parallel - all operations at the exact same time, roughly, in one clock. The issue is "one clock" doesn't necessarily mean parallel, to me at least... You can have a few sequential operations in one clock. E.g., parallel 1 2 3 4 - all at the same time - each no more than one clock sequential - one after another - total of one clock or less 1 - fraction of a clock 2 - fraction of a clock 3 - fraction of a clock 4 - fraction of a clock > Instructions in a MISC are typically selected to run in one clock > cycle and not take multiple steps. Ok. > > Joking aside, your design is synchronous, not asynchronous, because of > > the FPGA, yes? > > Yes, it is synchronous. There are only three ways to create an async > design, one is a custom chip, e.g. the GA144, two is an async FPGA, e.g. > Achronix which pretty much only supports the large players, or three > wire your own stuff from MSI chips. I won't be doing any of these > unless someone with deep pockets comes along asking me to. ... clock-less or self-clocking or free-running design with static ram for registers ... Are FPGA's required to be clocked? I would've assumed that it's more safe for the signals to clock, but not a requirement. Is there a tool that estimates the cycle time or frequency for your clock based on your circuit design? Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <markrobertwills@yahoo.co.uk> |
|---|---|
| Date | 2012-10-30 02:07 -0700 |
| Message-ID | <8d64f90b-b266-4c0c-9d2c-fe5d551942ea@c20g2000vbz.googlegroups.com> |
| In reply to | #16831 |
On Oct 29, 11:02 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> wrote: > I was asking if the operations are occuring sequentially per one clock - one > after another. But, you just stated that they're parallel - all operations > at the exact same time, roughly, in one clock. The issue is "one clock" > doesn't necessarily mean parallel, to me at least... > > You can have a few sequential operations in one clock. E.g., > > parallel > 1 2 3 4 - all at the same time - each no more than one clock > > sequential - one after another - total of one clock or less > 1 - fraction of a clock > 2 - fraction of a clock > 3 - fraction of a clock > 4 - fraction of a clock > Irrelevant. Even if they do happen *slightly* one-after-the-other due to gate propagation delays and the like, all the actions are atomic from all perspectives except the very circuitry on the chip itself.
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-30 09:36 -0400 |
| Message-ID | <k6oksq$34n$1@speranza.aioe.org> |
| In reply to | #16834 |
"Mark Wills" <markrobertwills@yahoo.co.uk> wrote in message news:8d64f90b-b266-4c0c-9d2c-fe5d551942ea@c20g2000vbz.googlegroups.com... > On Oct 29, 11:02 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> > wrote: ... > > [in response to statements by rickman, snipped by MW] > > > > I was asking if the operations are occuring sequentially per one clock - > > one after another. But, you just stated that they're parallel - all > > operations at the exact same time, roughly, in one clock. The issue > > is "one clock" doesn't necessarily mean parallel, to me at least... > > > > You can have a few sequential operations in one clock. E.g., > > > > parallel > > 1 2 3 4 - all at the same time - each no more than one clock > > > > sequential - one after another - total of one clock or less > > 1 - fraction of a clock > > 2 - fraction of a clock > > 3 - fraction of a clock > > 4 - fraction of a clock > > > > Irrelevant. Even if they do happen *slightly* one-after-the-other due > to gate propagation delays and the like, all the actions are atomic > from all perspectives except the very circuitry on the chip itself. Yes, it's entirely irrelevant to an end user and c.l.f. No, it's not irrelevant when discussing rickman's microprocessor design. RP
[toc] | [prev] | [next] | [standalone]
| From | Mark Wills <markrobertwills@yahoo.co.uk> |
|---|---|
| Date | 2012-10-30 08:09 -0700 |
| Message-ID | <40bc1480-4d21-4390-b48f-54298c65d165@y6g2000vbb.googlegroups.com> |
| In reply to | #16837 |
On Oct 30, 1:32 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> wrote: > "Mark Wills" <markrobertwi...@yahoo.co.uk> wrote in message > > news:8d64f90b-b266-4c0c-9d2c-fe5d551942ea@c20g2000vbz.googlegroups.com...> On Oct 29, 11:02 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> > > wrote: > > ... > > > > > > > > [in response to statements by rickman, snipped by MW] > > > > I was asking if the operations are occuring sequentially per one clock - > > > one after another. But, you just stated that they're parallel - all > > > operations at the exact same time, roughly, in one clock. The issue > > > is "one clock" doesn't necessarily mean parallel, to me at least... > > > > You can have a few sequential operations in one clock. E.g., > > > > parallel > > > 1 2 3 4 - all at the same time - each no more than one clock > > > > sequential - one after another - total of one clock or less > > > 1 - fraction of a clock > > > 2 - fraction of a clock > > > 3 - fraction of a clock > > > 4 - fraction of a clock > > > Irrelevant. Even if they do happen *slightly* one-after-the-other due > > to gate propagation delays and the like, all the actions are atomic > > from all perspectives except the very circuitry on the chip itself. > > Yes, it's entirely irrelevant to an end user and c.l.f. > > No, it's not irrelevant when discussing rickman's microprocessor design. > > RP- Hide quoted text - > > - Show quoted text - I disagree. The clock cycle is the smallest (atomic) unit of time. If he was using internal clock doublers and the like then he would have said so!
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-30 19:14 -0400 |
| Message-ID | <k6pmog$tg9$1@speranza.aioe.org> |
| In reply to | #16838 |
"Mark Wills" <markrobertwills@yahoo.co.uk> wrote in message news:40bc1480-4d21-4390-b48f-54298c65d165@y6g2000vbb.googlegroups.com... > On Oct 30, 1:32 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> > wrote: > > "Mark Wills" <markrobertwi...@yahoo.co.uk> wrote in message > news:8d64f90b-b266-4c0c-9d2c-fe5d551942ea@c20g2000vbz.googlegroups.com... > > > On Oct 29, 11:02 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm> > > > wrote: ... > > > > [in response to statements by rickman, snipped by MW] > > > > > I was asking if the operations are occuring sequentially per one clock - > > > > one after another. But, you just stated that they're parallel - all > > > > operations at the exact same time, roughly, in one clock. The issue > > > > is "one clock" doesn't necessarily mean parallel, to me at least... > > > > > You can have a few sequential operations in one clock. E.g., > > > > > parallel > > > > 1 2 3 4 - all at the same time - each no more than one clock > > > > > sequential - one after another - total of one clock or less > > > > 1 - fraction of a clock > > > > 2 - fraction of a clock > > > > 3 - fraction of a clock > > > > 4 - fraction of a clock > > > > Irrelevant. Even if they do happen *slightly* one-after-the-other due > > > to gate propagation delays and the like, all the actions are atomic > > > from all perspectives except the very circuitry on the chip itself. > > > Yes, it's entirely irrelevant to an end user and c.l.f. > > > No, it's not irrelevant when discussing rickman's microprocessor design. > > > I disagree. Ok. > The clock cycle is the smallest (atomic) unit of time. No, it depends on the design. > If he was using internal clock doublers and the like > then he would have said so! You're assuming a clock is needed to drive each stage of a sequential sequence operations for an instruction. That's not the case. Sometimes, the internal circuits of a microprocessor are self-clocking or self-propagating. Once one step is completed, the next starts. In such designs, a clock is only required at the start or end of the instruction sequence to ensure a stable, or latched state. Of course, this is older microprocessor theory, not FPGA microprocessor theory. So, it's entirely possible microprocessor using an FGPA may not be able to be designed that way. "Rather than totally removing the clock signal, some CPU designs allow certain portions of the device to be asynchronous, such as using asynchronous ALUs in conjunction with superscalar pipelining to achieve some arithmetic performance gains." http://en.wikipedia.org/wiki/Central_processing_unit#Clock_rate "Unlike a conventional processor, a clockless processor (asynchronous CPU) has no central clock to coordinate the progress of data through the pipeline." http://en.wikipedia.org/wiki/Asynchronous_circuit#Asynchronous_CPU Rod Pemberton
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-30 18:50 -0400 |
| Message-ID | <k6pljm$df7$1@dont-email.me> |
| In reply to | #16831 |
On 10/29/2012 7:06 PM, Rod Pemberton wrote: > "rickman"<gnuarm@gmail.com> wrote in message > news:k6ju0o$5ib$1@dont-email.me... >> On 10/28/2012 4:21 AM, Rod Pemberton wrote: >>> "rickman"<gnuarm@gmail.com> wrote in message >>> news:k6hq5t$s0f$1@dont-email.me... > ... > >>>> Everything happens in one clock cycle and is done. >>> >>> Are there still multiple steps or operations per clock? (Yes?) >> >> I'm not sure what you are asking. There are any number of things going >> on in parallel. For example, when a stack is pushed, the pointer is >> incremented, the memory is written, the next instruction is fetched and >> the PC being used is registered. But none of this is pipelined, it all >> happens in one clock. >> > > I was asking if the operations are occuring sequentially per one clock - one > after another. But, you just stated that they're parallel - all operations > at the exact same time, roughly, in one clock. The issue is "one clock" > doesn't necessarily mean parallel, to me at least... > > You can have a few sequential operations in one clock. E.g., > > parallel > 1 2 3 4 - all at the same time - each no more than one clock > > sequential - one after another - total of one clock or less > 1 - fraction of a clock > 2 - fraction of a clock > 3 - fraction of a clock > 4 - fraction of a clock > >> Instructions in a MISC are typically selected to run in one clock >> cycle and not take multiple steps. > > Ok. > >>> Joking aside, your design is synchronous, not asynchronous, because of >>> the FPGA, yes? >> >> Yes, it is synchronous. There are only three ways to create an async >> design, one is a custom chip, e.g. the GA144, two is an async FPGA, e.g. >> Achronix which pretty much only supports the large players, or three >> wire your own stuff from MSI chips. I won't be doing any of these >> unless someone with deep pockets comes along asking me to. > > ... clock-less or self-clocking or free-running design with static ram for > registers ... > > Are FPGA's required to be clocked? I would've assumed that it's more safe > for the signals to clock, but not a requirement. I can't say I understand the question. If a register isn't clocked, how does it work? Are you asking if the FPGA uses clocked registers or async registers with set/reset controls? Many years ago it was found that it was hard to produce a design using async logic and even harder to prove that it worked properly, especially regarding timing. Partly in order to make simulation easier and to make timing analysis easier, they came up with LSSD (Level Sensitive Static Design). I won't go into what that entails but it mostly means all registers use the data and clock enable inputs and a Q output. All logic connects between those points and/or the I/Os. Async FF inputs are not used. If you do this, the logic can be simulated in what is called "unit delay" simulation where it is assumed that all signals meet setup and hold timing at each FF and only the logic needs to be considered. The timing can be analyzed separately with what is called "static timing analysis". This has been the stalwart of FPGA and ASIC design for many years. Async design like the GA144 or a very small number of other devices are very distant outliers. There are few tools for designing async devices and very little foundry support. > Is there a tool that estimates the cycle time or frequency for your clock > based on your circuit design? In an FPGA, yes, this is called static timing analysis. You typically give it a clock rate, in fact, the clock rate is provided to the tools doing the place and route as well, and other timing constraints (clock to a given input or output for example). The tool tells you if your design meets the goals, which signal fail and an analysis of the detailed timing path. I guess I've been assuming I was talking to someone who knew a bit about this. I understand some of the misunderstandings in our conversation now. Rick
[toc] | [prev] | [next] | [standalone]
| From | daveyrotten <danw8804@gmail.com> |
|---|---|
| Date | 2012-10-31 07:37 -0700 |
| Message-ID | <3289bb1a-2b62-4b5e-b32f-55887176d7b9@googlegroups.com> |
| In reply to | #16846 |
> > > > I can't say I understand the question. If a register isn't clocked, how > > does it work? Are you asking if the FPGA uses clocked registers or > > async registers with set/reset controls? > > > > Many years ago it was found that it was hard to produce a design using > > async logic and even harder to prove that it worked properly, especially > > regarding timing. Partly in order to make simulation easier and to make > > timing analysis easier, they came up with LSSD (Level Sensitive Static > > Design). I won't go into what that entails but it mostly means all > > registers use the data and clock enable inputs and a Q output. All logic > > connects between those points and/or the I/Os. Async FF inputs are not > > used. > > > > If you do this, the logic can be simulated in what is called "unit > > delay" simulation where it is assumed that all signals meet setup and > > hold timing at each FF and only the logic needs to be considered. The > > timing can be analyzed separately with what is called "static timing > > analysis". This has been the stalwart of FPGA and ASIC design for many > > years. > > > > Async design like the GA144 or a very small number of other devices are > > very distant outliers. There are few tools for designing async devices > > and very little foundry support. > > > > > > > Is there a tool that estimates the cycle time or frequency for your clock > > > based on your circuit design? > > > > In an FPGA, yes, this is called static timing analysis. You typically > > give it a clock rate, in fact, the clock rate is provided to the tools > > doing the place and route as well, and other timing constraints (clock > > to a given input or output for example). The tool tells you if your > > design meets the goals, which signal fail and an analysis of the > > detailed timing path. > > > > I guess I've been assuming I was talking to someone who knew a bit about > > this. I understand some of the misunderstandings in our conversation now. > > > > Rick Nice to see so much discussion of FPGAs, and especially implementing soft core Forth-based CPUs like the b16 and j1 in FPGAs. I think approaches like this are the most promising arena for the application of Forth at the moment. With a $55 Xula-200 board, a flash memory Pmod board, keyboard, and VGA monitor you can make a complete Forth computer in which you have control of everything. I've done it and it's a lot of fun to use. No, it's not ANS Forth compliant (yet) but still a blast (and best of all completely independant of Intel, PIC, Microsoft, or anybody else). With all due respect, however, I don't get too excited about all the (what seems to me) nitpicking about which architecture is slightly faster than which other one. It seems to me that any processor which executes Forth opcodes directly as primitives, as the b16 and j1 do, has got to be an order of magnitude faster than a Forth system built on top of another language (even assembly language). Of course, in an FPGA you can hook in custom hardware anywhere you want to speed up the system. I would rather just build stuff than worry endlessly about whether it's faster than the next guy's approach. But that's just me.
[toc] | [prev] | [next] | [standalone]
| From | stephenXXX@mpeforth.com (Stephen Pelc) |
|---|---|
| Date | 2012-10-31 15:27 +0000 |
| Message-ID | <50914182.96209701@192.168.0.50> |
| In reply to | #16862 |
On Wed, 31 Oct 2012 07:37:18 -0700 (PDT), daveyrotten <danw8804@gmail.com> wrote: >It seems to me that any processor which executes Forth opcod= >es directly as primitives, as the b16 and j1 do, has got to be an order of >magnitude faster than a Forth system built on top of another language (even >assembly language). That's what most implementers thought before we started getting good at writing optimising Forth compilers. For CPU architectures that are HLL friendly, optimised Forth code can be shorter than the equivalent threaded code and will probably run at least ten times faster. Where silicon machines may score is in their low use of silicon resources. Stephen -- Stephen Pelc, stephenXXX@mpeforth.com MicroProcessor Engineering Ltd - More Real, Less Time 133 Hill Lane, Southampton SO15 5AF, England tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691 web: http://www.mpeforth.com - free VFX Forth downloads
[toc] | [prev] | [next] | [standalone]
| From | daveyrotten <danw8804@gmail.com> |
|---|---|
| Date | 2012-10-31 08:59 -0700 |
| Message-ID | <cd851925-0dbf-454d-a6f7-8cf75815bffa@googlegroups.com> |
| In reply to | #16868 |
On Wednesday, October 31, 2012 10:30:03 AM UTC-5, Stephen Pelc wrote: > On Wed, 31 Oct 2012 07:37:18 -0700 (PDT), daveyrotten wrote > > > > > >It seems to me that any processor which executes Forth opcod= > > >es directly as primitives, as the b16 and j1 do, has got to be an order of > > >magnitude faster than a Forth system built on top of another language (even > > >assembly language). > > > > That's what most implementers thought before we started getting > > good at writing optimising Forth compilers. For CPU architectures > > that are HLL friendly, optimised Forth code can be shorter than > > the equivalent threaded code and will probably run at least ten > > times faster. > > > > Where silicon machines may score is in their low use of silicon > > resources. > > > > Stephen > > > > -- > > Stephen Pelc, stephenXXX@mpeforth.com > > MicroProcessor Engineering Ltd - More Real, Less Time > > 133 Hill Lane, Southampton SO15 5AF, England > > tel: +44 (0)23 8063 1441, fax: +44 (0)23 8033 9691 > > web: http://www.mpeforth.com - free VFX Forth downloads I should probably apologize in advance for not knowing this, but by 'threaded code' do you mean that the execution of each word requires a dictionary search? Because my machine does not normally use the dictionary when executing words, only when compiling new words.
[toc] | [prev] | [next] | [standalone]
| From | Andrew Haley <andrew29@littlepinkcloud.invalid> |
|---|---|
| Date | 2012-10-31 11:18 -0500 |
| Message-ID | <arWdnZGOE5fQ0gzNnZ2dnUVZ8o-dnZ2d@supernews.com> |
| In reply to | #16871 |
daveyrotten <danw8804@gmail.com> wrote: > I should probably apologize in advance for not knowing this, but by > 'threaded code' do you mean that the execution of each word requires > a dictionary search? Because my machine does not normally use the > dictionary when executing words, only when compiling new words. No. http://en.wikipedia.org/wiki/Threaded_code Andrew.
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-31 13:49 -0400 |
| Message-ID | <k6robv$ck3$1@dont-email.me> |
| In reply to | #16871 |
On 10/31/2012 11:59 AM, daveyrotten wrote: > > I should probably apologize in advance for not knowing this, but by 'threaded code' do you mean that the execution of each word requires a dictionary search? Because my machine does not normally use the dictionary when executing words, only when compiling new words. Threaded code is the standard way of compiling Forth code into something the CPU can execute. There are several ways to thread code, but in all cases the treading adds overhead, even subroutine threaded code which uses subroutine calls to invoke word definitions. In highly optimized code some of the words are implemented with inline machine code so that it is as fast as assembly language coding potentially. The limitation of compiler optimization is that it has to be implemented for each CPU design. Stephen has done a great job optimizing the code for the x86 architecture, but I don't think you will find much compiler optimization being done for FPGA CPUs. Rick
[toc] | [prev] | [next] | [standalone]
| From | rickman <gnuarm@gmail.com> |
|---|---|
| Date | 2012-10-31 13:43 -0400 |
| Message-ID | <k6ro0m$9v7$1@dont-email.me> |
| In reply to | #16868 |
On 10/31/2012 11:27 AM, Stephen Pelc wrote: > On Wed, 31 Oct 2012 07:37:18 -0700 (PDT), daveyrotten > <danw8804@gmail.com> wrote: > >> It seems to me that any processor which executes Forth opcod= >> es directly as primitives, as the b16 and j1 do, has got to be an order of >> magnitude faster than a Forth system built on top of another language (even >> assembly language). > > That's what most implementers thought before we started getting > good at writing optimising Forth compilers. For CPU architectures > that are HLL friendly, optimised Forth code can be shorter than > the equivalent threaded code and will probably run at least ten > times faster. > > Where silicon machines may score is in their low use of silicon > resources. > > Stephen And power consumption. Using 600 LUTs in an FPGA will use less power than most ARM processors and run at similar speeds. I hope to benchmark some power consumption in the near future. The other issue is the cost of the processor. Using a processor and an FPGA is more expensive than just using an FPGA... of course with the processor chip you get a few things for free, a hunk of easy to use memory, a number of peripherals and various power controls. As to Davey's concern that we are "nitpicking about which architecture is slightly faster", I just found a 2x speedup with some instruction optimization in a VLIW MISC approach. The trick with making any processor fast is to find ways to use all of the execution units most of the time. The standard MISC approach has very tiny instructions that do very tiny operations. I think a VLIW approach may help to more fully utilize the processor for faster and more efficient execution. Rick
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <do_not_have@notemailnotz.cnm> |
|---|---|
| Date | 2012-10-31 12:03 -0400 |
| Subject | Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] |
| Message-ID | <k6rht8$57r$1@speranza.aioe.org> |
| In reply to | #16862 |
"daveyrotten" <danw8804@gmail.com> wrote in message news:3289bb1a-2b62-4b5e-b32f-55887176d7b9@googlegroups.com... ... > With a $55 Xula-200 board, a flash memory Pmod board, keyboard, > and VGA monitor you can make a complete Forth computer in > which you have control of everything. Is that a good buy? E.g., You can get all this for $39.99 USD: 512Mb Raspberry-Pi with 700Mhz ARM cpu, MPEG-2 and OpenGL GPU, and 10/100 Ethernet http://blog.makezine.com/2012/10/29/maker-shed-now-shipping-512mb-raspberry-pis Specifications and purchasing: http://www.makershed.com/Raspberry_Pi_Model_B_Revision_2_512MB_p/mkrpi2.htm Rod Pemberton
[toc] | [prev] | [next] | [standalone]
Page 2 of 9 — ← Prev page 1 [2] 3 4 5 6 7 8 9 Next page →
Back to top | Article view | comp.lang.forth
csiph-web