Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #16579 > unrolled thread

RTX2000 optimization

Started byBrad Eckert <hwfwguy@gmail.com>
First post2012-10-22 08:53 -0700
Last post2012-10-23 10:52 +0000
Articles 20 on this page of 172 — 22 participants

Back to article view | Back to comp.lang.forth


Contents

  RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
    Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
    Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
      Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
        Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
        Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
        Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
          Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
            Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
              Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
                Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
                  Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
              Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
                Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
                    Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
                            Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
                                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
                                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
                                    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
                                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
                                        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
                                          Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
                                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
                                              Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
                                                Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
                                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
                                            Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
                                              Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
                                                Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
                                                  Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
                                                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
                                                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
                                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
                                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
                                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
                                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
                                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
                                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
                                                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
                                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
                                                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
                                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
                                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
                                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
                                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
                                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
                                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
                                              Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
                                                Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
                                                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
                                                  Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
                                                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
                                                      Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
                                                    Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
                                                      Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
                                                        Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
                                                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
                                                      Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
                                                      Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
                                          Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
                                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
                                              Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
                                                Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
                                                  Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
                                                  Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
                                                    Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
                                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
                                    Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
                                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
                                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
                                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
                                  Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
                                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
          Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
    Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
      Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
        Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
          Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
          Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
              Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
                  Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
                      Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
                        Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
                  Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
              Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
              Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
            addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
                Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
                    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
                        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
    Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
        Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
            Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
        Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
          Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
            Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
              Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
                Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
                  Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
                    Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
              Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
    Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000

Page 6 of 9 — ← Prev page 1 2 3 4 5 [6] 7 8 9  Next page →


#16955

Fromdaveyrotten <danw8804@gmail.com>
Date2012-11-01 12:14 -0700
Message-ID<3837ed9d-dfb2-48d6-b356-2279ba093f19@googlegroups.com>
In reply to#16954
On Thursday, November 1, 2012 2:12:10 PM UTC-5, daveyrotten wrote:
> On Thursday, November 1, 2012 12:59:54 PM UTC-5, Brad Eckert wrote:
> 
> > 
> 
> > The horizontal encoding style (Novix, J1) is compelling. It needs a better compiler if you want to avoid assembly, but as the J1 showed, using assembly in a really simple processor isn't complex. Besides, portable code isn't as important. It's not like your CPU can be "end-of-life"ed.
> 
> > 
> 
> I don't understand what you mean. Are you talking about compiling from C or from Forth?  What are you considering assembly for the J1?
> 
> 
> 
> Does the following code use assembly for example?
> 
> 
> 
> : !  +! drop ;
> 
> 
> 
> This is how most of my programs for the J1 start.  Sorry if this is too pedestrian. I'm just trying to understand. I think much of the problem is getting used to the terms. I'm a hardware guy delving into the software world. Probably that's a dangerous thing.
> 
> 
> 
> I very much agree with the last statement in your paragraph above, however. That's a big part of what's really cool about simple soft core cpus in FPGAs: obsolesence is really not an issue; at least as long as synthesis from RTL is available in the tools.

Sorry, I meant:

: !  !+ drop ;

[toc] | [prev] | [next] | [standalone]


#16985

FromBrad Eckert <hwfwguy@gmail.com>
Date2012-11-02 10:51 -0700
Message-ID<c05e1b4e-66c0-4695-bdce-bf271fdd6557@googlegroups.com>
In reply to#16954
On Thursday, November 1, 2012 12:12:10 PM UTC-7, daveyrotten wrote:
> I don't understand what you mean. Are you talking about compiling from C or from Forth?  What are you considering assembly for the J1?
>
: drop      N                   d-1 alu ;  

> Does the following code use assembly for example?
> : !  +! drop ;
> 

No. In fact, your application may be able to do without assembly. But then when you need to speed something up, assembly is available. The rule of thumb in programming is get it working, then make it pretty.

In this example, ! would compile as a call to +! followed by a jump to drop. Drop is written in assembly. +! is high level code. Both are library stuff that you wouldn't mess with unless you changed the CPU. Which is a possibility, since the J1 can be tweaked in so many ways.

The cool thing about the J1's compiler/disassembler (crossj1.fs) is that you can understand it. I don't think you'll find a comparable C cross compiler in a single 10KB text file.

[toc] | [prev] | [next] | [standalone]


#16988

Fromdaveyrotten <danw8804@gmail.com>
Date2012-11-02 11:28 -0700
Message-ID<7e73a19e-7205-449f-8284-95fa179a79ca@googlegroups.com>
In reply to#16985
On Friday, November 2, 2012 12:51:10 PM UTC-5, Brad Eckert wrote:

> The cool thing about the J1's compiler/disassembler (crossj1.fs) is that you can understand it. I don't think you'll find a comparable C cross compiler in a single 10KB text file.

The other cool thing is that you can just write your own compiler, which is what I did. I guess that explains part of my confusion: I am not familiar yet with 
the contents of crossj1.fs.  Thanks for the info though.

[toc] | [prev] | [next] | [standalone]


#16953

Fromrickman <gnuarm@gmail.com>
Date2012-11-01 14:58 -0400
Message-ID<k6ugp9$lae$1@dont-email.me>
In reply to#16919
On 10/31/2012 8:08 PM, Rod Pemberton wrote:
> "Brad Eckert"<hwfwguy@gmail.com>  wrote in message
> news:54a7ded5-b4a1-4ecd-a025-6a39b7e51487@googlegroups.com...
>> On Monday, October 29, 2012 4:02:36 PM UTC-7, Rod Pemberton wrote:
> ...
>
> Hey, did anyone ever answer your initial question?  The thread has kind of
> been taken over by rickman...
>
> With a list of primitives and the processor instruction set and the
> processor's stack/register model, an optimal sequences can be designed.
>
> I know Alex McDonald is working on an optimizer.  I've been using a private,
> heavily modified version of Peter Sovietov's "Forth Wizard".  It's in
> Javascript and fairly easy to modify, if you know C.
>
> The original is here.
> http://forpost.sourceforge.net/forthwiz.html
>
> Koopman's chapter lists some of RTX 2000 instruction set:
> http://www.ece.cmu.edu/~koopman/stack_computers/sec4_5.html
>
> It appears that they are expected Forth sequences, perhaps output from a
> compiler.  I can't see an individual using them ...  Based on sequences for
> various Forth words, I can take guess at how some of them are used:
>
> -possibly used for +! like operation
> DUP @ SWAP
>
> -possibly enhanced DROP sequence:
> DROP nn inv
> DROP lit inv
> DROP inv shift
> nn G@ DROP inv
>
> -possibly enhanced DIP sequence:
> DROP DUP inv shift
>
> -possibly enhanced DUP sequence:
> DUP inv shift
> DUP lit op
> DUP nn G@ inv
>
> -possibly enhanced OVER sequence:
> nn OVER op
> OVER inv shift
> nn G@ OVER op
>
> -possibly enhanced NUP sequence:
> OVER SWAP op shift
> OVER SWAP ! inv
> OVER SWAP ! nn
> OVER SWAP @ op
>
> -possibly enhanced SWAP sequence:
> lit SWAP inv
> lit SWAP op
> nn SWAP op
> SWAP inv shift
> nn G@ SWAP op
>
> -possibly enhanced NIP sequence:
> SWAP DROP inv shift
> SWAP DROP @ nn
>
> -possibly enhanced DRIP sequence:
> SWAP DROP DUP inv shift
> SWAP DROP DUP @ nn ROT op
> SWAP DROP DUP @ SWAP
>
> -possibly enhanced TUCK sequence:
> SWAP OVER op shift
> SWAP OVER !

The RTX2000 is VLIW in the sense that the instructions have many fields 
that control the hardware directly.  So there are many combinations that 
may or may not be useful.  The documentation lists them all and it is up 
to you to choose which ones you wish to use.


>>> Are FPGA's required to be clocked?  I would've assumed that it's more
>>> safe for the signals to clock, but not a requirement.
>>>
>> Pretty much. Synchronous design is generally assumed. You might be able
>> to build a tool that converts your async CPU design (in your own async
>> design language) to a combination of Verilog/VHDL and timing constraint
>> files for a particular flow. Such a language may already exist as a
>> research project. My impression is that the tools have molded the design
>> process, leaving synchronous design the only practical game in town.
>
> rickman said basically the same thing in one of his replies.
>
>
> What I was trying to get to with rickman, was what I quoted for Andrew:
>
> "Rather than totally removing the clock signal, some CPU designs allow
> certain portions of the device to be asynchronous, such as using
> asynchronous ALUs in conjunction with superscalar pipelining to achieve
> some arithmetic performance gains."
> http://en.wikipedia.org/wiki/Central_processing_unit#Clock_rate
>
> I.e., that "one clock" can represent multiple, sequential, internal
> operations, it doesn't necessarily mean parallel or pipelined.

What it "can" represent is not what is typically meant.  If the 
operations of a unit can be run at speeds much faster than the clock, 
why no just use a faster clock and get better control over them in a 
single clock cycle?

In the really "old" days, they used to use multiple phased clocks.  This 
was because they used pass transistors to make up FFs and only one could 
be switched on at a time to preclude the logic shooting through to the 
next stage.

There are all sorts of things that could be done.  But often they are 
more complicated and make things harder to design.  No small part of 
conventional technology has to do with how easy it is to use.

But then there can be multiple, sequential operations in one clock 
cycle.  They just don't register the intermediate steps.  Like read a 
register, invert the data, add it to another register, shift the result 
and save it away.

Rick

[toc] | [prev] | [next] | [standalone]


#16805

FromBernd Paysan <bernd.paysan@gmx.de>
Date2012-10-28 14:45 +0100
Message-ID<2071965.HC6Hy0jY3M@sunwukong.fritz.box>
In reply to#16793
rickman wrote:
>> Partly, probably because you insist on calling your design
>> non-pipelined and your parallel instruction fetch unit not a
>> "prefetch", even though it is.
> 
> Ok, I would like to understand why you think it is a prefetch.  Each
> instruction fetches the next sequential instruction unless the
> instruction is a branch, call or return in which case the appropriate
> instruction is fetched.  No instruction is fetched ahead of time, no
> instruction is ever tossed away.  The instruction to be executed next
> is always fetched.

That's a prefetch.  Each instruction fetches the *next* instruction, 
ahead of time.  Pre=next.  I don't know why you are confused.

> I think the issue is where I put the register in the fetch loop.
> Instead of putting the register between the address generation and
> memory, the register is at the output of the memory because the block
> RAMs have a register there and I don't get a choice.  So I moved the
> register in the loop from the input of the memory to the output.  This
> doesn't change the way the single cycle machine works, so I don't call
> it a prefetch.

If you try to communicate, call things the way they are.  If it looks 
like a duck, quacks like a duck, swims like a duck, walks like a duck, 
flies like a duck, it is called "duck".  If you fetch the next 
instruction while you are still busy with the current instruction, it is 
a prefetch.

>> This is the definition of a prefetch.  You fetch the next instruction
>> before you complete with the current instruction.
> 
> Not really before, at the same time.

The diagram of a two-stage pipeline is

fetch | execute |
        fetch   | execute |

You do the fetch *before* you are *done* with execute (my words).  So 
*concurrent* with execute (your words).  Can you transform these two 
identical statements into each other so we can agree that we agree?

> The instruction fetch does not
> depend on the current instruction unless it is a "fetch" type
> instruction.

Yes.  That's why prefetch is a low-hanging fruit, and done most of the 
time.  As your example shows.

> I must have mis-stated something.  The design is NOT pipelined.
> Everything happens in one clock cycle and is done.  A pipelined
> operation is split over multiple cycles.

Well it is, according to your description of what happens.  See above.  
You fetch the instruction in one cycle and execute it in the next cycle.  
"one" + "next" = 2.

The confusion might be that you think of an instruction as doing two 
things:

a) execute itself
b) fetch the next instruction

These two mostly independent operations happen in parallel.  The 
conceptual entity however is *one* instruction, not two of them.  One 
instruction is executed by

a) fetching it
b) executing it

This is a sequential process, taking two cycles.

Your mental model as above makes it easier for you to deal with the 
paralleism, and make sure there are no pipeline stalls in any case.  No 
argument about that.  This is just a discussion about terms.

> My design simply moves the
> register in the cycle from the start of the fetch to the start of the
> decode.  This doesn't change the way the cycle works.

It moves the entire fetch operation by one cycle - start of decode is 
end of fetch.

> I've never liked the term, "syntactic sugar" because it has no defined
> meaning.

Of course "syntactic sugar" has a defined meaning.  It is a syntactic 
element that makes the program easier to read, while having no function 
of its own:

http://en.wikipedia.org/wiki/Syntactic_sugar

I.e. to get HP's syntax, you define

: endif  postpone then ; immediate
: then   postpone if ; immediate
: if     ; immediate

and then you can write something like

: simple. ( n -- ) if dup 3 > then drop ." many " else . endif ;
\ numbers are 1 2 3 many...

IF in this case is obvious syntactic sugar, as you can leave it away 
without changing the functionality of the program.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#16814

Fromrickman <gnuarm@gmail.com>
Date2012-10-28 15:04 -0400
Message-ID<k6jvkc$gk2$1@dont-email.me>
In reply to#16805
On 10/28/2012 9:45 AM, Bernd Paysan wrote:
> rickman wrote:
>>> Partly, probably because you insist on calling your design
>>> non-pipelined and your parallel instruction fetch unit not a
>>> "prefetch", even though it is.
>>
>> Ok, I would like to understand why you think it is a prefetch.  Each
>> instruction fetches the next sequential instruction unless the
>> instruction is a branch, call or return in which case the appropriate
>> instruction is fetched.  No instruction is fetched ahead of time, no
>> instruction is ever tossed away.  The instruction to be executed next
>> is always fetched.
>
> That's a prefetch.  Each instruction fetches the *next* instruction,
> ahead of time.  Pre=next.  I don't know why you are confused.

Aren't the instructions always fetched ahead of the decode?  No, 
sometimes they are fetched in parallel of the decode and execute, but 
wait, that is not the *next* instruction, that is the instruction that 
will execute after the instruction you just fetched which is not yet 
executing... why, because the CPU is pipelined.


>> I think the issue is where I put the register in the fetch loop.
>> Instead of putting the register between the address generation and
>> memory, the register is at the output of the memory because the block
>> RAMs have a register there and I don't get a choice.  So I moved the
>> register in the loop from the input of the memory to the output.  This
>> doesn't change the way the single cycle machine works, so I don't call
>> it a prefetch.
>
> If you try to communicate, call things the way they are.  If it looks
> like a duck, quacks like a duck, swims like a duck, walks like a duck,
> flies like a duck, it is called "duck".  If you fetch the next
> instruction while you are still busy with the current instruction, it is
> a prefetch.
>
>>> This is the definition of a prefetch.  You fetch the next instruction
>>> before you complete with the current instruction.
>>
>> Not really before, at the same time.
>
> The diagram of a two-stage pipeline is
>
> fetch | execute |
>          fetch   | execute |
>
> You do the fetch *before* you are *done* with execute (my words).  So
> *concurrent* with execute (your words).  Can you transform these two
> identical statements into each other so we can agree that we agree?
>
>> The instruction fetch does not
>> depend on the current instruction unless it is a "fetch" type
>> instruction.
>
> Yes.  That's why prefetch is a low-hanging fruit, and done most of the
> time.  As your example shows.
>
>> I must have mis-stated something.  The design is NOT pipelined.
>> Everything happens in one clock cycle and is done.  A pipelined
>> operation is split over multiple cycles.
>
> Well it is, according to your description of what happens.  See above.
> You fetch the instruction in one cycle and execute it in the next cycle.
> "one" + "next" = 2.
>
> The confusion might be that you think of an instruction as doing two
> things:
>
> a) execute itself
> b) fetch the next instruction
>
> These two mostly independent operations happen in parallel.  The
> conceptual entity however is *one* instruction, not two of them.  One
> instruction is executed by
>
> a) fetching it
> b) executing it
>
> This is a sequential process, taking two cycles.
>
> Your mental model as above makes it easier for you to deal with the
> paralleism, and make sure there are no pipeline stalls in any case.  No
> argument about that.  This is just a discussion about terms.
>
>> My design simply moves the
>> register in the cycle from the start of the fetch to the start of the
>> decode.  This doesn't change the way the cycle works.
>
> It moves the entire fetch operation by one cycle - start of decode is
> end of fetch.

I don't agree.  Prefetch is fetching instructions ahead of time when you 
are pipelined.  The result is that when the code has a branch the 
instruction fetched has to be thrown away.

By your definition all processors prefetch because the instruction is 
fetched BEFORE it is decoded or executed.

The classic CPU model has a PC register which is the start of each 
instruction: fetch, decode, execute with no registers in between.  In my 
design I "rotated" the registers a bit.  This is because I have to have 
an instruction register because the memory has this.  So the instruction 
cycle is...  decode, execute, fetch.  No pipelining and I don't see how 
this makes the fetch a "prefetch".  The fetch is done on the next 
instruction to be executed, not an instruction that is *anticipated* to 
be fetched.

If you insist on calling this a prefetch, it adds nothing to the 
conversation because the results that typically are connected to a 
prefetch are not attached to my design.


>> I've never liked the term, "syntactic sugar" because it has no defined
>> meaning.
>
> Of course "syntactic sugar" has a defined meaning.  It is a syntactic
> element that makes the program easier to read, while having no function
> of its own:
>
> http://en.wikipedia.org/wiki/Syntactic_sugar

Wikipedia doesn't define the language.  Regardless, calling something 
syntactic sugar has no utility to me.


> I.e. to get HP's syntax, you define
>
> : endif  postpone then ; immediate
> : then   postpone if ; immediate
> : if     ; immediate
>
> and then you can write something like
>
> : simple. ( n -- ) if dup 3>  then drop ." many " else . endif ;
> \ numbers are 1 2 3 many...
>
> IF in this case is obvious syntactic sugar, as you can leave it away
> without changing the functionality of the program.
>

Unfortunately you have snipped the context of the syntactic sugar 
statement so I have no idea what this is now about.

Rick

[toc] | [prev] | [next] | [standalone]


#16817

FromBernd Paysan <bernd.paysan@gmx.de>
Date2012-10-28 21:09 +0100
Message-ID<7510467.F46bDX5XoM@sunwukong.fritz.box>
In reply to#16814
rickman wrote:
>> It moves the entire fetch operation by one cycle - start of decode is
>> end of fetch.
> 
> I don't agree.  Prefetch is fetching instructions ahead of time when
> you are pipelined.  The result is that when the code has a branch the
> instruction fetched has to be thrown away.
> 
> By your definition all processors prefetch because the instruction is
> fetched BEFORE it is decoded or executed.

No, please, try to understand what you are doing.  A processor which 
doesn't prefetch does

fetch | execute | fetch | execute

This is a non-pipelined processor.  Each stage happens on its own, no 
overlap.  Your processor does

cyc1  | cyc2    | cyc3    | cyc4
fetch | exeucte |
        fetch   | execute
                  fetch   | execute

> The fetch is done on the next
> instruction to be executed, not an instruction that is *anticipated*
> to be fetched.

That's perfectly fine, in a 2-stage pipeline, you don't need to 
anticipate anything.  You are probably confusing this with a classical 
four-stage pipeline design where you have

fetch | decode+r | execute  | write
        fetch    | decode+r | execute | write

In this case, if you execute a branch, you have already fetched *one* 
instruction after the branch.  This one needs to be thrown away.

With a two-stage pipeline, you don't need to throw anything away.  I.e. 
you always fetch the rigth instruction.

> If you insist on calling this a prefetch, it adds nothing to the
> conversation because the results that typically are connected to a
> prefetch are not attached to my design.

The starting point of this lengthy conversation was that I said that 
MISC processors typically use prefetch, but no other optimization.  This 
is what you opposed to, but the only disagreement we have is over 
language, not over actual facts.  Fact is that you parallelize fetch and 
execute, which *is* the low-hanging fruit of optimization I was talking 
about.

If we disagree about the term ("instruction prefetch" is pretty 
generic), then well.  Let's say fetching the next instruction is a 
degenerated case of prefetch, because the instruction arrives just when 
you actually need it, i.e. no need for queuing it up.

There are single-cycle MISC CPUs which do not prefetch at all, e.g. the 
J1 CPU.  There, the cycle consists of fetching the instruction, and then 
asynchronously decoding it and selecting the result of the corresponding 
operations (which have completed by the time the instruction is 
fetched).  It's not the next instrcution which is fetched, it's the 
*current* instruction.

>>> I've never liked the term, "syntactic sugar" because it has no
>>> defined meaning.
>>
>> Of course "syntactic sugar" has a defined meaning.  It is a syntactic
>> element that makes the program easier to read, while having no
>> function of its own:
>>
>> http://en.wikipedia.org/wiki/Syntactic_sugar
> 
> Wikipedia doesn't define the language.  Regardless, calling something
> syntactic sugar has no utility to me.

*Nobody* defines language.  Language is made up by *usage*, and usage 
can be *observed*.  An encyclopedia is a collection of observed usage.  
Languages work by common understanding and using the same terms for the 
same things; if you are in doubt that the way you use your terms are 
understood by others, get a life and check reality.  Wikipedia can be a 
good start.

>> I.e. to get HP's syntax, you define
>>
>> : endif  postpone then ; immediate
>> : then   postpone if ; immediate
>> : if     ; immediate
>>
>> and then you can write something like
>>
>> : simple. ( n -- ) if dup 3>  then drop ." many " else . endif ;
>> \ numbers are 1 2 3 many...
>>
>> IF in this case is obvious syntactic sugar, as you can leave it away
>> without changing the functionality of the program.
>>
> 
> Unfortunately you have snipped the context of the syntactic sugar
> statement so I have no idea what this is now about.

The context of syntactic sugar was HP's modified IF THEN ELSE ENDIF, as 
you can read here.  The context is still there.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#16819

Fromrickman <gnuarm@gmail.com>
Date2012-10-28 16:59 -0400
Message-ID<k6k6cj$q98$1@dont-email.me>
In reply to#16817
On 10/28/2012 4:09 PM, Bernd Paysan wrote:
> rickman wrote:
>>> It moves the entire fetch operation by one cycle - start of decode is
>>> end of fetch.
>>
>> I don't agree.  Prefetch is fetching instructions ahead of time when
>> you are pipelined.  The result is that when the code has a branch the
>> instruction fetched has to be thrown away.
>>
>> By your definition all processors prefetch because the instruction is
>> fetched BEFORE it is decoded or executed.
>
> No, please, try to understand what you are doing.  A processor which
> doesn't prefetch does
>
> fetch | execute | fetch | execute

This is a retarded processor with pipeline registers, but no 
concurrency.  It should be...

fetch execute | fetch execute | fetch execute

Why would you draw a line between the fetch and the execute?


> This is a non-pipelined processor.  Each stage happens on its own, no
> overlap.  Your processor does
>
> cyc1  | cyc2    | cyc3    | cyc4
> fetch | exeucte |
>          fetch   | execute
>                    fetch   | execute

You are correct that the fetch and execute operate concurrently.  That 
is the only difference.  I simply put my register in a different place.


>> The fetch is done on the next
>> instruction to be executed, not an instruction that is *anticipated*
>> to be fetched.
>
> That's perfectly fine, in a 2-stage pipeline, you don't need to
> anticipate anything.  You are probably confusing this with a classical
> four-stage pipeline design where you have
>
> fetch | decode+r | execute  | write
>          fetch    | decode+r | execute | write
>
> In this case, if you execute a branch, you have already fetched *one*
> instruction after the branch.  This one needs to be thrown away.
>
> With a two-stage pipeline, you don't need to throw anything away.  I.e.
> you always fetch the rigth instruction.
>
>> If you insist on calling this a prefetch, it adds nothing to the
>> conversation because the results that typically are connected to a
>> prefetch are not attached to my design.
>
> The starting point of this lengthy conversation was that I said that
> MISC processors typically use prefetch, but no other optimization.  This
> is what you opposed to, but the only disagreement we have is over
> language, not over actual facts.  Fact is that you parallelize fetch and
> execute, which *is* the low-hanging fruit of optimization I was talking
> about.

Ok, I withdraw my objection to your original statement.  If you want to 
call what I do as pipelining, then ok, my design is MP, minimally 
pipelined.


> If we disagree about the term ("instruction prefetch" is pretty
> generic), then well.  Let's say fetching the next instruction is a
> degenerated case of prefetch, because the instruction arrives just when
> you actually need it, i.e. no need for queuing it up.
>
> There are single-cycle MISC CPUs which do not prefetch at all, e.g. the
> J1 CPU.  There, the cycle consists of fetching the instruction, and then
> asynchronously decoding it and selecting the result of the corresponding
> operations (which have completed by the time the instruction is
> fetched).  It's not the next instrcution which is fetched, it's the
> *current* instruction.

Ok, you are winning me over.  Here is the difference in what I am doing...

Typical 1 stage pipelining:

Fetch 1 | Execute 1 |
         | Fetch 2   | Execute 2 |
                     | Fetch 3   | Opps execute 2 was a jump!  I need 
Fetch N now...

This happens because the address of the Fetch 3 is from the PC which was 
registered at the end of Fetch 2.  I register the PC that was *used* for 
Fetch 2 and the Fetch 3 calculates the address before the actual fetch 
so if instruction 2 is a jump, the correct instruction for 3 is fetched 
always.


>>>> I've never liked the term, "syntactic sugar" because it has no
>>>> defined meaning.
>>>
>>> Of course "syntactic sugar" has a defined meaning.  It is a syntactic
>>> element that makes the program easier to read, while having no
>>> function of its own:
>>>
>>> http://en.wikipedia.org/wiki/Syntactic_sugar
>>
>> Wikipedia doesn't define the language.  Regardless, calling something
>> syntactic sugar has no utility to me.
>
> *Nobody* defines language.  Language is made up by *usage*, and usage
> can be *observed*.  An encyclopedia is a collection of observed usage.
> Languages work by common understanding and using the same terms for the
> same things; if you are in doubt that the way you use your terms are
> understood by others, get a life and check reality.  Wikipedia can be a
> good start.

Wikipedia is of dubious origins.  I never use it as a primary source, as 
recommended by *wikipedia*.  When I use wikipedia, I always check the 
references for anything important and I never use it to prove a point.

I don't *use* syntactic sugar because it doesn't mean anything to me. 
If you want to communicate a thought to me related to this, please use 
other terms.

Actually, I think I don't like the term because of the way others use 
it.  They use is as a label rather than discussing the real issue. 
Labels tend to simplify an issue and bypass real thought about it.


>>> I.e. to get HP's syntax, you define
>>>
>>> : endif  postpone then ; immediate
>>> : then   postpone if ; immediate
>>> : if     ; immediate
>>>
>>> and then you can write something like
>>>
>>> : simple. ( n -- ) if dup 3>   then drop ." many " else . endif ;
>>> \ numbers are 1 2 3 many...
>>>
>>> IF in this case is obvious syntactic sugar, as you can leave it away
>>> without changing the functionality of the program.
>>>
>>
>> Unfortunately you have snipped the context of the syntactic sugar
>> statement so I have no idea what this is now about.
>
> The context of syntactic sugar was HP's modified IF THEN ELSE ENDIF, as
> you can read here.  The context is still there.
>

Maybe the real issue is I don't know what this part of the conversation 
is about.  I think I said something about someone's objection to OR as 
the xor instruction and compared it to some of the other "issues" in 
Forth using unusual constructs like IF ELSE THEN.

I'm not sure I'm interested in exploring this part further.

Rick

[toc] | [prev] | [next] | [standalone]


#16798

Fromalbert@spenarnc.xs4all.nl (Albert van der Horst)
Date2012-10-28 07:59 +0000
Message-ID<508ce5de$0$3195$e4fe514c@dreader36.news.xs4all.nl>
In reply to#16789
In article <1762397.1u9z86ECTl@sunwukong.fritz.box>,
Bernd Paysan  <bernd.paysan@gmx.de> wrote:
<SNIP>
>
>Calling an XOR and XOR is really wise, and SWAP is a useful instruction.
>I don't know why Chuck wants to deliberately confuse people with his
>OR=XOR and -=INVERT.  If he wants to be compact, he could use ~ for
>invert, & for AND, | for OR, and ^ for XOR.  Nobody would complain.

If you try working with colorforth it becomes painfully obvious.
With the colorforth encoding you've only 80% chance to get the
right key. So typing OR succeeds in 64% of the cases, while XOR
would result in about 49 %. The word - you can type with an 80% hit rate
and the chance of getting INVERT correct becomes infinitesimal.
Remember there is no erase key, only an erase word!

<SNIP>

>--
>Bernd Paysan
>"If you want it done right, you have to do it yourself"
>http://bernd-paysan.de/
>

[toc] | [prev] | [next] | [standalone]


#16815

Fromrickman <gnuarm@gmail.com>
Date2012-10-28 15:06 -0400
Message-ID<k6jvnp$gk2$2@dont-email.me>
In reply to#16798
On 10/28/2012 3:59 AM, Albert van der Horst wrote:
> In article<1762397.1u9z86ECTl@sunwukong.fritz.box>,
> Bernd Paysan<bernd.paysan@gmx.de>  wrote:
> <SNIP>
>>
>> Calling an XOR and XOR is really wise, and SWAP is a useful instruction.
>> I don't know why Chuck wants to deliberately confuse people with his
>> OR=XOR and -=INVERT.  If he wants to be compact, he could use ~ for
>> invert,&  for AND, | for OR, and ^ for XOR.  Nobody would complain.
>
> If you try working with colorforth it becomes painfully obvious.
> With the colorforth encoding you've only 80% chance to get the
> right key. So typing OR succeeds in 64% of the cases, while XOR
> would result in about 49 %. The word - you can type with an 80% hit rate
> and the chance of getting INVERT correct becomes infinitesimal.
> Remember there is no erase key, only an erase word!

You may just be right!

Rick

[toc] | [prev] | [next] | [standalone]


#16709

Fromrickman <gnuarm@gmail.com>
Date2012-10-25 16:35 -0400
Message-ID<k6c7pr$7eo$1@dont-email.me>
In reply to#16682
On 10/25/2012 12:30 AM, Rod Pemberton wrote:
> "rickman"<gnuarm@gmail.com>  wrote in message
> news:k69fbv$a6n$1@dont-email.me...
>> On 10/24/2012 6:56 AM, Rod Pemberton wrote:
>>> "rickman"<gnuarm@gmail.com>   wrote in message
>>>> It seems that if the CPU isn't pipelined and blazing fast, it isn't
>>>> interesting to most FPGA users.
>>>
>>> Without pipelining, I doubt it's of interest to anyone.  The 6502
>>> proved how effective pipelining in a microprocessor is in increasing
>>> performance.
>>
>> I don't follow that.  There are lots of MCUs that aren't pipelined.  The
>> need for speed is a spectrum.  Often there are other things that are
>> more important requirements, like size.  How many LUTs is the Leon
>> processor, 5000?  Fast, but large.
>>
>
> Are we talking about the same thing?
>
> http://en.wikipedia.org/wiki/Instruction_pipeline
> http://en.wikipedia.org/wiki/Pipelining
>
> AFAIK, the 6502 was the first microprocessor with pipelining.  And, the Cray
> CDC 6600 had the first central processing unit with it.
>
> If we're talking about the same thing and a microprocessor design isn't
> using pipelining, how do you get any acceptable level of performance?  Do
> you just use faster logic and clock?

Before I can answer that question, you need to define "acceptable level 
of performance".  My entire point is that the performance needs of 
different applications vary across the spectrum.  There are costs 
associated with pipelining; it takes more gates, it uses more power and 
the performance speed up depends on the code being executed, there are 
pipeline stalls and flushes.  Pipelining is not a magic formula for 
"free" performance.  It is a tradeoff.

I think I already mentioned that my design goals had to do with 
minimizing the resources used balanced against performance.  Using one 
block RAM for the stacks, a second for instruction memory and a third 
for RAM, it only required a further 600 LUTs to implement a 16 bit 
processor.  That included a few bells and whistles that I would likely 
rip out with any further work.  This ran at 50 MHz in a 15 year old FPGA.

That level of performance was fine for the application I had.  I came 
close to using it again on a design a couple of years ago when a 
customer wanted me to upgrade an existing design.  But I was able to use 
the Lattice device at 90% without an impact on the performance so in the 
end adding the functions as logic worked out ok.

I'm looking at a design in a Lattice iCE40 device which would be the 
lowest power way of implementing this I believe.  I might use this 
processor rather than application specific logic.  We'll see.  It all 
depends on how small I can get the processor.  It may require an 8 bit 
version or even a redesign to 4 or 5 bit data paths.

Rick

[toc] | [prev] | [next] | [standalone]


#16619

Fromanton@mips.complang.tuwien.ac.at (Anton Ertl)
Date2012-10-23 12:44 +0000
Message-ID<2012Oct23.144415@mips.complang.tuwien.ac.at>
In reply to#16595
rickman <gnuarm@gmail.com> writes:
>Info that would be VERY useful to me is frequency of use of instructions 
>in combination.  This could be pulled from existing code by measuring 
>how often instructions are found adjacent to each other.  This may not 
>be a perfect measure, but it would be a great start.  Can anyone 
>generate a metric on this similar to Koopman's data on instruction use?

http://www.complang.tuwien.ac.at/forth/peep/

in particular:

http://www.complang.tuwien.ac.at/forth/peep/sorted

There is also a later paper that shows different kinds of data:

http://www.complang.tuwien.ac.at/anton/euroforth/ef01/gregg01.pdf

- anton
-- 
M. Anton Ertl  http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
     New standard: http://www.forth200x.org/forth200x.html
   EuroForth 2012: http://www.euroforth.org/ef12/

[toc] | [prev] | [next] | [standalone]


#16626

Fromrickman <gnuarm@gmail.com>
Date2012-10-23 17:32 -0400
Message-ID<k672d4$60o$1@dont-email.me>
In reply to#16619
On 10/23/2012 8:44 AM, Anton Ertl wrote:
> rickman<gnuarm@gmail.com>  writes:
>> Info that would be VERY useful to me is frequency of use of instructions
>> in combination.  This could be pulled from existing code by measuring
>> how often instructions are found adjacent to each other.  This may not
>> be a perfect measure, but it would be a great start.  Can anyone
>> generate a metric on this similar to Koopman's data on instruction use?
>
> http://www.complang.tuwien.ac.at/forth/peep/
>
> in particular:
>
> http://www.complang.tuwien.ac.at/forth/peep/sorted
>
> There is also a later paper that shows different kinds of data:
>
> http://www.complang.tuwien.ac.at/anton/euroforth/ef01/gregg01.pdf
>
> - anton

Thank you.  I'll take a look at this data.

Rick

[toc] | [prev] | [next] | [standalone]


#16596

Fromdaveyrotten <danw8804@gmail.com>
Date2012-10-22 13:54 -0700
Message-ID<df90307f-0d2e-463f-9b1c-a9790cb2dea6@googlegroups.com>
In reply to#16579
On Monday, October 22, 2012 10:53:24 AM UTC-5, Brad Eckert wrote:
> Hi All,
> 
> 
> 
> I've been thinking about Novix style processors like the RTX2000. There are many Forth sequences that can be compacted into one instruction, so with a good optimizer the chip can execute several Forth (source) primitives in one machine cycle. I suspect though that such optimization opportunities are the exception rather than the rule.
> 
> 
> 
> Does anyone here have a feel for the correspondence between Forth source primitives and generated code?
> 
> 
> 
> This kind of architecture is pretty good if you can get a wide instruction from code space every clock. You have calls and you have everything else, where everything else is a kind of compact VLIW. Implementation can be simple, as shown by James Bowman's J1.

Are you sure you're not confusing machine cycles with address fetches?  Paysan's B16 (taking a cue from Moore's work) packs something like 3 instructions in a single memory word. But each of those instructions still takes a complete processor clock cycle to execute (ie, they execute in sequence). So it saves memory space but doesn't really execute 3 times faster. Maybe I'm not understanding what you mean by machine cycle. I'm a huge fan of James Bowman's J1, but I don't believe it packs instructions in memory at all. It simply executes them very fast thru the use of dual port memory. I think of the B16 and J1 as two sides of the same coin. The B16 is very memory efficient, but slower than the J1. The J1 is very fast, but not as memory efficient as the B16.  The B16 is programmed directly in Forth, using about 32 built-in words. The J1 is similar but I believe has a few more possible Forth words as primitive instructions.

[toc] | [prev] | [next] | [standalone]


#16597

Fromvisualforth@rocketmail.com
Date2012-10-22 14:14 -0700
Message-ID<d87fa050-c403-4908-a96b-f8a35a17aed7@googlegroups.com>
In reply to#16596
On Monday, October 22, 2012 4:54:33 PM UTC-4, daveyrotten wrote:
> Paysan's B16 (taking a cue from Moore's work) packs something like 3 instructions in a single memory word. But each of those instructions still takes a complete processor clock cycle to execute (ie, they execute in sequence). So it saves memory space but doesn't really execute 3 times faster. 

The RTX2000 family of RISCs is able to execute several commands within one clock cycle. Only memory access needs two clock cycles: one to address memory, and one to fetch/store.
 

[toc] | [prev] | [next] | [standalone]


#16599

Fromdaveyrotten <danw8804@gmail.com>
Date2012-10-22 14:26 -0700
Message-ID<adaa3a9b-c35f-4ff3-9609-f79fea88f59e@googlegroups.com>
In reply to#16597
On Monday, October 22, 2012 4:14:19 PM UTC-5, visua...@rocketmail.com wrote:
> On Monday, October 22, 2012 4:54:33 PM UTC-4, daveyrotten wrote:
> 
> > Paysan's B16 (taking a cue from Moore's work) packs something like 3 instructions in a single memory word. But each of those instructions still takes a complete processor clock cycle to execute (ie, they execute in sequence). So it saves memory space but doesn't really execute 3 times faster. 
> 
> 
> 
> The RTX2000 family of RISCs is able to execute several commands within one clock cycle. Only memory access needs two clock cycles: one to address memory, and one to fetch/store.

Ok, sorry. I guess the compiler must pick instructions that have no interdependency to pack together. I'm not familiar with the RTX2000 itself.

[toc] | [prev] | [next] | [standalone]


#16601

Fromvisualforth@rocketmail.com
Date2012-10-22 16:00 -0700
Message-ID<6ff014de-e10a-4ae3-abaf-b2abdf916549@googlegroups.com>
In reply to#16599
On Monday, October 22, 2012 5:26:20 PM UTC-4, daveyrotten wrote:
> On Monday, October 22, 2012 4:14:19 PM UTC-5, visua...@rocketmail.com wrote: > On Monday, October 22, 2012 4:54:33 PM UTC-4, daveyrotten wrote: > > > Paysan's B16 (taking a cue from Moore's work) packs something like 3 instructions in a single memory word. But each of those instructions still takes a complete processor clock cycle to execute (ie, they execute in sequence). So it saves memory space but doesn't really execute 3 times faster. > > > > The RTX2000 family of RISCs is able to execute several commands within one clock cycle. Only memory access needs two clock cycles: one to address memory, and one to fetch/store. Ok, sorry. I guess the compiler must pick instructions that have no interdependency to pack together. I'm not familiar with the RTX2000 itself.

I am. I did several designs, 1988-1995:
http://www.somersetweb.com/BruehlConsult/Projects/RTX2000-MINI.html
http://www.somersetweb.com/BruehlConsult/Projects/mc-RISC-EMUF.html
http://www.somersetweb.com/BruehlConsult/Projects/S5-4MB-RTX2000.html
http://www.somersetweb.com/BruehlConsult/Projects/S5-F-1MBd-RTX2000.html

[toc] | [prev] | [next] | [standalone]


#16605

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-22 21:51 -0400
Message-ID<k64svk$njr$1@speranza.aioe.org>
In reply to#16579
"Brad Eckert" <hwfwguy@gmail.com> wrote in message
news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
> I've been thinking about Novix style processors like the RTX2000.
> There are many Forth sequences that can be compacted into one
> instruction, so with a good optimizer the chip can execute several
> Forth (source) primitives in one machine cycle. I suspect though
> that such optimization opportunities are the exception rather than the
> rule.
>

The question for both you (and Rick) is if you create new, faster, more
powerful, multiple operation instructions, how do you ensure they are used?
Without an optimizer, it's likely the instruction will have a low
instruction frequency.  I.e., a person is unlikely to use it.  In which
case, there is no point in using or implementing it.
(This is repeated later in a reply to Rick.)

> Does anyone here have a feel for the correspondence between Forth
> source primitives and generated code?

Generally, Forth's built using "primitives" or low-level words generally
need 30 to 40 or so.  I kept track of how many are needed for certain
Forths.  There are a few posts by me to c.l.f. with counts and specific
words used.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#16627

Fromrickman <gnuarm@gmail.com>
Date2012-10-23 17:36 -0400
Message-ID<k672lj$82o$1@dont-email.me>
In reply to#16605
On 10/22/2012 9:51 PM, Rod Pemberton wrote:
> "Brad Eckert"<hwfwguy@gmail.com>  wrote in message
> news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
>> I've been thinking about Novix style processors like the RTX2000.
>> There are many Forth sequences that can be compacted into one
>> instruction, so with a good optimizer the chip can execute several
>> Forth (source) primitives in one machine cycle. I suspect though
>> that such optimization opportunities are the exception rather than the
>> rule.
>>
>
> The question for both you (and Rick) is if you create new, faster, more
> powerful, multiple operation instructions, how do you ensure they are used?
> Without an optimizer, it's likely the instruction will have a low
> instruction frequency.  I.e., a person is unlikely to use it.  In which
> case, there is no point in using or implementing it.
> (This is repeated later in a reply to Rick.)

We aren't talking about Forth coding really.  We are talking about the 
assembly language for a machine.  I don't think instructions will go 
unused just because they are mapped to Forth in a more complicated way 
than 1 to 1 (or 1/2 to 1).


>> Does anyone here have a feel for the correspondence between Forth
>> source primitives and generated code?
>
> Generally, Forth's built using "primitives" or low-level words generally
> need 30 to 40 or so.  I kept track of how many are needed for certain
> Forths.  There are a few posts by me to c.l.f. with counts and specific
> words used.

Don't confuse Forth low level primitives (which are really HLL 
primitives selected to be convenient for the programmer writing a Forth) 
and assembly language which has to be selected in part based on what is 
practical and efficient to implement.  Chuck's machine only uses 32 
opcodes and you can get by with as few as 16.

Rick

[toc] | [prev] | [next] | [standalone]


#16646

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-24 08:47 -0400
Message-ID<k68npd$9f7$1@speranza.aioe.org>
In reply to#16627
"rickman" <gnuarm@gmail.com> wrote in message
news:k672lj$82o$1@dont-email.me...
> On 10/22/2012 9:51 PM, Rod Pemberton wrote:
> > "Brad Eckert"<hwfwguy@gmail.com>  wrote in message
> > news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
> >> I've been thinking about Novix style processors like the RTX2000.
> >> There are many Forth sequences that can be compacted into one
> >> instruction, so with a good optimizer the chip can execute several
> >> Forth (source) primitives in one machine cycle. I suspect though
> >> that such optimization opportunities are the exception rather than the
> >> rule.
> >>
> >
> > The question for both you (and Rick) is if you create new, faster, more
> > powerful, multiple operation instructions, how do you ensure they are
> > used?  Without an optimizer, it's likely the instruction will have a low
> > instruction frequency.  I.e., a person is unlikely to use it.  In which
> > case, there is no point in using or implementing it.
> > (This is repeated later in a reply to Rick.)
>
> We aren't talking about Forth coding really.  We are talking about the
> assembly language for a machine.  I don't think instructions will go
> unused just because they are mapped to Forth in a more complicated way
> than 1 to 1 (or 1/2 to 1).
>
> >> Does anyone here have a feel for the correspondence between Forth
> >> source primitives and generated code?
> >
> > Generally, Forth's built using "primitives" or low-level words generally
> > need 30 to 40 or so.  I kept track of how many are needed for certain
> > Forths.  There are a few posts by me to c.l.f. with counts and specific
> > words used.
>
> Don't confuse Forth low level primitives (which are really HLL
> primitives selected to be convenient for the programmer writing a Forth)
> and assembly language which has to be selected in part based on what is
> practical and efficient to implement.  Chuck's machine only uses 32
> opcodes and you can get by with as few as 16.
>

For temporary reasons, most of my Forth stack operators, like DUP SWAP etc,
aren't currently low-level Forth "primitives".  They're implemented in
high-level Forth using an even lower set of actual stack "primitives", which
I'll call sub-operators for this thread.  These sub-operators could be
considered to be "stack assembly instructions" for my Forth.  Currently,
only >R and R> are actually "primitives" coded in C.  Eventually, DUP DROP
SWAP OVER will be actual low-level Forth "primitives", as they once were.


So, a word like OVER can be coded in many ways.  I have 22 different
definitions just for OVER in a list, and one can construct many more.  E.g.,

 : OVER >R DUP R> SWAP ;
 : OVER 1 PICK ;
 : OVER SWAP TUCK ;
 : OVER SWAP DUP -ROT ;
 : OVER NUP SWAP ;
etc.

Which definition you choose affects how fast OVER executes and depends on
what operations you have available and on how fast each of those operations
are.

E.g., if >R >R DUP SWAP and -ROT are all very fast machine instructions for
your processor, then ">R DUP R> SWAP" and "SWAP DUP -ROT" should be fast
sequences.  And, one sequence will be faster than the other, depending on
how fast each instruction is.  But, if -ROT is not a machine instruction,
e.g., perhaps coded as "ROT ROT" or "SWAP >R SWAP R>", or -ROT is a very
slow machine instruction, then "SWAP DUP -ROT" is more expensive than ">R
DUP R> SWAP".


The instructions also need to be balanced acrossed many Forth words.
Originally, I had the 2xxx series words defined in terms of the simpler
stack operators, like SWAP DUP etc.  When I converted to sub-operators, the
number of operations per definition dropped dramatically for some of the
2xxx definitions.  Since the sub-operators are primitives, some of the 2xxx
definitions became faster.  However, after converting all of the simpler
stack operators to sub-operators, a few of the non-primitive simple stack
operators ended up with more operations per definition.  So some words
became much faster, while others became slightly slower.  And, the former
primitives became real slow, but they'll be converted back eventually.


Let's take a look at 2OVER for my Forth interpreter.  I could easily
implement it as a Forth low-level "primitive".  Or, I could define it in
high-level Forth. Or, I could define it in high-level Forth in terms of
sub-operators which are actually the low-level "primitives".

In high-level Forth, 2OVER can be defined:

 : 2OVER 2>R 2DUP 2R> 2SWAP ;

In terms of my sub-operators, my 2OVER definition has ten words it's
definition.  Clearly, that's many more words than the four in definition
above.  But, 2>R 2R> 2DUP and 2SWAP are also implemented in high-level
Forth using the same set of sub-operators.  So, 2>R 2R> 2DUP and 2SWAP
aren't "primitives" nor are they a single sub-operator each.

The 2>R sequence has six items.
The 2DUP sequence has six items.
The 2R> sequence has six items.
The 2SWAP sequence has eight items.

The 2OVER sequence has ten items.

So, 2OVER is only 10 items using a single sequence of sub-operators instead
of a total 28 items using four sequences of sub-operators for the four words
in the high-level definition.  If the 2>R 2DUP 2R> and 2SWAP words are
defined in terms of standard Forth words:

The 2>R sequence has three items.
The 2DUP sequence has ten items over multiple words.
The 2R> sequence has three items.
The 2SWAP sequence has twelve items over multiple words.

In this case, the counts changed, but it just happens that it's total is 28
also...  Usually, it's more.  Of course, you'd rather have a 2DUP of six
items instead of ten, i.e., balance.  If the instructions are unbalanced,
then heavy use of a single Forth word will slow the code speed way down.
I.e., many 2DUP's of ten items is much worse than many 2DUP's with six
items, even if other used words are made slightly slower, like 2>R and 2R>.
Of course, you don't know if your user's code will follow the measured
instruction frequencies or not.  But, at this point, you're the only user
...

As a primitive, 2OVER will be a small C routine which is compiled to
optimized, machine code.


Rod Pemberton


[toc] | [prev] | [next] | [standalone]


Page 6 of 9 — ← Prev page 1 2 3 4 5 [6] 7 8 9  Next page →

Back to top | Article view | comp.lang.forth


csiph-web