Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.forth > #16579 > unrolled thread

RTX2000 optimization

Started byBrad Eckert <hwfwguy@gmail.com>
First post2012-10-22 08:53 -0700
Last post2012-10-23 10:52 +0000
Articles 20 on this page of 172 — 22 participants

Back to article view | Back to comp.lang.forth


Contents

  RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-22 08:53 -0700
    Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-22 11:21 -0500
    Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 09:56 -0700
      Re: RTX2000 optimization Mark Wills <forthfreak@gmail.com> - 2012-10-22 12:10 -0700
        Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 12:55 -0700
        Re: RTX2000 optimization Coos Haak <chforth@hccnet.nl> - 2012-10-22 22:02 +0200
        Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-22 16:50 -0400
          Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:52 -0400
            Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-22 22:03 -0700
              Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-23 03:19 -0500
                Re: RTX2000 optimization vandys@vsta.org - 2012-10-23 17:54 +0000
                  Re: RTX2000 optimization Hugh Aguilar <hughaguilar96@yahoo.com> - 2012-10-24 08:16 -0700
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:19 -0400
              Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 06:56 -0400
                Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-24 12:26 +0000
                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:25 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:30 -0400
                    Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 22:10 -0700
                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:43 -0400
                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 17:55 +0200
                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:44 -0400
                            Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 02:17 +0200
                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 23:01 -0400
                                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-27 22:18 +0200
                                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 19:19 -0400
                                    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 04:21 -0400
                                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 14:36 -0400
                                        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:06 -0400
                                          Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 02:07 -0700
                                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 09:36 -0400
                                              Re: RTX2000 optimization Mark Wills <markrobertwills@yahoo.co.uk> - 2012-10-30 08:09 -0700
                                                Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-30 19:14 -0400
                                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 18:50 -0400
                                            Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 07:37 -0700
                                              Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-31 15:27 +0000
                                                Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-31 08:59 -0700
                                                  Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-10-31 11:18 -0500
                                                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:49 -0400
                                                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 13:43 -0400
                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 12:03 -0400
                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Mark Wills <forthfreak@gmail.com> - 2012-10-31 09:05 -0700
                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] daveyrotten <danw8804@gmail.com> - 2012-10-31 09:06 -0700
                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 09:25 -0700
                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] danw8804@gmail.com - 2012-10-31 09:38 -0700
                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:30 -0400
                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 14:27 -0400
                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 12:01 -0700
                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-10-31 16:31 -0400
                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-10-31 20:33 -0700
                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:05 -0400
                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 11:23 -0700
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-01 14:31 -0400
                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-02 22:11 -0700
                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 12:50 -0400
                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 10:42 -0700
                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 13:59 -0400
                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 12:10 -0700
                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 15:42 -0400
                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-03 15:56 -0700
                                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-03 20:53 -0400
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] stephenXXX@mpeforth.com (Stephen Pelc) - 2012-11-04 11:10 +0000
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-04 22:58 -0800
                                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 15:29 +0100
                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 14:40 +0000
                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 11:36 -0500
                                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 09:02 -0800
                                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:22 -0500
                                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-05 12:56 -0800
                                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:22 -0500
                                                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 09:29 -0800
                                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:53 -0500
                                                                                                    Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-06 10:00 -0800
                                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Elizabeth D. Rather" <erather@forth.com> - 2012-11-06 08:04 -1000
                                                                                                      Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 13:37 -0500
                                                                                        Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-05 21:06 +0100
                                                                                          Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-05 15:29 -0500
                                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-06 16:28 +0100
                                                                                              Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] rickman <gnuarm@gmail.com> - 2012-11-06 12:50 -0500
                                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-07 20:29 -0800
                                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-11-04 10:40 +0000
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-01 21:26 +0100
                                                                  Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] Paul Rubin <no.email@nospam.invalid> - 2012-11-01 13:44 -0700
                                                                Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-02 04:03 -0400
                                                            Re: Xula-200 or Raspberry-Pi a better buy?, was [Re: RTX2000 optimization] David Schultz <abuse@127.0.0.1> - 2012-11-04 17:17 -0600
                                              Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 12:48 -0700
                                                Re: RTX2000 optimization "Elizabeth D. Rather" <erather@forth.com> - 2012-11-02 10:25 -1000
                                                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-02 19:54 -0700
                                                  Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-03 17:47 +0100
                                                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-03 14:05 -0400
                                                      Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-11-05 11:55 +0000
                                                    Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-11-03 19:20 -0700
                                                      Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 16:08 +0200
                                                        Re: RTX2000 optimization mhx@iae.nl (Marcel Hendrix) - 2012-11-04 17:14 +0200
                                                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-11-04 18:40 +0100
                                                      Re: RTX2000 optimization Andrew Haley <andrew29@littlepinkcloud.invalid> - 2012-11-04 09:36 -0600
                                                      Re: RTX2000 optimization Andy Valencia <vandys@vsta.org> - 2012-11-04 23:49 +0000
                                          Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-31 12:11 -0700
                                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 20:08 -0400
                                              Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-01 10:59 -0700
                                                Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:12 -0700
                                                  Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-01 12:14 -0700
                                                  Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-11-02 10:51 -0700
                                                    Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-11-02 11:28 -0700
                                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-11-01 14:58 -0400
                                    Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 14:45 +0100
                                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:04 -0400
                                        Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-28 21:09 +0100
                                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 16:59 -0400
                                  Re: RTX2000 optimization albert@spenarnc.xs4all.nl (Albert van der Horst) - 2012-10-28 07:59 +0000
                                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:06 -0400
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 16:35 -0400
          Re: RTX2000 optimization anton@mips.complang.tuwien.ac.at (Anton Ertl) - 2012-10-23 12:44 +0000
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:32 -0400
    Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 13:54 -0700
      Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 14:14 -0700
        Re: RTX2000 optimization daveyrotten <danw8804@gmail.com> - 2012-10-22 14:26 -0700
          Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 16:00 -0700
    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-22 21:51 -0400
      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:36 -0400
        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-24 08:47 -0400
          Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-24 06:36 -0700
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:33 -0400
              Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:54 -0700
                Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 17:35 -0400
                  Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 15:27 -0700
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:10 -0400
                      Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 18:34 -0700
                        Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:03 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:56 -0400
                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:34 +0200
                  Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 22:27 +0200
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:27 -0400
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:17 -0400
                  Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 19:44 -0400
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-26 19:59 -0400
            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:42 -0400
              Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:55 -0400
              Re: RTX2000 optimization Alex McDonald <blog@rivadpm.com> - 2012-10-25 04:49 -0700
            addressable stack, was [Re: RTX2000 optimization] "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-11-05 14:04 -0500
          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 15:44 -0400
            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-25 00:50 -0400
              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 18:58 -0400
                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 16:07 -0700
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:20 -0400
                Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-26 20:57 -0400
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-27 15:43 -0400
                    Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-28 05:01 -0400
                      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-28 15:23 -0400
                        Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-29 19:32 -0400
                          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-30 19:00 -0400
                            Re: RTX2000 optimization "Rod Pemberton" <do_not_have@notemailnotz.cnm> - 2012-10-31 01:23 -0400
                              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-31 14:56 -0400
    Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-22 19:44 -0700
      Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 17:47 -0400
        Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 16:00 -0700
          Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-23 21:07 -0400
            Re: RTX2000 optimization visualforth@rocketmail.com - 2012-10-23 18:44 -0700
              Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:03 -0400
                Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:15 -0700
                  Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:25 -0400
                Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 03:15 +0200
        Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 09:55 -0700
          Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 10:04 -0700
            Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-24 12:00 -0700
              Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 23:24 -0700
                Re: RTX2000 optimization Brad Eckert <hwfwguy@gmail.com> - 2012-10-25 09:26 -0700
                  Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-25 10:39 -0700
                    Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-25 19:35 -0400
                    Re: RTX2000 optimization Bernd Paysan <bernd.paysan@gmx.de> - 2012-10-26 18:26 +0200
            Re: RTX2000 optimization rickman <gnuarm@gmail.com> - 2012-10-24 16:10 -0400
              Re: RTX2000 optimization Paul Rubin <no.email@nospam.invalid> - 2012-10-24 13:21 -0700
    Re: RTX2000 optimization stephenXXX@mpeforth.com (Stephen Pelc) - 2012-10-23 10:52 +0000

Page 7 of 9 — ← Prev page 1 2 3 4 5 6 [7] 8 9  Next page →


#16649

FromAlex McDonald <blog@rivadpm.com>
Date2012-10-24 06:36 -0700
Message-ID<172e6a33-bc52-4fe9-8867-88f69c2954f8@10g2000vbu.googlegroups.com>
In reply to#16646
On Oct 24, 1:43 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm>
wrote:
> "rickman" <gnu...@gmail.com> wrote in message
>
> news:k672lj$82o$1@dont-email.me...
>
>
>
>
>
>
>
>
>
> > On 10/22/2012 9:51 PM, Rod Pemberton wrote:
> > > "Brad Eckert"<hwfw...@gmail.com>  wrote in message
> > >news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
> > >> I've been thinking about Novix style processors like the RTX2000.
> > >> There are many Forth sequences that can be compacted into one
> > >> instruction, so with a good optimizer the chip can execute several
> > >> Forth (source) primitives in one machine cycle. I suspect though
> > >> that such optimization opportunities are the exception rather than the
> > >> rule.
>
> > > The question for both you (and Rick) is if you create new, faster, more
> > > powerful, multiple operation instructions, how do you ensure they are
> > > used?  Without an optimizer, it's likely the instruction will have a low
> > > instruction frequency.  I.e., a person is unlikely to use it.  In which
> > > case, there is no point in using or implementing it.
> > > (This is repeated later in a reply to Rick.)
>
> > We aren't talking about Forth coding really.  We are talking about the
> > assembly language for a machine.  I don't think instructions will go
> > unused just because they are mapped to Forth in a more complicated way
> > than 1 to 1 (or 1/2 to 1).
>
> > >> Does anyone here have a feel for the correspondence between Forth
> > >> source primitives and generated code?
>
> > > Generally, Forth's built using "primitives" or low-level words generally
> > > need 30 to 40 or so.  I kept track of how many are needed for certain
> > > Forths.  There are a few posts by me to c.l.f. with counts and specific
> > > words used.
>
> > Don't confuse Forth low level primitives (which are really HLL
> > primitives selected to be convenient for the programmer writing a Forth)
> > and assembly language which has to be selected in part based on what is
> > practical and efficient to implement.  Chuck's machine only uses 32
> > opcodes and you can get by with as few as 16.
>
> For temporary reasons, most of my Forth stack operators, like DUP SWAP etc,
> aren't currently low-level Forth "primitives".  They're implemented in
> high-level Forth using an even lower set of actual stack "primitives", which
> I'll call sub-operators for this thread.  These sub-operators could be
> considered to be "stack assembly instructions" for my Forth.  Currently,
> only >R and R> are actually "primitives" coded in C.  Eventually, DUP DROP
> SWAP OVER will be actual low-level Forth "primitives", as they once were.
>
> So, a word like OVER can be coded in many ways.  I have 22 different
> definitions just for OVER in a list, and one can construct many more.  E.g.,
>
>  : OVER >R DUP R> SWAP ;
>  : OVER 1 PICK ;
>  : OVER SWAP TUCK ;
>  : OVER SWAP DUP -ROT ;
>  : OVER NUP SWAP ;
> etc.
>
> Which definition you choose affects how fast OVER executes and depends on
> what operations you have available and on how fast each of those operations
> are.
>
> E.g., if >R >R DUP SWAP and -ROT are all very fast machine instructions for
> your processor, then ">R DUP R> SWAP" and "SWAP DUP -ROT" should be fast
> sequences.  And, one sequence will be faster than the other, depending on
> how fast each instruction is.  But, if -ROT is not a machine instruction,
> e.g., perhaps coded as "ROT ROT" or "SWAP >R SWAP R>", or -ROT is a very
> slow machine instruction, then "SWAP DUP -ROT" is more expensive than ">R
> DUP R> SWAP".
>
> The instructions also need to be balanced acrossed many Forth words.
> Originally, I had the 2xxx series words defined in terms of the simpler
> stack operators, like SWAP DUP etc.  When I converted to sub-operators, the
> number of operations per definition dropped dramatically for some of the
> 2xxx definitions.  Since the sub-operators are primitives, some of the 2xxx
> definitions became faster.  However, after converting all of the simpler
> stack operators to sub-operators, a few of the non-primitive simple stack
> operators ended up with more operations per definition.  So some words
> became much faster, while others became slightly slower.  And, the former
> primitives became real slow, but they'll be converted back eventually.
>
> Let's take a look at 2OVER for my Forth interpreter.  I could easily
> implement it as a Forth low-level "primitive".  Or, I could define it in
> high-level Forth. Or, I could define it in high-level Forth in terms of
> sub-operators which are actually the low-level "primitives".
>
> In high-level Forth, 2OVER can be defined:
>
>  : 2OVER 2>R 2DUP 2R> 2SWAP ;
>
> In terms of my sub-operators, my 2OVER definition has ten words it's
> definition.  Clearly, that's many more words than the four in definition
> above.  But, 2>R 2R> 2DUP and 2SWAP are also implemented in high-level
> Forth using the same set of sub-operators.  So, 2>R 2R> 2DUP and 2SWAP
> aren't "primitives" nor are they a single sub-operator each.
>
> The 2>R sequence has six items.
> The 2DUP sequence has six items.
> The 2R> sequence has six items.
> The 2SWAP sequence has eight items.
>
> The 2OVER sequence has ten items.
>
> So, 2OVER is only 10 items using a single sequence of sub-operators instead
> of a total 28 items using four sequences of sub-operators for the four words
> in the high-level definition.  If the 2>R 2DUP 2R> and 2SWAP words are
> defined in terms of standard Forth words:
>
> The 2>R sequence has three items.
> The 2DUP sequence has ten items over multiple words.
> The 2R> sequence has three items.
> The 2SWAP sequence has twelve items over multiple words.
>
> In this case, the counts changed, but it just happens that it's total is 28
> also...  Usually, it's more.  Of course, you'd rather have a 2DUP of six
> items instead of ten, i.e., balance.  If the instructions are unbalanced,
> then heavy use of a single Forth word will slow the code speed way down.
> I.e., many 2DUP's of ten items is much worse than many 2DUP's with six
> items, even if other used words are made slightly slower, like 2>R and 2R>.
> Of course, you don't know if your user's code will follow the measured
> instruction frequencies or not.  But, at this point, you're the only user
> ...
>
> As a primitive, 2OVER will be a small C routine which is compiled to
> optimized, machine code.
>
> Rod Pemberton

I believe the minimal set of primitives out of which all stack
juggling words can be built is DUP DROP SWAP >R R>.

: over         >r dup r> swap            ;
: nip          swap drop                 ;
: tuck         swap over                 ;
: rot          >r swap r> swap           ;
: -rot         swap >r swap r>           ;
: 2swap        rot >r rot r>             ;
: 2dup         over over                 ;
: 2over        2>r 2dup 2r> 2swap        ;
: 2drop        drop drop                 ;
: 2nip         2swap 2drop               ;
: 2rot         2>r 2swap 2r> 2swap       ;
: r@           r> dup >r                 ;
: 2>r          swap >r >r                ;
: 2r>          r> r> swap                ;
: 2r@          2r> 2dup 2>r              ;

In your "sub-operator" notation, such sequences can be greatly
simplified. Assuming an addressable stack (that is, we don't need to
POP to get entries off the stack to access them, such as on the x86)
there are 6 of these operations, 3 for each stack. SGET S[n] and RGET
R[n] fetch entries from the stacks by fixed offset from a stack
pointer, SPUT S[n] and RPUT R[n] store entries on the stacks, and SPTR
+ n and RPTR+ n adjust the stack pointers. In the example below, only
the S operators are shown, since the rstack operations and other
juggling have been optimised away. [reg] is a virtual register.

: 2OVER 2>R 2DUP 2R> 2SWAP ;
block ( S: 4 -- 6 R: 0 -- 0 )
[1] block-begin [2]
[5] sget S[2]
[4] sget S[3]
[5] sput S[-2]
[4] sput S[-1]
[3] sptr+ -2
[2] block-end [1]

[toc] | [prev] | [next] | [standalone]


#16665

Fromrickman <gnuarm@gmail.com>
Date2012-10-24 15:33 -0400
Message-ID<k69fr4$d5p$1@dont-email.me>
In reply to#16649
On 10/24/2012 9:36 AM, Alex McDonald wrote:
> On Oct 24, 1:43 pm, "Rod Pemberton"<do_not_h...@notemailnotz.cnm>
> wrote:
>> "rickman"<gnu...@gmail.com>  wrote in message
>>
>> news:k672lj$82o$1@dont-email.me...
>>
>>> We aren't talking about Forth coding really.  We are talking about the
>>> assembly language for a machine.  I don't think instructions will go
>>> unused just because they are mapped to Forth in a more complicated way
>>> than 1 to 1 (or 1/2 to 1).
>>
>>> Don't confuse Forth low level primitives (which are really HLL
>>> primitives selected to be convenient for the programmer writing a Forth)
>>> and assembly language which has to be selected in part based on what is
>>> practical and efficient to implement.  Chuck's machine only uses 32
>>> opcodes and you can get by with as few as 16.
>>
>> For temporary reasons, most of my Forth stack operators, like DUP SWAP etc,
>> aren't currently low-level Forth "primitives".  They're implemented in
>> high-level Forth using an even lower set of actual stack "primitives", which
>> I'll call sub-operators for this thread.  These sub-operators could be
>> considered to be "stack assembly instructions" for my Forth.  Currently,
>> only>R and R>  are actually "primitives" coded in C.  Eventually, DUP DROP
>> SWAP OVER will be actual low-level Forth "primitives", as they once were.
>>
>> So, a word like OVER can be coded in many ways.  I have 22 different
>> definitions just for OVER in a list, and one can construct many more.  E.g.,
>>
>>   : OVER>R DUP R>  SWAP ;
>>   : OVER 1 PICK ;
>>   : OVER SWAP TUCK ;
>>   : OVER SWAP DUP -ROT ;
>>   : OVER NUP SWAP ;
>> etc.
>>
>> Which definition you choose affects how fast OVER executes and depends on
>> what operations you have available and on how fast each of those operations
>> are.
>>
>> E.g., if>R>R DUP SWAP and -ROT are all very fast machine instructions for
>> your processor, then ">R DUP R>  SWAP" and "SWAP DUP -ROT" should be fast
>> sequences.  And, one sequence will be faster than the other, depending on
>> how fast each instruction is.  But, if -ROT is not a machine instruction,
>> e.g., perhaps coded as "ROT ROT" or "SWAP>R SWAP R>", or -ROT is a very
>> slow machine instruction, then "SWAP DUP -ROT" is more expensive than ">R
>> DUP R>  SWAP".
>>
>> The instructions also need to be balanced acrossed many Forth words.
>> Originally, I had the 2xxx series words defined in terms of the simpler
>> stack operators, like SWAP DUP etc.  When I converted to sub-operators, the
>> number of operations per definition dropped dramatically for some of the
>> 2xxx definitions.  Since the sub-operators are primitives, some of the 2xxx
>> definitions became faster.  However, after converting all of the simpler
>> stack operators to sub-operators, a few of the non-primitive simple stack
>> operators ended up with more operations per definition.  So some words
>> became much faster, while others became slightly slower.  And, the former
>> primitives became real slow, but they'll be converted back eventually.
>>
>> Let's take a look at 2OVER for my Forth interpreter.  I could easily
>> implement it as a Forth low-level "primitive".  Or, I could define it in
>> high-level Forth. Or, I could define it in high-level Forth in terms of
>> sub-operators which are actually the low-level "primitives".
>>
>> In high-level Forth, 2OVER can be defined:
>>
>>   : 2OVER 2>R 2DUP 2R>  2SWAP ;
>>
>> In terms of my sub-operators, my 2OVER definition has ten words it's
>> definition.  Clearly, that's many more words than the four in definition
>> above.  But, 2>R 2R>  2DUP and 2SWAP are also implemented in high-level
>> Forth using the same set of sub-operators.  So, 2>R 2R>  2DUP and 2SWAP
>> aren't "primitives" nor are they a single sub-operator each.
>>
>> The 2>R sequence has six items.
>> The 2DUP sequence has six items.
>> The 2R>  sequence has six items.
>> The 2SWAP sequence has eight items.
>>
>> The 2OVER sequence has ten items.
>>
>> So, 2OVER is only 10 items using a single sequence of sub-operators instead
>> of a total 28 items using four sequences of sub-operators for the four words
>> in the high-level definition.  If the 2>R 2DUP 2R>  and 2SWAP words are
>> defined in terms of standard Forth words:
>>
>> The 2>R sequence has three items.
>> The 2DUP sequence has ten items over multiple words.
>> The 2R>  sequence has three items.
>> The 2SWAP sequence has twelve items over multiple words.
>>
>> In this case, the counts changed, but it just happens that it's total is 28
>> also...  Usually, it's more.  Of course, you'd rather have a 2DUP of six
>> items instead of ten, i.e., balance.  If the instructions are unbalanced,
>> then heavy use of a single Forth word will slow the code speed way down.
>> I.e., many 2DUP's of ten items is much worse than many 2DUP's with six
>> items, even if other used words are made slightly slower, like 2>R and 2R>.
>> Of course, you don't know if your user's code will follow the measured
>> instruction frequencies or not.  But, at this point, you're the only user
>> ...
>>
>> As a primitive, 2OVER will be a small C routine which is compiled to
>> optimized, machine code.
>>
>> Rod Pemberton
>
> I believe the minimal set of primitives out of which all stack
> juggling words can be built is DUP DROP SWAP>R R>.
>
> : over>r dup r>  swap            ;
> : nip          swap drop                 ;
> : tuck         swap over                 ;
> : rot>r swap r>  swap           ;
> : -rot         swap>r swap r>            ;
> : 2swap        rot>r rot r>              ;
> : 2dup         over over                 ;
> : 2over        2>r 2dup 2r>  2swap        ;
> : 2drop        drop drop                 ;
> : 2nip         2swap 2drop               ;
> : 2rot         2>r 2swap 2r>  2swap       ;
> : r@           r>  dup>r                 ;
> : 2>r          swap>r>r                ;
> : 2r>           r>  r>  swap                ;
> : 2r@          2r>  2dup 2>r              ;
>
> In your "sub-operator" notation, such sequences can be greatly
> simplified. Assuming an addressable stack (that is, we don't need to
> POP to get entries off the stack to access them, such as on the x86)
> there are 6 of these operations, 3 for each stack. SGET S[n] and RGET
> R[n] fetch entries from the stacks by fixed offset from a stack
> pointer, SPUT S[n] and RPUT R[n] store entries on the stacks, and SPTR
> + n and RPTR+ n adjust the stack pointers. In the example below, only
> the S operators are shown, since the rstack operations and other
> juggling have been optimised away. [reg] is a virtual register.
>
> : 2OVER 2>R 2DUP 2R>  2SWAP ;
> block ( S: 4 -- 6 R: 0 -- 0 )
> [1] block-begin [2]
> [5] sget S[2]
> [4] sget S[3]
> [5] sput S[-2]
> [4] sput S[-1]
> [3] sptr+ -2
> [2] block-end [1]
>

Interesting.  Has anyone noticed that Chuck's GA144 does not include 
SWAP as a primitive?  I find it very odd, but workable.  In fact, the 
absence of this operator led me to discover that - . + - performs a 
subtraction without adding the 1 constant!

(for those of you who aren't familiar with the F18A assembly language, + 
is an add, . is a nop to let the carry propagate and - is a 1's 
complement,  not a 2's complement. )

It seems that in a lot of cases when you think you want a swap, you can 
do very well without it!

Rick

[toc] | [prev] | [next] | [standalone]


#16695

FromAlex McDonald <blog@rivadpm.com>
Date2012-10-25 04:54 -0700
Message-ID<73a7ee56-ed69-4993-a5f1-2295ef611f32@x21g2000vbg.googlegroups.com>
In reply to#16665
On Oct 24, 8:33 pm, rickman <gnu...@gmail.com> wrote:
[snip]
>
> Interesting.  Has anyone noticed that Chuck's GA144 does not include
> SWAP as a primitive?  I find it very odd, but workable.  In fact, the
> absence of this operator led me to discover that - . + - performs a
> subtraction without adding the 1 constant!
>
> (for those of you who aren't familiar with the F18A assembly language, +
> is an add, . is a nop to let the carry propagate and - is a 1's
> complement,  not a 2's complement. )
>
> It seems that in a lot of cases when you think you want a swap, you can
> do very well without it!
>
> Rick

It's possible to swap using XORs, but I'd be interested to see the
F18A equivalent. For some code sequences, I don't see how it can be
avoided; A B - is not the same as B A -

[toc] | [prev] | [next] | [standalone]


#16712

Fromrickman <gnuarm@gmail.com>
Date2012-10-25 17:35 -0400
Message-ID<k6cbc1$u8c$1@dont-email.me>
In reply to#16695
On 10/25/2012 7:54 AM, Alex McDonald wrote:
> On Oct 24, 8:33 pm, rickman<gnu...@gmail.com>  wrote:
> [snip]
>>
>> Interesting.  Has anyone noticed that Chuck's GA144 does not include
>> SWAP as a primitive?  I find it very odd, but workable.  In fact, the
>> absence of this operator led me to discover that - . + - performs a
>> subtraction without adding the 1 constant!
>>
>> (for those of you who aren't familiar with the F18A assembly language, +
>> is an add, . is a nop to let the carry propagate and - is a 1's
>> complement,  not a 2's complement. )
>>
>> It seems that in a lot of cases when you think you want a swap, you can
>> do very well without it!
>>
>> Rick
>
> It's possible to swap using XORs, but I'd be interested to see the
> F18A equivalent. For some code sequences, I don't see how it can be
> avoided; A B - is not the same as B A -

Yes, I recall there is a sequence using OVER and OR (xor) to do a swap. 
  I can't recall where I saw it now but I think it is not so optimal and 
still requires a temp register...  OVER DUP A! OR OR A@ or something 
like that.

No, A B - is not the same as B A -, in fact for all of the cases where I 
was doing a subtract, it was preceded by a swap, so this was a win-win. 
  Turns out the 1 constant is a very inefficient operation.  Not only 
does it occupy a full word of instruction memory, it takes an extra 
memory access (the slowest operation on the CPU, about three ALU ops) 
and kills the prefetch which hits performance (and power consumption) 
some more.  Anyone using the number 700 MIPS is kidding themselves. 
This CPU only does about half that in any but the most optimal 
situations and the "ops" are very basic so you need a lot more of them 
to do much.

That is the sort of stuff you need to learn to use the GA144.  I spent 
some two or three weeks wracking my brain to try to learn enough to code 
some tasks.  It was rewarding to find things like this that help to 
enhance the utility of the device.  But it just has so many hurdles that 
don't have "tricks" to get around.  It is still on my list of chips to 
consider, but the design I'm currently looking at doing requires lower 
power consumption than the GA144 is capable of.  I'm now looking at an 
iCE40 FPGA.

Rick

[toc] | [prev] | [next] | [standalone]


#16713

FromPaul Rubin <no.email@nospam.invalid>
Date2012-10-25 15:27 -0700
Message-ID<7x3912dw6x.fsf@ruckus.brouhaha.com>
In reply to#16712
rickman <gnuarm@gmail.com> writes:
> Yes, I recall there is a sequence using OVER and OR (xor) to do a
> swap. I can't recall where I saw it now but I think it is not so
> optimal and still requires a temp register...  OVER DUP A! OR OR A@ or
> something like that.

   push a! pop a

Or if you don't need the third thing on the stack:

   over over

Both of these are from  http://www.colorforth.com/inst.htm .

[toc] | [prev] | [next] | [standalone]


#16716

Fromrickman <gnuarm@gmail.com>
Date2012-10-25 19:10 -0400
Message-ID<k6cgu5$1pb$1@dont-email.me>
In reply to#16713
On 10/25/2012 6:27 PM, Paul Rubin wrote:
> rickman<gnuarm@gmail.com>  writes:
>> Yes, I recall there is a sequence using OVER and OR (xor) to do a
>> swap. I can't recall where I saw it now but I think it is not so
>> optimal and still requires a temp register...  OVER DUP A! OR OR A@ or
>> something like that.
>
>     push a! pop a

I don't consider this one very usable because of the number of side 
effects, using one of 10 levels of the return stack and clobbering the A 
register.  I'm not saying it can't be used, I just won't consider it 
often.

> Or if you don't need the third thing on the stack:
>
>     over over
>
> Both of these are from  http://www.colorforth.com/inst.htm .

Over over is a 2DUP, not a swap.  You could do an OVER, use the two 
operands and then do a DROP later in the code.  That is the most 
efficient when practical I suppose.

Personally, I'll just keep the SWAP as a primitive.

Rick

[toc] | [prev] | [next] | [standalone]


#16721

FromPaul Rubin <no.email@nospam.invalid>
Date2012-10-25 18:34 -0700
Message-ID<7xpq46m2x4.fsf@ruckus.brouhaha.com>
In reply to#16716
rickman <gnuarm@gmail.com> writes:
>>     push a! pop a
>
> I don't consider this one very usable because of the number of side
> effects, using one of 10 levels of the return stack and clobbering the
> A register.

I think F18 code won't normally have 10 levels of subroutine calls, so
storing temporaries on the return stack is fine.  You could even save
the A register there if you didn't want to clobber it.

> Over over is a 2DUP, not a swap.  

Oh of course you're right, I mis-read the page, should have said just
"over" instead of "over over".  Because of the circular stack, doing
this repeatedly won't cause overflow.  The issue is only if you've got
stuff past the 2nd element that needs to be preserved.

> You could do an OVER, use the two operands and then do a DROP later in
> the code.  That is the most efficient when practical I suppose.

You might not even need the DROP later.  

[toc] | [prev] | [next] | [standalone]


#16746

Fromrickman <gnuarm@gmail.com>
Date2012-10-26 19:03 -0400
Message-ID<k6f4s1$8vp$1@dont-email.me>
In reply to#16721
On 10/25/2012 9:34 PM, Paul Rubin wrote:
> rickman<gnuarm@gmail.com>  writes:
>>>      push a! pop a
>>
>> I don't consider this one very usable because of the number of side
>> effects, using one of 10 levels of the return stack and clobbering the
>> A register.
>
> I think F18 code won't normally have 10 levels of subroutine calls, so
> storing temporaries on the return stack is fine.  You could even save
> the A register there if you didn't want to clobber it.
>
>> Over over is a 2DUP, not a swap.
>
> Oh of course you're right, I mis-read the page, should have said just
> "over" instead of "over over".  Because of the circular stack, doing
> this repeatedly won't cause overflow.  The issue is only if you've got
> stuff past the 2nd element that needs to be preserved.
>
>> You could do an OVER, use the two operands and then do a DROP later in
>> the code.  That is the most efficient when practical I suppose.
>
> You might not even need the DROP later.

Now that is the right kind of thinking for programming the F18A!

Rick

[toc] | [prev] | [next] | [standalone]


#16753

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-26 19:56 -0400
Message-ID<k6f7of$70d$1@speranza.aioe.org>
In reply to#16712
"rickman" <gnuarm@gmail.com> wrote in message
news:k6cbc1$u8c$1@dont-email.me...
> On 10/25/2012 7:54 AM, Alex McDonald wrote:
> > On Oct 24, 8:33 pm, rickman<gnu...@gmail.com>  wrote:
> > [snip]

> > It's possible to swap using XORs, but I'd be interested to see the
> > F18A equivalent. For some code sequences, I don't see how it can be
> > avoided; A B - is not the same as B A -
>
> Yes, I recall there is a sequence using OVER and OR (xor) to do a swap.
>   I can't recall where I saw it now but I think it is not so optimal and
> still requires a temp register...  OVER DUP A! OR OR A@ or something
> like that.

Swapping two values using XOR's was a well known
programming "trick" in the 1980's ...


I'm not sure what Forth sequence you saw, but I have these in a file:

swap   over >r xor r@ xor r>
swap   over dup >r xor xor r>
swap   >r dup r> xor dup >r xor dup r> xor
swap   dup >r xor dup r> xor dup >r xor r>


They're much longer than other swap sequences:

swap   1 roll
swap   nup drip take
swap   tuck drop
swap   dup rot rot drop
swap   dup take
swap   over take
swap   over rot drop
swap   dup rot nip
swap   dup -rot drop
swap   over >r >r drop r> r>
swap   over >r nip r>
swap   over dup 3 roll drop drop
swap   0 pick take
swap   1 pick take
swap   >r >r 2r>
swap   2>r r> r>

etc.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#16741

FromBernd Paysan <bernd.paysan@gmx.de>
Date2012-10-26 18:34 +0200
Message-ID<1766437.ADrNJNi4aN@sunwukong.fritz.box>
In reply to#16695
Alex McDonald wrote:
> It's possible to swap using XORs, but I'd be interested to see the
> F18A equivalent. For some code sequences, I don't see how it can be
> avoided; A B - is not the same as B A -

But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r com 
r> +c".

Swap itself is "over >r nip r>".  You don't need to shuffle the stack 
that often.  In the 2k bytes Triceps program, there are in total 5 
swaps.  There are a lot more dups, overs (most of them as 2dups), >rs, 
and r>s.

Granted, I use swap a lot more in normal Forth programs, but as I know 
that it is an expensive operation on b16, I code around this fact, and 
it is possible.

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#16744

FromBernd Paysan <bernd.paysan@gmx.de>
Date2012-10-26 22:27 +0200
Message-ID<3819563.ke3Hixt9Hf@sunwukong.fritz.box>
In reply to#16741
Bernd Paysan wrote:

> Alex McDonald wrote:
>> It's possible to swap using XORs, but I'd be interested to see the
>> F18A equivalent. For some code sequences, I don't see how it can be
>> avoided; A B - is not the same as B A -
> 
> But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r
> com r> +c".

There's a faster swap -: com + com, on the F18A, that's written as

- . + -

-- 
Bernd Paysan
"If you want it done right, you have to do it yourself"
http://bernd-paysan.de/

[toc] | [prev] | [next] | [standalone]


#16748

Fromrickman <gnuarm@gmail.com>
Date2012-10-26 19:27 -0400
Message-ID<k6f69u$n3h$1@dont-email.me>
In reply to#16744
On 10/26/2012 4:27 PM, Bernd Paysan wrote:
> Bernd Paysan wrote:
>
>> Alex McDonald wrote:
>>> It's possible to swap using XORs, but I'd be interested to see the
>>> F18A equivalent. For some code sequences, I don't see how it can be
>>> avoided; A B - is not the same as B A -
>>
>> But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r
>> com r>  +c".
>
> There's a faster swap -: com + com, on the F18A, that's written as
>
> - . + -

Yup, that's what started this swap conversation.

Rick

[toc] | [prev] | [next] | [standalone]


#16747

Fromrickman <gnuarm@gmail.com>
Date2012-10-26 19:17 -0400
Message-ID<k6f5n2$k98$1@dont-email.me>
In reply to#16741
On 10/26/2012 12:34 PM, Bernd Paysan wrote:
> Alex McDonald wrote:
>> It's possible to swap using XORs, but I'd be interested to see the
>> F18A equivalent. For some code sequences, I don't see how it can be
>> avoided; A B - is not the same as B A -
>
> But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r com
> r>  +c".

There you go!  So on the F18A a subtract TOS from NOS can be done as 
push - pop . + -.  A lot faster and much smaller than the stock - . + 1 
. +.


> Swap itself is "over>r nip r>".  You don't need to shuffle the stack
> that often.  In the 2k bytes Triceps program, there are in total 5
> swaps.  There are a lot more dups, overs (most of them as 2dups),>rs,
> and r>s.

That may well be why the F18A doesn't have a swap.


> Granted, I use swap a lot more in normal Forth programs, but as I know
> that it is an expensive operation on b16, I code around this fact, and
> it is possible.

Indeed it is.

I have done my share of micro-code back in the day when we actually 
designed computers with boards rather than just a chip on a board. 
Micro-code is something you spend a lot of time on to optimize.  I guess 
low level forth is similar.

Rick

[toc] | [prev] | [next] | [standalone]


#16749

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-26 19:44 -0400
Message-ID<k6f710$5ie$1@speranza.aioe.org>
In reply to#16741
"Bernd Paysan" <bernd.paysan@gmx.de> wrote in message
news:1766437.ADrNJNi4aN@sunwukong.fritz.box...
> Alex McDonald wrote:
> > It's possible to swap using XORs, but I'd be interested to see the
> > F18A equivalent. For some code sequences, I don't see how it can be
> > avoided; A B - is not the same as B A -
>
> But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r com
> r> +c".
>
> Swap itself is "over >r nip r>".  You don't need to shuffle the stack
> that often.  In the 2k bytes Triceps program, there are in total 5
> swaps.  There are a lot more dups, overs (most of them as 2dups), >rs,
> and r>s.
>
> Granted, I use swap a lot more in normal Forth programs, but as I know
> that it is an expensive operation on b16, I code around this fact, and
> it is possible.
>

The question is if 'nip' is effective or useful enough to have it
as a primitive or MISC/RISC instruction.

Swap is a whole bunch of things ...

swap   1 roll
swap   nup drip take
swap   tuck drop
swap   dup rot rot drop
swap   dup take
swap   over take
swap   over rot drop
swap   dup rot nip
swap   dup -rot drop
swap   over >r >r drop r> r>
swap   over >r nip r>
swap   over dup 3 roll drop drop
swap   0 pick take
swap   1 pick take
swap   >r >r 2r>
swap   2>r r> r>
swap   over >r xor r@ xor r>
swap   over dup >r xor xor r>
swap   >r dup r> xor dup >r xor dup r> xor
swap   dup >r xor dup r> xor dup >r xor r>

etc.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#16754

Fromrickman <gnuarm@gmail.com>
Date2012-10-26 19:59 -0400
Message-ID<k6f84m$vpi$1@dont-email.me>
In reply to#16749
On 10/26/2012 7:44 PM, Rod Pemberton wrote:
> "Bernd Paysan"<bernd.paysan@gmx.de>  wrote in message
> news:1766437.ADrNJNi4aN@sunwukong.fritz.box...
>> Alex McDonald wrote:
>>> It's possible to swap using XORs, but I'd be interested to see the
>>> F18A equivalent. For some code sequences, I don't see how it can be
>>> avoided; A B - is not the same as B A -
>>
>> But there is no -.  On the b16, "-" is "com +c", and "swap -" is ">r com
>> r>  +c".
>>
>> Swap itself is "over>r nip r>".  You don't need to shuffle the stack
>> that often.  In the 2k bytes Triceps program, there are in total 5
>> swaps.  There are a lot more dups, overs (most of them as 2dups),>rs,
>> and r>s.
>>
>> Granted, I use swap a lot more in normal Forth programs, but as I know
>> that it is an expensive operation on b16, I code around this fact, and
>> it is possible.
>>
>
> The question is if 'nip' is effective or useful enough to have it
> as a primitive or MISC/RISC instruction.
>
> Swap is a whole bunch of things ...
>
> swap   1 roll
> swap   nup drip take
> swap   tuck drop
> swap   dup rot rot drop
> swap   dup take
> swap   over take
> swap   over rot drop
> swap   dup rot nip
> swap   dup -rot drop
> swap   over>r>r drop r>  r>
> swap   over>r nip r>
> swap   over dup 3 roll drop drop
> swap   0 pick take
> swap   1 pick take
> swap>r>r 2r>
> swap   2>r r>  r>
> swap   over>r xor r@ xor r>
> swap   over dup>r xor xor r>
> swap>r dup r>  xor dup>r xor dup r>  xor
> swap   dup>r xor dup r>  xor dup>r xor r>
>
> etc.
>
>
> Rod Pemberton

Yeah, thanks for the update.

I think the point is that nip is an instruction on the B16... I may give 
that a think...

Rick

[toc] | [prev] | [next] | [standalone]


#16683

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-25 00:42 -0400
Message-ID<k6afoh$laf$1@speranza.aioe.org>
In reply to#16649
"Alex McDonald" <blog@rivadpm.com> wrote in message
news:172e6a33-bc52-4fe9-8867-88f69c2954f8@10g2000vbu.googlegroups.com...
> On Oct 24, 1:43 pm, "Rod Pemberton" <do_not_h...@notemailnotz.cnm>
> wrote:
> > "rickman" <gnu...@gmail.com> wrote in message
> > news:k672lj$82o$1@dont-email.me...
> > On 10/22/2012 9:51 PM, Rod Pemberton wrote:
> > > "Brad Eckert"<hwfw...@gmail.com> wrote in message
> > >news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
...

> > > >> I've been thinking about Novix style processors like the RTX2000.
> > > >> There are many Forth sequences that can be compacted into one
> > > >> instruction, so with a good optimizer the chip can execute several
> > > >> Forth (source) primitives in one machine cycle. I suspect though
> > > >> that such optimization opportunities are the exception rather than
> > > >> the rule.
>
> > > > The question for both you (and Rick) is if you create new, faster,
> > > > more powerful, multiple operation instructions, how do you ensure
> > > > they are used? Without an optimizer, it's likely the instruction
> > > > will have a low instruction frequency. I.e., a person is unlikely to
> > > > use it.  In which case, there is no point in using or implementing
> > > > it. (This is repeated later in a reply to Rick.)
>
> > > We aren't talking about Forth coding really. We are talking about the
> > > assembly language for a machine. I don't think instructions will go
> > > unused just because they are mapped to Forth in a more complicated way
> > > than 1 to 1 (or 1/2 to 1).
>
> > > >> Does anyone here have a feel for the correspondence between Forth
> > > >> source primitives and generated code?
>
> > > > Generally, Forth's built using "primitives" or low-level words
> > > > generally need 30 to 40 or so. I kept track of how many are needed
> > > > for certain Forths. There are a few posts by me to c.l.f. with
> > > > counts and specific words used.
>
> > > Don't confuse Forth low level primitives (which are really HLL
> > > primitives selected to be convenient for the programmer writing a
> > > Forth) and assembly language which has to be selected in part based
> > > on what is practical and efficient to implement. Chuck's machine only
> > > uses 32 opcodes and you can get by with as few as 16.
>
> > For temporary reasons, most of my Forth stack operators, like DUP SWAP
> > etc, aren't currently low-level Forth "primitives". They're implemented
> > in high-level Forth using an even lower set of actual stack
> > "primitives", which I'll call sub-operators for this thread. These
> > sub-operators could be considered to be "stack assembly instructions"
> > for my Forth.  Currently, only >R and R> are actually "primitives" coded
> > in C. Eventually, DUP DROP SWAP OVER will be actual low-level Forth
> > "primitives", as they once were.
>
> > So, a word like OVER can be coded in many ways. I have 22 different
> > definitions just for OVER in a list, and one can construct many more.
> > E.g.,
>
> > : OVER >R DUP R> SWAP ;
> > : OVER 1 PICK ;
> > : OVER SWAP TUCK ;
> > : OVER SWAP DUP -ROT ;
> > : OVER NUP SWAP ;
> > etc.
>
> > Which definition you choose affects how fast OVER executes and depends
> > on what operations you have available and on how fast each of those
> > operations are.
>
> > E.g., if >R >R DUP SWAP and -ROT are all very fast machine instructions
> > for your processor, then ">R DUP R> SWAP" and "SWAP DUP -ROT" should
> > be fast sequences. And, one sequence will be faster than the other,
> > depending on how fast each instruction is. But, if -ROT is not a machine
> > instruction, e.g., perhaps coded as "ROT ROT" or "SWAP >R SWAP R>",
> > or -ROT is a very slow machine instruction, then "SWAP DUP -ROT" is
> > more expensive than ">R DUP R> SWAP".
>
> > The instructions also need to be balanced acrossed many Forth words.
> > Originally, I had the 2xxx series words defined in terms of the simpler
> > stack operators, like SWAP DUP etc. When I converted to sub-operators,
> > the number of operations per definition dropped dramatically for some of
> > the 2xxx definitions. Since the sub-operators are primitives, some of
> > the 2xxx definitions became faster. However, after converting all of the
> > simpler stack operators to sub-operators, a few of the non-primitive
> > simple stack operators ended up with more operations per definition.
> > So some words became much faster, while others became slightly slower.
> > And, the former primitives became real slow, but they'll be converted
> > back eventually.
>
> > Let's take a look at 2OVER for my Forth interpreter. I could easily
> > implement it as a Forth low-level "primitive". Or, I could define it in
> > high-level Forth. Or, I could define it in high-level Forth in terms of
> > sub-operators which are actually the low-level "primitives".
>
> > In high-level Forth, 2OVER can be defined:
>
> > : 2OVER 2>R 2DUP 2R> 2SWAP ;
>
> > In terms of my sub-operators, my 2OVER definition has ten words it's
> > definition. Clearly, that's many more words than the four in definition
> > above. But, 2>R 2R> 2DUP and 2SWAP are also implemented in high-level
> > Forth using the same set of sub-operators. So, 2>R 2R> 2DUP and 2SWAP
> > aren't "primitives" nor are they a single sub-operator each.
>
> > The 2>R sequence has six items.
> > The 2DUP sequence has six items.
> > The 2R> sequence has six items.
> > The 2SWAP sequence has eight items.
>
> > The 2OVER sequence has ten items.
>
> > So, 2OVER is only 10 items using a single sequence of sub-operators
> > instead of a total 28 items using four sequences of sub-operators for
> > the four words in the high-level definition. If the 2>R 2DUP 2R> and
> > 2SWAP words are defined in terms of standard Forth words:
>
> > The 2>R sequence has three items.
> > The 2DUP sequence has ten items over multiple words.
> > The 2R> sequence has three items.
> > The 2SWAP sequence has twelve items over multiple words.
>
> > In this case, the counts changed, but it just happens that it's total is
> > 28 also... Usually, it's more. Of course, you'd rather have a 2DUP of
> > six items instead of ten, i.e., balance. If the instructions are
> > unbalanced, then heavy use of a single Forth word will slow the code
> > speed way down.  I.e., many 2DUP's of ten items is much worse than
> > many 2DUP's with six items, even if other used words are made slightly
> > slower, like 2>R and 2R>. Of course, you don't know if your user's code
> > will follow the measured instruction frequencies or not. But, at this
> > point, you're the only user ...
>
> > As a primitive, 2OVER will be a small C routine which is compiled to
> > optimized, machine code.
>
>

IIRC, you're the guy who says I snip too much.  Well, it's all there.
Reformatted.  I don't know who is going to read through all that or
enjoy scrolling for six pages just to get to this, which is almost
another topic ...

> I believe the minimal set of primitives out of which all stack
> juggling words can be built is DUP DROP SWAP >R R>.
>
> : over         >r dup r> swap            ;
> : nip          swap drop                 ;
> : tuck         swap over                 ;
> : rot          >r swap r> swap           ;
> : -rot         swap >r swap r>           ;
> : 2swap        rot >r rot r>             ;
> : 2dup         over over                 ;
> : 2over        2>r 2dup 2r> 2swap        ;
> : 2drop        drop drop                 ;
> : 2nip         2swap 2drop               ;
> : 2rot         2>r 2swap 2r> 2swap       ;
> : r@           r> dup >r                 ;
> : 2>r          swap >r >r                ;
> : 2r>          r> r> swap                ;
> : 2r@          2r> 2dup 2>r              ;


Oops, I'm missing the 2r@ definition using 2xxx words ...

Yes.  But, the minimal set is not likely to be the fastest solution.

E.g., 2r@ calls 2r> which calls 3 words, 2>r which calls 3 words, and 2dup
which calls over twice which calls four other words for 8 words, plus the
overhead for calling DOCOL/ENTER SEMIS/EXIT.  A single, optimized
definition for 2r@ can be shorter and therefore faster:

: 2r@  r> r> dup >r swap dup >r ;

That's only trivially faster, one operation less, assuming DUP DROP SWAP
>R R> are primitives.  Unfortunately, I can't cite my 2r@ in sub-operators
as faster.  It's slightly worse.  Other definitions are improved more.  My
sub-operators improve some of the 2xxx words quite a bit.

> In your "sub-operator" notation, such sequences can be greatly
> simplified. Assuming an addressable stack (that is, we don't need to
> POP to get entries off the stack to access them, such as on the x86)
> there are 6 of these operations, 3 for each stack. SGET S[n] and RGET
> R[n] fetch entries from the stacks by fixed offset from a stack
> pointer, SPUT S[n] and RPUT R[n] store entries on the stacks, and SPTR
> + n and RPTR+ n adjust the stack pointers. In the example below, only
> the S operators are shown, since the rstack operations and other
> juggling have been optimised away. [reg] is a virtual register.
>
> : 2OVER 2>R 2DUP 2R> 2SWAP ;
> block ( S: 4 -- 6 R: 0 -- 0 )
> [1] block-begin [2]
> [5] sget S[2]
> [4] sget S[3]
> [5] sput S[-2]
> [4] sput S[-1]
> [3] sptr+ -2
> [2] block-end [1]

Well, I haven't done a 3-level analysis, but did do a 2-level direct
read/write of the data stack.  I'm using a highly modified version, by me,
of Peter Sovietov's Forth Wizard in Javascript.

Anyway, I concluded that the 2-level direct read/write wasn't all that
optimal.  I found a different combination of pushing/popping which I believe
to provide better overall results.  But, eventually, SWAP DROP DUP R> >R
will be primitives again.  So, this may be all for naught.

For Forthers here, SGET and SPUT would be like PICK and the uncommon
opposite operation: PLACE.


Rod Pemberton


[toc] | [prev] | [next] | [standalone]


#16685

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-10-25 00:55 -0400
Message-ID<k6agfv$mnr$1@speranza.aioe.org>
In reply to#16683
"Rod Pemberton" <do_not_have@notemailnotz.cnm> wrote in message
news:k6afoh$laf$1@speranza.aioe.org...

Correction:

> Oops, I'm missing the 2r@ definition using 2xxx words ...
>
> Yes.  But, the minimal set is not likely to be the fastest solution.
>
> E.g., 2r@ calls 2r> which calls 3 words, 2>r which calls 3 words, and 2dup
> which calls over twice which calls four other words for 8 words, plus the
> overhead for calling DOCOL/ENTER SEMIS/EXIT.  A single, optimized
> definition for 2r@ can be shorter and therefore faster:
>
> : 2r@  r> r> dup >r swap dup >r ;
>
> That's only trivially faster, one operation less, assuming DUP DROP SWAP
> >R R> are primitives.  Unfortunately, I can't cite my 2r@ in sub-operators
> as faster.  It's slightly worse.  Other definitions are improved more.  My
> sub-operators improve some of the 2xxx words quite a bit.

Sorry, I miscounted as 7 vs. 8.  It's 7 vs. 14.  So, it's 50% faster.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#16694

FromAlex McDonald <blog@rivadpm.com>
Date2012-10-25 04:49 -0700
Message-ID<afad494f-e1ea-43d6-b2a1-65e82111ef8d@g18g2000vbf.googlegroups.com>
In reply to#16683
On Oct 25, 5:38 am, "Rod Pemberton" <do_not_h...@notemailnotz.cnm>
wrote:
> "Alex McDonald" <b...@rivadpm.com> wrote in message
>
> news:172e6a33-bc52-4fe9-8867-88f69c2954f8@10g2000vbu.googlegroups.com...>

[snip]

> > I believe the minimal set of primitives out of which all stack
> > juggling words can be built is DUP DROP SWAP >R R>.
>
> > : over         >r dup r> swap            ;
> > : nip          swap drop                 ;
> > : tuck         swap over                 ;
> > : rot          >r swap r> swap           ;
> > : -rot         swap >r swap r>           ;
> > : 2swap        rot >r rot r>             ;
> > : 2dup         over over                 ;
> > : 2over        2>r 2dup 2r> 2swap        ;
> > : 2drop        drop drop                 ;
> > : 2nip         2swap 2drop               ;
> > : 2rot         2>r 2swap 2r> 2swap       ;
> > : r@           r> dup >r                 ;
> > : 2>r          swap >r >r                ;
> > : 2r>          r> r> swap                ;
> > : 2r@          2r> 2dup 2>r              ;
>
> Oops, I'm missing the 2r@ definition using 2xxx words ...
>
> Yes.  But, the minimal set is not likely to be the fastest solution.
>
> E.g., 2r@ calls 2r> which calls 3 words, 2>r which calls 3 words, and 2dup
> which calls over twice which calls four other words for 8 words, plus the
> overhead for calling DOCOL/ENTER SEMIS/EXIT.  A single, optimized
> definition for 2r@ can be shorter and therefore faster:
>
> : 2r@  r> r> dup >r swap dup >r ;
>
> That's only trivially faster, one operation less, assuming DUP DROP SWAP>R R> are primitives.  Unfortunately, I can't cite my 2r@ in sub-operators
>
> as faster.  It's slightly worse.  Other definitions are improved more.  My
> sub-operators improve some of the 2xxx words quite a bit.
>

The idea is to use the primitives. 2r@ above is

: 2r@  r> r> dup >r swap dup >r ;
block ( S: 0 -- 2 R: 2 -- 2 )
[1] block-begin [2]
[5] rget R[0]
[4] rget R[1]
[7] sput [5] S[-2]
[6] sput [4] S[-1]
[3] sptr+ -2
[2] block-end

Using the compounds 2r> and so on compiles to;

: 2r@ 2r> 2dup 2>r ;
     ^
Warning -4100 2r@ is redefined
block ( S: 0 -- 2 R: 2 -- 2 )
[1] block-begin [2]
[5] rget R[0]
[4] rget R[1]
[7] sput [5] S[-2]
[6] sput [4] S[-1]
[3] sptr+ -2
[2] block-end

i.e. exactly the same sequence of primitives.

What I'm suggesting is that you re-write all the operations in terms
of simpler, non-ANS Forth words like SGET, SPUT and so on. The
examples above were simply there to show that DUP DROP SWAP >R and R>
could be used to build all the others.


> > In your "sub-operator" notation, such sequences can be greatly
> > simplified. Assuming an addressable stack (that is, we don't need to
> > POP to get entries off the stack to access them, such as on the x86)
> > there are 6 of these operations, 3 for each stack. SGET S[n] and RGET
> > R[n] fetch entries from the stacks by fixed offset from a stack
> > pointer, SPUT S[n] and RPUT R[n] store entries on the stacks, and SPTR
> > + n and RPTR+ n adjust the stack pointers. In the example below, only
> > the S operators are shown, since the rstack operations and other
> > juggling have been optimised away. [reg] is a virtual register.
>
> > : 2OVER 2>R 2DUP 2R> 2SWAP ;
> > block ( S: 4 -- 6 R: 0 -- 0 )
> > [1] block-begin [2]
> > [5] sget S[2]
> > [4] sget S[3]
> > [5] sput S[-2]
> > [4] sput S[-1]
> > [3] sptr+ -2
> > [2] block-end [1]
>
> Well, I haven't done a 3-level analysis, but did do a 2-level direct
> read/write of the data stack.  I'm using a highly modified version, by me,
> of Peter Sovietov's Forth Wizard in Javascript.
>
> Anyway, I concluded that the 2-level direct read/write wasn't all that
> optimal.  I found a different combination of pushing/popping which I believe
> to provide better overall results.  But, eventually, SWAP DROP DUP R> >R
> will be primitives again.  So, this may be all for naught.
>
> For Forthers here, SGET and SPUT would be like PICK and the uncommon
> opposite operation: PLACE.

There's a difference between SGET and PICK.

: x 3 pick ;
block ( S: 3 -- 4 R: 0 -- 0 )
[1] block-begin [2]
[4] sget S[2]
[5] sput [4] S[-1]
[3] sptr+ -1
[2] block-end

n PICK puts its results on the stack; SGET and SPUT fetch and store
into virtual registers, which eventually get translated at the machine
code level into real registers; or they may be further optimised
away.

The output I'm showing here is from my experimental SSA based Forth
optimiser that I've been working on intermittently for a couple of
years, based in part on Anton Ertl's RAFTS papers. Progress is slow, I
must admit, but as you can see there are minor differences in the
output between my original post and this one, so work proceeds. Even
in this state it does an excellent and optimal job of removing all the
stack juggling as it refers to each stack entry only once, and it
reveals the issues you showed upthread with the various
implementations of OVER. They're obviously all equivalent, and reduce
to the same sequence of simple "sub-operators".

If the target supports PUSH and POP or if the target doesn't have an
addressable stack and only has PUSH and POP, the translation into SSA
could be changed. For instance, generate an SPOP that is the
equivalent of SGET S[0] SPTR+ 1 and an SPUSH that is SPUT S[-1] SPTR+
-1 . I use the intermediate form shown to allow addressing the stack
on an x86 as an array (hence the S[n] and R[n] notation for elements
of the stack). It makes the SSA analysis easier at the expense of a
slightly more complex code generator if PUSH and POP are to be
generated from these types of sequences.

>
> Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#17072 — addressable stack, was [Re: RTX2000 optimization]

From"Rod Pemberton" <do_not_have@notemailnotz.cnm>
Date2012-11-05 14:04 -0500
Subjectaddressable stack, was [Re: RTX2000 optimization]
Message-ID<k792cn$ba2$1@speranza.aioe.org>
In reply to#16649
"Alex McDonald" <blog@rivadpm.com> wrote in message
news:172e6a33-bc52-4fe9-8867-88f69c2954f8@10g2000vbu.googlegroups.com...

> [...] addressable stack [...]

Topic.

> [...] Assuming an addressable stack (that is, we don't need to
> POP to get entries off the stack to access them, such as on the x86)
> there are 6 of these operations, 3 for each stack. SGET S[n] and RGET
> R[n] fetch entries from the stacks by fixed offset from a stack
> pointer, SPUT S[n] and RPUT R[n] store entries on the stacks, and
> SPTR + n and RPTR+ n adjust the stack pointers. In the example below,
> only the S operators are shown, since the rstack operations and other
> juggling have been optimised away. [reg] is a virtual register.
>
> : 2OVER 2>R 2DUP 2R> 2SWAP ;
> block ( S: 4 -- 6 R: 0 -- 0 )
> [1] block-begin [2]
> [5] sget S[2]
> [4] sget S[3]
> [5] sput S[-2]
> [4] sput S[-1]
> [3] sptr+ -2
> [2] block-end [1]
>

It appears the optimizer you're using is only implementing the stack as
directly readable/writable memory.  It's not treating the stack as a LIFO
stack.  The optimizer is reducing everything into direct reads/writes and
stack adjustments.  Is this optimization/reducation causing the optimizer to
pass over some useful assembly instructions?  I.e., if reading or writing
the the 2nd, 3rd, 4th, etc stack item, a 'move' and adjustment of 'sp' are
required.  But, if writing the the 1st stack item, a 'push' (or 'pop') does
both the 'move' and adjustment of 'sp'.  How do you plan to emit pushes and
pops on the TOS (top of stack) for the situations where they are available?

I also noticed that some of the sequences you posted add a negative
adjustment to 'sp' at the _end_ of the sequence for allocating new items.  I
would think you would want to do negative 'sp' adjustments at the _start_ of
the sequence.


Rod Pemberton


[toc] | [prev] | [next] | [standalone]


#16666

Fromrickman <gnuarm@gmail.com>
Date2012-10-24 15:44 -0400
Message-ID<k69gf5$hic$1@dont-email.me>
In reply to#16646
On 10/24/2012 8:47 AM, Rod Pemberton wrote:
> "rickman"<gnuarm@gmail.com>  wrote in message
> news:k672lj$82o$1@dont-email.me...
>> On 10/22/2012 9:51 PM, Rod Pemberton wrote:
>>> "Brad Eckert"<hwfwguy@gmail.com>   wrote in message
>>> news:e1479cfa-a969-40ab-b20c-096d82601657@googlegroups.com...
>>>> I've been thinking about Novix style processors like the RTX2000.
>>>> There are many Forth sequences that can be compacted into one
>>>> instruction, so with a good optimizer the chip can execute several
>>>> Forth (source) primitives in one machine cycle. I suspect though
>>>> that such optimization opportunities are the exception rather than the
>>>> rule.
>>>>
>>>
>>> The question for both you (and Rick) is if you create new, faster, more
>>> powerful, multiple operation instructions, how do you ensure they are
>>> used?  Without an optimizer, it's likely the instruction will have a low
>>> instruction frequency.  I.e., a person is unlikely to use it.  In which
>>> case, there is no point in using or implementing it.
>>> (This is repeated later in a reply to Rick.)
>>
>> We aren't talking about Forth coding really.  We are talking about the
>> assembly language for a machine.  I don't think instructions will go
>> unused just because they are mapped to Forth in a more complicated way
>> than 1 to 1 (or 1/2 to 1).
>>
>>>> Does anyone here have a feel for the correspondence between Forth
>>>> source primitives and generated code?
>>>
>>> Generally, Forth's built using "primitives" or low-level words generally
>>> need 30 to 40 or so.  I kept track of how many are needed for certain
>>> Forths.  There are a few posts by me to c.l.f. with counts and specific
>>> words used.
>>
>> Don't confuse Forth low level primitives (which are really HLL
>> primitives selected to be convenient for the programmer writing a Forth)
>> and assembly language which has to be selected in part based on what is
>> practical and efficient to implement.  Chuck's machine only uses 32
>> opcodes and you can get by with as few as 16.
>>
>
> For temporary reasons, most of my Forth stack operators, like DUP SWAP etc,
> aren't currently low-level Forth "primitives".  They're implemented in
> high-level Forth using an even lower set of actual stack "primitives", which
> I'll call sub-operators for this thread.  These sub-operators could be
> considered to be "stack assembly instructions" for my Forth.  Currently,
> only>R and R>  are actually "primitives" coded in C.  Eventually, DUP DROP
> SWAP OVER will be actual low-level Forth "primitives", as they once were.

Coded in 'C'?  Perhaps I missed something.  I was talking about a Forth 
like CPU.  Are you referring to a Forth compiler on a PC?


> So, a word like OVER can be coded in many ways.  I have 22 different
> definitions just for OVER in a list, and one can construct many more.  E.g.,
>
>   : OVER>R DUP R>  SWAP ;
>   : OVER 1 PICK ;
>   : OVER SWAP TUCK ;
>   : OVER SWAP DUP -ROT ;
>   : OVER NUP SWAP ;
> etc.

Yes it can, but why?  Over is a very simple construct for either a Forth 
like CPU or a Forth compiler.


> Which definition you choose affects how fast OVER executes and depends on
> what operations you have available and on how fast each of those operations
> are.
>
> E.g., if>R>R DUP SWAP and -ROT are all very fast machine instructions for
> your processor, then ">R DUP R>  SWAP" and "SWAP DUP -ROT" should be fast
> sequences.  And, one sequence will be faster than the other, depending on
> how fast each instruction is.  But, if -ROT is not a machine instruction,
> e.g., perhaps coded as "ROT ROT" or "SWAP>R SWAP R>", or -ROT is a very
> slow machine instruction, then "SWAP DUP -ROT" is more expensive than ">R
> DUP R>  SWAP".

I don't follow that a four instruction sequence should be "fast" in any 
real sense.  I don't follow at all why you would have -ROT as a 
primitive and not OVER.


> The instructions also need to be balanced acrossed many Forth words.
> Originally, I had the 2xxx series words defined in terms of the simpler
> stack operators, like SWAP DUP etc.  When I converted to sub-operators, the
> number of operations per definition dropped dramatically for some of the
> 2xxx definitions.  Since the sub-operators are primitives, some of the 2xxx
> definitions became faster.  However, after converting all of the simpler
> stack operators to sub-operators, a few of the non-primitive simple stack
> operators ended up with more operations per definition.  So some words
> became much faster, while others became slightly slower.  And, the former
> primitives became real slow, but they'll be converted back eventually.

Not sure why you are looking at this.  If you want to optimize the 
implementation, why not look at what is used most often and start by 
optimizing those operators?  That is why I went to Koopman's book.  Not 
much return on optimizing things that aren't used so much.


> Let's take a look at 2OVER for my Forth interpreter.  I could easily
> implement it as a Forth low-level "primitive".  Or, I could define it in
> high-level Forth. Or, I could define it in high-level Forth in terms of
> sub-operators which are actually the low-level "primitives".
>
> In high-level Forth, 2OVER can be defined:
>
>   : 2OVER 2>R 2DUP 2R>  2SWAP ;
>
> In terms of my sub-operators, my 2OVER definition has ten words it's
> definition.  Clearly, that's many more words than the four in definition
> above.  But, 2>R 2R>  2DUP and 2SWAP are also implemented in high-level
> Forth using the same set of sub-operators.  So, 2>R 2R>  2DUP and 2SWAP
> aren't "primitives" nor are they a single sub-operator each.
>
> The 2>R sequence has six items.
> The 2DUP sequence has six items.
> The 2R>  sequence has six items.
> The 2SWAP sequence has eight items.
>
> The 2OVER sequence has ten items.
>
> So, 2OVER is only 10 items using a single sequence of sub-operators instead
> of a total 28 items using four sequences of sub-operators for the four words
> in the high-level definition.  If the 2>R 2DUP 2R>  and 2SWAP words are
> defined in terms of standard Forth words:
>
> The 2>R sequence has three items.
> The 2DUP sequence has ten items over multiple words.
> The 2R>  sequence has three items.
> The 2SWAP sequence has twelve items over multiple words.
>
> In this case, the counts changed, but it just happens that it's total is 28
> also...  Usually, it's more.  Of course, you'd rather have a 2DUP of six
> items instead of ten, i.e., balance.  If the instructions are unbalanced,
> then heavy use of a single Forth word will slow the code speed way down.
> I.e., many 2DUP's of ten items is much worse than many 2DUP's with six
> items, even if other used words are made slightly slower, like 2>R and 2R>.
> Of course, you don't know if your user's code will follow the measured
> instruction frequencies or not.  But, at this point, you're the only user
> ....
>
> As a primitive, 2OVER will be a small C routine which is compiled to
> optimized, machine code.

Ok, so this is running on some other computer.

Rick

[toc] | [prev] | [next] | [standalone]


Page 7 of 9 — ← Prev page 1 2 3 4 5 6 [7] 8 9  Next page →

Back to top | Article view | comp.lang.forth


csiph-web