Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > alt.os.development > #8747 > unrolled thread
| Started by | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| First post | 2015-09-11 19:11 -0400 |
| Last post | 2015-11-01 09:07 -0800 |
| Articles | 20 on this page of 142 — 9 participants |
Back to article view | Back to alt.os.development
C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-11 19:11 -0400
Re: C parser ramblings, language design, etc "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-12 05:10 -0700
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 04:41 -0400
Re: C parser ramblings, language design, etc "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 02:37 -0700
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 07:20 -0400
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-13 18:37 +0100
Re: C parser ramblings, language design, etc "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 21:03 -0700
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-12 13:49 +0100
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 06:27 -0400
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-13 19:02 +0100
Re: C parser ramblings, language design, etc "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 15:10 -0700
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 20:30 -0400
Re: C parser ramblings, language design, etc "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 20:07 -0700
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-14 22:29 +0100
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-20 18:48 -0400
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-22 20:04 +0100
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-30 20:32 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2015-12-27 13:51 +0000
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-12-27 17:35 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-16 17:09 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-16 14:30 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-16 21:40 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-18 07:34 -0500
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-18 08:37 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-18 23:35 +0000
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-19 12:52 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-19 13:09 +0000
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-19 21:46 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-21 11:15 +0000
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-21 21:48 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-22 15:16 +0000
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-22 23:02 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-01-23 17:04 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-27 15:27 +0000
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-02-29 09:26 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-02-29 19:02 +0000
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-03-02 20:48 +0000
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-03-09 22:29 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-03-17 16:41 +0000
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-03-29 09:00 +0200
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-03-09 22:29 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-03-17 17:36 +0000
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-03-29 09:01 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-02 10:13 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-04-10 19:33 +0200
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-04-15 16:27 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-26 08:10 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-03 20:14 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-04 09:17 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-21 21:10 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-29 08:44 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-29 12:05 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-29 12:28 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-06-01 19:22 +0200
OT: Obama in the UK (was: C parser ramblings, language design, etc) James Harris <james.harris.1@gmail.com> - 2016-04-26 08:46 +0100
Re: OT: Obama in the UK (was: C parser ramblings, language design, etc) Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-26 23:58 -0400
Re: OT: Obama in the UK (was: C parser ramblings, language design, etc) Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 00:08 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-04-27 06:40 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 05:25 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-04-27 17:45 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 18:25 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-02 23:44 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-03 16:56 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-04 09:23 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-04 17:22 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-06 11:47 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-06 19:34 -0400
Re: OT: Obama in the UK "James Harris" <james.harris.1@gmail.com> - 2016-05-07 12:51 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-07 19:20 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-08 01:46 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-07 21:43 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-08 09:36 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-08 15:34 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-10 11:25 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-10 17:32 -0400
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-10 18:25 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-15 00:06 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-07 21:32 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-06-02 07:41 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-06-02 16:21 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-20 14:59 +0100
Re: OT: Obama in the UK "wolfgang kern" <nowhere@never.at> - 2016-05-20 21:13 +0200
Re: OT: Obama in the UK "wolfgang kern" <nowhere@never.at> - 2016-05-20 21:17 +0200
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-21 08:04 +0100
Re: OT: Obama in the UK Bernhard Schornak <schornak@web.de> - 2016-05-21 21:10 +0200
Re: OT: Obama in the UK "wolfgang kern" <nowhere@never.at> - 2016-05-21 11:16 +0200
Re: OT: Obama in the UK Bernhard Schornak <schornak@web.de> - 2016-05-21 21:11 +0200
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-05-20 22:51 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-21 07:29 +0100
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-06-02 08:33 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-06-02 16:21 -0400
UK pronunciation was, [Re: OT: Obama in the UK, was C lang parser ramblings] Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-06-11 14:34 -0400
Re: UK pronunciation was, [Re: OT: Obama in the UK, was C lang parser ramblings] James Harris <james.harris.1@gmail.com> - 2016-07-24 18:28 +0100
Re: UK pronunciation was, [Re: OT: Obama in the UK, was C lang parser ramblings] Rod Pemberton <NoHaveNotOne@bcczxcfrze.cam> - 2016-07-25 21:54 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-07-24 18:21 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfrze.cam> - 2016-07-25 21:54 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-07-26 10:02 +0100
Re: OT: Obama in the UK Rod Pemberton <NoHaveNotOne@bcczxcfrze.cam> - 2016-07-26 07:56 -0400
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-07-30 12:26 +0100
Re: OT: Obama in the UK Bernhard Schornak <schornak@web.de> - 2016-05-03 20:14 +0200
Re: OT: Obama in the UK James Harris <james.harris.1@gmail.com> - 2016-05-04 09:25 +0100
Re: OT: Obama in the UK Bernhard Schornak <schornak@web.de> - 2016-05-21 21:11 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-26 07:59 +0100
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 00:01 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-27 09:28 +0100
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 05:42 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-27 17:16 +0100
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-04-27 18:05 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-27 09:29 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-04-27 09:42 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-03 20:14 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-07 14:34 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-21 21:10 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-29 11:27 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-06-01 19:22 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-07 16:34 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-05-21 21:10 +0200
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-05-29 12:10 +0100
Re: C parser ramblings, language design, etc Bernhard Schornak <schornak@web.de> - 2016-06-01 19:22 +0200
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-25 18:29 -0500
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-26 11:47 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-27 17:08 +0000
Re: C parser ramblings, language design, etc "wolfgang kern" <nowhere@never.at> - 2016-01-28 11:54 +0100
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-28 11:47 +0000
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-21 10:13 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-25 18:29 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-27 17:15 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-27 17:31 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-28 11:07 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-19 22:28 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-22 15:04 +0000
Re: C parser ramblings, language design, etc Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-25 17:48 -0500
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2016-01-27 17:16 +0000
Re: cows and bulls game, was [Re: C parser ramblings, language design, etc] "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-12-27 21:19 -0500
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-30 20:32 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2015-10-05 15:13 +0100
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-10-05 21:56 -0400
Re: C parser ramblings, language design, etc James Harris <james.harris.1@gmail.com> - 2015-12-27 13:56 +0000
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-10-21 19:46 -0400
Re: C parser ramblings, language design, etc "James Harris" <james.harris.1@gmail.com> - 2015-09-12 15:38 +0100
Re: C parser ramblings, language design, etc "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 07:10 -0400
Re: C parser ramblings, language design, etc "s_dubrovich@yahoo.com" <s_dubrovich@yahoo.com> - 2015-11-01 09:07 -0800
Page 1 of 8 [1] 2 3 4 5 6 7 8 Next page →
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-11 19:11 -0400 |
| Subject | C parser ramblings, language design, etc |
| Message-ID | <op.x4tmp0ssyfako5@localhost> |
You would think that C, being created long ago, would be very easy
to parse with truly simple techniques, but things like optional
braces, 'while' used as both 'while' and 'do-while', i.e., which
'while' goes with 'do' when nested, base types comprising multiple
words, e.g., "long long", implicit int's, no keyword for typedef
usage, conflict with implicit int's and typedef's, ... etc all
complicate very simple C parsers.
E.g., since braces are optional in C, there isn't much useful
information for parsing, from statements like this:
do while(x--); while(x--);
That's confusing for a human and a simple parser. It could be
recognized by both as either of these:
do { while(x--); } while(x--);
do {} while (x--); while(x--);
Which 'while' goes with the 'do'? Only the first one is correct,
but how do you determine that from the braceless syntax above?
You get a 'while' prior to the 'while' you need.
Technically, the second example needs a semicolon, but that
doesn't help any with properly parsing the braceless example.
'while' is overloaded or ambiguous. It means two different
things at different places, and somehow the parser or compiler
must distinquish and keep track of which 'while' goes with
which 'do'. Nesting can throw the dual-keyword do-while loop
tracking out of whack. Adding nested while just complicates
it further. You almost need a stack and a state-machine.
That would've been much simpler to parse and emit code for if a
different keyword was used at the bottom of the loop or if the
condition was placed at the top of the loop for a do-while.
I.e., this is much easier to parse, since the 'while' keyword
isn't overloaded and it's easier to match paired 'do' and 'end':
do end(x--); while(x--);
do while(x--); end(x--);
Now, you know that every end is at the same scope as the 'do'
and you're not going to have "spurious" 'end's' inbetween.
/* condition syntax at top of bottom test/exit loop */
do(x--); { while(x--); }
That's even simpler.
I think the goal for any language should not only be LALR(1)
but also be able to be compiled in a single-pass with no
backtracking and with exceptionally easy parsing. That's why
my language uses space delimited parsing, like Forth, with
character directed parsing. I.e., a character in front of
each keyword, operator, etc in the language, tells the parser
what is coming next.
E.g., the hello world program for my language:
^main
{
$"Hello_World!" ~puts
}
I.e., '^' for declaring a procedure, '~' for calling a
procedure, etc. The parser doesn't have to determine that
"main" is a procedure or an identifier or that it's being
declared here, or that "puts" is a procedure which is being
called, or that "Hello_World" is a string, etc. The character
directive tells it that. Save info. It's like a built-in
AST, somewhat. I also went to a little bit of extra work
to support curly braces without following that pattern, which
would've required something to follow a brace. Unfortunately,
I just realized that I should've provided for a C style
semicolon too ';' for multipled statements per block.
I also did some work on my larger C parser project. I merged
in a bunch of similar code from other projects which helped
to fill the program out somewhat more than it was. Now, it
emits nicely reformatted C, an AST as text, and it's emitting
some very rudimentary assembly code, but it has a long, ...
long way to go to properly compile C code. Fortunately,
I have a number of other simpler C projects.
I also recently made some great progress on my OS-like
environment for DOS that provides a console window for DOS.
It's coded for DJGPP (GCC) C and uses CWSDPMI features.
I now just need to find some use for it. At the moment it
doesn't provide any advantage over a standard DOS command
line, except I can watch PM interrupts, which are only called
by DJGPP apps, and install PM interrupts and keep them active.
So, no RM interrupts are being monitored at the moment. The
Windows 98/SE console windows was much faster than DOS. So,
improving speed is one potential option for a purpose.
Unfortunately, switching RM interrupts back to PM has overhead
which could reduce any real gains from 32-bit PM C code.
Unfortunately, CWSDPMI doesn't use v86, unless VCPI is
installed, which would've reflected all RM interrupts to PM.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
[toc] | [next] | [standalone]
| From | "Alexei A. Frounze" <alexfrunews@gmail.com> |
|---|---|
| Date | 2015-09-12 05:10 -0700 |
| Message-ID | <76dff2ed-c7a8-42fe-8eab-82d1dce5487f@googlegroups.com> |
| In reply to | #8747 |
On Friday, September 11, 2015 at 4:11:45 PM UTC-7, Rod Pemberton wrote:
> You would think that C, being created long ago, would be very easy
> to parse with truly simple techniques, but things like optional
> braces, 'while' used as both 'while' and 'do-while',
Just as in your language, there are things that have multiple
meanings, each applicable in its own context. But there's no
ambiguity here as long as you maintain minimal state while
parsing. Your brain does that (maintaining state) too when
you listen to speech or read text and it does more complex
things.
> i.e., which
> 'while' goes with 'do' when nested,
I see no problem with that in terms of parsing. If you get do,
you then expect to get while. Just like with parens and braces/
brackets. There's an opening/leading one and a closing/trailing
one.
> base types comprising multiple
> words, e.g., "long long",
Irritating, but straightforward.
> implicit int's,
Ugly and ill-conceived, IMO.
> no keyword for typedef
> usage,
What kind of usage?
> conflict with implicit int's and typedef's, ... etc all
> complicate very simple C parsers.
>
> E.g., since braces are optional in C, there isn't much useful
> information for parsing, from statements like this:
>
> do while(x--); while(x--);
>
> That's confusing for a human and a simple parser.
That is a very confusing way of saying "loop forever". :)
The parser (and the human) need to maintain some rather small
state to parse this out.
> It could be
> recognized by both as either of these:
>
> do { while(x--); } while(x--);
> do {} while (x--); while(x--);
>
> Which 'while' goes with the 'do'? Only the first one is correct,
> but how do you determine that from the braceless syntax above?
> You get a 'while' prior to the 'while' you need.
State!
> Technically, the second example needs a semicolon, but that
> doesn't help any with properly parsing the braceless example.
I think you've got all the semicolons in the above that are
required, so, I'm not sure why "the second example needs a
semicolon".
> 'while' is overloaded or ambiguous.
Not entirely ambiguous. If you look at the language, many
things have dual and even triple(?) use. There are some
factors: how many special symbols you have on your computer
(we now have 80+-100+ keys on our PC keyboards, but that
wasn't something everyone always had, hence trigraphs and
digraphs in the language). One could reserve more keywords
and make source code bulkier and compiler a tad mode larger
and complex too. That, however, also restricts the programmer's
choice of identifier names. So, both special symbols and
words are used with some balance, which some may find just
right, while others may find subjectively unreadable or less
readable than in other languages.
> It means two different
> things at different places, and somehow the parser or compiler
> must distinquish and keep track of which 'while' goes with
> which 'do'. Nesting can throw the dual-keyword do-while loop
> tracking out of whack. Adding nested while just complicates
> it further. You almost need a stack and a state-machine.
You do. Yo do need a state machine and some kind of stack.
Unless you're into esoteric stuff like brainfuck, which has no
form of recursion or nestedness in its syntax, you want your
language to provide facilities for nested and recursive things.
Even your mother tongue has it, despite not being designed
by wise engineers (even if there were wise engineers behind
spoken languages, the overall degree of chaos in every single
one of them indicates that wisdom and order weren't quite there
all the time). Some ideas and algorithms are most naturally
expressed in terms of recursion or induction. That's why
you have it, it's handy at times.
> That would've been much simpler to parse and emit code for if a
> different keyword was used at the bottom of the loop or if the
> condition was placed at the top of the loop for a do-while.
You could indeed use some sort of do/repeat-while() and
do/repeat-until(). But in any case you still need to parse
everything in between. And you are still not free to grab the
first while/until you see and use it to close the loop. If you
support comments and strings or string literals, you still need
to parse while maintaining state.
// this while is part of a comment
"this while is part of a string"
// "this while is part of a comment"
"this // while is part of a string"
Or do you propose to do something radical with strings and
comments as well so as to eliminate stateful parsing of them
and of everything around?
> I.e., this is much easier to parse, since the 'while' keyword
> isn't overloaded and it's easier to match paired 'do' and 'end':
>
> do end(x--); while(x--);
> do while(x--); end(x--);
>
> Now, you know that every end is at the same scope as the 'do'
> and you're not going to have "spurious" 'end's' inbetween.
This while/do issue must be a huge first world problem or
something. :) It's a rambleful one for sure. :)
> /* condition syntax at top of bottom test/exit loop */
> do(x--); { while(x--); }
>
> That's even simpler.
>
> I think the goal for any language should not only be LALR(1)
> but also be able to be compiled in a single-pass with no
> backtracking and with exceptionally easy parsing.
Certain improvements can be made and irregularities rooted
out. However, recursion is common and it requires special
treatment. Even your arithmetic expressions hide recursion.
> That's why
> my language uses space delimited parsing, like Forth, with
> character directed parsing. I.e., a character in front of
> each keyword, operator, etc in the language, tells the parser
> what is coming next.
>
> E.g., the hello world program for my language:
>
> ^main
> {
> $"Hello_World!" ~puts
> }
>
> I.e., '^' for declaring a procedure, '~' for calling a
> procedure, etc. The parser doesn't have to determine that
> "main" is a procedure or an identifier or that it's being
> declared here, or that "puts" is a procedure which is being
> called, or that "Hello_World" is a string, etc. The character
> directive tells it that. Save info. It's like a built-in
> AST, somewhat. I also went to a little bit of extra work
> to support curly braces without following that pattern, which
> would've required something to follow a brace. Unfortunately,
> I just realized that I should've provided for a C style
> semicolon too ';' for multipled statements per block.
What you describe is a way to simplify the language for the
compiler/interpreter developer. It may not necessarily be
helpful for the user of the language. If it involves
unnecessary typing or typing of characters whose typing cost
is higher than that of special characters in other languages,
it's just asking the programmer to do more work, which the
machine can do for him. When computers were big, expensive,
slow, with small memories and limited I/O, adapting things
for the machine made a lot more sense than adapting things
for the user. These days it's mostly the other way around
as I see it.
Alex
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-13 04:41 -0400 |
| Message-ID | <op.x4v7q5xkyfako5@localhost> |
| In reply to | #8758 |
On Sat, 12 Sep 2015 08:10:23 -0400, Alexei A. Frounze <alexfrunews@gmail.com> wrote:
> On Friday, September 11, 2015 at 4:11:45 PM UTC-7, Rod Pemberton wrote:
>> no keyword for typedef
>> usage,
>
> What kind of usage?
'union' and 'struct' keywords both declare the data type and declare
variables of said type. 'typedef' only declares the data type. So,
usage of a typedef without needing a keyword prior to it creates a
conflict between an identifier and typedef name in the grammar. This
also creates conflicts with implicit ints.
union abc{
...
}
union abc var0;
typedef int abc;
<here> abc var1;
There is no keyword at <here>, such as 'typedef'. That's the problem.
The grammar expects a keyword like 'union' or 'struct' there, but
there is nothing for a 'typedef'. This problem has to be solved by
some other method: expansion of 'abc' to 'int', symbol lookup tables,
state-machine to determine "missing" typedef, etc.
BTW, I've had this conversation with both you and James,
and multiple times with James going back *many* years both
here and comp.lang.misc: 2014, '13, '11, '8 ... There is
in-depth info, including links to discussions on the issue
by ANSI C X3J11 maintainers, in a couple of those threads.
>> 'while' is overloaded or ambiguous.
>
> Not entirely ambiguous. If you look at the language, many
> things have dual and even triple(?) use.
at least quadruple: 'static' keyword
>> It means two different
>> things at different places, and somehow the parser or compiler
>> must distinquish and keep track of which 'while' goes with
>> which 'do'. Nesting can throw the dual-keyword do-while loop
>> tracking out of whack. Adding nested while just complicates
>> it further. You almost need a stack and a state-machine.
>
> You do. Yo do need a state machine and some kind of stack.
I.e., it's not simple enough for such an old language.
Maybe, "trivial" is the word I'm looking for ...
>> That would've been much simpler to parse and emit code for if a
>> different keyword was used at the bottom of the loop or if the
>> condition was placed at the top of the loop for a do-while.
>
> You could indeed use some sort of do/repeat-while() and
> do/repeat-until(). But in any case you still need to parse
> everything in between. And you are still not free to grab the
> first while/until you see and use it to close the loop. If you
> support comments and strings or string literals, you still need
> to parse while maintaining state.
A simple counter can keep track of scope level. So, the next loop
ending keyword at the same scope level terminates the loop, if
the conflict between standalone 'while' and the 'while' of a do-while
didn't exist.
> // this while is part of a comment
> "this while is part of a string"
> // "this while is part of a comment"
> "this // while is part of a string"
>
> Or do you propose to do something radical with strings and
> comments as well so as to eliminate stateful parsing of them
> and of everything around?
An int flag for each (in comment, or in string) is sufficient.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
[toc] | [prev] | [next] | [standalone]
| From | "Alexei A. Frounze" <alexfrunews@gmail.com> |
|---|---|
| Date | 2015-09-13 02:37 -0700 |
| Message-ID | <44441c58-9675-43f0-a337-5e4f7f17947b@googlegroups.com> |
| In reply to | #8769 |
On Sunday, September 13, 2015 at 1:41:13 AM UTC-7, Rod Pemberton wrote:
> On Sat, 12 Sep 2015 08:10:23 -0400, Alexei A. Frounze <...@gmail.com> wrote:
>
> > On Friday, September 11, 2015 at 4:11:45 PM UTC-7, Rod Pemberton wrote:
>
> >> no keyword for typedef
> >> usage,
> >
> > What kind of usage?
>
> 'union' and 'struct' keywords both declare the data type and declare
> variables of said type. 'typedef' only declares the data type. So,
> usage of a typedef without needing a keyword prior to it creates a
> conflict between an identifier and typedef name in the grammar. This
> also creates conflicts with implicit ints.
>
> union abc{
> ...
> }
>
> union abc var0;
>
>
> typedef int abc;
>
> <here> abc var1;
>
> There is no keyword at <here>, such as 'typedef'. That's the problem.
> The grammar expects a keyword like 'union' or 'struct' there, but
> there is nothing for a 'typedef'. This problem has to be solved by
> some other method: expansion of 'abc' to 'int', symbol lookup tables,
> state-machine to determine "missing" typedef, etc.
Right. This is a method to create custom types or aliases for types.
And the problem is there because syntactically there's no
difference between a type name and a variable/function name and
either can occur at the beginning of a block (more freedom in C99
and C++ with their allowing to declare things almost everywhere).
If declarations had been designed with explicit delimiters (I think
you'd like that), this problem would've been long solved, e.g.:
a, b, c : blah;
or preferably for the compiler writer:
blah : a, b, c;
You don't need to know what blah is to determine that it is a
type.
There's a similar problem with sizeof. sizeof blah is sizeof
of an expression. sizeof(blah) can be either that same thing
(because parens don't alter expressions contained in them and
only specify order) or sizeof of a type, depending on what's
declared most recently, type blah or variable blah.
You need to fix sizeof as well.
> BTW, I've had this conversation with both you and James,
> and multiple times with James going back *many* years both
> here and comp.lang.misc: 2014, '13, '11, '8 ... There is
> in-depth info, including links to discussions on the issue
> by ANSI C X3J11 maintainers, in a couple of those threads.
>
> >> 'while' is overloaded or ambiguous.
> >
> > Not entirely ambiguous. If you look at the language, many
> > things have dual and even triple(?) use.
>
> at least quadruple: 'static' keyword
>
> >> It means two different
> >> things at different places, and somehow the parser or compiler
> >> must distinquish and keep track of which 'while' goes with
> >> which 'do'. Nesting can throw the dual-keyword do-while loop
> >> tracking out of whack. Adding nested while just complicates
> >> it further. You almost need a stack and a state-machine.
> >
> > You do. Yo do need a state machine and some kind of stack.
>
> I.e., it's not simple enough for such an old language.
> Maybe, "trivial" is the word I'm looking for ...
>
> >> That would've been much simpler to parse and emit code for if a
> >> different keyword was used at the bottom of the loop or if the
> >> condition was placed at the top of the loop for a do-while.
> >
> > You could indeed use some sort of do/repeat-while() and
> > do/repeat-until(). But in any case you still need to parse
> > everything in between. And you are still not free to grab the
> > first while/until you see and use it to close the loop. If you
> > support comments and strings or string literals, you still need
> > to parse while maintaining state.
>
> A simple counter can keep track of scope level. So, the next loop
> ending keyword at the same scope level terminates the loop, if
> the conflict between standalone 'while' and the 'while' of a do-while
> didn't exist.
There's no real conflict with while in terms of C syntax/grammar.
There's no need to guess or try/check anything complicated.
> > // this while is part of a comment
> > "this while is part of a string"
> > // "this while is part of a comment"
> > "this // while is part of a string"
> >
> > Or do you propose to do something radical with strings and
> > comments as well so as to eliminate stateful parsing of them
> > and of everything around?
>
> An int flag for each (in comment, or in string) is sufficient.
I'm not sure I get the idea. What I was trying to say is that
it seemed to me that you wanted to quickly locate the closing
while of do, before even parsing the statement that must exist
between do and while. If that's what you wanted, then you would
likely have a problem with comments or strings, which too can
contain while. To solve that problem with comments and strings
you'd employ some stateful parsing in order to note where a
comment or a string begins and ends. But then how's that
different or why should that be different from parsing
do statement while? Now, if you didn't want to fast forward from
do to while, what is the fuss about? while can be that statement
and all is well, except it's unlikely someone would want to
write that or be happy to read that.
Alex
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-13 07:20 -0400 |
| Message-ID | <op.x4we36pfyfako5@localhost> |
| In reply to | #8773 |
On Sun, 13 Sep 2015 05:37:49 -0400, Alexei A. Frounze <alexfrunews@gmail.com> wrote: > Now, if you didn't want to fast forward from > do to while, what is the fuss about? The 'while' as a statement is seen prior to the 'while' which goes with the 'do'. A simple parser/compiler has no way to determine which is which without braces for a compound statement. I.e., it can't determine what is or isn't a statement, just what are keywords and other language elements. Analysis of the AST or a grammar-based parser or some other solution would be required to eliminate the 'while' as a statement as being the 'while' paired with 'do'. I.e., C is too complicated for simple parsing and code generation from the language. Rod Pemberton -- Just how many texting and calendar apps does humanity need?
[toc] | [prev] | [next] | [standalone]
| From | "James Harris" <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-09-13 18:37 +0100 |
| Message-ID | <mt4c5n$voo$1@dont-email.me> |
| In reply to | #8780 |
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message news:op.x4we36pfyfako5@localhost... > On Sun, 13 Sep 2015 05:37:49 -0400, Alexei A. Frounze > <alexfrunews@gmail.com> wrote: > >> Now, if you didn't want to fast forward from >> do to while, what is the fuss about? > > The 'while' as a statement is seen prior to the 'while' > which goes with the 'do'. A simple parser/compiler has > no way to determine which is which without braces for > a compound statement. Yes, it can do this simply. I thought I already answered that. > I.e., it can't determine what > is or isn't a statement, just what are keywords and > other language elements. Analysis of the AST or a > grammar-based parser or some other solution would be > required to eliminate the 'while' as a statement as being > the 'while' paired with 'do'. I.e., C is too complicated > for simple parsing and code generation from the language. C is quite complicated to parse but not for the reasons you mention, IMO. James
[toc] | [prev] | [next] | [standalone]
| From | "Alexei A. Frounze" <alexfrunews@gmail.com> |
|---|---|
| Date | 2015-09-13 21:03 -0700 |
| Message-ID | <85913429-0fed-4492-a0b9-d2513a2f991c@googlegroups.com> |
| In reply to | #8773 |
On Sunday, September 13, 2015 at 2:37:50 AM UTC-7, Alexei A. Frounze wrote: ... > Right. This is a method to create custom types or aliases for types. > And the problem is there because syntactically there's no > difference between a type name and a variable/function name and > either can occur at the beginning of a block (more freedom in C99 > and C++ with their allowing to declare things almost everywhere). > If declarations had been designed with explicit delimiters (I think > you'd like that), this problem would've been long solved, e.g.: > > a, b, c : blah; > > or preferably for the compiler writer: > > blah : a, b, c; > > You don't need to know what blah is to determine that it is a > type. The latter form would need to be disambiguated from labels, however. It could be done as easily and as ugly as in DOS batch files, just put the colon before the label name: rem blah blah blah goto label rem blah blah blah :label Alex
[toc] | [prev] | [next] | [standalone]
| From | "James Harris" <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-09-12 13:49 +0100 |
| Message-ID | <mt16t1$bsu$1@dont-email.me> |
| In reply to | #8747 |
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
news:op.x4tmp0ssyfako5@localhost...
>
> You would think that C, being created long ago, would be very easy
> to parse with truly simple techniques, but things like optional
> braces, 'while' used as both 'while' and 'do-while', i.e., which
> 'while' goes with 'do' when nested,
Those are easy for a compiler to distinguish. Will explain what I mean
below.
> base types comprising multiple
> words, e.g., "long long",
Not too hard to handle, I would have thought.
> implicit int's,
I am not sure about how hard this would be to handle. It may be similar
to your long long case: just remember what's already been declared and
react accordingly.
Isn't C's rule that declaration mimics use much more of a problem for
parsing declarations? You don't necessarily know the type when you see
the identifier and the type can be split between typedef and where the
typedef is used.
For example,
typedef int T[4];
T a[5];
The declaration combines the [4] and the [5] together.
> no keyword for typedef usage,
My guess is that the compiler has to see the typedef first and,
effectively, adds the typedef name to its list of type-declaring
keywords. So after
typedef <type> T
it then regards "T" as it would "int" so when it sees the T in something
like
static T x
it knows that T is a type name. Not too hard to parse.
> conflict with implicit int's and typedef's, ... etc all
> complicate very simple C parsers.
>
> E.g., since braces are optional in C, there isn't much useful
> information for parsing, from statements like this:
>
> do while(x--); while(x--);
That's a good one. I had to stare at it for a while to work it out so I
have to agree that it's not easy for a human to parse.
> That's confusing for a human and a simple parser. It could be
> recognized by both as either of these:
>
> do { while(x--); } while(x--);
> do {} while (x--); while(x--);
>
> Which 'while' goes with the 'do'? Only the first one is correct,
> but how do you determine that from the braceless syntax above?
> You get a 'while' prior to the 'while' you need.
This should be easy for a compiler, IMO, as long as the compiler is
properly written. I say that because a compiler will work forward a
symbol at a time. After
do
the compiler will expect a statement. Note, a statement, not a {compound
statement}. Only *after* the statement will it expect a while clause.
What the compiler should look for, then, is
DO statement WHILE '(' expression ')' ';'
(That is from the link I will place at the end.)
As long as it sees a recognisable *statement* it will accept it, even if
that statement is a while loop. Easy.
> Technically, the second example needs a semicolon, but that
> doesn't help any with properly parsing the braceless example.
>
> 'while' is overloaded or ambiguous. It means two different
> things at different places, and somehow the parser or compiler
> must distinquish and keep track of which 'while' goes with
> which 'do'. Nesting can throw the dual-keyword do-while loop
> tracking out of whack. Adding nested while just complicates
> it further. You almost need a stack and a state-machine.
>
> That would've been much simpler to parse and emit code for if a
> different keyword was used at the bottom of the loop or if the
> condition was placed at the top of the loop for a do-while.
> I.e., this is much easier to parse, since the 'while' keyword
> isn't overloaded and it's easier to match paired 'do' and 'end':
>
> do end(x--); while(x--);
> do while(x--); end(x--);
>
> Now, you know that every end is at the same scope as the 'do'
> and you're not going to have "spurious" 'end's' inbetween.
>
> /* condition syntax at top of bottom test/exit loop */
> do(x--); { while(x--); }
>
> That's even simpler.
>
> I think the goal for any language should not only be LALR(1)
AIUI a *grammar* can be LALR(1) but not a language. There are many
grammars which can describe a given language.
> but also be able to be compiled in a single-pass
I am not even sure what single-pass or multiple-pass means any more.
When the term was first devised, AIUI, compilers could not keep all of
the source code in memory at a time. Consequently they read the source
code twice, say. The first time they read the source code they built
data structures from it. The second time they combined the data
structures with the source in order to emit the output code.
Those days are long gone. Compilers these days tend to read the source
once and then make multiple "passes" over internal data structures.
I guess you mean to emit code as soon as you see each part of the
source. That can cannot be done unless you have all the info you need at
each point. That might require the programmer to declare elements in a
given order such as constants and types first before code.
> with no backtracking
I think backtracking can be avoided with predictive parsing; that's what
parse tables are built to do (not that using parse tables is mandatory
or even necessarily a good idea; there are other ways to parse without
such tables).
> and with exceptionally easy parsing. That's why
> my language
Your own language is of interest. AFAIK this is the first time you have
written any specifics about it. This may have been more appropriate for
comp.lang.misc, though, and even although some people read both groups
you may get some further comments and interest there.
> uses space delimited parsing, like Forth, with
> character directed parsing. I.e., a character in front of
> each keyword, operator, etc in the language, tells the parser
> what is coming next.
Are there enough symbols in ASCII to express all the different syntactic
elements that you want to distinguish...? If you use too many won't the
code look ugly?
> E.g., the hello world program for my language:
>
> ^main
> {
> $"Hello_World!" ~puts
> }
>
> I.e., '^' for declaring a procedure, '~' for calling a
> procedure, etc. The parser doesn't have to determine that
> "main" is a procedure or an identifier or that it's being
> declared here, or that "puts" is a procedure which is being
> called, or that "Hello_World" is a string, etc. The character
> directive tells it that. Save info. It's like a built-in
> AST, somewhat.
Agreed. Once someone understands the symbols you use the above code
looks easy to read.
> I also went to a little bit of extra work
> to support curly braces without following that pattern, which
> would've required something to follow a brace. Unfortunately,
> I just realized that I should've provided for a C style
> semicolon too ';' for multipled statements per block.
>
> I also did some work on my larger C parser project. I merged
> in a bunch of similar code from other projects which helped
> to fill the program out somewhat more than it was. Now, it
> emits nicely reformatted C, an AST as text, and it's emitting
> some very rudimentary assembly code, but it has a long, ...
> long way to go to properly compile C code. Fortunately,
> I have a number of other simpler C projects.
It sounds as though you are making progress on it. There is a useful
grammar in
http://www.lysator.liu.se/%28nobg%29/c/ANSI-C-grammar-y.html
If you closely follow something like that it should make your life a lot
easier. There won't be any problem with where labels can go, what
statements can appear where, when braces are needed etc.
James
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-13 06:27 -0400 |
| Message-ID | <op.x4wcoua0yfako5@localhost> |
| In reply to | #8759 |
On Sat, 12 Sep 2015 08:49:21 -0400, James Harris <james.harris.1@gmail.com> wrote:
> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
> news:op.x4tmp0ssyfako5@localhost...
>> base types comprising multiple
>> words, e.g., "long long",
>
> Not too hard to handle, I would have thought.
That depends on the parser design. Most base types are a single word with
some qualifiers etc. But, both "long" and "long long" are base types.
"longlong" is not recognized, i.e., requiring whitespace in-between, without
regard for the length of or type of whitespace. You can end up with newlines
in-between. So, "long long" is fine for a "maximal munch" algorithm, but not
necessarily other parsers.
>> implicit int's,
>
> I am not sure about how hard this would be to handle. It may be similar
> to your long long case: just remember what's already been declared and
> react accordingly.
When typedef's were introduced, they created a conflict in the grammar.
The result was that typedef's were obsoleted from C. It's possible that
a non-grammar based parser may not have any issues.
> Isn't C's rule that declaration mimics use much more of a problem for
> parsing declarations? You don't necessarily know the type when you see
> the identifier and the type can be split between typedef and where the
> typedef is used.
>
> For example,
>
> typedef int T[4];
>
> T a[5];
>
> The declaration combines the [4] and the [5] together.
>
>> no keyword for typedef usage,
>
> My guess is that the compiler has to see the typedef first and,
> effectively, adds the typedef name to its list of type-declaring
> keywords. So after
>
> typedef <type> T
>
> it then regards "T" as it would "int" so when it sees the T in something
> like
>
> static T x
>
> it knows that T is a type name. Not too hard to parse.
See reply to Alexei.
Also, it requires symbol or lookup tables. For a lex and yacc style
parser, this creates problems because the lexer must communicate with
the parser, which they aren't designed to do.
>> conflict with implicit int's and typedef's, ... etc all
>> complicate very simple C parsers.
>>
>> E.g., since braces are optional in C, there isn't much useful
>> information for parsing, from statements like this:
>>
>> do while(x--); while(x--);
>
> That's a good one. I had to stare at it for a while to work it
> out so I have to agree that it's not easy for a human to parse.
:-)
Yeah, I had to look at it a few times even though I knew how I got
to it and what the correct interpretation was. If I saw that in
real code, I would pause for a moment. E.g., the do and first while
makes one think of an infinite for or while loop:
for(;;);
while(1);
do while(1); /* NOT a do-while infinite loop, incomplete */
do ; while(1); /* infinite loop, complete */
Is a space required before the semicolon after 'do' in the last
example? If not, then ...
do; while(1);
>> That's confusing for a human and a simple parser. It could be
>> recognized by both as either of these:
>>
>> do { while(x--); } while(x--);
>> do {} while (x--); while(x--);
>>
>> Which 'while' goes with the 'do'? Only the first one is correct,
>> but how do you determine that from the braceless syntax above?
>> You get a 'while' prior to the 'while' you need.
>
> This should be easy for a compiler, IMO, as long as the compiler is
> properly written. I say that because a compiler will work forward a
> symbol at a time.
Yes, 'long' then forward to another symbol 'long' ...
Where is the "long long"? ;-) I.e., the two 'longs' must be
combined in the AST, perhaps? This would've been easier to
parse with one keyword.
> After
>
> do
>
> the compiler will expect a statement. Note, a statement, not a {compound
> statement}. Only *after* the statement will it expect a while clause.
> What the compiler should look for, then, is
>
> DO statement WHILE '(' expression ')' ';'
>
> (That is from the link I will place at the end.)
>
> As long as it sees a recognisable *statement* it will accept it, even if
> that statement is a while loop. Easy.
That works for grammar based parsers. That make not work
for others which might detect a 'while' too early, without
an easy way to track scope level, such as braces for a compound
statement. They would need a method to track and pair do's
and while's correctly. Solutions like counters and flags
might not work, i.e., may need stack or state-machine.
>> and with exceptionally easy parsing. That's why
>> my language
>
> Your own language is of interest. AFAIK this is the first time you have
> written any specifics about it.
I've mentioned the techniques previously on c.l.m. I also recall posting
an assembly example that no one liked, probably a.l.a. It was RPN assembly
plus character directed parsing. I thought I posted a sample of the
higher-level language a while back, somewhere.
> This may have been more appropriate for
> comp.lang.misc, though, and even although some people read both groups
> you may get some further comments and interest there.
It's still ultra-primitive: if-else, loop, characters, strings, byte
integer, larger integer, Forth like functionality to access memory.
The interpreter version, instead of the compiled version, looks more
likely to be useful, at this point.
>> uses space delimited parsing, like Forth, with
>> character directed parsing. I.e., a character in front of
>> each keyword, operator, etc in the language, tells the parser
>> what is coming next.
>
> Are there enough symbols in ASCII to express all the different
> syntactic elements that you want to distinguish...?
That's an issue, but I haven't used too many so far.
I use this technique for a number of related projects.
The high-level language only uses fifteen. The interpreter
uses the same fifteen. The general x86 assembler uses thirteen.
It emits hex, not binary. The hex to binary app uses seventeen.
Another assembler only for the high-level language uses two.
It's a minimal assembler used specifically during development
with the high-level language until the general x86 assembler
becomes more complete ... or not.
> If you use too many won't the code look ugly?
Yes, that's an issue ... I changed a few around already.
They need to be both easily remembered and not too awkward to
view. I tried to use ones familiar to me from other languages,
where possible. If the language set becomes too large, then
this will become a problem, i.e., using a wrong char.
Obviously, this would be less noticeable for intermediate
stages of compiling, i.e., unseen assembly or for code
backends or even for CLIs with a small command set.
>> I also did some work on my larger C parser project.
>
> It sounds as though you are making progress on it. There is a useful
> grammar in
>
> [link]
>
> If you closely follow something like that it should make your life
> a lot easier. There won't be any problem with where labels can go,
> what statements can appear where, when braces are needed etc.
I have that grammar, and a version I updated. This doesn't follow
a grammar. It's more of a state-machine, based upon a few switch()
statements and numerous flags, perhaps more like two actually.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
[toc] | [prev] | [next] | [standalone]
| From | "James Harris" <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-09-13 19:02 +0100 |
| Message-ID | <mt4dje$5l3$1@dont-email.me> |
| In reply to | #8775 |
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
news:op.x4wcoua0yfako5@localhost...
> On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
> <james.harris.1@gmail.com> wrote:
>
>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>> news:op.x4tmp0ssyfako5@localhost...
...
>>> do while(x--); while(x--);
...
>>> That's confusing for a human and a simple parser. It could be
>>> recognized by both as either of these:
>>>
>>> do { while(x--); } while(x--);
>>> do {} while (x--); while(x--);
>>>
>>> Which 'while' goes with the 'do'? Only the first one is correct,
>>> but how do you determine that from the braceless syntax above?
>>> You get a 'while' prior to the 'while' you need.
>>
>> This should be easy for a compiler, IMO, as long as the compiler is
>> properly written. I say that because a compiler will work forward a
>> symbol at a time.
>
> Yes, 'long' then forward to another symbol 'long' ...
>
> Where is the "long long"? ;-) I.e., the two 'longs' must be
> combined in the AST, perhaps? This would've been easier to
> parse with one keyword.
A number of C types can be expressed by multiple keywords:
signed char
unsigned long
long double
And those can have qualifiers such as const and static and auto etc.
That's just something the compiler has to live with.
>> After
>>
>> do
>>
>> the compiler will expect a statement. Note, a statement, not a
>> {compound
>> statement}. Only *after* the statement will it expect a while clause.
>> What the compiler should look for, then, is
>>
>> DO statement WHILE '(' expression ')' ';'
>>
>> (That is from the link I will place at the end.)
>>
>> As long as it sees a recognisable *statement* it will accept it, even
>> if
>> that statement is a while loop. Easy.
>
> That works for grammar based parsers. That make not work
> for others which might detect a 'while' too early, without
> an easy way to track scope level, such as braces for a compound
> statement. They would need a method to track and pair do's
> and while's correctly. Solutions like counters and flags
> might not work, i.e., may need stack or state-machine.
There are areas where parsing C is difficult but this isn't one of them.
Consider a typical simple hand-written parser which includes the
following code.
if (token == DO) {
skip(); /* Skip the DO keyword */
statement_parse(); /* Parse the statement following DO */
if (token != WHILE) { /* Check for WHILE keyword */
.... exception handling code ....
} else {
skip(); /* Skip the WHILE keyword */
require(LPAREN);
expression_parse();
require(RPAREN);
etc.
In case it's not obvious the check (token != WHILE) on the fourth line
only takes place *after* statement_parse() returns. As such it would
find the *second* "while" in your example. The first "while" validly
begins a statement so it would be recognised in statement_parse() and
your example would parse properly and all without any special effort on
the part of the parser writer. I cannot understand why you think that is
hard to parse. You have mentioned a simple parser a few times. If you
still think there's a problem perhaps you can say what you mean by a
simple parser and why would it not parse the example as naturally as the
code above.
James
[toc] | [prev] | [next] | [standalone]
| From | "Alexei A. Frounze" <alexfrunews@gmail.com> |
|---|---|
| Date | 2015-09-13 15:10 -0700 |
| Message-ID | <4766fbce-135f-4957-991c-ba3d953026ec@googlegroups.com> |
| In reply to | #8790 |
On Sunday, September 13, 2015 at 11:02:08 AM UTC-7, James Harris wrote:
> "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message
> news:op.x4wcoua0yfako5@localhost...
> > On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
> > <...@gmail.com> wrote:
> >
> >> "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message
> >> news:op.x4tmp0ssyfako5@localhost...
>
> ...
>
> >>> do while(x--); while(x--);
>
> ...
>
> >>> That's confusing for a human and a simple parser. It could be
> >>> recognized by both as either of these:
> >>>
> >>> do { while(x--); } while(x--);
> >>> do {} while (x--); while(x--);
> >>>
> >>> Which 'while' goes with the 'do'? Only the first one is correct,
> >>> but how do you determine that from the braceless syntax above?
> >>> You get a 'while' prior to the 'while' you need.
> >>
> >> This should be easy for a compiler, IMO, as long as the compiler is
> >> properly written. I say that because a compiler will work forward a
> >> symbol at a time.
> >
> > Yes, 'long' then forward to another symbol 'long' ...
> >
> > Where is the "long long"? ;-) I.e., the two 'longs' must be
> > combined in the AST, perhaps? This would've been easier to
> > parse with one keyword.
>
> A number of C types can be expressed by multiple keywords:
>
> signed char
> unsigned long
> long double
>
> And those can have qualifiers such as const and static and auto etc.
> That's just something the compiler has to live with.
>
> >> After
> >>
> >> do
> >>
> >> the compiler will expect a statement. Note, a statement, not a
> >> {compound
> >> statement}. Only *after* the statement will it expect a while clause.
> >> What the compiler should look for, then, is
> >>
> >> DO statement WHILE '(' expression ')' ';'
> >>
> >> (That is from the link I will place at the end.)
> >>
> >> As long as it sees a recognisable *statement* it will accept it, even
> >> if
> >> that statement is a while loop. Easy.
> >
> > That works for grammar based parsers. That make not work
> > for others which might detect a 'while' too early, without
> > an easy way to track scope level, such as braces for a compound
> > statement. They would need a method to track and pair do's
> > and while's correctly. Solutions like counters and flags
> > might not work, i.e., may need stack or state-machine.
>
> There are areas where parsing C is difficult but this isn't one of them.
> Consider a typical simple hand-written parser which includes the
> following code.
>
> if (token == DO) {
> skip(); /* Skip the DO keyword */
> statement_parse(); /* Parse the statement following DO */
> if (token != WHILE) { /* Check for WHILE keyword */
> .... exception handling code ....
> } else {
> skip(); /* Skip the WHILE keyword */
> require(LPAREN);
> expression_parse();
> require(RPAREN);
> etc.
>
> In case it's not obvious the check (token != WHILE) on the fourth line
> only takes place *after* statement_parse() returns. As such it would
> find the *second* "while" in your example. The first "while" validly
> begins a statement so it would be recognised in statement_parse() and
> your example would parse properly and all without any special effort on
> the part of the parser writer. I cannot understand why you think that is
> hard to parse. You have mentioned a simple parser a few times. If you
> still think there's a problem perhaps you can say what you mean by a
> simple parser and why would it not parse the example as naturally as the
> code above.
And that's precisely what I do in Smaller C.
I have ParseStatement() that parses statements as defined in the
language standard. When ParseStatement() sees do, it calls itself
recursively and parses whatever the next statement there is,
including a while statement. Upon returning from this recursive
call, it just parses the trailing while for this do. I see no
problem here whatsoever. I don't understand why Rod can't handle
recursion here in the same fashion. If he doesn't want explicit
recursion, he can always use a stack data structure instead of
recursive calls on the CPU stack.
Alex
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-13 20:30 -0400 |
| Message-ID | <op.x4xfpfabyfako5@localhost> |
| In reply to | #8790 |
On Sun, 13 Sep 2015 14:02:06 -0400, James Harris <james.harris.1@gmail.com> wrote:
> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
> news:op.x4wcoua0yfako5@localhost...
>> On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
>> <james.harris.1@gmail.com> wrote:
>>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>>> news:op.x4tmp0ssyfako5@localhost...
>> That works for grammar based parsers. That make not work
>> for others which might detect a 'while' too early, without
>> an easy way to track scope level, such as braces for a compound
>> statement. They would need a method to track and pair do's
>> and while's correctly. Solutions like counters and flags
>> might not work, i.e., may need stack or state-machine.
>
> There are areas where parsing C is difficult but this isn't one of them.
> Consider a typical simple hand-written parser which includes the
> following code.
>
> if (token == DO) {
> skip(); /* Skip the DO keyword */
> statement_parse(); /* Parse the statement following DO */
> if (token != WHILE) { /* Check for WHILE keyword */
> .... exception handling code ....
> } else {
> skip(); /* Skip the WHILE keyword */
> require(LPAREN);
> expression_parse();
> require(RPAREN);
> etc.
>
> In case it's not obvious the check (token != WHILE) on the fourth line
> only takes place *after* statement_parse() returns. As such it would
> find the *second* "while" in your example. The first "while" validly
> begins a statement so it would be recognised in statement_parse() and
> your example would parse properly and all without any special effort on
> the part of the parser writer.
...
> I cannot understand why you think that is hard to parse.
Well, not all parsers work that way ... That's the point.
How do you parse C with a simple parser? That's not simple
enough, and it requires a certain type of design, i.e.,
grammar based and probably recursive descent, and uses
look-ahead.
So, your code example follows a grammar based design and looks
ahead by calling statement_parse() upon finding "token == DO".
Don't look ahead. So, if you can't check for "token == WHILE"
or "token != WHILE" in the "token == DO" section, and you can't
call statement_parse() there, what do you do for while, when
you reach the "token == WHILE" section? How do you determine
what type the 'while' is when you see 'while' or how do you
identify enough information prior to that point to know which
'while' it is at that point?
Think about single-pass where something is emitted immediately for each
'while' and only one token is known at a time. Also, you can't parse
ahead or look ahead as you're doing by calling statement_parse(), since
the code has already been parsed and stored as an AST of tokens, without
that being done. You're just emitting code for each token in the AST.
You can't backtrack or trace the AST either. So, when you come across
a 'while', how do you know what type of while and whether it's part of
a do-while or not? Assume this is a single-go, single-pass of the AST.
All you have are the prior 'do', any braces, 'while', and semicolons,
and/or the presence or absence of each thereof, to key off of, in order
to determine the type of 'while' each 'while' is.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
[toc] | [prev] | [next] | [standalone]
| From | "Alexei A. Frounze" <alexfrunews@gmail.com> |
|---|---|
| Date | 2015-09-13 20:07 -0700 |
| Message-ID | <db3e724a-4ac2-42af-882c-55c84f274c4a@googlegroups.com> |
| In reply to | #8798 |
On Sunday, September 13, 2015 at 5:30:36 PM UTC-7, Rod Pemberton wrote:
> On Sun, 13 Sep 2015 14:02:06 -0400, James Harris <...@gmail.com> wrote:
>
> > "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message
> > news:op.x4wcoua0yfako5@localhost...
> >> On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
> >> <...@gmail.com> wrote:
> >>> "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message
> >>> news:op.x4tmp0ssyfako5@localhost...
>
> >> That works for grammar based parsers. That make not work
> >> for others which might detect a 'while' too early, without
> >> an easy way to track scope level, such as braces for a compound
> >> statement. They would need a method to track and pair do's
> >> and while's correctly. Solutions like counters and flags
> >> might not work, i.e., may need stack or state-machine.
> >
> > There are areas where parsing C is difficult but this isn't one of them.
> > Consider a typical simple hand-written parser which includes the
> > following code.
> >
> > if (token == DO) {
> > skip(); /* Skip the DO keyword */
> > statement_parse(); /* Parse the statement following DO */
> > if (token != WHILE) { /* Check for WHILE keyword */
> > .... exception handling code ....
> > } else {
> > skip(); /* Skip the WHILE keyword */
> > require(LPAREN);
> > expression_parse();
> > require(RPAREN);
> > etc.
> >
> > In case it's not obvious the check (token != WHILE) on the fourth line
> > only takes place *after* statement_parse() returns. As such it would
> > find the *second* "while" in your example. The first "while" validly
> > begins a statement so it would be recognised in statement_parse() and
> > your example would parse properly and all without any special effort on
> > the part of the parser writer.
>
> ...
>
> > I cannot understand why you think that is hard to parse.
>
> Well, not all parsers work that way ... That's the point.
Then, maybe they should if you want them to be able to perform
equally well as in being able to parse the same code
unambiguously?
> How do you parse C with a simple parser? That's not simple
> enough, and it requires a certain type of design, i.e.,
> grammar based and probably recursive descent, and uses
> look-ahead.
>
> So, your code example follows a grammar based design and looks
> ahead by calling statement_parse() upon finding "token == DO".
> Don't look ahead.
Who looks ahead? Does consuming a token from the input stream
count as looking ahead? James is describing the same kind of
implementation as in Smaller C. There's no look ahead. You can
consume a token from the input stream of tokens and that's the
only thing you can look at until you consume another token.
There's no GetMe2ndFromCurrentToken() or anything like that.
However, it is virtually impossible to parse C code without
somehow storing/caching some tokens internally or altering
compiler state based on the tokens seen so far. I do not
consider that look ahead. That's more of look behind. But
for the sake of argument you can of course say that one
is just a time-shifted version of another. :) Are you trying
to say that?
> So, if you can't check for "token == WHILE"
> or "token != WHILE" in the "token == DO" section, and you can't
> call statement_parse() there, what do you do for while, when
> you reach the "token == WHILE" section?
Why can't you? You have to parse nested/recursive code. The
parser must be able to do that by itself being recursive in one
sense or another. It must maintain some state, which will not be
static but grow with every level of nesting. If you're trying to
avoid any form of recursion in the parser, you won't get far.
It's much like trying to parse HTML with regular expressions
alone.
Here's a good write-up with links:
http://blog.codinghorror.com/parsing-html-the-cthulhu-way/
IOW, if you're trying to avoid recursion, maybe you need
an almost non-recursive language instead of C and it would be
just fine? Like assembly or primitive dialects of BASIC?
> How do you determine
> what type the 'while' is when you see 'while' or how do you
> identify enough information prior to that point to know which
> 'while' it is at that point?
>
> Think about single-pass where something is emitted immediately for each
> 'while' and only one token is known at a time. Also, you can't parse
> ahead or look ahead as you're doing by calling statement_parse(), since
> the code has already been parsed and stored as an AST of tokens, without
> that being done. You're just emitting code for each token in the AST.
> You can't backtrack or trace the AST either. So, when you come across
> a 'while', how do you know what type of while and whether it's part of
> a do-while or not? Assume this is a single-go, single-pass of the AST.
> All you have are the prior 'do', any braces, 'while', and semicolons,
> and/or the presence or absence of each thereof, to key off of, in order
> to determine the type of 'while' each 'while' is.
You maintain state (whose size isn't fixed), AST or not, and consult it.
You can't do it any other way with C.
Alex
[toc] | [prev] | [next] | [standalone]
| From | "James Harris" <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-09-14 22:29 +0100 |
| Message-ID | <mt7e3r$u0k$1@dont-email.me> |
| In reply to | #8798 |
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
news:op.x4xfpfabyfako5@localhost...
> On Sun, 13 Sep 2015 14:02:06 -0400, James Harris
> <james.harris.1@gmail.com> wrote:
>
>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>> news:op.x4wcoua0yfako5@localhost...
>>> On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
>>> <james.harris.1@gmail.com> wrote:
>>>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>>>> news:op.x4tmp0ssyfako5@localhost...
>
>>> That works for grammar based parsers. That make not work
>>> for others which might detect a 'while' too early, without
>>> an easy way to track scope level, such as braces for a compound
>>> statement. They would need a method to track and pair do's
>>> and while's correctly. Solutions like counters and flags
>>> might not work, i.e., may need stack or state-machine.
>>
>> There are areas where parsing C is difficult but this isn't one of
>> them.
>> Consider a typical simple hand-written parser which includes the
>> following code.
>>
>> if (token == DO) {
>> skip(); /* Skip the DO keyword */
>> statement_parse(); /* Parse the statement following DO */
>> if (token != WHILE) { /* Check for WHILE keyword */
>> .... exception handling code ....
>> } else {
>> skip(); /* Skip the WHILE keyword */
>> require(LPAREN);
>> expression_parse();
>> require(RPAREN);
>> etc.
>>
>> In case it's not obvious the check (token != WHILE) on the fourth
>> line
>> only takes place *after* statement_parse() returns. As such it would
>> find the *second* "while" in your example. The first "while" validly
>> begins a statement so it would be recognised in statement_parse() and
>> your example would parse properly and all without any special effort
>> on
>> the part of the parser writer.
>
> ...
>
>> I cannot understand why you think that is hard to parse.
>
> Well, not all parsers work that way ... That's the point.
> How do you parse C with a simple parser? That's not simple
> enough, and it requires a certain type of design, i.e.,
> grammar based and probably recursive descent, and uses
> look-ahead.
That's about as simple as I imagine a parser could be and still parse C.
What parsing mechanism did you have in mind that's simpler? Detail,
please, as mentioned below.
> So, your code example follows a grammar based design and looks
> ahead by calling statement_parse() upon finding "token == DO".
No, the code shown doesn't look ahead at all. It just processes one
token at a time.
> Don't look ahead. So, if you can't check for "token == WHILE"
> or "token != WHILE" in the "token == DO" section, and you can't
> call statement_parse() there, what do you do for while, when
> you reach the "token == WHILE" section? How do you determine
> what type the 'while' is when you see 'while' or how do you
> identify enough information prior to that point to know which
> 'while' it is at that point?
After DO, C expects a statement. The language definition explains that.
So the parser has to look for a statement at that point.
Ah, perhaps you are thinking that the parser's logic should go
if symbol is FOR
handle for loop
else if symbol is SWITCH
handle switch construct
else if symbol is WHILE
parser is confused not knowing which WHILE it is
Is that what you have in mind? I have never seen a parser work that way.
I don't think it is feasible because elements of C, like other
languages, are context-sensitve. To make a simple example you similarly
cannot say
else if symbol is RIGHT_BRACE
parser is confused not knowing which LEFT_BRACE it pairs with
Parser's don't work that way, AFAIK, and I cannot see that they could.
They know the *context* and so can match closing tokens with the opening
ones.
Maybe that's not what you have in mind. If not could you show some
pseudocode to explain because I cannot see what you think is a problem.
> Think about single-pass where something is emitted immediately for
> each
> 'while' and only one token is known at a time.
I don't know much about single-pass compiling. Again, psome seudocode
(sic) would help to explain.
> Also, you can't parse
> ahead or look ahead as you're doing by calling statement_parse(),
> since
> the code has already been parsed and stored as an AST of tokens,
> without
> that being done.
Um, statement_parse() doesn't look ahead. It just processes the next
part of the source. When it reaches the end of a statement it returns.
That's all.
> You're just emitting code for each token in the AST.
No, the tokens in my example are not from the AST. They are what the
lexer reads from the source.
> You can't backtrack or trace the AST either. So, when you come across
> a 'while', how do you know what type of while and whether it's part of
> a do-while or not?
OK, I can think of a simple answer to that direct question:
A: If the WHILE is at the beginning of a statement it's a while loop.
Whereas if the WHILE comes after DO and a <statement> then it's the end
of a DO loop.
That's really the nub of it. *Where* the WHILE appears is the key but I
would stress that the code does *not* read the WHILE and wonder which
type it is. Rather, it knows where it is in a construct and looks for
specific tokens which are valid at that point.
> Assume this is a single-go, single-pass of the AST.
I didn't even posit an AST. The parse code was to process the source.
> All you have are the prior 'do', any braces, 'while', and semicolons,
> and/or the presence or absence of each thereof, to key off of, in
> order
> to determine the type of 'while' each 'while' is.
I am sure from past experience you won't be satisfied with the above.
:-( As I say, though, perhaps if you could explain more what you are
thinking of....
James
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-20 18:48 -0400 |
| Message-ID | <op.x499nwp5yfako5@localhost> |
| In reply to | #8805 |
On Mon, 14 Sep 2015 17:29:16 -0400, James Harris <james.harris.1@gmail.com> wrote:
> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
> news:op.x4xfpfabyfako5@localhost...
>> On Sun, 13 Sep 2015 14:02:06 -0400, James Harris
>> <james.harris.1@gmail.com> wrote:
>>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>>> news:op.x4wcoua0yfako5@localhost...
>>>> On Sat, 12 Sep 2015 08:49:21 -0400, James Harris
>>>> <james.harris.1@gmail.com> wrote:
>>>>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>>>>> news:op.x4tmp0ssyfako5@localhost...
>>>> That works for grammar based parsers. That make not work
>>>> for others which might detect a 'while' too early, without
>>>> an easy way to track scope level, such as braces for a compound
>>>> statement. They would need a method to track and pair do's
>>>> and while's correctly. Solutions like counters and flags
>>>> might not work, i.e., may need stack or state-machine.
>>>
>>> There are areas where parsing C is difficult but this isn't one of
>>> them.
>>> Consider a typical simple hand-written parser which includes the
>>> following code.
>>>
>>> if (token == DO) {
>>> skip(); /* Skip the DO keyword */
>>> statement_parse(); /* Parse the statement following DO */
>>> if (token != WHILE) { /* Check for WHILE keyword */
>>> .... exception handling code ....
>>> } else {
>>> skip(); /* Skip the WHILE keyword */
>>> require(LPAREN);
>>> expression_parse();
>>> require(RPAREN);
>>> etc.
>>>
>>> In case it's not obvious the check (token != WHILE) on the fourth
>>> line
>>> only takes place *after* statement_parse() returns. As such it would
>>> find the *second* "while" in your example. The first "while" validly
>>> begins a statement so it would be recognised in statement_parse() and
>>> your example would parse properly and all without any special effort
>>> on
>>> the part of the parser writer.
>>
>> ...
>>
>>> I cannot understand why you think that is hard to parse.
>>
>> Well, not all parsers work that way ... That's the point.
>> How do you parse C with a simple parser? That's not simple
>> enough, and it requires a certain type of design, i.e.,
>> grammar based and probably recursive descent, and uses
>> look-ahead.
>
> That's about as simple as I imagine a parser could be and still parse C.
> What parsing mechanism did you have in mind that's simpler? Detail,
> please, as mentioned below.
>
>> So, your code example follows a grammar based design and looks
>> ahead by calling statement_parse() upon finding "token == DO".
>
> No, the code shown doesn't look ahead at all. It just processes one
> token at a time.
Fine.
Your code proceeds to look at the next token, which may be a WHILE
token, while in the code construct for handling the DO token. What
if you can't proceed to the next token once a DO is found? You can
only look at DO. You can't parse a statement next. You must wait
for WHILE and then determine how it fits.
> After DO, C expects a statement. The language definition
> explains that. So the parser has to look for a statement
> at that point.
>
> Ah, perhaps you are thinking that the parser's logic should go
>
> if symbol is FOR
> handle for loop
> else if symbol is SWITCH
> handle switch construct
> else if symbol is WHILE
> parser is confused not knowing which WHILE it is
>
> Is that what you have in mind? I have never seen a parser
> work that way.
The logic is vastly different, but that's close enough
to the issue.
> I don't think it is feasible because elements of C, like other
> languages, are context-sensitve. To make a simple example you
> similarly cannot say
>
> else if symbol is RIGHT_BRACE
> parser is confused not knowing which LEFT_BRACE it pairs with
You don't need to know which brace goes with which brace as
long as all braces are properly paired. An up/down counter
can keep track of scope level and provide a count of missing
braces.
> Parser's don't work that way, AFAIK, and I cannot see that
> they could. They know the *context* and so can match closing
> tokens with the opening ones.
The goal is to determine the context from syntax as described.
> Maybe that's not what you have in mind. If not could you show some
> pseudocode to explain because I cannot see what you think is a problem.
I was hoping to use logic states, flags, rejection etc, to keep
track of where the 'statement' for the DO-WHILE is so that
intermediate WHILEs, those within the 'statement' can be rejected.
I believe this is only needed for the situtations without the
braces. There is very little syntax or presence/absence info to
key off of for an example like the one I posted:
do while(x++); while(x++);
___-----------____________
E.g., look at logic pulse above. Set a flag at start of
<statement>, and disable at the end. With braces present,
this can be done easily by detecting the presence of the
braces.
I.e., there is minimal information context at that point:
1) keyword do
2) keyword "other", i.e., first "while" here
3) whitespace
4) absence of braces prior to while
5) semicolon
So, in this case, the best I can figure is that I need
to detect the absence of the brace at a keyword after DO
but other than DO, then set a flag, disable at the
semicolon, and use a counter. If the counter is non-zero,
and a WHILE is found, and that WHILE is not in the rejected
region, i.e., the innermost braceless <statement>, then
the WHILE is part of a DO-WHILE.
... do do while(x++); while(x++); while(x++);
__________-----------___________________________
0 1 2 (rejected) 2 1 0
I haven't worked through this for other cases. I'm
just presenting this as an example of a possible way
to solve this. I don't believe this is needed for
when the braces are present.
>> You're just emitting code for each token in the AST.
>
> No, the tokens in my example are not from the AST.
> They are what the lexer reads from the source.
That sentence was part of me telling you about what to
do to mimic my situation. It was not describing my
understanding of your situation.
>> You can't backtrack or trace the AST either. So, when you come across
>> a 'while', how do you know what type of while and whether it's part of
>> a do-while or not?
>
> OK, I can think of a simple answer to that direct question:
>
> A: If the WHILE is at the beginning of a statement it's a while loop.
> Whereas if the WHILE comes after DO and a <statement> then it's the end
> of a DO loop.
Yes, that's true, but not the issue at hand.
> That's really the nub of it.
That's the "nub of it" for parsers designed the way you presented.
It's not the "nub of it" for the parser in question. The parser in
question has no ability to distinquish a WHILE in a <statement> from
the WHILE which comes after DO and a <statement>. It knows where
each while is, each do, but has no ability to link them. I.e.,
within the statement, if there is a WHILE, it needs an ability to
differentiate that while from the WHILE which is to come later and
goes with a DO. This parser needs to determine this purely from
the textual syntax, which is difficult to do without braces and
with overloaded uses of WHILE.
> *Where* the WHILE appears is the key but I would stress that the
> code does *not* read the WHILE and wonder which type it is. Rather,
> it knows where it is in a construct and looks for specific tokens
> which are valid at that point.
Ok... Think about construct-less parser. It doesn't know
where the constructs are and need to determine whether the next
WHILE found is within the <statement> of a DO-WHILE, the start
of a WHILE loop, or the end of a DO-WHILE, and which WHILE if
many are nested. This needs to be done without setting up a large
amount of tracking, stacks, binary trees, etc. You only have the
whitespace, keywords, and absence of braces etc that I mentioned
above to determine this. You can't check for a <statement> after
DO, but you need to identify and store some info in order to reject
any WHILE's found in the <statement> as being a match to the DO.
>> Assume this is a single-go, single-pass of the AST.
>
> I didn't even posit an AST. The parse code was to process the source.
Again, I was describing my situation.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
[toc] | [prev] | [next] | [standalone]
| From | "James Harris" <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-09-22 20:04 +0100 |
| Message-ID | <mts8k8$kgq$1@dont-email.me> |
| In reply to | #8823 |
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
news:op.x499nwp5yfako5@localhost...
> On Mon, 14 Sep 2015 17:29:16 -0400, James Harris
> <james.harris.1@gmail.com> wrote:
>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>> news:op.x4xfpfabyfako5@localhost...
...
>> That's about as simple as I imagine a parser could be and still parse
>> C.
>> What parsing mechanism did you have in mind that's simpler? Detail,
>> please, as mentioned below.
>>
>>> So, your code example follows a grammar based design and looks
>>> ahead by calling statement_parse() upon finding "token == DO".
>>
>> No, the code shown doesn't look ahead at all. It just processes one
>> token at a time.
>
> Fine.
>
> Your code proceeds to look at the next token, which may be a WHILE
> token, while in the code construct for handling the DO token. What
> if you can't proceed to the next token once a DO is found? You can
> only look at DO. You can't parse a statement next. You must wait
> for WHILE and then determine how it fits.
I have been trying to get my head round what you are trying to do - or,
more accurately, how you are trying to do it. I think I am getting
closer to understanding what you have in mind. I am not sure of all the
details, though. For example, when you see a WHILE token what other info
do you have available that you can look at?
Specifically, when you see WHILE are you thinking to scan
forwards/backwards for nearby tokens in the stream of tokens that
represents the source code?
You mentioned an AST earlier but IIUC from what you said below you are
not thinking of reading the WHILE token from a tree of some sort and
then looking for adjacent tree nodes to tell you what the context is.
...
>> Ah, perhaps you are thinking that the parser's logic should go
>>
>> if symbol is FOR
>> handle for loop
>> else if symbol is SWITCH
>> handle switch construct
>> else if symbol is WHILE
>> parser is confused not knowing which WHILE it is
>>
>> Is that what you have in mind? I have never seen a parser
>> work that way.
>
> The logic is vastly different, but that's close enough
> to the issue.
Noted.
>> I don't think it is feasible because elements of C, like other
>> languages, are context-sensitve. To make a simple example you
>> similarly cannot say
>>
>> else if symbol is RIGHT_BRACE
>> parser is confused not knowing which LEFT_BRACE it pairs with
>
> You don't need to know which brace goes with which brace as
> long as all braces are properly paired. An up/down counter
> can keep track of scope level and provide a count of missing
> braces.
Maybe a brace is a poor analog of the WHILE problem but even then
consider
do {
if ( condition }
The closing brace there cannot simply be taken as closing the prior left
brace construct. In other words, a compiler is there not just to parse
correct code but also needs to handle erroneous code. I don't know if
that affects your intentions or not.
>> Parser's don't work that way, AFAIK, and I cannot see that
>> they could. They know the *context* and so can match closing
>> tokens with the opening ones.
>
> The goal is to determine the context from syntax as described.
OK.
>> Maybe that's not what you have in mind. If not could you show some
>> pseudocode to explain because I cannot see what you think is a
>> problem.
>
> I was hoping to use logic states, flags, rejection etc, to keep
> track of where the 'statement' for the DO-WHILE is so that
> intermediate WHILEs, those within the 'statement' can be rejected.
OK. I think that can be done. As long as you keep track of context
somewhere it may be parseable.
> I believe this is only needed for the situtations without the
> braces. There is very little syntax or presence/absence info to
> key off of for an example like the one I posted:
>
> do while(x++); while(x++);
> ___-----------____________
>
> E.g., look at logic pulse above. Set a flag at start of
> <statement>, and disable at the end. With braces present,
> this can be done easily by detecting the presence of the
> braces.
Making up a variable "state" and an enum or set of defines could you not
set state = STATEMENT_START immediately after recognising the DO?
Then when you see WHILE you can see if state == STATEMENT_START. If it
does then the WHILE is a new statement.
After seeing the start of a statement you would have to set state to
something else.
Possibly after a statement you would set it to STATEMENT_END. Not sure
without working through it but in that case you would have
.... prior code in the if-elseif chain ....
} else if (token == WHILE) {
if (state == STATEMENT_START) {
/* Handle a while loop */
} else if (state == STATEMENT_END) {
/* Check that we are expecting to end a DO loop */
}
} else if (token == the next token type..... etc
> I.e., there is minimal information context at that point:
> 1) keyword do
> 2) keyword "other", i.e., first "while" here
> 3) whitespace
> 4) absence of braces prior to while
> 5) semicolon
>
> So, in this case, the best I can figure is that I need
> to detect the absence of the brace at a keyword after DO
> but other than DO, then set a flag, disable at the
> semicolon, and use a counter. If the counter is non-zero,
> and a WHILE is found, and that WHILE is not in the rejected
> region, i.e., the innermost braceless <statement>, then
> the WHILE is part of a DO-WHILE.
>
> ... do do while(x++); while(x++); while(x++);
> __________-----------___________________________
> 0 1 2 (rejected) 2 1 0
>
> I haven't worked through this for other cases. I'm
> just presenting this as an example of a possible way
> to solve this. I don't believe this is needed for
> when the braces are present.
C's syntax is a lot more complex than we normally use or allow for. I
think it would be a bad idea to just try to detect things like braces
and the absence thereof, or even to try to enumerate all of the possible
ways that a DO statement can continue. C will often throw up some other
virtually-never-used construct that is still perfectly valid.
For example, consider these
do for (;++i < 10;) while (j < 10) j++; while (0);
do if (c) whileloop: other: while (cond) x++; while (1);
In each case the inner while, if I have written them correctly, will be
a separate statement. No braces anywhere. Perhaps worse, braces could be
added that were nothing to do with the DO-WHILE loop.
...
>>> You can't backtrack or trace the AST either. So, when you come
>>> across
>>> a 'while', how do you know what type of while and whether it's part
>>> of
>>> a do-while or not?
>>
>> OK, I can think of a simple answer to that direct question:
>>
>> A: If the WHILE is at the beginning of a statement it's a while loop.
>> Whereas if the WHILE comes after DO and a <statement> then it's the
>> end
>> of a DO loop.
>
> Yes, that's true, but not the issue at hand.
>
>> That's really the nub of it.
>
> That's the "nub of it" for parsers designed the way you presented.
I don't agree. ISTM that that's the way C as a language is structured.
Nothing to do with a particular type of parser.
> It's not the "nub of it" for the parser in question. The parser in
> question has no ability to distinquish a WHILE in a <statement> from
> the WHILE which comes after DO and a <statement>. It knows where
> each while is, each do, but has no ability to link them. I.e.,
> within the statement, if there is a WHILE, it needs an ability to
> differentiate that while from the WHILE which is to come later and
> goes with a DO. This parser needs to determine this purely from
> the textual syntax, which is difficult to do without braces and
> with overloaded uses of WHILE.
>
>> *Where* the WHILE appears is the key but I would stress that the
>> code does *not* read the WHILE and wonder which type it is. Rather,
>> it knows where it is in a construct and looks for specific tokens
>> which are valid at that point.
>
> Ok... Think about construct-less parser. It doesn't know
> where the constructs are and need to determine whether the next
> WHILE found is within the <statement> of a DO-WHILE, the start
> of a WHILE loop, or the end of a DO-WHILE, and which WHILE if
> many are nested. This needs to be done without setting up a large
> amount of tracking, stacks, binary trees, etc. You only have the
> whitespace, keywords, and absence of braces etc that I mentioned
> above to determine this. You can't check for a <statement> after
> DO, but you need to identify and store some info in order to reject
> any WHILE's found in the <statement> as being a match to the DO.
I think you need to maintain an indication of state, as I suggest above,
especially because you cannot rely on the input source being
syntactically correct.
I think you will need to push and pop states so will need one stack data
structure.
James
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-09-30 20:32 -0400 |
| Message-ID | <op.x5sw4btxyfako5@localhost> |
| In reply to | #8837 |
On Tue, 22 Sep 2015 15:04:28 -0400, James Harris <james.harris.1@gmail.com> wrote:
> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
> news:op.x499nwp5yfako5@localhost...
>> On Mon, 14 Sep 2015 17:29:16 -0400, James Harris
>> <james.harris.1@gmail.com> wrote:
>>> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
>>> news:op.x4xfpfabyfako5@localhost...
>>> That's about as simple as I imagine a parser could be and still parse
>>> C.
>>> What parsing mechanism did you have in mind that's simpler? Detail,
>>> please, as mentioned below.
>>>
>>>> So, your code example follows a grammar based design and looks
>>>> ahead by calling statement_parse() upon finding "token == DO".
>>>
>>> No, the code shown doesn't look ahead at all. It just processes one
>>> token at a time.
>>
>> Fine.
>>
>> Your code proceeds to look at the next token, which may be a WHILE
>> token, while in the code construct for handling the DO token. What
>> if you can't proceed to the next token once a DO is found? You can
>> only look at DO. You can't parse a statement next. You must wait
>> for WHILE and then determine how it fits.
>
> I have been trying to get my head round what you are trying to do
> - or, more accurately, how you are trying to do it. I think I am
> getting closer to understanding what you have in mind. I am not
> sure of all the details, though.
...
> For example, when you see a WHILE token what other info
> do you have available that you can look at?
WHILE
and anything else I code up, such as flags.
Technically, I *could* walk the AST, but that is slows everything
down. I was attempting to do this without doing so, i.e., single
pass of AST to assembly, no walking or backtracking of ASTs whether
as a binary tree or linked-list.
> Specifically, when you see WHILE are you thinking to scan
> forwards/backwards for nearby tokens in the stream of tokens
> that represents the source code?
No, I'm attempting to *NOT* do that.
> You mentioned an AST earlier but IIUC from what you said below
> you are not thinking of reading the WHILE token from a tree of
> some sort and then looking for adjacent tree nodes to tell you
> what the context is.
The WHILE token already exists in a tree. This is to emit code
from the AST tree, without walking or backtracing the tree.
>>> I don't think it is feasible because elements of C, like other
>>> languages, are context-sensitve. To make a simple example you
>>> similarly cannot say
>>>
>>> else if symbol is RIGHT_BRACE
>>> parser is confused not knowing which LEFT_BRACE it pairs with
>>
>> You don't need to know which brace goes with which brace as
>> long as all braces are properly paired. An up/down counter
>> can keep track of scope level and provide a count of missing
>> braces.
>
> Maybe a brace is a poor analog of the WHILE problem but even
> then consider
I'm sending another post after this one on the braces issue.
> do {
> if ( condition }
>
> The closing brace there cannot simply be taken as closing the
> prior left brace construct. In other words, a compiler is there
> not just to parse correct code but also needs to handle erroneous
> code. I don't know if that affects your intentions or not.
I'm guessing that would break something somewhere ...
>>> Maybe that's not what you have in mind. If not could you show some
>>> pseudocode to explain because I cannot see what you think is a
>>> problem.
>>
>> I was hoping to use logic states, flags, rejection etc, to keep
>> track of where the 'statement' for the DO-WHILE is so that
>> intermediate WHILEs, those within the 'statement' can be rejected.
>
> OK. I think that can be done. As long as you keep track of context
> somewhere it may be parseable.
>
My other post on braces will demonstrate other solutions and mention
fails for the braces. I was thinking of something similar here.
>> I believe this is only needed for the situtations without the
>> braces. There is very little syntax or presence/absence info to
>> key off of for an example like the one I posted:
>>
>> do while(x++); while(x++);
>> ___-----------____________
>>
>> E.g., look at logic pulse above. Set a flag at start of
>> <statement>, and disable at the end. With braces present,
>> this can be done easily by detecting the presence of the
>> braces.
>
> Making up a variable "state" and an enum or set of defines
> could you not set state = STATEMENT_START immediately after
> recognising the DO?
That is one of the things in the list.
> Then when you see WHILE you can see if state ==
> STATEMENT_START. If it does then the WHILE is a new
> statement.
>
> After seeing the start of a statement you would have to set
> state to something else.
>
> Possibly after a statement you would set it to STATEMENT_END.
> Not sure without working through it but in that case you
> would have
>
> .... prior code in the if-elseif chain ....
> } else if (token == WHILE) {
> if (state == STATEMENT_START) {
> /* Handle a while loop */
> } else if (state == STATEMENT_END) {
> /* Check that we are expecting to end a DO loop */
> }
> } else if (token == the next token type..... etc
...
>> I.e., there is minimal information context at that point:
>> 1) keyword do
>> 2) keyword "other", i.e., first "while" here
>> 3) whitespace
>> 4) absence of braces prior to while
>> 5) semicolon
>>
>> So, in this case, the best I can figure is that I need
>> to detect the absence of the brace at a keyword after DO
>> but other than DO, then set a flag, disable at the
>> semicolon, and use a counter. If the counter is non-zero,
>> and a WHILE is found, and that WHILE is not in the rejected
>> region, i.e., the innermost braceless <statement>, then
>> the WHILE is part of a DO-WHILE.
>>
>> ... do do while(x++); while(x++); while(x++);
>> __________-----------___________________________
>> 0 1 2 (rejected) 2 1 0
>>
>> I haven't worked through this for other cases. I'm
>> just presenting this as an example of a possible way
>> to solve this. I don't believe this is needed for
>> when the braces are present.
>
> C's syntax is a lot more complex than we normally use
> or allow for. I think it would be a bad idea to just try
> to detect things like braces and the absence thereof, or
> even to try to enumerate all of the possible ways that
> a DO statement can continue. C will often throw up some
> other virtually-never-used construct that is still
> perfectly valid.
There are many C compilers with incorrect or incomplete
implementations. Even the old PCC compiler couldn't handle
structs correctly, or maybe it was pointers to structs.
> For example, consider these
>
> do for (;++i < 10;) while (j < 10) j++; while (0);
> do if (c) whileloop: other: while (cond) x++; while (1);
>
> In each case the inner while, if I have written them
> correctly, will be a separate statement. No braces anywhere.
> Perhaps worse, braces could be added that were nothing to
> do with the DO-WHILE loop.
That's right.
How do you detect the "implicit" but not-present braces?
You need to know where they are to generate branches and
branch targets.
>>>> You can't backtrack or trace the AST either. So,
>>>> when you come across a 'while', how do you know
>>>> what type of while and whether it's part of a
>>>> do-while or not?
>>>
>>> OK, I can think of a simple answer to that direct question:
>>>
>>> A: If the WHILE is at the beginning of a statement it's
>>> a while loop. Whereas if the WHILE comes after DO and
>>> a <statement> then it's the end of a DO loop.
>>
>> Yes, that's true, but not the issue at hand.
>>
>>> That's really the nub of it.
>>
>> That's the "nub of it" for parsers designed the way
>> you presented.
>
> I don't agree. ISTM that that's the way C as a language
> is structured. Nothing to do with a particular type of
> parser.
Your parser code is following the way C's grammar rules
are written. For mine, I'm attempting to implement other
solutions for the same issues.
Rod Pemberton
--
Just how many texting and calendar apps does humanity need?
Just how many food articles from neurotic millenials do we need?
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2015-12-27 13:51 +0000 |
| Message-ID | <n5oq88$qa9$1@dont-email.me> |
| In reply to | #8860 |
On 01/10/2015 01:32, Rod Pemberton wrote: > On Tue, 22 Sep 2015 15:04:28 -0400, James Harris > <james.harris.1@gmail.com> wrote: ... Replying after some absence. You may have already got past the points below. Feel free to ignore if you have. >> I have been trying to get my head round what you are trying to do >> - or, more accurately, how you are trying to do it. I think I am >> getting closer to understanding what you have in mind. I am not >> sure of all the details, though. > > ... > >> For example, when you see a WHILE token what other info >> do you have available that you can look at? > > WHILE > > and anything else I code up, such as flags. Does your approach boil down to needing to distinguishing WHILE in these three contexts? 1. at the start of a statement 2. at the conclusion of a DO construct 3. everywhere else (an error) If so you might get away with setting flags BUT because C is a recursive language you would need to save any such flag or flags on a stack if you made a call that could end up being recursive. > Technically, I *could* walk the AST, but that is slows everything > down. I was attempting to do this without doing so, i.e., single > pass of AST to assembly, no walking or backtracking of ASTs whether > as a binary tree or linked-list. If you can walk any sort of tree at this point then I have to suggest it would be better to use a normal parse mechanism. With the exception of a backtracking parser, parsing is one of the faster phases of a compiler so you may be ill advised to try to save time with your contextless approach. It may be a bit like trying to analyse a chess position based purely on the current board without forming a game tree: great and revolutionary if you could make it work but probably impractical, impossible to program, and more work in the long run. >> Specifically, when you see WHILE are you thinking to scan >> forwards/backwards for nearby tokens in the stream of tokens >> that represents the source code? > > No, I'm attempting to *NOT* do that. OK. >> You mentioned an AST earlier but IIUC from what you said below >> you are not thinking of reading the WHILE token from a tree of >> some sort and then looking for adjacent tree nodes to tell you >> what the context is. > > The WHILE token already exists in a tree. This is to emit code > from the AST tree, without walking or backtracing the tree. Normally an AST is the /output/ of a parse. If you have put a WHILE node in a tree and you don't know what type of WHILE it is then you may be making a rod for your own back when you try to distinguish such elements later. I think you would have had to have already worked out what type of WHILE it was before you could really call the tree an AST. ... >> For example, consider these >> >> do for (;++i < 10;) while (j < 10) j++; while (0); >> do if (c) whileloop: other: while (cond) x++; while (1); >> >> In each case the inner while, if I have written them >> correctly, will be a separate statement. No braces anywhere. >> Perhaps worse, braces could be added that were nothing to >> do with the DO-WHILE loop. > > That's right. > > How do you detect the "implicit" but not-present braces? > > You need to know where they are to generate branches and > branch targets. I would say that if you followed a correct C grammar then braces would appear naturally where a statement was expected. If a pair of braces was not present at a certain point then valid code would still have a statement there. Braces represent a compound statement but that's just one form of statement. ... > Your parser code is following the way C's grammar rules > are written. For mine, I'm attempting to implement other > solutions for the same issues. It's an interesting idea. You might get it to work. It reminds me, however, that when I was at school I tried to come up with an algorithm for marking a game that some people call bulls-and-cows. Rather than do it in a simple way I tried to do it a clever way that would be faster than the obvious solution. But it became so complicated that I eventually realised that it would be far better to do the marking in the most obvious way. Even if I had got my clever approach to work it would have been unmaintainable. James
[toc] | [prev] | [next] | [standalone]
| From | "Rod Pemberton" <boo@fasdfrewar.cdm> |
|---|---|
| Date | 2015-12-27 17:35 -0500 |
| Message-ID | <op.yabqczhayfako5@localhost> |
| In reply to | #9048 |
On Sun, 27 Dec 2015 08:51:24 -0500, James Harris <james.harris.1@gmail.com> wrote: > On 01/10/2015 01:32, Rod Pemberton wrote: >> On Tue, 22 Sep 2015 15:04:28 -0400, James Harris >> <james.harris.1@gmail.com> wrote: > Replying after some absence. You may have already got past > the points below. Feel free to ignore if you have. Nope ... Since the method I was using for branch labels is not correct, I'll probably rewrite the code to work with one of the two methods for generating them which I posted. I haven't decided as to which, yet, but one will be sufficient for known brace locations. I determined (and posted) the nesting levels in hope that it might help me decide ... Determining the locations of the "missing" or "implicit" braces, or which type of WHILE etc, will just be left for a future task since I don't plan on using a grammar rules based method or re-coding it at present. This part of an already rather large program that would mostly be discarded, if that was done. On merit alone, I'm hanging on to the code. Around the time you last posted, I figured out two paths I could take my command-line interpreter project, after realizing that half my OS code should be used in the command-line interpreter project and the other half should really be spun out as a pure, standalone OS. My initial path for my OS project was for an execution environment for DJGPP apps and deviated into the OS project. The first path for the command-line interpreters is towards a Windows 98/SE like OS, i.e., execution environment for DJGPP apps. I'm using "OS" loosely here. The other path would be towards a Linux kernel and/or shell on DOS. While the latter would be interesting, the DJGPP C compiler won't support full Linux apps on DOS. There is a bit of work in four or five areas I'd need to do to progress this. I _really_ haven't been motivated to do any programming since I figured out those paths for my projects ... I did do a new and small project with parsing and compiling BrainFuck, derived from my other BrainFuck and Forth and ITC projects. It combines ITC Brainfuck interpreter in C with the standard Brainfuck array in C. The ITC interpreter is a reduced, optimized, variation of my Forth interpreter. I already have other Brainfuck projects, like a single pass, no memory allocation, trivial optimizing, Brainfuck to pure C converter. Now, with a slightly different language front end or some enhancements to BrainFuck language, the ITC BF would create a very compact execution model, very embeddable. I did some work in the past on determining common BF sequences, which could provide slightly higher level language functionality. I'm wondering if I could combine it with another compact model, e.g., FSM, and whether I could find any purpose to do so, i.e., code reduction, code optimization. >>> I have been trying to get my head round what you are trying to do >>> - or, more accurately, how you are trying to do it. I think I am >>> getting closer to understanding what you have in mind. I am not >>> sure of all the details, though. >> >> ... >> >>> For example, when you see a WHILE token what other info >>> do you have available that you can look at? >> >> WHILE >> >> and anything else I code up, such as flags. > > Does your approach boil down to needing to distinguishing WHILE > in these three contexts? > > 1. at the start of a statement > 2. at the conclusion of a DO construct > 3. everywhere else (an error) (I'm not sure. The thread is somewhat stale now. My mental state on the issue is lost, but I seem to recall you asking something similar previously.) However, probably 1 and 2, but not 3. Issues with erroneous code would be for later, or as a result of failed code generation, or for the user, or for the public to fix, depending on my future time, effort, and/or desires. I.e., no proper syntax checking at the moment. E.g., I can always run my personal C code through other compilers to confirm correctness for now. In my own languages, I can simply use different keywords for each WHILE position, e.g., DO UNTIL, or use my personal preference, which is to have only one keyword, like LOOP , for an infinite looping construct. Of course, it would be breakable or exitable. > If so you might get away with setting flags BUT because C is > a recursive language you would need to save any such flag or > flags on a stack if you made a call that could end up being > recursive. ... >> Technically, I *could* walk the AST, but that is slows everything >> down. I was attempting to do this without doing so, i.e., single >> pass of AST to assembly, no walking or backtracking of ASTs whether >> as a binary tree or linked-list. > > If you can walk any sort of tree at this point then I have to suggest > it would be better to use a normal parse mechanism. With the exception > of a backtracking parser, parsing is one of the faster phases of a > compiler so you may be ill advised to try to save time with your > contextless approach. ... > It may be a bit like trying to analyse a chess > position based purely on the current board without forming a game > tree: great and revolutionary if you could make it work but probably > impractical, impossible to program, and more work in the long run. Why do you see this as any different from analyzing a chess game from the starting positions? i.e., no move history or game tree, unless artificially constructed for initial piece placement ... Depending on the pieces remaining and their positions, especially proximity, it may actually be much easier, i.e., fewer pieces and positions under attack. Positions under attack can be trimmed for computation, if the opponent is limited to short-move pieces, e.g., no queen, rook, bishop. The opponent may not be able to move into the broad attack positions of your queen, rook, bishop. IIRC, the end-games of chess up to a handful of pieces have all been solved. IIRC, you brought up a chess analogy before too. > It's an interesting idea. You might get it to work. It reminds me, > however, that when I was at school I tried to come up with an > algorithm for marking a game that some people call bulls-and-cows. > Rather than do it in a simple way I tried to do it a clever way > that would be faster than the obvious solution. But it became so > complicated that I eventually realised that it would be far better > to do the marking in the most obvious way. Even if I had got my > clever approach to work it would have been unmaintainable. Bulls and cows? ... (look up) Wikipedia says the game is similar to the commercial game Mastermind. I loved Mastermind as a kid for a year or so. However, I have no recollection as to how the game was played, or how I solved it, just that I seemed to win frequently, against other kids ... I still remember the board and pegs. The Wikipedia page on Mastermind says Donald Knuth determined an algorithm for solving it. Some mathematicians have developed their own algorithms. It seems they've determined the average game length. So, any algorithm for solving either game is measurable. Rod Pemberton -- The idea that sentient beings can be suppressed by a few simple rules is a farce. Even so, Isaac Asimov posited such rules for sentient artificial intelligence.
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-01-16 17:09 +0000 |
| Message-ID | <n7dtc2$sev$1@dont-email.me> |
| In reply to | #9052 |
On 27/12/2015 22:35, Rod Pemberton wrote: > On Sun, 27 Dec 2015 08:51:24 -0500, James Harris > <james.harris.1@gmail.com> wrote: > >> On 01/10/2015 01:32, Rod Pemberton wrote: ... > Around the time you last posted, I figured out two paths > I could take my command-line interpreter project, after > realizing that half my OS code should be used in the > command-line interpreter project and the other half should > really be spun out as a pure, standalone OS. My initial > path for my OS project was for an execution environment > for DJGPP apps and deviated into the OS project. The first > path for the command-line interpreters is towards a Windows > 98/SE like OS, i.e., execution environment for DJGPP apps. > I'm using "OS" loosely here. The other path would be > towards a Linux kernel and/or shell on DOS. While the > latter would be interesting, the DJGPP C compiler won't > support full Linux apps on DOS. There is a bit of work > in four or five areas I'd need to do to progress this. That seems to be traditional and a good design. IIRC the Unix model has a privileged kernel and an unprivileged command-line interpreter. The latter communicates with the former by (making library calls which get translated to) system calls. > I _really_ haven't been motivated to do any programming > since I figured out those paths for my projects ... I don't have the motivation I once had. Perhaps we should do these projects while we are younger! ... I've snipped your entire paragraph talking about a language that has an expletive in its name. If possible, it would be great if you could omit further reference to that language. Your choice, of course, but some people, including me, still find such words offensive and it's not something I would type or discuss, and is a word I try to avoid reading. ... >>> Technically, I *could* walk the AST, but that is slows everything >>> down. I was attempting to do this without doing so, i.e., single >>> pass of AST to assembly, no walking or backtracking of ASTs whether >>> as a binary tree or linked-list. >> >> If you can walk any sort of tree at this point then I have to suggest >> it would be better to use a normal parse mechanism. With the exception >> of a backtracking parser, parsing is one of the faster phases of a >> compiler so you may be ill advised to try to save time with your >> contextless approach. > > .... > >> It may be a bit like trying to analyse a chess >> position based purely on the current board without forming a game >> tree: great and revolutionary if you could make it work but probably >> impractical, impossible to program, and more work in the long run. > > Why do you see this as any different from analyzing a chess game > from the starting positions? i.e., no move history or game tree, > unless artificially constructed for initial piece placement ... I see them as similar. Both (traditionally) need a tree to be constructed. I was saying that parsing without constructing a tree seems to me similar to trying to analyse a chess position without constructing a tree. There may be a way to do both without using a tree but no one has discovered it yet and it's likely to be much easier in both cases just to follow the traditional process; you may be making life harder for yourself in the long run by trying to avoid what you call a grammar-based parse. ... > Bulls and cows? ... (look up) > > Wikipedia says the game is similar to the commercial game Mastermind. > > I loved Mastermind as a kid for a year or so. However, I have > no recollection as to how the game was played, or how I solved it, > just that I seemed to win frequently, against other kids ... I > still remember the board and pegs. I think the Mastermind board game was just a modern copy of Bulls and Cows. Like you I had fun playing it as a kid. In the UK there was a TV programme called Mastermind in which a questionmaster would grill contestants on specialist subjects and general knowledge. I think the Mastermind board game box graphic was based on it. It seems that the Mastermind TV show is still running or has been revived. I wouldn't be surprised if it was run in other countries too. I think I have finally caught up with any outstanding replies on this newsgroup. Sorry for the delay. Let me know if I have missed any you are aware of and I will get to them. James
[toc] | [prev] | [next] | [standalone]
Page 1 of 8 [1] 2 3 4 5 6 7 8 Next page →
Back to top | Article view | alt.os.development
csiph-web