Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compilers > #704 > unrolled thread

type or identifier fundamental parsing issue - Need help from parsing experts

Started byAD <hsad005@gmail.com>
First post2012-07-03 11:30 -0700
Last post2012-07-12 09:23 -0700
Articles 5 — 5 participants

Back to article view | Back to comp.compilers


Contents

  type or identifier fundamental parsing issue - Need help from parsing experts AD <hsad005@gmail.com> - 2012-07-03 11:30 -0700
    Re: type or identifier fundamental parsing issue - Need help from parsing experts George Neuner <gneuner2@comcast.net> - 2012-07-04 00:57 -0400
      Re: type or identifier fundamental parsing issue - Need help from parsing experts torbenm@diku.dk (Torben Ægidius Mogensen) - 2012-07-11 12:17 +0200
    Re: type or identifier fundamental parsing issue - Need help from parsing experts Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-07-04 11:43 +0100
      Re: C arcana, was type or identifier fundamental parsing issue - Need help from parsing experts "christian.bau" <christian.bau@cbau.wanadoo.co.uk> - 2012-07-12 09:23 -0700

#704 — type or identifier fundamental parsing issue - Need help from parsing experts

FromAD <hsad005@gmail.com>
Date2012-07-03 11:30 -0700
Subjecttype or identifier fundamental parsing issue - Need help from parsing experts
Message-ID<12-07-004@comp.compilers>
Greetings All,

I am dealing with a programming langauge that supports something like
sizeof(<typename>) as well as sizeof(<variable-name>) expression.

For parsing such a construct, one would need a parser/yacc rule somewhat like
the following:

SIZEOF_KEYWORD '(' IDENTIFIER ')'

The fundamemtal problem in the rules of the language (that I am dealing with)
is that its lookup/resolution rule *doesn't* allow me to check in symbol table
if that 'IDENTIFIER' is a type or non-type variable. Problem is, this language
supports certain constructs which can potentially/later make such early
lookup/resolution decisions wrong. In short, name resolutions in this langauge
(as per the langauge definition) can only be initiated after the entire source
code has been completely parsed/seen.

Given this restriction, I will probably have to delay/defer the decision of,
whether we saw a 'type' or a non-type variable (with the 'sizeof' operator) to
"semantic check phase".

On the other hand, some people/experts believe that such decisions of whether
something is a type or non-type idernfier has to be frozen/finished during
parsing and *SHOULD NOT* be deferred/delayed to 'semantic check phase'.

I am not an expert compiler researcher/scientist, so am seeking some opinion
here, if you happen to have a sound knowledge on this issue.

Many thanks.

Regards,
AD

[toc] | [next] | [standalone]


#706

FromGeorge Neuner <gneuner2@comcast.net>
Date2012-07-04 00:57 -0400
Message-ID<12-07-006@comp.compilers>
In reply to#704
On Tue, 3 Jul 2012 11:30:48 -0700 (PDT), AD <hsad005@gmail.com> wrote:

>I am dealing with a programming langauge that supports something like
>sizeof(<typename>) as well as sizeof(<variable-name>) expression.
>   :
>Problem is, this language supports certain constructs which can potentially
>later make such early lookup/resolution decisions wrong. In short, name
>resolutions in this langauge (as per the langauge definition) can only be
>initiated after the entire source code has been completely parsed/seen.
>
>Given this restriction, I will probably have to delay/defer the decision of,
>whether we saw a 'type' or a non-type variable (with the 'sizeof' operator) to
>"semantic check phase".
>
>On the other hand, some people/experts believe that such decisions of whether
>something is a type or non-type idernfier has to be frozen/finished during
>parsing and *SHOULD NOT* be deferred/delayed to 'semantic check phase'.

There are no hard and fast rules.  There typically is efficiency to be
gained by classifying identifiers as early as is practical, but the
notion that such classification *must* be done at parse time simply is
ridiculous.  Simply do whatever is most convenient.

That said, ambiguity such as you describe tends to make a language
complex and difficult for programmers to understand.  This is not
necessarily a bad thing, but it should be justified by a measurable
gain in expressive power.  Ambiguity introduced merely for brevity
(e.g., shorthand) or for the sake of clever features which have no
clearly beneficial use cases should be carefully examined to determine
whether the semantics provided are even useful or if they can be
expressed in a less confusing way.

George

[toc] | [prev] | [next] | [standalone]


#712

Fromtorbenm@diku.dk (Torben Ægidius Mogensen)
Date2012-07-11 12:17 +0200
Message-ID<12-07-012@comp.compilers>
In reply to#706
George Neuner <gneuner2@comcast.net> writes:


> There are no hard and fast rules.  There typically is efficiency to be
> gained by classifying identifiers as early as is practical, but the
> notion that such classification *must* be done at parse time simply is
> ridiculous.  Simply do whatever is most convenient.
>
> That said, ambiguity such as you describe tends to make a language
> complex and difficult for programmers to understand.  This is not
> necessarily a bad thing, but it should be justified by a measurable
> gain in expressive power.

I agree.  If making even a partial parse requires classification of
identifiers based on declarations that can be arbitrarily far way,
this is not so much a problem for the compiler (which can in most
cases easily remember all previous declarations and use these to make
this classification) but for the programmer (who can't).

This is a problem in C, where a*b; can be a declaration of b to be a
pointer to a value of type a or an expression statement that
multiplies two values depending on whether a is a type or a variable.
It is even worse in C++, where a<b,c>(d) can be either a call to a
template function or a comma expression consisting of two comparisons.

Even SML, which is otherwise a very clean design that is easy to parse
for humans, the lack of syntactic distinction between variables and
nullary constructors can make it hard to know if a pattern is a binding
instance of a variable or a constructor pattern, so I prefer the Haskell
approach that distinguishes variables and constructors by case: Upper
case indicates constructors and lower case indicate variables.

So my advice is to either make the syntax such that you don't need to
classify identifiers or make the classification local, such as by
upper/lower case, initial letter (like in Fortran), a suffixed $ or %
(like in BASIC) or some other feature of the name.

Similarly, if you allow declaration of infix operators with different
precedences, it can be hard for a reader of a program to parse an
expression without knowing the precedences.  If the precedence
declarations can be arbitrarily far away from the expression, this is a
problem for the programmer (but not the compiler).  An elegant solution
(IMO) is empliyed by O'Caml: Infix operators are built from a limited
set of symbols and the first symbol in an operator name indicates its
precedence: +=-< has the same precedence as +, <-=+ has the same
precedence as < and so on. So all you need to recall is the precendences
of the standard operators.  This can be bad enough if there are dozens
of operators with over a dozen different precedences (like in C++), but
if you keep the number modest, it is no problem.  I think Wirth went to
far in restricting the number of precedence levels in Pascal, but
anything over 8 is probably too many.

An IDE can, of course, help a human to parse a program text, but that
only works on a screen and it may take extra time (for the programmer)
to process the information provided by the IDE, which is often in the
form of mouse-pointer information, colouring or matching brackets.  So,
ideally, the program text should be easy to parse for a human without
computer assistance.

	Torben

[toc] | [prev] | [next] | [standalone]


#707

FromHans-Peter Diettrich <DrDiettrich1@aol.com>
Date2012-07-04 11:43 +0100
Message-ID<12-07-007@comp.compilers>
In reply to#704
AD schrieb:
> Greetings All,
>
> I am dealing with a programming langauge that supports something like
> sizeof(<typename>) as well as sizeof(<variable-name>) expression.
>
> For parsing such a construct, one would need a parser/yacc rule somewhat like
> the following:
>
> SIZEOF_KEYWORD '(' IDENTIFIER ')'
>
> The fundamemtal problem in the rules of the language (that I am dealing with)
> is that its lookup/resolution rule *doesn't* allow me to check in symbol table
> if that 'IDENTIFIER' is a type or non-type variable. Problem is, this language
> supports certain constructs which can potentially/later make such early
> lookup/resolution decisions wrong. In short, name resolutions in this langauge
> (as per the langauge definition) can only be initiated after the entire source
> code has been completely parsed/seen.

There exist pros and cons. The C preprocessor *requires* that sizeof
is a built-in *macro*, so that it can be used in #if conditions.

OTOH the size of structs depends heavily on the target environment,
alignments and other factors, so that I'd delay the evaluation until
code generation, where all involved factors are definitely known.

Anything between these extremes is possible as well. The real question
is: Do there exist reasons/situations, where sizeof *must* be evaluated
prior to the code generation/execution phase?

DoDi

[toc] | [prev] | [next] | [standalone]


#713 — Re: C arcana, was type or identifier fundamental parsing issue - Need help from parsing experts

From"christian.bau" <christian.bau@cbau.wanadoo.co.uk>
Date2012-07-12 09:23 -0700
SubjectRe: C arcana, was type or identifier fundamental parsing issue - Need help from parsing experts
Message-ID<12-07-013@comp.compilers>
In reply to#707
On Jul 4, 11:43 am, Hans-Peter Diettrich <DrDiettri...@aol.com> wrote:

> There exist pros and cons. The C preprocessor *requires* that sizeof
> is a built-in *macro*, so that it can be used in #if conditions.

Ahem... No, it doesn't. Unless you do something really perverse like

#define sizeof <whatever>

sizeof will not be defined as a macro, and any non-macro identifier
used within #if other than as an operand to "defined" will be replaced
by 0.
[After squinting at the various C standards and checking with
committee members, I found that the C preprocessor does not know
about any keywords at all, so it treats sizeof and int as ordinary
names.  This allows occasionaly useful kludges like #define short int
-John]

[toc] | [prev] | [standalone]


Back to top | Article view | comp.compilers


csiph-web