Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.compilers > #704 > unrolled thread
| Started by | AD <hsad005@gmail.com> |
|---|---|
| First post | 2012-07-03 11:30 -0700 |
| Last post | 2012-07-12 09:23 -0700 |
| Articles | 5 — 5 participants |
Back to article view | Back to comp.compilers
type or identifier fundamental parsing issue - Need help from parsing experts AD <hsad005@gmail.com> - 2012-07-03 11:30 -0700
Re: type or identifier fundamental parsing issue - Need help from parsing experts George Neuner <gneuner2@comcast.net> - 2012-07-04 00:57 -0400
Re: type or identifier fundamental parsing issue - Need help from parsing experts torbenm@diku.dk (Torben Ægidius Mogensen) - 2012-07-11 12:17 +0200
Re: type or identifier fundamental parsing issue - Need help from parsing experts Hans-Peter Diettrich <DrDiettrich1@aol.com> - 2012-07-04 11:43 +0100
Re: C arcana, was type or identifier fundamental parsing issue - Need help from parsing experts "christian.bau" <christian.bau@cbau.wanadoo.co.uk> - 2012-07-12 09:23 -0700
| From | AD <hsad005@gmail.com> |
|---|---|
| Date | 2012-07-03 11:30 -0700 |
| Subject | type or identifier fundamental parsing issue - Need help from parsing experts |
| Message-ID | <12-07-004@comp.compilers> |
Greetings All,
I am dealing with a programming langauge that supports something like
sizeof(<typename>) as well as sizeof(<variable-name>) expression.
For parsing such a construct, one would need a parser/yacc rule somewhat like
the following:
SIZEOF_KEYWORD '(' IDENTIFIER ')'
The fundamemtal problem in the rules of the language (that I am dealing with)
is that its lookup/resolution rule *doesn't* allow me to check in symbol table
if that 'IDENTIFIER' is a type or non-type variable. Problem is, this language
supports certain constructs which can potentially/later make such early
lookup/resolution decisions wrong. In short, name resolutions in this langauge
(as per the langauge definition) can only be initiated after the entire source
code has been completely parsed/seen.
Given this restriction, I will probably have to delay/defer the decision of,
whether we saw a 'type' or a non-type variable (with the 'sizeof' operator) to
"semantic check phase".
On the other hand, some people/experts believe that such decisions of whether
something is a type or non-type idernfier has to be frozen/finished during
parsing and *SHOULD NOT* be deferred/delayed to 'semantic check phase'.
I am not an expert compiler researcher/scientist, so am seeking some opinion
here, if you happen to have a sound knowledge on this issue.
Many thanks.
Regards,
AD
[toc] | [next] | [standalone]
| From | George Neuner <gneuner2@comcast.net> |
|---|---|
| Date | 2012-07-04 00:57 -0400 |
| Message-ID | <12-07-006@comp.compilers> |
| In reply to | #704 |
On Tue, 3 Jul 2012 11:30:48 -0700 (PDT), AD <hsad005@gmail.com> wrote: >I am dealing with a programming langauge that supports something like >sizeof(<typename>) as well as sizeof(<variable-name>) expression. > : >Problem is, this language supports certain constructs which can potentially >later make such early lookup/resolution decisions wrong. In short, name >resolutions in this langauge (as per the langauge definition) can only be >initiated after the entire source code has been completely parsed/seen. > >Given this restriction, I will probably have to delay/defer the decision of, >whether we saw a 'type' or a non-type variable (with the 'sizeof' operator) to >"semantic check phase". > >On the other hand, some people/experts believe that such decisions of whether >something is a type or non-type idernfier has to be frozen/finished during >parsing and *SHOULD NOT* be deferred/delayed to 'semantic check phase'. There are no hard and fast rules. There typically is efficiency to be gained by classifying identifiers as early as is practical, but the notion that such classification *must* be done at parse time simply is ridiculous. Simply do whatever is most convenient. That said, ambiguity such as you describe tends to make a language complex and difficult for programmers to understand. This is not necessarily a bad thing, but it should be justified by a measurable gain in expressive power. Ambiguity introduced merely for brevity (e.g., shorthand) or for the sake of clever features which have no clearly beneficial use cases should be carefully examined to determine whether the semantics provided are even useful or if they can be expressed in a less confusing way. George
[toc] | [prev] | [next] | [standalone]
| From | torbenm@diku.dk (Torben Ægidius Mogensen) |
|---|---|
| Date | 2012-07-11 12:17 +0200 |
| Message-ID | <12-07-012@comp.compilers> |
| In reply to | #706 |
George Neuner <gneuner2@comcast.net> writes: > There are no hard and fast rules. There typically is efficiency to be > gained by classifying identifiers as early as is practical, but the > notion that such classification *must* be done at parse time simply is > ridiculous. Simply do whatever is most convenient. > > That said, ambiguity such as you describe tends to make a language > complex and difficult for programmers to understand. This is not > necessarily a bad thing, but it should be justified by a measurable > gain in expressive power. I agree. If making even a partial parse requires classification of identifiers based on declarations that can be arbitrarily far way, this is not so much a problem for the compiler (which can in most cases easily remember all previous declarations and use these to make this classification) but for the programmer (who can't). This is a problem in C, where a*b; can be a declaration of b to be a pointer to a value of type a or an expression statement that multiplies two values depending on whether a is a type or a variable. It is even worse in C++, where a<b,c>(d) can be either a call to a template function or a comma expression consisting of two comparisons. Even SML, which is otherwise a very clean design that is easy to parse for humans, the lack of syntactic distinction between variables and nullary constructors can make it hard to know if a pattern is a binding instance of a variable or a constructor pattern, so I prefer the Haskell approach that distinguishes variables and constructors by case: Upper case indicates constructors and lower case indicate variables. So my advice is to either make the syntax such that you don't need to classify identifiers or make the classification local, such as by upper/lower case, initial letter (like in Fortran), a suffixed $ or % (like in BASIC) or some other feature of the name. Similarly, if you allow declaration of infix operators with different precedences, it can be hard for a reader of a program to parse an expression without knowing the precedences. If the precedence declarations can be arbitrarily far away from the expression, this is a problem for the programmer (but not the compiler). An elegant solution (IMO) is empliyed by O'Caml: Infix operators are built from a limited set of symbols and the first symbol in an operator name indicates its precedence: +=-< has the same precedence as +, <-=+ has the same precedence as < and so on. So all you need to recall is the precendences of the standard operators. This can be bad enough if there are dozens of operators with over a dozen different precedences (like in C++), but if you keep the number modest, it is no problem. I think Wirth went to far in restricting the number of precedence levels in Pascal, but anything over 8 is probably too many. An IDE can, of course, help a human to parse a program text, but that only works on a screen and it may take extra time (for the programmer) to process the information provided by the IDE, which is often in the form of mouse-pointer information, colouring or matching brackets. So, ideally, the program text should be easy to parse for a human without computer assistance. Torben
[toc] | [prev] | [next] | [standalone]
| From | Hans-Peter Diettrich <DrDiettrich1@aol.com> |
|---|---|
| Date | 2012-07-04 11:43 +0100 |
| Message-ID | <12-07-007@comp.compilers> |
| In reply to | #704 |
AD schrieb:
> Greetings All,
>
> I am dealing with a programming langauge that supports something like
> sizeof(<typename>) as well as sizeof(<variable-name>) expression.
>
> For parsing such a construct, one would need a parser/yacc rule somewhat like
> the following:
>
> SIZEOF_KEYWORD '(' IDENTIFIER ')'
>
> The fundamemtal problem in the rules of the language (that I am dealing with)
> is that its lookup/resolution rule *doesn't* allow me to check in symbol table
> if that 'IDENTIFIER' is a type or non-type variable. Problem is, this language
> supports certain constructs which can potentially/later make such early
> lookup/resolution decisions wrong. In short, name resolutions in this langauge
> (as per the langauge definition) can only be initiated after the entire source
> code has been completely parsed/seen.
There exist pros and cons. The C preprocessor *requires* that sizeof
is a built-in *macro*, so that it can be used in #if conditions.
OTOH the size of structs depends heavily on the target environment,
alignments and other factors, so that I'd delay the evaluation until
code generation, where all involved factors are definitely known.
Anything between these extremes is possible as well. The real question
is: Do there exist reasons/situations, where sizeof *must* be evaluated
prior to the code generation/execution phase?
DoDi
[toc] | [prev] | [next] | [standalone]
| From | "christian.bau" <christian.bau@cbau.wanadoo.co.uk> |
|---|---|
| Date | 2012-07-12 09:23 -0700 |
| Subject | Re: C arcana, was type or identifier fundamental parsing issue - Need help from parsing experts |
| Message-ID | <12-07-013@comp.compilers> |
| In reply to | #707 |
On Jul 4, 11:43 am, Hans-Peter Diettrich <DrDiettri...@aol.com> wrote: > There exist pros and cons. The C preprocessor *requires* that sizeof > is a built-in *macro*, so that it can be used in #if conditions. Ahem... No, it doesn't. Unless you do something really perverse like #define sizeof <whatever> sizeof will not be defined as a macro, and any non-macro identifier used within #if other than as an operand to "defined" will be replaced by 0. [After squinting at the various C standards and checking with committee members, I found that the C preprocessor does not know about any keywords at all, so it treats sizeof and int as ordinary names. This allows occasionaly useful kludges like #define short int -John]
[toc] | [prev] | [standalone]
Back to top | Article view | comp.compilers
csiph-web