Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.compilers > #1595 > unrolled thread
| Started by | Jens Kallup <jkallup@web.de> |
|---|---|
| First post | 2015-09-01 04:20 +0200 |
| Last post | 2015-09-07 17:11 +0200 |
| Articles | 8 — 4 participants |
Back to article view | Back to comp.compilers
lex - how to get current file position of parser file Jens Kallup <jkallup@web.de> - 2015-09-01 04:20 +0200
Re: lex - how to get current file position of parser file Kaz Kylheku <kaz@kylheku.com> - 2015-09-01 14:19 +0000
Re: lex - how to get current file position of parser file federation2005@netzero.com - 2015-09-03 15:47 -0700
Re: include files, was lex - how to get current file position of parser file Kaz Kylheku <kaz@kylheku.com> - 2015-09-04 05:26 +0000
Re: lex - how to get current file position of parser file Gene Wirchenko <genew@telus.net> - 2015-09-03 18:19 -0700
Re: lex - how to get current file position of parser file Jens Kallup <jkallup@web.de> - 2015-09-06 21:53 +0200
Re: lex - how to get current file position of parser file Kaz Kylheku <kaz@kylheku.com> - 2015-09-07 05:59 +0000
Re: lex - how to get current file position of parser file Jens Kallup <jkallup@web.de> - 2015-09-07 17:11 +0200
| From | Jens Kallup <jkallup@web.de> |
|---|---|
| Date | 2015-09-01 04:20 +0200 |
| Subject | lex - how to get current file position of parser file |
| Message-ID | <15-08-019@comp.compilers> |
Hello, I used flex, and bison under Linux. But how can I handle multiple input files? When I follow the examples of the documentation, I fail to handle tokens after include file. Example: PRINT "Before" SET PROCEDURE TO testproc PRINT "After" this is a abstract language, which does (should) open a file called "testproc". But I can not handle line 3. I try to "ftell" the position after testproc, but get the size of the parser file. So, now I have no glue, how to fix that. It would be great, if someone have a idea, and maybe a short code snippet. Thanks Jens [Flex buffers its input, so you can't just do an ftell(). Use the yy_create_buffer() and related routines. If you have a copy of my book "flex & bison" there's an example of nested include files in chapter 2. -John]
[toc] | [next] | [standalone]
| From | Kaz Kylheku <kaz@kylheku.com> |
|---|---|
| Date | 2015-09-01 14:19 +0000 |
| Message-ID | <15-09-001@comp.compilers> |
| In reply to | #1595 |
On 2015-09-01, Jens Kallup <jkallup@web.de> wrote: > Hello, > > I used flex, and bison under Linux. But how can I handle multiple > input files? When I follow the examples of the documentation, I fail > to handle tokens after include file. > > Example: > > PRINT "Before" > SET PROCEDURE TO testproc > PRINT "After" > this is a abstract language, which does (should) open a file called > "testproc". I think that in a dBase-type language, SET PROCEDURE doesn't necessarily have to open the file at parse time. According to current documentation that I'm able to find, it's actually become a complex module-loading command, and not parse-time textual inclusion. It can load compiled files, not only text, and cause non-persistent files to be unloaded. Loaded files have a refcount so that if you invoke SET PROCEDURE on some file N times, you have to CLOSE PROCEDURE that many times before it is unloaded. > But I can not handle line 3. > I try to "ftell" the position after testproc, > but get the size of the parser file. > > So, now I have no glue, how to fix that. It would be great, if > someone have a idea, and maybe a short code snippet. > > Thanks > Jens > [Flex buffers its input, so you can't just do an ftell(). Use the > yy_create_buffer() and related routines. If you have a copy of my > book "flex & bison" there's an example of nested include files in > chapter 2. -John] The benefit of yy_create_buffer et al is that you maintain the illusion that there is just a single stream of tokens, though multiple files are being shuffled around by the lexer, which maintains the "include stack". This multiple stream of tokens is handled by a single parse job (one big yyparse). Another approach (if you can design the language that way) is to specify reentrant parser and lexer. This lets you issue a new, recursive call to yyparse to handle the loaded/included file. The call has its own instance of the lexer, which has no interaction with the suspended lexer of the parent. If your parser builds an abstract syntax tree, then the phrase structure rule for the include construct pulls out the AST from the nested yyparse job, and integrates it into the surrounding AST. Of course, this can't be used if you're implementing "dumb" textual inclusion, such that the syntax matched by a single phrase structure rule can span across multiple files, and even start in one file and end in another.
[toc] | [prev] | [next] | [standalone]
| From | federation2005@netzero.com |
|---|---|
| Date | 2015-09-03 15:47 -0700 |
| Message-ID | <15-09-002@comp.compilers> |
| In reply to | #1596 |
On Thursday, September 3, 2015 at 10:02:20 AM UTC-5, Kaz Kylheku wrote: > Another approach (if you can design the language that way) is to specify > reentrant parser and lexer. If you can get a clean unambiguous cut this should work. But it bears to point out that the reason you don't see reentrancy very much is that the LR framework is not really compatible with it. Lookaheads have to all be lined up, for a clean descent to a sub-level. This issue is what underlies the frequent-occurrence of the "advanced entry point" hack (e.g. starting up a sub-parser for "expressions" in a typical Algol language), where the sub-level starts up one or more tokens past the starting point of the thing being parsed. Those are the very things factored-out up front in an LR table. So you end up having to compromise on the LR-ness of LR itself, with the hack or similar means, if you want separate parser levels.
[toc] | [prev] | [next] | [standalone]
| From | Kaz Kylheku <kaz@kylheku.com> |
|---|---|
| Date | 2015-09-04 05:26 +0000 |
| Subject | Re: include files, was lex - how to get current file position of parser file |
| Message-ID | <15-09-004@comp.compilers> |
| In reply to | #1597 |
On 2015-09-03, federation2005@netzero.com <federation2005@netzero.com> wrote:
> On Thursday, September 3, 2015 at 10:02:20 AM UTC-5, Kaz Kylheku wrote:
>> Another approach (if you can design the language that way) is to specify
>> reentrant parser and lexer.
>
> If you can get a clean unambiguous cut this should work. But it bears
> to point out that the reason you don't see reentrancy very much is
> that the LR framework is not really compatible with it. Lookaheads
> have to all be lined up, for a clean descent to a sub-level.
While that is true, it is not an impediment at all. Reentrant parsers
involve the use of completely separate parser/lexer instances operating
on separate streams. Each maintain their lookahead.
For instance, suppose that parser P0 (top level) reduces an include
construct:
include : INCLUDE file
{
/* Pseudo code */
parser_t P1; /* our custom parser type */
parser_init(&P1, $2); /* opens file */
yyparse(&P1);
$$ = P1.abstract_syntax_tree;
}
;
Even if P0 has read the next token past the "INCLUDE file" syntax
(and based on that token, in fact, it has decided to reduce
this rule), that doesn't affect anything going on in parser P1.
It has its own file open, own lexer with its own token stream, its
own Yacc stack, its own lookahead token.
When the nested yyparse is done, we have the parse; we can integrate
it into the outer parse, and keep going.
> This issue is what underlies the frequent-occurrence of the "advanced
> entry point" hack (e.g. starting up a sub-parser for "expressions" in
> a typical Algol language), where the sub-level starts up one or more
> tokens past the starting point of the thing being parsed.
Parsing a subset of the grammar like just an "expression" is a different issue,
though related because it behooves you to use reentrant parsers. However, the
use of reentrant parsers doesn't necessarily you are doing such a thing.
I have experience in this area. Look for SECRET_ESCAPE_E in this
Yacc file:
http://www.kylheku.com/cgit/txr/tree/parser.y
The E stands for expression. SECRET_ESCAPE_E is a token that is never
lexically analyzed. It is injected into the token stream via a "token unget"
type operation (like ungetc(stream) but for a token, not a character).
To see how this is used, look in this file:
http://www.kylheku.com/cgit/txr/tree/parser.c
for a function called "prime_parser". When we want to call yyparse to
read anothe expression from the stream, we must not only prime the parser
with the SECRET_ESCAPE_E token, but we must also inject the lookahead token
from the previously parsed expression!
For the sake of this, I support multiple tokens of pushback in the token
stream (up to four, but I only ever use two).
If there is a "yychar" from a recent parse (yychar is the Yacc name,
as you know, of the lookahead token), then we push that first.
Then we push the secret escape token.
Blam: call yyparse and it is fooled. The secret token guides it into the
correct area of the grammar to parse what we want and the pushed yychar
restores its continuation context. Like reloading the registers of a thread
and dispatching so it continues where it was.
I suppose we could call this ... Yacc/cc: Yacc with current continuation.
OMG kill me now. :)
[toc] | [prev] | [next] | [standalone]
| From | Gene Wirchenko <genew@telus.net> |
|---|---|
| Date | 2015-09-03 18:19 -0700 |
| Message-ID | <15-09-003@comp.compilers> |
| In reply to | #1596 |
On Tue, 1 Sep 2015 14:19:33 +0000 (UTC), Kaz Kylheku <kaz@kylheku.com>
wrote:
>On 2015-09-01, Jens Kallup <jkallup@web.de> wrote:
>> Hello,
>>
>> I used flex, and bison under Linux. But how can I handle multiple
>> input files? When I follow the examples of the documentation, I fail
>> to handle tokens after include file.
>>
>> Example:
>>
>> PRINT "Before"
>> SET PROCEDURE TO testproc
>> PRINT "After"
>
>> this is a abstract language, which does (should) open a file called
>> "testproc".
>
>I think that in a dBase-type language, SET PROCEDURE doesn't necessarily have
>to open the file at parse time. According to current documentation that
It does not have to, but if compiling a project, it will be
compiled as well.
>I'm able to find, it's actually become a complex module-loading command, and
>not parse-time textual inclusion. It can load compiled files, not only text,
Yes.
>and cause non-persistent files to be unloaded. Loaded files have a refcount so
>that if you invoke SET PROCEDURE on some file N times, you have to CLOSE
>PROCEDURE that many times before it is unloaded.
No. At least, this is not the case in Microsoft Visual FoxPro
9.0. You can not add a procedure file more than once. If you do, it
is just ignored. I tried using just the filename and then with the
path. The two were recognised as being the same. And, yes, I was
using the ADDITIVE keyword. CLOSE PROCEDURE releases *all* of the
procedure files.
SET PROCEDURE's exact meaning is determined by file contents at
run-time.
[snip]
Sincerely,
Gene Wirchenko
[toc] | [prev] | [next] | [standalone]
| From | Jens Kallup <jkallup@web.de> |
|---|---|
| Date | 2015-09-06 21:53 +0200 |
| Message-ID | <15-09-007@comp.compilers> |
| In reply to | #1598 |
Hello,
I start a testcase Project - see attachment(s).
Most of the source comes from John's book.
But I collect his code, and using the Qt5 framework,
to make an executable.
It can be created by "qmake", which makes a "Makefile".
And with "make", it can be compiled into exec file.
It will be run fine, except "set procedure to test.prg".
The applications display the content of initial code, as
given by "argv[1]" at start up.
But, if I try to include "test.prg", the application
display me a blank line, and jumps to next line in initial
file.
Have someone a glue?
Thanks
Jens
// ---%<--- main.cpp:
#include <stdio.h>
#include <string.h>
#include <stdlib.h>
extern int newfile(char*);
extern int yylex(void);
int main(int argc, char *argv[])
{
if (argc < 2) {
fprintf(stderr, "need filename\n");
return 1;
}
if (newfile(argv[1]))
return yylex();
}
// --->%--- EOF - main.cpp
// ---%<--- flex.l:
#include <string.h>
#include <stdlib.h>
#include <iostream>
#include <algorithm>
#include <vector>
#include "y.tab.h"
#include <QString>
#include <QVector>
QVector<QString> paths_vector;
struct bufstack {
struct bufstack *prev; /* previous entry */
YY_BUFFER_STATE bs; /* saved buffer */
int lineno; /* saved line number */
int parse_mode; /* compiler pass */
char *filename; /* name of this file */
FILE *f; /* current file */
} *curbs = 0;
char *curfilename; /* name of current input file */
int newfile(char *fn)
{
FILE *f = fopen(fn, "r");
struct bufstack *bs = (bufstack*) malloc(sizeof(struct bufstack));
/* die if no file or no room */
if (!f) { perror(fn); return 0; }
if (!bs) { perror("mallocas"); return 0; }
/* remember state */
if (curbs) curbs->lineno = yylineno;
bs->prev = curbs;
/* set up current entry */
bs->bs = yy_create_buffer(f, YY_BUF_SIZE);
bs->f = f;
bs->filename = fn;
yy_switch_to_buffer(bs->bs);
curbs = bs;
yylineno = 1;
curfilename = fn;
return 1;
}
int popfile(void)
{
struct bufstack *bs = curbs;
struct bufstack *prevbs;
if(!bs) return 0;
/* get rid of current entry */
fclose(bs->f);
yy_delete_buffer(bs->bs);
/* switch back to previous */
prevbs = bs->prev;
free(bs);
if (!prevbs) return 0;
yy_switch_to_buffer(prevbs->bs);
curbs = prevbs;
yylineno = curbs->lineno;
curfilename = curbs->filename;
return 1;
}
void yyerror(char* message)
{
printf("Error: '%s' in line: %d",message,yylineno);
exit(1);
}
extern "C" int yywrap(void) { return 1; }
%}
%%
"set"[ \t\n]*"path"[ \t\n]*"to"[ \t\n]*\"([^\"]*)\" {
QString str = yytext;
str = str.replace("set","");
str = str.replace("path","");
str = str.replace("to","");
str = str.replace("\"","");
str = str.replace("\n","");
str = str.replace("\t","");
str = str.replace(" ","");
paths_vector << QString(str.toStdString().c_str());
BEGIN(INITIAL);
}
"set"[ \t\n]*"procedure"[ \t\n]*"to"[ \t\n]*([0-9a-zA-Z_]+[0-9a-zA-Z_\.]*) {
QString str = yytext;
str = str.replace("procedure","");
str = str.replace("set","");
str = str.replace("to","");
str = str.replace("\n","");
str = str.replace("\t","");
str = str.replace(" ","");
QString s;
if (paths_vector.size() < 0) {
printf("hhh\n");
s = QString("./" + str);
if (!newfile((char*)s.toStdString().c_str())) {
printf("file not found!\n");
yyterminate();
}
}
else {
for (int i = 0; i < paths_vector.size(); i++)
{
s = paths_vector.at(i);
s = s + QString("/") + str;
if ((newfile((char*)s.toStdString().c_str()))) {
break;
}
}
}
BEGIN(INITIAL);
}
<<EOF>> { if(!popfile()) yyterminate(); }
^. { fprintf(yyout, "%4d %s", yylineno, yytext); }
^\n { fprintf(yyout, "%4d %s", yylineno++, yytext); }
\n { ECHO; yylineno++; }
. { ECHO; }
%%
// --->%--- EOF - flex.l
// ---%<--- yacc.y
%{
extern int yylex();
extern void yyerror(char*);
%}
%start program
%%
program : stmt_seq;
stmt_seq
: stmt_seq stmt { }
| stmt
;
stmt
: { /* empty */ }
;
%%
// --->%--- EOF - yacc.y
// ---%<--- qmake.pro
QT += core
QT -= gui
TARGET = predbase
CONFIG += console
CONFIG -= app_bundle
TEMPLATE = app
SOURCES += \
pre_main.cpp
OTHER_FILES += \
pre_pcode.l\
pre_pcode.y
FLEXSOURCES = pre_pcode.l
BISONSOURCES = pre_pcode.y
flex.commands = flex -i ${QMAKE_FILE_IN} && mv lex.yy.c lex.yy.cpp
flex.input = FLEXSOURCES
flex.output = lex.yy.cpp
flex.variable_out = SOURCES
flex.depends = y.tab.h
flex.name = flex
QMAKE_EXTRA_COMPILERS += flex
bison.commands = bison -d -t -y ${QMAKE_FILE_IN} && mv y.tab.c y.tab.cpp
bison.input = BISONSOURCES
bison.output = y.tab.cpp
bison.variable_out = SOURCES
bison.name = bison
QMAKE_EXTRA_COMPILERS += bison
bisonheader.commands = @true
bisonheader.input = BISONSOURCES
bisonheader.output = y.tab.h
bisonheader.variable_out = HEADERS
bisonheader.name = bison header
bisonheader.depends = y.tab.cpp
QMAKE_EXTRA_COMPILERS += bisonheader
QMAKE_EXTRA_TARGETS += flex bison
// --->%--- EOF - qmake.pro
[I haven't a clue. What ever happened to using debuggers
and breakpoints to debug your code? -John]
[toc] | [prev] | [next] | [standalone]
| From | Kaz Kylheku <kaz@kylheku.com> |
|---|---|
| Date | 2015-09-07 05:59 +0000 |
| Message-ID | <15-09-008@comp.compilers> |
| In reply to | #1602 |
On 2015-09-06, Jens Kallup <jkallup@web.de> wrote: > [I haven't a clue. What ever happened to using debuggers > and breakpoints to debug your code? -John] Lex and Yacc happened. :) [Well, yes, but I've put plenty of breakpoints in my lex and yacc semantic code. YYDEBUG helps a lot, too. -John]
[toc] | [prev] | [next] | [standalone]
| From | Jens Kallup <jkallup@web.de> |
|---|---|
| Date | 2015-09-07 17:11 +0200 |
| Message-ID | <15-09-009@comp.compilers> |
| In reply to | #1596 |
Hello @all, I have a codeing mistake. here comes the glue: // ---%<--- // a line comment /* block comment */ PRINT "hello" set path to "/home/bak/src/ui/dbase/predbase" set procedure to source.txt else OR 33 // --->%--- see "set path" I have to fix that, so, that the compiler looks into current folder, first. ok, thx John
[toc] | [prev] | [standalone]
Back to top | Article view | comp.compilers
csiph-web