Path: csiph.com!v102.xanadu-bbs.net!xanadu-bbs.net!news.glorb.com!news-out.readnews.com!news-xxxfer.readnews.com!news.misty.com!news.iecc.com!.POSTED!nerds-end From: bnrj.rudra@gmail.com Newsgroups: comp.compilers Subject: parsing bibtex file with flex/bison Date: Mon, 4 Mar 2013 15:48:56 -0800 (PST) Organization: Compilers Central Lines: 92 Sender: johnl@iecc.com Approved: comp.compilers@iecc.com Message-ID: <13-03-003@comp.compilers> NNTP-Posting-Host: news.iecc.com Mime-Version: 1.0 Content-Type: text/plain; charset=ISO-8859-1 X-Trace: leila.iecc.com 1362460433 22355 64.57.183.58 (5 Mar 2013 05:13:53 GMT) X-Complaints-To: abuse@iecc.com NNTP-Posting-Date: Tue, 5 Mar 2013 05:13:53 +0000 (UTC) Injection-Date: Mon, 04 Mar 2013 23:48:56 +0000 Keywords: lex, parse, design, comment Posted-Date: 05 Mar 2013 00:13:53 EST X-submission-address: compilers@iecc.com X-moderator-address: compilers-request@iecc.com X-FAQ-and-archives: http://compilers.iecc.com Xref: csiph.com comp.compilers:864 I want to parse bibtex file using flex/bison. A sample bibtex is: @Book{a1, author="amook", Title="ASR", Publisher="oxf", Year="2010", Add="UK", Edition="1", } @Article{a2, Author="Rudra Banerjee", Title={FeNiMo}, Publisher={P{\"R}B}, Issue="12", Page="36690", Year="2011", Add="UK", Edition="1", } (A new key may start in same line) Now, I have written a flex code: %{ #include #include %} %{ char yylval; int YEAR,i; //char array_author[1000]; %} %x author %x title %x pub %x year %% @ printf("\nNEWENTRY\n"); [a-zA-Z][a-zA-Z0-9]* {printf("%s",yytext); BEGIN(INITIAL);} author= {BEGIN(author);} \"[a-zA-Z\/.]+\" {printf("%s",yytext); BEGIN(INITIAL);} title= {BEGIN(title);} \"[a-zA-Z\/.]+\" {printf("%s",yytext); BEGIN(INITIAL);} publisher= {BEGIN(pub);} <pub>\"[a-zA-Z\/.]+\" {printf("%s",yytext); BEGIN(INITIAL);} [a-zA-Z0-9\/.-]+= printf("ENTRY TYPE "); \" printf("QUOTE "); \{ printf("LCB "); \} printf(" RCB"); ; printf("SEMICOLON "); \n printf("\n"); %% int main(){ yylex(); //char array_author[1000]; //printf("%d%s",&i,array_author[i]); i++; return 0; } while this is peeking up the few things, not all. Can anyone kindly help me with this? [My suggestion would be to do less in the lexer and more in the parser. In the lexer, responable tokens might be '@' '{' '}' '=' ',' word qstring (quoted string) Then you could write bison rules like this: clause: '@' word '{' word ',' attrlist '}' ; attrlist: attr | attr ',' attrlist ; attr: name '=' value ; value: word | qstring | nestlist : nestlist: '{' list '}' ; list: listitem | list listitem ; listitem: word | qstring | nestlist : And so forth. This isn't exactly right, but it should get you going in the right direction. The parser will recognize some invalid bibtex, e.g., words that aren't attribute names, which it's easier to check in semantic code rather than trying to stick laundry lists of keywords into the parser. -John]