Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compilers > #864 > unrolled thread

parsing bibtex file with flex/bison

Started bybnrj.rudra@gmail.com
First post2013-03-04 15:48 -0800
Last post2013-03-08 03:59 +0000
Articles 6 — 5 participants

Back to article view | Back to comp.compilers


Contents

  parsing bibtex file with flex/bison bnrj.rudra@gmail.com - 2013-03-04 15:48 -0800
    Re: parsing bibtex file with flex/bison Evangelos Drikos <drikosev@otenet.gr> - 2013-03-06 13:45 +0200
      Re: parsing bibtex file with flex/bison Rudra Banerjee <bnrj.rudra@gmail.com> - 2013-03-07 11:20 +0000
        Re: parsing bibtex file with flex/bison Rudra Banerjee <bnrj.rudra@gmail.com> - 2013-03-17 00:18 +0000
          Re: parsing bibtex file with flex/bison Torsten Eichstädt <torsten.eichstaedt@FernUni-Hagen.de> - 2013-03-25 13:41 +0100
    Re: parsing bibtex file with flex/bison glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2013-03-08 03:59 +0000

#864 — parsing bibtex file with flex/bison

Frombnrj.rudra@gmail.com
Date2013-03-04 15:48 -0800
Subjectparsing bibtex file with flex/bison
Message-ID<13-03-003@comp.compilers>
I want to parse bibtex file using flex/bison. A sample bibtex is:
@Book{a1,
author="amook",
Title="ASR",
Publisher="oxf",
Year="2010",
Add="UK",
Edition="1",
}
@Article{a2,
Author="Rudra Banerjee",
Title={FeNiMo},
Publisher={P{\"R}B},
Issue="12",
Page="36690",
Year="2011",
Add="UK",
Edition="1",
}
(A new key may start in same line)
Now, I have written a flex code:

%{
#include <stdio.h>
#include <stdlib.h>
%}

%{
char yylval;
int YEAR,i;
//char array_author[1000];
%}
%x author
%x title
%x pub
%x year
%%
@          			printf("\nNEWENTRY\n");
[a-zA-Z][a-zA-Z0-9]*   		{printf("%s",yytext);
					BEGIN(INITIAL);}
author=                 	{BEGIN(author);}
<author>\"[a-zA-Z\/.]+\"  	{printf("%s",yytext);
                           		BEGIN(INITIAL);}
title=                 		{BEGIN(title);}
<title>\"[a-zA-Z\/.]+\"  	{printf("%s",yytext);
                           		BEGIN(INITIAL);}
publisher=                 	{BEGIN(pub);}
<pub>\"[a-zA-Z\/.]+\"  		{printf("%s",yytext);
                           		BEGIN(INITIAL);}
[a-zA-Z0-9\/.-]+=        printf("ENTRY TYPE ");
\"                      printf("QUOTE ");
\{                      printf("LCB ");
\}                      printf(" RCB");
;                       printf("SEMICOLON ");
\n                      printf("\n");
%%

int main(){
  yylex();
//char array_author[1000];
//printf("%d%s",&i,array_author[i]);
i++;
return 0;
}

while this is peeking up the few things, not all.
Can anyone kindly help me with this?
[My suggestion would be to do less in the lexer and more in the parser.  In the
lexer, responable tokens might be '@' '{' '}' '=' ',' word
qstring (quoted string)

Then you could write bison rules like this:

clause: '@' word '{' word ',' attrlist '}' ;

attrlist: attr | attr ',' attrlist ;

attr: name '=' value ;

value: word | qstring | nestlist :

nestlist: '{' list '}' ;

list: listitem | list listitem ;

listitem: word | qstring | nestlist :

And so forth.  This isn't exactly right, but it should get you going
in the right direction. The parser will recognize some invalid bibtex,
e.g., words that aren't attribute names, which it's easier to check in
semantic code rather than trying to stick laundry lists of keywords
into the parser. -John]

[toc] | [next] | [standalone]


#865

FromEvangelos Drikos <drikosev@otenet.gr>
Date2013-03-06 13:45 +0200
Message-ID<13-03-004@comp.compilers>
In reply to#864
On 3/5/13 1:48 AM, bnrj.rudra@gmail.com wrote:
> I want to parse bibtex file using flex/bison...

A summary of BibTex can be found here:

http://maverick.inria.fr/~Xavier.Decoret/resources/xdkbibtex/bibtex_summary.html#file_format

According to the documentation found in the link above
-you can have arbitrarily nested pairs of braces.
-but braces must also be balanced inside quotes

Provided the documentation above is accurate,one cannot describe with a
regular grammar the quoted strings used in BibTex.

Ev. Drikos
[Quite right.  The lexer would recognize them in chunks and the
parser, which uses a pushdown automaton, puts them together.  Or the
other usual approach is a kludge with code in the lexer to count the
braces and set start states that control what's returned. -John]

[toc] | [prev] | [next] | [standalone]


#867

FromRudra Banerjee <bnrj.rudra@gmail.com>
Date2013-03-07 11:20 +0000
Message-ID<13-03-006@comp.compilers>
In reply to#865
> [Quite right.  The lexer would recognize them in chunks and the
> parser, which uses a pushdown automaton, puts them together.  Or the
> other usual approach is a kludge with code in the lexer to count the
> braces and set start states that control what's returned. -John]
>
Hi John and Evangelos,
Thanks for your comments.
I am new in this parsing business, and may not ever need it once this
project is over(personal project, for fun).
According to your comments, can you kindly help me in one thing:

 I *NEED* to parse the bibtex file. So, is flex+bison is the best option
I have or am I already in a wrong way?

If I am in right path, I don't mind learning, but if there is any
alternative, please suggest(oh...its not btparse I am looking for)
[I see perl BibTeX::Parser, python https://launchpad.net/pybtex/,
and java https://code.google.com/p/javabib/.  Unless you really really
need your parser to be written in C, I'd start with one of those. -John]

[toc] | [prev] | [next] | [standalone]


#874

FromRudra Banerjee <bnrj.rudra@gmail.com>
Date2013-03-17 00:18 +0000
Message-ID<13-03-013@comp.compilers>
In reply to#867
I have managed considerable progress in parsing the bib file, but the
next step is quite tough for my present level of understanding.
I have created bison and flex code, that parses the bib file correctly
(some exceptions may still be there):
$cat bib.y
%{
#include <stdio.h>
%}

// Symbols.
%union
{
	char	*sval;
};
%token <sval> VALUE
%token <sval> KEY
%token OBRACE
%token EBRACE
%token QUOTE
%token SEMICOLON

%start Input
%%
Input:
     /* empty */
     | Input Entry ;  /* input is zero or more entires */
Entry:
     '@' KEY '{' KEY ','{ printf("===========\n%s : %s\n",$2, $4); }
     KeyVals '}'
     ;
KeyVals:
       /* empty */
       | KeyVals KeyVal ; /* zero or more keyvals */
KeyVal:
      KEY '=' VALUE ',' { printf("%s : %s\n",$1, $3); };

%%

int yyerror(char *s) {
  printf("yyerror : %s\n",s);
}

int main(void) {
  yyparse();
}

and
$cat bib.l
%{
#include "bib.tab.h"
%}

%%
[A-Za-z][A-Za-z0-9]*      { yylval.sval = strdup(yytext); return KEY; }
\"([^\"]|\\.)*\"|\{([^\"]|\\.)*\}	  { yylval.sval = strdup(yytext);
return VALUE; }
[ \t\n]                   ; /* ignore whitespace */
[{}@=,]                   { return *yytext; }
.                         { fprintf(stderr, "Unrecognized character %c
in input\n", *yytext); }
%%


I want to have those values in a container. For last few days, I read
the vast documentation of glib and came out with hash container as most
suitable for my case.
Below is a basic hash code, where it have the hashes correctly, once the
values are put in the array keys and vals.
#include <glib.h>
#define slen 1024

int main(gint argc, gchar** argv)
{
  char *keys[] = {"id", "type", "author", "year",NULL};
  char *vals[] = {"one",  "Book",  "RB", "2013", NULL};
  gint i;
  GHashTable* table = g_hash_table_new(g_str_hash, g_str_equal);
  GHashTableIter iter;
  g_hash_table_iter_init (&iter, table);
  for (i= 0; i<=3; i++)
  {
    g_hash_table_insert(table, keys[i],vals[i]);
    g_printf("%d=>%s:%s
\n",i,keys[i],g_hash_table_lookup(table,keys[i]));
  }
}

The problem is, how I integrate this two code, i.e. put the $1, $3 of "
KEY '=' VALUE ',' { printf("%s : %s\n",$1, $3); };" in the hash table.


Any kind help is appreciated.

[If you want to put it in the hash table, you do so, something like
  g_hash_table_insert(table, $1, $3)
-John]

[toc] | [prev] | [next] | [standalone]


#877

FromTorsten Eichstädt <torsten.eichstaedt@FernUni-Hagen.de>
Date2013-03-25 13:41 +0100
Message-ID<13-03-016@comp.compilers>
In reply to#874
Rudra Banerjee wrote:
> The problem is, how I integrate this two code, i.e. put the $1, $3 of "
> KEY '=' VALUE ',' { printf("%s : %s\n",$1, $3); };" in the hash table.
RTFM man scanf ???

And for a more complete "real" program you may want to insert error handling
in your code, i.e. detect syntax errors and things like file not found etc..

Do I remember right that bison/flex info docs contain good hints how
to do that?  Or there is a good HOWTO or docs in
/usr/share/doc/flex/bison or you search the net (some years past since
I wrote a parser w/ flex/bison).  From my experience it is good to do
so (latest) once you have a working outline (frame), and it looks like
you have one.  Do not make it too detailed, though, to not fog your
"real" code too much. Very likely you can subsume error cases for a
simple program.

Personally, I like the approach to start writing test cases and then fill in
the real code to fulfill the requirements.  This helps you to make sure you
do not miss s/th and in many cases unveils simple flaws _early_ that we
usually insert when programming.  I.e. you start w/ a program that fails --
but it tells you why!  Then, step by step, your program solves the
requirements until you're satisfied (according to the test
cases/requirements).  E.g.
/* very (too) simple implementation */
#ifdef TESTING
check( requirement_1 )
#endif
--
=|o)

[toc] | [prev] | [next] | [standalone]


#868

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2013-03-08 03:59 +0000
Message-ID<13-03-007@comp.compilers>
In reply to#864
bnrj.rudra@gmail.com wrote:
> I want to parse bibtex file using flex/bison. A sample bibtex is:
> @Book{a1,
> author="amook",
> Title="ASR",
> Publisher="oxf",
> Year="2010",
> Add="UK",
> Edition="1",
> }

(snip)

You might try writing bibtex macros that would parse them, then
write them out in an easier to parse by you form.

If you are only doing it once, you don't need the full syntax that
bibtex allows, but only what yours use. That might allow for a
simpler parser.

I presume, for example, that TeX macros could be used in the
bibtex file, which you would otherwise have to parse.

-- glen

[toc] | [prev] | [standalone]


Back to top | Article view | comp.compilers


csiph-web