Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.compilers > #1620 > unrolled thread

Natural Language Parser

Started bySeima Rao <seimarao@gmail.com>
First post2015-09-29 06:15 +0530
Last post2015-10-06 12:54 -0700
Articles 7 — 6 participants

Back to article view | Back to comp.compilers


Contents

  Natural Language Parser Seima Rao <seimarao@gmail.com> - 2015-09-29 06:15 +0530
    Re: Natural Language Parser George Neuner <gneuner2@comcast.net> - 2015-09-29 17:02 -0400
      Re: Natural Language Parser rpw3@rpw3.org (Rob Warnock) - 2015-09-30 11:05 +0000
        Re: unnatural natural language, was Natural Language Parser George Neuner <gneuner2@comcast.net> - 2015-10-01 12:41 -0400
    Re: Natural Language Parser Quinn Jackson <thothic.quinn@gmail.com> - 2015-09-29 13:11 -0300
      Re: Natural Language Parser BGB <cr88192@hotmail.com> - 2015-10-02 17:53 -0500
    Re: Natural Language Parser Gene Wirchenko <genew@telus.net> - 2015-10-06 12:54 -0700

#1620 — Natural Language Parser

FromSeima Rao <seimarao@gmail.com>
Date2015-09-29 06:15 +0530
SubjectNatural Language Parser
Message-ID<15-09-025@comp.compilers>
Hi,

    I am looking for a C++ API based English Language Parser.

    The specific task I want to do on a *regular basis* is to
    parse English Language Documents and arrive
    at a Mathematical artefact.

   The Mathematical artefact will (have) built a Syntax Tree
   of the English text that is input to the NLP parser.

   This is my only requirement. I dont need any semanticizing
    artefacts. So, my requirement is limited to parsing
   english language documents and arriving at a tree(
   or any other mathematical structure that binds the
   English grammar to the input document aka syntax trees).

   My guess is that the specific C++ API based NL parser
   will be using some english dictionary inside the tool
    to do all the jobs that are advertised.

    Can readers of this forum direct me to a stable and
    active C++ API based NL parser ?

    I have zero experience in natural language parsing(compiling)
    and zero experience in using such tools.

    However, I intend to maintain internally the source code
    of the toolkit via whatever version control software
   that is used by the developers of the tool so that
   I am able to get regular updates and not break
   anything.

Sincerely,
Seima Rao.
[You might start with this parser from Stanford:
http://nlp.stanford.edu/software/lex-parser.shtml
Or this one in python:
http://spacy.io/
Parsing English or any natural language is very hard, and you'll never
get more than an approximate result.  Modern language translation
systems don't even try and use machine learning on large corpora.
-John]

[toc] | [next] | [standalone]


#1623

FromGeorge Neuner <gneuner2@comcast.net>
Date2015-09-29 17:02 -0400
Message-ID<15-09-028@comp.compilers>
In reply to#1620
On Tue, 29 Sep 2015 06:15:22 +0530, Seima Rao <seimarao@gmail.com>
wrote:

>    I am looking for a C++ API based English Language Parser.

Sorry, I don't know of one offhand.  Most I seen have been written
either in Lisp or in Prolog.

The Stanford and Berkeley projects have parsers available in Java, but
they are very complex and may be difficult to port to C++ (if you even
would want to).


>    The specific task I want to do on a *regular basis* is to
>    parse English Language Documents and arrive
>    at a Mathematical artefact.

There's little "mathematical" about natural language - it's a jumbled
heap of garbage though which we dig for nuggets of meaning.  That it
can be mechanically understood well enough to be badly translated is
itself something of a miracle.


>   The Mathematical artefact will (have) built a Syntax Tree
>   of the English text that is input to the NLP parser.
>
>   This is my only requirement. I dont need any semanticizing
>    artefacts. So, my requirement is limited to parsing
>   english language documents and arriving at a tree(
>   or any other mathematical structure that binds the
>   English grammar to the input document aka syntax trees).

You need to be aware that natural languages *can't* be parsed without
semantics - i.e. without considering "parts of speech".

Depending on context, the same word may represent, e.g., a verb, an
adverb, or even a (type of) noun.  What part of speech the word
represents in context determines the ultimate meaning of the sentence.
Even which part of speech a word represents may be controversial. Most
human writing (and speaking) is quite imprecise: most sentences can be
parsed in more than one way, and the meanings of the various parsings
may be very different.


>   My guess is that the specific C++ API based NL parser
>   will be using some english dictionary inside the tool
>    to do all the jobs that are advertised.
>
>    Can readers of this forum direct me to a stable and
>    active C++ API based NL parser ?

As John mentioned already, Stanford has a good NLP project.  Berkeley
also has good project.

IMO, Princeton's Wordnet is the most comprehensive English language
POS database available.  It has a C library interface.  There are some
toy projects that demonstrate using it, but Princeton primarily has
focused on the database itself rather than on developing tools that
use it.
see https://wordnet.princeton.edu


>    I have zero experience in natural language parsing(compiling)
>    and zero experience in using such tools.
>
>   However, I intend to maintain internally the source code
>   of the toolkit via whatever version control software
>   that is used by the developers of the tool so that
>   I am able to get regular updates and not break  anything.

Most so-called "NLP" systems depend on operating within a very limited
scope - e.g., needing to "understand" only the commands and data
objects of a single program.  General purpose NLP is supercomputer
territory [think IBM's Watson] ... Siri and Cortana, etc. - systems
which _appear_ to understand language - really are keyword driven and
don't actually understand anything at all.

If you intend for this to be some kind of general purpose tool, then
you need to do a LOT more research before you start.

George

[toc] | [prev] | [next] | [standalone]


#1627

Fromrpw3@rpw3.org (Rob Warnock)
Date2015-09-30 11:05 +0000
Message-ID<15-09-032@comp.compilers>
In reply to#1623
George Neuner  <gneuner2@comcast.net> wrote:
+---------------
| Seima Rao <seimarao@gmail.com> wrote:
| > I am looking for a C++ API based English Language Parser.
...
| > This is my only requirement. I dont need any semanticizing
| > artefacts.
|
| You need to be aware that natural languages *can't* be parsed
| without semantics - i.e. without considering "parts of speech".
|
| Depending on context, the same word may represent, e.g., a verb, an
| adverb, or even a (type of) noun.  What part of speech the word
| represents in context determines the ultimate meaning of the sentence.
| Even which part of speech a word represents may be controversial. Most
| human writing (and speaking) is quite imprecise: most sentences can be
| parsed in more than one way, and the meanings of the various parsings
| may be very different.
+---------------

Exactly. Consider the classical "Buffalo" example:

    https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffalo_buffalo_buffalo_Buffalo_buffalo
    "Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo"
    is a grammatically correct sentence in the English language,
    used as an example of how homonyms and homophones can be used
    to create complicated linguistic constructs.
    ...
    Thomas Tymoczko has pointed out that there is nothing special
    about eight "buffalos"; any sentence consisting solely of the
    word "buffalo" repeated any number of times is grammatically
    correct.
    ...
    Versions of the linguistic oddity can be constructed with
    other words which similarly simultaneously serve as collective
    noun, adjective, and verb, some of which need no capitalization
    (such as "police").


-Rob

-----
Rob Warnock		<rpw3@rpw3.org>
627 26th Avenue		<http://rpw3.org/>
San Mateo, CA 94403

[toc] | [prev] | [next] | [standalone]


#1628 — Re: unnatural natural language, was Natural Language Parser

FromGeorge Neuner <gneuner2@comcast.net>
Date2015-10-01 12:41 -0400
SubjectRe: unnatural natural language, was Natural Language Parser
Message-ID<15-10-001@comp.compilers>
In reply to#1627
>| > I am looking for a C++ API based English Language Parser. ....
>|
>| You need to be aware that natural languages *can't* be parsed
>| without semantics - i.e. without considering "parts of speech".

>Exactly. Consider the classical "Buffalo" example:
>
>    https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffalo_buffalo_buffalo_Buffalo_buffalo
>    "Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo"

My family is from Buffalo.

That particular example actually gets even more perverse than the
article indicates.   The staff at the Buffalo Zoo claim that their
animals have a unique method of harassing each other which they call
the "Buffalo buffalo".

If the animals really do "Buffalo buffalo" each other, then it follows
that "Buffalo buffalo Buffalo buffalo Buffalo buffalo Buffalo buffalo
Buffalo buffalo".  [That's 10 "buffalos" if you are counting.]
<grin>


FWIW:
  1) The animals in the Buffalo Zoo actually are bison.

  2) Bison are related to buffalo, but are a different species.

  3) Natives pronounce the city's name as "buf4 lo",
      accent on the 1st syllable and without any 'a' sound.
      There are competing  theories regarding the evolution
      of this pronunciation: one claim that it is a mangled
      Seneca Indian name, and two different claims that it
      is mangled French.

YMMV,
George
[I think it's time to get back to computer languages, perhaps after we
discuss the derivation of beef on weck. -John]

[toc] | [prev] | [next] | [standalone]


#1626

FromQuinn Jackson <thothic.quinn@gmail.com>
Date2015-09-29 13:11 -0300
Message-ID<15-09-031@comp.compilers>
In reply to#1620
On Mon, Sep 28, 2015 at 9:45 PM, Seima Rao <seimarao@gmail.com> wrote:
>
>  I am looking for a C++ API based English Language Parser.

I sit in front of one of these beasts every day.

And yes, per John: "Parsing English or any natural language is very hard..."

But hey -- someone has to do it. ;-)

--
Quinn Jackson

LinkedIn:            http://ca.linkedin.com/in/quinnjackson/
ResearchGate:  http://researchgate.net/profile/Quinn_Jackson/
[Unfortunately, he says it's proprietary. -John]

[toc] | [prev] | [next] | [standalone]


#1629

FromBGB <cr88192@hotmail.com>
Date2015-10-02 17:53 -0500
Message-ID<15-10-002@comp.compilers>
In reply to#1626
On 9/29/2015 11:11 AM, Quinn Jackson wrote:
> On Mon, Sep 28, 2015 at 9:45 PM, Seima Rao <seimarao@gmail.com> wrote:
>>
>>   I am looking for a C++ API based English Language Parser.
>
> I sit in front of one of these beasts every day.
>
> And yes, per John: "Parsing English or any natural language is very hard..."
>
> But hey -- someone has to do it. ;-)

IIRC, I once did a parser for an English subset, mostly by creating
the subset where each word only had a single word type.

Likewise, a finite dictionary was used, and constraints were put on
the grammatical constructs allowed. I remember it having taken some
information from Basic English, but I forget the specifics (IIRC, it
was mostly word lists and other things, but I still had to do a little
work to figure out the parsing rules for the grammar).

With this much constraining, it wasn't really all that much different
from parsing a something like a programming language, and I could use
a fairly straightforward recursive descent parser.

Previously (before this point), I had done similar with an Esperanto
variant, where one can (more or less) rely on the word suffixes to
disambiguate the word-types and desired syntax tree. I realized though
that if you know the word types from a dictionary, the suffixes (and
the use of non-English vocabulary) is unnecessary.

IIRC, some of the metadata for this, was sort of an English/Esperanto
mash-up (and there was a notation for sticking suffixes onto words).


From what I remember, what killed the effort at the time, was that I
couldn't figure out any good semantic model to map this onto. the
problem was that, language without semantics isn't particularly
useful.

You can do crude machine translation or similar, but nearly anything
"interesting" you could do would require a semantic model and some
form of rudimentary "intelligence".

other basic forms of "AI" don't really need grammar trees, either
Responding to keywords, or to temporal associations between words
(with no respect paid to grammatical structure).

So, it all goes in my "stuff that can be done but lacks any obvious
use-case" bin (sort of like me trying to find a use-case for neural-nets
which isn't better served via more conventional strategies, *, and I
still can't really make object and speech recognition work to a usable
level).

Similarly, I have had issues when it comes to getting particularly
intelligible results from text-to-speech (my best results were by
mixing together diphone synthesis with a recorded list of common
words, my past attempts at formant synthesis generally not working so
well and producing mostly unintelligible results).

*: Given CPU power is a finite resource, and NNs tend to boil down
mostly to rather inefficiently implemented signal filters. like with
genetic programming, any "intelligent" behavior is elusive, and GP is
mostly good at either finding mediocre patterns and/or a way to break
the test (though, is at least does "ok" at finding and fine-tuning
heuristics for signal filters).


Likewise, a fixed-grammar parser wont really give sane results though if
given free-form natural language: it would be necessary to write in the
subset of the language that the parser is able to understand.

If writing for such a parser, such a subset isn't particularly difficult
apart from the tendency of one to forget about it and write phases
outside those allowed by the grammar.

[toc] | [prev] | [next] | [standalone]


#1630

FromGene Wirchenko <genew@telus.net>
Date2015-10-06 12:54 -0700
Message-ID<15-10-003@comp.compilers>
In reply to#1620
On Tue, 29 Sep 2015 06:15:22 +0530, Seima Rao <seimarao@gmail.com>
wrote:

[snip]

>[You might start with this parser from Stanford:
>http://nlp.stanford.edu/software/lex-parser.shtml
>Or this one in python:
>http://spacy.io/
>Parsing English or any natural language is very hard, and you'll never
>get more than an approximate result.  Modern language translation
>systems don't even try and use machine learning on large corpora.
>-John]

     Another area to try is some of the interactive fiction (also
known as text adventures) languages.

Sincerely,

Gene Wirchenko

[toc] | [prev] | [standalone]


Back to top | Article view | comp.compilers


csiph-web