Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.compilers > #1620 > unrolled thread
| Started by | Seima Rao <seimarao@gmail.com> |
|---|---|
| First post | 2015-09-29 06:15 +0530 |
| Last post | 2015-10-06 12:54 -0700 |
| Articles | 7 — 6 participants |
Back to article view | Back to comp.compilers
Natural Language Parser Seima Rao <seimarao@gmail.com> - 2015-09-29 06:15 +0530
Re: Natural Language Parser George Neuner <gneuner2@comcast.net> - 2015-09-29 17:02 -0400
Re: Natural Language Parser rpw3@rpw3.org (Rob Warnock) - 2015-09-30 11:05 +0000
Re: unnatural natural language, was Natural Language Parser George Neuner <gneuner2@comcast.net> - 2015-10-01 12:41 -0400
Re: Natural Language Parser Quinn Jackson <thothic.quinn@gmail.com> - 2015-09-29 13:11 -0300
Re: Natural Language Parser BGB <cr88192@hotmail.com> - 2015-10-02 17:53 -0500
Re: Natural Language Parser Gene Wirchenko <genew@telus.net> - 2015-10-06 12:54 -0700
| From | Seima Rao <seimarao@gmail.com> |
|---|---|
| Date | 2015-09-29 06:15 +0530 |
| Subject | Natural Language Parser |
| Message-ID | <15-09-025@comp.compilers> |
Hi,
I am looking for a C++ API based English Language Parser.
The specific task I want to do on a *regular basis* is to
parse English Language Documents and arrive
at a Mathematical artefact.
The Mathematical artefact will (have) built a Syntax Tree
of the English text that is input to the NLP parser.
This is my only requirement. I dont need any semanticizing
artefacts. So, my requirement is limited to parsing
english language documents and arriving at a tree(
or any other mathematical structure that binds the
English grammar to the input document aka syntax trees).
My guess is that the specific C++ API based NL parser
will be using some english dictionary inside the tool
to do all the jobs that are advertised.
Can readers of this forum direct me to a stable and
active C++ API based NL parser ?
I have zero experience in natural language parsing(compiling)
and zero experience in using such tools.
However, I intend to maintain internally the source code
of the toolkit via whatever version control software
that is used by the developers of the tool so that
I am able to get regular updates and not break
anything.
Sincerely,
Seima Rao.
[You might start with this parser from Stanford:
http://nlp.stanford.edu/software/lex-parser.shtml
Or this one in python:
http://spacy.io/
Parsing English or any natural language is very hard, and you'll never
get more than an approximate result. Modern language translation
systems don't even try and use machine learning on large corpora.
-John]
[toc] | [next] | [standalone]
| From | George Neuner <gneuner2@comcast.net> |
|---|---|
| Date | 2015-09-29 17:02 -0400 |
| Message-ID | <15-09-028@comp.compilers> |
| In reply to | #1620 |
On Tue, 29 Sep 2015 06:15:22 +0530, Seima Rao <seimarao@gmail.com> wrote: > I am looking for a C++ API based English Language Parser. Sorry, I don't know of one offhand. Most I seen have been written either in Lisp or in Prolog. The Stanford and Berkeley projects have parsers available in Java, but they are very complex and may be difficult to port to C++ (if you even would want to). > The specific task I want to do on a *regular basis* is to > parse English Language Documents and arrive > at a Mathematical artefact. There's little "mathematical" about natural language - it's a jumbled heap of garbage though which we dig for nuggets of meaning. That it can be mechanically understood well enough to be badly translated is itself something of a miracle. > The Mathematical artefact will (have) built a Syntax Tree > of the English text that is input to the NLP parser. > > This is my only requirement. I dont need any semanticizing > artefacts. So, my requirement is limited to parsing > english language documents and arriving at a tree( > or any other mathematical structure that binds the > English grammar to the input document aka syntax trees). You need to be aware that natural languages *can't* be parsed without semantics - i.e. without considering "parts of speech". Depending on context, the same word may represent, e.g., a verb, an adverb, or even a (type of) noun. What part of speech the word represents in context determines the ultimate meaning of the sentence. Even which part of speech a word represents may be controversial. Most human writing (and speaking) is quite imprecise: most sentences can be parsed in more than one way, and the meanings of the various parsings may be very different. > My guess is that the specific C++ API based NL parser > will be using some english dictionary inside the tool > to do all the jobs that are advertised. > > Can readers of this forum direct me to a stable and > active C++ API based NL parser ? As John mentioned already, Stanford has a good NLP project. Berkeley also has good project. IMO, Princeton's Wordnet is the most comprehensive English language POS database available. It has a C library interface. There are some toy projects that demonstrate using it, but Princeton primarily has focused on the database itself rather than on developing tools that use it. see https://wordnet.princeton.edu > I have zero experience in natural language parsing(compiling) > and zero experience in using such tools. > > However, I intend to maintain internally the source code > of the toolkit via whatever version control software > that is used by the developers of the tool so that > I am able to get regular updates and not break anything. Most so-called "NLP" systems depend on operating within a very limited scope - e.g., needing to "understand" only the commands and data objects of a single program. General purpose NLP is supercomputer territory [think IBM's Watson] ... Siri and Cortana, etc. - systems which _appear_ to understand language - really are keyword driven and don't actually understand anything at all. If you intend for this to be some kind of general purpose tool, then you need to do a LOT more research before you start. George
[toc] | [prev] | [next] | [standalone]
| From | rpw3@rpw3.org (Rob Warnock) |
|---|---|
| Date | 2015-09-30 11:05 +0000 |
| Message-ID | <15-09-032@comp.compilers> |
| In reply to | #1623 |
George Neuner <gneuner2@comcast.net> wrote:
+---------------
| Seima Rao <seimarao@gmail.com> wrote:
| > I am looking for a C++ API based English Language Parser.
...
| > This is my only requirement. I dont need any semanticizing
| > artefacts.
|
| You need to be aware that natural languages *can't* be parsed
| without semantics - i.e. without considering "parts of speech".
|
| Depending on context, the same word may represent, e.g., a verb, an
| adverb, or even a (type of) noun. What part of speech the word
| represents in context determines the ultimate meaning of the sentence.
| Even which part of speech a word represents may be controversial. Most
| human writing (and speaking) is quite imprecise: most sentences can be
| parsed in more than one way, and the meanings of the various parsings
| may be very different.
+---------------
Exactly. Consider the classical "Buffalo" example:
https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffalo_buffalo_buffalo_Buffalo_buffalo
"Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo"
is a grammatically correct sentence in the English language,
used as an example of how homonyms and homophones can be used
to create complicated linguistic constructs.
...
Thomas Tymoczko has pointed out that there is nothing special
about eight "buffalos"; any sentence consisting solely of the
word "buffalo" repeated any number of times is grammatically
correct.
...
Versions of the linguistic oddity can be constructed with
other words which similarly simultaneously serve as collective
noun, adjective, and verb, some of which need no capitalization
(such as "police").
-Rob
-----
Rob Warnock <rpw3@rpw3.org>
627 26th Avenue <http://rpw3.org/>
San Mateo, CA 94403
[toc] | [prev] | [next] | [standalone]
| From | George Neuner <gneuner2@comcast.net> |
|---|---|
| Date | 2015-10-01 12:41 -0400 |
| Subject | Re: unnatural natural language, was Natural Language Parser |
| Message-ID | <15-10-001@comp.compilers> |
| In reply to | #1627 |
>| > I am looking for a C++ API based English Language Parser. ....
>|
>| You need to be aware that natural languages *can't* be parsed
>| without semantics - i.e. without considering "parts of speech".
>Exactly. Consider the classical "Buffalo" example:
>
> https://en.wikipedia.org/wiki/Buffalo_buffalo_Buffalo_buffalo_buffalo_buffalo_Buffalo_buffalo
> "Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo"
My family is from Buffalo.
That particular example actually gets even more perverse than the
article indicates. The staff at the Buffalo Zoo claim that their
animals have a unique method of harassing each other which they call
the "Buffalo buffalo".
If the animals really do "Buffalo buffalo" each other, then it follows
that "Buffalo buffalo Buffalo buffalo Buffalo buffalo Buffalo buffalo
Buffalo buffalo". [That's 10 "buffalos" if you are counting.]
<grin>
FWIW:
1) The animals in the Buffalo Zoo actually are bison.
2) Bison are related to buffalo, but are a different species.
3) Natives pronounce the city's name as "buf4 lo",
accent on the 1st syllable and without any 'a' sound.
There are competing theories regarding the evolution
of this pronunciation: one claim that it is a mangled
Seneca Indian name, and two different claims that it
is mangled French.
YMMV,
George
[I think it's time to get back to computer languages, perhaps after we
discuss the derivation of beef on weck. -John]
[toc] | [prev] | [next] | [standalone]
| From | Quinn Jackson <thothic.quinn@gmail.com> |
|---|---|
| Date | 2015-09-29 13:11 -0300 |
| Message-ID | <15-09-031@comp.compilers> |
| In reply to | #1620 |
On Mon, Sep 28, 2015 at 9:45 PM, Seima Rao <seimarao@gmail.com> wrote: > > I am looking for a C++ API based English Language Parser. I sit in front of one of these beasts every day. And yes, per John: "Parsing English or any natural language is very hard..." But hey -- someone has to do it. ;-) -- Quinn Jackson LinkedIn: http://ca.linkedin.com/in/quinnjackson/ ResearchGate: http://researchgate.net/profile/Quinn_Jackson/ [Unfortunately, he says it's proprietary. -John]
[toc] | [prev] | [next] | [standalone]
| From | BGB <cr88192@hotmail.com> |
|---|---|
| Date | 2015-10-02 17:53 -0500 |
| Message-ID | <15-10-002@comp.compilers> |
| In reply to | #1626 |
On 9/29/2015 11:11 AM, Quinn Jackson wrote: > On Mon, Sep 28, 2015 at 9:45 PM, Seima Rao <seimarao@gmail.com> wrote: >> >> I am looking for a C++ API based English Language Parser. > > I sit in front of one of these beasts every day. > > And yes, per John: "Parsing English or any natural language is very hard..." > > But hey -- someone has to do it. ;-) IIRC, I once did a parser for an English subset, mostly by creating the subset where each word only had a single word type. Likewise, a finite dictionary was used, and constraints were put on the grammatical constructs allowed. I remember it having taken some information from Basic English, but I forget the specifics (IIRC, it was mostly word lists and other things, but I still had to do a little work to figure out the parsing rules for the grammar). With this much constraining, it wasn't really all that much different from parsing a something like a programming language, and I could use a fairly straightforward recursive descent parser. Previously (before this point), I had done similar with an Esperanto variant, where one can (more or less) rely on the word suffixes to disambiguate the word-types and desired syntax tree. I realized though that if you know the word types from a dictionary, the suffixes (and the use of non-English vocabulary) is unnecessary. IIRC, some of the metadata for this, was sort of an English/Esperanto mash-up (and there was a notation for sticking suffixes onto words). From what I remember, what killed the effort at the time, was that I couldn't figure out any good semantic model to map this onto. the problem was that, language without semantics isn't particularly useful. You can do crude machine translation or similar, but nearly anything "interesting" you could do would require a semantic model and some form of rudimentary "intelligence". other basic forms of "AI" don't really need grammar trees, either Responding to keywords, or to temporal associations between words (with no respect paid to grammatical structure). So, it all goes in my "stuff that can be done but lacks any obvious use-case" bin (sort of like me trying to find a use-case for neural-nets which isn't better served via more conventional strategies, *, and I still can't really make object and speech recognition work to a usable level). Similarly, I have had issues when it comes to getting particularly intelligible results from text-to-speech (my best results were by mixing together diphone synthesis with a recorded list of common words, my past attempts at formant synthesis generally not working so well and producing mostly unintelligible results). *: Given CPU power is a finite resource, and NNs tend to boil down mostly to rather inefficiently implemented signal filters. like with genetic programming, any "intelligent" behavior is elusive, and GP is mostly good at either finding mediocre patterns and/or a way to break the test (though, is at least does "ok" at finding and fine-tuning heuristics for signal filters). Likewise, a fixed-grammar parser wont really give sane results though if given free-form natural language: it would be necessary to write in the subset of the language that the parser is able to understand. If writing for such a parser, such a subset isn't particularly difficult apart from the tendency of one to forget about it and write phases outside those allowed by the grammar.
[toc] | [prev] | [next] | [standalone]
| From | Gene Wirchenko <genew@telus.net> |
|---|---|
| Date | 2015-10-06 12:54 -0700 |
| Message-ID | <15-10-003@comp.compilers> |
| In reply to | #1620 |
On Tue, 29 Sep 2015 06:15:22 +0530, Seima Rao <seimarao@gmail.com>
wrote:
[snip]
>[You might start with this parser from Stanford:
>http://nlp.stanford.edu/software/lex-parser.shtml
>Or this one in python:
>http://spacy.io/
>Parsing English or any natural language is very hard, and you'll never
>get more than an approximate result. Modern language translation
>systems don't even try and use machine learning on large corpora.
>-John]
Another area to try is some of the interactive fiction (also
known as text adventures) languages.
Sincerely,
Gene Wirchenko
[toc] | [prev] | [standalone]
Back to top | Article view | comp.compilers
csiph-web