Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #266281

Re: counting commas

From Greg Wooledge <greg@wooledge.org>
Newsgroups linux.debian.user
Subject Re: counting commas
Date 2024-01-19 17:50 +0100
Message-ID <HY0e6-4aNp-5@gated-at.bofh.it> (permalink)
References <HXWk9-48pp-1@gated-at.bofh.it> <HXYYF-4a8y-5@gated-at.bofh.it> <HXZ8m-4abx-7@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


On Fri, Jan 19, 2024 at 03:30:17PM +0000, fxkl47BF@protonmail.com wrote:
> >> But at this point, we have to wonder what the *actual* goal is.
> 
> to exclude phrases with commas for seperate examination

Parsing natural language text is going to be tricky.  I can only talk
about English, and not about whatever language your text is actually
written in.

Let's look at a few example English sentences:

    Good morning, John.

    I went to the store with Mary, Paul, Susan and Ralph.

    I won, and you lost.

    The bear, who was hungry, looked for food.

    Oh, that's interesting.

These are five different examples of comma usage in English.  Do you
happen to know in advance that your text will *only* contain samples
that use the fourth style above?  Let's assume this.  Let's then form
a template:

    STUFF, ASIDE, MORE STUFF, ASIDE, STILL MORE STUFF.

I.e. given a sentence which conforms to expectation, we should see
an even number of commas (is *THIS* why you were counting them??) and
we should extract the ASIDEs from in between the first and second, then
the third and fourth, and so on.

So... uh, I guess my next question is: are you *pre-filtering* the
sentences and keeping only the ones which have an even number of
commas?  Or have you already *done* that, and now you're asking how
to extract the ASIDEs?

I really don't think I'd try this with shell scripts.  The tools just
aren't designed for this.  You really want tools that are custom built
for natural language processing, or a language that lets you run
through a large string character by character in a fast, efficient
way (C comes to mind) if you're trying to build your tools from the
ground up.

The "obvious" algorithm for extracting the ASIDEs would be use a
simple finite state machine, and march through the sentence
character by character.  When you encounter a comma, change state.
Otherwise, if you're in the "ASIDE" state, copy the character to your
output buffer.  When you leave the "ASIDE" state, terminate the current
output buffer and move to the next one.  That's how I'd do it in C.
Add whitespace trimming and so on.

Also note that breaking a piece of natural language text *into*
sentences in the first place is extraordinarily difficult.  If you
haven't already got a way to do that, you're probably screwed.
Seriously, asking the debian-user list how to count the number of
commas in a text file is *not* a good sign if you're dealing with a
masters-degree-level problem in natural language analysis.

Back to linux.debian.user | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

counting commas fxkl47BF@protonmail.com - 2024-01-19 08:20 +0100
  Re: counting commas John Crawley <john@bunsenlabs.org> - 2024-01-19 08:30 +0100
  Re: counting commas <tomas@tuxteam.de> - 2024-01-19 09:10 +0100
  Re: counting commas "Thomas Schmitt" <scdbackup@gmx.net> - 2024-01-19 09:30 +0100
    Re: counting commas Michael Grant <mgrant@grant.org> - 2024-01-19 12:20 +0100
    Re: counting commas Greg Wooledge <greg@wooledge.org> - 2024-01-19 13:40 +0100
      Re: counting commas "Thomas Schmitt" <scdbackup@gmx.net> - 2024-01-19 16:30 +0100
        Re: counting commas fxkl47BF@protonmail.com - 2024-01-19 16:40 +0100
          Re: counting commas Greg Wooledge <greg@wooledge.org> - 2024-01-19 17:50 +0100
            Re: counting commas fxkl47BF@protonmail.com - 2024-01-19 18:30 +0100
            Re: counting commas debian-user@howorth.org.uk - 2024-01-19 18:30 +0100
              Re: counting commas John Hasler <john@sugarbit.com> - 2024-01-19 21:10 +0100
                Re: counting commas Nicholas Geovanis <nickgeovanis@gmail.com> - 2024-01-20 04:00 +0100
                Re: counting commas Nicholas Geovanis <nickgeovanis@gmail.com> - 2024-01-20 05:00 +0100
              Re: counting commas David Wright <deblis@lionunicorn.co.uk> - 2024-01-20 00:10 +0100
                Re: counting commas Cindy Sue Causey <butterflybytes@gmail.com> - 2024-01-20 01:10 +0100
                Re: counting commas Peter Hillier-Brook <phb@hbsys.plus.com> - 2024-01-20 01:10 +0100
                Re: counting commas Curt <curty@free.fr> - 2024-01-20 18:20 +0100
                Re: counting commas Greg Wooledge <greg@wooledge.org> - 2024-01-20 18:30 +0100
                Re: counting commas fxkl47BF@protonmail.com - 2024-01-20 18:40 +0100
                Re: counting commas Curt <curty@free.fr> - 2024-01-20 18:40 +0100
                Re: counting commas David Wright <deblis@lionunicorn.co.uk> - 2024-01-20 21:10 +0100
        Re: counting commas Nicholas Geovanis <nickgeovanis@gmail.com> - 2024-01-20 03:50 +0100
          Re: counting commas gene heskett <gheskett@shentel.net> - 2024-01-20 06:50 +0100
          Re: counting commas gene heskett <gheskett@shentel.net> - 2024-01-20 07:00 +0100
          Re: counting commas "Roy J. Tellason, Sr." <roy@rtellason.com> - 2024-01-20 21:40 +0100
            Re: counting commas gene heskett <gheskett@shentel.net> - 2024-01-20 22:10 +0100
            Re: counting commas John Hasler <john@sugarbit.com> - 2024-01-21 01:10 +0100
              Re: counting commas gene heskett <gheskett@shentel.net> - 2024-01-21 01:20 +0100
                Re: counting commas John Hasler <john@sugarbit.com> - 2024-01-21 01:30 +0100
                Re: counting commas gene heskett <gheskett@shentel.net> - 2024-01-21 02:10 +0100

csiph-web