Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #3855 > unrolled thread

preg_match() oddities and question

Started bySandman <mr@sandman.net>
First post2011-11-22 12:21 +0100
Last post2011-11-26 11:21 +0100
Articles 14 on this page of 34 — 7 participants

Back to article view | Back to comp.lang.php


Contents

  preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:21 +0100
    Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 11:26 +0000
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:36 +0100
      Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-22 07:22 -0500
    Re: preg_match() oddities and question tony@mountifield.org (Tony Mountifield) - 2011-11-22 11:47 +0000
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:12 +0100
    Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 13:30 +0100
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:55 +0100
        Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 17:56 +0100
          Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 17:30 +0000
            Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-22 17:20 -0600
              Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 23:59 +0000
              Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 01:59 +0100
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:58 +0100
                  Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 22:02 +0100
                    Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:20 +0100
                      Re: preg_match() oddities and question Denis McMahon <denismfmcmahon@gmail.com> - 2011-11-24 12:55 +0000
                        Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:36 +0100
                      Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-24 22:41 +0100
                        Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:26 +0100
                          Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 15:44 +0100
                            Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 16:34 +0100
                              Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 23:23 +0100
                Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 09:35 +0000
          Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:55 +0100
            Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 07:53 -0600
              Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 19:01 +0100
                Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 18:54 +0000
                  Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 20:23 +0100
                Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 12:58 -0600
                  Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:28 +0100
    SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 10:32 +0100
      Re: SOLVED: Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-25 18:55 -0500
        Re: SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-26 11:21 +0100

Page 2 of 2 — ← Prev page 1 [2]


#3933

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-25 15:44 +0100
Message-ID<1438797.UceJUlZ0hu@PointedEars.de>
In reply to#3922
Sandman wrote:

>  Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> >> With separate controls they can be sure where to enter what;
>> >> it is accessible, and you have no problem processing the data.  With
>> >> one control, neither applies.
>> > Well, I have been doing this for about ten years now,
>> I have been doing this for about fourteen years now.  So what?
> 
> You have monitored swedish address search terms for fourteen years? We
> should compare notes.

You are missing the point.  The kind of structured data that you need to 
enter in a form does not matter.  Using cursor keys to move the text cursor 
between delimiters in a running text always causes has more accessibility 
and usability problems, and consequently information processing problems, 
than tabbing (or otherwise moving the focus) from one control to another.
 
>> There are basic accessibility guidelines that no amount of
>> development experience can substitute (although studying usability,
>> as I did, can help).  Many of which must be followed per
>> legislation in some countries.
> 
> This.. has nothing to do with the topic at hand.

You are just not seeing how much it has to do with the topic at hand.  
Changing the way people put in data towards one that is *actually* easier 
for them solves, at least, three problems at once, including the one that 
you have been asking about.  Because when structured data is being entered 
in a structured way, you have almost no trouble storing it in a structured 
way.  (BTW, in case you do not know, moving the focus from one input control 
to the next while typing text can be assisted with client-side scripting.)

You should read, for example, <http://www.useit.com/> which contains many 
ideas on how you can make Web sites better usable and therefore, almost in 
passing, more accessible.

>> > When I say it's inconvenient for the end user, it's not something I
>> > make up on the spot to be obnoxious.
>> Nevertheless, your logic is flawed.
> 
> Well, as long as you're merely saying that instead of actually, you 
> know, substantiate that opinion, I have no idea what you expect me 
> to do with it.
> 
> Words are easy :)

This is about as much as I will discuss this here because your *actual* 
problem has nothing to do with PHP, and little to do with Regular 
Expressions.


HTH

PointedEars
-- 
var bugRiddenCrashPronePieceOfJunk = (
    navigator.userAgent.indexOf('MSIE 5') != -1
    && navigator.userAgent.indexOf('Mac') != -1
)  // Plone, register_function.js:16

[toc] | [prev] | [next] | [standalone]


#3934

FromSandman <mr@sandman.net>
Date2011-11-25 16:34 +0100
Message-ID<mr-685156.16344225112011@News.Individual.NET>
In reply to#3933
In article <1438797.UceJUlZ0hu@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> > You have monitored swedish address search terms for fourteen years? We
> > should compare notes.
> 
> You are missing the point.

I disagree. :)

> The kind of structured data that you need to 
> enter in a form does not matter.  Using cursor keys to move the text cursor 
> between delimiters in a running text always causes has more accessibility 
> and usability problems, and consequently information processing problems, 
> than tabbing (or otherwise moving the focus) from one control to another.

Whatever gave you the idea that I am proposing a solution where the 
user would have to "use cursor keys to move the text cursor between 
delimiter in a running text"? I literally have no idea what you are 
talking about, and it has absolutely nothing to do with this thread.

> >> There are basic accessibility guidelines that no amount of
> >> development experience can substitute (although studying usability,
> >> as I did, can help).  Many of which must be followed per
> >> legislation in some countries.
> > 
> > This.. has nothing to do with the topic at hand.
> 
> You are just not seeing how much it has to do with the topic at hand.  

And you seem to be failing to explain how it does :)

> Changing the way people put in data towards one that is *actually* easier 
> for them solves, at least, three problems at once, including the one that 
> you have been asking about.

Adding form fields to make it more cumbersome for them to search for 
their address, however, does not.

My *experience* (i.e. not guessing, but rather - having done it 
exactly the way you propose) tells me that adding form fields to this 
situation adds more illegal search terms. Users tend to make more 
mistakes the more details you expect them to provide. Users often 
entered "X", "?" or "-" in the letter search field, or tried to type 
"none". Most of the time, however, they continued to type their entire 
address in the first field and then hit return.

When you find yourself having to educate or expect the user to provide 
data in a specific way is when you fail as a software engineer. You 
have to make it as easy as possible for them to search for their 
address - and the thing you're dealing with here is *Google*. People 
know how to Google, they Google whatever shit they can and Google just 
always manages to figure out pretty much exactly what they need - with 
one input field. That's the level your visitors are on. 

They shouldn't have to read labels or instructions to search for their 
address. Also the meaning of the street letter may be different 
depending on what kind of property you live in and whether your'e a 
company or a private person. All that has to be explained for these 
users, and my experience (i.e. actual facts provided by years and 
years on monitoring exactly this) shows me that this is not the 
correct way to deal with input data.

You are free to disagree all you want, and perhaps visitors to your 
applications and your digested search term analysis show you something 
else, but mine does not. When I wrote the OP and provided the examples 
it wasn't something out of the blue.

> >> > When I say it's inconvenient for the end user, it's not something I
> >> > make up on the spot to be obnoxious.
> >> Nevertheless, your logic is flawed.
> > 
> > Well, as long as you're merely saying that instead of actually, you 
> > know, substantiate that opinion, I have no idea what you expect me 
> > to do with it.
> > 
> > Words are easy :)
> 
> This is about as much as I will discuss this here because your *actual* 
> problem has nothing to do with PHP, and little to do with Regular 
> Expressions.

If you saw my followup to my OP you may have seen that I found the 
solution elsewhere and that it indeed was solved using PHp and regular 
expressions - just as I knew it could be :)





-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3935

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-25 23:23 +0100
Message-ID<7038554.eTKvZm4JqM@PointedEars.de>
In reply to#3934
Sandman wrote:

> In article <1438797.UceJUlZ0hu@PointedEars.de>,
>> The kind of structured data that you need to enter in a form does not
>> matter.  Using cursor keys to move the text cursor between delimiters in
>> a running text always causes has more accessibility and usability
>> problems, and consequently information processing problems, than tabbing
>> (or otherwise moving the focus) from one control to another.
> 
> Whatever gave you the idea that I am proposing a solution where the
> user would have to "use cursor keys to move the text cursor between
> delimiter in a running text"? I literally have no idea what you are
> talking about, and it has absolutely nothing to do with this thread.

I am imagining a visitor of your Web site entering "foo street 34" (perhaps 
as part of a longer text; your next question about phone numbers indicates 
just that).  Seeing that the "34" needed to be "42" instead, they would tab 
back, upon which the entire text field is usually selected, and they have to 
press the End key (or Alt+Arrow Right on Macs, IIRC) to get to the end.  
Then they would perhaps press the Backspace key twice and type "42"  Or 
perhaps they needed it to be "134", so they would perhaps press the Arrow 
Left or Backspace key twice (or Ctrl/Compose+Left) before they could make 
their edit.  (It would be comparably tedious with a pointing device, so do 
not get me started on that.)

Now, it would be so much easier for them to work with the form, and easier 
for you to process the data they entered, if you had the street name and the 
house number in separate controls.  Then they would tab back to the house 
number control, which text would be selected, and they could immediately fix 
their mistake.  In addition, you could set up access keys and labels for the 
controls with pure HTML so that users can use either way to focus the 
control they want to edit.  This would work with a graphical browser, a text 
browser, a screen reader etc.  Lately, it would work best with a mobile 
device (anyone who has experienced first-hand how tedious it is to position 
the text cursor with a mobile device, especially with a touchscreen, knows 
what I mean).  You cannot do that with one control.

This is just a simple example, of course, but it shows rather clearly the 
benefits associated with providing separate form controls for the components 
of structured data.  Another good example are separate inputs for the 
components of a date where you can, regardless of date format and component 
order in that date format, always be sure what is the date, the month, and 
the year *as meant by the user*; it is also one instance where client-side 
scripts can assist the user in entering data in several ways.


EOD here or, if you want to, F'up2 comp.infosystems.www.authoring.misc

PointedEars
-- 
    realism:    HTML 4.01 Strict
    evangelism: XHTML 1.0 Strict
    madness:    XHTML 1.1 as application/xhtml+xml
                                                    -- Bjoern Hoehrmann

[toc] | [prev] | [next] | [standalone]


#3886

FromThe Natural Philosopher <tnp@invalid.invalid>
Date2011-11-23 09:35 +0000
Message-ID<jaieoj$irv$1@news.albasani.net>
In reply to#3880
Thomas 'PointedEars' Lahn wrote:
> Peter H. Coffin wrote:
> 
>> On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote:
>>> Quite right. Is worse than you can possibly iagine at leats here in te
>>> UK, where addresses can be as little as 2 lines long or up to 6..
>>>
>>> So
>>>
>>> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
>>>
>>> Out in the sticks you might get
>>>
>>> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr
>>> Stonehouse, Gloucestershire GL13 6AH
>>>
>>>
>>> And if that comes at you without commas, god help you.
>>>
>>> I have spent DAYS taking name/address fields and parsing them *manually*
>>> into structured tables...
>> It is at this point that most people that have an actual need to solve
>> these kinds of problems turn to the available commercial software and
>> decide to solve it with money instead of manpower.
> 
> Where the question must be allowed: How came that the data has not been 
> requested and stored in a structured form to begin with?  That is, for 
> example, why only an address field in a form – why not a street, house
> number aso. field?  ISTM that we are seeing here an example of a mistake 
> made at the beginning which overall cost naturally grows larger and larger 
> as the project is nearing completion.
> 
IME this happens when you move from a crappy old database system to a 
properly designed one, and the data migration begins...


> 
> PointedEars

[toc] | [prev] | [next] | [standalone]


#3882

FromSandman <mr@sandman.net>
Date2011-11-23 09:55 +0100
Message-ID<mr-FF0348.09552223112011@News.Individual.NET>
In reply to#3871
In article <3004614.SPkdTlGXAF@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> >> > So I have this regexp:
> >> > 
> >> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> >> >     $streetname = uc_words($m[1]);
> >> >     $streetnumber = trim($m[2]);
> >> >     $streetletter = strtoupper($m[3]);
> >> >     $search = trim($streetname . SPACE . $streetnumber .
> >> > $streetletter);
> >> > }
> >> > 
> >> > The desired result is taki9ng the input ($search) and split it into
> >> > its parts as an address, right? $search can be, for example, "foo
> >> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
> >> 
> >> "10 East 42nd Street, New York, NY 10017, USA".
> > 
> > That wouldn't be a normal swedish address, no. :)
> 
> You had not limited the country or the language of your street addresses.

Well, to my defense, the subject line was "preg_match() and swedish 
characters" until I changed it. I hadn't changed it when I wrote my 
examples.

> My point is that parsing a street name and a house number from a street 
> address is a hard problem that cannot be solved only by applying one regular 
> expression.

Right, but your example is not a valid argument for that conclusion. 
My examples contained the variations of addresses that I wanted to 
match. Or are you saying that there is no way to use regular 
expressions to catch the examples I gave? Because I have a hard time 
believing that.



-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3890

From"Peter H. Coffin" <hellsop@ninehells.com>
Date2011-11-23 07:53 -0600
Message-ID<slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>
In reply to#3882
On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote:

> In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars'
> Lahn <PointedEars@web.de> wrote:
>
>> >> "10 East 42nd Street, New York, NY 10017, USA".
>> >
>> > That wouldn't be a normal swedish address, no. :)
>>
>> You had not limited the country or the language of your street
>> addresses.
>
> Well, to my defense, the subject line was "preg_match() and swedish
> characters" until I changed it. I hadn't changed it when I wrote my
> examples.
>
>> My point is that parsing a street name and a house number from a
>> street address is a hard problem that cannot be solved only by
>> applying one regular expression.
>
> Right, but your example is not a valid argument for that conclusion.
> My examples contained the variations of addresses that I wanted
> to match. Or are you saying that there is no way to use regular
> expressions to catch the examples I gave? Because I have a hard time
> believing that.

Address-matching is a hard task. I did that for a decade professionally
(as part of a job, not the sole function), and it's not easy to do well
for even one postal system, and trying to write a generalized one is
basically impossible to manage in one lifetime. The best *simple* way
to manage it is to take a field, blow it out into individual words,
standardize all the words you can find without trying to sort out
what they are (which is the Very Hard part of that task), throw the
alphabetic ones into soundex or nysiis, make a loose match by a chunk of
postal code or city code or province, then pick the item(s) that have
the greatest number of matches between incoming and loose-match record
of the numeric and nysiis-encoded alphabetical elements. If you weight
things like "numeric match = 1, plaintext that's in a dictionary that
matches when nysiis = 2, nondictionary text that matches nysiis = 3",
and do that for NAME as well as ADDRESS, you get about as good as you
can get without buying someone else's work. And that's STILL a lot of
effort to write. Regexp alone for address matching is a snipe-hunt. It
looks obviously right and you can spend a lot of time playing with it,
but it ends up being a dead end.

-- 
_  o
 |/)

[toc] | [prev] | [next] | [standalone]


#3893

FromSandman <mr@sandman.net>
Date2011-11-23 19:01 +0100
Message-ID<mr-F6EC15.19011423112011@News.Individual.NET>
In reply to#3890
In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>,
 "Peter H. Coffin" <hellsop@ninehells.com> wrote:

> On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote:
> 
> > In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars'
> > Lahn <PointedEars@web.de> wrote:
> >
> >> >> "10 East 42nd Street, New York, NY 10017, USA".
> >> >
> >> > That wouldn't be a normal swedish address, no. :)
> >>
> >> You had not limited the country or the language of your street
> >> addresses.
> >
> > Well, to my defense, the subject line was "preg_match() and swedish
> > characters" until I changed it. I hadn't changed it when I wrote my
> > examples.
> >
> >> My point is that parsing a street name and a house number from a
> >> street address is a hard problem that cannot be solved only by
> >> applying one regular expression.
> >
> > Right, but your example is not a valid argument for that conclusion.
> > My examples contained the variations of addresses that I wanted
> > to match. Or are you saying that there is no way to use regular
> > expressions to catch the examples I gave? Because I have a hard time
> > believing that.
> 
> Address-matching is a hard task. I did that for a decade professionally
> (as part of a job, not the sole function), and it's not easy to do well
> for even one postal system, and trying to write a generalized one is
> basically impossible to manage in one lifetime. The best *simple* way
> to manage it is to take a field, blow it out into individual words,
> standardize all the words you can find without trying to sort out
> what they are (which is the Very Hard part of that task), throw the
> alphabetic ones into soundex or nysiis, make a loose match by a chunk of
> postal code or city code or province, then pick the item(s) that have
> the greatest number of matches between incoming and loose-match record
> of the numeric and nysiis-encoded alphabetical elements. If you weight
> things like "numeric match = 1, plaintext that's in a dictionary that
> matches when nysiis = 2, nondictionary text that matches nysiis = 3",
> and do that for NAME as well as ADDRESS, you get about as good as you
> can get without buying someone else's work. And that's STILL a lot of
> effort to write. Regexp alone for address matching is a snipe-hunt. It
> looks obviously right and you can spend a lot of time playing with it,
> but it ends up being a dead end.

I thank you for your input, but I still maintain that my examples 
could be parsed by using a regular expression, and unless explicitly 
told so by using examples will I admit otherwise :-D

No offense, though.


-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3895

FromThe Natural Philosopher <tnp@invalid.invalid>
Date2011-11-23 18:54 +0000
Message-ID<jajfi0$pmj$3@news.albasani.net>
In reply to#3893
Sandman wrote:
> In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>,
>  "Peter H. Coffin" <hellsop@ninehells.com> wrote:
> 
>> On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote:
>>
>>> In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars'
>>> Lahn <PointedEars@web.de> wrote:
>>>
>>>>>> "10 East 42nd Street, New York, NY 10017, USA".
>>>>> That wouldn't be a normal swedish address, no. :)
>>>> You had not limited the country or the language of your street
>>>> addresses.
>>> Well, to my defense, the subject line was "preg_match() and swedish
>>> characters" until I changed it. I hadn't changed it when I wrote my
>>> examples.
>>>
>>>> My point is that parsing a street name and a house number from a
>>>> street address is a hard problem that cannot be solved only by
>>>> applying one regular expression.
>>> Right, but your example is not a valid argument for that conclusion.
>>> My examples contained the variations of addresses that I wanted
>>> to match. Or are you saying that there is no way to use regular
>>> expressions to catch the examples I gave? Because I have a hard time
>>> believing that.
>> Address-matching is a hard task. I did that for a decade professionally
>> (as part of a job, not the sole function), and it's not easy to do well
>> for even one postal system, and trying to write a generalized one is
>> basically impossible to manage in one lifetime. The best *simple* way
>> to manage it is to take a field, blow it out into individual words,
>> standardize all the words you can find without trying to sort out
>> what they are (which is the Very Hard part of that task), throw the
>> alphabetic ones into soundex or nysiis, make a loose match by a chunk of
>> postal code or city code or province, then pick the item(s) that have
>> the greatest number of matches between incoming and loose-match record
>> of the numeric and nysiis-encoded alphabetical elements. If you weight
>> things like "numeric match = 1, plaintext that's in a dictionary that
>> matches when nysiis = 2, nondictionary text that matches nysiis = 3",
>> and do that for NAME as well as ADDRESS, you get about as good as you
>> can get without buying someone else's work. And that's STILL a lot of
>> effort to write. Regexp alone for address matching is a snipe-hunt. It
>> looks obviously right and you can spend a lot of time playing with it,
>> but it ends up being a dead end.
> 
> I thank you for your input, but I still maintain that my examples 
> could be parsed by using a regular expression, and unless explicitly 
> told so by using examples will I admit otherwise :-D
> 
> No offense, though.
> 
> 
after three days, you could have done the data conversion by hand...

[toc] | [prev] | [next] | [standalone]


#3897

FromSandman <mr@sandman.net>
Date2011-11-23 20:23 +0100
Message-ID<mr-B941EF.20230523112011@News.Individual.NET>
In reply to#3895
In article <jajfi0$pmj$3@news.albasani.net>,
 The Natural Philosopher <tnp@invalid.invalid> wrote:

> >> Address-matching is a hard task. I did that for a decade professionally
> >> (as part of a job, not the sole function), and it's not easy to do well
> >> for even one postal system, and trying to write a generalized one is
> >> basically impossible to manage in one lifetime. The best *simple* way
> >> to manage it is to take a field, blow it out into individual words,
> >> standardize all the words you can find without trying to sort out
> >> what they are (which is the Very Hard part of that task), throw the
> >> alphabetic ones into soundex or nysiis, make a loose match by a chunk of
> >> postal code or city code or province, then pick the item(s) that have
> >> the greatest number of matches between incoming and loose-match record
> >> of the numeric and nysiis-encoded alphabetical elements. If you weight
> >> things like "numeric match = 1, plaintext that's in a dictionary that
> >> matches when nysiis = 2, nondictionary text that matches nysiis = 3",
> >> and do that for NAME as well as ADDRESS, you get about as good as you
> >> can get without buying someone else's work. And that's STILL a lot of
> >> effort to write. Regexp alone for address matching is a snipe-hunt. It
> >> looks obviously right and you can spend a lot of time playing with it,
> >> but it ends up being a dead end.
> > 
> > I thank you for your input, but I still maintain that my examples 
> > could be parsed by using a regular expression, and unless explicitly 
> > told so by using examples will I admit otherwise :-D
> > 
> > No offense, though.
>
> after three days, you could have done the data conversion by hand...

Huh? What data conversion? The data to be searched does not need to be 
converted into anything. It is neatly separated and also kept in a 
combined form. It's the *in-data* (i.e. the search terms provided by 
visitors to the sites). that I need to massage :)

And, three days wouldn't have gotten me far, the database contains 
over 600,000 posts of addresses :)

But, as I said, the database is neatly formatted. 



-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3898

From"Peter H. Coffin" <hellsop@ninehells.com>
Date2011-11-23 12:58 -0600
Message-ID<slrnjcqghr.85q.hellsop@nibelheim.ninehells.com>
In reply to#3893
On Wed, 23 Nov 2011 19:01:14 +0100, Sandman wrote:
> In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>,
>  "Peter H. Coffin" <hellsop@ninehells.com> wrote:
>
>> On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote:
>> 
>> > Right, but your example is not a valid argument for that conclusion.
>> > My examples contained the variations of addresses that I wanted
>> > to match. Or are you saying that there is no way to use regular
>> > expressions to catch the examples I gave? Because I have a hard time
>> > believing that.
>> 
>> Address-matching is a hard task. I did that for a decade professionally
>> (as part of a job, not the sole function), and it's not easy to do well
>> for even one postal system, and trying to write a generalized one is
>> basically impossible to manage in one lifetime. The best *simple* way
>> to manage it is to take a field, blow it out into individual words,
>> standardize all the words you can find without trying to sort out
>> what they are (which is the Very Hard part of that task), throw the
>> alphabetic ones into soundex or nysiis, make a loose match by a chunk of
>> postal code or city code or province, then pick the item(s) that have
>> the greatest number of matches between incoming and loose-match record
>> of the numeric and nysiis-encoded alphabetical elements. If you weight
>> things like "numeric match = 1, plaintext that's in a dictionary that
>> matches when nysiis = 2, nondictionary text that matches nysiis = 3",
>> and do that for NAME as well as ADDRESS, you get about as good as you
>> can get without buying someone else's work. And that's STILL a lot of
>> effort to write. Regexp alone for address matching is a snipe-hunt. It
>> looks obviously right and you can spend a lot of time playing with it,
>> but it ends up being a dead end.
>
> I thank you for your input, but I still maintain that my examples 
> could be parsed by using a regular expression, and unless explicitly 
> told so by using examples will I admit otherwise :-D

*grin* Any given (note: given) example set can be parsed with a
sufficiently complicated regexp. If your task is small enough and clean
enough, it might even not be THAT hard to accomplish. It's impossible to
provide advice about it, though, without having that complete example
set as well. The incoming data, however, is almost always going to
contain data that is not clean enough and will also probably end up
containing stuff that does not match your parsing rules, in a "because
fools are so ingenious" sense. 

And, at that point, you'll want to be looking at how you handle those
exeptions: reject, pass, send for clerical review, and what those
categories mean for your process. 

> No offense, though.

None to take. 

-- 
58. If it becomes necessary to escape, I will never stop to pose 
    dramatically and toss off a one-liner.
	--Peter Anspach's list of things to do as an Evil Overlord

[toc] | [prev] | [next] | [standalone]


#3904

FromSandman <mr@sandman.net>
Date2011-11-24 08:28 +0100
Message-ID<mr-DF03FA.08284024112011@News.Individual.NET>
In reply to#3898
In article <slrnjcqghr.85q.hellsop@nibelheim.ninehells.com>,
 "Peter H. Coffin" <hellsop@ninehells.com> wrote:

> > I thank you for your input, but I still maintain that my examples 
> > could be parsed by using a regular expression, and unless explicitly 
> > told so by using examples will I admit otherwise :-D
> 
> *grin* Any given (note: given) example set can be parsed with a
> sufficiently complicated regexp. If your task is small enough and clean
> enough, it might even not be THAT hard to accomplish. It's impossible to
> provide advice about it, though, without having that complete example
> set as well.

Huh? It was given in my OP. 

> The incoming data, however, is almost always going to
> contain data that is not clean enough and will also probably end up
> containing stuff that does not match your parsing rules, in a "because
> fools are so ingenious" sense.

Sure, but then there will be no matches of course. I'm trying to deal 
with the 99.9% of "correctly" formed search terms, though :)

> And, at that point, you'll want to be looking at how you handle those
> exeptions: reject, pass, send for clerical review, and what those
> categories mean for your process. 

Nah, if a user searchs for "34 Vikavägen B" they will get no hits, 
because no addresses in Sweden look like that. In such cases, we 
encourage the user to check their spelling of the adress and try 
again, of course :)

Basically, I have a number of examples. I want to parse them with a 
regexp. Someone here helped me solve why swedish characters broke it, 
but none yet have even tried to look at my regexp and suggest 
alterations to fit my examples.

Not that I could *expect* anything, it was just a comment on how 
common it is in groups like these for the readers to assume stupidity 
on the part of the OP and make edge cases (some more wild than others) 
where the proposed scenario would break.

I'm not stupid though, and I have a database of over 1 million 
historical search terms and looking through that I know *exactly* how 
and what people search for and now I'm looking to make sure that I can 
parse it better and with more certainity.

As a but of background, the DB that is to be searched has four 
relevant fields, one for street name, one for street number and one 
for street letter, and then one where the three are combined. It's 
this combined field I've so far used for loose string matching on the 
incoming search field, but there are still searches that won't match 
that should, so I need to parse and massage the incoming search terms.

Which brings us full circle :-D



-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3925 — SOLVED: Re: preg_match() oddities and question

FromSandman <mr@sandman.net>
Date2011-11-25 10:32 +0100
SubjectSOLVED: Re: preg_match() oddities and question
Message-ID<mr-4B661B.10320625112011@News.Individual.NET>
In reply to#3855
In article <mr-5B96D1.12212022112011@News.Individual.NET>,
 Sandman <mr@sandman.net> wrote:

> So I have this regexp:
> 
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>     $streetname = uc_words($m[1]);
>     $streetnumber = trim($m[2]);
>     $streetletter = strtoupper($m[3]);
>     $search = trim($streetname . SPACE . $streetnumber . 
> $streetletter);
> }

It's funny how asking for help on the net works. Being an "oldie" I 
usually turn to IRC first, mostly because it's the most direct medium, 
knowledgable people may be online right there and then.

My next option usually is USENET, because I like the format and it's a 
lot like mail.

But I think I need to change all of that thinking. The last few years 
I've gotten the most effective help from sites such as stackoverflow, 
really.

So with my problem above, I didn't find any regexp wizard online on 
IRC so I came here asking it, using examples. That was two days ago. I 
recived plenty of responses, some really helpful, some not so much. I 
wouldn't expect immediate help and salvation from CLP really, but it's 
not totally uncommon for people to at least wanting to help in some 
way.

Since most replies were about the in-data or the database being 
incorrect in the first place, instead of just focusing on the actual 
question asked, I turned to stackoverflow. I posted a question today 
at 10 am

At 10:19am, I got a clean cut, no frills solution to my actual problem:

preg_match("/([a-zA-Z ]+) ?([0-9]+)? ?([a-zA-Z]+)?/"...)

That captures *all* my examples in the OP, and formats it *exactly* 
like I wanted to.

I'm not trying to disrespect anyone here, but there is a bit too much 
"elitism" and "You're doing it wrong" mentality here, and have always 
been. And some times it's justified, when pure newbies come here and 
asks how to create a guestbook in php. But it seems this mentality 
spills over to posters like me that aren't newbies but still need help.

I've been in this group for about ten years:

<http://groups.google.com/groups/profile?show=more&enc_user=YhtCLA4AAAB
X2xnmqiyRi8RpNstwOyTw&group=comp.lang.php>

But I'm proficient enough in PHP that I don't post all that often 
about needing help so I'm not a "regular" here and probably easily 
mistaken for a newbie.

See this as a pleed to think beyond your own preconceptions about the 
proficiency of a poster, you know the entire "innocent until proven 
guilty" :)

Or, just ignore it altogether. I got my solution so I'm happy either 
way :)


-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3936 — Re: SOLVED: Re: preg_match() oddities and question

FromJerry Stuckle <jstucklex@attglobal.net>
Date2011-11-25 18:55 -0500
SubjectRe: SOLVED: Re: preg_match() oddities and question
Message-ID<jap9ti$onj$1@dont-email.me>
In reply to#3925
On 11/25/2011 4:32 AM, Sandman wrote:
> In article<mr-5B96D1.12212022112011@News.Individual.NET>,
>   Sandman<mr@sandman.net>  wrote:
>
>> So I have this regexp:
>>
>> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>>      $streetname = uc_words($m[1]);
>>      $streetnumber = trim($m[2]);
>>      $streetletter = strtoupper($m[3]);
>>      $search = trim($streetname . SPACE . $streetnumber .
>> $streetletter);
>> }
>
> It's funny how asking for help on the net works. Being an "oldie" I
> usually turn to IRC first, mostly because it's the most direct medium,
> knowledgable people may be online right there and then.
>
> My next option usually is USENET, because I like the format and it's a
> lot like mail.
>
> But I think I need to change all of that thinking. The last few years
> I've gotten the most effective help from sites such as stackoverflow,
> really.
>
> So with my problem above, I didn't find any regexp wizard online on
> IRC so I came here asking it, using examples. That was two days ago. I
> recived plenty of responses, some really helpful, some not so much. I
> wouldn't expect immediate help and salvation from CLP really, but it's
> not totally uncommon for people to at least wanting to help in some
> way.
>
> Since most replies were about the in-data or the database being
> incorrect in the first place, instead of just focusing on the actual
> question asked, I turned to stackoverflow. I posted a question today
> at 10 am
>
> At 10:19am, I got a clean cut, no frills solution to my actual problem:
>
> preg_match("/([a-zA-Z ]+) ?([0-9]+)? ?([a-zA-Z]+)?/"...)
>
> That captures *all* my examples in the OP, and formats it *exactly*
> like I wanted to.
>
> I'm not trying to disrespect anyone here, but there is a bit too much
> "elitism" and "You're doing it wrong" mentality here, and have always
> been. And some times it's justified, when pure newbies come here and
> asks how to create a guestbook in php. But it seems this mentality
> spills over to posters like me that aren't newbies but still need help.
>
> I've been in this group for about ten years:
>
> <http://groups.google.com/groups/profile?show=more&enc_user=YhtCLA4AAAB
> X2xnmqiyRi8RpNstwOyTw&group=comp.lang.php>
>
> But I'm proficient enough in PHP that I don't post all that often
> about needing help so I'm not a "regular" here and probably easily
> mistaken for a newbie.
>
> See this as a pleed to think beyond your own preconceptions about the
> proficiency of a poster, you know the entire "innocent until proven
> guilty" :)
>
> Or, just ignore it altogether. I got my solution so I'm happy either
> way :)
>
>

Look at who most of your answers were from - TNP and "Pointed Head".

Both known trolls in this newsgroup.

I would have tried to help, but I'm not that great in regex's myself. 
Normally I try to find other ways of solving my problem.  Usually I can 
do so.

Please don't condemn usenet because of a couple of well-known trolls.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
JDS Computer Training Corp.
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#3942 — Re: SOLVED: Re: preg_match() oddities and question

FromSandman <mr@sandman.net>
Date2011-11-26 11:21 +0100
SubjectRe: SOLVED: Re: preg_match() oddities and question
Message-ID<mr-A71241.11213026112011@News.Individual.NET>
In reply to#3936
In article <jap9ti$onj$1@dont-email.me>,
 Jerry Stuckle <jstucklex@attglobal.net> wrote:

> > Or, just ignore it altogether. I got my solution so I'm happy either
> > way :)
> 
> Look at who most of your answers were from - TNP and "Pointed Head".
>
> Both known trolls in this newsgroup.

One drawback of not being a regular is not knowing who is a troll or 
not. But the point still stands - I've yet to run into a troll on 
stackoverflow. 

> I would have tried to help, but I'm not that great in regex's myself. 
> Normally I try to find other ways of solving my problem.  Usually I can 
> do so.
> 
> Please don't condemn usenet because of a couple of well-known trolls.

It was not my intention to do that, and I apologize if that's how it 
could be interpreted. It was directed at those that did reply but were 
unable to help or unable to stay on topic. I have no problems 
accepting that these people were trolls. 



-- 
Sandman[.net]

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | comp.lang.php


csiph-web