Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.php > #3855 > unrolled thread
| Started by | Sandman <mr@sandman.net> |
|---|---|
| First post | 2011-11-22 12:21 +0100 |
| Last post | 2011-11-26 11:21 +0100 |
| Articles | 14 on this page of 34 — 7 participants |
Back to article view | Back to comp.lang.php
preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:21 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 11:26 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:36 +0100
Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-22 07:22 -0500
Re: preg_match() oddities and question tony@mountifield.org (Tony Mountifield) - 2011-11-22 11:47 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:12 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 13:30 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:55 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 17:56 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 17:30 +0000
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-22 17:20 -0600
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 23:59 +0000
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 01:59 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:58 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 22:02 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:20 +0100
Re: preg_match() oddities and question Denis McMahon <denismfmcmahon@gmail.com> - 2011-11-24 12:55 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:36 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-24 22:41 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:26 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 15:44 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 16:34 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 23:23 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 09:35 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:55 +0100
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 07:53 -0600
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 19:01 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 18:54 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 20:23 +0100
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 12:58 -0600
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:28 +0100
SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 10:32 +0100
Re: SOLVED: Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-25 18:55 -0500
Re: SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-26 11:21 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-25 15:44 +0100 |
| Message-ID | <1438797.UceJUlZ0hu@PointedEars.de> |
| In reply to | #3922 |
Sandman wrote:
> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> >> With separate controls they can be sure where to enter what;
>> >> it is accessible, and you have no problem processing the data. With
>> >> one control, neither applies.
>> > Well, I have been doing this for about ten years now,
>> I have been doing this for about fourteen years now. So what?
>
> You have monitored swedish address search terms for fourteen years? We
> should compare notes.
You are missing the point. The kind of structured data that you need to
enter in a form does not matter. Using cursor keys to move the text cursor
between delimiters in a running text always causes has more accessibility
and usability problems, and consequently information processing problems,
than tabbing (or otherwise moving the focus) from one control to another.
>> There are basic accessibility guidelines that no amount of
>> development experience can substitute (although studying usability,
>> as I did, can help). Many of which must be followed per
>> legislation in some countries.
>
> This.. has nothing to do with the topic at hand.
You are just not seeing how much it has to do with the topic at hand.
Changing the way people put in data towards one that is *actually* easier
for them solves, at least, three problems at once, including the one that
you have been asking about. Because when structured data is being entered
in a structured way, you have almost no trouble storing it in a structured
way. (BTW, in case you do not know, moving the focus from one input control
to the next while typing text can be assisted with client-side scripting.)
You should read, for example, <http://www.useit.com/> which contains many
ideas on how you can make Web sites better usable and therefore, almost in
passing, more accessible.
>> > When I say it's inconvenient for the end user, it's not something I
>> > make up on the spot to be obnoxious.
>> Nevertheless, your logic is flawed.
>
> Well, as long as you're merely saying that instead of actually, you
> know, substantiate that opinion, I have no idea what you expect me
> to do with it.
>
> Words are easy :)
This is about as much as I will discuss this here because your *actual*
problem has nothing to do with PHP, and little to do with Regular
Expressions.
HTH
PointedEars
--
var bugRiddenCrashPronePieceOfJunk = (
navigator.userAgent.indexOf('MSIE 5') != -1
&& navigator.userAgent.indexOf('Mac') != -1
) // Plone, register_function.js:16
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-25 16:34 +0100 |
| Message-ID | <mr-685156.16344225112011@News.Individual.NET> |
| In reply to | #3933 |
In article <1438797.UceJUlZ0hu@PointedEars.de>, Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: > > You have monitored swedish address search terms for fourteen years? We > > should compare notes. > > You are missing the point. I disagree. :) > The kind of structured data that you need to > enter in a form does not matter. Using cursor keys to move the text cursor > between delimiters in a running text always causes has more accessibility > and usability problems, and consequently information processing problems, > than tabbing (or otherwise moving the focus) from one control to another. Whatever gave you the idea that I am proposing a solution where the user would have to "use cursor keys to move the text cursor between delimiter in a running text"? I literally have no idea what you are talking about, and it has absolutely nothing to do with this thread. > >> There are basic accessibility guidelines that no amount of > >> development experience can substitute (although studying usability, > >> as I did, can help). Many of which must be followed per > >> legislation in some countries. > > > > This.. has nothing to do with the topic at hand. > > You are just not seeing how much it has to do with the topic at hand. And you seem to be failing to explain how it does :) > Changing the way people put in data towards one that is *actually* easier > for them solves, at least, three problems at once, including the one that > you have been asking about. Adding form fields to make it more cumbersome for them to search for their address, however, does not. My *experience* (i.e. not guessing, but rather - having done it exactly the way you propose) tells me that adding form fields to this situation adds more illegal search terms. Users tend to make more mistakes the more details you expect them to provide. Users often entered "X", "?" or "-" in the letter search field, or tried to type "none". Most of the time, however, they continued to type their entire address in the first field and then hit return. When you find yourself having to educate or expect the user to provide data in a specific way is when you fail as a software engineer. You have to make it as easy as possible for them to search for their address - and the thing you're dealing with here is *Google*. People know how to Google, they Google whatever shit they can and Google just always manages to figure out pretty much exactly what they need - with one input field. That's the level your visitors are on. They shouldn't have to read labels or instructions to search for their address. Also the meaning of the street letter may be different depending on what kind of property you live in and whether your'e a company or a private person. All that has to be explained for these users, and my experience (i.e. actual facts provided by years and years on monitoring exactly this) shows me that this is not the correct way to deal with input data. You are free to disagree all you want, and perhaps visitors to your applications and your digested search term analysis show you something else, but mine does not. When I wrote the OP and provided the examples it wasn't something out of the blue. > >> > When I say it's inconvenient for the end user, it's not something I > >> > make up on the spot to be obnoxious. > >> Nevertheless, your logic is flawed. > > > > Well, as long as you're merely saying that instead of actually, you > > know, substantiate that opinion, I have no idea what you expect me > > to do with it. > > > > Words are easy :) > > This is about as much as I will discuss this here because your *actual* > problem has nothing to do with PHP, and little to do with Regular > Expressions. If you saw my followup to my OP you may have seen that I found the solution elsewhere and that it indeed was solved using PHp and regular expressions - just as I knew it could be :) -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-25 23:23 +0100 |
| Message-ID | <7038554.eTKvZm4JqM@PointedEars.de> |
| In reply to | #3934 |
Sandman wrote:
> In article <1438797.UceJUlZ0hu@PointedEars.de>,
>> The kind of structured data that you need to enter in a form does not
>> matter. Using cursor keys to move the text cursor between delimiters in
>> a running text always causes has more accessibility and usability
>> problems, and consequently information processing problems, than tabbing
>> (or otherwise moving the focus) from one control to another.
>
> Whatever gave you the idea that I am proposing a solution where the
> user would have to "use cursor keys to move the text cursor between
> delimiter in a running text"? I literally have no idea what you are
> talking about, and it has absolutely nothing to do with this thread.
I am imagining a visitor of your Web site entering "foo street 34" (perhaps
as part of a longer text; your next question about phone numbers indicates
just that). Seeing that the "34" needed to be "42" instead, they would tab
back, upon which the entire text field is usually selected, and they have to
press the End key (or Alt+Arrow Right on Macs, IIRC) to get to the end.
Then they would perhaps press the Backspace key twice and type "42" Or
perhaps they needed it to be "134", so they would perhaps press the Arrow
Left or Backspace key twice (or Ctrl/Compose+Left) before they could make
their edit. (It would be comparably tedious with a pointing device, so do
not get me started on that.)
Now, it would be so much easier for them to work with the form, and easier
for you to process the data they entered, if you had the street name and the
house number in separate controls. Then they would tab back to the house
number control, which text would be selected, and they could immediately fix
their mistake. In addition, you could set up access keys and labels for the
controls with pure HTML so that users can use either way to focus the
control they want to edit. This would work with a graphical browser, a text
browser, a screen reader etc. Lately, it would work best with a mobile
device (anyone who has experienced first-hand how tedious it is to position
the text cursor with a mobile device, especially with a touchscreen, knows
what I mean). You cannot do that with one control.
This is just a simple example, of course, but it shows rather clearly the
benefits associated with providing separate form controls for the components
of structured data. Another good example are separate inputs for the
components of a date where you can, regardless of date format and component
order in that date format, always be sure what is the date, the month, and
the year *as meant by the user*; it is also one instance where client-side
scripts can assist the user in entering data in several ways.
EOD here or, if you want to, F'up2 comp.infosystems.www.authoring.misc
PointedEars
--
realism: HTML 4.01 Strict
evangelism: XHTML 1.0 Strict
madness: XHTML 1.1 as application/xhtml+xml
-- Bjoern Hoehrmann
[toc] | [prev] | [next] | [standalone]
| From | The Natural Philosopher <tnp@invalid.invalid> |
|---|---|
| Date | 2011-11-23 09:35 +0000 |
| Message-ID | <jaieoj$irv$1@news.albasani.net> |
| In reply to | #3880 |
Thomas 'PointedEars' Lahn wrote: > Peter H. Coffin wrote: > >> On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote: >>> Quite right. Is worse than you can possibly iagine at leats here in te >>> UK, where addresses can be as little as 2 lines long or up to 6.. >>> >>> So >>> >>> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address >>> >>> Out in the sticks you might get >>> >>> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr >>> Stonehouse, Gloucestershire GL13 6AH >>> >>> >>> And if that comes at you without commas, god help you. >>> >>> I have spent DAYS taking name/address fields and parsing them *manually* >>> into structured tables... >> It is at this point that most people that have an actual need to solve >> these kinds of problems turn to the available commercial software and >> decide to solve it with money instead of manpower. > > Where the question must be allowed: How came that the data has not been > requested and stored in a structured form to begin with? That is, for > example, why only an address field in a form – why not a street, house > number aso. field? ISTM that we are seeing here an example of a mistake > made at the beginning which overall cost naturally grows larger and larger > as the project is nearing completion. > IME this happens when you move from a crappy old database system to a properly designed one, and the data migration begins... > > PointedEars
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-23 09:55 +0100 |
| Message-ID | <mr-FF0348.09552223112011@News.Individual.NET> |
| In reply to | #3871 |
In article <3004614.SPkdTlGXAF@PointedEars.de>,
Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
> >> > So I have this regexp:
> >> >
> >> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> >> > $streetname = uc_words($m[1]);
> >> > $streetnumber = trim($m[2]);
> >> > $streetletter = strtoupper($m[3]);
> >> > $search = trim($streetname . SPACE . $streetnumber .
> >> > $streetletter);
> >> > }
> >> >
> >> > The desired result is taki9ng the input ($search) and split it into
> >> > its parts as an address, right? $search can be, for example, "foo
> >> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
> >>
> >> "10 East 42nd Street, New York, NY 10017, USA".
> >
> > That wouldn't be a normal swedish address, no. :)
>
> You had not limited the country or the language of your street addresses.
Well, to my defense, the subject line was "preg_match() and swedish
characters" until I changed it. I hadn't changed it when I wrote my
examples.
> My point is that parsing a street name and a house number from a street
> address is a hard problem that cannot be solved only by applying one regular
> expression.
Right, but your example is not a valid argument for that conclusion.
My examples contained the variations of addresses that I wanted to
match. Or are you saying that there is no way to use regular
expressions to catch the examples I gave? Because I have a hard time
believing that.
--
Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | "Peter H. Coffin" <hellsop@ninehells.com> |
|---|---|
| Date | 2011-11-23 07:53 -0600 |
| Message-ID | <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com> |
| In reply to | #3882 |
On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote: > In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars' > Lahn <PointedEars@web.de> wrote: > >> >> "10 East 42nd Street, New York, NY 10017, USA". >> > >> > That wouldn't be a normal swedish address, no. :) >> >> You had not limited the country or the language of your street >> addresses. > > Well, to my defense, the subject line was "preg_match() and swedish > characters" until I changed it. I hadn't changed it when I wrote my > examples. > >> My point is that parsing a street name and a house number from a >> street address is a hard problem that cannot be solved only by >> applying one regular expression. > > Right, but your example is not a valid argument for that conclusion. > My examples contained the variations of addresses that I wanted > to match. Or are you saying that there is no way to use regular > expressions to catch the examples I gave? Because I have a hard time > believing that. Address-matching is a hard task. I did that for a decade professionally (as part of a job, not the sole function), and it's not easy to do well for even one postal system, and trying to write a generalized one is basically impossible to manage in one lifetime. The best *simple* way to manage it is to take a field, blow it out into individual words, standardize all the words you can find without trying to sort out what they are (which is the Very Hard part of that task), throw the alphabetic ones into soundex or nysiis, make a loose match by a chunk of postal code or city code or province, then pick the item(s) that have the greatest number of matches between incoming and loose-match record of the numeric and nysiis-encoded alphabetical elements. If you weight things like "numeric match = 1, plaintext that's in a dictionary that matches when nysiis = 2, nondictionary text that matches nysiis = 3", and do that for NAME as well as ADDRESS, you get about as good as you can get without buying someone else's work. And that's STILL a lot of effort to write. Regexp alone for address matching is a snipe-hunt. It looks obviously right and you can spend a lot of time playing with it, but it ends up being a dead end. -- _ o |/)
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-23 19:01 +0100 |
| Message-ID | <mr-F6EC15.19011423112011@News.Individual.NET> |
| In reply to | #3890 |
In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>, "Peter H. Coffin" <hellsop@ninehells.com> wrote: > On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote: > > > In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars' > > Lahn <PointedEars@web.de> wrote: > > > >> >> "10 East 42nd Street, New York, NY 10017, USA". > >> > > >> > That wouldn't be a normal swedish address, no. :) > >> > >> You had not limited the country or the language of your street > >> addresses. > > > > Well, to my defense, the subject line was "preg_match() and swedish > > characters" until I changed it. I hadn't changed it when I wrote my > > examples. > > > >> My point is that parsing a street name and a house number from a > >> street address is a hard problem that cannot be solved only by > >> applying one regular expression. > > > > Right, but your example is not a valid argument for that conclusion. > > My examples contained the variations of addresses that I wanted > > to match. Or are you saying that there is no way to use regular > > expressions to catch the examples I gave? Because I have a hard time > > believing that. > > Address-matching is a hard task. I did that for a decade professionally > (as part of a job, not the sole function), and it's not easy to do well > for even one postal system, and trying to write a generalized one is > basically impossible to manage in one lifetime. The best *simple* way > to manage it is to take a field, blow it out into individual words, > standardize all the words you can find without trying to sort out > what they are (which is the Very Hard part of that task), throw the > alphabetic ones into soundex or nysiis, make a loose match by a chunk of > postal code or city code or province, then pick the item(s) that have > the greatest number of matches between incoming and loose-match record > of the numeric and nysiis-encoded alphabetical elements. If you weight > things like "numeric match = 1, plaintext that's in a dictionary that > matches when nysiis = 2, nondictionary text that matches nysiis = 3", > and do that for NAME as well as ADDRESS, you get about as good as you > can get without buying someone else's work. And that's STILL a lot of > effort to write. Regexp alone for address matching is a snipe-hunt. It > looks obviously right and you can spend a lot of time playing with it, > but it ends up being a dead end. I thank you for your input, but I still maintain that my examples could be parsed by using a regular expression, and unless explicitly told so by using examples will I admit otherwise :-D No offense, though. -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | The Natural Philosopher <tnp@invalid.invalid> |
|---|---|
| Date | 2011-11-23 18:54 +0000 |
| Message-ID | <jajfi0$pmj$3@news.albasani.net> |
| In reply to | #3893 |
Sandman wrote: > In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>, > "Peter H. Coffin" <hellsop@ninehells.com> wrote: > >> On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote: >> >>> In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars' >>> Lahn <PointedEars@web.de> wrote: >>> >>>>>> "10 East 42nd Street, New York, NY 10017, USA". >>>>> That wouldn't be a normal swedish address, no. :) >>>> You had not limited the country or the language of your street >>>> addresses. >>> Well, to my defense, the subject line was "preg_match() and swedish >>> characters" until I changed it. I hadn't changed it when I wrote my >>> examples. >>> >>>> My point is that parsing a street name and a house number from a >>>> street address is a hard problem that cannot be solved only by >>>> applying one regular expression. >>> Right, but your example is not a valid argument for that conclusion. >>> My examples contained the variations of addresses that I wanted >>> to match. Or are you saying that there is no way to use regular >>> expressions to catch the examples I gave? Because I have a hard time >>> believing that. >> Address-matching is a hard task. I did that for a decade professionally >> (as part of a job, not the sole function), and it's not easy to do well >> for even one postal system, and trying to write a generalized one is >> basically impossible to manage in one lifetime. The best *simple* way >> to manage it is to take a field, blow it out into individual words, >> standardize all the words you can find without trying to sort out >> what they are (which is the Very Hard part of that task), throw the >> alphabetic ones into soundex or nysiis, make a loose match by a chunk of >> postal code or city code or province, then pick the item(s) that have >> the greatest number of matches between incoming and loose-match record >> of the numeric and nysiis-encoded alphabetical elements. If you weight >> things like "numeric match = 1, plaintext that's in a dictionary that >> matches when nysiis = 2, nondictionary text that matches nysiis = 3", >> and do that for NAME as well as ADDRESS, you get about as good as you >> can get without buying someone else's work. And that's STILL a lot of >> effort to write. Regexp alone for address matching is a snipe-hunt. It >> looks obviously right and you can spend a lot of time playing with it, >> but it ends up being a dead end. > > I thank you for your input, but I still maintain that my examples > could be parsed by using a regular expression, and unless explicitly > told so by using examples will I admit otherwise :-D > > No offense, though. > > after three days, you could have done the data conversion by hand...
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-23 20:23 +0100 |
| Message-ID | <mr-B941EF.20230523112011@News.Individual.NET> |
| In reply to | #3895 |
In article <jajfi0$pmj$3@news.albasani.net>, The Natural Philosopher <tnp@invalid.invalid> wrote: > >> Address-matching is a hard task. I did that for a decade professionally > >> (as part of a job, not the sole function), and it's not easy to do well > >> for even one postal system, and trying to write a generalized one is > >> basically impossible to manage in one lifetime. The best *simple* way > >> to manage it is to take a field, blow it out into individual words, > >> standardize all the words you can find without trying to sort out > >> what they are (which is the Very Hard part of that task), throw the > >> alphabetic ones into soundex or nysiis, make a loose match by a chunk of > >> postal code or city code or province, then pick the item(s) that have > >> the greatest number of matches between incoming and loose-match record > >> of the numeric and nysiis-encoded alphabetical elements. If you weight > >> things like "numeric match = 1, plaintext that's in a dictionary that > >> matches when nysiis = 2, nondictionary text that matches nysiis = 3", > >> and do that for NAME as well as ADDRESS, you get about as good as you > >> can get without buying someone else's work. And that's STILL a lot of > >> effort to write. Regexp alone for address matching is a snipe-hunt. It > >> looks obviously right and you can spend a lot of time playing with it, > >> but it ends up being a dead end. > > > > I thank you for your input, but I still maintain that my examples > > could be parsed by using a regular expression, and unless explicitly > > told so by using examples will I admit otherwise :-D > > > > No offense, though. > > after three days, you could have done the data conversion by hand... Huh? What data conversion? The data to be searched does not need to be converted into anything. It is neatly separated and also kept in a combined form. It's the *in-data* (i.e. the search terms provided by visitors to the sites). that I need to massage :) And, three days wouldn't have gotten me far, the database contains over 600,000 posts of addresses :) But, as I said, the database is neatly formatted. -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | "Peter H. Coffin" <hellsop@ninehells.com> |
|---|---|
| Date | 2011-11-23 12:58 -0600 |
| Message-ID | <slrnjcqghr.85q.hellsop@nibelheim.ninehells.com> |
| In reply to | #3893 |
On Wed, 23 Nov 2011 19:01:14 +0100, Sandman wrote:
> In article <slrnjcpunb.85q.hellsop@nibelheim.ninehells.com>,
> "Peter H. Coffin" <hellsop@ninehells.com> wrote:
>
>> On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote:
>>
>> > Right, but your example is not a valid argument for that conclusion.
>> > My examples contained the variations of addresses that I wanted
>> > to match. Or are you saying that there is no way to use regular
>> > expressions to catch the examples I gave? Because I have a hard time
>> > believing that.
>>
>> Address-matching is a hard task. I did that for a decade professionally
>> (as part of a job, not the sole function), and it's not easy to do well
>> for even one postal system, and trying to write a generalized one is
>> basically impossible to manage in one lifetime. The best *simple* way
>> to manage it is to take a field, blow it out into individual words,
>> standardize all the words you can find without trying to sort out
>> what they are (which is the Very Hard part of that task), throw the
>> alphabetic ones into soundex or nysiis, make a loose match by a chunk of
>> postal code or city code or province, then pick the item(s) that have
>> the greatest number of matches between incoming and loose-match record
>> of the numeric and nysiis-encoded alphabetical elements. If you weight
>> things like "numeric match = 1, plaintext that's in a dictionary that
>> matches when nysiis = 2, nondictionary text that matches nysiis = 3",
>> and do that for NAME as well as ADDRESS, you get about as good as you
>> can get without buying someone else's work. And that's STILL a lot of
>> effort to write. Regexp alone for address matching is a snipe-hunt. It
>> looks obviously right and you can spend a lot of time playing with it,
>> but it ends up being a dead end.
>
> I thank you for your input, but I still maintain that my examples
> could be parsed by using a regular expression, and unless explicitly
> told so by using examples will I admit otherwise :-D
*grin* Any given (note: given) example set can be parsed with a
sufficiently complicated regexp. If your task is small enough and clean
enough, it might even not be THAT hard to accomplish. It's impossible to
provide advice about it, though, without having that complete example
set as well. The incoming data, however, is almost always going to
contain data that is not clean enough and will also probably end up
containing stuff that does not match your parsing rules, in a "because
fools are so ingenious" sense.
And, at that point, you'll want to be looking at how you handle those
exeptions: reject, pass, send for clerical review, and what those
categories mean for your process.
> No offense, though.
None to take.
--
58. If it becomes necessary to escape, I will never stop to pose
dramatically and toss off a one-liner.
--Peter Anspach's list of things to do as an Evil Overlord
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-24 08:28 +0100 |
| Message-ID | <mr-DF03FA.08284024112011@News.Individual.NET> |
| In reply to | #3898 |
In article <slrnjcqghr.85q.hellsop@nibelheim.ninehells.com>, "Peter H. Coffin" <hellsop@ninehells.com> wrote: > > I thank you for your input, but I still maintain that my examples > > could be parsed by using a regular expression, and unless explicitly > > told so by using examples will I admit otherwise :-D > > *grin* Any given (note: given) example set can be parsed with a > sufficiently complicated regexp. If your task is small enough and clean > enough, it might even not be THAT hard to accomplish. It's impossible to > provide advice about it, though, without having that complete example > set as well. Huh? It was given in my OP. > The incoming data, however, is almost always going to > contain data that is not clean enough and will also probably end up > containing stuff that does not match your parsing rules, in a "because > fools are so ingenious" sense. Sure, but then there will be no matches of course. I'm trying to deal with the 99.9% of "correctly" formed search terms, though :) > And, at that point, you'll want to be looking at how you handle those > exeptions: reject, pass, send for clerical review, and what those > categories mean for your process. Nah, if a user searchs for "34 Vikavägen B" they will get no hits, because no addresses in Sweden look like that. In such cases, we encourage the user to check their spelling of the adress and try again, of course :) Basically, I have a number of examples. I want to parse them with a regexp. Someone here helped me solve why swedish characters broke it, but none yet have even tried to look at my regexp and suggest alterations to fit my examples. Not that I could *expect* anything, it was just a comment on how common it is in groups like these for the readers to assume stupidity on the part of the OP and make edge cases (some more wild than others) where the proposed scenario would break. I'm not stupid though, and I have a database of over 1 million historical search terms and looking through that I know *exactly* how and what people search for and now I'm looking to make sure that I can parse it better and with more certainity. As a but of background, the DB that is to be searched has four relevant fields, one for street name, one for street number and one for street letter, and then one where the three are combined. It's this combined field I've so far used for loose string matching on the incoming search field, but there are still searches that won't match that should, so I need to parse and massage the incoming search terms. Which brings us full circle :-D -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-25 10:32 +0100 |
| Subject | SOLVED: Re: preg_match() oddities and question |
| Message-ID | <mr-4B661B.10320625112011@News.Individual.NET> |
| In reply to | #3855 |
In article <mr-5B96D1.12212022112011@News.Individual.NET>,
Sandman <mr@sandman.net> wrote:
> So I have this regexp:
>
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> $streetname = uc_words($m[1]);
> $streetnumber = trim($m[2]);
> $streetletter = strtoupper($m[3]);
> $search = trim($streetname . SPACE . $streetnumber .
> $streetletter);
> }
It's funny how asking for help on the net works. Being an "oldie" I
usually turn to IRC first, mostly because it's the most direct medium,
knowledgable people may be online right there and then.
My next option usually is USENET, because I like the format and it's a
lot like mail.
But I think I need to change all of that thinking. The last few years
I've gotten the most effective help from sites such as stackoverflow,
really.
So with my problem above, I didn't find any regexp wizard online on
IRC so I came here asking it, using examples. That was two days ago. I
recived plenty of responses, some really helpful, some not so much. I
wouldn't expect immediate help and salvation from CLP really, but it's
not totally uncommon for people to at least wanting to help in some
way.
Since most replies were about the in-data or the database being
incorrect in the first place, instead of just focusing on the actual
question asked, I turned to stackoverflow. I posted a question today
at 10 am
At 10:19am, I got a clean cut, no frills solution to my actual problem:
preg_match("/([a-zA-Z ]+) ?([0-9]+)? ?([a-zA-Z]+)?/"...)
That captures *all* my examples in the OP, and formats it *exactly*
like I wanted to.
I'm not trying to disrespect anyone here, but there is a bit too much
"elitism" and "You're doing it wrong" mentality here, and have always
been. And some times it's justified, when pure newbies come here and
asks how to create a guestbook in php. But it seems this mentality
spills over to posters like me that aren't newbies but still need help.
I've been in this group for about ten years:
<http://groups.google.com/groups/profile?show=more&enc_user=YhtCLA4AAAB
X2xnmqiyRi8RpNstwOyTw&group=comp.lang.php>
But I'm proficient enough in PHP that I don't post all that often
about needing help so I'm not a "regular" here and probably easily
mistaken for a newbie.
See this as a pleed to think beyond your own preconceptions about the
proficiency of a poster, you know the entire "innocent until proven
guilty" :)
Or, just ignore it altogether. I got my solution so I'm happy either
way :)
--
Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Jerry Stuckle <jstucklex@attglobal.net> |
|---|---|
| Date | 2011-11-25 18:55 -0500 |
| Subject | Re: SOLVED: Re: preg_match() oddities and question |
| Message-ID | <jap9ti$onj$1@dont-email.me> |
| In reply to | #3925 |
On 11/25/2011 4:32 AM, Sandman wrote:
> In article<mr-5B96D1.12212022112011@News.Individual.NET>,
> Sandman<mr@sandman.net> wrote:
>
>> So I have this regexp:
>>
>> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>> $streetname = uc_words($m[1]);
>> $streetnumber = trim($m[2]);
>> $streetletter = strtoupper($m[3]);
>> $search = trim($streetname . SPACE . $streetnumber .
>> $streetletter);
>> }
>
> It's funny how asking for help on the net works. Being an "oldie" I
> usually turn to IRC first, mostly because it's the most direct medium,
> knowledgable people may be online right there and then.
>
> My next option usually is USENET, because I like the format and it's a
> lot like mail.
>
> But I think I need to change all of that thinking. The last few years
> I've gotten the most effective help from sites such as stackoverflow,
> really.
>
> So with my problem above, I didn't find any regexp wizard online on
> IRC so I came here asking it, using examples. That was two days ago. I
> recived plenty of responses, some really helpful, some not so much. I
> wouldn't expect immediate help and salvation from CLP really, but it's
> not totally uncommon for people to at least wanting to help in some
> way.
>
> Since most replies were about the in-data or the database being
> incorrect in the first place, instead of just focusing on the actual
> question asked, I turned to stackoverflow. I posted a question today
> at 10 am
>
> At 10:19am, I got a clean cut, no frills solution to my actual problem:
>
> preg_match("/([a-zA-Z ]+) ?([0-9]+)? ?([a-zA-Z]+)?/"...)
>
> That captures *all* my examples in the OP, and formats it *exactly*
> like I wanted to.
>
> I'm not trying to disrespect anyone here, but there is a bit too much
> "elitism" and "You're doing it wrong" mentality here, and have always
> been. And some times it's justified, when pure newbies come here and
> asks how to create a guestbook in php. But it seems this mentality
> spills over to posters like me that aren't newbies but still need help.
>
> I've been in this group for about ten years:
>
> <http://groups.google.com/groups/profile?show=more&enc_user=YhtCLA4AAAB
> X2xnmqiyRi8RpNstwOyTw&group=comp.lang.php>
>
> But I'm proficient enough in PHP that I don't post all that often
> about needing help so I'm not a "regular" here and probably easily
> mistaken for a newbie.
>
> See this as a pleed to think beyond your own preconceptions about the
> proficiency of a poster, you know the entire "innocent until proven
> guilty" :)
>
> Or, just ignore it altogether. I got my solution so I'm happy either
> way :)
>
>
Look at who most of your answers were from - TNP and "Pointed Head".
Both known trolls in this newsgroup.
I would have tried to help, but I'm not that great in regex's myself.
Normally I try to find other ways of solving my problem. Usually I can
do so.
Please don't condemn usenet because of a couple of well-known trolls.
--
==================
Remove the "x" from my email address
Jerry Stuckle
JDS Computer Training Corp.
jstucklex@attglobal.net
==================
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-26 11:21 +0100 |
| Subject | Re: SOLVED: Re: preg_match() oddities and question |
| Message-ID | <mr-A71241.11213026112011@News.Individual.NET> |
| In reply to | #3936 |
In article <jap9ti$onj$1@dont-email.me>, Jerry Stuckle <jstucklex@attglobal.net> wrote: > > Or, just ignore it altogether. I got my solution so I'm happy either > > way :) > > Look at who most of your answers were from - TNP and "Pointed Head". > > Both known trolls in this newsgroup. One drawback of not being a regular is not knowing who is a troll or not. But the point still stands - I've yet to run into a troll on stackoverflow. > I would have tried to help, but I'm not that great in regex's myself. > Normally I try to find other ways of solving my problem. Usually I can > do so. > > Please don't condemn usenet because of a couple of well-known trolls. It was not my intention to do that, and I apologize if that's how it could be interpreted. It was directed at those that did reply but were unable to help or unable to stay on topic. I have no problems accepting that these people were trolls. -- Sandman[.net]
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | comp.lang.php
csiph-web