Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.php > #3855 > unrolled thread
| Started by | Sandman <mr@sandman.net> |
|---|---|
| First post | 2011-11-22 12:21 +0100 |
| Last post | 2011-11-26 11:21 +0100 |
| Articles | 20 on this page of 34 — 7 participants |
Back to article view | Back to comp.lang.php
preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:21 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 11:26 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:36 +0100
Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-22 07:22 -0500
Re: preg_match() oddities and question tony@mountifield.org (Tony Mountifield) - 2011-11-22 11:47 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:12 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 13:30 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:55 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 17:56 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 17:30 +0000
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-22 17:20 -0600
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 23:59 +0000
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 01:59 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:58 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 22:02 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:20 +0100
Re: preg_match() oddities and question Denis McMahon <denismfmcmahon@gmail.com> - 2011-11-24 12:55 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:36 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-24 22:41 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:26 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 15:44 +0100
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 16:34 +0100
Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 23:23 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 09:35 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:55 +0100
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 07:53 -0600
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 19:01 +0100
Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 18:54 +0000
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 20:23 +0100
Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 12:58 -0600
Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:28 +0100
SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 10:32 +0100
Re: SOLVED: Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-25 18:55 -0500
Re: SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-26 11:21 +0100
Page 1 of 2 [1] 2 Next page →
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-22 12:21 +0100 |
| Subject | preg_match() oddities and question |
| Message-ID | <mr-5B96D1.12212022112011@News.Individual.NET> |
So I have this regexp:
if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
$streetname = uc_words($m[1]);
$streetnumber = trim($m[2]);
$streetletter = strtoupper($m[3]);
$search = trim($streetname . SPACE . $streetnumber .
$streetletter);
}
The desired result is taki9ng the input ($search) and split it into
its parts as an address, right? $search can be, for example, "foo
street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
So, if I print_r($m) with different input I get:
Array
(
[0] => foo street 34
[1] => foo street
[2] => 34
[3] =>
)
Array
(
[0] => longstreet 45b
[1] => longstreet
[2] => 45
[3] => b
)
Array
(
[0] => longstreet 45 b
[1] => longstreet
[2] => 45
[3] => b
)
You get the idea. But problems arise when I search for the streetname
alone:
Array
(
[0] => longstreet
[1] =>
[2] =>
[3] => longstreet
)
As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire
search term since there are no digits and the first group is
non-greedy. And if I make the first group greedy, "longstreet" is
matched correctly, but it also catches the entire "longstreet 45b"
when searching for that.
Also, when searching for a term in swedish characters, I get this:
Array
(
[0] => vikavägen
[1] => vikavä
[2] =>
[3] => gen
)
Which is quite odd to me, why isn't "vikavägen" matched the same
(undesired) way that "oongstreet". I have tried the /u modifier, and
made sure that it was utf8-encoded, but it didn't make a difference
(incoming encoding is ISO 8859-1).
Why the difference, and how do I correctly parse out parts as needed?
Any help is appreciated.
--
Sandman[.net]
[toc] | [next] | [standalone]
| From | The Natural Philosopher <tnp@invalid.invalid> |
|---|---|
| Date | 2011-11-22 11:26 +0000 |
| Message-ID | <jag0sr$49h$1@news.albasani.net> |
| In reply to | #3855 |
Sandman wrote: > > Why the difference, and how do I correctly parse out parts as needed? > I have always found establishing the correct regexp expression to take longer than writing my own filters in whatever language I happened to be using.... Life is too short for regexps. > Any help is appreciated. > > > > >
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-22 12:36 +0100 |
| Message-ID | <mr-7CF721.12361222112011@News.Individual.NET> |
| In reply to | #3856 |
In article <jag0sr$49h$1@news.albasani.net>, The Natural Philosopher <tnp@invalid.invalid> wrote: > > Why the difference, and how do I correctly parse out parts as needed? > > I have always found establishing the correct regexp expression to take > longer than writing my own filters in whatever language I happened to > be using.... > > Life is too short for regexps. I've never had much problem (time-wise) with regexps. I'm just stumped about the difference in execution of this one, and need a little help figuring out the syntax for the other part. In short, regexps are rarely a problem for me, and I don't know how I would solve my situation without using one, if you have any suggestions, please share :) -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Jerry Stuckle <jstucklex@attglobal.net> |
|---|---|
| Date | 2011-11-22 07:22 -0500 |
| Message-ID | <jag45m$5um$3@dont-email.me> |
| In reply to | #3856 |
On 11/22/2011 6:26 AM, The Natural Philosopher wrote: > Sandman wrote: > >> >> Why the difference, and how do I correctly parse out parts as needed? >> > > I have always found establishing the correct regexp expression to take > longer than writing my own filters in whatever language I happened to be > using.... > > > Life is too short for regexps. > That's because regex's take intelligence - which you don't have. -- ================== Remove the "x" from my email address Jerry Stuckle JDS Computer Training Corp. jstucklex@attglobal.net ==================
[toc] | [prev] | [next] | [standalone]
| From | tony@mountifield.org (Tony Mountifield) |
|---|---|
| Date | 2011-11-22 11:47 +0000 |
| Message-ID | <jag24m$4nj$1@softins.clara.co.uk> |
| In reply to | #3855 |
In article <mr-5B96D1.12212022112011@News.Individual.NET>,
Sandman <mr@sandman.net> wrote:
> So I have this regexp:
>
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
You don't need the commas in the character class, unless you want to
match a literal comma, in which case you only need it once.
> $streetname = uc_words($m[1]);
> $streetnumber = trim($m[2]);
> $streetletter = strtoupper($m[3]);
> $search = trim($streetname . SPACE . $streetnumber .
> $streetletter);
> }
>
> The desired result is taki9ng the input ($search) and split it into
> its parts as an address, right? $search can be, for example, "foo
> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
What about "foo street"? (i.e. with a space, but no number)
> So, if I print_r($m) with different input I get:
>
> Array
> (
> [0] => foo street 34
> [1] => foo street
> [2] => 34
> [3] =>
> )
> Array
> (
> [0] => longstreet 45b
> [1] => longstreet
> [2] => 45
> [3] => b
> )
> Array
> (
> [0] => longstreet 45 b
> [1] => longstreet
> [2] => 45
> [3] => b
> )
>
> You get the idea. But problems arise when I search for the streetname
> alone:
>
> Array
> (
> [0] => longstreet
> [1] =>
> [2] =>
> [3] => longstreet
> )
And you would also get:
Array
(
[0] => foo street
[1] => foo
[2] =>
[3] => street
)
> As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire
> search term since there are no digits and the first group is
> non-greedy. And if I make the first group greedy, "longstreet" is
> matched correctly, but it also catches the entire "longstreet 45b"
> when searching for that.
Yes, you need to define your rules more closely. Not at the regex level,
but actually at the logic/decision level. If you can make rules that
can unambiguously specify how all kinds of input should be parsed,
then you can look at how to represent that in regexes. You might need
some additional logic to operate on the parsed result.
> Also, when searching for a term in swedish characters, I get this:
>
> Array
> (
> [0] => vikavägen
> [1] => vikavä
> [2] =>
> [3] => gen
> )
>
> Which is quite odd to me, why isn't "vikavägen" matched the same
> (undesired) way that "oongstreet". I have tried the /u modifier, and
> made sure that it was utf8-encoded, but it didn't make a difference
> (incoming encoding is ISO 8859-1).
>
> Why the difference, and how do I correctly parse out parts as needed?
That's because ä is not in the set A-Za-z. If you want a character class
that properly recognises locale-specific letters, you need to change your
character class above to this:
[[:alpha:]\-]
Hope this helps!
Tony
--
Tony Mountifield
Work: tony@softins.co.uk - http://www.softins.co.uk
Play: tony@mountifield.org - http://tony.mountifield.org
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-22 13:12 +0100 |
| Message-ID | <mr-C40BF0.13123922112011@News.Individual.NET> |
| In reply to | #3858 |
In article <jag24m$4nj$1@softins.clara.co.uk>,
tony@mountifield.org (Tony Mountifield) wrote:
> > So I have this regexp:
> >
> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>
> You don't need the commas in the character class, unless you want to
> match a literal comma, in which case you only need it once.
Right, thanks :)
> > $streetname = uc_words($m[1]);
> > $streetnumber = trim($m[2]);
> > $streetletter = strtoupper($m[3]);
> > $search = trim($streetname . SPACE . $streetnumber .
> > $streetletter);
> > }
> >
> > The desired result is taki9ng the input ($search) and split it into
> > its parts as an address, right? $search can be, for example, "foo
> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>
> What about "foo street"? (i.e. with a space, but no number)
Exactly, that gets this:
Array
(
[0] => foo street
[1] => foo
[2] =>
[3] => street
)
Which is incorrect. IN fact, the last group SHOULD be defined as
([A-Za-z]{0,1}) but that still messes it up like:
Array
(
[0] => foo street
[1] => foo stree
[2] =>
[3] => t
)
So I've tried variations for that as well.
<snip>
> And you would also get:
>
> Array
> (
> [0] => foo street
> [1] => foo
> [2] =>
> [3] => street
> )
>
> > As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire
> > search term since there are no digits and the first group is
> > non-greedy. And if I make the first group greedy, "longstreet" is
> > matched correctly, but it also catches the entire "longstreet 45b"
> > when searching for that.
>
> Yes, you need to define your rules more closely. Not at the regex level,
> but actually at the logic/decision level. If you can make rules that
> can unambiguously specify how all kinds of input should be parsed,
> then you can look at how to represent that in regexes. You might need
> some additional logic to operate on the parsed result.
What you're basically suggesting is a series of regexp to find out
what "style" an adress is given in, and then parse out the parts?
Because I'm not sure how I would be able to do it without a series if
if/else preg_match():es?
> > Also, when searching for a term in swedish characters, I get this:
> >
> > Array
> > (
> > [0] => vikavÀgen
> > [1] => vikavÀ
> > [2] =>
> > [3] => gen
> > )
> >
> > Which is quite odd to me, why isn't "vikavÀgen" matched the same
> > (undesired) way that "oongstreet". I have tried the /u modifier, and
> > made sure that it was utf8-encoded, but it didn't make a difference
> > (incoming encoding is ISO 8859-1).
> >
> > Why the difference, and how do I correctly parse out parts as needed?
>
> That's because À is not in the set A-Za-z. If you want a character class
> that properly recognises locale-specific letters, you need to change your
> character class above to this:
>
> [[:alpha:]\-]
>
> Hope this helps!
That explains the difference, thank you very much for that. Now I
still need to figure out a global parse routine or criteria for
parsing out the address parts...
--
Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-22 13:30 +0100 |
| Message-ID | <1670168.aK4W3vaeNJ@PointedEars.de> |
| In reply to | #3855 |
Sandman wrote:
> So I have this regexp:
>
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> $streetname = uc_words($m[1]);
> $streetnumber = trim($m[2]);
> $streetletter = strtoupper($m[3]);
> $search = trim($streetname . SPACE . $streetnumber .
> $streetletter);
> }
>
> The desired result is taki9ng the input ($search) and split it into
> its parts as an address, right? $search can be, for example, "foo
> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
"10 East 42nd Street, New York, NY 10017, USA".
PointedEars
--
When all you know is jQuery, every problem looks $(olvable).
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-22 13:55 +0100 |
| Message-ID | <mr-52D11B.13550322112011@News.Individual.NET> |
| In reply to | #3863 |
In article <1670168.aK4W3vaeNJ@PointedEars.de>,
Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
> Sandman wrote:
>
> > So I have this regexp:
> >
> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> > $streetname = uc_words($m[1]);
> > $streetnumber = trim($m[2]);
> > $streetletter = strtoupper($m[3]);
> > $search = trim($streetname . SPACE . $streetnumber .
> > $streetletter);
> > }
> >
> > The desired result is taki9ng the input ($search) and split it into
> > its parts as an address, right? $search can be, for example, "foo
> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>
> "10 East 42nd Street, New York, NY 10017, USA".
That wouldn't be a normal swedish address, no. :)
--
Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-22 17:56 +0100 |
| Message-ID | <3004614.SPkdTlGXAF@PointedEars.de> |
| In reply to | #3865 |
Sandman wrote:
> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> Sandman wrote:
>> > So I have this regexp:
>> >
>> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>> > $streetname = uc_words($m[1]);
>> > $streetnumber = trim($m[2]);
>> > $streetletter = strtoupper($m[3]);
>> > $search = trim($streetname . SPACE . $streetnumber .
>> > $streetletter);
>> > }
>> >
>> > The desired result is taki9ng the input ($search) and split it into
>> > its parts as an address, right? $search can be, for example, "foo
>> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>>
>> "10 East 42nd Street, New York, NY 10017, USA".
>
> That wouldn't be a normal swedish address, no. :)
You had not limited the country or the language of your street addresses.
My point is that parsing a street name and a house number from a street
address is a hard problem that cannot be solved only by applying one regular
expression.
PointedEars
--
realism: HTML 4.01 Strict
evangelism: XHTML 1.0 Strict
madness: XHTML 1.1 as application/xhtml+xml
-- Bjoern Hoehrmann
[toc] | [prev] | [next] | [standalone]
| From | The Natural Philosopher <tnp@invalid.invalid> |
|---|---|
| Date | 2011-11-22 17:30 +0000 |
| Message-ID | <jagm86$kus$1@news.albasani.net> |
| In reply to | #3871 |
Thomas 'PointedEars' Lahn wrote:
> Sandman wrote:
>
>> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>>> Sandman wrote:
>>>> So I have this regexp:
>>>>
>>>> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>>>> $streetname = uc_words($m[1]);
>>>> $streetnumber = trim($m[2]);
>>>> $streetletter = strtoupper($m[3]);
>>>> $search = trim($streetname . SPACE . $streetnumber .
>>>> $streetletter);
>>>> }
>>>>
>>>> The desired result is taki9ng the input ($search) and split it into
>>>> its parts as an address, right? $search can be, for example, "foo
>>>> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>>> "10 East 42nd Street, New York, NY 10017, USA".
>> That wouldn't be a normal swedish address, no. :)
>
> You had not limited the country or the language of your street addresses.
>
> My point is that parsing a street name and a house number from a street
> address is a hard problem that cannot be solved only by applying one regular
> expression.
>
Quite right. Is worse than you can possibly iagine at leats here in te
UK, where addresses can be as little as 2 lines long or up to 6..
So
10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
Out in the sticks you might get
Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr
Stonehouse, Gloucestershire GL13 6AH
And if that comes at you without commas, god help you.
I have spent DAYS taking name/address fields and parsing them *manually*
into structured tables...
>
> PointedEars
[toc] | [prev] | [next] | [standalone]
| From | "Peter H. Coffin" <hellsop@ninehells.com> |
|---|---|
| Date | 2011-11-22 17:20 -0600 |
| Message-ID | <slrnjcobhs.85q.hellsop@nibelheim.ninehells.com> |
| In reply to | #3873 |
On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote:
> Quite right. Is worse than you can possibly iagine at leats here in te
> UK, where addresses can be as little as 2 lines long or up to 6..
>
> So
>
> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
>
> Out in the sticks you might get
>
> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr
> Stonehouse, Gloucestershire GL13 6AH
>
>
> And if that comes at you without commas, god help you.
>
> I have spent DAYS taking name/address fields and parsing them *manually*
> into structured tables...
It is at this point that most people that have an actual need to solve
these kinds of problems turn to the available commercial software and
decide to solve it with money instead of manpower.
--
They got rid of it because they judged it more trouble than it was
worth. (And considering they'd gone to great lengths to minimize its
worth, I suppose they were right.)
-- J. D. Baldwin
[toc] | [prev] | [next] | [standalone]
| From | The Natural Philosopher <tnp@invalid.invalid> |
|---|---|
| Date | 2011-11-22 23:59 +0000 |
| Message-ID | <jahd15$666$3@news.albasani.net> |
| In reply to | #3878 |
Peter H. Coffin wrote: > On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote: >> Quite right. Is worse than you can possibly iagine at leats here in te >> UK, where addresses can be as little as 2 lines long or up to 6.. >> >> So >> >> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address >> >> Out in the sticks you might get >> >> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr >> Stonehouse, Gloucestershire GL13 6AH >> >> >> And if that comes at you without commas, god help you. >> >> I have spent DAYS taking name/address fields and parsing them *manually* >> into structured tables... > > It is at this point that most people that have an actual need to solve > these kinds of problems turn to the available commercial software and > decide to solve it with money instead of manpower. > there is no AI that can match a human brain in decoding human idiocy...yet
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-23 01:59 +0100 |
| Message-ID | <1780410.D2KPVSPYKU@PointedEars.de> |
| In reply to | #3878 |
Peter H. Coffin wrote: > On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote: >> Quite right. Is worse than you can possibly iagine at leats here in te >> UK, where addresses can be as little as 2 lines long or up to 6.. >> >> So >> >> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address >> >> Out in the sticks you might get >> >> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr >> Stonehouse, Gloucestershire GL13 6AH >> >> >> And if that comes at you without commas, god help you. >> >> I have spent DAYS taking name/address fields and parsing them *manually* >> into structured tables... > > It is at this point that most people that have an actual need to solve > these kinds of problems turn to the available commercial software and > decide to solve it with money instead of manpower. Where the question must be allowed: How came that the data has not been requested and stored in a structured form to begin with? That is, for example, why only an address field in a form – why not a street, house number aso. field? ISTM that we are seeing here an example of a mistake made at the beginning which overall cost naturally grows larger and larger as the project is nearing completion. PointedEars -- Danny Goodman's books are out of date and teach practices that are positively harmful for cross-browser scripting. -- Richard Cornford, cljs, <cife6q$253$1$8300dec7@news.demon.co.uk> (2004)
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-23 09:58 +0100 |
| Message-ID | <mr-59B44E.09581623112011@News.Individual.NET> |
| In reply to | #3880 |
In article <1780410.D2KPVSPYKU@PointedEars.de>, Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: > > It is at this point that most people that have an actual need to solve > > these kinds of problems turn to the available commercial software and > > decide to solve it with money instead of manpower. > > Where the question must be allowed: How came that the data has not been > requested and stored in a structured form to begin with? That is, for > example, why only an address field in a form – why not a street, house > number aso. field? Convenience for the user, of course. This is a form that says "Are you connected to the citynet?" and then you just enter your address to search the database. If the user has to provide street name, street number and street letter in separate fields, it's inconvenient for them. -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-23 22:02 +0100 |
| Message-ID | <4778042.ypaU67uLZW@PointedEars.de> |
| In reply to | #3883 |
Sandman wrote: > Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: >> > It is at this point that most people that have an actual need to solve >> > these kinds of problems turn to the available commercial software and >> > decide to solve it with money instead of manpower. >> >> Where the question must be allowed: How came that the data has not been >> requested and stored in a structured form to begin with? That is, for >> example, why only an address field in a form – why not a street, house >> number aso. field? > > Convenience for the user, of course. You can't be serious. > This is a form that says "Are you connected to the citynet?" and then > you just enter your address to search the database. If the user has to > provide street name, street number and street letter in separate > fields, it's inconvenient for them. No, it's not. With separate controls they can be sure where to enter what; it is accessible, and you have no problem processing the data. With one control, neither applies. PointedEars -- Use any version of Microsoft Frontpage to create your site. (This won't prevent people from viewing your source, but no one will want to steal it.) -- from <http://www.vortex-webdesign.com/help/hidesource.htm> (404-comp.)
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-24 08:20 +0100 |
| Message-ID | <mr-8BFE1F.08204424112011@News.Individual.NET> |
| In reply to | #3900 |
In article <4778042.ypaU67uLZW@PointedEars.de>, Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: > >> > It is at this point that most people that have an actual need to solve > >> > these kinds of problems turn to the available commercial software and > >> > decide to solve it with money instead of manpower. > >> > >> Where the question must be allowed: How came that the data has not been > >> requested and stored in a structured form to begin with? That is, for > >> example, why only an address field in a form – why not a street, house > >> number aso. field? > > > > Convenience for the user, of course. > > You can't be serious. I can :) > > This is a form that says "Are you connected to the citynet?" and then > > you just enter your address to search the database. If the user has to > > provide street name, street number and street letter in separate > > fields, it's inconvenient for them. > > No, it's not. Actually, yes it is :) > With separate controls they can be sure where to enter what; > it is accessible, and you have no problem processing the data. With one > control, neither applies. Well, I have been doing this for about ten years now, and recived tons of feedback from my clients on things like this. When I say it's inconvenient for the end user, it's not something I make up on the spot to be obnoxious. Just like with my examples, I have a pretty clear picture of what problem I need to solve. I find it curious that no one in CLP even attempt to look at that, and instead trying to find other examples, or claiming that the frontend should be changed. Makes me think that you guys deem the examples I gave as unsolvable, which of course I refuse to agree with. :) No offense though. -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Denis McMahon <denismfmcmahon@gmail.com> |
|---|---|
| Date | 2011-11-24 12:55 +0000 |
| Message-ID | <4ece3eb7$0$28723$a8266bb1@newsreader.readnews.com> |
| In reply to | #3903 |
On Thu, 24 Nov 2011 08:20:45 +0100, Sandman wrote:
> Well, I have been doing this for about ten years now, and recived tons
> of feedback from my clients on things like this. When I say it's
> inconvenient for the end user, it's not something I make up on the spot
> to be obnoxious.
It's a point that a lot of website UI designers overlook - what we think
is the best and most obvious layout, and even "the way everyone does it",
isn't always what the website user finds most user friendly.
As to the OPs problem, perhaps it needs a different approach:
Presumably for any given city you can obtain a list of valid streets?
Why not:
foreach ($street_name in $streets_of_city) {
if (($street_start = strpos($address, $street_name)) !== false) {
foreach ($district_name in $districts_of_city) {
if (($district_start = strpos($address, $district_name,
$street_start)) !== false) {
break;
}
}
break;
}
}
You may need to allow for "district creep" and will need to allow for
streets crossing district boundaries.
Another option, and one that is becoming popular in some areas, might be
to allow entry of the postcode, and then generate a drop down list,
perhaps populated using an XMLHttpRequest, of addresses matching the
postcode - this would require access to a copy of the relevant postcode
database which might involve some cost, and of course you need to
implement a mechanism for handling updates to that database.
Possibly a similar sort of approach, based on a select menu for district,
then a select menu for street names, and then a text entry field for
"apartment number, building name and / or number, or house number as
appropriate".
I don't think you're ever going to parse a single line address entry with
regex, because there are too many different formats that might get thrown
at you. You either need to get the data in a different format, or find a
different method of processing the data that you are getting.
Rgds
Denis McMahon
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-25 09:36 +0100 |
| Message-ID | <mr-1BAC45.09360625112011@News.Individual.NET> |
| In reply to | #3916 |
In article <4ece3eb7$0$28723$a8266bb1@newsreader.readnews.com>,
Denis McMahon <denismfmcmahon@gmail.com> wrote:
> > Well, I have been doing this for about ten years now, and recived tons
> > of feedback from my clients on things like this. When I say it's
> > inconvenient for the end user, it's not something I make up on the spot
> > to be obnoxious.
>
> It's a point that a lot of website UI designers overlook - what we think
> is the best and most obvious layout, and even "the way everyone does it",
> isn't always what the website user finds most user friendly.
Couldn't agree with you more here.
> As to the OPs problem, perhaps it needs a different approach:
(I am the OP, just for your information) :)
> Presumably for any given city you can obtain a list of valid streets?
No, not really. I have a database of addresses that the search should
match against. The relevant DB fields are these:
streetname "Stora gatan"
streetnumber "34"
streetletter "B"
address "Stora gatan 34B"
In-data variants I am concerned with are:
"Stora gatan"
"Stora gatan 34"
"Stora gatan 34b"
"Stora gatan 34 b"
And I need to build a regexp to extract the three parts from all these
in-data versions to match agains the "address" field (or, maybe even
against the three discrete fields).
This is where I'm stuck, since the regexp I use doesn't adequately
match the different versions of how the street letter is sent.
> Another option, and one that is becoming popular in some areas, might be
> to allow entry of the postcode, and then generate a drop down list,
> perhaps populated using an XMLHttpRequest, of addresses matching the
> postcode - this would require access to a copy of the relevant postcode
> database which might involve some cost, and of course you need to
> implement a mechanism for handling updates to that database.
I have all the post codes, but that's not helping me here. In short,
if the same address exists in two post codes, I would show both and
the user would select which one.
It's the address part that is my current problem. And splitting the
search box into three boxes (one for name, number and letter) is not a
desirable option for this application, unfortunately.
> I don't think you're ever going to parse a single line address entry with
> regex, because there are too many different formats that might get thrown
> at you.
Yeah, people here keep claiming that while ignoring that I am only
interested in capturing the above formats, that make out the vast vast
vast majority of all searches being done to this database. If I don't
match "34 Storgatan b", that's just fine by me. I have a very specific
case of searches that currently fail that I feel strongly could be
averted by a regular expression on the receiving end.
--
Sandman[.net]
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2011-11-24 22:41 +0100 |
| Message-ID | <4340380.gzYE47dVVZ@PointedEars.de> |
| In reply to | #3903 |
Sandman wrote: > Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: >> >> > It is at this point that most people that have an actual need to >> >> > solve these kinds of problems turn to the available commercial >> >> > software and decide to solve it with money instead of manpower. >> >> >> >> Where the question must be allowed: How came that the data has not >> >> been >> >> requested and stored in a structured form to begin with? That is, for >> >> example, why only an address field in a form – why not a street, house >> >> number aso. field? >> > >> > Convenience for the user, of course. >> >> You can't be serious. > > I can :) Apparently. >> > This is a form that says "Are you connected to the citynet?" and then >> > you just enter your address to search the database. If the user has to >> > provide street name, street number and street letter in separate >> > fields, it's inconvenient for them. >> >> No, it's not. > > Actually, yes it is :) No, it is not. >> With separate controls they can be sure where to enter what; >> it is accessible, and you have no problem processing the data. With one >> control, neither applies. > > Well, I have been doing this for about ten years now, I have been doing this for about fourteen years now. So what? There are basic accessibility guidelines that no amount of development experience can substitute (although studying usability, as I did, can help). Many of which must be followed per legislation in some countries. > and recived tons of feedback from my clients on things like this. […] There remains the possibility that you did it wrong in another way all the time. > When I say it's inconvenient for the end user, it's not something I make > up on the spot to be obnoxious. Nevertheless, your logic is flawed. PointedEars -- Use any version of Microsoft Frontpage to create your site. (This won't prevent people from viewing your source, but no one will want to steal it.) -- from <http://www.vortex-webdesign.com/help/hidesource.htm> (404-comp.)
[toc] | [prev] | [next] | [standalone]
| From | Sandman <mr@sandman.net> |
|---|---|
| Date | 2011-11-25 09:26 +0100 |
| Message-ID | <mr-07C3E2.09260025112011@News.Individual.NET> |
| In reply to | #3920 |
In article <4340380.gzYE47dVVZ@PointedEars.de>, Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote: > >> >> > It is at this point that most people that have an actual need to > >> >> > solve these kinds of problems turn to the available commercial > >> >> > software and decide to solve it with money instead of manpower. > >> >> > >> >> Where the question must be allowed: How came that the data has not > >> >> been > >> >> requested and stored in a structured form to begin with? That is, for > >> >> example, why only an address field in a form – why not a street, house > >> >> number aso. field? > >> > > >> > Convenience for the user, of course. > >> > >> You can't be serious. > > > > I can :) > > Apparently. Indeed. :) > >> > This is a form that says "Are you connected to the citynet?" and then > >> > you just enter your address to search the database. If the user has to > >> > provide street name, street number and street letter in separate > >> > fields, it's inconvenient for them. > >> > >> No, it's not. > > > > Actually, yes it is :) > > No, it is not. Actually, yes it is :) > >> With separate controls they can be sure where to enter what; > >> it is accessible, and you have no problem processing the data. With one > >> control, neither applies. > > > > Well, I have been doing this for about ten years now, > > I have been doing this for about fourteen years now. So what? You have monitored swedish address search terms for fourteen years? We should compare notes. > There are basic accessibility guidelines that no amount of > development experience can substitute (although studying usability, > as I did, can help). Many of which must be followed per > legislation in some countries. This.. has nothing to do with the topic at hand. > > and recived tons of feedback from my clients on things like this. […] > > There remains the possibility that you did it wrong in another way all the > time. Also, there is a possibility that you don't have enough information about the details of my situation to make any judgmental comments at all. > > When I say it's inconvenient for the end user, it's not something I make > > up on the spot to be obnoxious. > > Nevertheless, your logic is flawed. Well, as long as you're merely *saying* that instead of actually, you know, *substantiate* that opinion, I have no idea what you expect me to do with it. Words are easy :) -- Sandman[.net]
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | comp.lang.php
csiph-web