Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #3855 > unrolled thread

preg_match() oddities and question

Started bySandman <mr@sandman.net>
First post2011-11-22 12:21 +0100
Last post2011-11-26 11:21 +0100
Articles 20 on this page of 34 — 7 participants

Back to article view | Back to comp.lang.php


Contents

  preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:21 +0100
    Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 11:26 +0000
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:36 +0100
      Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-22 07:22 -0500
    Re: preg_match() oddities and question tony@mountifield.org (Tony Mountifield) - 2011-11-22 11:47 +0000
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:12 +0100
    Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 13:30 +0100
      Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:55 +0100
        Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 17:56 +0100
          Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 17:30 +0000
            Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-22 17:20 -0600
              Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 23:59 +0000
              Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 01:59 +0100
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:58 +0100
                  Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 22:02 +0100
                    Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:20 +0100
                      Re: preg_match() oddities and question Denis McMahon <denismfmcmahon@gmail.com> - 2011-11-24 12:55 +0000
                        Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:36 +0100
                      Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-24 22:41 +0100
                        Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:26 +0100
                          Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 15:44 +0100
                            Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 16:34 +0100
                              Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 23:23 +0100
                Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 09:35 +0000
          Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:55 +0100
            Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 07:53 -0600
              Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 19:01 +0100
                Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 18:54 +0000
                  Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 20:23 +0100
                Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 12:58 -0600
                  Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:28 +0100
    SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 10:32 +0100
      Re: SOLVED: Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-25 18:55 -0500
        Re: SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-26 11:21 +0100

Page 1 of 2  [1] 2  Next page →


#3855 — preg_match() oddities and question

FromSandman <mr@sandman.net>
Date2011-11-22 12:21 +0100
Subjectpreg_match() oddities and question
Message-ID<mr-5B96D1.12212022112011@News.Individual.NET>
So I have this regexp:

if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
    $streetname = uc_words($m[1]);
    $streetnumber = trim($m[2]);
    $streetletter = strtoupper($m[3]);
    $search = trim($streetname . SPACE . $streetnumber . 
$streetletter);
}

The desired result is taki9ng the input ($search) and split it into 
its parts as an address, right? $search can be, for example, "foo 
street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".

So, if I print_r($m) with different input I get:

Array
(
    [0] => foo street 34
    [1] => foo street
    [2] => 34
    [3] => 
)
Array
(
    [0] => longstreet 45b
    [1] => longstreet
    [2] => 45
    [3] => b
)
Array
(
    [0] => longstreet 45 b
    [1] => longstreet
    [2] => 45
    [3] => b
)

You get the idea. But problems arise when I search for the streetname 
alone:

Array
(
    [0] => longstreet
    [1] => 
    [2] => 
    [3] => longstreet
)

As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire 
search term since there are no digits and the first group is 
non-greedy. And if I make the first group greedy, "longstreet" is 
matched correctly, but it also catches the entire "longstreet 45b" 
when searching for that.

Also, when searching for a term in swedish characters, I get this:

Array
(
    [0] => vikavägen
    [1] => vikavä
    [2] => 
    [3] => gen
)

Which is quite odd to me, why isn't "vikavägen" matched the same 
(undesired) way that "oongstreet". I have tried the /u modifier, and 
made sure that it was utf8-encoded, but it didn't make a difference 
(incoming encoding is ISO 8859-1).

Why the difference, and how do I correctly parse out parts as needed?

Any help is appreciated. 





-- 
Sandman[.net]

[toc] | [next] | [standalone]


#3856

FromThe Natural Philosopher <tnp@invalid.invalid>
Date2011-11-22 11:26 +0000
Message-ID<jag0sr$49h$1@news.albasani.net>
In reply to#3855
Sandman wrote:

> 
> Why the difference, and how do I correctly parse out parts as needed?
> 

I have always found establishing the correct regexp expression to take 
longer than writing my own filters in whatever language  I happened to 
be using....


Life is too short for regexps.


> Any help is appreciated. 
> 
> 
> 
> 
> 

[toc] | [prev] | [next] | [standalone]


#3857

FromSandman <mr@sandman.net>
Date2011-11-22 12:36 +0100
Message-ID<mr-7CF721.12361222112011@News.Individual.NET>
In reply to#3856
In article <jag0sr$49h$1@news.albasani.net>,
 The Natural Philosopher <tnp@invalid.invalid> wrote:

> > Why the difference, and how do I correctly parse out parts as needed?
> 
> I have always found establishing the correct regexp expression to take 
> longer than writing my own filters in whatever language  I happened to 
> be using....
> 
> Life is too short for regexps.

I've never had much problem (time-wise) with regexps. I'm just stumped 
about the difference in execution of this one, and need a little help 
figuring out the syntax for the other part. In short, regexps are 
rarely a problem for me, and I don't know how I would solve my 
situation without using one, if you have any suggestions, please share 
:)



-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3862

FromJerry Stuckle <jstucklex@attglobal.net>
Date2011-11-22 07:22 -0500
Message-ID<jag45m$5um$3@dont-email.me>
In reply to#3856
On 11/22/2011 6:26 AM, The Natural Philosopher wrote:
> Sandman wrote:
>
>>
>> Why the difference, and how do I correctly parse out parts as needed?
>>
>
> I have always found establishing the correct regexp expression to take
> longer than writing my own filters in whatever language I happened to be
> using....
>
>
> Life is too short for regexps.
>

That's because regex's take intelligence - which you don't have.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
JDS Computer Training Corp.
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#3858

Fromtony@mountifield.org (Tony Mountifield)
Date2011-11-22 11:47 +0000
Message-ID<jag24m$4nj$1@softins.clara.co.uk>
In reply to#3855
In article <mr-5B96D1.12212022112011@News.Individual.NET>,
Sandman  <mr@sandman.net> wrote:
> So I have this regexp:
> 
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){

You don't need the commas in the character class, unless you want to
match a literal comma, in which case you only need it once.

>     $streetname = uc_words($m[1]);
>     $streetnumber = trim($m[2]);
>     $streetletter = strtoupper($m[3]);
>     $search = trim($streetname . SPACE . $streetnumber . 
> $streetletter);
> }
> 
> The desired result is taki9ng the input ($search) and split it into 
> its parts as an address, right? $search can be, for example, "foo 
> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".

What about "foo street"? (i.e. with a space, but no number)

> So, if I print_r($m) with different input I get:
> 
> Array
> (
>     [0] => foo street 34
>     [1] => foo street
>     [2] => 34
>     [3] => 
> )
> Array
> (
>     [0] => longstreet 45b
>     [1] => longstreet
>     [2] => 45
>     [3] => b
> )
> Array
> (
>     [0] => longstreet 45 b
>     [1] => longstreet
>     [2] => 45
>     [3] => b
> )
> 
> You get the idea. But problems arise when I search for the streetname 
> alone:
> 
> Array
> (
>     [0] => longstreet
>     [1] => 
>     [2] => 
>     [3] => longstreet
> )

And you would also get:

Array
(
    [0] => foo street
    [1] => foo
    [2] => 
    [3] => street
)

> As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire 
> search term since there are no digits and the first group is 
> non-greedy. And if I make the first group greedy, "longstreet" is 
> matched correctly, but it also catches the entire "longstreet 45b" 
> when searching for that.

Yes, you need to define your rules more closely. Not at the regex level,
but actually at the logic/decision level. If you can make rules that
can unambiguously specify how all kinds of input should be parsed,
then you can look at how to represent that in regexes. You might need
some additional logic to operate on the parsed result.

> Also, when searching for a term in swedish characters, I get this:
> 
> Array
> (
>     [0] => vikavägen
>     [1] => vikavä
>     [2] => 
>     [3] => gen
> )
> 
> Which is quite odd to me, why isn't "vikavägen" matched the same 
> (undesired) way that "oongstreet". I have tried the /u modifier, and 
> made sure that it was utf8-encoded, but it didn't make a difference 
> (incoming encoding is ISO 8859-1).
> 
> Why the difference, and how do I correctly parse out parts as needed?

That's because ä is not in the set A-Za-z. If you want a character class
that properly recognises locale-specific letters, you need to change your
character class above to this:

[[:alpha:]\-]

Hope this helps!
Tony
-- 
Tony Mountifield
Work: tony@softins.co.uk - http://www.softins.co.uk
Play: tony@mountifield.org - http://tony.mountifield.org

[toc] | [prev] | [next] | [standalone]


#3859

FromSandman <mr@sandman.net>
Date2011-11-22 13:12 +0100
Message-ID<mr-C40BF0.13123922112011@News.Individual.NET>
In reply to#3858
In article <jag24m$4nj$1@softins.clara.co.uk>,
 tony@mountifield.org (Tony Mountifield) wrote:

> > So I have this regexp:
> > 
> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> 
> You don't need the commas in the character class, unless you want to
> match a literal comma, in which case you only need it once.

Right, thanks :)

> >     $streetname = uc_words($m[1]);
> >     $streetnumber = trim($m[2]);
> >     $streetletter = strtoupper($m[3]);
> >     $search = trim($streetname . SPACE . $streetnumber . 
> > $streetletter);
> > }
> > 
> > The desired result is taki9ng the input ($search) and split it into 
> > its parts as an address, right? $search can be, for example, "foo 
> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
> 
> What about "foo street"? (i.e. with a space, but no number)

Exactly, that gets this:

Array
(
    [0] => foo street
    [1] => foo
    [2] => 
    [3] => street
)

Which is incorrect. IN fact, the last group SHOULD be defined as 
([A-Za-z]{0,1}) but that still messes it up like:

Array
(
    [0] => foo street
    [1] => foo stree
    [2] => 
    [3] => t
)

So I've tried variations for that as well.

<snip>

> And you would also get:
> 
> Array
> (
>     [0] => foo street
>     [1] => foo
>     [2] => 
>     [3] => street
> )
> 
> > As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire 
> > search term since there are no digits and the first group is 
> > non-greedy. And if I make the first group greedy, "longstreet" is 
> > matched correctly, but it also catches the entire "longstreet 45b" 
> > when searching for that.
> 
> Yes, you need to define your rules more closely. Not at the regex level,
> but actually at the logic/decision level. If you can make rules that
> can unambiguously specify how all kinds of input should be parsed,
> then you can look at how to represent that in regexes. You might need
> some additional logic to operate on the parsed result.

What you're basically suggesting is a series of regexp to find out 
what "style" an adress is given in, and then parse out the parts? 
Because I'm not sure how I would be able to do it without a series if 
if/else preg_match():es?

> > Also, when searching for a term in swedish characters, I get this:
> > 
> > Array
> > (
> >     [0] => vikavÀgen
> >     [1] => vikavÀ
> >     [2] => 
> >     [3] => gen
> > )
> > 
> > Which is quite odd to me, why isn't "vikavÀgen" matched the same 
> > (undesired) way that "oongstreet". I have tried the /u modifier, and 
> > made sure that it was utf8-encoded, but it didn't make a difference 
> > (incoming encoding is ISO 8859-1).
> > 
> > Why the difference, and how do I correctly parse out parts as needed?
> 
> That's because À is not in the set A-Za-z. If you want a character class
> that properly recognises locale-specific letters, you need to change your
> character class above to this:
> 
> [[:alpha:]\-]
> 
> Hope this helps!

That explains the difference, thank you very much for that. Now I 
still need to figure out a global parse routine or criteria for 
parsing out the address parts...







-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3863

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-22 13:30 +0100
Message-ID<1670168.aK4W3vaeNJ@PointedEars.de>
In reply to#3855
Sandman wrote:

> So I have this regexp:
> 
> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>     $streetname = uc_words($m[1]);
>     $streetnumber = trim($m[2]);
>     $streetletter = strtoupper($m[3]);
>     $search = trim($streetname . SPACE . $streetnumber .
> $streetletter);
> }
> 
> The desired result is taki9ng the input ($search) and split it into
> its parts as an address, right? $search can be, for example, "foo
> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".

"10 East 42nd Street, New York, NY 10017, USA".


PointedEars
-- 
When all you know is jQuery, every problem looks $(olvable).

[toc] | [prev] | [next] | [standalone]


#3865

FromSandman <mr@sandman.net>
Date2011-11-22 13:55 +0100
Message-ID<mr-52D11B.13550322112011@News.Individual.NET>
In reply to#3863
In article <1670168.aK4W3vaeNJ@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> Sandman wrote:
> 
> > So I have this regexp:
> > 
> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> >     $streetname = uc_words($m[1]);
> >     $streetnumber = trim($m[2]);
> >     $streetletter = strtoupper($m[3]);
> >     $search = trim($streetname . SPACE . $streetnumber .
> > $streetletter);
> > }
> > 
> > The desired result is taki9ng the input ($search) and split it into
> > its parts as an address, right? $search can be, for example, "foo
> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
> 
> "10 East 42nd Street, New York, NY 10017, USA".

That wouldn't be a normal swedish address, no. :)







-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3871

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-22 17:56 +0100
Message-ID<3004614.SPkdTlGXAF@PointedEars.de>
In reply to#3865
Sandman wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> Sandman wrote:
>> > So I have this regexp:
>> > 
>> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>> >     $streetname = uc_words($m[1]);
>> >     $streetnumber = trim($m[2]);
>> >     $streetletter = strtoupper($m[3]);
>> >     $search = trim($streetname . SPACE . $streetnumber .
>> > $streetletter);
>> > }
>> > 
>> > The desired result is taki9ng the input ($search) and split it into
>> > its parts as an address, right? $search can be, for example, "foo
>> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>> 
>> "10 East 42nd Street, New York, NY 10017, USA".
> 
> That wouldn't be a normal swedish address, no. :)

You had not limited the country or the language of your street addresses.

My point is that parsing a street name and a house number from a street 
address is a hard problem that cannot be solved only by applying one regular 
expression.


PointedEars
-- 
    realism:    HTML 4.01 Strict
    evangelism: XHTML 1.0 Strict
    madness:    XHTML 1.1 as application/xhtml+xml
                                                    -- Bjoern Hoehrmann

[toc] | [prev] | [next] | [standalone]


#3873

FromThe Natural Philosopher <tnp@invalid.invalid>
Date2011-11-22 17:30 +0000
Message-ID<jagm86$kus$1@news.albasani.net>
In reply to#3871
Thomas 'PointedEars' Lahn wrote:
> Sandman wrote:
> 
>> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>>> Sandman wrote:
>>>> So I have this regexp:
>>>>
>>>> if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
>>>>     $streetname = uc_words($m[1]);
>>>>     $streetnumber = trim($m[2]);
>>>>     $streetletter = strtoupper($m[3]);
>>>>     $search = trim($streetname . SPACE . $streetnumber .
>>>> $streetletter);
>>>> }
>>>>
>>>> The desired result is taki9ng the input ($search) and split it into
>>>> its parts as an address, right? $search can be, for example, "foo
>>>> street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
>>> "10 East 42nd Street, New York, NY 10017, USA".
>> That wouldn't be a normal swedish address, no. :)
> 
> You had not limited the country or the language of your street addresses.
> 
> My point is that parsing a street name and a house number from a street 
> address is a hard problem that cannot be solved only by applying one regular 
> expression.
> 

Quite right. Is worse than you can possibly iagine at leats here in te 
UK, where addresses can be as little as 2 lines long or up to 6..

So

10 Wonkers place, LONDON EC3 7QY is a typical TOWN address

Out in the sticks you might get

Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr 
Stonehouse, Gloucestershire GL13 6AH


And if that comes at you without commas, god help you.

I have spent DAYS taking name/address fields and parsing them *manually* 
into structured tables...

> 
> PointedEars

[toc] | [prev] | [next] | [standalone]


#3878

From"Peter H. Coffin" <hellsop@ninehells.com>
Date2011-11-22 17:20 -0600
Message-ID<slrnjcobhs.85q.hellsop@nibelheim.ninehells.com>
In reply to#3873
On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote:
> Quite right. Is worse than you can possibly iagine at leats here in te 
> UK, where addresses can be as little as 2 lines long or up to 6..
>
> So
>
> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
>
> Out in the sticks you might get
>
> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr 
> Stonehouse, Gloucestershire GL13 6AH
>
>
> And if that comes at you without commas, god help you.
>
> I have spent DAYS taking name/address fields and parsing them *manually* 
> into structured tables...

It is at this point that most people that have an actual need to solve
these kinds of problems turn to the available commercial software and
decide to solve it with money instead of manpower.

-- 
They got rid of it because they judged it more trouble than it was 
worth.  (And considering they'd gone to great lengths to minimize its 
worth, I suppose they were right.)
              -- J. D. Baldwin

[toc] | [prev] | [next] | [standalone]


#3879

FromThe Natural Philosopher <tnp@invalid.invalid>
Date2011-11-22 23:59 +0000
Message-ID<jahd15$666$3@news.albasani.net>
In reply to#3878
Peter H. Coffin wrote:
> On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote:
>> Quite right. Is worse than you can possibly iagine at leats here in te 
>> UK, where addresses can be as little as 2 lines long or up to 6..
>>
>> So
>>
>> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
>>
>> Out in the sticks you might get
>>
>> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr 
>> Stonehouse, Gloucestershire GL13 6AH
>>
>>
>> And if that comes at you without commas, god help you.
>>
>> I have spent DAYS taking name/address fields and parsing them *manually* 
>> into structured tables...
> 
> It is at this point that most people that have an actual need to solve
> these kinds of problems turn to the available commercial software and
> decide to solve it with money instead of manpower.
> 
there is no AI that can match a human brain in decoding human idiocy...yet

[toc] | [prev] | [next] | [standalone]


#3880

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-23 01:59 +0100
Message-ID<1780410.D2KPVSPYKU@PointedEars.de>
In reply to#3878
Peter H. Coffin wrote:

> On Tue, 22 Nov 2011 17:30:45 +0000, The Natural Philosopher wrote:
>> Quite right. Is worse than you can possibly iagine at leats here in te
>> UK, where addresses can be as little as 2 lines long or up to 6..
>>
>> So
>>
>> 10 Wonkers place, LONDON EC3 7QY is a typical TOWN address
>>
>> Out in the sticks you might get
>>
>> Apartment 4b, the Old Town House, Shire Lane, Recketts Green, Nr
>> Stonehouse, Gloucestershire GL13 6AH
>>
>>
>> And if that comes at you without commas, god help you.
>>
>> I have spent DAYS taking name/address fields and parsing them *manually*
>> into structured tables...
> 
> It is at this point that most people that have an actual need to solve
> these kinds of problems turn to the available commercial software and
> decide to solve it with money instead of manpower.

Where the question must be allowed: How came that the data has not been 
requested and stored in a structured form to begin with?  That is, for 
example, why only an address field in a form – why not a street, house
number aso. field?  ISTM that we are seeing here an example of a mistake 
made at the beginning which overall cost naturally grows larger and larger 
as the project is nearing completion.


PointedEars
-- 
Danny Goodman's books are out of date and teach practices that are
positively harmful for cross-browser scripting.
  -- Richard Cornford, cljs, <cife6q$253$1$8300dec7@news.demon.co.uk> (2004)

[toc] | [prev] | [next] | [standalone]


#3883

FromSandman <mr@sandman.net>
Date2011-11-23 09:58 +0100
Message-ID<mr-59B44E.09581623112011@News.Individual.NET>
In reply to#3880
In article <1780410.D2KPVSPYKU@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> > It is at this point that most people that have an actual need to solve
> > these kinds of problems turn to the available commercial software and
> > decide to solve it with money instead of manpower.
> 
> Where the question must be allowed: How came that the data has not been 
> requested and stored in a structured form to begin with?  That is, for 
> example, why only an address field in a form – why not a street, house
> number aso. field? 

Convenience for the user, of course.

This is a form that says "Are you connected to the citynet?" and then 
you just enter your address to search the database. If the user has to 
provide street name, street number and street letter in separate 
fields, it's inconvenient for them.


-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3900

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-23 22:02 +0100
Message-ID<4778042.ypaU67uLZW@PointedEars.de>
In reply to#3883
Sandman wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> > It is at this point that most people that have an actual need to solve
>> > these kinds of problems turn to the available commercial software and
>> > decide to solve it with money instead of manpower.
>> 
>> Where the question must be allowed: How came that the data has not been
>> requested and stored in a structured form to begin with?  That is, for
>> example, why only an address field in a form – why not a street, house
>> number aso. field?
> 
> Convenience for the user, of course.

You can't be serious.

> This is a form that says "Are you connected to the citynet?" and then
> you just enter your address to search the database. If the user has to
> provide street name, street number and street letter in separate
> fields, it's inconvenient for them.

No, it's not.  With separate controls they can be sure where to enter what; 
it is accessible, and you have no problem processing the data.  With one 
control, neither applies.


PointedEars
-- 
Use any version of Microsoft Frontpage to create your site.
(This won't prevent people from viewing your source, but no one
will want to steal it.)
  -- from <http://www.vortex-webdesign.com/help/hidesource.htm> (404-comp.)

[toc] | [prev] | [next] | [standalone]


#3903

FromSandman <mr@sandman.net>
Date2011-11-24 08:20 +0100
Message-ID<mr-8BFE1F.08204424112011@News.Individual.NET>
In reply to#3900
In article <4778042.ypaU67uLZW@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> >> > It is at this point that most people that have an actual need to solve
> >> > these kinds of problems turn to the available commercial software and
> >> > decide to solve it with money instead of manpower.
> >> 
> >> Where the question must be allowed: How came that the data has not been
> >> requested and stored in a structured form to begin with?  That is, for
> >> example, why only an address field in a form – why not a street, house
> >> number aso. field?
> > 
> > Convenience for the user, of course.
> 
> You can't be serious.

I can :)

> > This is a form that says "Are you connected to the citynet?" and then
> > you just enter your address to search the database. If the user has to
> > provide street name, street number and street letter in separate
> > fields, it's inconvenient for them.
> 
> No, it's not.

Actually, yes it is :)

> With separate controls they can be sure where to enter what; 
> it is accessible, and you have no problem processing the data.  With one 
> control, neither applies.

Well, I have been doing this for about ten years now, and recived tons 
of feedback from my clients on things like this. When I say it's 
inconvenient for the end user, it's not something I make up on the 
spot to be obnoxious.

Just like with my examples, I have a pretty clear picture of what 
problem I need to solve. I find it curious that no one in CLP even 
attempt to look at that, and instead trying to find other examples, or 
claiming that the frontend should be changed.

Makes me think that you guys deem the examples I gave as unsolvable, 
which of course I refuse to agree with. :)

No offense though.



-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3916

FromDenis McMahon <denismfmcmahon@gmail.com>
Date2011-11-24 12:55 +0000
Message-ID<4ece3eb7$0$28723$a8266bb1@newsreader.readnews.com>
In reply to#3903
On Thu, 24 Nov 2011 08:20:45 +0100, Sandman wrote:

> Well, I have been doing this for about ten years now, and recived tons
> of feedback from my clients on things like this. When I say it's
> inconvenient for the end user, it's not something I make up on the spot
> to be obnoxious.

It's a point that a lot of website UI designers overlook - what we think 
is the best and most obvious layout, and even "the way everyone does it", 
isn't always what the website user finds most user friendly.

As to the OPs problem, perhaps it needs a different approach:

Presumably for any given city you can obtain a list of valid streets?

Why not:

foreach ($street_name in $streets_of_city) {
  if (($street_start = strpos($address, $street_name)) !== false) {
    foreach ($district_name in $districts_of_city) {
      if (($district_start = strpos($address, $district_name, 
$street_start)) !== false) {
      break;
      }
    }
  break;
  }
}

You may need to allow for "district creep" and will need to allow for 
streets crossing district boundaries.

Another option, and one that is becoming popular in some areas, might be 
to allow entry of the postcode, and then generate a drop down list, 
perhaps populated using an XMLHttpRequest, of addresses matching the 
postcode - this would require access to a copy of the relevant postcode 
database which might involve some cost, and of course you need to 
implement a mechanism for handling updates to that database.

Possibly a similar sort of approach, based on a select menu for district, 
then a select menu for street names, and then a text entry field for 
"apartment number, building name and / or number, or house number as 
appropriate".

I don't think you're ever going to parse a single line address entry with 
regex, because there are too many different formats that might get thrown 
at you. You either need to get the data in a different format, or find a 
different method of processing the data that you are getting.

Rgds

Denis McMahon

[toc] | [prev] | [next] | [standalone]


#3923

FromSandman <mr@sandman.net>
Date2011-11-25 09:36 +0100
Message-ID<mr-1BAC45.09360625112011@News.Individual.NET>
In reply to#3916
In article <4ece3eb7$0$28723$a8266bb1@newsreader.readnews.com>,
 Denis McMahon <denismfmcmahon@gmail.com> wrote:

> > Well, I have been doing this for about ten years now, and recived tons
> > of feedback from my clients on things like this. When I say it's
> > inconvenient for the end user, it's not something I make up on the spot
> > to be obnoxious.
> 
> It's a point that a lot of website UI designers overlook - what we think 
> is the best and most obvious layout, and even "the way everyone does it", 
> isn't always what the website user finds most user friendly.

Couldn't agree with you more here. 

> As to the OPs problem, perhaps it needs a different approach:

(I am the OP, just for your information) :)

> Presumably for any given city you can obtain a list of valid streets?

No, not really. I have a database of addresses that the search should 
match against. The relevant DB fields are these:

    streetname      "Stora gatan"
    streetnumber    "34"
    streetletter    "B"
    address         "Stora gatan 34B"

In-data variants I am concerned with are:

    "Stora gatan"
    "Stora gatan 34"
    "Stora gatan 34b"
    "Stora gatan 34 b"

And I need to build a regexp to extract the three parts from all these 
in-data versions to match agains the "address" field (or, maybe even 
against the three discrete fields).

This is where I'm stuck, since the regexp I use doesn't adequately 
match the different versions of how the street letter is sent.

> Another option, and one that is becoming popular in some areas, might be 
> to allow entry of the postcode, and then generate a drop down list, 
> perhaps populated using an XMLHttpRequest, of addresses matching the 
> postcode - this would require access to a copy of the relevant postcode 
> database which might involve some cost, and of course you need to 
> implement a mechanism for handling updates to that database.

I have all the post codes, but that's not helping me here. In short, 
if the same address exists in two post codes, I would show both and 
the user would select which one.

It's the address part that is my current problem. And splitting the 
search box into three boxes (one for name, number and letter) is not a 
desirable option for this application, unfortunately. 

> I don't think you're ever going to parse a single line address entry with 
> regex, because there are too many different formats that might get thrown 
> at you.

Yeah, people here keep claiming that while ignoring that I am only 
interested in capturing the above formats, that make out the vast vast 
vast majority of all searches being done to this database. If I don't 
match "34 Storgatan b", that's just fine by me. I have a very specific 
case of searches that currently fail that I feel strongly could be 
averted by a regular expression on the receiving end.




-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


#3920

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2011-11-24 22:41 +0100
Message-ID<4340380.gzYE47dVVZ@PointedEars.de>
In reply to#3903
Sandman wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:
>> >> > It is at this point that most people that have an actual need to
>> >> > solve these kinds of problems turn to the available commercial
>> >> > software and decide to solve it with money instead of manpower.
>> >> 
>> >> Where the question must be allowed: How came that the data has not
>> >> been
>> >> requested and stored in a structured form to begin with?  That is, for
>> >> example, why only an address field in a form – why not a street, house
>> >> number aso. field?
>> > 
>> > Convenience for the user, of course.
>> 
>> You can't be serious.
> 
> I can :)

Apparently.

>> > This is a form that says "Are you connected to the citynet?" and then
>> > you just enter your address to search the database. If the user has to
>> > provide street name, street number and street letter in separate
>> > fields, it's inconvenient for them.
>> 
>> No, it's not.
> 
> Actually, yes it is :)

No, it is not.
 
>> With separate controls they can be sure where to enter what;
>> it is accessible, and you have no problem processing the data.  With one
>> control, neither applies.
> 
> Well, I have been doing this for about ten years now,

I have been doing this for about fourteen years now.  So what?  There are 
basic accessibility guidelines that no amount of development experience can 
substitute (although studying usability, as I did, can help).  Many of which 
must be followed per legislation in some countries.

> and recived tons of feedback from my clients on things like this. […]

There remains the possibility that you did it wrong in another way all the 
time.

> When I say it's inconvenient for the end user, it's not something I make
> up on the spot to be obnoxious.

Nevertheless, your logic is flawed.


PointedEars
-- 
Use any version of Microsoft Frontpage to create your site.
(This won't prevent people from viewing your source, but no one
will want to steal it.)
  -- from <http://www.vortex-webdesign.com/help/hidesource.htm> (404-comp.)

[toc] | [prev] | [next] | [standalone]


#3922

FromSandman <mr@sandman.net>
Date2011-11-25 09:26 +0100
Message-ID<mr-07C3E2.09260025112011@News.Individual.NET>
In reply to#3920
In article <4340380.gzYE47dVVZ@PointedEars.de>,
 Thomas 'PointedEars' Lahn <PointedEars@web.de> wrote:

> >> >> > It is at this point that most people that have an actual need to
> >> >> > solve these kinds of problems turn to the available commercial
> >> >> > software and decide to solve it with money instead of manpower.
> >> >> 
> >> >> Where the question must be allowed: How came that the data has not
> >> >> been
> >> >> requested and stored in a structured form to begin with?  That is, for
> >> >> example, why only an address field in a form – why not a street, house
> >> >> number aso. field?
> >> > 
> >> > Convenience for the user, of course.
> >> 
> >> You can't be serious.
> > 
> > I can :)
> 
> Apparently.

Indeed. :)

> >> > This is a form that says "Are you connected to the citynet?" and then
> >> > you just enter your address to search the database. If the user has to
> >> > provide street name, street number and street letter in separate
> >> > fields, it's inconvenient for them.
> >> 
> >> No, it's not.
> > 
> > Actually, yes it is :)
> 
> No, it is not.

Actually, yes it is :)

> >> With separate controls they can be sure where to enter what;
> >> it is accessible, and you have no problem processing the data.  With one
> >> control, neither applies.
> > 
> > Well, I have been doing this for about ten years now,
> 
> I have been doing this for about fourteen years now.  So what?

You have monitored swedish address search terms for fourteen years? We 
should compare notes. 

> There are basic accessibility guidelines that no amount of 
> development experience can substitute (although studying usability, 
> as I did, can help).  Many of which must be followed per 
> legislation in some countries.

This.. has nothing to do with the topic at hand. 

> > and recived tons of feedback from my clients on things like this. […]
> 
> There remains the possibility that you did it wrong in another way all the 
> time.

Also, there is a possibility that you don't have enough information 
about the details of my situation to make any judgmental comments at 
all. 

> > When I say it's inconvenient for the end user, it's not something I make
> > up on the spot to be obnoxious.
> 
> Nevertheless, your logic is flawed.

Well, as long as you're merely *saying* that instead of actually, you 
know, *substantiate* that opinion, I have no idea what you expect me 
to do with it.

Words are easy :)




-- 
Sandman[.net]

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | comp.lang.php


csiph-web