Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #3859

Re: preg_match() oddities and question

From Sandman <mr@sandman.net>
Newsgroups comp.lang.php
Subject Re: preg_match() oddities and question
Date 2011-11-22 13:12 +0100
Message-ID <mr-C40BF0.13123922112011@News.Individual.NET> (permalink)
References <mr-5B96D1.12212022112011@News.Individual.NET> <jag24m$4nj$1@softins.clara.co.uk>

Show all headers | View raw


In article <jag24m$4nj$1@softins.clara.co.uk>,
 tony@mountifield.org (Tony Mountifield) wrote:

> > So I have this regexp:
> > 
> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){
> 
> You don't need the commas in the character class, unless you want to
> match a literal comma, in which case you only need it once.

Right, thanks :)

> >     $streetname = uc_words($m[1]);
> >     $streetnumber = trim($m[2]);
> >     $streetletter = strtoupper($m[3]);
> >     $search = trim($streetname . SPACE . $streetnumber . 
> > $streetletter);
> > }
> > 
> > The desired result is taki9ng the input ($search) and split it into 
> > its parts as an address, right? $search can be, for example, "foo 
> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet".
> 
> What about "foo street"? (i.e. with a space, but no number)

Exactly, that gets this:

Array
(
    [0] => foo street
    [1] => foo
    [2] => 
    [3] => street
)

Which is incorrect. IN fact, the last group SHOULD be defined as 
([A-Za-z]{0,1}) but that still messes it up like:

Array
(
    [0] => foo street
    [1] => foo stree
    [2] => 
    [3] => t
)

So I've tried variations for that as well.

<snip>

> And you would also get:
> 
> Array
> (
>     [0] => foo street
>     [1] => foo
>     [2] => 
>     [3] => street
> )
> 
> > As you can see, the last group "([A-Z,a-z,-]*?)" matches the entire 
> > search term since there are no digits and the first group is 
> > non-greedy. And if I make the first group greedy, "longstreet" is 
> > matched correctly, but it also catches the entire "longstreet 45b" 
> > when searching for that.
> 
> Yes, you need to define your rules more closely. Not at the regex level,
> but actually at the logic/decision level. If you can make rules that
> can unambiguously specify how all kinds of input should be parsed,
> then you can look at how to represent that in regexes. You might need
> some additional logic to operate on the parsed result.

What you're basically suggesting is a series of regexp to find out 
what "style" an adress is given in, and then parse out the parts? 
Because I'm not sure how I would be able to do it without a series if 
if/else preg_match():es?

> > Also, when searching for a term in swedish characters, I get this:
> > 
> > Array
> > (
> >     [0] => vikavÀgen
> >     [1] => vikavÀ
> >     [2] => 
> >     [3] => gen
> > )
> > 
> > Which is quite odd to me, why isn't "vikavÀgen" matched the same 
> > (undesired) way that "oongstreet". I have tried the /u modifier, and 
> > made sure that it was utf8-encoded, but it didn't make a difference 
> > (incoming encoding is ISO 8859-1).
> > 
> > Why the difference, and how do I correctly parse out parts as needed?
> 
> That's because À is not in the set A-Za-z. If you want a character class
> that properly recognises locale-specific letters, you need to change your
> character class above to this:
> 
> [[:alpha:]\-]
> 
> Hope this helps!

That explains the difference, thank you very much for that. Now I 
still need to figure out a global parse routine or criteria for 
parsing out the address parts...







-- 
Sandman[.net]

Back to comp.lang.php | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:21 +0100
  Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 11:26 +0000
    Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 12:36 +0100
    Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-22 07:22 -0500
  Re: preg_match() oddities and question tony@mountifield.org (Tony Mountifield) - 2011-11-22 11:47 +0000
    Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:12 +0100
  Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 13:30 +0100
    Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-22 13:55 +0100
      Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-22 17:56 +0100
        Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 17:30 +0000
          Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-22 17:20 -0600
            Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-22 23:59 +0000
            Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 01:59 +0100
              Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:58 +0100
                Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-23 22:02 +0100
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:20 +0100
                Re: preg_match() oddities and question Denis McMahon <denismfmcmahon@gmail.com> - 2011-11-24 12:55 +0000
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:36 +0100
                Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-24 22:41 +0100
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 09:26 +0100
                Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 15:44 +0100
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 16:34 +0100
                Re: preg_match() oddities and question Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2011-11-25 23:23 +0100
              Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 09:35 +0000
        Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 09:55 +0100
          Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 07:53 -0600
            Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 19:01 +0100
              Re: preg_match() oddities and question The Natural Philosopher <tnp@invalid.invalid> - 2011-11-23 18:54 +0000
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-23 20:23 +0100
              Re: preg_match() oddities and question "Peter H. Coffin" <hellsop@ninehells.com> - 2011-11-23 12:58 -0600
                Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-24 08:28 +0100
  SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-25 10:32 +0100
    Re: SOLVED: Re: preg_match() oddities and question Jerry Stuckle <jstucklex@attglobal.net> - 2011-11-25 18:55 -0500
      Re: SOLVED: Re: preg_match() oddities and question Sandman <mr@sandman.net> - 2011-11-26 11:21 +0100

csiph-web