Path: csiph.com!x330-a1.tempe.blueboxinc.net!usenet.pasdenom.info!gegeweb.42!gegeweb.eu!nntpfeed.proxad.net!proxad.net!feeder2-2.proxad.net!newsfeed.arcor.de!newsspool3.arcor-online.net!news.arcor.de.POSTED!not-for-mail Content-Type: text/plain; charset="ISO-8859-1" Message-ID: <3004614.SPkdTlGXAF@PointedEars.de> From: Thomas 'PointedEars' Lahn Reply-To: Thomas 'PointedEars' Lahn Organization: PointedEars Software (PES) Date: Tue, 22 Nov 2011 17:56:51 +0100 User-Agent: KNode/4.4.11 Content-Transfer-Encoding: 7Bit Subject: Re: preg_match() oddities and question Newsgroups: comp.lang.php References: <1670168.aK4W3vaeNJ@PointedEars.de> Followup-To: comp.lang.php MIME-Version: 1.0 Lines: 35 NNTP-Posting-Date: 22 Nov 2011 17:56:52 CET NNTP-Posting-Host: a6c6833d.newsspool3.arcor-online.net X-Trace: DXC=m]E:h@Q>CA>RLigj];iP=8McF=Q^Z^V384Fo<]lROoR18kFTgTK X-Complaints-To: usenet-abuse@arcor.de Xref: x330-a1.tempe.blueboxinc.net comp.lang.php:3871 Sandman wrote: > Thomas 'PointedEars' Lahn wrote: >> Sandman wrote: >> > So I have this regexp: >> > >> > if (preg_match("/^(.*?)\s*(\d*?)\s*([A-Z,a-z,-]*?)$/", $search, $m)){ >> > $streetname = uc_words($m[1]); >> > $streetnumber = trim($m[2]); >> > $streetletter = strtoupper($m[3]); >> > $search = trim($streetname . SPACE . $streetnumber . >> > $streetletter); >> > } >> > >> > The desired result is taki9ng the input ($search) and split it into >> > its parts as an address, right? $search can be, for example, "foo >> > street 34", "longstreet 45b", "longstreet 45 b" or just "longstreet". >> >> "10 East 42nd Street, New York, NY 10017, USA". > > That wouldn't be a normal swedish address, no. :) You had not limited the country or the language of your street addresses. My point is that parsing a street name and a house number from a street address is a hard problem that cannot be solved only by applying one regular expression. PointedEars -- realism: HTML 4.01 Strict evangelism: XHTML 1.0 Strict madness: XHTML 1.1 as application/xhtml+xml -- Bjoern Hoehrmann