Path: csiph.com!x330-a1.tempe.blueboxinc.net!newsfeed.hal-mli.net!feeder3.hal-mli.net!newsfeed.hal-mli.net!feeder1.hal-mli.net!de-l.enfer-du-nord.net!feeder2.enfer-du-nord.net!fu-berlin.de!uni-berlin.de!individual.net!not-for-mail From: Sandman Newsgroups: comp.lang.php Subject: Re: preg_match() oddities and question Date: Wed, 23 Nov 2011 19:01:14 +0100 Lines: 57 Message-ID: References: <1670168.aK4W3vaeNJ@PointedEars.de> <3004614.SPkdTlGXAF@PointedEars.de> Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Trace: individual.net 1iZge02Ow5x5I2/d+berpgkQxu4wvkCcHycwCgMY+Yt0jg4ag= X-Orig-Path: mr Cancel-Lock: sha1:9J3K5P3yeYyNinQDVVPiIzUFvjY= User-Agent: MT-NewsWatcher/3.5.2 (Intel Mac OS X) X-Face: $@,Vfa$,)%=Qa7L]y)&oZj_\EiHc}}Af0Bei"4a_%)"c6TQ+P/:53>;PNGuWUmkqyeN-qM65foJ[;T_(k;>]&G\T4Lhm:2 ujye2_,iUJFE;NZn>y;.|-hl7g~bIOF1qG\o, "Peter H. Coffin" wrote: > On Wed, 23 Nov 2011 09:55:22 +0100, Sandman wrote: > > > In article <3004614.SPkdTlGXAF@PointedEars.de>, Thomas 'PointedEars' > > Lahn wrote: > > > >> >> "10 East 42nd Street, New York, NY 10017, USA". > >> > > >> > That wouldn't be a normal swedish address, no. :) > >> > >> You had not limited the country or the language of your street > >> addresses. > > > > Well, to my defense, the subject line was "preg_match() and swedish > > characters" until I changed it. I hadn't changed it when I wrote my > > examples. > > > >> My point is that parsing a street name and a house number from a > >> street address is a hard problem that cannot be solved only by > >> applying one regular expression. > > > > Right, but your example is not a valid argument for that conclusion. > > My examples contained the variations of addresses that I wanted > > to match. Or are you saying that there is no way to use regular > > expressions to catch the examples I gave? Because I have a hard time > > believing that. > > Address-matching is a hard task. I did that for a decade professionally > (as part of a job, not the sole function), and it's not easy to do well > for even one postal system, and trying to write a generalized one is > basically impossible to manage in one lifetime. The best *simple* way > to manage it is to take a field, blow it out into individual words, > standardize all the words you can find without trying to sort out > what they are (which is the Very Hard part of that task), throw the > alphabetic ones into soundex or nysiis, make a loose match by a chunk of > postal code or city code or province, then pick the item(s) that have > the greatest number of matches between incoming and loose-match record > of the numeric and nysiis-encoded alphabetical elements. If you weight > things like "numeric match = 1, plaintext that's in a dictionary that > matches when nysiis = 2, nondictionary text that matches nysiis = 3", > and do that for NAME as well as ADDRESS, you get about as good as you > can get without buying someone else's work. And that's STILL a lot of > effort to write. Regexp alone for address matching is a snipe-hunt. It > looks obviously right and you can spend a lot of time playing with it, > but it ends up being a dead end. I thank you for your input, but I still maintain that my examples could be parsed by using a regular expression, and unless explicitly told so by using examples will I admit otherwise :-D No offense, though. -- Sandman[.net]