Path: csiph.com!usenet.pasdenom.info!weretis.net!feeder4.news.weretis.net!news.mixmin.net!news2.arglkargh.de!news.karotte.org!uucp.gnuu.de!newsfeed.arcor.de!newsspool4.arcor-online.net!news.arcor.de.POSTED!not-for-mail Content-Type: text/plain; charset="UTF-8" Message-ID: <3150175.r8Iz2dskB7@PointedEars.de> From: Thomas 'PointedEars' Lahn Reply-To: Thomas 'PointedEars' Lahn Organization: PointedEars Software (PES) Date: Wed, 07 Nov 2012 23:21:20 +0100 User-Agent: KNode/4.4.11 Content-Transfer-Encoding: 8Bit X-Face: %i>XG-yXR'\"2P/C_aO%~;2o~?g0pPKmbOw^=NT`tprDEf++D.m7"}HW6.#=U:?2GGctkL,f89@H46O$ASoW&?s}.k+&. <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com> <48mi98trj215lkcssp9hq412ddggfnijed@4ax.com> <168bd0ef-ba1a-4448-ba80-8dac4e4c26ad@g8g2000yqp.googlegroups.com> Followup-To: comp.lang.javascript MIME-Version: 1.0 Lines: 163 NNTP-Posting-Date: 07 Nov 2012 23:21:21 CET NNTP-Posting-Host: 486db8fe.newsspool2.arcor-online.net X-Trace: DXC=YncmKc^KScJ@k=MdN::NBIA9EHlD;3YcB4Fo<]lROoRA8kFK@e^bhg0G;bbMl[2aEXcUnDG X-Complaints-To: usenet-abuse@arcor.de Xref: csiph.com comp.lang.javascript:17068 Patricia Shanahan wrote: > Scott Sauyet wrote: >> I personally can't see it. My example was intentionally verbose and >> as explicit as I could comfortably be, so as to remove the regex >> terseness factor from the equation and see if Tim still had >> objections. The objections many have to Perl are similar to one main >> objection to regexes: they are often simply unreadable: write-only >> code. Your abbreviations are shortened enough to no longer be clear. >> None of the keywords except "OR" are really obvious. It really >> doesn't look much clearer to me than `/^\s+|\s+$/g `. > > I think there is a significant difference between Perl and regex. In > Perl, terseness is a choice. One can use meaningful identifiers, white > space, and comments. I have written maintainable Perl code. > > In regex, there is no option for meaningful identifiers, or, as Gene > pointed out, white space. Comments have to be either before or after the > entire regex, not interleaved with it. Apples and oranges. Perl is a *programming language*; regular expressions are part of a *technique* (pattern matching) usable *in* programming languages. And apparently you are not aware that *especially* Perl's implementation of regular expressions, and Perl-Compatible Regular Expressions (PCRE) as supported by e.g. PHP, allow what you are asking for with the `x' (PCRE_EXTENDED) flag: my $leadingOrTrailingWhitespace = / ^ # (start of input # followed by \s+ # at least one whitespace) | # or \s+ # (at least one whitespace # followed by $ # end of input) /x; In PHP (with PCRE): In ECMAScript implementations (with JSX:regexp.js [1], revisions 272 and later): var leadingOrTrailingWhitespace = new jsx.regexp.RegExp( [ "^ # (start of input", " # followed by", "\\s+ # at least one whitespace)", "| # or", "\\s+ # (at least one whitespace", " # followed by", "$ # end of input)" ].join("\n"), "x" ); or (standards-compliant since Ed. 5, proprietarily available before): var leadingOrTrailingWhitespace = new jsx.regexp.RegExp( "^ # (start of input \n\ # followed by \n\ \\s+ # at least one whitespace) \n\ | # or \n\ \\s+ # (at least one whitespace \n\ # followed by \n\ $ # end of input", "x" ); > I would be much happier if one could build regular expressions up in > pieces, using identifiers to reference components: > > leadingSpaces = /^\s+/ > trailingSpaces = /\s+$/ > leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/ > > s = s.replace(/{leadingOrTrailingSpaces}/g, ""); Built-in (ECMA-262-5.1, ยง15.10.4): var leadingSpaces = "^\\s+"; var trailingSpaces = "\\s+$"; var leadingOrTrailingSpaces = new RegExp(leadingSpaces + "|" + trailingSpaces, "g"); var s = s.replace(leadingOrTrailingSpaces, ""); or var leadingSpaces = "^\\s+"; var trailingSpaces = "\\s+$"; var leadingOrTrailingSpaces = new RegExp([leadingSpaces, trailingSpaces].join("|"), "g"); var s = s.replace(leadingOrTrailingSpaces, ""); (You can find more examples in JSX:regexp.js where I am building the expressions to implement, for example, Unicode character property classes.) With JSX:regexp.js (revisions 275 and later): var leadingSpaces = /^\s+/; var trailingSpaces = /\s+$/; var leadingOrTrailingSpaces = jsx.regexp.concat(leadingSpaces, "|", trailingSpaces); s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), ""); or var leadingSpaces = /^\s+/; var trailingSpaces = /\s+$/; var leadingOrTrailingSpaces = leadingSpaces.concat("|", trailingSpaces); s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), ""); I am considering var leadingSpaces = /^\s+/; var trailingSpaces = /\s+$/; var leadingOrTrailingSpaces = jsx.regexp.alternate(leadingSpaces, trailingSpaces); s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), ""); and var leadingSpaces = /^\s+/; var trailingSpaces = /\s+$/; var leadingOrTrailingSpaces = leadingSpaces.alternate(trailingSpaces); s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), ""); What do you think? PointedEars ___________ [1] -- Sometimes, what you learn is wrong. If those wrong ideas are close to the root of the knowledge tree you build on a particular subject, pruning the bad branches can sometimes cause the whole tree to collapse. -- Mike Duffy in cljs,