Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.javascript > #17006 > unrolled thread

Good use of bitwise NOT "~" or unnecessary obfuscation?

Started byRobG <rgqld@iinet.net.au>
First post2012-11-01 18:32 -0700
Last post2012-11-07 11:22 -0800
Articles 20 on this page of 46 — 14 participants

Back to article view | Back to comp.lang.javascript


Contents

  Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-01 18:32 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Denis McMahon <denismfmcmahon@gmail.com> - 2012-11-02 06:14 +0000
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 02:28 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-02 10:22 +0100
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 10:37 +0100
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 06:25 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:35 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-02 14:58 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 08:49 -0700
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:52 +0100
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:49 +0000
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-05 13:18 +0100
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:48 -0800
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-06 15:04 +0000
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-06 10:43 -0800
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 11:22 -0800
                      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-06 12:16 -0800
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 12:53 -0800
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-07 17:34 +0100
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-07 23:21 +0100
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-07 23:23 +0100
                          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-07 17:00 -0800
                            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-08 07:16 +0100
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-07 17:50 +0000
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-05 19:26 +0000
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Stefan Weiss <krewecherl@gmail.com> - 2012-11-07 22:13 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 09:51 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 18:41 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 12:30 -0700
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 20:33 +0000
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:57 +0100
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:44 +0000
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Andrew Poulos <ap_prog@hotmail.com> - 2012-11-05 15:41 +1100
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-05 09:07 +0000
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 22:10 +0100
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 15:22 -0700
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 10:19 +0100
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-03 13:05 -0700
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 21:39 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 12:26 +0100
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Jukka K. Korpela" <jkorpela@cs.tut.fi> - 2012-11-02 12:28 +0200
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 10:00 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:15 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:32 -0800
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-06 18:23 -0800
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-07 11:22 -0800

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#17069

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2012-11-07 23:23 +0100
Message-ID<XnsA104EDE4957A2eejj99@194.109.133.133>
In reply to#17056
Patricia Shanahan wrote on 06 nov 2012 in comp.lang.javascript:

> leadingSpaces = /^\s+/
> trailingSpaces = /\s+$/
> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
> 
> s = s.replace(/{leadingOrTrailingSpaces}/g, "");

You will need the .source attribute and make a new RegExp.

Try this [Chrome tested]:

<script type='text/javascript'>

var t = '   a b   ';

var s1 = /(^\s+)/;
var s2 = /|/;
var s3 = /(\s+$)/g;

var r = new RegExp(s1.source+s2.source+s3.source, (s3.global)?'g':'');

document.write('=' + t + '=');
document.write('<br>');
document.write('=' + t.replace(r,'') + '=');

// explanation by example:
document.write('<br><br>');
document.write(r);
document.write('<br>');
document.write(r.source);
document.write('<br>');
document.write(r.global);
document.write('<br>');
document.write(r.ignoreCase);

</script>




-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#17077

FromGene Wirchenko <genew@ocis.net>
Date2012-11-07 17:00 -0800
Message-ID<4v0m989ipopoqeqvdlj4l92jk9tonfc690@4ax.com>
In reply to#17069
On Wed, 07 Nov 2012 23:23:08 +0100, "Evertjan."
<exxjxw.hannivoort@inter.nl.net> wrote:

>Patricia Shanahan wrote on 06 nov 2012 in comp.lang.javascript:
>
>> leadingSpaces = /^\s+/
>> trailingSpaces = /\s+$/
>> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>> 
>> s = s.replace(/{leadingOrTrailingSpaces}/g, "");
>
>You will need the .source attribute and make a new RegExp.
>
>Try this [Chrome tested]:

     It ran fine with Firefox 16.0.2 on my XP box, too.

><script type='text/javascript'>
>
>var t = '   a b   ';
>
>var s1 = /(^\s+)/;
>var s2 = /|/;
>var s3 = /(\s+$)/g;
>
>var r = new RegExp(s1.source+s2.source+s3.source, (s3.global)?'g':'');

     I did not know about being able concatenate regexes.  Thank you.

[snip]

Sincerely,

Gene Wirchenko

[toc] | [prev] | [next] | [standalone]


#17079

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2012-11-08 07:16 +0100
Message-ID<XnsA1054A075E9BEeejj99@194.109.133.133>
In reply to#17077
Gene Wirchenko wrote on 08 nov 2012 in comp.lang.javascript:

>>var r = new RegExp(s1.source+s2.source+s3.source, (s3.global)?'g':'');
> 
>      I did not know about being able concatenate regexes.  Thank you.

This is concatenating of simple strings.

The nice part, imho, is the use of the regex properties:

.source text of the RegExp pattern
.global if the "g" modifier is set
.ignoreCase if the "i" modifier is set
.multiline if the "m" modifier is set

other:

.input string against which a regular expression is matched
.lastMatch the last matched characters
.lastParen last matched parenthesized substring
.valueOf returns a primitive value
.lastIndex The index at which to start the next match

There is also a .compile method, which does not seem usefull to me.

<http://www.devguru.com/technologies/ecmascript/quickref/regexp.html>
<http://www.w3schools.com/jsref/jsref_obj_regexp.asp>


-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#17072

FromDr J R Stockton <reply1245@merlyn.demon.co.uk.invalid>
Date2012-11-07 17:50 +0000
Message-ID<H5J3V4GO9pmQFw5L@invalid.uk.co.demon.merlyn.invalid>
In reply to#17054
In comp.lang.javascript message <48mi98trj215lkcssp9hq412ddggfnijed@4ax.
com>, Tue, 6 Nov 2012 10:43:52, Gene Wirchenko <genew@ocis.net> posted:

>
>     Thelackofspacesorseparatingpunctuationmakeregexeshardertoread.
>
>     If spaces could be added, it could be clearer.
>          rx = /^ \s+ | \s+ $/ g


Rather than using a RegExp literal /.../ you could use a constructor,
new RegExp("...").  Then you can PreProcess the string, using
new RegExp(PP("...")), where PP removes actual spaces from the string
and spaces wanted in the RegExp are represented as \s.

-- 
 (c) John Stockton, nr London UK               Reply address via Home Page.
   news:comp.lang.javascript FAQ <http://www.jibbering.com/faq/index.html>.
   <http://www.merlyn.demon.co.uk/js-index.htm> jscr maths, dates, sources.
   <http://www.merlyn.demon.co.uk/> TP/BP/Delphi/jscr/&c, FAQ items, links.

[toc] | [prev] | [next] | [standalone]


#17045

FromDr J R Stockton <reply1245@merlyn.demon.co.uk.invalid>
Date2012-11-05 19:26 +0000
Message-ID<ZClAbbI$LBmQFw47@invalid.uk.co.demon.merlyn.invalid>
In reply to#17034
In comp.lang.javascript message <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.
com>, Sun, 4 Nov 2012 21:52:07, Hans-Georg Michna <hans-
georgNoEmailPlease@michna.com> posted:

>On Fri, 02 Nov 2012 08:49:33 -0700, Patricia Shanahan wrote:
>
>>On 11/2/2012 6:58 AM, Hans-Georg Michna wrote:
>
>>> I sometimes see code that inefficiently loops through strings or
>>> character arrays, apparently because the programmer did not know
>>> regular expressions.
>
>>Or perhaps the programmer thought the case was not performance critical,
>>and the explicit loop code was clearer and more readable than the
>>regular expression equivalent.
>
>At least a simple regular expression is clear and very readable
>for someone who knows the language, probably much clearer and
>easier to read than some multi-line code in any high-level
>language.
>
>So it depends on the programmer who reads the code. Win some,
>lose some.
>
>s = s.replace(/^\s+|\s+$/g, "");
>
>Is that easy to read? For me it is. It comes naturally. (:-)
>This is JavaScript's trim function, by the way.

Since a literal space - Unicode 0020 - in a RegExp literal can always be
represented as \s, and similarly for tab perhaps there should be a
vari-ant form of the literal in which spaces and tabs are not
significant.  In that form, expressions could be written with spaces to
indicate their structure.  That one could be /^ \s+ | \s+ $/ for
example.  The means of expressing the vari-ance is left as an exercise
for the reader; maybe by using character "`" as the bounding delimiter.

It is easy enough to learn enough about RegExps for them to be used on
coding, though it is hard to learn enough to be able to write the
optimal RegExp for any job and to read worst-case RegExps.

-- 
 (c) John Stockton, nr London UK. Mail, see Homepage. BP7, Delphi 3 & 2006.
   <http://www.merlyn.demon.co.uk/> TP/BP/Delphi/&c., FAQqy topics & links;
   <http://www.bancoems.com/CompLangPascalDelphiMisc-MiniFAQ.htm> clpdmFAQ;
   NOT <http://support.codegear.com/newsgroups/>: news:borland.* Guidelines

[toc] | [prev] | [next] | [standalone]


#17066

FromStefan Weiss <krewecherl@gmail.com>
Date2012-11-07 22:13 +0100
Message-ID<k7eitr$b68$2@news.albasani.net>
In reply to#17045
On 2012-11-05 20:26, Dr J R Stockton wrote:
> In comp.lang.javascript message <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.
> com>, Sun, 4 Nov 2012 21:52:07, Hans-Georg Michna <hans-
> georgNoEmailPlease@michna.com> posted:
> 
>>s = s.replace(/^\s+|\s+$/g, "");
>>
>>Is that easy to read? For me it is. It comes naturally. (:-)
>>This is JavaScript's trim function, by the way.
> 
> Since a literal space - Unicode 0020 - in a RegExp literal can always be
> represented as \s, and similarly for tab perhaps there should be a
> vari-ant form of the literal in which spaces and tabs are not
> significant.  In that form, expressions could be written with spaces to
> indicate their structure.  That one could be /^ \s+ | \s+ $/ for
> example.  The means of expressing the vari-ance is left as an exercise
> for the reader; maybe by using character "`" as the bounding delimiter.

Perl has a way to do that: the x modifier, which permits embedded
whitespace and even comments. For example, a simple formal date input
check like this -

    my $regex = qr/^\s*(\d{4})-(\d\d)-(\d\d)\s*$/;

- could be written like this:

    my $regex = qr/
        ^           # start of string

        \s*         # optional whitespace

        (           # start of first capturing group ($1 = year)
            \d{4}   # exactly four digits
        )           # end first group
        -
        (           # start second capturing group ($2 = month)
            \d\d    # exactly two digits
        )           # end second group
        -
        (           # start third capturing group ($3 = day)
            \d\d    # exactly two digits
        )           # end third group

        \s*         # optional whitespace

        $           # end of string
    /x;

That's overkill for such a simple regular expression, of course, but it
can dramatically improve readability for more complex patterns.

The x modifier has been proposed as a new feature for a future release
of the ECMAScript standard, but AFAIK there has been little progress on
that front lately.
http://wiki.ecmascript.org/doku.php?id=proposals:extend_regexps#x_flag

PCRE-based implementations usually also support the x modifier (for
example, PHP does). From what I read in this group, Thomas Lahn's
regexp.js apparently also supports it.


- stefan

[toc] | [prev] | [next] | [standalone]


#17018

FromGene Wirchenko <genew@ocis.net>
Date2012-11-02 09:51 -0700
Message-ID<e7u798528fq4tt4rs4jtambmr00b30rot2@4ax.com>
In reply to#17016
On Fri, 02 Nov 2012 14:58:38 +0100, Hans-Georg Michna
<hans-georgNoEmailPlease@michna.com> wrote:

>On Fri, 02 Nov 2012 06:25:08 -0700, Patricia Shanahan wrote:
>
>>If I'm prepared to look at it long enough, I can generally work out what
>>a Regex does, but Regex does not seem to me to be a very human-friendly,
>>smoothly readable language.
>
>I agree, but in my experience regular expressions are too
>frequently the best solution to string processing problems to
>ignore them.

     Agreed, but I add the proviso that regexes can get rather
complicated if you try to do a lot with them in one regex.  I have
mixed small regexes with other processing code, and that works nicely.

>I sometimes see code that inefficiently loops through strings or
>character arrays, apparently because the programmer did not know
>regular expressions.

     I will sometimes do this if the code is likely to change or is
not suited to regexes.  Non-trivial regexes are rather write-only.

>My personal advice to every JavaScript (and other) programmer is
>to learn at least the most basic parts of the regular expression
>language.

     Quite.  Regexes are a very useful tool.

Sincerely,

Gene Wirchenko

[toc] | [prev] | [next] | [standalone]


#17020

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-02 18:41 +0100
Message-ID<1805042.BzeuYpbCYS@PointedEars.de>
In reply to#17014
Patricia Shanahan wrote:

> On 11/2/2012 2:22 AM, Evertjan. wrote:
> ...
>> if ( /foo/.test(someString) ) {..}
>>
>> I trust not understanding Regex is not a valid counterargument.
> 
> If I'm prepared to look at it long enough, I can generally work out what
> a Regex does, but Regex does not seem to me to be a very human-friendly,
> smoothly readable language.

That depends on the flavor and your experience with them.
 
> For this particular case, it looks OK if the probe really is a literal
> such as "foo".

In exactly that case using a Regular Expression is overkill.  You need to 
consider that for every RegExp literal in the code (and per ES 5.x in every 
use of that literal), a new RegExp instance is being created.

> Could you show me the code you would use for this if the probe were an
> actual parameter or variable with unknown contents?
> 
> That is, what would you use to replace:
> 
> if(someString.indexOf(someOtherString)) != -1)

It is possible with

  if ((new RegExp(someOtherString)).test(someString))

but I would not use that.  However, the following can be useful and 
necessary:

  if ((new RegExp(someOtherString, "i")).test(someString))

The caveat in both cases is that any special characters in the value of 
`someOtherString' that should not assume their special meaning need to be 
escaped.  This can be accomplished with a previous replace() –

  if ((new RegExp(someOtherString.replace(/…/g), "i")).test(someString))

– or with a user-defined method:

  if ((new RegExp(jsx.regexp.escape(someOtherString), "i"))
      .test(someString))

One could also augment the ECMAScript String prototype object:

  if ((new RegExp(someOtherString.regexpEscape(), "i"))
      .test(someString))
  
See jsx.dom.css.addClassName() for a working example of a variable RegExp:

<http://PointedEars.de/websvn/filedetails.php?repname=JSX&path=%2Ftrunk%2Fdom%2Fcss.js>


PointedEars
-- 
Danny Goodman's books are out of date and teach practices that are
positively harmful for cross-browser scripting.
  -- Richard Cornford, cljs, <cife6q$253$1$8300dec7@news.demon.co.uk> (2004)

[toc] | [prev] | [next] | [standalone]


#17021

FromPatricia Shanahan <pats@acm.org>
Date2012-11-02 12:30 -0700
Message-ID<7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>
In reply to#17020
On 11/2/2012 10:41 AM, Thomas 'PointedEars' Lahn wrote:
> Patricia Shanahan wrote:
>
>> On 11/2/2012 2:22 AM, Evertjan. wrote:
>> ...
>>> if ( /foo/.test(someString) ) {..}
>>>
>>> I trust not understanding Regex is not a valid counterargument.
>>
>> If I'm prepared to look at it long enough, I can generally work out what
>> a Regex does, but Regex does not seem to me to be a very human-friendly,
>> smoothly readable language.
>
> That depends on the flavor and your experience with them.

Well, I don't quite have 30 years of experience with regular expressions
yet, but it's getting close.

>
>> For this particular case, it looks OK if the probe really is a literal
>> such as "foo".
>
> In exactly that case using a Regular Expression is overkill.  You need to
> consider that for every RegExp literal in the code (and per ES 5.x in every
> use of that literal), a new RegExp instance is being created.
>
>> Could you show me the code you would use for this if the probe were an
>> actual parameter or variable with unknown contents?
>>
>> That is, what would you use to replace:
>>
>> if(someString.indexOf(someOtherString)) != -1)
>
> It is possible with
>
>    if ((new RegExp(someOtherString)).test(someString))
>
> but I would not use that.  However, the following can be useful and
> necessary:
>
>    if ((new RegExp(someOtherString, "i")).test(someString))
>
> The caveat in both cases is that any special characters in the value of
> `someOtherString' that should not assume their special meaning need to be
> escaped.  This can be accomplished with a previous replace() –
>
>    if ((new RegExp(someOtherString.replace(/…/g), "i")).test(someString))
>
> – or with a user-defined method:
>
>    if ((new RegExp(jsx.regexp.escape(someOtherString), "i"))
>        .test(someString))
...

Do you feel this is an improvement on basing the contains test on the
String indexOf function? If so, why? Or is this just an illustration of
how you would do it with regular expressions if indexOf did not exist?

My general view of regular expressions is that they have a few uses, but
I've often seen code, in many languages, that seems to use them for the
sake of using them, at the expense of making the code less readable.

Patricia

[toc] | [prev] | [next] | [standalone]


#17022

FromTim Streater <timstreater@greenbee.net>
Date2012-11-02 20:33 +0000
Message-ID<timstreater-D8955D.20331302112012@news.individual.net>
In reply to#17021
In article <7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>,
 Patricia Shanahan <pats@acm.org> wrote:

> My general view of regular expressions is that they have a few uses, but
> I've often seen code, in many languages, that seems to use them for the
> sake of using them, at the expense of making the code less readable.

This is code written by smartarses who like to show how clever they are.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17035

FromHans-Georg Michna <hans-georgNoEmailPlease@michna.com>
Date2012-11-04 21:57 +0100
Message-ID<ahld981ren55qc0kssq621554fkh2nc20c@4ax.com>
In reply to#17022
On Fri, 02 Nov 2012 20:33:14 +0000, Tim Streater wrote:

>In article <7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>,
> Patricia Shanahan <pats@acm.org> wrote:

>> My general view of regular expressions is that they have a few uses, but
>> I've often seen code, in many languages, that seems to use them for the
>> sake of using them, at the expense of making the code less readable.

>This is code written by smartarses who like to show how clever they are.

But what if you happen to find some complex RegExp code in a
program that was not meant to be seen by anybody else but its
original programmer?

The problem in this discussion is that a regular expression can
look strange and unintelligible to somebody who does not master
the language, but can at the same time look simple,
straightforward, and efficient to somebody who does.

Hans-Georg

[toc] | [prev] | [next] | [standalone]


#17036

FromTim Streater <timstreater@greenbee.net>
Date2012-11-04 22:44 +0000
Message-ID<timstreater-579A11.22441204112012@news.individual.net>
In reply to#17035
In article <ahld981ren55qc0kssq621554fkh2nc20c@4ax.com>,
 Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote:

> On Fri, 02 Nov 2012 20:33:14 +0000, Tim Streater wrote:
> 
> >In article <7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>,
> > Patricia Shanahan <pats@acm.org> wrote:
> 
> >> My general view of regular expressions is that they have a few uses, but
> >> I've often seen code, in many languages, that seems to use them for the
> >> sake of using them, at the expense of making the code less readable.
> 
> >This is code written by smartarses who like to show how clever they are.
> 
> But what if you happen to find some complex RegExp code in a
> program that was not meant to be seen by anybody else but its
> original programmer?

There is no such program - in practice, at any rate.

> The problem in this discussion is that a regular expression can
> look strange and unintelligible to somebody who does not master
> the language, but can at the same time look simple,
> straightforward, and efficient to somebody who does.

Why should I "master" it? I ignore edlin, TECO, emacs, make, and many 
others, for much the same reasons.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17039

FromAndrew Poulos <ap_prog@hotmail.com>
Date2012-11-05 15:41 +1100
Message-ID<f5-dnXsJgokF3grNnZ2dnUVZ_q-dnZ2d@westnet.com.au>
In reply to#17036
On 5/11/2012 9:44 AM, Tim Streater wrote:
> In article <ahld981ren55qc0kssq621554fkh2nc20c@4ax.com>,
> Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote:
>
>> On Fri, 02 Nov 2012 20:33:14 +0000, Tim Streater wrote:
>>
>> >In article <7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>,
>> > Patricia Shanahan <pats@acm.org> wrote:
>>
>> >> My general view of regular expressions is that they have a few
>> uses, but
>> >> I've often seen code, in many languages, that seems to use them for
>> the
>> >> sake of using them, at the expense of making the code less readable.
>>
>> >This is code written by smartarses who like to show how clever they are.
>>
>> But what if you happen to find some complex RegExp code in a
>> program that was not meant to be seen by anybody else but its
>> original programmer?
>
> There is no such program - in practice, at any rate.
>
>> The problem in this discussion is that a regular expression can
>> look strange and unintelligible to somebody who does not master
>> the language, but can at the same time look simple,
>> straightforward, and efficient to somebody who does.
>
> Why should I "master" it? I ignore edlin, TECO, emacs, make, and many
> others, for much the same reasons.

Should one post a resume listing all the things they've not bothered to 
learn as, by the sounds of it, there must be lots of companies searching 
for people without just the right qualifications?

Andrew Poulos

[toc] | [prev] | [next] | [standalone]


#17042

FromTim Streater <timstreater@greenbee.net>
Date2012-11-05 09:07 +0000
Message-ID<timstreater-E59DAF.09075505112012@news.individual.net>
In reply to#17039
In article <f5-dnXsJgokF3grNnZ2dnUVZ_q-dnZ2d@westnet.com.au>,
 Andrew Poulos <ap_prog@hotmail.com> wrote:

> On 5/11/2012 9:44 AM, Tim Streater wrote:
> > In article <ahld981ren55qc0kssq621554fkh2nc20c@4ax.com>,
> > Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote:
> >
> >> On Fri, 02 Nov 2012 20:33:14 +0000, Tim Streater wrote:
> >>
> >> >In article <7MidnYXvqeq8ggnNnZ2dnUVZ_uGdnZ2d@earthlink.com>,
> >> > Patricia Shanahan <pats@acm.org> wrote:
> >>
> >> >> My general view of regular expressions is that they have a few
> >> uses, but
> >> >> I've often seen code, in many languages, that seems to use them for
> >> the
> >> >> sake of using them, at the expense of making the code less readable.
> >>
> >> >This is code written by smartarses who like to show how clever they are.
> >>
> >> But what if you happen to find some complex RegExp code in a
> >> program that was not meant to be seen by anybody else but its
> >> original programmer?
> >
> > There is no such program - in practice, at any rate.
> >
> >> The problem in this discussion is that a regular expression can
> >> look strange and unintelligible to somebody who does not master
> >> the language, but can at the same time look simple,
> >> straightforward, and efficient to somebody who does.
> >
> > Why should I "master" it? I ignore edlin, TECO, emacs, make, and many
> > others, for much the same reasons.
> 
> Should one post a resume listing all the things they've not bothered to 
> learn as, by the sounds of it, there must be lots of companies searching 
> for people without just the right qualifications?

Depends what you focus on, really.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17023

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-02 22:10 +0100
Message-ID<1466051.Cd1scNln3u@PointedEars.de>
In reply to#17021
Patricia Shanahan wrote:

> On 11/2/2012 10:41 AM, Thomas 'PointedEars' Lahn wrote:
>> Patricia Shanahan wrote:
>>> On 11/2/2012 2:22 AM, Evertjan. wrote:
>>> ...
>>>> if ( /foo/.test(someString) ) {..}
>>>>
>>>> I trust not understanding Regex is not a valid counterargument.
>>>
>>> If I'm prepared to look at it long enough, I can generally work out what
>>> a Regex does, but Regex does not seem to me to be a very human-friendly,
>>> smoothly readable language.
>>
>> That depends on the flavor and your experience with them.
> 
> Well, I don't quite have 30 years of experience with regular expressions
> yet, but it's getting close.

Which flavors?

>>> For this particular case, it looks OK if the probe really is a literal
>>> such as "foo".
>>
>> In exactly that case using a Regular Expression is overkill.  You need to
>> consider that for every RegExp literal in the code (and per ES 5.x in
>> every use of that literal), a new RegExp instance is being created.
>>
>>> Could you show me the code you would use for this if the probe were an
>>> actual parameter or variable with unknown contents?
>>>
>>> That is, what would you use to replace:
>>>
>>> if(someString.indexOf(someOtherString)) != -1)
>>
>> It is possible with
>>
>>    if ((new RegExp(someOtherString)).test(someString))
>>
>> but I would not use that.  However, the following can be useful and
>> necessary:
>>
>>    if ((new RegExp(someOtherString, "i")).test(someString))
>>
>> The caveat in both cases is that any special characters in the value of
>> `someOtherString' that should not assume their special meaning need to be
>> escaped.  This can be accomplished with a previous replace() –
>>
>>    if ((new RegExp(someOtherString.replace(/…/g), "i")).test(someString))
>>
>> – or with a user-defined method:
>>
>>    if ((new RegExp(jsx.regexp.escape(someOtherString), "i"))
>>        .test(someString))
> ...
> 
> Do you feel this is an improvement on basing the contains test on the
> String indexOf function?

The latter variants certainly are.

> If so, why?

Case-insensitive matching.

> Or is this just an illustration of how you would do it with regular
> expressions if indexOf did not exist?

Yes.
 
> My general view of regular expressions is that they have a few uses, but
> I've often seen code, in many languages, that seems to use them for the
> sake of using them, at the expense of making the code less readable.

The longer and better I know them, the more valid reasons I find to use 
them; not only *in* source code, but also *for writing* source code.  AISB, 
the readability of (code using) regular expressions depends very much on the 
flavor of regular expressions that is used (that the language/IDE is capable 
of) and the way you (can) write them in a programming language/IDE.

Because RegExp *literals* need to be delimited with `/' in ECMAScript 
implementations, some people think they need to escape `/' in string 
literals passed to the RegExp constructor (or elsewhere), or all special 
characters in character classes; actually they do not.  Attempting to 
simplify a regular expression often goes a long way towards better 
understanding them, and vice-versa.  You can find a lot of examples of that 
in my follow-ups in comp.lang.ALL and de.comp.lang.ALL.

In that sense I have devised RegExp.prototype.concat() (or 
jsx.regexp.concat()) so that a regular expression definition can span more 
than one line, and can have variable parts, *without* the need to use a 
string literal throughout (and the extra escaping that is required then) – 
or to wait for TC39 and implementors to catch up;

  var s = "baz";
  var rx =
    /foo/
    .concat(/bar/i)
    .concat(s)
    .concat(/bla/gm);

being equivalent to

  var rx = /foobarbazbla/gim;

And I had devised (independently of similar approaches) jsx.regexp.RegExp() 
so that you can use at least some features of Perl and Perl-Compatible 
Regular Expressions in ECMAScript implementations, many of which make 
regular expressions more powerful, some of which make them better readable 
and understandable.  For example,

  var rx = jsx.regexp.RegExp(
    /\s+ \w+ \s+  # Word delimited by whitespace/, "x");

is equivalent to

  var rx = /\s+\w+\s+/;

(Existing approaches are informing the implementation of new features now.)

See also: <http://PointedEars.de/scripts/test/regexp>


PointedEars
-- 
Danny Goodman's books are out of date and teach practices that are
positively harmful for cross-browser scripting.
  -- Richard Cornford, cljs, <cife6q$253$1$8300dec7@news.demon.co.uk> (2004)

[toc] | [prev] | [next] | [standalone]


#17024

FromPatricia Shanahan <pats@acm.org>
Date2012-11-02 15:22 -0700
Message-ID<IcidnRJ5-sYx2gnNnZ2dnUVZ_hKdnZ2d@earthlink.com>
In reply to#17023
On 11/2/2012 2:10 PM, Thomas 'PointedEars' Lahn wrote:
> Patricia Shanahan wrote:
>
>> On 11/2/2012 10:41 AM, Thomas 'PointedEars' Lahn wrote:
>>> Patricia Shanahan wrote:
>>>> On 11/2/2012 2:22 AM, Evertjan. wrote:
>>>> ...
>>>>> if ( /foo/.test(someString) ) {..}
>>>>>
>>>>> I trust not understanding Regex is not a valid counterargument.
>>>>
>>>> If I'm prepared to look at it long enough, I can generally work out what
>>>> a Regex does, but Regex does not seem to me to be a very human-friendly,
>>>> smoothly readable language.
>>>
>>> That depends on the flavor and your experience with them.
>>
>> Well, I don't quite have 30 years of experience with regular expressions
>> yet, but it's getting close.
>
> Which flavors?

Assorted UNIX shell tools (lex, grep, egrep, vi, ed, sed, awk etc.),
Perl, PHP, and Java.

Patricia

[toc] | [prev] | [next] | [standalone]


#17026

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-03 10:19 +0100
Message-ID<6563475.n3XdYolgyi@PointedEars.de>
In reply to#17024
[I had written a longer reply with examples, but apparently it has not 
reached Usenet.  So I will keep this one shorter and expand on it upon 
request.]

Patricia Shanahan wrote:

> On 11/2/2012 2:10 PM, Thomas 'PointedEars' Lahn wrote:
>> Patricia Shanahan wrote:
>>> On 11/2/2012 10:41 AM, Thomas 'PointedEars' Lahn wrote:
>>>> Patricia Shanahan wrote:
>>>>> On 11/2/2012 2:22 AM, Evertjan. wrote:
>>>>> ...
>>>>>> if ( /foo/.test(someString) ) {..}
>>>>>>
>>>>>> I trust not understanding Regex is not a valid counterargument.
>>>>> If I'm prepared to look at it long enough, I can generally work out
>>>>> what a Regex does, but Regex does not seem to me to be a very
>>>>> human-friendly, smoothly readable language.
>>>> That depends on the flavor and your experience with them.
>>> Well, I don't quite have 30 years of experience with regular expressions
>>> yet, but it's getting close.
>> Which flavors?
> 
> Assorted UNIX shell tools (lex, grep, egrep, vi, ed, sed, awk etc.),
> Perl, PHP, and Java.

As I see it, implementations of regular expressions are lacking readability 
because of two factors: insufficient expressiveness and stringly typing.

Insufficient expressiveness or excessive verbosity means you have to write 
longer expressions for rather simply concepts.  Consider, for example, POSIX 
`[[:space:]]' vs. Perl/PCRE `\s'.  Consider POSIX Basic Regular Expressions 
(BRE) `\{0,1\}' vs. POSIX Extended Regular Expressions (ERE)'s and 
Perl's/Perl-Compatible Regular Expressions (PCRE)'s `?'.  Consider POSIX BRE 
`\{1,\}' vs. ERE's/Perl's/PCRE's `+'.

Stringly typing, i. e. expressing data in string literals instead of in 
regular expression literals, causes regular expressions to be less readable, 
and in turn code that uses regular expressions to be less readable, due to 
the fact that string literals have their own escaping mechanism.  

I have not done much lex.  But as you probably know, POSIX grep, vi(m), ed, 
sed, awk & friends only support POSIX Basic Regular Expressions.  BRE both 
lack expressiveness and are stringly typed in direct use.  So they cannot be 
shining examples of readable regular expressions.

ERE as supported by POSIX egrep and GNU grep improve on that slightly by 
reversing expression logic (so you have to escape what you do *not* want to 
be special instead), by adding shortcuts for `{0,1}' and `{1,}' and 
alternation.  But they still have the excessive verbosity and stringly 
typing problem of POSIX REs.

Perl RE and PCRE (the latter is supported by GNU grep) improve on 
readability again by providing regular expression literals with and, among 
other powerful features such as interpolation, flags to improve readability 
specifically (such as `/x').  But you have to be aware of that.

While PHP eventually supports PCRE and allows a wider range of delimiters, 
its RE implementation suffers from the fact that there are no regular 
expression literals.  It steps back to stringly typing.  So does Java, but 
Java essentially allows only one delimiter and no multi-line syntax; another 
two steps back.  Java's RE implementation, in addition to that, suffers from 
Java's static typing and the excessive verbosity in implementation that 
brings.  So both of them do not provide good examples either.

You have not mentioned Python.  Python's implementation is a step back from 
PHP and Java in that it does not support PCRE, but its own PCRE-inspired 
flavor.  However, it is a step forward from that because of Python's 
implicit string concatenation, raw-strings to alleviate the escaping 
problem, and new RE features.

By comparison, ECMAScript and its implementations have regular expression 
literals, but those have no variable delimiter, no interpolation, no multi-
line syntax, and they do not support PCRE but their own flavor yet again.  
But on the plus side the languages are dynamic enough so that you can work 
around those shortcomings (as I did in JSX:regexp.js).

(There are also other RE implementations, like that of Microsoft, which 
primarily suffer from the fact that they are very different to the common 
aforementioned ones, or have limited use.  If you ever tried to use RE to 
search with Visual Studio, you know what I mean.)

Finally, as a third factor to readability of code using regular expressions, 
here comes in the person responsible for writing it.  Many people use 
regular expressions where they are not strictly necessary.  And they write 
needlessly complicated regular expressions because they do not know better 
or do not care.  Specifically for ECMAScript implementations (but you can 
also find similar bloat code elsewhere), people escape `/' in *string* 
literals passed to the RegExp constructor because `/' is the delimiter of 
*RegExp* literals.  They needlessly escape special characters in character 
classes.  They use alternation where character classes would have sufficed.  
And so on.

Attempting to avoid regular expressions where they are not necessary, and 
simplifying regular expressions where they are, thereby improving 
readability, goes a long way towards understanding them, and vice-versa.  
You can find examples of that in many of my follow-ups in comp.lang.ALL and 
de.comp.lang.ALL.


PointedEars
-- 
When all you know is jQuery, every problem looks $(olvable).

[toc] | [prev] | [next] | [standalone]


#17029

FromPatricia Shanahan <pats@acm.org>
Date2012-11-03 13:05 -0700
Message-ID<0aWdnSkO4e2z5AjNnZ2dnUVZ_oidnZ2d@earthlink.com>
In reply to#17026
On 11/3/2012 2:19 AM, Thomas 'PointedEars' Lahn wrote:
...
> Insufficient expressiveness or excessive verbosity means you have to write
> longer expressions for rather simply concepts.  Consider, for example, POSIX
> `[[:space:]]' vs. Perl/PCRE `\s'.  Consider POSIX Basic Regular Expressions
> (BRE) `\{0,1\}' vs. POSIX Extended Regular Expressions (ERE)'s and
> Perl's/Perl-Compatible Regular Expressions (PCRE)'s `?'.  Consider POSIX BRE
> `\{1,\}' vs. ERE's/Perl's/PCRE's `+'.
...

Why do you prefer terse expressions? To me, `[[:space:]]' says what it
means far more clearly than `\s'.

Patricia

[toc] | [prev] | [next] | [standalone]


#17031

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-03 21:39 +0100
Message-ID<1469555.Cl2dHhCXWe@PointedEars.de>
In reply to#17029
Patricia Shanahan wrote:

> On 11/3/2012 2:19 AM, Thomas 'PointedEars' Lahn wrote:
>> Insufficient expressiveness or excessive verbosity means you have to
>> write longer expressions for rather simply concepts.  Consider, for
                                            ^e
>> example, POSIX `[[:space:]]' vs. Perl/PCRE `\s'.  Consider POSIX Basic
>> Regular Expressions (BRE) `\{0,1\}' vs. POSIX Extended Regular
>> Expressions (ERE)'s and Perl's/Perl-Compatible Regular Expressions
>> (PCRE)'s `?'.  Consider POSIX BRE `\{1,\}' vs. ERE's/Perl's/PCRE's `+'.
>> [...]
> 
> Why do you prefer terse expressions? To me, `[[:space:]]' says what it
> means far more clearly than `\s'.

I do prefer

  m{\s?(Captain's\s+Log)\s?}

over

  '[[:space:]]\{0,1\}\(Captain'\''s[[:space:]]\{1,\}Log\)[[:space:]]\{0,1\}'

YMMV.


PointedEars
-- 
Use any version of Microsoft Frontpage to create your site.
(This won't prevent people from viewing your source, but no one
will want to steal it.)
  -- from <http://www.vortex-webdesign.com/help/hidesource.htm> (404-comp.)

[toc] | [prev] | [next] | [standalone]


#17027

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-03 12:26 +0100
Message-ID<1795049.zBDcGp28eQ@PointedEars.de>
In reply to#17020
Thomas 'PointedEars' Lahn wrote:

> […] However, the following can be useful and necessary:
> 
>   if ((new RegExp(someOtherString, "i")).test(someString))
> 
> The caveat in both cases is that any special characters in the value of
> `someOtherString' that should not assume their special meaning need to be
> escaped.  This can be accomplished with a previous replace() –
> 
>   if ((new RegExp(someOtherString.replace(/…/g), "i")).test(someString))

Should be

  if ((new RegExp(someOtherString.replace(/[…]/g, "\\$&"), "i"))
      .test(someString))
 
where `…' specifies the (empty-string delimited) list of (special) 
characters to be escaped (escaped themselves if necessary).  The later 
mentioned user-defined methods are preferable to that, of course.

> […]


PointedEars
-- 
Anyone who slaps a 'this page is best viewed with Browser X' label on
a Web page appears to be yearning for the bad old days, before the Web,
when you had very little chance of reading a document written on another
computer, another word processor, or another network. -- Tim Berners-Lee

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.lang.javascript


csiph-web