Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.javascript > #17006 > unrolled thread

Good use of bitwise NOT "~" or unnecessary obfuscation?

Started byRobG <rgqld@iinet.net.au>
First post2012-11-01 18:32 -0700
Last post2012-11-07 11:22 -0800
Articles 20 on this page of 46 — 14 participants

Back to article view | Back to comp.lang.javascript


Contents

  Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-01 18:32 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Denis McMahon <denismfmcmahon@gmail.com> - 2012-11-02 06:14 +0000
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 02:28 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-02 10:22 +0100
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 10:37 +0100
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 06:25 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:35 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-02 14:58 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 08:49 -0700
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:52 +0100
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:49 +0000
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-05 13:18 +0100
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:48 -0800
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-06 15:04 +0000
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-06 10:43 -0800
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 11:22 -0800
                      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-06 12:16 -0800
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 12:53 -0800
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-07 17:34 +0100
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-07 23:21 +0100
                        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-07 23:23 +0100
                          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-07 17:00 -0800
                            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-08 07:16 +0100
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-07 17:50 +0000
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-05 19:26 +0000
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Stefan Weiss <krewecherl@gmail.com> - 2012-11-07 22:13 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 09:51 -0700
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 18:41 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 12:30 -0700
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 20:33 +0000
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:57 +0100
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:44 +0000
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Andrew Poulos <ap_prog@hotmail.com> - 2012-11-05 15:41 +1100
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-05 09:07 +0000
            Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 22:10 +0100
              Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 15:22 -0700
                Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 10:19 +0100
                  Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-03 13:05 -0700
                    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 21:39 +0100
          Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 12:26 +0100
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Jukka K. Korpela" <jkorpela@cs.tut.fi> - 2012-11-02 12:28 +0200
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 10:00 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:15 -0700
    Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:32 -0800
      Re: Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-06 18:23 -0800
        Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-07 11:22 -0800

Page 1 of 3  [1] 2 3  Next page →


#17006 — Good use of bitwise NOT "~" or unnecessary obfuscation?

FromRobG <rgqld@iinet.net.au>
Date2012-11-01 18:32 -0700
SubjectGood use of bitwise NOT "~" or unnecessary obfuscation?
Message-ID<50330d14-3a30-4885-9833-18c75717a04d@googlegroups.com>
I frequently see:

    if (someString.indexOf('foo') != -1)

but it looks ugly and isn't all that intuitive. Adding one to the result means if the string isn't found, the result is zero or +ve if it is found, so:

    if (someString.indexOf('foo') + 1)

but that doesn't look much better. But the bitwise NOT comes to the rescue:

    if (~someString.indexOf('foo'))

now if the string isn't found, the result is 0 and some positive integer if it is. And I didn't have to create a regular expression to get similar semantics.

Can anyone see issues with the above? Other than perhaps some not seeing the bitwise NOT operator "~" so misunderstanding the test?


-- 
Rob

[toc] | [next] | [standalone]


#17007

FromDenis McMahon <denismfmcmahon@gmail.com>
Date2012-11-02 06:14 +0000
Message-ID<k6vobo$216$1@dont-email.me>
In reply to#17006
On Thu, 01 Nov 2012 18:32:59 -0700, RobG wrote:

>     if (~someString.indexOf('foo'))
>
> Can anyone see issues with the above? Other than perhaps some not seeing
> the bitwise NOT operator "~" so misunderstanding the test?

You're making assumptions about the bitwise representation of -1 and 0 
that might be implementation dependant.

I can't think of any reason why anyone would write a javascript 
implementation where two's complement wasn't used to store integers, but 
unless there's a specification somewhere that mandates two's complement 
for integers, then you're taking a risk based on an assumption about how 
everyone implements integer math that might not always be valid.

Bitfields and integers, although appearing to be the the same, may not 
always behave in the same way. In this case, I'd say that just because in 
every system that you (and I) am familiar with this would be ok, doesn't 
mean that there isn't a platform out there somewhere, now or in the 
future, where it won't work.

Rgds

Denis McMahon

[toc] | [prev] | [next] | [standalone]


#17010

FromPatricia Shanahan <pats@acm.org>
Date2012-11-02 02:28 -0700
Message-ID<KvGdnVxSk7C8Dw7NnZ2dnUVZ_j-dnZ2d@earthlink.com>
In reply to#17007
On 11/1/2012 11:14 PM, Denis McMahon wrote:
> On Thu, 01 Nov 2012 18:32:59 -0700, RobG wrote:
>
>>      if (~someString.indexOf('foo'))
>>
>> Can anyone see issues with the above? Other than perhaps some not seeing
>> the bitwise NOT operator "~" so misunderstanding the test?
>
> You're making assumptions about the bitwise representation of -1 and 0
> that might be implementation dependant.
>
> I can't think of any reason why anyone would write a javascript
> implementation where two's complement wasn't used to store integers, but
> unless there's a specification somewhere that mandates two's complement
> for integers, then you're taking a risk based on an assumption about how
> everyone implements integer math that might not always be valid.
>
> Bitfields and integers, although appearing to be the the same, may not
> always behave in the same way. In this case, I'd say that just because in
> every system that you (and I) am familiar with this would be ok, doesn't
> mean that there isn't a platform out there somewhere, now or in the
> future, where it won't work.

The use of 32 bit 2's complement for bitwise NOT is required by ECMA-262.

Because of the wrap-around in the 32 bit arithmetic, using **
for exponent, string indexOf results 2**32-1, 2**33-1, 2**34-1, ...,
would also map to bit pattern ffffffff on the required conversion.

I don't know how practical implementations deal with attempts to
construct strings length 2**32 and longer.

On the question in the subject, is there any way to understand
"~someString.indexOf('foo')" without mentally converting it to
"someString.indexOf('foo') != -1 [mod 2**32]"? If not, I think it is
unnecessary obfuscation.

Patricia

[toc] | [prev] | [next] | [standalone]


#17009

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2012-11-02 10:22 +0100
Message-ID<XnsA0FF6988CAC95eejj99@194.109.133.133>
In reply to#17006
RobG wrote on 02 nov 2012 in comp.lang.javascript:

> I frequently see:
> 
>     if (someString.indexOf('foo') != -1)
> 
> but it looks ugly and isn't all that intuitive. 

I don't know about your personal taste of beauty and your intuition,
I doubt these can be generalized.

The above is far easier [my subjectiveness]
and more versatile [easily shown if you like] met by using Regex:

if ( /foo/.test(someString) ) {..}

I trust not understanding Regex is not a valid counterargument.

-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#17012

FromTim Streater <timstreater@greenbee.net>
Date2012-11-02 10:37 +0100
Message-ID<timstreater-3F119F.10370502112012@news.individual.net>
In reply to#17009
In article <XnsA0FF6988CAC95eejj99@194.109.133.133>,
 "Evertjan." <exxjxw.hannivoort@inter.nl.net> wrote:

> RobG wrote on 02 nov 2012 in comp.lang.javascript:
> 
> > I frequently see:
> > 
> >     if (someString.indexOf('foo') != -1)
> > 
> > but it looks ugly and isn't all that intuitive. 
> 
> I don't know about your personal taste of beauty and your intuition,
> I doubt these can be generalized.
> 
> The above is far easier [my subjectiveness]
> and more versatile [easily shown if you like] met by using Regex:
> 
> if ( /foo/.test(someString) ) {..}
> 
> I trust not understanding Regex is not a valid counterargument.

Oh but it is. Regex, in general, is just line-noise.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17014

FromPatricia Shanahan <pats@acm.org>
Date2012-11-02 06:25 -0700
Message-ID<BMSdnRSFG8QIVA7NnZ2dnUVZ_vGdnZ2d@earthlink.com>
In reply to#17009
On 11/2/2012 2:22 AM, Evertjan. wrote:
...
> if ( /foo/.test(someString) ) {..}
>
> I trust not understanding Regex is not a valid counterargument.
>

If I'm prepared to look at it long enough, I can generally work out what
a Regex does, but Regex does not seem to me to be a very human-friendly,
smoothly readable language.

For this particular case, it looks OK if the probe really is a literal
such as "foo". Could you show me the code you would use for this if the
probe were an actual parameter or variable with unknown contents?

That is, what would you use to replace:

if(someString.indexOf(someOtherString)) != -1)

Thanks,

Patricia

[toc] | [prev] | [next] | [standalone]


#17015

FromAsen Bozhilov <asen.bozhilov@gmail.com>
Date2012-11-02 06:35 -0700
Message-ID<a0445966-aad0-4a66-b281-21db9430956f@q4g2000vbg.googlegroups.com>
In reply to#17014
Patricia Shanahan wrote:

> For this particular case, it looks OK if the probe really is a literal
> such as "foo". Could you show me the code you would use for this if the
> probe were an actual parameter or variable with unknown contents?

Right. RegExp solution would work with literal values. With runtime
evaluating values, it should need compilation of RegExp e.g.

new RegExp(value);

Which is inappropriate at that case.

> That is, what would you use to replace:
>
> if(someString.indexOf(someOtherString)) != -1)

Typo

[toc] | [prev] | [next] | [standalone]


#17016

FromHans-Georg Michna <hans-georgNoEmailPlease@michna.com>
Date2012-11-02 14:58 +0100
Message-ID<f7k798117mcgu610jd2i8ca98s770f9uc5@4ax.com>
In reply to#17014
On Fri, 02 Nov 2012 06:25:08 -0700, Patricia Shanahan wrote:

>If I'm prepared to look at it long enough, I can generally work out what
>a Regex does, but Regex does not seem to me to be a very human-friendly,
>smoothly readable language.

I agree, but in my experience regular expressions are too
frequently the best solution to string processing problems to
ignore them.

I sometimes see code that inefficiently loops through strings or
character arrays, apparently because the programmer did not know
regular expressions.

My personal advice to every JavaScript (and other) programmer is
to learn at least the most basic parts of the regular expression
language.

Hans-Georg

[toc] | [prev] | [next] | [standalone]


#17017

FromPatricia Shanahan <pats@acm.org>
Date2012-11-02 08:49 -0700
Message-ID<Ceidnfe5X9M_dg7NnZ2dnUVZ_hWdnZ2d@earthlink.com>
In reply to#17016
On 11/2/2012 6:58 AM, Hans-Georg Michna wrote:
> On Fri, 02 Nov 2012 06:25:08 -0700, Patricia Shanahan wrote:
>
>> If I'm prepared to look at it long enough, I can generally work out what
>> a Regex does, but Regex does not seem to me to be a very human-friendly,
>> smoothly readable language.
>
> I agree, but in my experience regular expressions are too
> frequently the best solution to string processing problems to
> ignore them.

I agree with not ignoring them, but also think there is a tendency to
over-use them, making code much more complicated, and less readable,
than it could be.

Using regular expression matching to do a simple string contains seems
to me to be an example. At a minimum, the reader has to check the probe
string for RegExp special characters.

>
> I sometimes see code that inefficiently loops through strings or
> character arrays, apparently because the programmer did not know
> regular expressions.

Or perhaps the programmer thought the case was not performance critical,
and the explicit loop code was clearer and more readable than the
regular expression equivalent.

>
> My personal advice to every JavaScript (and other) programmer is
> to learn at least the most basic parts of the regular expression
> language.

Been there, done that - in 1983 when I was learning ed, vi, grep, awk etc.

Patricia

[toc] | [prev] | [next] | [standalone]


#17034

FromHans-Georg Michna <hans-georgNoEmailPlease@michna.com>
Date2012-11-04 21:52 +0100
Message-ID<32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com>
In reply to#17017
On Fri, 02 Nov 2012 08:49:33 -0700, Patricia Shanahan wrote:

>On 11/2/2012 6:58 AM, Hans-Georg Michna wrote:

>> I sometimes see code that inefficiently loops through strings or
>> character arrays, apparently because the programmer did not know
>> regular expressions.

>Or perhaps the programmer thought the case was not performance critical,
>and the explicit loop code was clearer and more readable than the
>regular expression equivalent.

At least a simple regular expression is clear and very readable
for someone who knows the language, probably much clearer and
easier to read than some multi-line code in any high-level
language.

So it depends on the programmer who reads the code. Win some,
lose some.

s = s.replace(/^\s+|\s+$/g, "");

Is that easy to read? For me it is. It comes naturally. (:-)
This is JavaScript's trim function, by the way.

Hans-Georg

[toc] | [prev] | [next] | [standalone]


#17037

FromTim Streater <timstreater@greenbee.net>
Date2012-11-04 22:49 +0000
Message-ID<timstreater-72DE22.22494904112012@news.individual.net>
In reply to#17034
In article <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com>,
 Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote:

> On Fri, 02 Nov 2012 08:49:33 -0700, Patricia Shanahan wrote:
> 
> >On 11/2/2012 6:58 AM, Hans-Georg Michna wrote:
> 
> >> I sometimes see code that inefficiently loops through strings or
> >> character arrays, apparently because the programmer did not know
> >> regular expressions.
> 
> >Or perhaps the programmer thought the case was not performance critical,
> >and the explicit loop code was clearer and more readable than the
> >regular expression equivalent.
> 
> At least a simple regular expression is clear and very readable
> for someone who knows the language, probably much clearer and
> easier to read than some multi-line code in any high-level
> language.
> 
> So it depends on the programmer who reads the code. Win some,
> lose some.
> 
> s = s.replace(/^\s+|\s+$/g, "");
> 
> Is that easy to read? For me it is. It comes naturally. (:-)
> This is JavaScript's trim function, by the way.

Very likely it is; I'm not going to bother to check. Equally, I never 
bother to remember any except the most common keyboard shortcuts. Doing 
more is a waste of effort, as I'll forget them 23.6 seconds later.

s = s.trim ();   // Rather more readable.

For the same sorts of reasons, if I'm programming at the machine level 
I'll write in assembler rather than hex.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17043

FromHans-Georg Michna <hans-georgNoEmailPlease@michna.com>
Date2012-11-05 13:18 +0100
Message-ID<5vaf98189bn1ks9ikmdkp42mgs9n282hjd@4ax.com>
In reply to#17037
On Sun, 04 Nov 2012 22:49:49 +0000, Tim Streater wrote:

>In article <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com>,
> Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote:

>> s = s.replace(/^\s+|\s+$/g, "");
>> 
>> Is that easy to read? For me it is. It comes naturally. (:-)
>> This is JavaScript's trim function, by the way.

>s = s.trim ();   // Rather more readable.

Good example. It does not exist in the JavaScript version most
common today.

More interestingly, a trim() function also has to be learned.
Then you have to introduce ltrim and rtrim. But all these
functions may or may not also compress embedded white space. You
would also have to learn that for the functions.

You would be going the PHP way, have a very large number of
functions for everything the language designer has thought of.

Then what would you do if you simply wanted to remove tab
characters and turn them into spaces? Would you be totally lost?

s.replace(/\t/g, " ");

Hans-Georg

[toc] | [prev] | [next] | [standalone]


#17049

FromScott Sauyet <scott.sauyet@gmail.com>
Date2012-11-06 06:48 -0800
Message-ID<d725e235-3ad0-4e29-939b-d59d5f02beb3@y8g2000yqy.googlegroups.com>
In reply to#17037
Tim Streater  wrote:
> Hans-Georg Michna  wrote:

>> s = s.replace(/^\s+|\s+$/g, "");
>
>> Is that easy to read? For me it is. It comes naturally. (:-)
>> This is JavaScript's trim function, by the way.
>
> Very likely it is; I'm not going to bother to check. Equally, I never
> bother to remember any except the most common keyboard shortcuts. Doing
> more is a waste of effort, as I'll forget them 23.6 seconds later.
>
> s = s.trim ();   // Rather more readable.

I'm assuming that you understand that the line above would be the
*implementation* of such a `trim` function.

> For the same sorts of reasons, if I'm programming at the machine level
> I'll write in assembler rather than hex.

Is your objection only to the terseness of the syntax or so something
more fundamental?  That is, if it were written in a syntax more like
the following would your objections dissolve?:

    var r2 = RegExp2;
    // not: rx = /^\s+|\s+$/g
    var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
                    r2.ANY_WHITESPACE + r2.END_OF_STRING);

    s = rx.replaceAll(s, "");

Or perhaps

    var rx = new r2({
            [r2.START_OF_STRING, r2.ANY_WHITESPACE],
            [r2.ANY_WHITESPACE, r2.END_OF_STRING]
    });

I don't have any such implementation, nor am I interested in creating
one.  I'm just curious as to the source of your objection.  I
appreciate certain objections to regexes, but have at times found them
extremely useful and would hate to remove them entirely from my
toolkit.

  -- Scott

[toc] | [prev] | [next] | [standalone]


#17050

FromTim Streater <timstreater@greenbee.net>
Date2012-11-06 15:04 +0000
Message-ID<timstreater-E71989.15045606112012@news.individual.net>
In reply to#17049
In article 
<d725e235-3ad0-4e29-939b-d59d5f02beb3@y8g2000yqy.googlegroups.com>,
 Scott Sauyet <scott.sauyet@gmail.com> wrote:

> Tim Streater  wrote:

> > s = s.trim ();   // Rather more readable.

> I don't have any such implementation, nor am I interested in creating
> one.  I'm just curious as to the source of your objection.

Largely that they are generally unreadable, and, like TECO, perl, vi, 
and keyboard shortcuts, use arbitrary characters to represent things. I 
object to this as a programming approach. I agree that sometimes 
assembler isn't much better than hex (360 assembler comes to mind), but 
in many instances it is, and in those where is isn't, leads to pressure 
to create high-level assemblers whose code translates more or less on a 
one-to-one basis into machine code. I was able, for example, to write 
hardware bootstrap code in PL-11 for the PDP-11 using the exact same 
amount of memory as if I'd done in in assembler.

In my last job I used, occasionally, to have to create packet filters on 
Juniper router, usually when debugging some router configuration. These 
used a regex-like syntax for some filter conditions, and I had to refer 
to the manual every time.

-- 
Tim

"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted"  --  Bill of Rights 1689

[toc] | [prev] | [next] | [standalone]


#17054

FromGene Wirchenko <genew@ocis.net>
Date2012-11-06 10:43 -0800
Message-ID<48mi98trj215lkcssp9hq412ddggfnijed@4ax.com>
In reply to#17049
On Tue, 6 Nov 2012 06:48:56 -0800 (PST), Scott Sauyet
<scott.sauyet@gmail.com> wrote:

[snip]

>Is your objection only to the terseness of the syntax or so something
>more fundamental?  That is, if it were written in a syntax more like
>the following would your objections dissolve?:
>
>    var r2 = RegExp2;
>    // not: rx = /^\s+|\s+$/g
>    var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
>                    r2.ANY_WHITESPACE + r2.END_OF_STRING);
>
>    s = rx.replaceAll(s, "");
>
>Or perhaps
>
>    var rx = new r2({
>            [r2.START_OF_STRING, r2.ANY_WHITESPACE],
>            [r2.ANY_WHITESPACE, r2.END_OF_STRING]
>    });

     The examples are too verbose.  If you wanted to go that route,
then something like
          var rx=new r2(r2.SOS+r2.ANYWS+r2.OR+r2.ANYWS+r2.EOS);

>I don't have any such implementation, nor am I interested in creating
>one.  I'm just curious as to the source of your objection.  I
>appreciate certain objections to regexes, but have at times found them
>extremely useful and would hate to remove them entirely from my
>toolkit.

     Thelackofspacesorseparatingpunctuationmakeregexeshardertoread.

     If spaces could be added, it could be clearer.
          rx = /^ \s+ | \s+ $/ g

Sincerely,

Gene Wirchenko

[toc] | [prev] | [next] | [standalone]


#17055

FromScott Sauyet <scott.sauyet@gmail.com>
Date2012-11-06 11:22 -0800
Message-ID<168bd0ef-ba1a-4448-ba80-8dac4e4c26ad@g8g2000yqp.googlegroups.com>
In reply to#17054
Gene Wirchenko  wrote:
> Scott Sauyet wrote:

>> Is your objection only to the terseness of the syntax or so something
>> more fundamental?  That is, if it were written in a syntax more like
>> the following would your objections dissolve?:
>
>>    var r2 = RegExp2;
>>    // not: rx = /^\s+|\s+$/g
>>    var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
>>                    r2.ANY_WHITESPACE + r2.END_OF_STRING);
>
>>    s = rx.replaceAll(s, "");
>
>> Or perhaps
>
>>    var rx = new r2({
>>            [r2.START_OF_STRING, r2.ANY_WHITESPACE],
>>            [r2.ANY_WHITESPACE, r2.END_OF_STRING]
>>    });
>
>      The examples are too verbose.  If you wanted to go that route,
> then something like
>           var rx=new r2(r2.SOS+r2.ANYWS+r2.OR+r2.ANYWS+r2.EOS);

Are you saying that you share Tim's objections to regexes, but would
be willing to use them with a syntax such as that?

I personally can't see it.  My example was intentionally verbose and
as explicit as I could comfortably be, so as to remove the regex
terseness factor from the equation and see if Tim still had
objections.  The objections many have to Perl are similar to one main
objection to regexes: they are often simply unreadable: write-only
code.  Your abbreviations are shortened enough to no longer be clear.
None of the keywords except "OR" are really obvious.  It really
doesn't look much clearer to me than `/^\s+|\s+$/g `.

>> [ ... ]
>
>      Thelackofspacesorseparatingpunctuationmakeregexeshardertoread.
>
>      If spaces could be added, it could be clearer.
>           rx = /^ \s+ | \s+ $/ g

Yes, but that's only a small part of the readability problem.

  -- Scott

[toc] | [prev] | [next] | [standalone]


#17056

FromPatricia Shanahan <pats@acm.org>
Date2012-11-06 12:16 -0800
Message-ID<a7GdnW4cw5bI7QTNnZ2dnUVZ_ridnZ2d@earthlink.com>
In reply to#17055
On 11/6/2012 11:22 AM, Scott Sauyet wrote:
...
> I personally can't see it.  My example was intentionally verbose and
> as explicit as I could comfortably be, so as to remove the regex
> terseness factor from the equation and see if Tim still had
> objections.  The objections many have to Perl are similar to one main
> objection to regexes: they are often simply unreadable: write-only
> code.  Your abbreviations are shortened enough to no longer be clear.
> None of the keywords except "OR" are really obvious.  It really
> doesn't look much clearer to me than `/^\s+|\s+$/g `.
...

I think there is a significant difference between Perl and regex. In
Perl, terseness is a choice. One can use meaningful identifiers, white
space, and comments. I have written maintainable Perl code.

In regex, there is no option for meaningful identifiers, or, as Gene
pointed out, white space. Comments have to be either before or after the
entire regex, not interleaved with it.

I would be much happier if one could build regular expressions up in
pieces, using identifiers to reference components:

leadingSpaces = /^\s+/
trailingSpaces = /\s+$/
leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/

s = s.replace(/{leadingOrTrailingSpaces}/g, "");

Patricia

[toc] | [prev] | [next] | [standalone]


#17057

FromScott Sauyet <scott.sauyet@gmail.com>
Date2012-11-06 12:53 -0800
Message-ID<08bd3715-86f9-4e0d-a8ce-4269cac272f2@q16g2000yqc.googlegroups.com>
In reply to#17056
Patricia Shanahan  wrote:
> Scott Sauyet wrote:
>> [ ... ]  The objections many have to Perl are similar to one main
>> objection to regexes: they are often simply unreadable: write-only
>> code. [ ... ]
>
> I think there is a significant difference between Perl and regex. In
> Perl, terseness is a choice. One can use meaningful identifiers, white
> space, and comments. I have written maintainable Perl code.

Right, but the argument that people make against Perl certainly
carries over to regex, where in many flavors, including the built-in
Javascript one, one simply can't do so.


> In regex, there is no option for meaningful identifiers, or, as Gene
> pointed out, white space. Comments have to be either before or after the
> entire regex, not interleaved with it.
>
> I would be much happier if one could build regular expressions up in
> pieces, using identifiers to reference components:
>
> leadingSpaces = /^\s+/
> trailingSpaces = /\s+$/
> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>
> s = s.replace(/{leadingOrTrailingSpaces}/g, "");

Agreed.  That would be a significant improvement.  And although some
of that can be done with string concatenation and the RegExp
constructor, it's certainly not as convenient.

But do look at Thomas Lahn's implementation, where he includes a
similar feature.  I looked at it for another reason altogether and am
very interested, but have not had cause to try to use it yet.  It was
discussed earlier in this thread:

    <news:1466051.Cd1scNln3u@PointedEars.de>
    or <https://groups.google.com/group/comp.lang.javascript/msg/
52859f0fca841deb>

His unit tests are at

    <http://PointedEars.de/scripts/test/regexp>

  -- Scott

[toc] | [prev] | [next] | [standalone]


#17063

FromHans-Georg Michna <hans-georgNoEmailPlease@michna.com>
Date2012-11-07 17:34 +0100
Message-ID<st1l981go8tiiagk6pmte0oc09l69o23n3@4ax.com>
In reply to#17056
On Tue, 06 Nov 2012 12:16:59 -0800, Patricia Shanahan wrote:

>leadingSpaces = /^\s+/
>trailingSpaces = /\s+$/
>leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>
>s = s.replace(/{leadingOrTrailingSpaces}/g, "");

And that would already confuse an uninitiated reader, because
the regular expression does not select spaces. It selects white
space characters.

Your idea could easily be materialized with the help of any
macro processor of the kind that was standard in C environments
and many others.

I doubt that it would catch on, however. Regular expressions are
already firmly established. Too late to change. The only
decision is to learn or not to learn.

Hans-Georg

[toc] | [prev] | [next] | [standalone]


#17068

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2012-11-07 23:21 +0100
Message-ID<3150175.r8Iz2dskB7@PointedEars.de>
In reply to#17056
Patricia Shanahan wrote:

> Scott Sauyet wrote:
>> I personally can't see it.  My example was intentionally verbose and
>> as explicit as I could comfortably be, so as to remove the regex
>> terseness factor from the equation and see if Tim still had
>> objections.  The objections many have to Perl are similar to one main
>> objection to regexes: they are often simply unreadable: write-only
>> code.  Your abbreviations are shortened enough to no longer be clear.
>> None of the keywords except "OR" are really obvious.  It really
>> doesn't look much clearer to me than `/^\s+|\s+$/g `.
> 
> I think there is a significant difference between Perl and regex. In
> Perl, terseness is a choice. One can use meaningful identifiers, white
> space, and comments. I have written maintainable Perl code.
> 
> In regex, there is no option for meaningful identifiers, or, as Gene
> pointed out, white space. Comments have to be either before or after the
> entire regex, not interleaved with it.

Apples and oranges.  Perl is a *programming language*; regular expressions 
are part of a *technique* (pattern matching) usable *in* programming 
languages.

And apparently you are not aware that *especially* Perl's implementation of 
regular expressions, and Perl-Compatible Regular Expressions (PCRE) as 
supported by e.g. PHP, allow what you are asking for with the `x' 
(PCRE_EXTENDED) flag:

  my $leadingOrTrailingWhitespace =
    /
      ^    # (start of input
           # followed by
      \s+  # at least one whitespace)
      |    # or
      \s+  # (at least one whitespace
           # followed by
      $    # end of input)
    /x;

In PHP (with PCRE):

<?php
  $leadingOrTrailingWhitespace =
    '/
       ^    # (start of input
            # followed by
       \s+  # at least one whitespace)
       |    # or
       \s+  # (at least one whitespace
            # followed by
       $    # end of input)
     /x';

  $s = preg_replace($leadingOrTrailingWhitespace, '', '  foo  ');
?>

In ECMAScript implementations (with JSX:regexp.js [1], revisions 272 and 
later):

  var leadingOrTrailingWhitespace = new jsx.regexp.RegExp(
    [
      "^     # (start of input",
      "      # followed by",
      "\\s+  # at least one whitespace)",
      "|     # or",
      "\\s+  # (at least one whitespace",
      "      # followed by",
      "$     # end of input)"
    ].join("\n"),
    "x"
  );

or (standards-compliant since Ed. 5, proprietarily available before):

  var leadingOrTrailingWhitespace = new jsx.regexp.RegExp(
    "^     # (start of input \n\
           # followed by \n\
     \\s+  # at least one whitespace) \n\
     |     # or \n\
     \\s+  # (at least one whitespace \n\
           # followed by \n\
     $     # end of input",
    "x"
  ); 

> I would be much happier if one could build regular expressions up in
> pieces, using identifiers to reference components:
> 
> leadingSpaces = /^\s+/
> trailingSpaces = /\s+$/
> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
> 
> s = s.replace(/{leadingOrTrailingSpaces}/g, "");

Built-in (ECMA-262-5.1, §15.10.4):

  var leadingSpaces = "^\\s+";
  var trailingSpaces = "\\s+$";
  var leadingOrTrailingSpaces =
    new RegExp(leadingSpaces + "|" + trailingSpaces, "g");

  var s = s.replace(leadingOrTrailingSpaces, "");

or

  var leadingSpaces = "^\\s+";
  var trailingSpaces = "\\s+$";
  var leadingOrTrailingSpaces =
    new RegExp([leadingSpaces, trailingSpaces].join("|"), "g");

  var s = s.replace(leadingOrTrailingSpaces, "");

(You can find more examples in JSX:regexp.js where I am building the 
expressions to implement, for example, Unicode character property classes.)

With JSX:regexp.js (revisions 275 and later):

  var leadingSpaces = /^\s+/;
  var trailingSpaces = /\s+$/;
  var leadingOrTrailingSpaces =
    jsx.regexp.concat(leadingSpaces, "|", trailingSpaces);

  s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");

or

  var leadingSpaces = /^\s+/;
  var trailingSpaces = /\s+$/;
  var leadingOrTrailingSpaces = leadingSpaces.concat("|", trailingSpaces);

  s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");

I am considering

  var leadingSpaces = /^\s+/;
  var trailingSpaces = /\s+$/;
  var leadingOrTrailingSpaces =
    jsx.regexp.alternate(leadingSpaces, trailingSpaces);

  s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");

and

  var leadingSpaces = /^\s+/;
  var trailingSpaces = /\s+$/;
  var leadingOrTrailingSpaces = leadingSpaces.alternate(trailingSpaces);

  s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");

What do you think?


PointedEars
___________
[1] 
<http://PointedEars.de/websvn/filedetails.php?repname=JSX&path=%2Ftrunk%2Fregexp.js>
<http://PointedEars.de/scripts/test/regexp>
-- 
Sometimes, what you learn is wrong. If those wrong ideas are close to the 
root of the knowledge tree you build on a particular subject, pruning the 
bad branches can sometimes cause the whole tree to collapse.
  -- Mike Duffy in cljs, <news:Xns9FB6521286DB8invalidcom@94.75.214.39>

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | comp.lang.javascript


csiph-web