Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.javascript > #17006 > unrolled thread
| Started by | RobG <rgqld@iinet.net.au> |
|---|---|
| First post | 2012-11-01 18:32 -0700 |
| Last post | 2012-11-07 11:22 -0800 |
| Articles | 20 on this page of 46 — 14 participants |
Back to article view | Back to comp.lang.javascript
Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-01 18:32 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Denis McMahon <denismfmcmahon@gmail.com> - 2012-11-02 06:14 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 02:28 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-02 10:22 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 10:37 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 06:25 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:35 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-02 14:58 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 08:49 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:52 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:49 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-05 13:18 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:48 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-06 15:04 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-06 10:43 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 11:22 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-06 12:16 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 12:53 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-07 17:34 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-07 23:21 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-07 23:23 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-07 17:00 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2012-11-08 07:16 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-07 17:50 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Dr J R Stockton <reply1245@merlyn.demon.co.uk.invalid> - 2012-11-05 19:26 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Stefan Weiss <krewecherl@gmail.com> - 2012-11-07 22:13 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 09:51 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 18:41 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 12:30 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-02 20:33 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> - 2012-11-04 21:57 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-04 22:44 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Andrew Poulos <ap_prog@hotmail.com> - 2012-11-05 15:41 +1100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Tim Streater <timstreater@greenbee.net> - 2012-11-05 09:07 +0000
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-02 22:10 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-02 15:22 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 10:19 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Patricia Shanahan <pats@acm.org> - 2012-11-03 13:05 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 21:39 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2012-11-03 12:26 +0100
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? "Jukka K. Korpela" <jkorpela@cs.tut.fi> - 2012-11-02 12:28 +0200
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Gene Wirchenko <genew@ocis.net> - 2012-11-02 10:00 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Asen Bozhilov <asen.bozhilov@gmail.com> - 2012-11-02 06:15 -0700
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-06 06:32 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? RobG <rgqld@iinet.net.au> - 2012-11-06 18:23 -0800
Re: Good use of bitwise NOT "~" or unnecessary obfuscation? Scott Sauyet <scott.sauyet@gmail.com> - 2012-11-07 11:22 -0800
Page 1 of 3 [1] 2 3 Next page →
| From | RobG <rgqld@iinet.net.au> |
|---|---|
| Date | 2012-11-01 18:32 -0700 |
| Subject | Good use of bitwise NOT "~" or unnecessary obfuscation? |
| Message-ID | <50330d14-3a30-4885-9833-18c75717a04d@googlegroups.com> |
I frequently see:
if (someString.indexOf('foo') != -1)
but it looks ugly and isn't all that intuitive. Adding one to the result means if the string isn't found, the result is zero or +ve if it is found, so:
if (someString.indexOf('foo') + 1)
but that doesn't look much better. But the bitwise NOT comes to the rescue:
if (~someString.indexOf('foo'))
now if the string isn't found, the result is 0 and some positive integer if it is. And I didn't have to create a regular expression to get similar semantics.
Can anyone see issues with the above? Other than perhaps some not seeing the bitwise NOT operator "~" so misunderstanding the test?
--
Rob
[toc] | [next] | [standalone]
| From | Denis McMahon <denismfmcmahon@gmail.com> |
|---|---|
| Date | 2012-11-02 06:14 +0000 |
| Message-ID | <k6vobo$216$1@dont-email.me> |
| In reply to | #17006 |
On Thu, 01 Nov 2012 18:32:59 -0700, RobG wrote:
> if (~someString.indexOf('foo'))
>
> Can anyone see issues with the above? Other than perhaps some not seeing
> the bitwise NOT operator "~" so misunderstanding the test?
You're making assumptions about the bitwise representation of -1 and 0
that might be implementation dependant.
I can't think of any reason why anyone would write a javascript
implementation where two's complement wasn't used to store integers, but
unless there's a specification somewhere that mandates two's complement
for integers, then you're taking a risk based on an assumption about how
everyone implements integer math that might not always be valid.
Bitfields and integers, although appearing to be the the same, may not
always behave in the same way. In this case, I'd say that just because in
every system that you (and I) am familiar with this would be ok, doesn't
mean that there isn't a platform out there somewhere, now or in the
future, where it won't work.
Rgds
Denis McMahon
[toc] | [prev] | [next] | [standalone]
| From | Patricia Shanahan <pats@acm.org> |
|---|---|
| Date | 2012-11-02 02:28 -0700 |
| Message-ID | <KvGdnVxSk7C8Dw7NnZ2dnUVZ_j-dnZ2d@earthlink.com> |
| In reply to | #17007 |
On 11/1/2012 11:14 PM, Denis McMahon wrote:
> On Thu, 01 Nov 2012 18:32:59 -0700, RobG wrote:
>
>> if (~someString.indexOf('foo'))
>>
>> Can anyone see issues with the above? Other than perhaps some not seeing
>> the bitwise NOT operator "~" so misunderstanding the test?
>
> You're making assumptions about the bitwise representation of -1 and 0
> that might be implementation dependant.
>
> I can't think of any reason why anyone would write a javascript
> implementation where two's complement wasn't used to store integers, but
> unless there's a specification somewhere that mandates two's complement
> for integers, then you're taking a risk based on an assumption about how
> everyone implements integer math that might not always be valid.
>
> Bitfields and integers, although appearing to be the the same, may not
> always behave in the same way. In this case, I'd say that just because in
> every system that you (and I) am familiar with this would be ok, doesn't
> mean that there isn't a platform out there somewhere, now or in the
> future, where it won't work.
The use of 32 bit 2's complement for bitwise NOT is required by ECMA-262.
Because of the wrap-around in the 32 bit arithmetic, using **
for exponent, string indexOf results 2**32-1, 2**33-1, 2**34-1, ...,
would also map to bit pattern ffffffff on the required conversion.
I don't know how practical implementations deal with attempts to
construct strings length 2**32 and longer.
On the question in the subject, is there any way to understand
"~someString.indexOf('foo')" without mentally converting it to
"someString.indexOf('foo') != -1 [mod 2**32]"? If not, I think it is
unnecessary obfuscation.
Patricia
[toc] | [prev] | [next] | [standalone]
| From | "Evertjan." <exxjxw.hannivoort@inter.nl.net> |
|---|---|
| Date | 2012-11-02 10:22 +0100 |
| Message-ID | <XnsA0FF6988CAC95eejj99@194.109.133.133> |
| In reply to | #17006 |
RobG wrote on 02 nov 2012 in comp.lang.javascript:
> I frequently see:
>
> if (someString.indexOf('foo') != -1)
>
> but it looks ugly and isn't all that intuitive.
I don't know about your personal taste of beauty and your intuition,
I doubt these can be generalized.
The above is far easier [my subjectiveness]
and more versatile [easily shown if you like] met by using Regex:
if ( /foo/.test(someString) ) {..}
I trust not understanding Regex is not a valid counterargument.
--
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2012-11-02 10:37 +0100 |
| Message-ID | <timstreater-3F119F.10370502112012@news.individual.net> |
| In reply to | #17009 |
In article <XnsA0FF6988CAC95eejj99@194.109.133.133>,
"Evertjan." <exxjxw.hannivoort@inter.nl.net> wrote:
> RobG wrote on 02 nov 2012 in comp.lang.javascript:
>
> > I frequently see:
> >
> > if (someString.indexOf('foo') != -1)
> >
> > but it looks ugly and isn't all that intuitive.
>
> I don't know about your personal taste of beauty and your intuition,
> I doubt these can be generalized.
>
> The above is far easier [my subjectiveness]
> and more versatile [easily shown if you like] met by using Regex:
>
> if ( /foo/.test(someString) ) {..}
>
> I trust not understanding Regex is not a valid counterargument.
Oh but it is. Regex, in general, is just line-noise.
--
Tim
"That excessive bail ought not to be required, nor excessive fines imposed,
nor cruel and unusual punishments inflicted" -- Bill of Rights 1689
[toc] | [prev] | [next] | [standalone]
| From | Patricia Shanahan <pats@acm.org> |
|---|---|
| Date | 2012-11-02 06:25 -0700 |
| Message-ID | <BMSdnRSFG8QIVA7NnZ2dnUVZ_vGdnZ2d@earthlink.com> |
| In reply to | #17009 |
On 11/2/2012 2:22 AM, Evertjan. wrote:
...
> if ( /foo/.test(someString) ) {..}
>
> I trust not understanding Regex is not a valid counterargument.
>
If I'm prepared to look at it long enough, I can generally work out what
a Regex does, but Regex does not seem to me to be a very human-friendly,
smoothly readable language.
For this particular case, it looks OK if the probe really is a literal
such as "foo". Could you show me the code you would use for this if the
probe were an actual parameter or variable with unknown contents?
That is, what would you use to replace:
if(someString.indexOf(someOtherString)) != -1)
Thanks,
Patricia
[toc] | [prev] | [next] | [standalone]
| From | Asen Bozhilov <asen.bozhilov@gmail.com> |
|---|---|
| Date | 2012-11-02 06:35 -0700 |
| Message-ID | <a0445966-aad0-4a66-b281-21db9430956f@q4g2000vbg.googlegroups.com> |
| In reply to | #17014 |
Patricia Shanahan wrote: > For this particular case, it looks OK if the probe really is a literal > such as "foo". Could you show me the code you would use for this if the > probe were an actual parameter or variable with unknown contents? Right. RegExp solution would work with literal values. With runtime evaluating values, it should need compilation of RegExp e.g. new RegExp(value); Which is inappropriate at that case. > That is, what would you use to replace: > > if(someString.indexOf(someOtherString)) != -1) Typo
[toc] | [prev] | [next] | [standalone]
| From | Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> |
|---|---|
| Date | 2012-11-02 14:58 +0100 |
| Message-ID | <f7k798117mcgu610jd2i8ca98s770f9uc5@4ax.com> |
| In reply to | #17014 |
On Fri, 02 Nov 2012 06:25:08 -0700, Patricia Shanahan wrote: >If I'm prepared to look at it long enough, I can generally work out what >a Regex does, but Regex does not seem to me to be a very human-friendly, >smoothly readable language. I agree, but in my experience regular expressions are too frequently the best solution to string processing problems to ignore them. I sometimes see code that inefficiently loops through strings or character arrays, apparently because the programmer did not know regular expressions. My personal advice to every JavaScript (and other) programmer is to learn at least the most basic parts of the regular expression language. Hans-Georg
[toc] | [prev] | [next] | [standalone]
| From | Patricia Shanahan <pats@acm.org> |
|---|---|
| Date | 2012-11-02 08:49 -0700 |
| Message-ID | <Ceidnfe5X9M_dg7NnZ2dnUVZ_hWdnZ2d@earthlink.com> |
| In reply to | #17016 |
On 11/2/2012 6:58 AM, Hans-Georg Michna wrote: > On Fri, 02 Nov 2012 06:25:08 -0700, Patricia Shanahan wrote: > >> If I'm prepared to look at it long enough, I can generally work out what >> a Regex does, but Regex does not seem to me to be a very human-friendly, >> smoothly readable language. > > I agree, but in my experience regular expressions are too > frequently the best solution to string processing problems to > ignore them. I agree with not ignoring them, but also think there is a tendency to over-use them, making code much more complicated, and less readable, than it could be. Using regular expression matching to do a simple string contains seems to me to be an example. At a minimum, the reader has to check the probe string for RegExp special characters. > > I sometimes see code that inefficiently loops through strings or > character arrays, apparently because the programmer did not know > regular expressions. Or perhaps the programmer thought the case was not performance critical, and the explicit loop code was clearer and more readable than the regular expression equivalent. > > My personal advice to every JavaScript (and other) programmer is > to learn at least the most basic parts of the regular expression > language. Been there, done that - in 1983 when I was learning ed, vi, grep, awk etc. Patricia
[toc] | [prev] | [next] | [standalone]
| From | Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> |
|---|---|
| Date | 2012-11-04 21:52 +0100 |
| Message-ID | <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com> |
| In reply to | #17017 |
On Fri, 02 Nov 2012 08:49:33 -0700, Patricia Shanahan wrote: >On 11/2/2012 6:58 AM, Hans-Georg Michna wrote: >> I sometimes see code that inefficiently loops through strings or >> character arrays, apparently because the programmer did not know >> regular expressions. >Or perhaps the programmer thought the case was not performance critical, >and the explicit loop code was clearer and more readable than the >regular expression equivalent. At least a simple regular expression is clear and very readable for someone who knows the language, probably much clearer and easier to read than some multi-line code in any high-level language. So it depends on the programmer who reads the code. Win some, lose some. s = s.replace(/^\s+|\s+$/g, ""); Is that easy to read? For me it is. It comes naturally. (:-) This is JavaScript's trim function, by the way. Hans-Georg
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2012-11-04 22:49 +0000 |
| Message-ID | <timstreater-72DE22.22494904112012@news.individual.net> |
| In reply to | #17034 |
In article <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com>, Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote: > On Fri, 02 Nov 2012 08:49:33 -0700, Patricia Shanahan wrote: > > >On 11/2/2012 6:58 AM, Hans-Georg Michna wrote: > > >> I sometimes see code that inefficiently loops through strings or > >> character arrays, apparently because the programmer did not know > >> regular expressions. > > >Or perhaps the programmer thought the case was not performance critical, > >and the explicit loop code was clearer and more readable than the > >regular expression equivalent. > > At least a simple regular expression is clear and very readable > for someone who knows the language, probably much clearer and > easier to read than some multi-line code in any high-level > language. > > So it depends on the programmer who reads the code. Win some, > lose some. > > s = s.replace(/^\s+|\s+$/g, ""); > > Is that easy to read? For me it is. It comes naturally. (:-) > This is JavaScript's trim function, by the way. Very likely it is; I'm not going to bother to check. Equally, I never bother to remember any except the most common keyboard shortcuts. Doing more is a waste of effort, as I'll forget them 23.6 seconds later. s = s.trim (); // Rather more readable. For the same sorts of reasons, if I'm programming at the machine level I'll write in assembler rather than hex. -- Tim "That excessive bail ought not to be required, nor excessive fines imposed, nor cruel and unusual punishments inflicted" -- Bill of Rights 1689
[toc] | [prev] | [next] | [standalone]
| From | Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> |
|---|---|
| Date | 2012-11-05 13:18 +0100 |
| Message-ID | <5vaf98189bn1ks9ikmdkp42mgs9n282hjd@4ax.com> |
| In reply to | #17037 |
On Sun, 04 Nov 2012 22:49:49 +0000, Tim Streater wrote: >In article <32ld985i9rkvi86ul4eqiv1gsdhr78sd6g@4ax.com>, > Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> wrote: >> s = s.replace(/^\s+|\s+$/g, ""); >> >> Is that easy to read? For me it is. It comes naturally. (:-) >> This is JavaScript's trim function, by the way. >s = s.trim (); // Rather more readable. Good example. It does not exist in the JavaScript version most common today. More interestingly, a trim() function also has to be learned. Then you have to introduce ltrim and rtrim. But all these functions may or may not also compress embedded white space. You would also have to learn that for the functions. You would be going the PHP way, have a very large number of functions for everything the language designer has thought of. Then what would you do if you simply wanted to remove tab characters and turn them into spaces? Would you be totally lost? s.replace(/\t/g, " "); Hans-Georg
[toc] | [prev] | [next] | [standalone]
| From | Scott Sauyet <scott.sauyet@gmail.com> |
|---|---|
| Date | 2012-11-06 06:48 -0800 |
| Message-ID | <d725e235-3ad0-4e29-939b-d59d5f02beb3@y8g2000yqy.googlegroups.com> |
| In reply to | #17037 |
Tim Streater wrote:
> Hans-Georg Michna wrote:
>> s = s.replace(/^\s+|\s+$/g, "");
>
>> Is that easy to read? For me it is. It comes naturally. (:-)
>> This is JavaScript's trim function, by the way.
>
> Very likely it is; I'm not going to bother to check. Equally, I never
> bother to remember any except the most common keyboard shortcuts. Doing
> more is a waste of effort, as I'll forget them 23.6 seconds later.
>
> s = s.trim (); // Rather more readable.
I'm assuming that you understand that the line above would be the
*implementation* of such a `trim` function.
> For the same sorts of reasons, if I'm programming at the machine level
> I'll write in assembler rather than hex.
Is your objection only to the terseness of the syntax or so something
more fundamental? That is, if it were written in a syntax more like
the following would your objections dissolve?:
var r2 = RegExp2;
// not: rx = /^\s+|\s+$/g
var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
r2.ANY_WHITESPACE + r2.END_OF_STRING);
s = rx.replaceAll(s, "");
Or perhaps
var rx = new r2({
[r2.START_OF_STRING, r2.ANY_WHITESPACE],
[r2.ANY_WHITESPACE, r2.END_OF_STRING]
});
I don't have any such implementation, nor am I interested in creating
one. I'm just curious as to the source of your objection. I
appreciate certain objections to regexes, but have at times found them
extremely useful and would hate to remove them entirely from my
toolkit.
-- Scott
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2012-11-06 15:04 +0000 |
| Message-ID | <timstreater-E71989.15045606112012@news.individual.net> |
| In reply to | #17049 |
In article <d725e235-3ad0-4e29-939b-d59d5f02beb3@y8g2000yqy.googlegroups.com>, Scott Sauyet <scott.sauyet@gmail.com> wrote: > Tim Streater wrote: > > s = s.trim (); // Rather more readable. > I don't have any such implementation, nor am I interested in creating > one. I'm just curious as to the source of your objection. Largely that they are generally unreadable, and, like TECO, perl, vi, and keyboard shortcuts, use arbitrary characters to represent things. I object to this as a programming approach. I agree that sometimes assembler isn't much better than hex (360 assembler comes to mind), but in many instances it is, and in those where is isn't, leads to pressure to create high-level assemblers whose code translates more or less on a one-to-one basis into machine code. I was able, for example, to write hardware bootstrap code in PL-11 for the PDP-11 using the exact same amount of memory as if I'd done in in assembler. In my last job I used, occasionally, to have to create packet filters on Juniper router, usually when debugging some router configuration. These used a regex-like syntax for some filter conditions, and I had to refer to the manual every time. -- Tim "That excessive bail ought not to be required, nor excessive fines imposed, nor cruel and unusual punishments inflicted" -- Bill of Rights 1689
[toc] | [prev] | [next] | [standalone]
| From | Gene Wirchenko <genew@ocis.net> |
|---|---|
| Date | 2012-11-06 10:43 -0800 |
| Message-ID | <48mi98trj215lkcssp9hq412ddggfnijed@4ax.com> |
| In reply to | #17049 |
On Tue, 6 Nov 2012 06:48:56 -0800 (PST), Scott Sauyet
<scott.sauyet@gmail.com> wrote:
[snip]
>Is your objection only to the terseness of the syntax or so something
>more fundamental? That is, if it were written in a syntax more like
>the following would your objections dissolve?:
>
> var r2 = RegExp2;
> // not: rx = /^\s+|\s+$/g
> var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
> r2.ANY_WHITESPACE + r2.END_OF_STRING);
>
> s = rx.replaceAll(s, "");
>
>Or perhaps
>
> var rx = new r2({
> [r2.START_OF_STRING, r2.ANY_WHITESPACE],
> [r2.ANY_WHITESPACE, r2.END_OF_STRING]
> });
The examples are too verbose. If you wanted to go that route,
then something like
var rx=new r2(r2.SOS+r2.ANYWS+r2.OR+r2.ANYWS+r2.EOS);
>I don't have any such implementation, nor am I interested in creating
>one. I'm just curious as to the source of your objection. I
>appreciate certain objections to regexes, but have at times found them
>extremely useful and would hate to remove them entirely from my
>toolkit.
Thelackofspacesorseparatingpunctuationmakeregexeshardertoread.
If spaces could be added, it could be clearer.
rx = /^ \s+ | \s+ $/ g
Sincerely,
Gene Wirchenko
[toc] | [prev] | [next] | [standalone]
| From | Scott Sauyet <scott.sauyet@gmail.com> |
|---|---|
| Date | 2012-11-06 11:22 -0800 |
| Message-ID | <168bd0ef-ba1a-4448-ba80-8dac4e4c26ad@g8g2000yqp.googlegroups.com> |
| In reply to | #17054 |
Gene Wirchenko wrote:
> Scott Sauyet wrote:
>> Is your objection only to the terseness of the syntax or so something
>> more fundamental? That is, if it were written in a syntax more like
>> the following would your objections dissolve?:
>
>> var r2 = RegExp2;
>> // not: rx = /^\s+|\s+$/g
>> var rx = new r2(r2.START_OF_STRING + r2.ANY_WHITESPACE + r2.OR +
>> r2.ANY_WHITESPACE + r2.END_OF_STRING);
>
>> s = rx.replaceAll(s, "");
>
>> Or perhaps
>
>> var rx = new r2({
>> [r2.START_OF_STRING, r2.ANY_WHITESPACE],
>> [r2.ANY_WHITESPACE, r2.END_OF_STRING]
>> });
>
> The examples are too verbose. If you wanted to go that route,
> then something like
> var rx=new r2(r2.SOS+r2.ANYWS+r2.OR+r2.ANYWS+r2.EOS);
Are you saying that you share Tim's objections to regexes, but would
be willing to use them with a syntax such as that?
I personally can't see it. My example was intentionally verbose and
as explicit as I could comfortably be, so as to remove the regex
terseness factor from the equation and see if Tim still had
objections. The objections many have to Perl are similar to one main
objection to regexes: they are often simply unreadable: write-only
code. Your abbreviations are shortened enough to no longer be clear.
None of the keywords except "OR" are really obvious. It really
doesn't look much clearer to me than `/^\s+|\s+$/g `.
>> [ ... ]
>
> Thelackofspacesorseparatingpunctuationmakeregexeshardertoread.
>
> If spaces could be added, it could be clearer.
> rx = /^ \s+ | \s+ $/ g
Yes, but that's only a small part of the readability problem.
-- Scott
[toc] | [prev] | [next] | [standalone]
| From | Patricia Shanahan <pats@acm.org> |
|---|---|
| Date | 2012-11-06 12:16 -0800 |
| Message-ID | <a7GdnW4cw5bI7QTNnZ2dnUVZ_ridnZ2d@earthlink.com> |
| In reply to | #17055 |
On 11/6/2012 11:22 AM, Scott Sauyet wrote:
...
> I personally can't see it. My example was intentionally verbose and
> as explicit as I could comfortably be, so as to remove the regex
> terseness factor from the equation and see if Tim still had
> objections. The objections many have to Perl are similar to one main
> objection to regexes: they are often simply unreadable: write-only
> code. Your abbreviations are shortened enough to no longer be clear.
> None of the keywords except "OR" are really obvious. It really
> doesn't look much clearer to me than `/^\s+|\s+$/g `.
...
I think there is a significant difference between Perl and regex. In
Perl, terseness is a choice. One can use meaningful identifiers, white
space, and comments. I have written maintainable Perl code.
In regex, there is no option for meaningful identifiers, or, as Gene
pointed out, white space. Comments have to be either before or after the
entire regex, not interleaved with it.
I would be much happier if one could build regular expressions up in
pieces, using identifiers to reference components:
leadingSpaces = /^\s+/
trailingSpaces = /\s+$/
leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
s = s.replace(/{leadingOrTrailingSpaces}/g, "");
Patricia
[toc] | [prev] | [next] | [standalone]
| From | Scott Sauyet <scott.sauyet@gmail.com> |
|---|---|
| Date | 2012-11-06 12:53 -0800 |
| Message-ID | <08bd3715-86f9-4e0d-a8ce-4269cac272f2@q16g2000yqc.googlegroups.com> |
| In reply to | #17056 |
Patricia Shanahan wrote:
> Scott Sauyet wrote:
>> [ ... ] The objections many have to Perl are similar to one main
>> objection to regexes: they are often simply unreadable: write-only
>> code. [ ... ]
>
> I think there is a significant difference between Perl and regex. In
> Perl, terseness is a choice. One can use meaningful identifiers, white
> space, and comments. I have written maintainable Perl code.
Right, but the argument that people make against Perl certainly
carries over to regex, where in many flavors, including the built-in
Javascript one, one simply can't do so.
> In regex, there is no option for meaningful identifiers, or, as Gene
> pointed out, white space. Comments have to be either before or after the
> entire regex, not interleaved with it.
>
> I would be much happier if one could build regular expressions up in
> pieces, using identifiers to reference components:
>
> leadingSpaces = /^\s+/
> trailingSpaces = /\s+$/
> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>
> s = s.replace(/{leadingOrTrailingSpaces}/g, "");
Agreed. That would be a significant improvement. And although some
of that can be done with string concatenation and the RegExp
constructor, it's certainly not as convenient.
But do look at Thomas Lahn's implementation, where he includes a
similar feature. I looked at it for another reason altogether and am
very interested, but have not had cause to try to use it yet. It was
discussed earlier in this thread:
<news:1466051.Cd1scNln3u@PointedEars.de>
or <https://groups.google.com/group/comp.lang.javascript/msg/
52859f0fca841deb>
His unit tests are at
<http://PointedEars.de/scripts/test/regexp>
-- Scott
[toc] | [prev] | [next] | [standalone]
| From | Hans-Georg Michna <hans-georgNoEmailPlease@michna.com> |
|---|---|
| Date | 2012-11-07 17:34 +0100 |
| Message-ID | <st1l981go8tiiagk6pmte0oc09l69o23n3@4ax.com> |
| In reply to | #17056 |
On Tue, 06 Nov 2012 12:16:59 -0800, Patricia Shanahan wrote:
>leadingSpaces = /^\s+/
>trailingSpaces = /\s+$/
>leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>
>s = s.replace(/{leadingOrTrailingSpaces}/g, "");
And that would already confuse an uninitiated reader, because
the regular expression does not select spaces. It selects white
space characters.
Your idea could easily be materialized with the help of any
macro processor of the kind that was standard in C environments
and many others.
I doubt that it would catch on, however. Regular expressions are
already firmly established. Too late to change. The only
decision is to learn or not to learn.
Hans-Georg
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2012-11-07 23:21 +0100 |
| Message-ID | <3150175.r8Iz2dskB7@PointedEars.de> |
| In reply to | #17056 |
Patricia Shanahan wrote:
> Scott Sauyet wrote:
>> I personally can't see it. My example was intentionally verbose and
>> as explicit as I could comfortably be, so as to remove the regex
>> terseness factor from the equation and see if Tim still had
>> objections. The objections many have to Perl are similar to one main
>> objection to regexes: they are often simply unreadable: write-only
>> code. Your abbreviations are shortened enough to no longer be clear.
>> None of the keywords except "OR" are really obvious. It really
>> doesn't look much clearer to me than `/^\s+|\s+$/g `.
>
> I think there is a significant difference between Perl and regex. In
> Perl, terseness is a choice. One can use meaningful identifiers, white
> space, and comments. I have written maintainable Perl code.
>
> In regex, there is no option for meaningful identifiers, or, as Gene
> pointed out, white space. Comments have to be either before or after the
> entire regex, not interleaved with it.
Apples and oranges. Perl is a *programming language*; regular expressions
are part of a *technique* (pattern matching) usable *in* programming
languages.
And apparently you are not aware that *especially* Perl's implementation of
regular expressions, and Perl-Compatible Regular Expressions (PCRE) as
supported by e.g. PHP, allow what you are asking for with the `x'
(PCRE_EXTENDED) flag:
my $leadingOrTrailingWhitespace =
/
^ # (start of input
# followed by
\s+ # at least one whitespace)
| # or
\s+ # (at least one whitespace
# followed by
$ # end of input)
/x;
In PHP (with PCRE):
<?php
$leadingOrTrailingWhitespace =
'/
^ # (start of input
# followed by
\s+ # at least one whitespace)
| # or
\s+ # (at least one whitespace
# followed by
$ # end of input)
/x';
$s = preg_replace($leadingOrTrailingWhitespace, '', ' foo ');
?>
In ECMAScript implementations (with JSX:regexp.js [1], revisions 272 and
later):
var leadingOrTrailingWhitespace = new jsx.regexp.RegExp(
[
"^ # (start of input",
" # followed by",
"\\s+ # at least one whitespace)",
"| # or",
"\\s+ # (at least one whitespace",
" # followed by",
"$ # end of input)"
].join("\n"),
"x"
);
or (standards-compliant since Ed. 5, proprietarily available before):
var leadingOrTrailingWhitespace = new jsx.regexp.RegExp(
"^ # (start of input \n\
# followed by \n\
\\s+ # at least one whitespace) \n\
| # or \n\
\\s+ # (at least one whitespace \n\
# followed by \n\
$ # end of input",
"x"
);
> I would be much happier if one could build regular expressions up in
> pieces, using identifiers to reference components:
>
> leadingSpaces = /^\s+/
> trailingSpaces = /\s+$/
> leadingOrTrailingSpaces = /{leadingSpaces} | {trailingSpaces}/
>
> s = s.replace(/{leadingOrTrailingSpaces}/g, "");
Built-in (ECMA-262-5.1, §15.10.4):
var leadingSpaces = "^\\s+";
var trailingSpaces = "\\s+$";
var leadingOrTrailingSpaces =
new RegExp(leadingSpaces + "|" + trailingSpaces, "g");
var s = s.replace(leadingOrTrailingSpaces, "");
or
var leadingSpaces = "^\\s+";
var trailingSpaces = "\\s+$";
var leadingOrTrailingSpaces =
new RegExp([leadingSpaces, trailingSpaces].join("|"), "g");
var s = s.replace(leadingOrTrailingSpaces, "");
(You can find more examples in JSX:regexp.js where I am building the
expressions to implement, for example, Unicode character property classes.)
With JSX:regexp.js (revisions 275 and later):
var leadingSpaces = /^\s+/;
var trailingSpaces = /\s+$/;
var leadingOrTrailingSpaces =
jsx.regexp.concat(leadingSpaces, "|", trailingSpaces);
s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");
or
var leadingSpaces = /^\s+/;
var trailingSpaces = /\s+$/;
var leadingOrTrailingSpaces = leadingSpaces.concat("|", trailingSpaces);
s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");
I am considering
var leadingSpaces = /^\s+/;
var trailingSpaces = /\s+$/;
var leadingOrTrailingSpaces =
jsx.regexp.alternate(leadingSpaces, trailingSpaces);
s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");
and
var leadingSpaces = /^\s+/;
var trailingSpaces = /\s+$/;
var leadingOrTrailingSpaces = leadingSpaces.alternate(trailingSpaces);
s = s.replace(new jsx.regexp.RegExp(leadingOrTrailingSpaces, "g"), "");
What do you think?
PointedEars
___________
[1]
<http://PointedEars.de/websvn/filedetails.php?repname=JSX&path=%2Ftrunk%2Fregexp.js>
<http://PointedEars.de/scripts/test/regexp>
--
Sometimes, what you learn is wrong. If those wrong ideas are close to the
root of the knowledge tree you build on a particular subject, pruning the
bad branches can sometimes cause the whole tree to collapse.
-- Mike Duffy in cljs, <news:Xns9FB6521286DB8invalidcom@94.75.214.39>
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | comp.lang.javascript
csiph-web