Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.javascript > #18020 > unrolled thread
| Started by | "Mel Smith" <med_cutout_syntel@aol.com> |
|---|---|
| First post | 2013-01-08 10:44 -0700 |
| Last post | 2013-01-12 02:31 +0100 |
| Articles | 15 — 5 participants |
Back to article view | Back to comp.lang.javascript
ScreenName validation "Mel Smith" <med_cutout_syntel@aol.com> - 2013-01-08 10:44 -0700
Re: ScreenName validation Scott Sauyet <scott.sauyet@gmail.com> - 2013-01-08 11:02 -0800
Re: ScreenName validation "Mel Smith" <med_cutout_syntel@aol.com> - 2013-01-08 13:05 -0700
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-11 01:28 +0100
Re: ScreenName validation Jeff North <jnorthau@yahoo.com.au> - 2013-01-11 12:23 +1100
Re: ScreenName validation Stefan Weiss <krewecherl@gmail.com> - 2013-01-11 03:38 +0100
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-11 20:35 +0100
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-11 21:22 +0100
Re: ScreenName validation Stefan Weiss <krewecherl@gmail.com> - 2013-01-11 22:09 +0100
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-11 23:23 +0100
Re: ScreenName validation Stefan Weiss <krewecherl@gmail.com> - 2013-01-12 00:26 +0100
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-12 00:54 +0100
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-12 00:22 +0100
Re: ScreenName validation Scott Sauyet <scott.sauyet@gmail.com> - 2013-01-11 13:29 -0800
Re: ScreenName validation Eric Bednarz <bednarz@fahr-zur-hoelle.org> - 2013-01-12 02:31 +0100
| From | "Mel Smith" <med_cutout_syntel@aol.com> |
|---|---|
| Date | 2013-01-08 10:44 -0700 |
| Subject | ScreenName validation |
| Message-ID | <al347sFepchU1@mid.individual.net> |
Hi:
On a new site I'm building, I require the client (U.S., or Canadian) to
enter his chosen 'Screen Name'.
I wish to allow >= 2 and <= 12 characters to be input.
I want only lowercase alpha chars along with a dot ( . ), dash( - ), and
underscore( _ ) to be allowed. But no whitespace should be allowed.
At the server, I will add a 4-digit value to the chosen screenname,
giving the screen name some uniquity, and return the corrected screen name
to the client.
If the user leaves the screenname field blank, I will create a
screenname from the given and surnames (which are mandatory input fields).
These names will be checked against a national Names database for
verification.
The final screen name will be something like:
mel_s1234 // for a blank input
or
snow.bird1234 // where the user has entered " SNOW Bird"
Will someone suggest a Regex function for this input field.
Thanks,
--
Mel Smith
[toc] | [next] | [standalone]
| From | Scott Sauyet <scott.sauyet@gmail.com> |
|---|---|
| Date | 2013-01-08 11:02 -0800 |
| Message-ID | <53ccd13b-a933-4bbc-ba0a-7061cc9064cd@4g2000yqv.googlegroups.com> |
| In reply to | #18020 |
Mel Smith wrote:
> On a new site I'm building, I require the client (U.S., or Canadian) to
> enter his chosen 'Screen Name'.
>
> I wish to allow >= 2 and <= 12 characters to be input.
>
> I want only lowercase alpha chars along with a dot ( . ), dash( - ), and
> underscore( _ ) to be allowed. But no whitespace should be allowed.
This should be easy. And a regex to match that might be something
like:
/^[a-z\.\_\-]{2,12}$/
> At the server, I will add a 4-digit value to the chosen screenname,
> giving the screen name some uniquity, and return the corrected screen name
> to the client.
Is "uniquity" the combination of "uniqueness" and "iniquity"? :-)
Is this mandatory or just a suggested alternative in case the chosen
name is taken? If mandatory, I'd suggest you might want to rethink
this. People really don't want to have to remember extra arbitrary
numbers.
> If the user leaves the screenname field blank, I will create a
> screenname from the given and surnames (which are mandatory input fields).
> These names will be checked against a national Names database for
> verification.
Irrelevant, I think to your question, but ok.
> The final screen name will be something like:
>
> mel_s1234 // for a blank input
>
> or
>
> snow.bird1234 // where the user has entered " SNOW Bird"
Now you're confusing me. You're allowing the user to enter blank
space and uppercase converting it to your preferred form? You're
moving away from the sweet spot of simple regex. And it's not clear
if you're talking about something that you want to do client-side or
server-side. This is straightforward:
var f = function(s) {
return s.replace(/^\s+|\s+$/g, "").replace(/\s/g,
".").toLowerCase();
}
f(" SNOW bird") // "snow.bird";
but it's not clear to me if this is something that should be done
server-side anyway.
-- Scott
[toc] | [prev] | [next] | [standalone]
| From | "Mel Smith" <med_cutout_syntel@aol.com> |
|---|---|
| Date | 2013-01-08 13:05 -0700 |
| Message-ID | <al3chaFgp06U1@mid.individual.net> |
| In reply to | #18021 |
Scott said:
This should be easy. And a regex to match that might be something
like:
/^[a-z\.\_\-]{2,12}$/
Thanks for the Regex. I'll test it on my system (and also convert it to my
server language too
> At the server, I will add a 4-digit value to the chosen screenname,
> giving the screen name some uniquity, and return the corrected screen name
> to the client.
Is "uniquity" the combination of "uniqueness" and "iniquity"? :-)
Is this mandatory or just a suggested alternative in case the chosen
name is taken? If mandatory, I'd suggest you might want to rethink
this. People really don't want to have to remember extra arbitrary
numbers.
> If the user leaves the screenname field blank, I will create a
> screenname from the given and surnames (which are mandatory input fields).
> These names will be checked against a national Names database for
> verification.
Irrelevant, I think to your question, but ok.
> The final screen name will be something like:
>
> mel_s1234 // for a blank input
>
> or
>
> snow.bird1234 // where the user has entered " SNOW Bird"
Now you're confusing me. You're allowing the user to enter blank
space and uppercase converting it to your preferred form? You're
moving away from the sweet spot of simple regex. And it's not clear
if you're talking about something that you want to do client-side or
server-side. This is straightforward:
var f = function(s) {
return s.replace(/^\s+|\s+$/g, "").replace(/\s/g,
".").toLowerCase();
}
f(" SNOW bird") // "snow.bird";
but it's not clear to me if this is something that should be done
server-side anyway.
HI Scott:
As noted above, I'll use the validation technique on the server too
before passing back the corrected modified screenname. I'd like to allow
the client to completely pick his own screen name but quickly there will be
dups, obscene, grotesque, etc names abounding. My upper bound of the number
of registrants should be well below 130K people.
The 4 digits represent a code from the client's login IP Address (which
I'm saving in my database of registrants). This allows *many* Joe Smiths
to show on the screen, and allows my server programs to detect muiltple
registrations from spammers. I take a combination of the 1st digits of each
of the octets of the IP address then mix and stir to create this 4-digit
code.
I'm still not settled in my mind as to how to do this, but I know (from
past experience) I'm going to be attacked by the loonies, and I have to
start somewhere to set up my defences :)
Thanks again for the Regex and the straight-forward function above !
-Mel
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-11 01:28 +0100 |
| Message-ID | <m28v807ec9.fsf@nntp.bednarz.nl> |
| In reply to | #18021 |
Scott Sauyet <scott.sauyet@gmail.com> writes:
> /^[a-z\.\_\-]{2,12}$/
Uh, what's up with all those backslashes?
[toc] | [prev] | [next] | [standalone]
| From | Jeff North <jnorthau@yahoo.com.au> |
|---|---|
| Date | 2013-01-11 12:23 +1100 |
| Message-ID | <r1que89eftmfejrieqstmo7bp1ra3e92rv@4ax.com> |
| In reply to | #18062 |
On Fri, 11 Jan 2013 01:28:22 +0100, in comp.lang.javascript Eric
Bednarz <bednarz@fahr-zur-hoelle.org>
<m28v807ec9.fsf@nntp.bednarz.nl> wrote:
>| Scott Sauyet <scott.sauyet@gmail.com> writes:
>|
>| > /^[a-z\.\_\-]{2,12}$/
>|
>| Uh, what's up with all those backslashes?
The dot ( .* ) and hyphen ( [a-Z] ) characters are used within regexp.
To check for the actual characters they need to be escaped (preceded
with backslash).
[toc] | [prev] | [next] | [standalone]
| From | Stefan Weiss <krewecherl@gmail.com> |
|---|---|
| Date | 2013-01-11 03:38 +0100 |
| Message-ID | <kcnu00$dnm$1@news.albasani.net> |
| In reply to | #18063 |
On 2013-01-11 02:23, Jeff North wrote:
> On Fri, 11 Jan 2013 01:28:22 +0100, in comp.lang.javascript Eric
> Bednarz <bednarz@fahr-zur-hoelle.org>
> <m28v807ec9.fsf@nntp.bednarz.nl> wrote:
>
>>| Scott Sauyet <scott.sauyet@gmail.com> writes:
>>|
>>| > /^[a-z\.\_\-]{2,12}$/
>>|
>>| Uh, what's up with all those backslashes?
>
> The dot ( .* ) and hyphen ( [a-Z] ) characters are used within regexp.
> To check for the actual characters they need to be escaped (preceded
> with backslash).
Not in a character class (the hyphen may need to be escaped if it's not
at the start or end of the class). /^[a-z._-]{2,12}$/ does exactly the same.
Last time I checked, JSLint required the redundant backslash escapes for
some reason.
- stefan
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-11 20:35 +0100 |
| Message-ID | <m2hamnqzr4.fsf@nntp.bednarz.nl> |
| In reply to | #18063 |
Jeff North <jnorthau@yahoo.com.au> writes:
> Eric Bednarz <bednarz@fahr-zur-hoelle.org> wrote:
>
>>| Scott Sauyet <scott.sauyet@gmail.com> writes:
>>|
>>| > /^[a-z\.\_\-]{2,12}$/
>>|
>>| Uh, what's up with all those backslashes?
I'm afraid that was a rethorical question.
> The dot ( .* ) and hyphen ( [a-Z] ) characters are used within regexp.
> To check for the actual characters they need to be escaped (preceded
> with backslash).
// TODO homework
The only character that *must* to be escaped in a character class is the
escape character itself. Unescaped square brackets tend to confuse
newbies, so I would make an exception there, but otherwise regular
expressions are hard enough to read without user generated noise.
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-11 21:22 +0100 |
| Message-ID | <m2obgvii6g.fsf@nntp.bednarz.nl> |
| In reply to | #18072 |
Eric Bednarz <bednarz@fahr-zur-hoelle.org> writes: > The only character that *must* to be escaped in a character class is the > escape character itself. While this is true in most modern regular expression implementations, I just realized that it is actually false in ECMAScript, which inexplicably features empty character classes, so a closing square bracket literal actually has to be escaped. Except for JScript < 10. Them damn implementations of. :-)
[toc] | [prev] | [next] | [standalone]
| From | Stefan Weiss <krewecherl@gmail.com> |
|---|---|
| Date | 2013-01-11 22:09 +0100 |
| Message-ID | <kcpv2o$epv$1@news.albasani.net> |
| In reply to | #18074 |
On 2013-01-11 21:22, Eric Bednarz wrote:
> Eric Bednarz <bednarz@fahr-zur-hoelle.org> writes:
>
>> The only character that *must* to be escaped in a character class is the
>> escape character itself.
>
> While this is true in most modern regular expression implementations, I
> just realized that it is actually false in ECMAScript, which
> inexplicably features empty character classes, so a closing square
> bracket literal actually has to be escaped.
Could you give an example of when the literal ] character doesn't have
to be escaped in a character class?
The empty character class is supported in many (but not all) PCRE-like
regex implementations. It matches any character that belongs to the set
of no characters - in other words, it will never match. I'm not sure
where that would be useful. The inverse case, [^], matches any
character. It actually has a use, because it also matches newlines
(which are not matched by the "." special character).
In addition to "\", "]", and "-", some languages also require the regex
delimiter to be escaped in regex literals.
/[/]/.test("/") // true in ECMAScript
"/" =~ /[/]/ # syntax error in Perl
/[/]/.match("/") # syntax error in Ruby
preg_match("/[/]/", "/") // syntax error in PHP
ECMAScript is the odd one out here. For that reason, and because it can
mess up syntax highlighting, I'd also escape a slash in a character class.
> Except for JScript < 10. Them damn implementations of. :-)
JScript behaving differently from the rest? Shocking.
- stefan
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-11 23:23 +0100 |
| Message-ID | <m2txqnicks.fsf@nntp.bednarz.nl> |
| In reply to | #18075 |
Stefan Weiss <krewecherl@gmail.com> writes:
> Could you give an example of when the literal ] character doesn't have
> to be escaped in a character class?
I'm not sure what you are asking. Examples in other languages?
$ php -r 'echo preg_match("/^[]]$/", "]"), "\n";'
1
$ irb
>> /^[]]$/ =~ "]"
(irb):1: warning: character class has `]' without escape
=> 0
$ python
[...]
>>> import re
>>> re.compile('^[]]$').match(']')
<_sre.SRE_Match object at 0x109e39168>
$ groovysh
[...]
groovy:000> "]" ==~ /^[]]$/
===> true
> The empty character class is supported in many (but not all) PCRE-like
> regex implementations.
$ perl -e 'print "]" =~ /^[]]$/, "\n";'
1
> In addition to "\", "]", and "-", some languages also require the regex
> delimiter to be escaped in regex literals.
Well, some languages require you to .NOT escape the underscore. :-)
[toc] | [prev] | [next] | [standalone]
| From | Stefan Weiss <krewecherl@gmail.com> |
|---|---|
| Date | 2013-01-12 00:26 +0100 |
| Message-ID | <kcq72e$upp$1@news.albasani.net> |
| In reply to | #18077 |
On 2013-01-11 23:23, Eric Bednarz wrote:
> Stefan Weiss <krewecherl@gmail.com> writes:
>
>> Could you give an example of when the literal ] character doesn't have
>> to be escaped in a character class?
>
> I'm not sure what you are asking.
I may have misunderstood what you wrote. I interpreted it as: "the
literal ] has to be escaped because empty character classes are legal in
ECMAScript".
You did identify a case where a literal "]" in a character class doesn't
have to be escaped, and that is when
a) it immediately follows the opening "[", and
b) the language does not support the empty class []
So, I think it would be better to say that ] always has to be escaped,
except in this one specific edge case (which cannot occur in ECMAScript).
> Examples in other languages?
>
> $ php -r 'echo preg_match("/^[]]$/", "]"), "\n";'
> 1
Yes, but
$ php -r 'echo preg_match("/^[b]]$/", "b"), "\n";'
0
because the first ] ends the class. The same goes for the [snipped]
examples from other languages.
>> The empty character class is supported in many (but not all) PCRE-like
>> regex implementations.
Actually, now I can't think of a language other than ECMAScript that
supports empty classes, so "many implementations" was probably wrong.
- stefan
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-12 00:54 +0100 |
| Message-ID | <m2ip73i8cy.fsf@nntp.bednarz.nl> |
| In reply to | #18079 |
Stefan Weiss <krewecherl@gmail.com> writes: > On 2013-01-11 23:23, Eric Bednarz wrote: >> Stefan Weiss <krewecherl@gmail.com> writes: >> >>> Could you give an example of when the literal ] character doesn't have >>> to be escaped in a character class? >> >> I'm not sure what you are asking. > > I may have misunderstood what you wrote. I interpreted it as: "the > literal ] has to be escaped because empty character classes are legal in > ECMAScript". Don't you agree with that in your final statement: |> Actually, now I can't think of a language other than ECMAScript that |> supports empty classes, so "many implementations" was probably wrong. ? > So, I think it would be better to say that ] always has to be escaped, > except in this one specific edge case (which cannot occur in ECMAScript). As far as I'm concerned, it's no more of an edge case than not escaping `^' if it's not at the start of a character class, and not escaping `-' if it's at the beginning or end of a character class, or right after the negator. It just looks more obscure, but if you're not in the two problems camp and think about REs in general, I would consider the `pattern' itself common knowledge.
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-12 00:22 +0100 |
| Message-ID | <m2mwwfi9ta.fsf@nntp.bednarz.nl> |
| In reply to | #18075 |
Stefan Weiss <krewecherl@gmail.com> writes: > On 2013-01-11 21:22, Eric Bednarz wrote: >> [in ECMAScript] a closing square >> bracket literal actually has to be escaped [in a character (class|set)]. [...] >> Except for JScript < 10. Them damn implementations of. :-) > > JScript behaving differently from the rest? Shocking. The real shocker is that I shouldn't have talked about JScript versions, but Internet Explorer document modes. Internet Explorer <= 8 should never need the escape character. IE >= 9 does need it, unless the document mode is <= 8. Gotcha: in those cases @_jscript_version still evaluates to its native value, suggesting that conditional compilation isn't even useful to tag ECMAScript feature tests. So when I wrote "JScript < 10", I looked at an old test case with a pretty old document type declaration, resulting in * IE 9: quirks mode, no escape char necessary, JScript 9 * IE 10: quirks mode, escape char necessary, JScript 10 Turns out IE 10's quirks mode is not the same as IE 9's quirks mode, that's what it has the new IE 5 quirks mode for.
[toc] | [prev] | [next] | [standalone]
| From | Scott Sauyet <scott.sauyet@gmail.com> |
|---|---|
| Date | 2013-01-11 13:29 -0800 |
| Message-ID | <b21628e3-149b-41d5-b2a4-5c7efe933984@h2g2000yqa.googlegroups.com> |
| In reply to | #18072 |
Eric Bednarz wrote:
> Jeff North writes:
>> Eric Bednarz wrote:
>
>>>| Scott Sauyet <scott.sau...@gmail.com> writes:
>>>|
>>>| > /^[a-z\.\_\-]{2,12}$/
>>>|
>>>| Uh, what's up with all those backslashes?
Brain fart, mostly.
> I'm afraid that was a rethorical question.
>
>> The dot ( .* ) and hyphen ( [a-Z] ) characters are used within regexp.
>> To check for the actual characters they need to be escaped (preceded
>> with backslash).
>
> // TODO homework
>
> The only character that *must* to be escaped in a character class is the
> escape character itself. Unescaped square brackets tend to confuse
> newbies, so I would make an exception there, but otherwise regular
> expressions are hard enough to read without user generated noise.
Although this is true, you also have to be careful about the placement
of hyphens, and carets. And as Stefan pointed out, JSLint takes
exception to certain unescaped characters in character classes
(possibly only the hyphen.)
While this is equivalent:
/^[a-z._-]{2,12}$/
I would still prefer this:
/^[a-z._\-]{2,12}$/
especially as it doesn't fall apart if you switch the position of the
last two character identifiers:
/^[a-z.\-_]{2,12}$/
which would happen with the entirely unescaped one:
/^[a-z.-_]{2,12}$/
But really it was just a brain fart. I use regexes when I need them,
but am generally in the "two problems" camp.
-- Scott
[toc] | [prev] | [next] | [standalone]
| From | Eric Bednarz <bednarz@fahr-zur-hoelle.org> |
|---|---|
| Date | 2013-01-12 02:31 +0100 |
| Message-ID | <m2d2xbi3ub.fsf@nntp.bednarz.nl> |
| In reply to | #18076 |
Scott Sauyet <scott.sauyet@gmail.com> writes:
[redundant regexp character class escape sequences]
> Although this is true, you also have to be careful about the placement
> of hyphens, and carets.
You have to be careful about anything anyway because regular expressions
are hard to grok. Because:
* they are difficult to read anyway (especially long ones in
implementations without a /x matching mode)
* implementation differences can make your head spin, especially if you
don't use a particular language (version[1]) exclusively
* it's very easy to make grave conceptual mistakes even for what appears
to be a trivial task (e.g. match an attribute specification within an
HTML start tag), with resulting problems ranging from false
positives to potential performance nightmares (depending on all those
unforeseen subjects, of course)
[1] take ECMAScript 3 vs 5: are regexp literals cached or not, and why
(not)?
> And as Stefan pointed out, JSLint takes
> exception to certain unescaped characters in character classes
> (possibly only the hyphen.)
I'm all for catering (semi-)popular tools as long as it doesn't hurt
your brain, e.g. (function () {} ()) vs (function () {})(), who cares,
but I'm also all against writing silly code to satisfy challenged code
analysis (be it JSLint, one's pet-IDE or whatnot). I know I've done that
a lot, and it's not going anywhere, really fast. The only answer to that
is convincing bug tracker reports, or emigration.
> While this is equivalent:
>
> /^[a-z._-]{2,12}$/
>
> I would still prefer this:
>
> /^[a-z._\-]{2,12}$/
>
> especially as it doesn't fall apart if you switch the position of the
> last two character identifiers:
>
> /^[a-z.\-_]{2,12}$/
As much as I like things to be portable, I don't really buy that,
because it still falls apart if you only move one character around; if
you are smart enough to know that you need to move the escape sequence
as a whole, why wouldn't you just leave the unescaped hyphen/minus at
the end of the character class, as the average style guide (your locale
may vary) suggests?.
> But really it was just a brain fart. I use regexes when I need them,
> but am generally in the "two problems" camp.
That's a rather biased source, don't you think? Using (and remembering)
elisp regexp syntax in the (X)Emacs echo area is a lot like meeting
Kurtz at the end of the Congo River.
[toc] | [prev] | [standalone]
Back to top | Article view | comp.lang.javascript
csiph-web