Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming > #2123 > unrolled thread

Regular expressions and rusty knowledge

Started byNey André de Mello Zunino <zunino@softplan.com.br>
First post2012-08-29 14:56 -0300
Last post2012-08-30 00:12 +0100
Articles 3 — 3 participants

Back to article view | Back to comp.programming


Contents

  Regular expressions and rusty knowledge Ney André de Mello Zunino <zunino@softplan.com.br> - 2012-08-29 14:56 -0300
    Re: Regular expressions and rusty knowledge Daniel Pitts <newsgroup.nospam@virtualinfinity.net> - 2012-08-29 11:19 -0700
    Re: Regular expressions and rusty knowledge Ben Bacarisse <ben.usenet@bsb.me.uk> - 2012-08-30 00:12 +0100

#2123 — Regular expressions and rusty knowledge

FromNey André de Mello Zunino <zunino@softplan.com.br>
Date2012-08-29 14:56 -0300
SubjectRegular expressions and rusty knowledge
Message-ID<k1ll3j$jsa$1@speranza.aioe.org>
Hello.

Yesterday at work, I was presented with strings like the following:

a) foo
b) foo=bar
c) foo=bar;rate=high
d) foo;rate=high
e) foo;rate
f) foo=bar;rate=high;pos
g) foo=bar;rate=high;pos;brick=many
.
.
.

They are basically semicolon-separated pairs of key=value, where values 
may or may not be present. I thought I would try and come up with a 
regular expression to recognize and allow me to extract tokens from 
them. Here's a more detailed list of the restrictions of the language:

   1. values are optional
   2. the number of key/value pairs is not specified
   3. no string can begin or end with ';'
   4. no string can begin or end with '='
   5. if a key is followed by '=', a corresponding value must be present
   6. for simplicity, assume keys and values are only made of [a-z]

This is what I came up with:

[a-z]+(=([a-z]+))?(;[a-z]+(=([a-z]+))?)*

It seems to adhere to the restrictions and recognize valid strings as 
expected. However, I'm unsure whether my idea of using a RE like that to 
extract tokens is actually feasible (e.g. the value 'high' associated 
with the second key on sample string g). For a moment, I even questioned 
myself whether that was a regular language.

In the end, I suspect the years that have passed since I got my CS 
degree have begun to show and that I'm simply missing something the 
basics here. Anyway, I'd be thankful if somebody could shed some light 
on this and help me remove the rust from my understanding of this subject.

Thank you and regards,

-- 
Ney André de Mello Zunino

[toc] | [next] | [standalone]


#2124

FromDaniel Pitts <newsgroup.nospam@virtualinfinity.net>
Date2012-08-29 11:19 -0700
Message-ID<E2t%r.2849$gZ.706@newsfe20.iad>
In reply to#2123
On 8/29/12 10:56 AM, Ney André de Mello Zunino wrote:
> Hello.
>
> Yesterday at work, I was presented with strings like the following:
>
> a) foo
> b) foo=bar
> c) foo=bar;rate=high
> d) foo;rate=high
> e) foo;rate
> f) foo=bar;rate=high;pos
> g) foo=bar;rate=high;pos;brick=many
> .
> .
> .
>
> They are basically semicolon-separated pairs of key=value, where values
> may or may not be present. I thought I would try and come up with a
> regular expression to recognize and allow me to extract tokens from
> them. Here's a more detailed list of the restrictions of the language:
>
>    1. values are optional
>    2. the number of key/value pairs is not specified
>    3. no string can begin or end with ';'
>    4. no string can begin or end with '='
>    5. if a key is followed by '=', a corresponding value must be present
>    6. for simplicity, assume keys and values are only made of [a-z]
>
> This is what I came up with:
>
> [a-z]+(=([a-z]+))?(;[a-z]+(=([a-z]+))?)*
>
> It seems to adhere to the restrictions and recognize valid strings as
> expected. However, I'm unsure whether my idea of using a RE like that to
> extract tokens is actually feasible (e.g. the value 'high' associated
> with the second key on sample string g). For a moment, I even questioned
> myself whether that was a regular language.
>
> In the end, I suspect the years that have passed since I got my CS
> degree have begun to show and that I'm simply missing something the
> basics here. Anyway, I'd be thankful if somebody could shed some light
> on this and help me remove the rust from my understanding of this subject.
>
> Thank you and regards,
>

Don't use regex for this.  Split on ;, then iterate through that and 
split on =, limit 2.  This is very similar to query HTTP parameter 
parsing, where they use & instead of ";", and then they are also encoded 
where necessary.

If you need something faster, you can hand-code something fairly easily. 
  Regex is a hammer, and this isn't a nail.

[toc] | [prev] | [next] | [standalone]


#2125

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2012-08-30 00:12 +0100
Message-ID<0.66fe6e417dd173846aa7.20120830001248BST.871uipe1hr.fsf@bsb.me.uk>
In reply to#2123
Ney André de Mello Zunino <zunino@softplan.com.br> writes:

> Yesterday at work, I was presented with strings like the following:
>
> a) foo
> b) foo=bar
> c) foo=bar;rate=high
> d) foo;rate=high
> e) foo;rate
> f) foo=bar;rate=high;pos
> g) foo=bar;rate=high;pos;brick=many
> .
> .
> .
>
> They are basically semicolon-separated pairs of key=value, where
> values may or may not be present. I thought I would try and come up
> with a regular expression to recognize and allow me to extract tokens
> from them. Here's a more detailed list of the restrictions of the
> language:
>
>   1. values are optional
>   2. the number of key/value pairs is not specified
>   3. no string can begin or end with ';'
>   4. no string can begin or end with '='
>   5. if a key is followed by '=', a corresponding value must be present
>   6. for simplicity, assume keys and values are only made of [a-z]
>
> This is what I came up with:
>
> [a-z]+(=([a-z]+))?(;[a-z]+(=([a-z]+))?)*
>
> It seems to adhere to the restrictions and recognize valid strings as
> expected. However, I'm unsure whether my idea of using a RE like that
> to extract tokens is actually feasible (e.g. the value 'high'
> associated with the second key on sample string g). For a moment, I
> even questioned myself whether that was a regular language.
>
> In the end, I suspect the years that have passed since I got my CS
> degree have begun to show and that I'm simply missing something the
> basics here. Anyway, I'd be thankful if somebody could shed some light
> on this and help me remove the rust from my understanding of this
> subject.

Daniel Pitts is right.  REs are too much here.  For a good reason why,
consider what you might do when you don't get a match.  Your RE will
just fail, but your software might very well want to be able to say why,
or at least read those parameters that follow one that fails the
pattern.

There might be situations where any failure is so severe that the whole
string must be considered unsafe.  In those cases there is some reason
to consider a regular expression, but even there it seem something of a
stretch, given the simplicity of the alternative.

-- 
Ben.

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming


csiph-web