Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.python > #198071 > unrolled thread

Pythonic way to combine regular expressions?

Started byJohann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid>
First post2026-09-26 02:54 +0800
Last post2026-09-26 12:22 +0000
Articles 6 — 2 participants

Back to article view | Back to comp.lang.python


Contents

  Pythonic way to combine regular expressions? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-26 02:54 +0800
    Re: Pythonic way to combine regular expressions? ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-25 19:08 +0000
      Re: Pythonic way to combine regular expressions? ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-25 19:19 +0000
    Re: Pythonic way to combine regular expressions? ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-26 10:04 +0000
      Re: Pythonic way to combine regular expressions? Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-09-26 20:09 +0800
        Re: Pythonic way to combine regular expressions? ram@zedat.fu-berlin.de (Stefan Ram) - 2026-09-26 12:22 +0000

#198071 — Pythonic way to combine regular expressions?

FromJohann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid>
Date2026-09-26 02:54 +0800
SubjectPythonic way to combine regular expressions?
Message-ID<MrztS.208529$Nn1.7474@fx18.ams4>
Dear comp.lang.python, and not comp.compilers; this should be too easy.

What is a good way to combine regular expressions in Python?  We are
cogitating the most /Pythonic/ way to make this happen.

As an /inquisiting Eisenhorn/ we tackle the problem as follows.  First,
we check /Stack Overflow/ and see that the solutions presented there are
in fact /not Pythonic enough/.  So now we come to attempt to up-Python-
ate the solutions there, and play Heretic at the same time.  This will
be /fun/.

We are considering the /number/ in S.V.G. path elements, the /d/ attri-
bute, which has the following E.B.N.F. grammar, conveniently pasted in
from w3.org, since that site might disappear at any moment, and Usenet
is eternal.

    number:
        sign? integer-constant
        | sign? floating-point-constant
    integer-constant:
        digit-sequence
    floating-point-constant:
        fractional-constant exponent?
        | digit-sequence exponent
    fractional-constant:
        digit-sequence? "." digit-sequence
        | digit-sequence "."
    exponent:
        ( "e" | "E" ) sign? digit-sequence
    sign:
        "+" | "-"
    digit-sequence:
        digit
        | digit digit-sequence
    digit:
        "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"

Then we turn to /flex & bison/, lowercase as per title page, by the re-
markable John Levine, and cross post to alt.folklore.computers, since
they are out of topics, and John can just completely ignore this post.

We turn to page 21, and look at the section called /Regular Expression
Examples/, which conveniently teaches us to identify floating point con-
stants in Fortran 77.  We shall therefore use that knowledge, and see
what we get.

exp = r'(E|e)[-+]?[0-9]+' ## Completely untested.
digits_with_optional_dot = r'([0-9]*\.?[0-9]+)' ## Again, completely un-
                                                 ## tested.
digits_before_dot = r'[0-9]+[.]' ## Also utterly untested.
integer_constant = r'[0-9]+'     ## As above.

So now we need to combine this in the most Pythonic way possible, but
alas, the above is not all that useful, so we try again.

   import re
   number = re.compile( '|'.join( p for p in [ r'([0-9]*\.?[0-9]+)',
                                               r'[0-9]+[.],
                                               r'(E|e)[-+]?[0-9]+',
                                               r'[0-9]+' ] )

Alas, there is a bug.  Mea culpa, as Julius Caesar used to say.  How
do we get the exponent to attach optionally to the two previous reg-
ular expressions, and not to the last one?  The idea here, is to do
this completely undocumented, just like Kent Beck teaches us to do in
his book /eXtreme Programming/.  We are after all, using the /modern/
Pythonic programming language, and don't want to be old fashioned.

We also do not want to introduce /any/ extra identifiers, nor paren-
thesis.  Is our very own Python influencer, Lawrence D'Oliveiro up to
the task?

There is also the question of whether we can just completely skip all
this, and jump straight to re.compile(), without the extra .join() op-
erator?

Does Python deal well with really long constant regular expressions or
does Guido van Rossum run out of memory?  And how do we do this even
more misunderstandable?  Is there an up-complication we can apply to
this problem, or do we need to call on Gerald Sussman to write (apply)
for us?


All these questions need to be cogitated, but since Eisenhorn will not
read this until something like 40,000 thousand years later, we have time
on our hands.  What do you say?
-- 
Johann | email: invalid -> com | http://www.myrkraverk.com/blog/
I'm not from the Internet, I just work there. | via Easynews.com
https://bsky.app/profile/myrkraverk.bsky.social | for ( ;; ) _:;
Federated at https://fed.brid.gy/bsky/myrkraverk.bsky.social

[toc] | [next] | [standalone]


#198072

Fromram@zedat.fu-berlin.de (Stefan Ram)
Date2026-09-25 19:08 +0000
Message-ID<pattern-20260925200755@ram.dialup.fu-berlin.de>
In reply to#198071
Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>What is a good way to combine regular expressions in Python?  We are
>cogitating the most /Pythonic/ way to make this happen.
>    number:
>        sign? integer-constant
>        | sign? floating-point-constant
>    integer-constant:
>        digit-sequence
>    floating-point-constant:
>        fractional-constant exponent?
>        | digit-sequence exponent
>    fractional-constant:
>        digit-sequence? "." digit-sequence
>        | digit-sequence "."
>    exponent:
>        ( "e" | "E" ) sign? digit-sequence
>    sign:
>        "+" | "-"
>    digit-sequence:
>        digit
>        | digit digit-sequence
>    digit:
>        "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"

  Maybe something like this?

import re
from typing import Final

DIGIT: Final[ str ]= r"[0-9]"

DIGIT_SEQUENCE: Final[ str ]= rf"{DIGIT}+"

SIGN: Final[ str ]= r"[+-]"

EXPONENT: Final[ str ]= rf"[eE]{SIGN}?{DIGIT_SEQUENCE}"

FRACTIONAL_CONSTANT: Final[ str ]= \
rf"(?:(?:{DIGIT_SEQUENCE})?\.{DIGIT_SEQUENCE}|{DIGIT_SEQUENCE}\.)"

INTEGER_CONSTANT: Final[ str ]= DIGIT_SEQUENCE

FLOATING_POINT_CONSTANT: Final[ str ]= \
rf"(?:{FRACTIONAL_CONSTANT}(?:{EXPONENT})?|{DIGIT_SEQUENCE}{EXPONENT})"

NUMBER_PATTERN: Final[ str ]= \
rf"^{SIGN}?(?:{INTEGER_CONSTANT}|{FLOATING_POINT_CONSTANT})$"

NUMBER_VALIDATOR: Final[ re.Pattern[ str ]] = \
re.compile( NUMBER_PATTERN )

example = "123"
is_ok = bool( NUMBER_VALIDATOR.match( example ))
print( is_ok )

  (Actually, the Pythonic pattern uses "(" where I wrote "\".)

Newsgroups: comp.lang.python,alt.folklore.computers
Followup-To: comp.lang.python

[toc] | [prev] | [next] | [standalone]


#198073

Fromram@zedat.fu-berlin.de (Stefan Ram)
Date2026-09-25 19:19 +0000
Message-ID<order-20260925201802@ram.dialup.fu-berlin.de>
In reply to#198072
ram@zedat.fu-berlin.de (Stefan Ram) wrote or quoted:
>Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>>What is a good way to combine regular expressions in Python?  We are
>>cogitating the most /Pythonic/ way to make this happen.
>>    number:
>>        sign? integer-constant
>>        | sign? floating-point-constant
>>    integer-constant:
>>        digit-sequence
>>    floating-point-constant:
>>        fractional-constant exponent?
>>        | digit-sequence exponent
>>    fractional-constant:
>>        digit-sequence? "." digit-sequence
>>        | digit-sequence "."
>>    exponent:
>>        ( "e" | "E" ) sign? digit-sequence
>>    sign:
>>        "+" | "-"
>>    digit-sequence:
>>        digit
>>        | digit digit-sequence
>>    digit:
>>        "0" | "1" | "2" | "3" | "4" | "5" | "6" | "7" | "8" | "9"
>Maybe something like this?

  We can also write the definitions in the same order as in the EBNF.

import re
from typing import Final

NUMBER = lambda: rf"^{SIGN()}?(?:{INTEGER_CONSTANT()}|{FLOATING_POINT_CONSTANT()})$"

INTEGER_CONSTANT = lambda: DIGIT_SEQUENCE()

FLOATING_POINT_CONSTANT = lambda: \
rf"(?:{FRACTIONAL_CONSTANT()}(?:{EXPONENT()})?|{DIGIT_SEQUENCE()}{EXPONENT()})"

FRACTIONAL_CONSTANT = lambda: \
rf"(?:(?:{DIGIT_SEQUENCE()})?\.{DIGIT_SEQUENCE()}|{DIGIT_SEQUENCE()}\.)"

EXPONENT = lambda: rf"[eE]{SIGN()}?{DIGIT_SEQUENCE()}"

SIGN = lambda: r"[+-]"

DIGIT_SEQUENCE = lambda: rf"{DIGIT()}+"

DIGIT = lambda: r"[0-9]"

NUMBER_VALIDATOR: Final[ re.Pattern[ str ]] = re.compile( NUMBER() )

example = "123"
is_ok = bool( NUMBER_VALIDATOR.match( example ))
print( is_ok )

Newsgroups: comp.lang.python,alt.folklore.computers
Followup-To: comp.lang.python

[toc] | [prev] | [next] | [standalone]


#198078

Fromram@zedat.fu-berlin.de (Stefan Ram)
Date2026-09-26 10:04 +0000
Message-ID<names-20260926110359@ram.dialup.fu-berlin.de>
In reply to#198071
Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>We also do not want to introduce /any/ extra identifiers, nor paren-
>thesis.

  Here's without names and parens:

^[+-]?[0-9]+$|^[+-]?[0-9]+\.[0-9]+[eE][+-]?[0-9]+$|^[+-]?[0-9]+\.[0-9]+$|^[+-]?\.[0-9]+[eE][+-]?[0-9]+$|^[+-]?\.[0-9]+$|^[+-]?[0-9]+\.[eE][+-]?[0-9]+$|^[+-]?[0-9]+\.$|^[+-]?[0-9]+[eE][+-]?[0-9]+$

Newsgroups: comp.lang.python,alt.folklore.computers
Followup-To: comp.lang.python

[toc] | [prev] | [next] | [standalone]


#198079

FromJohann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid>
Date2026-09-26 20:09 +0800
Message-ID<TBOtS.247604$4yM7.183409@fx08.ams4>
In reply to#198078
On 9/26/2026 6:04 PM, Stefan Ram wrote:
> Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>> We also do not want to introduce /any/ extra identifiers, nor paren-
>> thesis.
> 
>    Here's without names and parens:
> 
> ^[+-]?[0-9]+$|^[+-]?[0-9]+\.[0-9]+[eE][+-]?[0-9]+$|^[+-]?[0-9]+\.[0-9]+$|^[+-]?\.[0-9]+[eE][+-]?[0-9]+$|^[+-]?\.[0-9]+$|^[+-]?[0-9]+\.[eE][+-]?[0-9]+$|^[+-]?[0-9]+\.$|^[+-]?[0-9]+[eE][+-]?[0-9]+$
> 
> Newsgroups: comp.lang.python,alt.folklore.computers
> Followup-To: comp.lang.python

Thank you.  I'm sure Kent Beck of /eXtreme Programming/ would be proud
of this regular expression if he bothered to read comp.lang.python.

It puts the write-only programming language of Perl to shame.  Do you
think he'd be well received on Pern?  We don't have alt.fantasy.anne-
mccaffrey, so I've only included alt.fantasy for further commentary on
Pernian Perl programming perennial illegibility.
-- 
Johann | email: invalid -> com | http://www.myrkraverk.com/blog/
I'm not from the Internet, I just work there. | via Easynews.com
https://bsky.app/profile/myrkraverk.bsky.social | for ( ;; ) _:;
Federated at https://fed.brid.gy/bsky/myrkraverk.bsky.social

[toc] | [prev] | [next] | [standalone]


#198080

Fromram@zedat.fu-berlin.de (Stefan Ram)
Date2026-09-26 12:22 +0000
Message-ID<VERBOSE-20260926132029@ram.dialup.fu-berlin.de>
In reply to#198079
Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> wrote or quoted:
>Thank you.  I'm sure Kent Beck of /eXtreme Programming/ would be proud
>of this regular expression if he bothered to read comp.lang.python.

  That regular expression can be made more readable in Python
  using verbose regular expressions while still avoiding names
  and parentheses within the regular expression.

exp = re.compile(
  r"""
  ^ [+-]? [0-9]+ $                             # Integer (123, -45)
  |
  ^ [+-]? [0-9]+ \. [0-9]+ [eE] [+-]? [0-9]+ $ # Float with "E"
  |
  ^ [+-]? [0-9]+ \. [0-9]+ $                   # Float without "E"
  |
  ^ [+-]? \. [0-9]+ [eE] [+-]? [0-9]+ $        # Starts with ".", has E
  |
  ^ [+-]? \. [0-9]+ $                          # Starts with ".", no "E"
  |
  ^ [+-]? [0-9]+ \. [eE] [+-]? [0-9]+ $        # Ends with ".", has "E"
  |
  ^ [+-]? [0-9]+ \. $                          # Ends with ".", no "E"
  |
  ^ [+-]? [0-9]+ [eE] [+-]? [0-9]+ $           # No ".", but "E"
  """,
  re.VERBOSE,
)

Newsgroups: comp.lang.python,alt.fantasy
Followup-To: comp.lang.python

[toc] | [prev] | [standalone]


Back to top | Article view | comp.lang.python


csiph-web