Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.javascript > #18082 > unrolled thread

Need regexp help

Started byjustaguy <lichunshen84@gmail.com>
First post2013-01-11 17:01 -0800
Last post2013-01-12 23:11 +0100
Articles 13 — 6 participants

Back to article view | Back to comp.lang.javascript


Contents

  Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-11 17:01 -0800
    Re: Need regexp help Richard Yates <richard@yatesguitar.com> - 2013-01-11 17:14 -0800
      Re: Need regexp help Richard Yates <richard@yatesguitar.com> - 2013-01-11 21:00 -0800
        Re: Need regexp help Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2013-01-12 16:16 +0100
          Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 16:27 +0100
          Re: Need regexp help Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2013-01-13 00:49 +0100
    Re: Need regexp help Luc Yen <luc@goal.tw> - 2013-01-11 22:10 -0800
    Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 09:25 +0100
      Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 06:33 -0800
    Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 10:22 -0800
      Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 10:48 -0800
        Re: Need regexp help SAM <stephanemoriaux.NoAdmin@wanadoo.fr.invalid> - 2013-01-13 22:02 +0100
      Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 23:11 +0100

#18082 — Need regexp help

Fromjustaguy <lichunshen84@gmail.com>
Date2013-01-11 17:01 -0800
SubjectNeed regexp help
Message-ID<eed9c2f1-0e66-4502-849d-9d538eddd1a8@googlegroups.com>
Hi,

How could use regexp (replace) to find all occurences of the following CSS
code in a long string and remove them?   Thanks.

<xmp>
body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 

td {
font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
</xmp>

[toc] | [next] | [standalone]


#18083

FromRichard Yates <richard@yatesguitar.com>
Date2013-01-11 17:14 -0800
Message-ID<d1e1f8d6shakdd5sjro0vcudj1utmsa4q7@4ax.com>
In reply to#18082
On Fri, 11 Jan 2013 17:01:25 -0800 (PST), justaguy
<lichunshen84@gmail.com> wrote:

>Hi,
>
>How could use regexp (replace) to find all occurences of the following CSS
>code in a long string and remove them?   Thanks.
>
><xmp>
>body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
>
>td {
>font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
></xmp>

Not sure why it's a regexp problem. Use indexOf() to find the location
of the section you want to remove, and then a couple substring() to
close up the gap.

[toc] | [prev] | [next] | [standalone]


#18086

FromRichard Yates <richard@yatesguitar.com>
Date2013-01-11 21:00 -0800
Message-ID<vcr1f8ptnnqptopbi07esct071uuo38htl@4ax.com>
In reply to#18083
On Fri, 11 Jan 2013 17:14:16 -0800, Richard Yates
<richard@yatesguitar.com> wrote:

>On Fri, 11 Jan 2013 17:01:25 -0800 (PST), justaguy
><lichunshen84@gmail.com> wrote:
>
>>Hi,
>>
>>How could use regexp (replace) to find all occurences of the following CSS
>>code in a long string and remove them?   Thanks.
>>
>><xmp>
>>body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
>>
>>td {
>>font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
>></xmp>
>
>Not sure why it's a regexp problem. Use indexOf() to find the location
>of the section you want to remove, and then a couple substring() to
>close up the gap.

var ind = haystack.indexOf(needle);
var newstring =
haystack.substring(0,ind)+haystack.substring(ind+needle.length);

[toc] | [prev] | [next] | [standalone]


#18090

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2013-01-12 16:16 +0100
Message-ID<16779241.CEys5BfS3R@PointedEars.de>
In reply to#18086
Richard Yates wrote:

> On Fri, 11 Jan 2013 17:14:16 -0800, Richard Yates
> <richard@yatesguitar.com> wrote:
>> justaguy <lichunshen84@gmail.com> wrote:
>>> How could use regexp (replace) to find all occurences of the following
                                           ^^^
>>> CSS code in a long string and remove them?   Thanks.
>>>
>>> <xmp>
>>> body { font-family : "Courier New", Courier, monospace; font-size : 9pt;
>>> valign : top; text-align : left; line-height: 9pt }
>>>
>>> td {
>>> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign
>>> : top; text-align : left; line-height: 9pt } </xmp>
>>
>> Not sure why it's a regexp problem. Use indexOf() to find the location
>> of the section you want to remove, and then a couple substring() to
>> close up the gap.
> 
> var ind = haystack.indexOf(needle);
> var newstring =
> haystack.substring(0,ind)+haystack.substring(ind+needle.length);

The key word here is “all”.  Your code would replace only *one* occurrence, 
so to replace *all* occurrences you would need to run it in a loop.  We had 
to do comparably inefficient stuff like that before regular expressions were 
introduced in ECMAScript Edition 3.  That was a little more than 13 years 
ago (1999-12).  (But even then we used String.prototype.replace() already.)

That said, one should not use regular expressions on a context-free language 
unless one really knows what one is doing.  In this exceptional case a 
*single* regular expression would be appropriate to parse CSS because the 
probability of a false positive for pattern matching can be reduced to close 
to zero.  Watch for word-wrap (or instead use RegExp.prototype.concat(), 
provided by jsx.regexp.concat() [1]):

  css = css.replace(/(^|\s*)body\s*\{\s*font-family\s*:\s*"Courier New",
\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;\s*valign\s*:\s*top;\s*text-
align\s*:\s*left;\s*line-height:\s*9pt\s*\}\s*td\s*\{\s*font-family\s*:
\s*"Courier New",\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;
\s*valign\s*:\s*top;text-align\s*:\s*left;\s*line-height:\s*9pt\s*\}/g, 
"$1");

This begs the question, however, why the OP thinks this would be necessary 
here in the first place.  I certainly would not recommend it.

BTW, the code to remove is not Valid (there is no “valign” CSS property to 
begin with; the property name is “vertical-align” [2]).  But that could be 
one reason why it is to be removed.

_________
[1] <http://PointedEars.de/scripts/test/regexp>
[2] <http://www.w3.org/TR/CSS21/>
-- 
PointedEars

Twitter: @PointedEars2
Please do not Cc: me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#18091

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2013-01-12 16:27 +0100
Message-ID<XnsA146A779FFBAeejj99@194.109.133.133>
In reply to#18090
Thomas 'PointedEars' Lahn wrote on 12 jan 2013 in comp.lang.javascript:

>> var ind = haystack.indexOf(needle);
>> var newstring =
>> haystack.substring(0,ind)+haystack.substring(ind+needle.length);
> 
> The key word here is ƒ oallƒ  .  Your code would replace only *one*
> occurrence, so to replace *all* occurrences you would need to run it
> in a loop.  

Probably a simpler way is to extract the part you want, 
instead of trying to delete all the stuf you do not want.

If you only want the body's content, 
and you know the body has both a single and clean <body>,
and a </body>:

s = s.split(/<body>/)[1].split(/<\/body>/)[0];

Not tested.

-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#18097

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2013-01-13 00:49 +0100
Message-ID<1373174.86QkPN8Rvz@PointedEars.de>
In reply to#18090
Thomas 'PointedEars' Lahn wrote:

> […]  In this exceptional case a *single* regular expression would be
> appropriate to parse CSS because the probability of a false positive for
> pattern matching can be reduced to close to zero.  Watch for word-wrap (or
> instead use RegExp.prototype.concat(), provided by jsx.regexp.concat()
> [1]):
> 
>   css = css.replace(/(^|\s*)body\s*\{\s*font-family\s*:\s*"Courier New",
> \s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;\s*valign\s*:\s*top;
\s*text-
> align\s*:\s*left;\s*line-height:\s*9pt\s*\}\s*td\s*\{\s*font-family\s*:
> \s*"Courier New",\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;
> \s*valign\s*:\s*top;text-align\s*:\s*left;\s*line-height:\s*9pt\s*\}/g,
> "$1");
> 
> This begs the question, however, why the OP thinks this would be necessary
> here in the first place.  I certainly would not recommend it.

A better readable version of this:

  var css = …;

  var needle = 'body { font-family : "Courier New", Courier, monospace;'
             + " font-size : 9pt; valign : top; text-align : left;"
             + " line-height: 9pt }"
             + ' td { font-family : "Courier New", Courier, monospace;'
             + " font-size : 9pt; valign : top; text-align : left;"
             + " line-height: 9pt }";

  var rxNeedle = new RegExp(
    "(^|\\s*)"
    + needle.replace(/("[^"]*")|\s+/g, function (match, doubleQuoted) {
        return doubleQuoted || "\\s*";
      }).replace(/[{}]/g, "\\$&"),
    "g");

  css = css.replace(rxNeedle, "");

See also RegExp.prototyp.escape() provided by jsx.regexp.escape(). [1]

> _________
> [1] <http://PointedEars.de/scripts/test/regexp>
-- 
PointedEars

Twitter: @PointedEars2
Please do not Cc: me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#18087

FromLuc Yen <luc@goal.tw>
Date2013-01-11 22:10 -0800
Message-ID<2a63dd21-aa61-41d1-82cb-d242305e8819@googlegroups.com>
In reply to#18082
justaguy於 2013年1月12日星期六UTC+8上午9時01分25秒寫道:
> Hi,
> 
> 
> 
> How could use regexp (replace) to find all occurences of the following CSS
> 
> code in a long string and remove them?   Thanks.
> 
> 
> 
> <xmp>
> 
> body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> 
> 
> td {
> 
> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> </xmp>

Not very clear about your intention. If you need to stripped out above body/td rules, maybe you can try:

/(body|td)\s*?{[^}]*}/gm

Sample code here:
var firstxmp = document.getElementsByTagName("xmp")[0];
firstxmp.innerHTML = firstxmp.innerHTML.replace(/(body|td)\s*?{[^}]*}/gm, "");

* css comments contain '}' character is not applied here *

[toc] | [prev] | [next] | [standalone]


#18088

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2013-01-12 09:25 +0100
Message-ID<XnsA1465FDD94A88eejj99@194.109.133.133>
In reply to#18082
justaguy wrote on 12 jan 2013 in comp.lang.javascript:

> Hi,
> 
> How could use regexp (replace) to find all occurences of the following
> CSS code in a long string and remove them?   Thanks.

Lacking your definition of "css-code", I will presume that is anything 
between { and }, stipulating those characters don't appear elsewhere in 
your string.

Without that stipulation your Q is unanswerable, 
as css-code is just normal text.

> <xmp>
> body { font-family : "Courier New", Courier, monospace; font-size :
> 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> td {
> font-family : "Courier New", Courier, monospace; font-size : 9pt;
> valign : top; text-align : left; line-height: 9pt } </xmp>

s = s.replace( /{[}]+}/g , '{}' );

-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#18089

Fromjustaguy <lichunshen84@gmail.com>
Date2013-01-12 06:33 -0800
Message-ID<86e2a9de-8651-49ed-969e-94068a6271c7@googlegroups.com>
In reply to#18088
Thanks all.

Well, all these CSS code, html code, and possibly dynamically generated javascript is from fedex incoming email, that is, in the BODY part of a fedex email. Btw, the XMP tag is what I added for server side output debugging.

And the reason I need to remove all these code is because it messed up my existing HTML code for server side output.

>On Saturday, January 12, 2013 3:25:26 AM UTC-5, Evertjan. wrote:
> justaguy wrote on 12 jan 2013 in comp.lang.javascript:
> 
> 
> 
> > Hi,
> 
> > 
> 
> > How could use regexp (replace) to find all occurences of the following
> 
> > CSS code in a long string and remove them?   Thanks.
> 
> 
> 
> Lacking your definition of "css-code", I will presume that is anything 
> 
> between { and }, stipulating those characters don't appear elsewhere in 
> 
> your string.
> 
> 
> 
> Without that stipulation your Q is unanswerable, 
> 
> as css-code is just normal text.
> 
> 
> 
> > <xmp>
> 
> > body { font-family : "Courier New", Courier, monospace; font-size :
> 
> > 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> > 
> 
> > td {
> 
> > font-family : "Courier New", Courier, monospace; font-size : 9pt;
> 
> > valign : top; text-align : left; line-height: 9pt } </xmp>
> 
> 
> 
> s = s.replace( /{[}]+}/g , '{}' );
> 
> 
> 
> -- 
> 
> Evertjan.
> 
> The Netherlands.
> 
> (Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [next] | [standalone]


#18092

Fromjustaguy <lichunshen84@gmail.com>
Date2013-01-12 10:22 -0800
Message-ID<f24e36f0-9161-4a1b-babe-73f58992c521@googlegroups.com>
In reply to#18082
Interesting idea.  In the meantime, I'd prefer not to use javascript syntax to process the lengthy data string, that is, the body content of an incoming email. So, in the case, I can't use split, which seems to be a js syntax.  Is there similar similar but just regexp?  Thanks.

>On Friday, January 11, 2013 8:01:25 PM UTC-5, justaguy wrote:
> Hi,
> 
> 
> 
> How could use regexp (replace) to find all occurences of the following CSS
> 
> code in a long string and remove them?   Thanks.
> 
> 
> 
> <xmp>
> 
> body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> 
> 
> td {
> 
> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> </xmp>

[toc] | [prev] | [next] | [standalone]


#18093

Fromjustaguy <lichunshen84@gmail.com>
Date2013-01-12 10:48 -0800
Message-ID<c148d94c-bdc4-4a79-a08f-39bda29a0eb1@googlegroups.com>
In reply to#18092
Some sample email body content below (I added xmp tag), thanks.

<xmp>
 monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
<table width="600" border="0" cellspacing="0" cellpadding="1" valign="top" align="left" bordercolor="#660099"> <br>
<table width="300" cellspacing="0" cellpadding="3"> 
<td width="%50">Ship (P/U) Date: 
<td width="%50">01/04/2013
 </table>

</xmp>

>On Saturday, January 12, 2013 1:22:16 PM UTC-5, justaguy wrote:
> Interesting idea.  In the meantime, I'd prefer not to use javascript syntax to process the lengthy data string, that is, the body content of an incoming email. So, in the case, I can't use split, which seems to be a js syntax.  Is there similar similar but just regexp?  Thanks.
> 
> 
> 
> >On Friday, January 11, 2013 8:01:25 PM UTC-5, justaguy wrote:
> 
> > Hi,
> 
> > 
> 
> > 
> 
> > 
> 
> > How could use regexp (replace) to find all occurences of the following CSS
> 
> > 
> 
> > code in a long string and remove them?   Thanks.
> 
> > 
> 
> > 
> 
> > 
> 
> > <xmp>
> 
> > 
> 
> > body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> > 
> 
> > 
> 
> > 
> 
> > td {
> 
> > 
> 
> > font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } 
> 
> > 
> 
> > </xmp>

[toc] | [prev] | [next] | [standalone]


#18100

FromSAM <stephanemoriaux.NoAdmin@wanadoo.fr.invalid>
Date2013-01-13 22:02 +0100
Message-ID<50f320e9$0$1374$ba4acef3@reader.news.orange.fr>
In reply to#18093
Le 12/01/13 19:48, justaguy a écrit :
> Some sample email body content below (I added xmp tag), thanks.
>
> <xmp>
>   monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
> <table width="600" border="0" cellspacing="0" cellpadding="1" valign="top" align="left" bordercolor="#660099"> <br>

???? What is this soup ?

> <table width="300" cellspacing="0" cellpadding="3">
> <td width="%50">Ship (P/U) Date:
> <td width="%50">01/04/2013
>   </table>
>
> </xmp>


-- 
Stéphane Moriaux avec/with iMac-intel

[toc] | [prev] | [next] | [standalone]


#18096

From"Evertjan." <exxjxw.hannivoort@inter.nl.net>
Date2013-01-12 23:11 +0100
Message-ID<XnsA146EBF5A9812eejj99@194.109.133.133>
In reply to#18092
justaguy wrote on 12 jan 2013 in comp.lang.javascript:

> Interesting idea.

What is? 
What idea you responding on?

> In the meantime, I'd prefer not to use javascript
> syntax to process the lengthy data string, that is, the body content
> of an incoming email. So, in the case, I can't use split, which seems
> to be a js syntax.  Is there similar similar but just regexp?  Thanks. 

May I ask what you are doing on this NG, 
if you do not want a javascript solution?

Non-javascript regex questions are off topic here.

-- 
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)

[toc] | [prev] | [standalone]


Back to top | Article view | comp.lang.javascript


csiph-web