Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.javascript > #18082 > unrolled thread
| Started by | justaguy <lichunshen84@gmail.com> |
|---|---|
| First post | 2013-01-11 17:01 -0800 |
| Last post | 2013-01-12 23:11 +0100 |
| Articles | 13 — 6 participants |
Back to article view | Back to comp.lang.javascript
Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-11 17:01 -0800
Re: Need regexp help Richard Yates <richard@yatesguitar.com> - 2013-01-11 17:14 -0800
Re: Need regexp help Richard Yates <richard@yatesguitar.com> - 2013-01-11 21:00 -0800
Re: Need regexp help Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2013-01-12 16:16 +0100
Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 16:27 +0100
Re: Need regexp help Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2013-01-13 00:49 +0100
Re: Need regexp help Luc Yen <luc@goal.tw> - 2013-01-11 22:10 -0800
Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 09:25 +0100
Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 06:33 -0800
Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 10:22 -0800
Re: Need regexp help justaguy <lichunshen84@gmail.com> - 2013-01-12 10:48 -0800
Re: Need regexp help SAM <stephanemoriaux.NoAdmin@wanadoo.fr.invalid> - 2013-01-13 22:02 +0100
Re: Need regexp help "Evertjan." <exxjxw.hannivoort@inter.nl.net> - 2013-01-12 23:11 +0100
| From | justaguy <lichunshen84@gmail.com> |
|---|---|
| Date | 2013-01-11 17:01 -0800 |
| Subject | Need regexp help |
| Message-ID | <eed9c2f1-0e66-4502-849d-9d538eddd1a8@googlegroups.com> |
Hi,
How could use regexp (replace) to find all occurences of the following CSS
code in a long string and remove them? Thanks.
<xmp>
body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
td {
font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
</xmp>
[toc] | [next] | [standalone]
| From | Richard Yates <richard@yatesguitar.com> |
|---|---|
| Date | 2013-01-11 17:14 -0800 |
| Message-ID | <d1e1f8d6shakdd5sjro0vcudj1utmsa4q7@4ax.com> |
| In reply to | #18082 |
On Fri, 11 Jan 2013 17:01:25 -0800 (PST), justaguy
<lichunshen84@gmail.com> wrote:
>Hi,
>
>How could use regexp (replace) to find all occurences of the following CSS
>code in a long string and remove them? Thanks.
>
><xmp>
>body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
>td {
>font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
></xmp>
Not sure why it's a regexp problem. Use indexOf() to find the location
of the section you want to remove, and then a couple substring() to
close up the gap.
[toc] | [prev] | [next] | [standalone]
| From | Richard Yates <richard@yatesguitar.com> |
|---|---|
| Date | 2013-01-11 21:00 -0800 |
| Message-ID | <vcr1f8ptnnqptopbi07esct071uuo38htl@4ax.com> |
| In reply to | #18083 |
On Fri, 11 Jan 2013 17:14:16 -0800, Richard Yates
<richard@yatesguitar.com> wrote:
>On Fri, 11 Jan 2013 17:01:25 -0800 (PST), justaguy
><lichunshen84@gmail.com> wrote:
>
>>Hi,
>>
>>How could use regexp (replace) to find all occurences of the following CSS
>>code in a long string and remove them? Thanks.
>>
>><xmp>
>>body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>>
>>td {
>>font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>></xmp>
>
>Not sure why it's a regexp problem. Use indexOf() to find the location
>of the section you want to remove, and then a couple substring() to
>close up the gap.
var ind = haystack.indexOf(needle);
var newstring =
haystack.substring(0,ind)+haystack.substring(ind+needle.length);
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2013-01-12 16:16 +0100 |
| Message-ID | <16779241.CEys5BfS3R@PointedEars.de> |
| In reply to | #18086 |
Richard Yates wrote:
> On Fri, 11 Jan 2013 17:14:16 -0800, Richard Yates
> <richard@yatesguitar.com> wrote:
>> justaguy <lichunshen84@gmail.com> wrote:
>>> How could use regexp (replace) to find all occurences of the following
^^^
>>> CSS code in a long string and remove them? Thanks.
>>>
>>> <xmp>
>>> body { font-family : "Courier New", Courier, monospace; font-size : 9pt;
>>> valign : top; text-align : left; line-height: 9pt }
>>>
>>> td {
>>> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign
>>> : top; text-align : left; line-height: 9pt } </xmp>
>>
>> Not sure why it's a regexp problem. Use indexOf() to find the location
>> of the section you want to remove, and then a couple substring() to
>> close up the gap.
>
> var ind = haystack.indexOf(needle);
> var newstring =
> haystack.substring(0,ind)+haystack.substring(ind+needle.length);
The key word here is “all”. Your code would replace only *one* occurrence,
so to replace *all* occurrences you would need to run it in a loop. We had
to do comparably inefficient stuff like that before regular expressions were
introduced in ECMAScript Edition 3. That was a little more than 13 years
ago (1999-12). (But even then we used String.prototype.replace() already.)
That said, one should not use regular expressions on a context-free language
unless one really knows what one is doing. In this exceptional case a
*single* regular expression would be appropriate to parse CSS because the
probability of a false positive for pattern matching can be reduced to close
to zero. Watch for word-wrap (or instead use RegExp.prototype.concat(),
provided by jsx.regexp.concat() [1]):
css = css.replace(/(^|\s*)body\s*\{\s*font-family\s*:\s*"Courier New",
\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;\s*valign\s*:\s*top;\s*text-
align\s*:\s*left;\s*line-height:\s*9pt\s*\}\s*td\s*\{\s*font-family\s*:
\s*"Courier New",\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;
\s*valign\s*:\s*top;text-align\s*:\s*left;\s*line-height:\s*9pt\s*\}/g,
"$1");
This begs the question, however, why the OP thinks this would be necessary
here in the first place. I certainly would not recommend it.
BTW, the code to remove is not Valid (there is no “valign” CSS property to
begin with; the property name is “vertical-align” [2]). But that could be
one reason why it is to be removed.
_________
[1] <http://PointedEars.de/scripts/test/regexp>
[2] <http://www.w3.org/TR/CSS21/>
--
PointedEars
Twitter: @PointedEars2
Please do not Cc: me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | "Evertjan." <exxjxw.hannivoort@inter.nl.net> |
|---|---|
| Date | 2013-01-12 16:27 +0100 |
| Message-ID | <XnsA146A779FFBAeejj99@194.109.133.133> |
| In reply to | #18090 |
Thomas 'PointedEars' Lahn wrote on 12 jan 2013 in comp.lang.javascript: >> var ind = haystack.indexOf(needle); >> var newstring = >> haystack.substring(0,ind)+haystack.substring(ind+needle.length); > > The key word here is ƒ oallƒ . Your code would replace only *one* > occurrence, so to replace *all* occurrences you would need to run it > in a loop. Probably a simpler way is to extract the part you want, instead of trying to delete all the stuf you do not want. If you only want the body's content, and you know the body has both a single and clean <body>, and a </body>: s = s.split(/<body>/)[1].split(/<\/body>/)[0]; Not tested. -- Evertjan. The Netherlands. (Please change the x'es to dots in my emailaddress)
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2013-01-13 00:49 +0100 |
| Message-ID | <1373174.86QkPN8Rvz@PointedEars.de> |
| In reply to | #18090 |
Thomas 'PointedEars' Lahn wrote:
> […] In this exceptional case a *single* regular expression would be
> appropriate to parse CSS because the probability of a false positive for
> pattern matching can be reduced to close to zero. Watch for word-wrap (or
> instead use RegExp.prototype.concat(), provided by jsx.regexp.concat()
> [1]):
>
> css = css.replace(/(^|\s*)body\s*\{\s*font-family\s*:\s*"Courier New",
> \s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;\s*valign\s*:\s*top;
\s*text-
> align\s*:\s*left;\s*line-height:\s*9pt\s*\}\s*td\s*\{\s*font-family\s*:
> \s*"Courier New",\s*Courier,\s*monospace;\s*font-size\s*:\s*9pt;
> \s*valign\s*:\s*top;text-align\s*:\s*left;\s*line-height:\s*9pt\s*\}/g,
> "$1");
>
> This begs the question, however, why the OP thinks this would be necessary
> here in the first place. I certainly would not recommend it.
A better readable version of this:
var css = …;
var needle = 'body { font-family : "Courier New", Courier, monospace;'
+ " font-size : 9pt; valign : top; text-align : left;"
+ " line-height: 9pt }"
+ ' td { font-family : "Courier New", Courier, monospace;'
+ " font-size : 9pt; valign : top; text-align : left;"
+ " line-height: 9pt }";
var rxNeedle = new RegExp(
"(^|\\s*)"
+ needle.replace(/("[^"]*")|\s+/g, function (match, doubleQuoted) {
return doubleQuoted || "\\s*";
}).replace(/[{}]/g, "\\$&"),
"g");
css = css.replace(rxNeedle, "");
See also RegExp.prototyp.escape() provided by jsx.regexp.escape(). [1]
> _________
> [1] <http://PointedEars.de/scripts/test/regexp>
--
PointedEars
Twitter: @PointedEars2
Please do not Cc: me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Luc Yen <luc@goal.tw> |
|---|---|
| Date | 2013-01-11 22:10 -0800 |
| Message-ID | <2a63dd21-aa61-41d1-82cb-d242305e8819@googlegroups.com> |
| In reply to | #18082 |
justaguy於 2013年1月12日星期六UTC+8上午9時01分25秒寫道:
> Hi,
>
>
>
> How could use regexp (replace) to find all occurences of the following CSS
>
> code in a long string and remove them? Thanks.
>
>
>
> <xmp>
>
> body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
>
>
> td {
>
> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
> </xmp>
Not very clear about your intention. If you need to stripped out above body/td rules, maybe you can try:
/(body|td)\s*?{[^}]*}/gm
Sample code here:
var firstxmp = document.getElementsByTagName("xmp")[0];
firstxmp.innerHTML = firstxmp.innerHTML.replace(/(body|td)\s*?{[^}]*}/gm, "");
* css comments contain '}' character is not applied here *
[toc] | [prev] | [next] | [standalone]
| From | "Evertjan." <exxjxw.hannivoort@inter.nl.net> |
|---|---|
| Date | 2013-01-12 09:25 +0100 |
| Message-ID | <XnsA1465FDD94A88eejj99@194.109.133.133> |
| In reply to | #18082 |
justaguy wrote on 12 jan 2013 in comp.lang.javascript:
> Hi,
>
> How could use regexp (replace) to find all occurences of the following
> CSS code in a long string and remove them? Thanks.
Lacking your definition of "css-code", I will presume that is anything
between { and }, stipulating those characters don't appear elsewhere in
your string.
Without that stipulation your Q is unanswerable,
as css-code is just normal text.
> <xmp>
> body { font-family : "Courier New", Courier, monospace; font-size :
> 9pt; valign : top; text-align : left; line-height: 9pt }
>
> td {
> font-family : "Courier New", Courier, monospace; font-size : 9pt;
> valign : top; text-align : left; line-height: 9pt } </xmp>
s = s.replace( /{[}]+}/g , '{}' );
--
Evertjan.
The Netherlands.
(Please change the x'es to dots in my emailaddress)
[toc] | [prev] | [next] | [standalone]
| From | justaguy <lichunshen84@gmail.com> |
|---|---|
| Date | 2013-01-12 06:33 -0800 |
| Message-ID | <86e2a9de-8651-49ed-969e-94068a6271c7@googlegroups.com> |
| In reply to | #18088 |
Thanks all.
Well, all these CSS code, html code, and possibly dynamically generated javascript is from fedex incoming email, that is, in the BODY part of a fedex email. Btw, the XMP tag is what I added for server side output debugging.
And the reason I need to remove all these code is because it messed up my existing HTML code for server side output.
>On Saturday, January 12, 2013 3:25:26 AM UTC-5, Evertjan. wrote:
> justaguy wrote on 12 jan 2013 in comp.lang.javascript:
>
>
>
> > Hi,
>
> >
>
> > How could use regexp (replace) to find all occurences of the following
>
> > CSS code in a long string and remove them? Thanks.
>
>
>
> Lacking your definition of "css-code", I will presume that is anything
>
> between { and }, stipulating those characters don't appear elsewhere in
>
> your string.
>
>
>
> Without that stipulation your Q is unanswerable,
>
> as css-code is just normal text.
>
>
>
> > <xmp>
>
> > body { font-family : "Courier New", Courier, monospace; font-size :
>
> > 9pt; valign : top; text-align : left; line-height: 9pt }
>
> >
>
> > td {
>
> > font-family : "Courier New", Courier, monospace; font-size : 9pt;
>
> > valign : top; text-align : left; line-height: 9pt } </xmp>
>
>
>
> s = s.replace( /{[}]+}/g , '{}' );
>
>
>
> --
>
> Evertjan.
>
> The Netherlands.
>
> (Please change the x'es to dots in my emailaddress)
[toc] | [prev] | [next] | [standalone]
| From | justaguy <lichunshen84@gmail.com> |
|---|---|
| Date | 2013-01-12 10:22 -0800 |
| Message-ID | <f24e36f0-9161-4a1b-babe-73f58992c521@googlegroups.com> |
| In reply to | #18082 |
Interesting idea. In the meantime, I'd prefer not to use javascript syntax to process the lengthy data string, that is, the body content of an incoming email. So, in the case, I can't use split, which seems to be a js syntax. Is there similar similar but just regexp? Thanks.
>On Friday, January 11, 2013 8:01:25 PM UTC-5, justaguy wrote:
> Hi,
>
>
>
> How could use regexp (replace) to find all occurences of the following CSS
>
> code in a long string and remove them? Thanks.
>
>
>
> <xmp>
>
> body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
>
>
> td {
>
> font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
> </xmp>
[toc] | [prev] | [next] | [standalone]
| From | justaguy <lichunshen84@gmail.com> |
|---|---|
| Date | 2013-01-12 10:48 -0800 |
| Message-ID | <c148d94c-bdc4-4a79-a08f-39bda29a0eb1@googlegroups.com> |
| In reply to | #18092 |
Some sample email body content below (I added xmp tag), thanks.
<xmp>
monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
<table width="600" border="0" cellspacing="0" cellpadding="1" valign="top" align="left" bordercolor="#660099"> <br>
<table width="300" cellspacing="0" cellpadding="3">
<td width="%50">Ship (P/U) Date:
<td width="%50">01/04/2013
</table>
</xmp>
>On Saturday, January 12, 2013 1:22:16 PM UTC-5, justaguy wrote:
> Interesting idea. In the meantime, I'd prefer not to use javascript syntax to process the lengthy data string, that is, the body content of an incoming email. So, in the case, I can't use split, which seems to be a js syntax. Is there similar similar but just regexp? Thanks.
>
>
>
> >On Friday, January 11, 2013 8:01:25 PM UTC-5, justaguy wrote:
>
> > Hi,
>
> >
>
> >
>
> >
>
> > How could use regexp (replace) to find all occurences of the following CSS
>
> >
>
> > code in a long string and remove them? Thanks.
>
> >
>
> >
>
> >
>
> > <xmp>
>
> >
>
> > body { font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
> >
>
> >
>
> >
>
> > td {
>
> >
>
> > font-family : "Courier New", Courier, monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt }
>
> >
>
> > </xmp>
[toc] | [prev] | [next] | [standalone]
| From | SAM <stephanemoriaux.NoAdmin@wanadoo.fr.invalid> |
|---|---|
| Date | 2013-01-13 22:02 +0100 |
| Message-ID | <50f320e9$0$1374$ba4acef3@reader.news.orange.fr> |
| In reply to | #18093 |
Le 12/01/13 19:48, justaguy a écrit : > Some sample email body content below (I added xmp tag), thanks. > > <xmp> > monospace; font-size : 9pt; valign : top; text-align : left; line-height: 9pt } > <table width="600" border="0" cellspacing="0" cellpadding="1" valign="top" align="left" bordercolor="#660099"> <br> ???? What is this soup ? > <table width="300" cellspacing="0" cellpadding="3"> > <td width="%50">Ship (P/U) Date: > <td width="%50">01/04/2013 > </table> > > </xmp> -- Stéphane Moriaux avec/with iMac-intel
[toc] | [prev] | [next] | [standalone]
| From | "Evertjan." <exxjxw.hannivoort@inter.nl.net> |
|---|---|
| Date | 2013-01-12 23:11 +0100 |
| Message-ID | <XnsA146EBF5A9812eejj99@194.109.133.133> |
| In reply to | #18092 |
justaguy wrote on 12 jan 2013 in comp.lang.javascript: > Interesting idea. What is? What idea you responding on? > In the meantime, I'd prefer not to use javascript > syntax to process the lengthy data string, that is, the body content > of an incoming email. So, in the case, I can't use split, which seems > to be a js syntax. Is there similar similar but just regexp? Thanks. May I ask what you are doing on this NG, if you do not want a javascript solution? Non-javascript regex questions are off topic here. -- Evertjan. The Netherlands. (Please change the x'es to dots in my emailaddress)
[toc] | [prev] | [standalone]
Back to top | Article view | comp.lang.javascript
csiph-web