Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #16497 > unrolled thread

Any recommendations on using PHP to parse marked-up text?

Started byJames Harris <james.harris.1@gmail.com>
First post2016-02-18 12:47 +0000
Last post2016-02-27 15:58 +1300
Articles 20 on this page of 50 — 8 participants

Back to article view | Back to comp.lang.php


Contents

  Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-18 12:47 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:52 +0100
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:28 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 14:56 +0000
      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:14 +0100
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 21:45 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 23:09 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 01:15 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-19 02:45 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 12:03 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? Tim Streater <timstreater@greenbee.net> - 2016-02-19 12:17 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-19 14:18 +0100
                    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 14:02 +0000
                      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:31 +0100
                        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 01:39 +0000
                          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:33 +0100
                            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 12:53 +0000
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:29 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:32 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:43 -0500
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:51 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 20:08 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:22 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 23:40 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:39 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 13:09 +0000
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 20:24 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 22:35 +0000
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:37 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 10:55 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:25 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:50 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:49 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:05 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:45 -0500
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 09:01 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:46 +0100
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 12:36 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:39 +0100
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:40 +0000
                    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:46 +0100
                      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 14:57 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:10 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:50 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:08 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-21 18:35 +0100
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:55 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-22 16:38 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:06 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Ian Collins <ian-news@hotmail.com> - 2016-02-27 15:58 +1300

Page 1 of 3  [1] 2 3  Next page →


#16497 — Any recommendations on using PHP to parse marked-up text?

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-18 12:47 +0000
SubjectAny recommendations on using PHP to parse marked-up text?
Message-ID<na4eau$at9$1@dont-email.me>
Quick question: If someone wanted to parse marked-up text (and to 
convert it to HTML) do you have any generic recommendation on a good 
approach to use in PHP?

Specifically, the text could be parsed character-by-character or line by 
line. Or perhaps it could be split with explode() and processed in pieces.

The input text could be ASCII or Unicode.

The marking up would be by tags of some sort. The form of a tag is not 
yet defined; it could be chosen to make the parsing easier, if 
necessary, as long as it was easy for a human to work with.

James

[toc] | [next] | [standalone]


#16500

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-18 15:52 +0100
Message-ID<56C5DAAB.1070307@arnowelzel.de>
In reply to#16497
James Harris schrieb am 2016-02-18 um 13:47:

> Quick question: If someone wanted to parse marked-up text (and to 
> convert it to HTML) do you have any generic recommendation on a good 
> approach to use in PHP?

No.

> Specifically, the text could be parsed character-by-character or line by 
> line. Or perhaps it could be split with explode() and processed in pieces.

Then go on and use explode() or split or whatever.

> The input text could be ASCII or Unicode.
> 
> The marking up would be by tags of some sort. The form of a tag is not 
> yet defined; it could be chosen to make the parsing easier, if 
> necessary, as long as it was easy for a human to work with.

Which sounds like "the input can be anything and I need to parse it".

Decide what format you want to use and maybe someone knows a good parser
for that format. If you create your own format you have to build your
own parser for it - and no, there is no thing like a "universal text
parser" in PHP.


-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


#16528

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 18:28 +0000
Message-ID<naab2e$mmd$1@dont-email.me>
In reply to#16500
On 18/02/2016 14:52, Arno Welzel wrote:
> James Harris schrieb am 2016-02-18 um 13:47:

...

>> The input text could be ASCII or Unicode.
>>
>> The marking up would be by tags of some sort. The form of a tag is not
>> yet defined; it could be chosen to make the parsing easier, if
>> necessary, as long as it was easy for a human to work with.
>
> Which sounds like "the input can be anything and I need to parse it".

Yep. The markup format can be chosen to make it easier to parse, as long 
as it is also easy for a human to work with.

> Decide what format you want to use and maybe someone knows a good parser
> for that format. If you create your own format you have to build your
> own parser for it - and no, there is no thing like a "universal text
> parser" in PHP.

Although there is no predefined markup form there are some universal 
requirements. Some/most markup tags will start a new context (which 
other tags will terminate). They are ones which would be converted to 
HTML opening and closing tags such as <h2>...</h2>.

But other markup tags would be self-contained and will not need 
corresponding closing tags. They would be converted to such as <hr>.

James

[toc] | [prev] | [next] | [standalone]


#16501

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-18 14:56 +0000
Message-ID<87twl66p9h.fsf@bsb.me.uk>
In reply to#16497
James Harris <james.harris.1@gmail.com> writes:

> Quick question: If someone wanted to parse marked-up text (and to
> convert it to HTML) do you have any generic recommendation on a good
> approach to use in PHP?
>
> Specifically, the text could be parsed character-by-character or line
> by line. Or perhaps it could be split with explode() and processed in
> pieces.

I would process it with preg_replace_callback -- possibly more that
once.  That's just about the most general way, but it means you are less
likely to get stuck if your markup gets more complicated than originally
thought.

It's general enough to be handle nested constructs because Perl's REs
are, in practice, much more powerful than REs should be.  It's good to
about nested construct (which I'd do by repeated scanning from the
inside out) but it's there for those cases where nothing else will do.

You might want to consider using one of the existing schemes (even f you
have to extend it) like markdown.

<snip>
-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16513

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-18 22:14 +0100
Message-ID<3816086.GMaWPedH5Q@PointedEars.de>
In reply to#16501
Ben Bacarisse wrote:

> James Harris <james.harris.1@gmail.com> writes:
>> Quick question: If someone wanted to parse marked-up text (and to
>> convert it to HTML) do you have any generic recommendation on a good
>> approach to use in PHP?
>> […]
> 
> I would process it with preg_replace_callback -- possibly more that
                           ^^^
> once.  That's just about the most general way, but it means you are less
> likely to get stuck if your markup gets more complicated than originally
> thought.

Sigh. [psf 10.1]

<https://en.wikipedia.org/wiki/Chomsky_hierarchy>
 
-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16515

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-18 21:45 +0000
Message-ID<87bn7d7kvl.fsf@bsb.me.uk>
In reply to#16513
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> James Harris <james.harris.1@gmail.com> writes:
>>> Quick question: If someone wanted to parse marked-up text (and to
>>> convert it to HTML) do you have any generic recommendation on a good
>>> approach to use in PHP?
>>> […]
>> 
>> I would process it with preg_replace_callback -- possibly more that
>                            ^^^
>> once.  That's just about the most general way, but it means you are less
>> likely to get stuck if your markup gets more complicated than originally
>> thought.
>
> Sigh. [psf 10.1]
>
> <https://en.wikipedia.org/wiki/Chomsky_hierarchy>

Are these blank references supposed to convey some useful point?

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16516

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-18 23:09 +0100
Message-ID<1615449.BjJPyqVnO6@PointedEars.de>
In reply to#16515
Ben Bacarisse wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>> Ben Bacarisse wrote:
>>> James Harris <james.harris.1@gmail.com> writes:
>>>> Quick question: If someone wanted to parse marked-up text (and to
>>>> convert it to HTML) do you have any generic recommendation on a good
>>>> approach to use in PHP?
>>>> […]
>>> I would process it with preg_replace_callback -- possibly more that
>>                           ^^^
>>> once.  That's just about the most general way, but it means you are less
>>> likely to get stuck if your markup gets more complicated than originally
>>> thought.
>>
>> Sigh. [psf 10.1]
>>
>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy>
> 
> Are these blank references

I would not call them “blank”.

> supposed to convey some useful point?

Yes.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16520

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-19 01:15 +0000
Message-ID<87lh6h5wlb.fsf@bsb.me.uk>
In reply to#16516
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>>> Ben Bacarisse wrote:
>>>> James Harris <james.harris.1@gmail.com> writes:
>>>>> Quick question: If someone wanted to parse marked-up text (and to
>>>>> convert it to HTML) do you have any generic recommendation on a good
>>>>> approach to use in PHP?
>>>>> […]
>>>> I would process it with preg_replace_callback -- possibly more that
>>>                           ^^^
>>>> once.  That's just about the most general way, but it means you are less
>>>> likely to get stuck if your markup gets more complicated than originally
>>>> thought.
>>>
>>> Sigh. [psf 10.1]
>>>
>>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy>
>> 
>> Are these blank references
>
> I would not call them “blank”.

psf 10.1 is blank to me, and the second simply summaries things that I
would assume almost everyone here knows (I certainly do).  There's no
hint at what you intend to say by posting it.  Maybe you are simply
suggesting some people might like to read it?

>> supposed to convey some useful point?
>
> Yes.

Good.  I hope whoever it was directed at understood it.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16521

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-19 02:45 +0100
Message-ID<2353152.leaMV68clc@PointedEars.de>
In reply to#16520
Ben Bacarisse wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>> Ben Bacarisse wrote:
>>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>>>> Sigh. [psf 10.1]
>>>>
>>>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy>
>>> 
>>> Are these blank references
>>
>> I would not call them “blank”.
> 
> psf 10.1 is blank to me,

STFW.

> and the second simply summaries things that I would assume almost everyone
> here knows (I certainly do). 

Maybe you know it, but you do not understand either it, or how 
preg_replace_callback() works, or both.

> There's no hint at what you intend to say by posting it.

But there is.  I have also marked the most relevant part of your posting.  
Unfortunately, I cannot quote it again *and* post it.

> Maybe you are simply suggesting some people might like to read it?

Yes, and draw the correct conclusions from it – whereas “some people” 
includes you.
 
>>> supposed to convey some useful point?
>> Yes.
> 
> Good.  I hope whoever it was directed at understood it.

It was directed at you, of course.
 
-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16522

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-19 12:03 +0000
Message-ID<874md46h6a.fsf@bsb.me.uk>
In reply to#16521
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>>> Ben Bacarisse wrote:
>>>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>>>>> Sigh. [psf 10.1]
>>>>>
>>>>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy>
>>>> 
>>>> Are these blank references
>>>
>>> I would not call them “blank”.
>> 
>> psf 10.1 is blank to me,
>
> STFW.

Now, why do you think I didn't?  You make hae no respect for the
knowledge and experience of others, but I do.  Why you say something I
try to understand it.  I could attach no meaning to those letters and
symbols.  That maybe DuckDuckGo's problem, but since you are so keen on
citations and good references I'm surprised.

>> and the second simply summaries things that I would assume almost everyone
>> here knows (I certainly do). 
>
> Maybe you know it, but you do not understand either it, or how 
> preg_replace_callback() works, or both.

Maybe.  Since you won't say what you mean by pointing me at that page,
we'll never know.  Maybe it you who doesn't understand one or other of
those things.  Where the confusion lies seems likely to remain a mystery
to both of us.

>> There's no hint at what you intend to say by posting it.
>
> But there is.  I have also marked the most relevant part of your posting.  
> Unfortunately, I cannot quote it again *and* post it.

Yes, it was obvious you were making some reference regular languages and
regular expressions (please don't take me for a fool) but what was the
point?  Were you agreeing with me?  Were you disagreeing?  Did I make
some mistake you are helpful correcting?  Was my advice bad for some
reason that "psf 10.1" or the WP page explains?

>> Maybe you are simply suggesting some people might like to read it?
>
> Yes, and draw the correct conclusions from it – whereas “some people” 
> includes you.
>  
>>>> supposed to convey some useful point?
>>> Yes.
>> 
>> Good.  I hope whoever it was directed at understood it.
>
> It was directed at you, of course.

Unless you are prepared to actually say what you mean we have to leave
it at that.  I read the page, I searched for psf 10.1, and I still have
no idea what you want to say.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16523

FromTim Streater <timstreater@greenbee.net>
Date2016-02-19 12:17 +0000
Message-ID<190220161217359137%timstreater@greenbee.net>
In reply to#16522
In article <874md46h6a.fsf@bsb.me.uk>, Ben Bacarisse
<ben.usenet@bsb.me.uk> wrote:

>Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>
>> Ben Bacarisse wrote:

>>> psf 10.1 is blank to me,
>>
>> STFW.
>
>Now, why do you think I didn't?  You may have no respect for the
>knowledge and experience of others, but I do.  When you say something I
>try to understand it.  I could attach no meaning to those letters and
>symbols.  That maybe DuckDuckGo's problem, but since you are so keen on
>citations and good references I'm surprised.

Just giving an acronym in that manner tells the other person that you
really can't be bothered to deal with them. Not the sort of person who
should be a teacher of any kind. It's just plain rude.

[snip]

>>> Good.  I hope whoever it was directed at understood it.
>>
>> It was directed at you, of course.
>
>Unless you are prepared to actually say what you mean we have to leave
>it at that.  I read the page, I searched for psf 10.1, and I still have
>no idea what you want to say.

Are you sure he actually wants to help people? I'm not.

-- 
New Socialism consists essentially in being seen to have your heart in 
the right place whilst your head is in the clouds and your hand is in 
someone else's pocket.

[toc] | [prev] | [next] | [standalone]


#16524

From"Christoph M. Becker" <cmbecker69@arcor.de>
Date2016-02-19 14:18 +0100
Message-ID<na74nd$4v6$1@solani.org>
In reply to#16522
Ben Bacarisse wrote:

> Unless you are prepared to actually say what you mean we have to leave
> it at that.  I read the page, I searched for psf 10.1, and I still have
> no idea what you want to say.

Googling for "psf 10.1" brought the desired information:

<http://pointedears.de/psf/index.en#sigh>

Anyhow, what Thomas most likely wanted to convey is that
preg_replace_callback() is not general enough to process all markup
languages.

-- 
Christoph M. Becker

[toc] | [prev] | [next] | [standalone]


#16525

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-19 14:02 +0000
Message-ID<87wpq04x3b.fsf@bsb.me.uk>
In reply to#16524
"Christoph M. Becker" <cmbecker69@arcor.de> writes:

> Ben Bacarisse wrote:
>
>> Unless you are prepared to actually say what you mean we have to leave
>> it at that.  I read the page, I searched for psf 10.1, and I still have
>> no idea what you want to say.
>
> Googling for "psf 10.1" brought the desired information:
>
> <http://pointedears.de/psf/index.en#sigh>

Oh!  That did not show up in my search engine.  Why on Earth would he
not just cite that link -- he complains enough about other people not
citing things properly?  So the "psf 10.1" was just to show me that he
has listed and numbered things he often says?  Blimey.
That's... well... it's very Usenet.

> Anyhow, what Thomas most likely wanted to convey is that
> preg_replace_callback() is not general enough to process all markup
> languages.

Seems unlikely.  Since his reference was to a theoretical article, the
answer would be equally theoretical, and there is no more powerful model
than even the simplest recogniser linked with a callback to a Turing
complete language.  Anyway, there are lots of things he might have
meant.  I wonder if he'll ever say?

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16533

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-20 20:31 +0100
Message-ID<1504490.9PMmeY6QeD@PointedEars.de>
In reply to#16525
Ben Bacarisse wrote:

> "Christoph M. Becker" <cmbecker69@arcor.de> writes:
>> Anyhow, what Thomas most likely wanted to convey is that
>> preg_replace_callback() is not general enough to process all markup
>> languages.
> 
> Seems unlikely. 

How did you get that idea?  Christoph has understood me correctly.

> Since his reference was to a theoretical article, the answer would be
> equally theoretical, […]

It is theoretical, but it has important implications in practice.  A 
mathematician (or, at least an individual that – the record shows – appears 
to have much interest in mathematics) like yourself should be able to 
appreciate that.  And I had expected such a person to understand the 
implications without needing to be spelled out for them.  Do not blame me 
for your obtuseness.

> I wonder if he'll ever say?

Since some people need to have it spelled out for them, the not-so-new 
argument (not even new in this newsgroup) goes as follows:

  0. The Chomsky hierarchy classifies formal languages into four types,
     depending on the type of formal grammars that they can be produced
     with: Recursively enumerable (0), context-sensitive (1),
     context-free (2), and regular (3).

  1. (It can be showed by application of the pumping lemma for context-free
     languages that) markup languages are context-free languages.

     For example, the simplest markup language (that IMO deserves that
     designation) consists of words in Latin letters that are surrounded by 
     pairs of matching parentheses (also called a “bracket language”;
     it helps to think of those parentheses as start-tag and end-tag of
     an element, and the words as its content, respectively).  A context
     free grammar that produces it contains the following productions:

       S → (S)
       S → SS
       S → a
       …
       S → z

     [Example application: S → (S) → (SS) → (SSS) → (foo)]

  2. From the Chomsky hierarchy follows by applying *set theory*: While 
     “every regular language is context-free” (ibid.), not every
     context-free language is regular.

  3. *Regular* expressions are (WLOG) a way to describe *regular*
     languages *only*.

  4. It can be showed by application of the pumping lemma for regular
     languages that there are markup (and programming) languages that
     are not regular.  For example, XML, despite its well-formedness,
     is not a regular language.

  5. Therefore, *regular* expressions are in general not suited to parse 
     arbitrary *markup* (or programming) languages; instead, implementations
     of non-deterministic push-down automata (stacks) are.

  6. preg_replace_callback(), as all preg_*() functions in PHP, supports 
     Perl-Compatible *Regular* Expressions (PCRE).  <http://php.net/pcre>

  7. While PCRE support recursion to some extent (which implements a stack),
     and therefore can to some extent solve the problem that regular
     expressions when applied to arbitrary bracket languages always either
     match too much (greedy) or too little (non-greedy), a *single* PCRE
     fails to parse an arbitrary markup language at a fundamental
     level: For a non-trivial markup language you would have to use 
     alternation, and the *first* *next* match wins, not the *next* 
     *longest* one, different to what would be required by a markup parser.

With that said, regular expressions can be useful in parsing markup 
languages (BTDT).  But they should not be the first consideration when 
attempting to do that.  There are tools better suited to the purpose.

See also <http://stackoverflow.com/a/1732454/855543> :)

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16552

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-21 01:39 +0000
Message-ID<87mvqu2655.fsf@bsb.me.uk>
In reply to#16533
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> "Christoph M. Becker" <cmbecker69@arcor.de> writes:
>>> Anyhow, what Thomas most likely wanted to convey is that
>>> preg_replace_callback() is not general enough to process all markup
>>> languages.
>> 
>> Seems unlikely. 
>
> How did you get that idea?  Christoph has understood me correctly.

Hmm... I wonder why you just cut, then, without comment my explanation
of why that's wrong.  Do you have nothing to say about that part?

>> Since his reference was to a theoretical article, the answer would be
>> equally theoretical, […]
>
> It is theoretical, but it has important implications in practice.  A 
> mathematician (or, at least an individual that – the record shows – appears 
> to have much interest in mathematics) like yourself should be able to 
> appreciate that.  And I had expected such a person to understand the 
> implications without needing to be spelled out for them.  Do not blame me 
> for your obtuseness.
>
>> I wonder if he'll ever say?
>
> Since some people need to have it spelled out for them, the not-so-new 
> argument (not even new in this newsgroup) goes as follows:
>
>   0. The Chomsky hierarchy classifies formal languages into four types,
>      depending on the type of formal grammars that they can be produced
>      with: Recursively enumerable (0), context-sensitive (1),
>      context-free (2), and regular (3).
>
>   1. (It can be showed by application of the pumping lemma for context-free
>      languages that) markup languages are context-free languages.
>
>      For example, the simplest markup language (that IMO deserves that
>      designation) consists of words in Latin letters that are surrounded by 
>      pairs of matching parentheses (also called a “bracket language”;
>      it helps to think of those parentheses as start-tag and end-tag of
>      an element, and the words as its content, respectively).  A context
>      free grammar that produces it contains the following productions:
>
>        S → (S)
>        S → SS
>        S → a
>        …
>        S → z
>
>      [Example application: S → (S) → (SS) → (SSS) → (foo)]

And yet '/( [a-z]+ | \( (?R) \) )+/x' matches the strings of this
language.  Maybe you just thought that saying "pumping lemma" would
sound intimidating?  I may have got the details wrong -- it's late --
but it is well known that a PCRE can match the strings (and only the
strings) of your example language.

>   2. From the Chomsky hierarchy follows by applying *set theory*: While 
>      “every regular language is context-free” (ibid.), not every
>      context-free language is regular.
>
>   3. *Regular* expressions are (WLOG) a way to describe *regular*
>      languages *only*.
>
>   4. It can be showed by application of the pumping lemma for regular
>      languages that there are markup (and programming) languages that
>      are not regular.  For example, XML, despite its well-formedness,
>      is not a regular language.
>
>   5. Therefore, *regular* expressions are in general not suited to parse 
>      arbitrary *markup* (or programming) languages; instead, implementations
>      of non-deterministic push-down automata (stacks) are.
>
>   6. preg_replace_callback(), as all preg_*() functions in PHP, supports 
>      Perl-Compatible *Regular* Expressions (PCRE).
>      <http://php.net/pcre>

All this is just bluster.  Regular languages are not relevant here.
What regular language does '/([ab]*)\1/' match?  Can you even say where
in Chomsky's hierarchy the languages matched by PCREs lie?

>   7. While PCRE support recursion to some extent (which implements a stack),
>      and therefore can to some extent solve the problem that regular
>      expressions when applied to arbitrary bracket languages always either
>      match too much (greedy) or too little (non-greedy), a *single* PCRE
>      fails to parse an arbitrary markup language at a fundamental
>      level: For a non-trivial markup language you would have to use 
>      alternation, and the *first* *next* match wins, not the *next* 
>      *longest* one, different to what would be required by a markup
>      parser.

And this just shows you didn't read what I wrote, or, if you did, you
loaded it up with your own inferences.  There was no suggestion that
somehow a single PCRE was to be used.  Nor was there any suggestion that
all the parsing should be done by the matching engine -- that's where
the callback comes in.

> With that said, regular expressions can be useful in parsing markup 
> languages (BTDT).  But they should not be the first consideration when 
> attempting to do that.  There are tools better suited to the purpose.

That's so vague as to be almost content-free.  The only concrete remark
is either wrong or assumes facts not yet established.  PCREs have been,
for some time, my first consideration when processing some kinds of
markup -- particularly when, as here, I can choose the notation myself.
As you say, BTDT.  Obviously YMMV.

These posts all occurred within a specific context but you often like to
strip remarks from their natural home.  My reply was to a post that
suggested using explode or character-by-character processing.  That does
not suggest that a complex grammar was going to be involved.  Using
multiple PCREs with callbacks is a good way to process simple languages,
but as soon as it because clear the something more sophisticated might
be envisaged, I suggested using XML and leveraging existing tools.  Even
if James goes with a new syntax that is a bit XML-like, processing with
the PCRE functions might still be a good way forward since the tags
don't need to nest (note: tags not elements, and XML-like not XML).

> See also <http://stackoverflow.com/a/1732454/855543> :)

No thanks.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16557

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-21 11:33 +0100
Message-ID<1615686.hQ19P4Vg7h@PointedEars.de>
In reply to#16552
Ben Bacarisse wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>> Ben Bacarisse wrote:
>>> "Christoph M. Becker" <cmbecker69@arcor.de> writes:
>>>> Anyhow, what Thomas most likely wanted to convey is that
>>>> preg_replace_callback() is not general enough to process all markup
>>>> languages.
>>> Seems unlikely.
>> How did you get that idea?  Christoph has understood me correctly.
> 
> Hmm... I wonder why you just cut, then, without comment my explanation
> of why that's wrong.  Do you have nothing to say about that part?

The only thing that I can see that you said definitively in this thread 
about using preg_replace_callback() in the way you suggested is this:

>>>>> It's general enough to be handle nested constructs because Perl's REs
>>>>> are, in practice, much more powerful than REs should be.  It's good to
>>>>> about nested construct (which I'd do by repeated scanning from the
>>>>> inside out) but it's there for those cases where nothing else will do.

And that is wrong.  First of all, you are proceeding from a false 
assumption: preg_*() do _not_ support Perl’s RE; they support Perl-
*Compatible* RE which are not as powerful as Perl’s RE:

<http://php.net/manual/en/reference.pcre.pattern.differences.php>

Second, I have explicitly commented on your assertion that because recursion 
is possible with PCRE, they would be “general enough” for parsing markup 
languages.  I have given you a lot of good reasons, both theoretical and 
practical, why that is wrong (you derided them).  It is also wrong insofar 
as “general enough” is just weasel words:

,-<http://php.net/manual/en/regexp.reference.recursive.php>
| 
| […] Perl 5.6 has provided an experimental facility that allows regular 
                               ^^^^^^^^^^^^^^^^^^^^^
| expressions to recurse (among other things). The special item (?R) is 
| provided for the specific case of recursion. This PCRE pattern solves the 
| parentheses problem (assume the PCRE_EXTENDED option is set so that white 
| space is ignored): \( ( (?>[^()]+) | (?R) )* \)
| 
| […]
| The values set for any capturing subpatterns are those from the outermost 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| level of the recursion at which the subpattern value is set. If the 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| pattern above is matched against (ab(cd)ef) the value for the capturing 
| parentheses is "ef", which is the last value taken on at the top level. If 
| additional parentheses are added, giving \( ( ( (?>[^()]+) | (?R) )* ) \) 
| then the string they capture is "ab(cd)ef", the contents of the top level 
| parentheses. If there are more than 15 capturing parentheses in a pattern, 
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| PCRE has to obtain extra memory to store data during a recursion, which it 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| does by using pcre_malloc, freeing it via pcre_free afterwards. If no
                                                                  ^^^^^ 
| memory can be obtained, it saves data for the first 15 capturing 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| parentheses only, as there is no way to give an out-of-memory error from 
  ^^^^^^^^^^^^^^^^
| within a recursion.
| 
| […]
| The maximum length of a subject string is the largest positive number that 
| an integer variable can hold. However, PCRE uses recursion to handle 
                                         ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| subpatterns and indefinite repetition. This means that the available stack 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| space may limit the size of a subject string that can be processed by 
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| certain patterns.
  ^^^^^^^^^^^^^^^^^

Third, cases “where nothing else will do”, as you implied, do not exist.  
Markup parsers exist, and can be written with PHP.  Most notably, PHP 
includes at least four ways to parse HTML and XML with libxml:

DOMDocument
<http://php.net/manual/en/domdocument.load.php>
<http://php.net/manual/en/domdocument.loadhtml.php>
<http://php.net/manual/en/domdocument.loadhtmlfile.php>
<http://php.net/manual/en/domdocument.loadxml.php>

SimpleXML
<http://php.net/manual/en/book.simplexml.php>

XML Parser
<http://php.net/manual/en/book.xml.php>

XMLReader
<http://php.net/manual/en/book.xmlreader.php>

See also: <http://php.net/manual/en/book.libxml.php>


The rest of your posting appears to be just full quote, /ad hominem/, 
misconceptions, and blissful ignorance, so I do not find it worth my 
precious time to be commented on.  (Just in case you are wondering.)

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16562

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-21 12:53 +0000
Message-ID<8760xi1ayz.fsf@bsb.me.uk>
In reply to#16557
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

<snip>
> The rest of your posting appears to be just full quote, /ad hominem/, 
> misconceptions, and blissful ignorance, so I do not find it worth my 
> precious time to be commented on.  (Just in case you are wondering.)

I am much relieved.  I seems I have found the solution I've been missing
for so long.  Will full quoting always bring about this happy outcome?

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16529

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 18:29 +0000
Message-ID<naab55$mmd$2@dont-email.me>
In reply to#16501
On 18/02/2016 14:56, Ben Bacarisse wrote:
> James Harris <james.harris.1@gmail.com> writes:
>
>> Quick question: If someone wanted to parse marked-up text (and to
>> convert it to HTML) do you have any generic recommendation on a good
>> approach to use in PHP?

...

> I would process it with preg_replace_callback -- possibly more that
> once.  That's just about the most general way, but it means you are less
> likely to get stuck if your markup gets more complicated than originally
> thought.

Thanks, I can see that that is a good option.

James

[toc] | [prev] | [next] | [standalone]


#16534

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-20 20:32 +0100
Message-ID<5829000.BqkcIiuHTp@PointedEars.de>
In reply to#16529
James Harris wrote:

> On 18/02/2016 14:56, Ben Bacarisse wrote:
>> James Harris <james.harris.1@gmail.com> writes:
>>> Quick question: If someone wanted to parse marked-up text (and to
>>> convert it to HTML) do you have any generic recommendation on a good
>>> approach to use in PHP?
>> […]
>> I would process it with preg_replace_callback -- possibly more that
>> once.  That's just about the most general way, but it means you are less
>> likely to get stuck if your markup gets more complicated than originally
>> thought.
> 
> Thanks, I can see that that is a good option.

It is not.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16503

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-18 10:43 -0500
Message-ID<na4oki$igp$1@jstuckle.eternal-september.org>
In reply to#16497
On 2/18/2016 7:47 AM, James Harris wrote:
> Quick question: If someone wanted to parse marked-up text (and to
> convert it to HTML) do you have any generic recommendation on a good
> approach to use in PHP?
> 
> Specifically, the text could be parsed character-by-character or line by
> line. Or perhaps it could be split with explode() and processed in pieces.
> 
> The input text could be ASCII or Unicode.
> 
> The marking up would be by tags of some sort. The form of a tag is not
> yet defined; it could be chosen to make the parsing easier, if
> necessary, as long as it was easy for a human to work with.
> 
> James
> 

Way to vague a question for an intelligent answer, and any answer at
this juncture is just a guess.  It all depends on how the text is going
to be marked up.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | comp.lang.php


csiph-web