Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.php > #16497 > unrolled thread
| Started by | James Harris <james.harris.1@gmail.com> |
|---|---|
| First post | 2016-02-18 12:47 +0000 |
| Last post | 2016-02-27 15:58 +1300 |
| Articles | 20 on this page of 50 — 8 participants |
Back to article view | Back to comp.lang.php
Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-18 12:47 +0000
Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:52 +0100
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:28 +0000
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 14:56 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:14 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 21:45 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 23:09 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 01:15 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-19 02:45 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 12:03 +0000
Re: Any recommendations on using PHP to parse marked-up text? Tim Streater <timstreater@greenbee.net> - 2016-02-19 12:17 +0000
Re: Any recommendations on using PHP to parse marked-up text? "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-19 14:18 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 14:02 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:31 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 01:39 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:33 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 12:53 +0000
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:29 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:32 +0100
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:43 -0500
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:51 +0000
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 20:08 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:22 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 23:40 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:39 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 13:09 +0000
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 20:24 +0000
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 22:35 +0000
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:37 +0000
Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 10:55 +0000
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:25 +0000
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:50 -0500
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:49 +0000
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:05 -0500
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:45 -0500
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 09:01 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:46 +0100
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 12:36 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:39 +0100
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:40 +0000
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:46 +0100
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 14:57 +0000
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:10 -0500
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:50 +0000
Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:08 -0500
Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-21 18:35 +0100
Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:55 +0000
Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-22 16:38 +0100
Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:06 +0100
Re: Any recommendations on using PHP to parse marked-up text? Ian Collins <ian-news@hotmail.com> - 2016-02-27 15:58 +1300
Page 1 of 3 [1] 2 3 Next page →
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-18 12:47 +0000 |
| Subject | Any recommendations on using PHP to parse marked-up text? |
| Message-ID | <na4eau$at9$1@dont-email.me> |
Quick question: If someone wanted to parse marked-up text (and to convert it to HTML) do you have any generic recommendation on a good approach to use in PHP? Specifically, the text could be parsed character-by-character or line by line. Or perhaps it could be split with explode() and processed in pieces. The input text could be ASCII or Unicode. The marking up would be by tags of some sort. The form of a tag is not yet defined; it could be chosen to make the parsing easier, if necessary, as long as it was easy for a human to work with. James
[toc] | [next] | [standalone]
| From | Arno Welzel <usenet@arnowelzel.de> |
|---|---|
| Date | 2016-02-18 15:52 +0100 |
| Message-ID | <56C5DAAB.1070307@arnowelzel.de> |
| In reply to | #16497 |
James Harris schrieb am 2016-02-18 um 13:47: > Quick question: If someone wanted to parse marked-up text (and to > convert it to HTML) do you have any generic recommendation on a good > approach to use in PHP? No. > Specifically, the text could be parsed character-by-character or line by > line. Or perhaps it could be split with explode() and processed in pieces. Then go on and use explode() or split or whatever. > The input text could be ASCII or Unicode. > > The marking up would be by tags of some sort. The form of a tag is not > yet defined; it could be chosen to make the parsing easier, if > necessary, as long as it was easy for a human to work with. Which sounds like "the input can be anything and I need to parse it". Decide what format you want to use and maybe someone knows a good parser for that format. If you create your own format you have to build your own parser for it - and no, there is no thing like a "universal text parser" in PHP. -- Arno Welzel http://arnowelzel.de http://de-rec-fahrrad.de http://fahrradzukunft.de
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-20 18:28 +0000 |
| Message-ID | <naab2e$mmd$1@dont-email.me> |
| In reply to | #16500 |
On 18/02/2016 14:52, Arno Welzel wrote: > James Harris schrieb am 2016-02-18 um 13:47: ... >> The input text could be ASCII or Unicode. >> >> The marking up would be by tags of some sort. The form of a tag is not >> yet defined; it could be chosen to make the parsing easier, if >> necessary, as long as it was easy for a human to work with. > > Which sounds like "the input can be anything and I need to parse it". Yep. The markup format can be chosen to make it easier to parse, as long as it is also easy for a human to work with. > Decide what format you want to use and maybe someone knows a good parser > for that format. If you create your own format you have to build your > own parser for it - and no, there is no thing like a "universal text > parser" in PHP. Although there is no predefined markup form there are some universal requirements. Some/most markup tags will start a new context (which other tags will terminate). They are ones which would be converted to HTML opening and closing tags such as <h2>...</h2>. But other markup tags would be self-contained and will not need corresponding closing tags. They would be converted to such as <hr>. James
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-18 14:56 +0000 |
| Message-ID | <87twl66p9h.fsf@bsb.me.uk> |
| In reply to | #16497 |
James Harris <james.harris.1@gmail.com> writes: > Quick question: If someone wanted to parse marked-up text (and to > convert it to HTML) do you have any generic recommendation on a good > approach to use in PHP? > > Specifically, the text could be parsed character-by-character or line > by line. Or perhaps it could be split with explode() and processed in > pieces. I would process it with preg_replace_callback -- possibly more that once. That's just about the most general way, but it means you are less likely to get stuck if your markup gets more complicated than originally thought. It's general enough to be handle nested constructs because Perl's REs are, in practice, much more powerful than REs should be. It's good to about nested construct (which I'd do by repeated scanning from the inside out) but it's there for those cases where nothing else will do. You might want to consider using one of the existing schemes (even f you have to extend it) like markdown. <snip> -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-18 22:14 +0100 |
| Message-ID | <3816086.GMaWPedH5Q@PointedEars.de> |
| In reply to | #16501 |
Ben Bacarisse wrote:
> James Harris <james.harris.1@gmail.com> writes:
>> Quick question: If someone wanted to parse marked-up text (and to
>> convert it to HTML) do you have any generic recommendation on a good
>> approach to use in PHP?
>> […]
>
> I would process it with preg_replace_callback -- possibly more that
^^^
> once. That's just about the most general way, but it means you are less
> likely to get stuck if your markup gets more complicated than originally
> thought.
Sigh. [psf 10.1]
<https://en.wikipedia.org/wiki/Chomsky_hierarchy>
--
PointedEars
Zend Certified PHP Engineer
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-18 21:45 +0000 |
| Message-ID | <87bn7d7kvl.fsf@bsb.me.uk> |
| In reply to | #16513 |
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: > Ben Bacarisse wrote: > >> James Harris <james.harris.1@gmail.com> writes: >>> Quick question: If someone wanted to parse marked-up text (and to >>> convert it to HTML) do you have any generic recommendation on a good >>> approach to use in PHP? >>> […] >> >> I would process it with preg_replace_callback -- possibly more that > ^^^ >> once. That's just about the most general way, but it means you are less >> likely to get stuck if your markup gets more complicated than originally >> thought. > > Sigh. [psf 10.1] > > <https://en.wikipedia.org/wiki/Chomsky_hierarchy> Are these blank references supposed to convey some useful point? -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-18 23:09 +0100 |
| Message-ID | <1615449.BjJPyqVnO6@PointedEars.de> |
| In reply to | #16515 |
Ben Bacarisse wrote: > Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >> Ben Bacarisse wrote: >>> James Harris <james.harris.1@gmail.com> writes: >>>> Quick question: If someone wanted to parse marked-up text (and to >>>> convert it to HTML) do you have any generic recommendation on a good >>>> approach to use in PHP? >>>> […] >>> I would process it with preg_replace_callback -- possibly more that >> ^^^ >>> once. That's just about the most general way, but it means you are less >>> likely to get stuck if your markup gets more complicated than originally >>> thought. >> >> Sigh. [psf 10.1] >> >> <https://en.wikipedia.org/wiki/Chomsky_hierarchy> > > Are these blank references I would not call them “blank”. > supposed to convey some useful point? Yes. -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-19 01:15 +0000 |
| Message-ID | <87lh6h5wlb.fsf@bsb.me.uk> |
| In reply to | #16516 |
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: > Ben Bacarisse wrote: > >> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >>> Ben Bacarisse wrote: >>>> James Harris <james.harris.1@gmail.com> writes: >>>>> Quick question: If someone wanted to parse marked-up text (and to >>>>> convert it to HTML) do you have any generic recommendation on a good >>>>> approach to use in PHP? >>>>> […] >>>> I would process it with preg_replace_callback -- possibly more that >>> ^^^ >>>> once. That's just about the most general way, but it means you are less >>>> likely to get stuck if your markup gets more complicated than originally >>>> thought. >>> >>> Sigh. [psf 10.1] >>> >>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy> >> >> Are these blank references > > I would not call them “blank”. psf 10.1 is blank to me, and the second simply summaries things that I would assume almost everyone here knows (I certainly do). There's no hint at what you intend to say by posting it. Maybe you are simply suggesting some people might like to read it? >> supposed to convey some useful point? > > Yes. Good. I hope whoever it was directed at understood it. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-19 02:45 +0100 |
| Message-ID | <2353152.leaMV68clc@PointedEars.de> |
| In reply to | #16520 |
Ben Bacarisse wrote: > Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >> Ben Bacarisse wrote: >>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >>>> Sigh. [psf 10.1] >>>> >>>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy> >>> >>> Are these blank references >> >> I would not call them “blank”. > > psf 10.1 is blank to me, STFW. > and the second simply summaries things that I would assume almost everyone > here knows (I certainly do). Maybe you know it, but you do not understand either it, or how preg_replace_callback() works, or both. > There's no hint at what you intend to say by posting it. But there is. I have also marked the most relevant part of your posting. Unfortunately, I cannot quote it again *and* post it. > Maybe you are simply suggesting some people might like to read it? Yes, and draw the correct conclusions from it – whereas “some people” includes you. >>> supposed to convey some useful point? >> Yes. > > Good. I hope whoever it was directed at understood it. It was directed at you, of course. -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-19 12:03 +0000 |
| Message-ID | <874md46h6a.fsf@bsb.me.uk> |
| In reply to | #16521 |
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: > Ben Bacarisse wrote: > >> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >>> Ben Bacarisse wrote: >>>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: >>>>> Sigh. [psf 10.1] >>>>> >>>>> <https://en.wikipedia.org/wiki/Chomsky_hierarchy> >>>> >>>> Are these blank references >>> >>> I would not call them “blank”. >> >> psf 10.1 is blank to me, > > STFW. Now, why do you think I didn't? You make hae no respect for the knowledge and experience of others, but I do. Why you say something I try to understand it. I could attach no meaning to those letters and symbols. That maybe DuckDuckGo's problem, but since you are so keen on citations and good references I'm surprised. >> and the second simply summaries things that I would assume almost everyone >> here knows (I certainly do). > > Maybe you know it, but you do not understand either it, or how > preg_replace_callback() works, or both. Maybe. Since you won't say what you mean by pointing me at that page, we'll never know. Maybe it you who doesn't understand one or other of those things. Where the confusion lies seems likely to remain a mystery to both of us. >> There's no hint at what you intend to say by posting it. > > But there is. I have also marked the most relevant part of your posting. > Unfortunately, I cannot quote it again *and* post it. Yes, it was obvious you were making some reference regular languages and regular expressions (please don't take me for a fool) but what was the point? Were you agreeing with me? Were you disagreeing? Did I make some mistake you are helpful correcting? Was my advice bad for some reason that "psf 10.1" or the WP page explains? >> Maybe you are simply suggesting some people might like to read it? > > Yes, and draw the correct conclusions from it – whereas “some people” > includes you. > >>>> supposed to convey some useful point? >>> Yes. >> >> Good. I hope whoever it was directed at understood it. > > It was directed at you, of course. Unless you are prepared to actually say what you mean we have to leave it at that. I read the page, I searched for psf 10.1, and I still have no idea what you want to say. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2016-02-19 12:17 +0000 |
| Message-ID | <190220161217359137%timstreater@greenbee.net> |
| In reply to | #16522 |
In article <874md46h6a.fsf@bsb.me.uk>, Ben Bacarisse <ben.usenet@bsb.me.uk> wrote: >Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: > >> Ben Bacarisse wrote: >>> psf 10.1 is blank to me, >> >> STFW. > >Now, why do you think I didn't? You may have no respect for the >knowledge and experience of others, but I do. When you say something I >try to understand it. I could attach no meaning to those letters and >symbols. That maybe DuckDuckGo's problem, but since you are so keen on >citations and good references I'm surprised. Just giving an acronym in that manner tells the other person that you really can't be bothered to deal with them. Not the sort of person who should be a teacher of any kind. It's just plain rude. [snip] >>> Good. I hope whoever it was directed at understood it. >> >> It was directed at you, of course. > >Unless you are prepared to actually say what you mean we have to leave >it at that. I read the page, I searched for psf 10.1, and I still have >no idea what you want to say. Are you sure he actually wants to help people? I'm not. -- New Socialism consists essentially in being seen to have your heart in the right place whilst your head is in the clouds and your hand is in someone else's pocket.
[toc] | [prev] | [next] | [standalone]
| From | "Christoph M. Becker" <cmbecker69@arcor.de> |
|---|---|
| Date | 2016-02-19 14:18 +0100 |
| Message-ID | <na74nd$4v6$1@solani.org> |
| In reply to | #16522 |
Ben Bacarisse wrote: > Unless you are prepared to actually say what you mean we have to leave > it at that. I read the page, I searched for psf 10.1, and I still have > no idea what you want to say. Googling for "psf 10.1" brought the desired information: <http://pointedears.de/psf/index.en#sigh> Anyhow, what Thomas most likely wanted to convey is that preg_replace_callback() is not general enough to process all markup languages. -- Christoph M. Becker
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-19 14:02 +0000 |
| Message-ID | <87wpq04x3b.fsf@bsb.me.uk> |
| In reply to | #16524 |
"Christoph M. Becker" <cmbecker69@arcor.de> writes: > Ben Bacarisse wrote: > >> Unless you are prepared to actually say what you mean we have to leave >> it at that. I read the page, I searched for psf 10.1, and I still have >> no idea what you want to say. > > Googling for "psf 10.1" brought the desired information: > > <http://pointedears.de/psf/index.en#sigh> Oh! That did not show up in my search engine. Why on Earth would he not just cite that link -- he complains enough about other people not citing things properly? So the "psf 10.1" was just to show me that he has listed and numbered things he often says? Blimey. That's... well... it's very Usenet. > Anyhow, what Thomas most likely wanted to convey is that > preg_replace_callback() is not general enough to process all markup > languages. Seems unlikely. Since his reference was to a theoretical article, the answer would be equally theoretical, and there is no more powerful model than even the simplest recogniser linked with a callback to a Turing complete language. Anyway, there are lots of things he might have meant. I wonder if he'll ever say? -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-20 20:31 +0100 |
| Message-ID | <1504490.9PMmeY6QeD@PointedEars.de> |
| In reply to | #16525 |
Ben Bacarisse wrote:
> "Christoph M. Becker" <cmbecker69@arcor.de> writes:
>> Anyhow, what Thomas most likely wanted to convey is that
>> preg_replace_callback() is not general enough to process all markup
>> languages.
>
> Seems unlikely.
How did you get that idea? Christoph has understood me correctly.
> Since his reference was to a theoretical article, the answer would be
> equally theoretical, […]
It is theoretical, but it has important implications in practice. A
mathematician (or, at least an individual that – the record shows – appears
to have much interest in mathematics) like yourself should be able to
appreciate that. And I had expected such a person to understand the
implications without needing to be spelled out for them. Do not blame me
for your obtuseness.
> I wonder if he'll ever say?
Since some people need to have it spelled out for them, the not-so-new
argument (not even new in this newsgroup) goes as follows:
0. The Chomsky hierarchy classifies formal languages into four types,
depending on the type of formal grammars that they can be produced
with: Recursively enumerable (0), context-sensitive (1),
context-free (2), and regular (3).
1. (It can be showed by application of the pumping lemma for context-free
languages that) markup languages are context-free languages.
For example, the simplest markup language (that IMO deserves that
designation) consists of words in Latin letters that are surrounded by
pairs of matching parentheses (also called a “bracket language”;
it helps to think of those parentheses as start-tag and end-tag of
an element, and the words as its content, respectively). A context
free grammar that produces it contains the following productions:
S → (S)
S → SS
S → a
…
S → z
[Example application: S → (S) → (SS) → (SSS) → (foo)]
2. From the Chomsky hierarchy follows by applying *set theory*: While
“every regular language is context-free” (ibid.), not every
context-free language is regular.
3. *Regular* expressions are (WLOG) a way to describe *regular*
languages *only*.
4. It can be showed by application of the pumping lemma for regular
languages that there are markup (and programming) languages that
are not regular. For example, XML, despite its well-formedness,
is not a regular language.
5. Therefore, *regular* expressions are in general not suited to parse
arbitrary *markup* (or programming) languages; instead, implementations
of non-deterministic push-down automata (stacks) are.
6. preg_replace_callback(), as all preg_*() functions in PHP, supports
Perl-Compatible *Regular* Expressions (PCRE). <http://php.net/pcre>
7. While PCRE support recursion to some extent (which implements a stack),
and therefore can to some extent solve the problem that regular
expressions when applied to arbitrary bracket languages always either
match too much (greedy) or too little (non-greedy), a *single* PCRE
fails to parse an arbitrary markup language at a fundamental
level: For a non-trivial markup language you would have to use
alternation, and the *first* *next* match wins, not the *next*
*longest* one, different to what would be required by a markup parser.
With that said, regular expressions can be useful in parsing markup
languages (BTDT). But they should not be the first consideration when
attempting to do that. There are tools better suited to the purpose.
See also <http://stackoverflow.com/a/1732454/855543> :)
--
PointedEars
Zend Certified PHP Engineer
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-21 01:39 +0000 |
| Message-ID | <87mvqu2655.fsf@bsb.me.uk> |
| In reply to | #16533 |
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: > Ben Bacarisse wrote: > >> "Christoph M. Becker" <cmbecker69@arcor.de> writes: >>> Anyhow, what Thomas most likely wanted to convey is that >>> preg_replace_callback() is not general enough to process all markup >>> languages. >> >> Seems unlikely. > > How did you get that idea? Christoph has understood me correctly. Hmm... I wonder why you just cut, then, without comment my explanation of why that's wrong. Do you have nothing to say about that part? >> Since his reference was to a theoretical article, the answer would be >> equally theoretical, […] > > It is theoretical, but it has important implications in practice. A > mathematician (or, at least an individual that – the record shows – appears > to have much interest in mathematics) like yourself should be able to > appreciate that. And I had expected such a person to understand the > implications without needing to be spelled out for them. Do not blame me > for your obtuseness. > >> I wonder if he'll ever say? > > Since some people need to have it spelled out for them, the not-so-new > argument (not even new in this newsgroup) goes as follows: > > 0. The Chomsky hierarchy classifies formal languages into four types, > depending on the type of formal grammars that they can be produced > with: Recursively enumerable (0), context-sensitive (1), > context-free (2), and regular (3). > > 1. (It can be showed by application of the pumping lemma for context-free > languages that) markup languages are context-free languages. > > For example, the simplest markup language (that IMO deserves that > designation) consists of words in Latin letters that are surrounded by > pairs of matching parentheses (also called a “bracket language”; > it helps to think of those parentheses as start-tag and end-tag of > an element, and the words as its content, respectively). A context > free grammar that produces it contains the following productions: > > S → (S) > S → SS > S → a > … > S → z > > [Example application: S → (S) → (SS) → (SSS) → (foo)] And yet '/( [a-z]+ | \( (?R) \) )+/x' matches the strings of this language. Maybe you just thought that saying "pumping lemma" would sound intimidating? I may have got the details wrong -- it's late -- but it is well known that a PCRE can match the strings (and only the strings) of your example language. > 2. From the Chomsky hierarchy follows by applying *set theory*: While > “every regular language is context-free” (ibid.), not every > context-free language is regular. > > 3. *Regular* expressions are (WLOG) a way to describe *regular* > languages *only*. > > 4. It can be showed by application of the pumping lemma for regular > languages that there are markup (and programming) languages that > are not regular. For example, XML, despite its well-formedness, > is not a regular language. > > 5. Therefore, *regular* expressions are in general not suited to parse > arbitrary *markup* (or programming) languages; instead, implementations > of non-deterministic push-down automata (stacks) are. > > 6. preg_replace_callback(), as all preg_*() functions in PHP, supports > Perl-Compatible *Regular* Expressions (PCRE). > <http://php.net/pcre> All this is just bluster. Regular languages are not relevant here. What regular language does '/([ab]*)\1/' match? Can you even say where in Chomsky's hierarchy the languages matched by PCREs lie? > 7. While PCRE support recursion to some extent (which implements a stack), > and therefore can to some extent solve the problem that regular > expressions when applied to arbitrary bracket languages always either > match too much (greedy) or too little (non-greedy), a *single* PCRE > fails to parse an arbitrary markup language at a fundamental > level: For a non-trivial markup language you would have to use > alternation, and the *first* *next* match wins, not the *next* > *longest* one, different to what would be required by a markup > parser. And this just shows you didn't read what I wrote, or, if you did, you loaded it up with your own inferences. There was no suggestion that somehow a single PCRE was to be used. Nor was there any suggestion that all the parsing should be done by the matching engine -- that's where the callback comes in. > With that said, regular expressions can be useful in parsing markup > languages (BTDT). But they should not be the first consideration when > attempting to do that. There are tools better suited to the purpose. That's so vague as to be almost content-free. The only concrete remark is either wrong or assumes facts not yet established. PCREs have been, for some time, my first consideration when processing some kinds of markup -- particularly when, as here, I can choose the notation myself. As you say, BTDT. Obviously YMMV. These posts all occurred within a specific context but you often like to strip remarks from their natural home. My reply was to a post that suggested using explode or character-by-character processing. That does not suggest that a complex grammar was going to be involved. Using multiple PCREs with callbacks is a good way to process simple languages, but as soon as it because clear the something more sophisticated might be envisaged, I suggested using XML and leveraging existing tools. Even if James goes with a new syntax that is a bit XML-like, processing with the PCRE functions might still be a good way forward since the tags don't need to nest (note: tags not elements, and XML-like not XML). > See also <http://stackoverflow.com/a/1732454/855543> :) No thanks. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-21 11:33 +0100 |
| Message-ID | <1615686.hQ19P4Vg7h@PointedEars.de> |
| In reply to | #16552 |
Ben Bacarisse wrote:
> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:
>> Ben Bacarisse wrote:
>>> "Christoph M. Becker" <cmbecker69@arcor.de> writes:
>>>> Anyhow, what Thomas most likely wanted to convey is that
>>>> preg_replace_callback() is not general enough to process all markup
>>>> languages.
>>> Seems unlikely.
>> How did you get that idea? Christoph has understood me correctly.
>
> Hmm... I wonder why you just cut, then, without comment my explanation
> of why that's wrong. Do you have nothing to say about that part?
The only thing that I can see that you said definitively in this thread
about using preg_replace_callback() in the way you suggested is this:
>>>>> It's general enough to be handle nested constructs because Perl's REs
>>>>> are, in practice, much more powerful than REs should be. It's good to
>>>>> about nested construct (which I'd do by repeated scanning from the
>>>>> inside out) but it's there for those cases where nothing else will do.
And that is wrong. First of all, you are proceeding from a false
assumption: preg_*() do _not_ support Perl’s RE; they support Perl-
*Compatible* RE which are not as powerful as Perl’s RE:
<http://php.net/manual/en/reference.pcre.pattern.differences.php>
Second, I have explicitly commented on your assertion that because recursion
is possible with PCRE, they would be “general enough” for parsing markup
languages. I have given you a lot of good reasons, both theoretical and
practical, why that is wrong (you derided them). It is also wrong insofar
as “general enough” is just weasel words:
,-<http://php.net/manual/en/regexp.reference.recursive.php>
|
| […] Perl 5.6 has provided an experimental facility that allows regular
^^^^^^^^^^^^^^^^^^^^^
| expressions to recurse (among other things). The special item (?R) is
| provided for the specific case of recursion. This PCRE pattern solves the
| parentheses problem (assume the PCRE_EXTENDED option is set so that white
| space is ignored): \( ( (?>[^()]+) | (?R) )* \)
|
| […]
| The values set for any capturing subpatterns are those from the outermost
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| level of the recursion at which the subpattern value is set. If the
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| pattern above is matched against (ab(cd)ef) the value for the capturing
| parentheses is "ef", which is the last value taken on at the top level. If
| additional parentheses are added, giving \( ( ( (?>[^()]+) | (?R) )* ) \)
| then the string they capture is "ab(cd)ef", the contents of the top level
| parentheses. If there are more than 15 capturing parentheses in a pattern,
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| PCRE has to obtain extra memory to store data during a recursion, which it
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| does by using pcre_malloc, freeing it via pcre_free afterwards. If no
^^^^^
| memory can be obtained, it saves data for the first 15 capturing
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| parentheses only, as there is no way to give an out-of-memory error from
^^^^^^^^^^^^^^^^
| within a recursion.
|
| […]
| The maximum length of a subject string is the largest positive number that
| an integer variable can hold. However, PCRE uses recursion to handle
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| subpatterns and indefinite repetition. This means that the available stack
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| space may limit the size of a subject string that can be processed by
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
| certain patterns.
^^^^^^^^^^^^^^^^^
Third, cases “where nothing else will do”, as you implied, do not exist.
Markup parsers exist, and can be written with PHP. Most notably, PHP
includes at least four ways to parse HTML and XML with libxml:
DOMDocument
<http://php.net/manual/en/domdocument.load.php>
<http://php.net/manual/en/domdocument.loadhtml.php>
<http://php.net/manual/en/domdocument.loadhtmlfile.php>
<http://php.net/manual/en/domdocument.loadxml.php>
SimpleXML
<http://php.net/manual/en/book.simplexml.php>
XML Parser
<http://php.net/manual/en/book.xml.php>
XMLReader
<http://php.net/manual/en/book.xmlreader.php>
See also: <http://php.net/manual/en/book.libxml.php>
The rest of your posting appears to be just full quote, /ad hominem/,
misconceptions, and blissful ignorance, so I do not find it worth my
precious time to be commented on. (Just in case you are wondering.)
--
PointedEars
Zend Certified PHP Engineer
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2016-02-21 12:53 +0000 |
| Message-ID | <8760xi1ayz.fsf@bsb.me.uk> |
| In reply to | #16557 |
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: <snip> > The rest of your posting appears to be just full quote, /ad hominem/, > misconceptions, and blissful ignorance, so I do not find it worth my > precious time to be commented on. (Just in case you are wondering.) I am much relieved. I seems I have found the solution I've been missing for so long. Will full quoting always bring about this happy outcome? -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-20 18:29 +0000 |
| Message-ID | <naab55$mmd$2@dont-email.me> |
| In reply to | #16501 |
On 18/02/2016 14:56, Ben Bacarisse wrote: > James Harris <james.harris.1@gmail.com> writes: > >> Quick question: If someone wanted to parse marked-up text (and to >> convert it to HTML) do you have any generic recommendation on a good >> approach to use in PHP? ... > I would process it with preg_replace_callback -- possibly more that > once. That's just about the most general way, but it means you are less > likely to get stuck if your markup gets more complicated than originally > thought. Thanks, I can see that that is a good option. James
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-20 20:32 +0100 |
| Message-ID | <5829000.BqkcIiuHTp@PointedEars.de> |
| In reply to | #16529 |
James Harris wrote: > On 18/02/2016 14:56, Ben Bacarisse wrote: >> James Harris <james.harris.1@gmail.com> writes: >>> Quick question: If someone wanted to parse marked-up text (and to >>> convert it to HTML) do you have any generic recommendation on a good >>> approach to use in PHP? >> […] >> I would process it with preg_replace_callback -- possibly more that >> once. That's just about the most general way, but it means you are less >> likely to get stuck if your markup gets more complicated than originally >> thought. > > Thanks, I can see that that is a good option. It is not. -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Jerry Stuckle <jstucklex@attglobal.net> |
|---|---|
| Date | 2016-02-18 10:43 -0500 |
| Message-ID | <na4oki$igp$1@jstuckle.eternal-september.org> |
| In reply to | #16497 |
On 2/18/2016 7:47 AM, James Harris wrote: > Quick question: If someone wanted to parse marked-up text (and to > convert it to HTML) do you have any generic recommendation on a good > approach to use in PHP? > > Specifically, the text could be parsed character-by-character or line by > line. Or perhaps it could be split with explode() and processed in pieces. > > The input text could be ASCII or Unicode. > > The marking up would be by tags of some sort. The form of a tag is not > yet defined; it could be chosen to make the parsing easier, if > necessary, as long as it was easy for a human to work with. > > James > Way to vague a question for an intelligent answer, and any answer at this juncture is just a guess. It all depends on how the text is going to be marked up. -- ================== Remove the "x" from my email address Jerry Stuckle jstucklex@attglobal.net ==================
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | comp.lang.php
csiph-web