Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #16497 > unrolled thread

Any recommendations on using PHP to parse marked-up text?

Started byJames Harris <james.harris.1@gmail.com>
First post2016-02-18 12:47 +0000
Last post2016-02-27 15:58 +1300
Articles 10 on this page of 50 — 8 participants

Back to article view | Back to comp.lang.php


Contents

  Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-18 12:47 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:52 +0100
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:28 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 14:56 +0000
      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:14 +0100
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 21:45 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 23:09 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 01:15 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-19 02:45 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 12:03 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? Tim Streater <timstreater@greenbee.net> - 2016-02-19 12:17 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-19 14:18 +0100
                    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 14:02 +0000
                      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:31 +0100
                        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 01:39 +0000
                          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:33 +0100
                            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 12:53 +0000
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:29 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:32 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:43 -0500
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:51 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 20:08 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:22 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 23:40 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:39 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 13:09 +0000
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 20:24 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 22:35 +0000
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:37 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 10:55 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:25 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:50 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:49 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:05 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:45 -0500
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 09:01 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:46 +0100
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 12:36 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:39 +0100
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:40 +0000
                    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:46 +0100
                      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 14:57 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:10 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:50 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:08 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-21 18:35 +0100
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:55 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-22 16:38 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:06 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Ian Collins <ian-news@hotmail.com> - 2016-02-27 15:58 +1300

Page 3 of 3 — ← Prev page 1 2 [3]


#16568

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-21 14:46 +0100
Message-ID<4842430.bZfGOjMuyE@PointedEars.de>
In reply to#16566
James Harris wrote:

> On 21/02/2016 13:39, Thomas 'PointedEars' Lahn wrote:
>> I have no idea what you are talking about.

Which does not mean that I do not understand the subject.

>> (Do you?)
> 
> Yes.

Somehow, with your habit of jumping to conclusions, I find it hard to 
believe you.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16570

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 14:57 +0000
Message-ID<nacj3o$sjv$1@dont-email.me>
In reply to#16568
On 21/02/2016 13:46, Thomas 'PointedEars' Lahn wrote:
> James Harris wrote:
>
>> On 21/02/2016 13:39, Thomas 'PointedEars' Lahn wrote:
>>> I have no idea what you are talking about.
>
> Which does not mean that I do not understand the subject.
>
>>> (Do you?)
>>
>> Yes.
>
> Somehow, with your habit of jumping to conclusions, I find it hard to
> believe you.

You don't see the irony in that? I think you lost track of what I was 
talking about a few posts back! If you don't understand people it's no 
wonder you think they are jumping to conclusions.

James

[toc] | [prev] | [next] | [standalone]


#16574

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-21 10:10 -0500
Message-ID<nacjs0$cd$1@jstuckle.eternal-september.org>
In reply to#16555
On 2/21/2016 4:01 AM, James Harris wrote:
> On 20/02/2016 22:45, Jerry Stuckle wrote:
> 
> ... (comments about markup and tag forms snipped)
> 
>> But that's what you have to define before you can pick a way of parsing
>> it.  Different markups will require different parsing techniques.  If
>> you can define the markup, you should define it according to how you
>> want to parse it.
> 
> I am puzzled by this as you seem to say two different things: 1) define
> the markup before the parse method, 2) define the markup to suit (i.e.
> after) the intended parse method.
> 

I'm saying you need to do both.  Define a markup, then look at how to
parse it.  Modify your markup design as necessary.  They should go hand
in hand.

> I did say before that the markup could be chosen to be easy to parse as
> long as it was also easy for a human to work with. I see no reason why a
> markup should be defined first.
> 

Forcing people to learn something completely new is NOT "easy for a
human to work with".

> For example, some markup schemes use different symbols to indicate
> different tags, such as
> 
> == Heading at level 2 ==
>> Indent
> * Bulleted list element
> | Cells | in a | table |
> 
> ISTM that such a markup would require character-by-character parsing and
> would have loads of contexts to manage, making it a nightmare to parse.
> By contrast, a single rarely used character or character sequence to
> begin each tag would be much easier to locate.
> 

That's correct.  But my question is - what does it matter?  And if you
aren't going to hamstring the user (another reason not to use your
code), you'll have to handle these contexts, no matter what markup you use.

>> And it even can define which language you should be
>> using - PHP may or may not be the best.
> 
> True. I assume PHP is the right language to use on server-side code
> because it is so widespread. Are you saying that other languages are
> available server-side?
> 
> James
> 

Many.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16576

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 17:50 +0000
Message-ID<nact8c$67s$1@dont-email.me>
In reply to#16574
On 21/02/2016 15:10, Jerry Stuckle wrote:
> On 2/21/2016 4:01 AM, James Harris wrote:
>> On 20/02/2016 22:45, Jerry Stuckle wrote:
>>
>> ... (comments about markup and tag forms snipped)
>>
>>> But that's what you have to define before you can pick a way of parsing
>>> it.  Different markups will require different parsing techniques.  If
>>> you can define the markup, you should define it according to how you
>>> want to parse it.
>>
>> I am puzzled by this as you seem to say two different things: 1) define
>> the markup before the parse method, 2) define the markup to suit (i.e.
>> after) the intended parse method.
>>
>
> I'm saying you need to do both.  Define a markup, then look at how to
> parse it.  Modify your markup design as necessary.  They should go hand
> in hand.

Odd. That's what I was saying too. Never mind. Good to end a discussion 
with agreement!

James

[toc] | [prev] | [next] | [standalone]


#16578

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-21 14:08 -0500
Message-ID<nad1po$p7i$1@jstuckle.eternal-september.org>
In reply to#16576
On 2/21/2016 12:50 PM, James Harris wrote:
> On 21/02/2016 15:10, Jerry Stuckle wrote:
>> On 2/21/2016 4:01 AM, James Harris wrote:
>>> On 20/02/2016 22:45, Jerry Stuckle wrote:
>>>
>>> ... (comments about markup and tag forms snipped)
>>>
>>>> But that's what you have to define before you can pick a way of parsing
>>>> it.  Different markups will require different parsing techniques.  If
>>>> you can define the markup, you should define it according to how you
>>>> want to parse it.
>>>
>>> I am puzzled by this as you seem to say two different things: 1) define
>>> the markup before the parse method, 2) define the markup to suit (i.e.
>>> after) the intended parse method.
>>>
>>
>> I'm saying you need to do both.  Define a markup, then look at how to
>> parse it.  Modify your markup design as necessary.  They should go hand
>> in hand.
> 
> Odd. That's what I was saying too. Never mind. Good to end a discussion
> with agreement!
> 
> James
> 

The problem is - you're asking "what's the best way to parse...".  What
I'm saying is - you can't say what the "best way" is until you have
defined the format of the markup to be parsed.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16575

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-21 18:35 +0100
Message-ID<56C9F573.6060406@arnowelzel.de>
In reply to#16530
James Harris schrieb am 2016-02-20 um 19:51:

> On 18/02/2016 15:43, Jerry Stuckle wrote:
[...]
>> Way to vague a question for an intelligent answer, and any answer at
>> this juncture is just a guess.  It all depends on how the text is going
>> to be marked up.
> 
> The markup can be chosen. As long as it is easy to parse and easy for a 
> human to work with it would be suitable. Most such markup would be 
> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding 
> closing tags. There could also be self-contained tags which get 
> converted to <br>, <hr> etc.
> 
> To illustrate, here is one possible scheme for opening and closing tags:
> 
> @(tagname, parameter, parameter)text@(/tagname)

And why using this special markup and not HTML or XML? This would not be
much different.

> Here is another option:
> 
> @tagname:parameter, parameter;text@/tagname;

This even looks worse than HTML in respect of "understandable for
humans". Keep in mind that HTML was designed to be easy to understand
for humans. So maybe you try to re-invent the wheel.

If you intent to make the markup *easier* for people than HTML, then you
should start with something which is well documented and not your own
creation and it should be *easy* to use.

There are already parsers for popular markup formats:

Markdown: <http://parsedown.org/>

Textile: <https://github.com/textile/php-textile>



-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


#16577

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 17:55 +0000
Message-ID<nactgu$7nu$1@dont-email.me>
In reply to#16575
On 21/02/2016 17:35, Arno Welzel wrote:
> James Harris schrieb am 2016-02-20 um 19:51:
>
>> On 18/02/2016 15:43, Jerry Stuckle wrote:
> [...]
>>> Way to vague a question for an intelligent answer, and any answer at
>>> this juncture is just a guess.  It all depends on how the text is going
>>> to be marked up.
>>
>> The markup can be chosen. As long as it is easy to parse and easy for a
>> human to work with it would be suitable. Most such markup would be
>> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
>> closing tags. There could also be self-contained tags which get
>> converted to <br>, <hr> etc.
>>
>> To illustrate, here is one possible scheme for opening and closing tags:
>>
>> @(tagname, parameter, parameter)text@(/tagname)
>
> And why using this special markup and not HTML or XML? This would not be
> much different.

For a few reasons. I think I have mentioned most of them in other 
subthreads but don't worry about them.

>> Here is another option:
>>
>> @tagname:parameter, parameter;text@/tagname;
>
> This even looks worse than HTML in respect of "understandable for
> humans". Keep in mind that HTML was designed to be easy to understand
> for humans. So maybe you try to re-invent the wheel.

Yes. It served to help illustrate the kinds of tags needed but it is too 
ugly to be used as it is.

James

[toc] | [prev] | [next] | [standalone]


#16590

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-22 16:38 +0100
Message-ID<56CB2B60.1060703@arnowelzel.de>
In reply to#16577
James Harris schrieb am 2016-02-21 um 18:55:

> On 21/02/2016 17:35, Arno Welzel wrote:
>> James Harris schrieb am 2016-02-20 um 19:51:
>>
>>> On 18/02/2016 15:43, Jerry Stuckle wrote:
>> [...]
>>>> Way to vague a question for an intelligent answer, and any answer at
>>>> this juncture is just a guess.  It all depends on how the text is going
>>>> to be marked up.
>>>
>>> The markup can be chosen. As long as it is easy to parse and easy for a
>>> human to work with it would be suitable. Most such markup would be
>>> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
>>> closing tags. There could also be self-contained tags which get
>>> converted to <br>, <hr> etc.
>>>
>>> To illustrate, here is one possible scheme for opening and closing tags:
>>>
>>> @(tagname, parameter, parameter)text@(/tagname)
>>
>> And why using this special markup and not HTML or XML? This would not be
>> much different.
> 
> For a few reasons. I think I have mentioned most of them in other 
> subthreads but don't worry about them.

Let's see one of them:

"I see what you mean but in this case the markup is partly for security.
I need to ensure that the person who writes the markup does not have
access to general HTML facilities. But at the same time I need to
provide enough to allow for basic markup of text and probably a few
other things."

Avoiding "general HTML facilities" is easy - just create whitelist of
allowed tags and attributes and filter everything else out. This is how
most CMS handle HTML input to avoid embedded JavaScript or breaking the
site layout using local CSS etc.

And then:

"For example, I may have a /conditional/ link tag that, for local
targets, will check that the target exists. If the target does not exist
then the link will render as non-clickable text. If the target does
exist then the text will be clickable."

Well - checking the existence of a linked target hast is nothing to do
with markup at all. You may use HTML as a basis and either define that
links will only be kept if the target exists. No user will understand,
that he or she has to use a special "conditional link" element for local
targets which may not exist.

Maybe you should just try existing systems like DokuWiki to get a
feeling for how this kind of content management works.


-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


#16512

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-18 22:06 +0100
Message-ID<1628012.948ngoYJoA@PointedEars.de>
In reply to#16497
James Harris wrote:

> Quick question: If someone wanted to parse marked-up text (and to
> convert it to HTML) do you have any generic recommendation on a good
> approach to use in PHP?

Yes.
 
> Specifically, the text could be parsed character-by-character or line by
> line. Or perhaps it could be split with explode() and processed in pieces.

All those approaches are unwise.
 
> The input text could be ASCII or Unicode.

Again, please read the Unicode FAQ.  All of it.
 
> The marking up would be by tags of some sort. The form of a tag is not
> yet defined; it could be chosen to make the parsing easier, if
> necessary, as long as it was easy for a human to work with.

OK.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16604

FromIan Collins <ian-news@hotmail.com>
Date2016-02-27 15:58 +1300
Message-ID<djchmeFrc31U1@mid.individual.net>
In reply to#16497
James Harris wrote:
> Quick question: If someone wanted to parse marked-up text (and to
> convert it to HTML) do you have any generic recommendation on a good
> approach to use in PHP?
>
> Specifically, the text could be parsed character-by-character or line by
> line. Or perhaps it could be split with explode() and processed in pieces.
>
> The input text could be ASCII or Unicode.
>
> The marking up would be by tags of some sort. The form of a tag is not
> yet defined; it could be chosen to make the parsing easier, if
> necessary, as long as it was easy for a human to work with.

Write the text in Open/Libre Office and parse the XML with the native 
PHP parsers.

-- 
Ian Collins

[toc] | [prev] | [standalone]


Page 3 of 3 — ← Prev page 1 2 [3]

Back to top | Article view | comp.lang.php


csiph-web