Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #16497 > unrolled thread

Any recommendations on using PHP to parse marked-up text?

Started byJames Harris <james.harris.1@gmail.com>
First post2016-02-18 12:47 +0000
Last post2016-02-27 15:58 +1300
Articles 20 on this page of 50 — 8 participants

Back to article view | Back to comp.lang.php


Contents

  Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-18 12:47 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:52 +0100
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:28 +0000
    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 14:56 +0000
      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:14 +0100
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-18 21:45 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 23:09 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 01:15 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-19 02:45 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 12:03 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? Tim Streater <timstreater@greenbee.net> - 2016-02-19 12:17 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-19 14:18 +0100
                    Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-19 14:02 +0000
                      Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:31 +0100
                        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 01:39 +0000
                          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:33 +0100
                            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 12:53 +0000
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:29 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 20:32 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:43 -0500
      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 18:51 +0000
        Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 20:08 +0000
          Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:22 +0100
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 23:40 +0000
              Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:39 +0100
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 13:09 +0000
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-20 20:24 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-20 22:35 +0000
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:37 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Ben Bacarisse <ben.usenet@bsb.me.uk> - 2016-02-21 10:55 +0000
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:25 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:50 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 08:49 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:05 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-20 17:45 -0500
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 09:01 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 11:46 +0100
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 12:36 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:39 +0100
                  Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 13:40 +0000
                    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:46 +0100
                      Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 14:57 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 10:10 -0500
              Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:50 +0000
                Re: Any recommendations on using PHP to parse marked-up text? Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:08 -0500
        Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-21 18:35 +0100
          Re: Any recommendations on using PHP to parse marked-up text? James Harris <james.harris.1@gmail.com> - 2016-02-21 17:55 +0000
            Re: Any recommendations on using PHP to parse marked-up text? Arno Welzel <usenet@arnowelzel.de> - 2016-02-22 16:38 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:06 +0100
    Re: Any recommendations on using PHP to parse marked-up text? Ian Collins <ian-news@hotmail.com> - 2016-02-27 15:58 +1300

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#16530

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 18:51 +0000
Message-ID<naacdf$s6p$1@dont-email.me>
In reply to#16503
On 18/02/2016 15:43, Jerry Stuckle wrote:
> On 2/18/2016 7:47 AM, James Harris wrote:
>> Quick question: If someone wanted to parse marked-up text (and to
>> convert it to HTML) do you have any generic recommendation on a good
>> approach to use in PHP?
>>
>> Specifically, the text could be parsed character-by-character or line by
>> line. Or perhaps it could be split with explode() and processed in pieces.
>>
>> The input text could be ASCII or Unicode.
>>
>> The marking up would be by tags of some sort. The form of a tag is not
>> yet defined; it could be chosen to make the parsing easier, if
>> necessary, as long as it was easy for a human to work with.

...

> Way to vague a question for an intelligent answer, and any answer at
> this juncture is just a guess.  It all depends on how the text is going
> to be marked up.

The markup can be chosen. As long as it is easy to parse and easy for a 
human to work with it would be suitable. Most such markup would be 
converted to HTML elements: <hN>, <p>, <a>, etc and corresponding 
closing tags. There could also be self-contained tags which get 
converted to <br>, <hr> etc.

To illustrate, here is one possible scheme for opening and closing tags:

@(tagname, parameter, parameter)text@(/tagname)

Here is another option:

@tagname:parameter, parameter;text@/tagname;

In either case, "text" is not part of a tag.

There would additionally need to be a way to express a self-contained 
tag (i.e. one that does not need a following closing tag). Perhaps

@tagname:parameter, parameter/

Given that the above all use an at sign (@) to indicate the start of a 
tag there would need to be a convenient way to include an at sign in the 
output. Perhaps

@at/

As I say, the above are just examples of the kind of elements a markup 
would have to express, and the tag forms could be chosen to help parsing.

James

[toc] | [prev] | [next] | [standalone]


#16537

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-20 20:08 +0000
Message-ID<8760xj2lhx.fsf@bsb.me.uk>
In reply to#16530
James Harris <james.harris.1@gmail.com> writes:
<snip>
> The markup can be chosen. As long as it is easy to parse and easy for
> a human to work with it would be suitable. Most such markup would be
> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
> closing tags. There could also be self-contained tags which get
> converted to <br>, <hr> etc.
>
> To illustrate, here is one possible scheme for opening and closing tags:
>
> @(tagname, parameter, parameter)text@(/tagname)
>
> Here is another option:
>
> @tagname:parameter, parameter;text@/tagname;

That seems almost as hard to write as HTML or XML.  If you are happy (or
you are forced) to write comparatively fussy markup, then a well-trodden
path would be to use XML and an XML processor to convert your custom
markup to HTML.  The only thing that's simpler about your syntax is that
you have positional parameters rather then being named attributes and
many would say that's not really simpler to write.

I though you were doing something like markdown where the markup is made
very simple indeed, but your use-case may oblige you to use a very much
more expressive markup.  Much as I dislike it, XML has much to offer
when doing this sort of thing.

<snip>
-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16539

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-20 21:22 +0100
Message-ID<1893469.h9tBEbt3fS@PointedEars.de>
In reply to#16537
Ben Bacarisse wrote:

> I though you were doing something like markdown where the markup is made
> very simple indeed, but your use-case may oblige you to use a very much
> more expressive markup.  Much as I dislike it, XML has much to offer
> when doing this sort of thing.

And one does _not_ parse it with (Perl-compatible) regular expressions as 
you suggested before, but with an XML *parser* like libxml, which is 
supported by PHP.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16551

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-20 23:40 +0000
Message-ID<87twl30x48.fsf@bsb.me.uk>
In reply to#16539
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> I though you were doing something like markdown where the markup is made
>> very simple indeed, but your use-case may oblige you to use a very much
>> more expressive markup.  Much as I dislike it, XML has much to offer
>> when doing this sort of thing.
>
> And one does _not_ parse it with (Perl-compatible) regular expressions as 
> you suggested before, but with an XML *parser* like libxml, which is 
> supported by PHP.

When I suggest XML, I advised using an established XML processor.  In
fact it's the existing tools that I've stressed as being the main reason
to consider XML.  I'm sure James is not confused on this point, but you
may not have been following the whole thread.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16558

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-21 11:39 +0100
Message-ID<1469853.zbu1769jS1@PointedEars.de>
In reply to#16551
Ben Bacarisse wrote:

> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: 
>> Ben Bacarisse wrote:
>>> I though you were doing something like markdown where the markup is made
>>> very simple indeed, but your use-case may oblige you to use a very much
>>> more expressive markup.  Much as I dislike it, XML has much to offer
>>> when doing this sort of thing.
>> And one does _not_ parse it with (Perl-compatible) regular expressions as
>> you suggested before, but with an XML *parser* like libxml, which is
>> supported by PHP.
> 
> When I suggest XML, I advised using an established XML processor.  In
> fact it's the existing tools that I've stressed as being the main reason
> to consider XML.  I'm sure James is not confused on this point, but you
> may not have been following the whole thread.

No, *you* have not been following.  One should *always* prefer a *parser* 
over plain application of regular expressions with non-regular languages
to be parsed.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16563

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-21 13:09 +0000
Message-ID<87ziuuyztn.fsf@bsb.me.uk>
In reply to#16558
Thomas 'PointedEars' Lahn <PointedEars@web.de> writes:

> Ben Bacarisse wrote:
>
>> Thomas 'PointedEars' Lahn <PointedEars@web.de> writes: 
>>> Ben Bacarisse wrote:
>>>> I though you were doing something like markdown where the markup is made
>>>> very simple indeed, but your use-case may oblige you to use a very much
>>>> more expressive markup.  Much as I dislike it, XML has much to offer
>>>> when doing this sort of thing.
>>> And one does _not_ parse it with (Perl-compatible) regular expressions as
>>> you suggested before, but with an XML *parser* like libxml, which is
>>> supported by PHP.
>> 
>> When I suggest XML, I advised using an established XML processor.  In
>> fact it's the existing tools that I've stressed as being the main reason
>> to consider XML.  I'm sure James is not confused on this point, but you
>> may not have been following the whole thread.
>
> No, *you* have not been following.  One should *always* prefer a *parser* 
> over plain application of regular expressions with non-regular languages
> to be parsed.

Fortunately no one but you have even suggested the "plain application of
regular expressions".  But if you are referring to what I wrote about
using PHP's PCREs then you are wrong.  There are simple non-regular
markup schemes that can be handled very simply with preg_replace_callback.
That remains true even if you force the use of the the strictly regular
subset of patterns because of the power gained from the callback.  The
non-regularity of the markup is not the key deciding point.  I'd give
you an example but I am hoping that by quoting in full you won't waste
your precious time commenting on this.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16540

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 20:24 +0000
Message-ID<naahtk$i3u$1@dont-email.me>
In reply to#16537
On 20/02/2016 20:08, Ben Bacarisse wrote:
> James Harris <james.harris.1@gmail.com> writes:
> <snip>
>> The markup can be chosen. As long as it is easy to parse and easy for
>> a human to work with it would be suitable. Most such markup would be
>> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
>> closing tags. There could also be self-contained tags which get
>> converted to <br>, <hr> etc.
>>
>> To illustrate, here is one possible scheme for opening and closing tags:
>>
>> @(tagname, parameter, parameter)text@(/tagname)
>>
>> Here is another option:
>>
>> @tagname:parameter, parameter;text@/tagname;
>
> That seems almost as hard to write as HTML or XML.

Maybe it looks worse than it should! I was just trying to illustrate the 
elements I thought would be needed in markup tags and to convey the idea 
that any suitable markup would do.

Here's another attempt to make it look a bit better. This may not be any 
more successful....

(@tagname, parameter, parameter)text(@/tagname)

Maybe that's slightly easier on the eye...?

> If you are happy (or
> you are forced) to write comparatively fussy markup, then a well-trodden
> path would be to use XML and an XML processor to convert your custom
> markup to HTML.

I see what you mean but in this case the markup is partly for security. 
I need to ensure that the person who writes the markup does not have 
access to general HTML facilities. But at the same time I need to 
provide enough to allow for basic markup of text and probably a few 
other things.

For example, I may have a /conditional/ link tag that, for local 
targets, will check that the target exists. If the target does not exist 
then the link will render as non-clickable text. If the target does 
exist then the text will be clickable.

 > The only thing that's simpler about your syntax is that
 > you have positional parameters rather then being named attributes and
 > many would say that's not really simpler to write.

Again, this was just an example. The parameters could be keyword but 
most of them would be optional. I was just trying to show the elements 
that I think I might need.

> I though you were doing something like markdown where the markup is made
> very simple indeed, but your use-case may oblige you to use a very much
> more expressive markup.  Much as I dislike it, XML has much to offer
> when doing this sort of thing.

One thing I may take from other markup schemes is to allow paragraphs to 
be separated by blank lines.

James

[toc] | [prev] | [next] | [standalone]


#16547

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-20 22:35 +0000
Message-ID<87ziuv103n.fsf@bsb.me.uk>
In reply to#16540
James Harris <james.harris.1@gmail.com> writes:

> On 20/02/2016 20:08, Ben Bacarisse wrote:
>> James Harris <james.harris.1@gmail.com> writes:
>> <snip>
>>> The markup can be chosen. As long as it is easy to parse and easy for
>>> a human to work with it would be suitable. Most such markup would be
>>> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
>>> closing tags. There could also be self-contained tags which get
>>> converted to <br>, <hr> etc.
>>>
>>> To illustrate, here is one possible scheme for opening and closing tags:
>>>
>>> @(tagname, parameter, parameter)text@(/tagname)
>>>
>>> Here is another option:
>>>
>>> @tagname:parameter, parameter;text@/tagname;
>>
>> That seems almost as hard to write as HTML or XML.
>
> Maybe it looks worse than it should! I was just trying to illustrate
> the elements I thought would be needed in markup tags and to convey
> the idea that any suitable markup would do.
>
> Here's another attempt to make it look a bit better. This may not be
> any more successful....
>
> (@tagname, parameter, parameter)text(@/tagname)
>
> Maybe that's slightly easier on the eye...?

It's not about the eye (for me) but that it's not very far from, say:

<tagname p="parameter,parameter">text</tagname>

and there are lots of existing tool to help write XML, validate and
process XML.  (I'm not sure what the parameters are, but it might even
be better to be able to name them: <tagname fill="left" pad="0">.

As I've said, I'm not much of a fan, but unless the syntax is very much
simpler there is some value in leveraging the XML ecosystem.  (Did I
really say that?)

>> If you are happy (or
>> you are forced) to write comparatively fussy markup, then a well-trodden
>> path would be to use XML and an XML processor to convert your custom
>> markup to HTML.
>
> I see what you mean but in this case the markup is partly for
> security. I need to ensure that the person who writes the markup does
> not have access to general HTML facilities. But at the same time I
> need to provide enough to allow for basic markup of text and probably
> a few other things.

I understand that.  I was never suggesting you HTML.  You would remove
as invalid and tags that don't match the ones you know how to convert.

> For example, I may have a /conditional/ link tag that, for local
> targets, will check that the target exists. If the target does not
> exist then the link will render as non-clickable text. If the target
> does exist then the text will be clickable.

Sure.  I'm talking syntax.  How you process the XML would be up to you.

Anyway, please don't think I'm pushing XML, I'm just saying that you
might be 9/10ths of the way there already and it is not without its
advantages.

<snip>
-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16553

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 08:37 +0000
Message-ID<nabsrm$f7r$1@dont-email.me>
In reply to#16547
On 20/02/2016 22:35, Ben Bacarisse wrote:
> James Harris <james.harris.1@gmail.com> writes:

...

>> Here's another attempt to make it look a bit better. This may not be
>> any more successful....
>>
>> (@tagname, parameter, parameter)text(@/tagname)
>>
>> Maybe that's slightly easier on the eye...?
>
> It's not about the eye (for me) but that it's not very far from, say:
>
> <tagname p="parameter,parameter">text</tagname>
>
> and there are lots of existing tool to help write XML, validate and
> process XML.  (I'm not sure what the parameters are, but it might even
> be better to be able to name them: <tagname fill="left" pad="0">.
>
> As I've said, I'm not much of a fan, but unless the syntax is very much
> simpler there is some value in leveraging the XML ecosystem.  (Did I
> really say that?)

Sorry, Ben, I didn't mean to ignore part of your previous post but I am 
afraid I did at least subconsciously dismiss the XML-type idea. To try 
to explain why, ISTM that: XML is too heavy for this, is hard for a 
human to edit, and is hard to read. I don't want to have to use tools 
(other than a text editor) to work with the marked-up text. Nor do I 
want to use pre-existing tools to parse it. In fact, I think it would be 
harder for me to restrict an existing XML parser than it would be to 
write a parser from scratch, at least for the relatively simple app I 
have in mind.

If I were to try to parse the example you give myself, I would have to 
distinguish between <tagname and < used in other contexts. For this, at 
least, angle brackets are too widely used. That's why I used the @ sign 
or the ~ or the pair (@ in the examples I gave. (And there was a way to 
generate the meta character(s) if needed.)

As I see it, the parsing of this should be simple. However the tags are 
recognised, structurally I would need only a stack of contexts. As an 
opening tag is encountered it would be pushed as a new context. As a 
closing tag is encountered it would be checked against the current 
context and then the stack will be popped. When the parse encounters a 
tag which is self contained, it will not change the context so the stack 
stays unchanged.

Other than that, pretty much all I have to do is translate input tags to 
HTML ones and do a bit of other processing which is specific to the tag.

That parse approach seems to be simple. It will not have to do any error 
recovery. If a closing tag does not match the context (i.e. does not 
match the current opening tag) then it is an error and I can stop 
parsing. If the parse gets to the end of the file and there is something 
on the stack then it is also an error. That's about it.

At least that's what I have in mind. I may have missed something important.

James

[toc] | [prev] | [next] | [standalone]


#16560

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2016-02-21 10:55 +0000
Message-ID<87h9h21gef.fsf@bsb.me.uk>
In reply to#16553
James Harris <james.harris.1@gmail.com> writes:

> On 20/02/2016 22:35, Ben Bacarisse wrote:
>> James Harris <james.harris.1@gmail.com> writes:
>
> ...
>
>>> Here's another attempt to make it look a bit better. This may not be
>>> any more successful....
>>>
>>> (@tagname, parameter, parameter)text(@/tagname)
>>>
>>> Maybe that's slightly easier on the eye...?
>>
>> It's not about the eye (for me) but that it's not very far from, say:
>>
>> <tagname p="parameter,parameter">text</tagname>
>>
>> and there are lots of existing tool to help write XML, validate and
>> process XML.  (I'm not sure what the parameters are, but it might even
>> be better to be able to name them: <tagname fill="left" pad="0">.
>>
>> As I've said, I'm not much of a fan, but unless the syntax is very much
>> simpler there is some value in leveraging the XML ecosystem.  (Did I
>> really say that?)
>
> Sorry, Ben, I didn't mean to ignore part of your previous post but I
> am afraid I did at least subconsciously dismiss the XML-type idea. To
> try to explain why, ISTM that: XML is too heavy for this, is hard for
> a human to edit, and is hard to read.

That's fine.  I hope you can see how giving an example syntax that looks
just as heavy but is harder to edit (my editor has a helpful XML mode)
and not obviously simpler to read lead me astray!

> I don't want to have to use
> tools (other than a text editor) to work with the marked-up text. Nor
> do I want to use pre-existing tools to parse it.

Eh?  Other than PHP, you mean?  And only some subset of PHP's functions
(i.e. PCRE ones are OK but XML ones are not)?

> In fact, I think it
> would be harder for me to restrict an existing XML parser than it
> would be to write a parser from scratch, at least for the relatively
> simple app I have in mind.

I'm not sure if that's true but there is a lot unknown at this stage.
Anyway, you will hear no more about XML from me.

<snip>
-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#16564

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 13:25 +0000
Message-ID<nacdnn$841$1@dont-email.me>
In reply to#16560
On 21/02/2016 10:55, Ben Bacarisse wrote:
> James Harris <james.harris.1@gmail.com> writes:

...

>> I don't want to have to use
>> tools (other than a text editor) to work with the marked-up text. Nor
>> do I want to use pre-existing tools to parse it.
>
> Eh?  Other than PHP, you mean?  And only some subset of PHP's functions
> (i.e. PCRE ones are OK but XML ones are not)?

I meant that a text editor would have to be enough to create and modify 
markup. I hadn't realised that PHP has XML-handling functions inbuilt 
and was thinking that to parse it you meant having to modify someone 
else's parsing code. I see the inbuilt functions now. Nevertheless, the 
editing and appearance of XML rule it out for me.

James

[toc] | [prev] | [next] | [standalone]


#16550

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-20 17:50 -0500
Message-ID<naaqed$i39$1@jstuckle.eternal-september.org>
In reply to#16540
On 2/20/2016 3:24 PM, James Harris wrote:
> On 20/02/2016 20:08, Ben Bacarisse wrote:
>> James Harris <james.harris.1@gmail.com> writes:
>> <snip>
>>> The markup can be chosen. As long as it is easy to parse and easy for
>>> a human to work with it would be suitable. Most such markup would be
>>> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
>>> closing tags. There could also be self-contained tags which get
>>> converted to <br>, <hr> etc.
>>>
>>> To illustrate, here is one possible scheme for opening and closing tags:
>>>
>>> @(tagname, parameter, parameter)text@(/tagname)
>>>
>>> Here is another option:
>>>
>>> @tagname:parameter, parameter;text@/tagname;
>>
>> That seems almost as hard to write as HTML or XML.
> 
> Maybe it looks worse than it should! I was just trying to illustrate the
> elements I thought would be needed in markup tags and to convey the idea
> that any suitable markup would do.
> 
> Here's another attempt to make it look a bit better. This may not be any
> more successful....
> 
> (@tagname, parameter, parameter)text(@/tagname)
> 
> Maybe that's slightly easier on the eye...?
> 
>> If you are happy (or
>> you are forced) to write comparatively fussy markup, then a well-trodden
>> path would be to use XML and an XML processor to convert your custom
>> markup to HTML.
> 
> I see what you mean but in this case the markup is partly for security.
> I need to ensure that the person who writes the markup does not have
> access to general HTML facilities. But at the same time I need to
> provide enough to allow for basic markup of text and probably a few
> other things.
>

It's impossible to ensure the person doesn't have access to general HTML
facilities.  But then why do you care?

> For example, I may have a /conditional/ link tag that, for local
> targets, will check that the target exists. If the target does not exist
> then the link will render as non-clickable text. If the target does
> exist then the text will be clickable.
> 

Or, just do like everyone else does.  If the user provides a link, it's
clickable.  If the link is bad, that's not your problem - it's the
person who provided the link.

And that's not at all.  Just because the link can be accessed doesn't
mean anything.  For instance, just a couple of days ago someone posted a
link in a social network.  But the story was behind a paywall.

>> The only thing that's simpler about your syntax is that
>> you have positional parameters rather then being named attributes and
>> many would say that's not really simpler to write.
> 
> Again, this was just an example. The parameters could be keyword but
> most of them would be optional. I was just trying to show the elements
> that I think I might need.
>

I think you're worrying about too many things here, and as a result you
are overcomplicating things.  Don't expect users to use a special markup
just for your site.  It won't happen.

>> I though you were doing something like markdown where the markup is made
>> very simple indeed, but your use-case may oblige you to use a very much
>> more expressive markup.  Much as I dislike it, XML has much to offer
>> when doing this sort of thing.
> 
> One thing I may take from other markup schemes is to allow paragraphs to
> be separated by blank lines.
> 
> James
> 

That's a style consideration.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16554

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 08:49 +0000
Message-ID<nabtht$h7u$1@dont-email.me>
In reply to#16550
On 20/02/2016 22:50, Jerry Stuckle wrote:
> On 2/20/2016 3:24 PM, James Harris wrote:
>> On 20/02/2016 20:08, Ben Bacarisse wrote:

...

>>> If you are happy (or
>>> you are forced) to write comparatively fussy markup, then a well-trodden
>>> path would be to use XML and an XML processor to convert your custom
>>> markup to HTML.
>>
>> I see what you mean but in this case the markup is partly for security.
>> I need to ensure that the person who writes the markup does not have
>> access to general HTML facilities. But at the same time I need to
>> provide enough to allow for basic markup of text and probably a few
>> other things.
>>
>
> It's impossible to ensure the person doesn't have access to general HTML
> facilities.  But then why do you care?

I am not sure what you are thinking of but because the input is to be 
marked-up text I will be able to ensure the person who wrote it does not 
have access to full HTML.

As for why I care about that, while most HTML would be OK it is too 
permissive. It would be far too much work to write an HTML parser and 
just recognise the dodgy bits such as embedded scripts.

Also, the app is to do a little more than can be done with HTML alone. 
For example, the conditional links mentioned below and adding a table of 
contents. The site styling and page headers and footers would also be 
maintained separately.

>> For example, I may have a /conditional/ link tag that, for local
>> targets, will check that the target exists. If the target does not exist
>> then the link will render as non-clickable text. If the target does
>> exist then the text will be clickable.
>>
>
> Or, just do like everyone else does.  If the user provides a link, it's
> clickable.  If the link is bad, that's not your problem - it's the
> person who provided the link.

I think it would be easier for a person to build up a site if he or she 
can liberally include links for topics that should be present, and then 
go back later and create those link targets.

> And that's not at all.  Just because the link can be accessed doesn't
> mean anything.  For instance, just a couple of days ago someone posted a
> link in a social network.  But the story was behind a paywall.

That well illustrates the annoyance of dead or useless links.

James

[toc] | [prev] | [next] | [standalone]


#16573

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-21 10:05 -0500
Message-ID<nacjhq$uop$1@jstuckle.eternal-september.org>
In reply to#16554
On 2/21/2016 3:49 AM, James Harris wrote:
> On 20/02/2016 22:50, Jerry Stuckle wrote:
>> On 2/20/2016 3:24 PM, James Harris wrote:
>>> On 20/02/2016 20:08, Ben Bacarisse wrote:
> 
> ...
> 
>>>> If you are happy (or
>>>> you are forced) to write comparatively fussy markup, then a
>>>> well-trodden
>>>> path would be to use XML and an XML processor to convert your custom
>>>> markup to HTML.
>>>
>>> I see what you mean but in this case the markup is partly for security.
>>> I need to ensure that the person who writes the markup does not have
>>> access to general HTML facilities. But at the same time I need to
>>> provide enough to allow for basic markup of text and probably a few
>>> other things.
>>>
>>
>> It's impossible to ensure the person doesn't have access to general HTML
>> facilities.  But then why do you care?
> 
> I am not sure what you are thinking of but because the input is to be
> marked-up text I will be able to ensure the person who wrote it does not
> have access to full HTML.
> 
> As for why I care about that, while most HTML would be OK it is too
> permissive. It would be far too much work to write an HTML parser and
> just recognise the dodgy bits such as embedded scripts.
>

It's not that hard at all.  Virtually every CMS has it. They even limit
the types of html tags which are allowed.  It would be much easier than
creating a whole new parsing language.  And don't expect any users to
learn it; can you imaging what it would be like if every social website
required someone to learn a new markup?  There's a reason they don't.

> Also, the app is to do a little more than can be done with HTML alone.
> For example, the conditional links mentioned below and adding a table of
> contents. The site styling and page headers and footers would also be
> maintained separately.
> 

That can all be done from HTML, also.

>>> For example, I may have a /conditional/ link tag that, for local
>>> targets, will check that the target exists. If the target does not exist
>>> then the link will render as non-clickable text. If the target does
>>> exist then the text will be clickable.
>>>
>>
>> Or, just do like everyone else does.  If the user provides a link, it's
>> clickable.  If the link is bad, that's not your problem - it's the
>> person who provided the link.
> 
> I think it would be easier for a person to build up a site if he or she
> can liberally include links for topics that should be present, and then
> go back later and create those link targets.
>

Yes, but that's not a problem.  I don't publish a page until the links
are satisfied, anyway.

>> And that's not at all.  Just because the link can be accessed doesn't
>> mean anything.  For instance, just a couple of days ago someone posted a
>> link in a social network.  But the story was behind a paywall.
> 
> That well illustrates the annoyance of dead or useless links.
>

Which your idea will never catch.

> James
> 

I just think you're going about this the wrong way.  I *want* to be able
to create HTML for what I want.  I *want* to be able to use CSS's to
style my pages.  And so does everyone I know who creates websites - even
"hobby" sites.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16548

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-20 17:45 -0500
Message-ID<naaq3o$guv$1@jstuckle.eternal-september.org>
In reply to#16530
On 2/20/2016 1:51 PM, James Harris wrote:
> On 18/02/2016 15:43, Jerry Stuckle wrote:
>> On 2/18/2016 7:47 AM, James Harris wrote:
>>> Quick question: If someone wanted to parse marked-up text (and to
>>> convert it to HTML) do you have any generic recommendation on a good
>>> approach to use in PHP?
>>>
>>> Specifically, the text could be parsed character-by-character or line by
>>> line. Or perhaps it could be split with explode() and processed in
>>> pieces.
>>>
>>> The input text could be ASCII or Unicode.
>>>
>>> The marking up would be by tags of some sort. The form of a tag is not
>>> yet defined; it could be chosen to make the parsing easier, if
>>> necessary, as long as it was easy for a human to work with.
> 
> ...
> 
>> Way to vague a question for an intelligent answer, and any answer at
>> this juncture is just a guess.  It all depends on how the text is going
>> to be marked up.
> 
> The markup can be chosen. As long as it is easy to parse and easy for a
> human to work with it would be suitable. Most such markup would be
> converted to HTML elements: <hN>, <p>, <a>, etc and corresponding
> closing tags. There could also be self-contained tags which get
> converted to <br>, <hr> etc.
> 
> To illustrate, here is one possible scheme for opening and closing tags:
> 
> @(tagname, parameter, parameter)text@(/tagname)
> 
> Here is another option:
> 
> @tagname:parameter, parameter;text@/tagname;
> 
> In either case, "text" is not part of a tag.
> 
> There would additionally need to be a way to express a self-contained
> tag (i.e. one that does not need a following closing tag). Perhaps
> 
> @tagname:parameter, parameter/
> 
> Given that the above all use an at sign (@) to indicate the start of a
> tag there would need to be a convenient way to include an at sign in the
> output. Perhaps
> 
> @at/
> 
> As I say, the above are just examples of the kind of elements a markup
> would have to express, and the tag forms could be chosen to help parsing.
> 
> James
> 

But that's what you have to define before you can pick a way of parsing
it.  Different markups will require different parsing techniques.  If
you can define the markup, you should define it according to how you
want to parse it.  And it even can define which language you should be
using - PHP may or may not be the best.

But here, I agree with Ben.  If your markup is as complicated as you
want, you should just use XML or similar which already has parsing
libraries.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16555

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 09:01 +0000
Message-ID<nabu7k$jbn$1@dont-email.me>
In reply to#16548
On 20/02/2016 22:45, Jerry Stuckle wrote:

... (comments about markup and tag forms snipped)

> But that's what you have to define before you can pick a way of parsing
> it.  Different markups will require different parsing techniques.  If
> you can define the markup, you should define it according to how you
> want to parse it.

I am puzzled by this as you seem to say two different things: 1) define 
the markup before the parse method, 2) define the markup to suit (i.e. 
after) the intended parse method.

I did say before that the markup could be chosen to be easy to parse as 
long as it was also easy for a human to work with. I see no reason why a 
markup should be defined first.

For example, some markup schemes use different symbols to indicate 
different tags, such as

== Heading at level 2 ==
 > Indent
* Bulleted list element
| Cells | in a | table |

ISTM that such a markup would require character-by-character parsing and 
would have loads of contexts to manage, making it a nightmare to parse. 
By contrast, a single rarely used character or character sequence to 
begin each tag would be much easier to locate.

 > And it even can define which language you should be
 > using - PHP may or may not be the best.

True. I assume PHP is the right language to use on server-side code 
because it is so widespread. Are you saying that other languages are 
available server-side?

James

[toc] | [prev] | [next] | [standalone]


#16559

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-21 11:46 +0100
Message-ID<4598235.NEp8bqMrk4@PointedEars.de>
In reply to#16555
James Harris wrote:

> For example, some markup schemes use different symbols to indicate
> different tags, such as
> 
> == Heading at level 2 ==
>  > Indent
> * Bulleted list element
> | Cells | in a | table |
> 
> ISTM that such a markup would require character-by-character parsing

How did you get that idea?

> and would have loads of contexts to manage, making it a nightmare to
> parse. 

Yeah, that must be why it is used :->

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16561

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 12:36 +0000
Message-ID<nacarh$t85$1@dont-email.me>
In reply to#16559
On 21/02/2016 10:46, Thomas 'PointedEars' Lahn wrote:
> James Harris wrote:
>
>> For example, some markup schemes use different symbols to indicate
>> different tags, such as
>>
>> == Heading at level 2 ==
>>   > Indent
>> * Bulleted list element
>> | Cells | in a | table |
>>
>> ISTM that such a markup would require character-by-character parsing
>
> How did you get that idea?

I mean of the simple options I mentioned earlier. By contrast, if all 
tags begin with the same string (of length 1 or more) and that string 
was not permitted elsewhere (but could be built when needed) then tags 
would be very easy to locate.

James

[toc] | [prev] | [next] | [standalone]


#16565

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-21 14:39 +0100
Message-ID<6419381.RGxWzRkdQ9@PointedEars.de>
In reply to#16561
James Harris wrote:

> On 21/02/2016 10:46, Thomas 'PointedEars' Lahn wrote:
>> James Harris wrote:
>>> For example, some markup schemes use different symbols to indicate
>>> different tags, such as
>>>
>>> == Heading at level 2 ==
>>>   > Indent
>>> * Bulleted list element
>>> | Cells | in a | table |
>>>
>>> ISTM that such a markup would require character-by-character parsing
>>
>> How did you get that idea?
> 
> I mean of the simple options I mentioned earlier.

I have no idea what you are talking about.  (Do you?)  You have claimed that 
the markup above (Markdown) “would require character-by-character parsing”.  
It does not, because the Markdown rules are rather simple.

<https://en.wikipedia.org/wiki/Markdown>

> By contrast, if all tags begin with the same string (of length 1 or more)
> and that string was not permitted elsewhere (but could be built when
> needed) then tags would be very easy to locate.

Markdown begins “with the same string (of length 1 or more)”.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16566

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-21 13:40 +0000
Message-ID<nacej9$att$1@dont-email.me>
In reply to#16565
On 21/02/2016 13:39, Thomas 'PointedEars' Lahn wrote:

...

> I have no idea what you are talking about.  (Do you?)

Yes.

James

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | comp.lang.php


csiph-web