Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.php > #16438 > unrolled thread

PHP processing steps to apply to a URL to make it safe

Started byJames Harris <james.harris.1@gmail.com>
First post2016-02-17 09:37 +0000
Last post2016-03-22 05:23 -0700
Articles 20 on this page of 86 — 9 participants

Back to article view | Back to comp.lang.php


Contents

  PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 09:37 +0000
    Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 11:05 +0100
      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 15:04 +0000
        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:54 -0500
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:09 +0000
            Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 21:12 -0500
              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:07 +0000
                Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:07 +0100
                  Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 20:35 +0000
                    Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:40 +0100
                      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 21:02 +0000
                Re: PHP processing steps to apply to a URL to make it safe gordonb.lkcah@burditt.org (Gordon Burditt) - 2016-02-29 21:29 -0600
        Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:33 +0100
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:12 +0000
            Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 21:48 +0100
          Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:05 +0100
            Re: PHP processing steps to apply to a URL to make it safe Tim Streater <timstreater@greenbee.net> - 2016-02-18 09:59 +0000
              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 10:29 +0000
                Re: PHP processing steps to apply to a URL to make it safe Tim Streater <timstreater@greenbee.net> - 2016-02-18 11:36 +0000
              Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:44 +0100
            Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 21:47 +0100
        Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:02 +0100
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:29 +0000
            Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:37 +0100
              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 21:37 +0000
                Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 15:02 +0100
                  Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 15:02 +0000
    Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 08:19 -0500
      Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 14:43 +0100
        Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 15:04 +0100
          Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 15:17 +0100
            Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 15:47 +0100
              Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 15:57 +0100
                Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:57 -0500
                  Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 17:10 +0100
                    Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:11 -0500
                      Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 18:51 +0100
                        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:55 -0500
                Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:06 +0100
                  Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 20:19 +0100
                    Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:18 +0100
                      Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-18 09:20 +0100
                        Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:48 +0100
                          Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-18 16:25 +0100
                            Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 17:00 +0100
                              Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:26 +0100
                          Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:48 -0500
              Re: PHP processing steps to apply to a URL to make it safe "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-17 16:38 +0100
                Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:56 -0500
                Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:08 +0100
                Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:35 +0100
            Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 09:56 -0500
              Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 16:44 +0100
                Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:12 -0500
                  Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 19:04 +0100
        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 09:55 -0500
          Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 16:39 +0100
            Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 11:00 -0500
      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 15:15 +0000
        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 11:05 -0500
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:18 +0000
        Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:14 +0100
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 00:19 +0000
        Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:36 +0100
    Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:26 +0100
      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 00:36 +0000
        Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:03 +0100
          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:38 +0000
            Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:17 +0100
              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 20:56 +0000
                Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 23:48 +0100
                  Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 09:37 +0000
                    Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:42 +0100
                      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 15:03 +0000
                        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:09 -0500
                          Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 19:34 +0000
                            Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:49 -0500
                              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 19:56 +0000
                                Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 20:29 -0500
                                  Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-22 07:30 +0000
                                    Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-22 08:13 -0500
                                      Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-22 15:01 +0000
                                        Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-22 10:31 -0500
                            Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-22 00:10 +0100
                              Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 23:45 +0000
    Re: PHP processing steps to apply to a URL to make it safe pittendrigh <Sandy.Pittendrigh@gmail.com> - 2016-03-22 05:23 -0700

Page 1 of 5  [1] 2 3 4 5  Next page →


#16438 — PHP processing steps to apply to a URL to make it safe

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-17 09:37 +0000
SubjectPHP processing steps to apply to a URL to make it safe
Message-ID<na1es3$rsd$1@dont-email.me>
In line with the principle of a program vetting its input, this is about 
vetting the supplied URL.

In order to process a URL or parts of a URL in PHP what processing ought 
to be applied to it? I have been using some steps which I will write 
below but I am becoming increasingly uncertain that I have covered all 
the bases.

The main concern is safety - ensuring that the PHP code is protected 
against anything that might otherwise trip it up. The second concern is 
correctness in even corner cases such as odd browsers or extended 
character sets.

I have been working with $_SERVER["REQUEST_URI"], in case that is relevant.

Steps so far:

   urldecode to convert %nn etc

   trim "/" from the ends in order to normalise

   check each character is from an allowed set

   check there are no // parts

   check there are no .. parts

All needed? Any not needed? Latest firefox strips out any .. entries 
before sending the URL but I am not sure that all earlier browsers would.

In addition to the above, could the URL be in Unicode form? I was 
thinking to change to index through it with

   $uri[n]

but I gather that will index bytes and not characters. Could the URL 
string that PHP receives be stored as Unicode and should something like 
mb_substr be used instead?

That's all the steps I have come up with so far. It seems a lot. 
Presumably you guys have steps you use yourselves. Are all the things I 
have listed necessary? Is there anything else that should be done to 
process URLs (or components thereof) safely and correctly?

James

[toc] | [next] | [standalone]


#16440

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-17 11:05 +0100
Message-ID<56C445F6.1090008@arnowelzel.de>
In reply to#16438
James Harris schrieb am 2016-02-17 um 10:37:

> In line with the principle of a program vetting its input, this is about 
> vetting the supplied URL.
> 
> In order to process a URL or parts of a URL in PHP what processing ought 
> to be applied to it? I have been using some steps which I will write 
> below but I am becoming increasingly uncertain that I have covered all 
> the bases.

Well - it depends for which reasons you want to process the URL.

> The main concern is safety - ensuring that the PHP code is protected 
> against anything that might otherwise trip it up. The second concern is 
> correctness in even corner cases such as odd browsers or extended 
> character sets.

One of the main reasons for unsafe code in PHP is using eval() and SQL
statements with values you don't control on your own. You should also
never process XML without some sanity checks.

Further reading:
<http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html>

> I have been working with $_SERVER["REQUEST_URI"], in case that is relevant.
> 
> Steps so far:
> 
>    urldecode to convert %nn etc
> 
>    trim "/" from the ends in order to normalise

There is no "normal" URL - in fact the "/" at the end is part of the URL.

>    check each character is from an allowed set

Why?

>    check there are no // parts

Why?

>    check there are no .. parts

Why?

> All needed? Any not needed? Latest firefox strips out any .. entries 
> before sending the URL but I am not sure that all earlier browsers would.

And what would happen, if your script get's an URL with ".." and "//"
inside?

> In addition to the above, could the URL be in Unicode form? I was 
> thinking to change to index through it with

See RFC 1808:

<https://tools.ietf.org/html/rfc1808>

[...]
> That's all the steps I have come up with so far. It seems a lot. 
> Presumably you guys have steps you use yourselves. Are all the things I 
> have listed necessary? Is there anything else that should be done to 
> process URLs (or components thereof) safely and correctly?

As I said: It depends on what you want to achieve - do you make sure the
URL you get is valid? Do you want to parse the URL in any way?


-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


#16451

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-17 15:04 +0000
Message-ID<na220i$27e$1@dont-email.me>
In reply to#16440
On 17/02/2016 10:05, Arno Welzel wrote:
> James Harris schrieb am 2016-02-17 um 10:37:
>
>> In line with the principle of a program vetting its input, this is about
>> vetting the supplied URL.
>>
>> In order to process a URL or parts of a URL in PHP what processing ought
>> to be applied to it? I have been using some steps which I will write
>> below but I am becoming increasingly uncertain that I have covered all
>> the bases.
>
> Well - it depends for which reasons you want to process the URL.

It will go through a series of steps to change it and it will end up 
identifying something in the file system or something in a database.

>> The main concern is safety - ensuring that the PHP code is protected
>> against anything that might otherwise trip it up. The second concern is
>> correctness in even corner cases such as odd browsers or extended
>> character sets.
>
> One of the main reasons for unsafe code in PHP is using eval() and SQL
> statements with values you don't control on your own. You should also
> never process XML without some sanity checks.

Acknowledged that they are big issues but IMO all inputs should be checked.

> Further reading:
> <http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html>
>
>> I have been working with $_SERVER["REQUEST_URI"], in case that is relevant.
>>
>> Steps so far:
>>
>>     urldecode to convert %nn etc
>>
>>     trim "/" from the ends in order to normalise
>
> There is no "normal" URL - in fact the "/" at the end is part of the URL.

The docs for request_uri (which I'll use as shorthand for the full 
expression) did not specify whether leading or trailing slashes would be 
present/maintained.

>>     check each character is from an allowed set
>
> Why?

My own insecurity, perhaps!

>>     check there are no // parts
>
> Why?

Because such would designate an empty path component as between "the" 
and "bad" in http://site.com/this/is/the//bad/part.

>>     check there are no .. parts
>
> Why?

Because that indicates the parent directory and can be used to move 'up' 
a level in the path hierarchy. It is thus insecure.

>> All needed? Any not needed? Latest firefox strips out any .. entries
>> before sending the URL but I am not sure that all earlier browsers would.
>
> And what would happen, if your script get's an URL with ".." and "//"
> inside?

At the moment I report them as errors.

>> In addition to the above, could the URL be in Unicode form? I was
>> thinking to change to index through it with
>
> See RFC 1808:
>
> <https://tools.ietf.org/html/rfc1808>

I looked through it but could not see anything about Unicode. The info 
on relative/partial URLs was useful, however.

I since found that PHP has functions such as parse_url but that would 
not help me because while it can extract the piece I am interested in I 
also want to vet the whole string so extracting a part is not useful.

> [...]
>> That's all the steps I have come up with so far. It seems a lot.
>> Presumably you guys have steps you use yourselves. Are all the things I
>> have listed necessary? Is there anything else that should be done to
>> process URLs (or components thereof) safely and correctly?
>
> As I said: It depends on what you want to achieve - do you make sure the
> URL you get is valid? Do you want to parse the URL in any way?

The URL component (request_uri) will be mapped to a file or a database 
entry in this case.

James

[toc] | [prev] | [next] | [standalone]


#16456

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-17 10:54 -0500
Message-ID<na24tj$nkl$1@jstuckle.eternal-september.org>
In reply to#16451
On 2/17/2016 10:04 AM, James Harris wrote:
> On 17/02/2016 10:05, Arno Welzel wrote:
>> James Harris schrieb am 2016-02-17 um 10:37:
>>
>>> In line with the principle of a program vetting its input, this is about
>>> vetting the supplied URL.
>>>
>>> In order to process a URL or parts of a URL in PHP what processing ought
>>> to be applied to it? I have been using some steps which I will write
>>> below but I am becoming increasingly uncertain that I have covered all
>>> the bases.
>>
>> Well - it depends for which reasons you want to process the URL.
> 
> It will go through a series of steps to change it and it will end up
> identifying something in the file system or something in a database.
>

That can be a very complicated process and prone to errors.

>>> The main concern is safety - ensuring that the PHP code is protected
>>> against anything that might otherwise trip it up. The second concern is
>>> correctness in even corner cases such as odd browsers or extended
>>> character sets.
>>
>> One of the main reasons for unsafe code in PHP is using eval() and SQL
>> statements with values you don't control on your own. You should also
>> never process XML without some sanity checks.
> 
> Acknowledged that they are big issues but IMO all inputs should be checked.
> 

Yes, you need to check all input from the user.  However, WHAT you check
and HOW you check it depend in a large way as to how the data will be used.

>> Further reading:
>> <http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html>
>>
>>> I have been working with $_SERVER["REQUEST_URI"], in case that is
>>> relevant.
>>>
>>> Steps so far:
>>>
>>>     urldecode to convert %nn etc
>>>
>>>     trim "/" from the ends in order to normalise
>>
>> There is no "normal" URL - in fact the "/" at the end is part of the URL.
> 
> The docs for request_uri (which I'll use as shorthand for the full
> expression) did not specify whether leading or trailing slashes would be
> present/maintained.
> 

It has whatever was passed from the web server.  PHP does not change it.

>>>     check each character is from an allowed set
>>
>> Why?
> 
> My own insecurity, perhaps!
> 

It can only be in the charset allowed by URLs - or it wouldn't have
gotten this far.

>>>     check there are no // parts
>>
>> Why?
> 
> Because such would designate an empty path component as between "the"
> and "bad" in http://site.com/this/is/the//bad/part.
> 

See above.

>>>     check there are no .. parts
>>
>> Why?
> 
> Because that indicates the parent directory and can be used to move 'up'
> a level in the path hierarchy. It is thus insecure.
>

Not past the DOCUMENT_ROOT.  The web server will not allow it.

>>> All needed? Any not needed? Latest firefox strips out any .. entries
>>> before sending the URL but I am not sure that all earlier browsers
>>> would.
>>
>> And what would happen, if your script get's an URL with ".." and "//"
>> inside?
> 
> At the moment I report them as errors.
> 

They are perfectly valid.

>>> In addition to the above, could the URL be in Unicode form? I was
>>> thinking to change to index through it with
>>
>> See RFC 1808:
>>
>> <https://tools.ietf.org/html/rfc1808>
> 
> I looked through it but could not see anything about Unicode. The info
> on relative/partial URLs was useful, however.
> 

That's because only the ASCII subset of characters (and not all of them)
are allowed in a URL.

> I since found that PHP has functions such as parse_url but that would
> not help me because while it can extract the piece I am interested in I
> also want to vet the whole string so extracting a part is not useful.
> 
>> [...]
>>> That's all the steps I have come up with so far. It seems a lot.
>>> Presumably you guys have steps you use yourselves. Are all the things I
>>> have listed necessary? Is there anything else that should be done to
>>> process URLs (or components thereof) safely and correctly?
>>
>> As I said: It depends on what you want to achieve - do you make sure the
>> URL you get is valid? Do you want to parse the URL in any way?
> 
> The URL component (request_uri) will be mapped to a file or a database
> entry in this case.
> 
> James
> 

I still don't understand exactly what you're trying to do.  Are all page
requests to your website to be redirected to this script?  If so, you'll
have to validate each part of the path - not just the entire uri.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16485

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-18 01:09 +0000
Message-ID<na35fm$qc4$1@dont-email.me>
In reply to#16456
On 17/02/2016 15:54, Jerry Stuckle wrote:
> On 2/17/2016 10:04 AM, James Harris wrote:
>> On 17/02/2016 10:05, Arno Welzel wrote:
>>> James Harris schrieb am 2016-02-17 um 10:37:

...

>>>> I have been working with $_SERVER["REQUEST_URI"], in case that is
>>>> relevant.
>>>>
>>>> Steps so far:
>>>>
>>>>      urldecode to convert %nn etc
>>>>
>>>>      trim "/" from the ends in order to normalise
>>>
>>> There is no "normal" URL - in fact the "/" at the end is part of the URL.
>>
>> The docs for request_uri (which I'll use as shorthand for the full
>> expression) did not specify whether leading or trailing slashes would be
>> present/maintained.
>>
>
> It has whatever was passed from the web server.  PHP does not change it.

It seems right, then, to remove the slash from the front, if there is 
one. I will, however, give the potential trailing slash some more 
thought. I may be best to leave that in place, if there is one, and 
report it as an error. Thanks for pointing that out.

>>>>      check each character is from an allowed set
>>>
>>> Why?
>>
>> My own insecurity, perhaps!
>>
>
> It can only be in the charset allowed by URLs - or it wouldn't have
> gotten this far.

My bad. I should have said that I do the charset checking *after* using 
urldecode(). The idea is as far as possible to vet the string that the 
user sees in his browser's location box. I guess browsers differ but 
from tests on recent Firefox and IE they display the glyphs in the 
location bar but pass them encoded in the request URI. I need to check 
the characters after converting them back to the glyphs.

>>>>      check there are no // parts
>>>
>>> Why?
>>
>> Because such would designate an empty path component as between "the"
>> and "bad" in http://site.com/this/is/the//bad/part.
>>
>
> See above.

Sorry, not sure which bit above, but I don't want a URL with a double 
slash in the path to work. Perhaps N slashes could be merged into a 
single slash but that would make two URLs identify the same resource on 
the site. That is not how this is intended to work. It's a bit 
complicated but to try to explain

   http://site.com/folder/file
   http://site.com/folder//file

Those two URLs would identify a virtual folder and a virtual file on the 
site. The PHP code maps them to a real resource - which could be in a 
database or it could be a file. The URL structure is intended to be 
independent of and to outlive any changes to how the files or resources 
are stored. I therefore want to make sure that users use the single 
correct URL from day 1.

>>>>      check there are no .. parts
>>>
>>> Why?
>>
>> Because that indicates the parent directory and can be used to move 'up'
>> a level in the path hierarchy. It is thus insecure.
>>
>
> Not past the DOCUMENT_ROOT.  The web server will not allow it.

Nevertheless, for the same reasons as above, the URL structure does not 
necessarily mean a part of a filesystem. To extend the earlier example,

   http://site.com/folder/file
   http://site.com/folder/dummy/../file

If both of those specified a part of the file system then the two URLs 
could refer to the same file. But it is not intended to be a file path. 
The .. entry, therefore, needs to be reported as an error and not used. 
It cannot be used to backspace over the dummy folder name that precedes it.

When I tested this, browsers automatically removed .. entries from their 
location bars but I did get at least one browser to send .. by using the 
percent code for a dot (%2c, IIRC) twice.

>>>> All needed? Any not needed? Latest firefox strips out any .. entries
>>>> before sending the URL but I am not sure that all earlier browsers
>>>> would.
>>>
>>> And what would happen, if your script get's an URL with ".." and "//"
>>> inside?
>>
>> At the moment I report them as errors.
>>
>
> They are perfectly valid.

I prohibit them for the reasons mentioned above.

...

>>>> That's all the steps I have come up with so far. It seems a lot.
>>>> Presumably you guys have steps you use yourselves. Are all the things I
>>>> have listed necessary? Is there anything else that should be done to
>>>> process URLs (or components thereof) safely and correctly?
>>>
>>> As I said: It depends on what you want to achieve - do you make sure the
>>> URL you get is valid? Do you want to parse the URL in any way?
>>
>> The URL component (request_uri) will be mapped to a file or a database
>> entry in this case.
>>
>> James
>>
>
> I still don't understand exactly what you're trying to do.  Are all page
> requests to your website to be redirected to this script?  If so, you'll
> have to validate each part of the path - not just the entire uri.

Yes, at least all the relevant ones will go to this script.

What do you mean, "validate each part of the path - not just the entire 
uri"? To make an example,

    http://site.com:port/folder1/folder2/file

In that, my script will only be activated if http://site.com:port is 
correct. So AFAICS, I only need to validate what follows it.

James

[toc] | [prev] | [next] | [standalone]


#16488

FromJerry Stuckle <jstucklex@attglobal.net>
Date2016-02-17 21:12 -0500
Message-ID<na3940$3qd$1@jstuckle.eternal-september.org>
In reply to#16485
On 2/17/2016 8:09 PM, James Harris wrote:
> On 17/02/2016 15:54, Jerry Stuckle wrote:
>> On 2/17/2016 10:04 AM, James Harris wrote:
>>> On 17/02/2016 10:05, Arno Welzel wrote:
>>>> James Harris schrieb am 2016-02-17 um 10:37:
> 
> ...
> 
>>>>> I have been working with $_SERVER["REQUEST_URI"], in case that is
>>>>> relevant.
>>>>>
>>>>> Steps so far:
>>>>>
>>>>>      urldecode to convert %nn etc
>>>>>
>>>>>      trim "/" from the ends in order to normalise
>>>>
>>>> There is no "normal" URL - in fact the "/" at the end is part of the
>>>> URL.
>>>
>>> The docs for request_uri (which I'll use as shorthand for the full
>>> expression) did not specify whether leading or trailing slashes would be
>>> present/maintained.
>>>
>>
>> It has whatever was passed from the web server.  PHP does not change it.
> 
> It seems right, then, to remove the slash from the front, if there is
> one. I will, however, give the potential trailing slash some more
> thought. I may be best to leave that in place, if there is one, and
> report it as an error. Thanks for pointing that out.
> 

It is NOT an error to have a trailing slash in a URL.

>>>>>      check each character is from an allowed set
>>>>
>>>> Why?
>>>
>>> My own insecurity, perhaps!
>>>
>>
>> It can only be in the charset allowed by URLs - or it wouldn't have
>> gotten this far.
> 
> My bad. I should have said that I do the charset checking *after* using
> urldecode(). The idea is as far as possible to vet the string that the
> user sees in his browser's location box. I guess browsers differ but
> from tests on recent Firefox and IE they display the glyphs in the
> location bar but pass them encoded in the request URI. I need to check
> the characters after converting them back to the glyphs.
> 
>>>>>      check there are no // parts
>>>>
>>>> Why?
>>>
>>> Because such would designate an empty path component as between "the"
>>> and "bad" in http://site.com/this/is/the//bad/part.
>>>
>>
>> See above.
> 
> Sorry, not sure which bit above, but I don't want a URL with a double
> slash in the path to work. Perhaps N slashes could be merged into a
> single slash but that would make two URLs identify the same resource on
> the site. That is not how this is intended to work. It's a bit
> complicated but to try to explain
> 
>   http://site.com/folder/file
>   http://site.com/folder//file
> 
> Those two URLs would identify a virtual folder and a virtual file on the
> site. The PHP code maps them to a real resource - which could be in a
> database or it could be a file. The URL structure is intended to be
> independent of and to outlive any changes to how the files or resources
> are stored. I therefore want to make sure that users use the single
> correct URL from day 1.
>

You need to be careful here.  You shouldn't be changing the rules as to
what is a valid or invalid URL.

>>>>>      check there are no .. parts
>>>>
>>>> Why?
>>>
>>> Because that indicates the parent directory and can be used to move 'up'
>>> a level in the path hierarchy. It is thus insecure.
>>>
>>
>> Not past the DOCUMENT_ROOT.  The web server will not allow it.
> 
> Nevertheless, for the same reasons as above, the URL structure does not
> necessarily mean a part of a filesystem. To extend the earlier example,
> 
>   http://site.com/folder/file
>   http://site.com/folder/dummy/../file
> 
> If both of those specified a part of the file system then the two URLs
> could refer to the same file. But it is not intended to be a file path.
> The .. entry, therefore, needs to be reported as an error and not used.
> It cannot be used to backspace over the dummy folder name that precedes it.
> 
> When I tested this, browsers automatically removed .. entries from their
> location bars but I did get at least one browser to send .. by using the
> percent code for a dot (%2c, IIRC) twice.
> 

Yes, and you can get it with embedded objects, also, such as an image
which resides in a higher level directory (or a subdirectory therein).
Remember that embedded objects can be referenced relative to the
directory the page is loaded from.

>>>>> All needed? Any not needed? Latest firefox strips out any .. entries
>>>>> before sending the URL but I am not sure that all earlier browsers
>>>>> would.
>>>>
>>>> And what would happen, if your script get's an URL with ".." and "//"
>>>> inside?
>>>
>>> At the moment I report them as errors.
>>>
>>
>> They are perfectly valid.
> 
> I prohibit them for the reasons mentioned above.
> 
> ...
>

And you should not, for reasons mentioned above.

>>>>> That's all the steps I have come up with so far. It seems a lot.
>>>>> Presumably you guys have steps you use yourselves. Are all the
>>>>> things I
>>>>> have listed necessary? Is there anything else that should be done to
>>>>> process URLs (or components thereof) safely and correctly?
>>>>
>>>> As I said: It depends on what you want to achieve - do you make sure
>>>> the
>>>> URL you get is valid? Do you want to parse the URL in any way?
>>>
>>> The URL component (request_uri) will be mapped to a file or a database
>>> entry in this case.
>>>
>>> James
>>>
>>
>> I still don't understand exactly what you're trying to do.  Are all page
>> requests to your website to be redirected to this script?  If so, you'll
>> have to validate each part of the path - not just the entire uri.
> 
> Yes, at least all the relevant ones will go to this script.
> 
> What do you mean, "validate each part of the path - not just the entire
> uri"? To make an example,
> 
>    http://site.com:port/folder1/folder2/file
> 
> In that, my script will only be activated if http://site.com:port is
> correct. So AFAICS, I only need to validate what follows it.
> 
> James
> 

Yes, and you need to validate each part following the :port, separately
and together.

What you are doing is NOT simple, and you need to be very careful.
Allow all legitimate URLs while rejecting invalid ones can be very
difficult.

-- 
==================
Remove the "x" from my email address
Jerry Stuckle
jstucklex@attglobal.net
==================

[toc] | [prev] | [next] | [standalone]


#16531

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 19:07 +0000
Message-ID<naadca$bq$1@dont-email.me>
In reply to#16488
On 18/02/2016 02:12, Jerry Stuckle wrote:
> On 2/17/2016 8:09 PM, James Harris wrote:
>> On 17/02/2016 15:54, Jerry Stuckle wrote:
>>> On 2/17/2016 10:04 AM, James Harris wrote:
>>>> On 17/02/2016 10:05, Arno Welzel wrote:
>>>>> James Harris schrieb am 2016-02-17 um 10:37:

...

>>>>>>       check there are no // parts
>>>>>
>>>>> Why?
>>>>
>>>> Because such would designate an empty path component as between "the"
>>>> and "bad" in http://site.com/this/is/the//bad/part.
>>>>
>>>
>>> See above.
>>
>> Sorry, not sure which bit above, but I don't want a URL with a double
>> slash in the path to work. Perhaps N slashes could be merged into a
>> single slash but that would make two URLs identify the same resource on
>> the site. That is not how this is intended to work. It's a bit
>> complicated but to try to explain
>>
>>    http://site.com/folder/file
>>    http://site.com/folder//file
>>
>> Those two URLs would identify a virtual folder and a virtual file on the
>> site. The PHP code maps them to a real resource - which could be in a
>> database or it could be a file. The URL structure is intended to be
>> independent of and to outlive any changes to how the files or resources
>> are stored. I therefore want to make sure that users use the single
>> correct URL from day 1.
>>
>
> You need to be careful here.  You shouldn't be changing the rules as to
> what is a valid or invalid URL.

Sorry. I think I have been at fault in not being clear that although I 
am getting data from a URL my application is really using the URL path 
as a unique key. What the URL contains after the site name is taken as a 
hierarchical sequence of key values. In

   http://site.com/folder/subfolder/file

the "folder" is a key within the site; the "subfolder" is a key within 
the first key, etc.

As such, any // element is invalid because it specifies a null key. 
Also, /../ is invalid. It does not indicate a parent folder but the ".." 
key. And a trailing slash is (or at least could be seen as) invalid 
because it falsely indicates that another key is to follow.

So I am applying rules on top of those that apply to URLs, not taking or 
having to vet all URL rules as they stand.

James

[toc] | [prev] | [next] | [standalone]


#16536

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-20 21:07 +0100
Message-ID<20227200.nrqMxn3Nlc@PointedEars.de>
In reply to#16531
James Harris wrote:

> Sorry. I think I have been at fault in not being clear that although I
> am getting data from a URL my application is really using the URL path
> as a unique key. What the URL contains after the site name is taken as a
> hierarchical sequence of key values. In
> 
>    http://site.com/folder/subfolder/file

Please use the “example” TLD for examples instead:

  http://site.example/folder/subfolder/file

See also <http://tools.ietf.org/html/rfc2606>.
 
> the "folder" is a key within the site; the "subfolder" is a key within
> the first key, etc.
> 
> As such, any // element is invalid because it specifies a null key.
> Also, /../ is invalid. It does not indicate a parent folder but the ".."
> key. And a trailing slash is (or at least could be seen as) invalid
> because it falsely indicates that another key is to follow.
>  
> So I am applying rules on top of those that apply to URLs, not taking or
> having to vet all URL rules as they stand.

I stand corrected with regard to my earlier statement that it would be 
guaranteed that $_SERVER['REQUEST_URI'] would always start with a “/”.
I have confirmed that it can also be an absolute URI instead:

----------------------------------------------------------------------
$ echo '<?= $_SERVER["REQUEST_URI"] . "\n" ?>' > tmp/server.php
$ php -S localhost:1337 -t tmp/ &
[1] 25339
$ PHP 5.6.14-0+deb8u1 Development Server started at Sat Feb 20 21:03:50 2016
Listening on http://localhost:1337
Document root is /home/pelinux/tmp
Press Ctrl-C to quit.

$ telnet localhost 1337
Trying ::1...
Connected to localhost.
Escape character is '^]'.
GET http://localhost:1337/server.php HTTP/1.0

HTTP/1.0 200 OK
Connection: close
X-Powered-By: PHP/5.6.14-0+deb8u1
Content-type: text/html; charset=UTF-8

http://localhost:1337/server.php
Connection closed by foreign host.
----------------------------------------------------------------------

See also <http://tools.ietf.org/html/rfc2616#section-5.1.2>.

But what reason do you have to think that someone would make such an HTTP 
request to your server?  Also, what reason do you have to think that someone 
would make an HTTP request to your server that contains “..”, i.e. that this 
would not be resolved by the HTTP client already?

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16541

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 20:35 +0000
Message-ID<naaii9$kf5$1@dont-email.me>
In reply to#16536
On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote:
> James Harris wrote:
>
>> Sorry. I think I have been at fault in not being clear that although I
>> am getting data from a URL my application is really using the URL path
>> as a unique key. What the URL contains after the site name is taken as a
>> hierarchical sequence of key values. In
>>
>>     http://site.com/folder/subfolder/file
>
> Please use the “example” TLD for examples instead:
>
>    http://site.example/folder/subfolder/file

I see there is http://example.com as well. I might use that instead as 
people could find the .example TLD a bit odd.

> See also <http://tools.ietf.org/html/rfc2606>.

Thanks.

>> the "folder" is a key within the site; the "subfolder" is a key within
>> the first key, etc.
>>
>> As such, any // element is invalid because it specifies a null key.
>> Also, /../ is invalid. It does not indicate a parent folder but the ".."
>> key. And a trailing slash is (or at least could be seen as) invalid
>> because it falsely indicates that another key is to follow.
>>
>> So I am applying rules on top of those that apply to URLs, not taking or
>> having to vet all URL rules as they stand.

...

> But what reason do you have to think that someone would make such an HTTP
> request to your server?  Also, what reason do you have to think that someone
> would make an HTTP request to your server that contains “..”, i.e. that this
> would not be resolved by the HTTP client already?

I don't think they would. I think they might be able to. IMO vetting 
inputs is not about what sensible people might do but about what those 
with bad intent could do. And there are so many different versions of 
browser out there that we cannot test on anything like all of them.

James

[toc] | [prev] | [next] | [standalone]


#16543

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-20 21:40 +0100
Message-ID<2843221.kHc23yZ0jg@PointedEars.de>
In reply to#16541
James Harris wrote:

> On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote:
>> But what reason do you have to think that someone would make [an HTTP
>> request with an absolute URI] to your server?  Also, what reason do you
>> have to think that someone would make an HTTP request to your server that
>> contains “..”, i.e. that this would not be resolved by the HTTP client
>> already?
> 
> I don't think they would. I think they might be able to. IMO vetting
> inputs is not about what sensible people might do but about what those
> with bad intent could do. And there are so many different versions of
> browser out there that we cannot test on anything like all of them.

The “people with bad intent” would literally GET *nothing* in this case.
What is the harm in that?

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16545

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-20 21:02 +0000
Message-ID<naak44$q4o$1@dont-email.me>
In reply to#16543
On 20/02/2016 20:40, Thomas 'PointedEars' Lahn wrote:
> James Harris wrote:
>
>> On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote:
>>> But what reason do you have to think that someone would make [an HTTP
>>> request with an absolute URI] to your server?  Also, what reason do you
>>> have to think that someone would make an HTTP request to your server that
>>> contains “..”, i.e. that this would not be resolved by the HTTP client
>>> already?
>>
>> I don't think they would. I think they might be able to. IMO vetting
>> inputs is not about what sensible people might do but about what those
>> with bad intent could do. And there are so many different versions of
>> browser out there that we cannot test on anything like all of them.
>
> The “people with bad intent” would literally GET *nothing* in this case.
> What is the harm in that?

I am not sure what that means. We may be talking about different things. 
No problem.

James

[toc] | [prev] | [next] | [standalone]


#16607

Fromgordonb.lkcah@burditt.org (Gordon Burditt)
Date2016-02-29 21:29 -0600
Message-ID<uLydnfP7no0pkUjLnZ2dnUU7-KHNnZ2d@posted.internetamerica>
In reply to#16531
A URL is a string.  If you don't do anything potentially dangerous
with a string, there is no security issue.  Contrary to popular
opinion, there is no computer equivalent to the "killer joke", which
kills people who hear the joke, if the computer doesn't process it.

You haven't said whether or not the URL points to YOUR servers or
not.  Example: the vast majority of the URLs in Google's search
engine point to somewhere besides Google.  That's the whole point
of a worldwide search engine, right?  It's also not at all uncommon
for ads on a website to have URLs that point to the advertiser's
website, not the website you're viewing.  A PHP page might be set up
to select a random ad, output a link to it, and log the page view,
so subsequent runs prefer ads with fewer views.

Some potentially dangerous things to do with a string in PHP, especially 
if it contains user-supplied data, or comes from a database that might 
have unvetted user-supplied data:
	(1) Use parts of the string to reference the file system.
	    (bypassing web server permission restrictions)  Quoting doesn't
	    really work here - you need to reject attempts to access outside
	    the section of the filesystem.
	(2) Use parts of the string in a SQL query without proper quoting.
	    (SQL injection)
	(3) Use the string in content passed to eval() without proper quoting.
	    (Executing arbitrary code)
	(4) Use the string in content passed to a shell without proper quoting.
	    (Executing arbitrary code, `rm -rf /` being the standard bad example)
	    WARNING:  DO NOT USE THIS POST AS A SHELL "HERE" DOCUMENT.
	(5) Use the string as an email address passed to a mail transport service.
	    (Spamming, injecting extra destinations into email headers.)
	(6) Output the string to a web page. (Possible XSS attack.)
	(7) Validate the string using only Javascript, which can be turned off,
	    or with HTML (such as input length limits), which can be bypassed,
	    say, by telnet to the HTTP port, with a URL manually typed in.
	    (Bots tend not to use real browsers anyway.)  The string should also 
	    be validated on the server side, although letting the user know of a 
	    problem BEFORE pressing submit is more user-friendly.  
	(8) Using a non-constant string (subject to variable substitution) in 
	    a filesystem (or, *MUCH WORSE*, URL) reference, such as include, 
	    require, etc.
	(9) Using even a constant string that refers to a website you don't 
	    control in a reference that executes the output as code, such as 
	    include, require, etc.  DNS spoofing or playing games with ARP
	    could make even references to *YOUR* sites dangerous.
	(10) It's generally a bad idea to alter data in response to a HTTP GET
	    request, such as transferring money, ordering merchandise, or
	    deleting records (think about what a webcrawler might do to your
	    database!).  Use HTTP POST for that.  Logging hits and page view
	    counts is an exception.
	(11) The data fetched from a URL should be treated as user-supplied.
	Lots of other stuff I forgot to mention.
	
For example:
	http://www.google.com/news/../../../../../etc/passwd
is not dangerous because:
	(1) It does not refer to YOUR server, so it's not YOUR security problem.
	(2) If it *DOES* refer to YOUR server, rejecting the string doesn't
	    improve the situation much if a direct request to your web server
	    processes it anyway.  However, you need to avoid introducing new
	    problems of this type by copying a string that is part of a URL
	    reference to a filesystem reference.

> Sorry. I think I have been at fault in not being clear that although I 
> am getting data from a URL my application is really using the URL path 
> as a unique key. 

You appear to be using it as part of a *reference into your
filesystem*, which is dangerous (but it can be made safe with
appropriate checking), and that should suggest what checks are
necessary.  You also may have to deal with the "unique key" issue
of not really being unique if your filesystem is case-insensitive
but the URL is case-sensitive.

If the URL is supposed to refer to YOUR site, you can put arbitrary
restrictions on what is valid.  For example, you might allow / (if
it's first, and only one of them), and 0, O, o, i, I, L, l, 1, !,
and | (Note:  look carefully - there are no repeat characters in
that list, which intentionally contains characters that are hard
to distinguish from each other) (and absolutely *NO* other characters).
This might not make a whole lot of sense, but it's allowed.  That
ends the problem with ../ .  It's no longer just a URL, it's a
reference into your filesystem.  Or database, or whatever.
../ may be a problem in a filesystem reference but not in a database reference.


> What the URL contains after the site name is taken as a 
> hierarchical sequence of key values. In

Sequence of key values, or filesystem reference?  There's a difference.
A sequence of key values doesn't have an issue with ../ other than that
it's probably an invalid key.  It has other meanings in the filesystem.

>   http://site.com/folder/subfolder/file
> 
> the "folder" is a key within the site; the "subfolder" is a key within 
> the first key, etc.
> 
> As such, any // element is invalid because it specifies a null key. 
> Also, /../ is invalid. It does not indicate a parent folder but the ".." 
> key. And a trailing slash is (or at least could be seen as) invalid 
> because it falsely indicates that another key is to follow.
> 
> So I am applying rules on top of those that apply to URLs, not taking or 
> having to vet all URL rules as they stand.

[toc] | [prev] | [next] | [standalone]


#16477

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-17 23:33 +0100
Message-ID<3115586.UhRzssAi4L@PointedEars.de>
In reply to#16451
James Harris wrote:

> On 17/02/2016 10:05, Arno Welzel wrote:
>> James Harris schrieb am 2016-02-17 um 10:37:
>>>     check there are no // parts
>> Why?
> 
> Because such would designate an empty path component as between "the"
> and "bad" in http://site.com/this/is/the//bad/part.

How did you get the idea that this would be a problem?
 
>>>     check there are no .. parts
>>
>> Why?
> 
> Because that indicates the parent directory and can be used to move 'up'
> a level in the path hierarchy. It is thus insecure.

Why do you use user input verbatim in your code in the first place?
 
>>> In addition to the above, could the URL be in Unicode form? I was
>>> thinking to change to index through it with
>>
>> See RFC 1808:
>>
>> <https://tools.ietf.org/html/rfc1808>
> 
> I looked through it but could not see anything about Unicode.

It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains 
something about Unicode, in particular it defines UTF-8 percent-encoding.

> The info on relative/partial URLs was useful, however.

Do not rely on obsolete RFCs.

> The URL component (request_uri) will be mapped to a file or a database
> entry in this case.

BAD.  Broken as designed.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16486

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-18 01:12 +0000
Message-ID<na35l0$qc4$2@dont-email.me>
In reply to#16477
On 17/02/2016 22:33, Thomas 'PointedEars' Lahn wrote:

...

> Why do you use user input verbatim in your code in the first place?

I don't. The whole point of this query is about how to vet user input 
and I did explain that I use the user input to locate a resource.

James

[toc] | [prev] | [next] | [standalone]


#16510

FromThomas 'PointedEars' Lahn <PointedEars@web.de>
Date2016-02-18 21:48 +0100
Message-ID<1460842.BEhbNg33hS@PointedEars.de>
In reply to#16486
James Harris wrote:

> On 17/02/2016 22:33, Thomas 'PointedEars' Lahn wrote:
>> Why do you use user input verbatim in your code in the first place?
> 
> I don't. The whole point of this query is about how to vet user input
> and I did explain that I use the user input to locate a resource.

But ISTM that you are using and sanitizing the wrong kind of user input.

-- 
PointedEars
Zend Certified PHP Engineer 
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.

[toc] | [prev] | [next] | [standalone]


#16490

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-18 08:05 +0100
Message-ID<56C56D36.2060805@arnowelzel.de>
In reply to#16477
Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33:
> James Harris wrote:
> 
>> On 17/02/2016 10:05, Arno Welzel wrote:
>>> James Harris schrieb am 2016-02-17 um 10:37:
[...]
>>>> In addition to the above, could the URL be in Unicode form? I was
>>>> thinking to change to index through it with
>>>
>>> See RFC 1808:
>>>
>>> <https://tools.ietf.org/html/rfc1808>
>>
>> I looked through it but could not see anything about Unicode.
> 
> It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains 
> something about Unicode, in particular it defines UTF-8 percent-encoding.

That's the problem with RFCs - nothing in the old RFC indicates the
replacement ;-)

But thanks for the hint.



-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


#16494

FromTim Streater <timstreater@greenbee.net>
Date2016-02-18 09:59 +0000
Message-ID<180220160959369722%timstreater@greenbee.net>
In reply to#16490
In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel
<usenet@arnowelzel.de> wrote:

>Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33:
>> James Harris wrote:
>> 
>>> On 17/02/2016 10:05, Arno Welzel wrote:
>>>> James Harris schrieb am 2016-02-17 um 10:37:
>[...]
>>>>> In addition to the above, could the URL be in Unicode form? I was
>>>>> thinking to change to index through it with
>>>>
>>>> See RFC 1808:
>>>>
>>>> <https://tools.ietf.org/html/rfc1808>
>>>
>>> I looked through it but could not see anything about Unicode.
>> 
>> It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains 
>> something about Unicode, in particular it defines UTF-8 percent-encoding.
>
>That's the problem with RFCs - nothing in the old RFC indicates the
>replacement ;-)

Yes it does. Look at the top of the old RFC (assuming you didn't
download it on stone tablets in AD 54) and it will have an "Obsoleted
by:" line indicating which RFC replaces it.

-- 
"... you must remember that if you're trying to propagate a creed of 
 poverty, gentleness and tolerance,  you need a very rich, powerful,
 authoritarian organisation to do it."             - Vice-Pope Eric

[toc] | [prev] | [next] | [standalone]


#16495

FromJames Harris <james.harris.1@gmail.com>
Date2016-02-18 10:29 +0000
Message-ID<na468j$b44$1@dont-email.me>
In reply to#16494
On 18/02/2016 09:59, Tim Streater wrote:
> In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel
> <usenet@arnowelzel.de> wrote:

...

>> That's the problem with RFCs - nothing in the old RFC indicates the
>> replacement ;-)
>
> Yes it does. Look at the top of the old RFC (assuming you didn't
> download it on stone tablets in AD 54) and it will have an "Obsoleted
> by:" line indicating which RFC replaces it.

AIUI, Arno is right in that RFCs may not be modified once published. I 
think you mean that the "obsoleted by" header appears in certain 
presentations of an RFC - e.g. an HTML presentation.

Because of the immutable nature of RFCs, such a line cannot be added to 
the original. Compare these two.

   https://www.ietf.org/rfc/rfc1808.txt
   https://tools.ietf.org/html/rfc1808

James

[toc] | [prev] | [next] | [standalone]


#16496

FromTim Streater <timstreater@greenbee.net>
Date2016-02-18 11:36 +0000
Message-ID<180220161136429284%timstreater@greenbee.net>
In reply to#16495
In article <na468j$b44$1@dont-email.me>, James Harris
<james.harris.1@gmail.com> wrote:

>On 18/02/2016 09:59, Tim Streater wrote:
>> In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel
>> <usenet@arnowelzel.de> wrote:
>
>...
>
>>> That's the problem with RFCs - nothing in the old RFC indicates the
>>> replacement ;-)
>>
>> Yes it does. Look at the top of the old RFC (assuming you didn't
>> download it on stone tablets in AD 54) and it will have an "Obsoleted
>> by:" line indicating which RFC replaces it.
>
>AIUI, Arno is right in that RFCs may not be modified once published. I 
>think you mean that the "obsoleted by" header appears in certain 
>presentations of an RFC - e.g. an HTML presentation.
>
>Because of the immutable nature of RFCs, such a line cannot be added to 
>the original. Compare these two.
>
>   https://www.ietf.org/rfc/rfc1808.txt
>   https://tools.ietf.org/html/rfc1808

Well indeed but why look at the text version of an RFC? Not just for
the reason above but also the html versions allow easier navigation
within the document. Thass why I have the tools site you cite here on
my favourites list. Just a click away.

-- 
Lady Astor: "If you were my husband I'd give you poison." Churchill: "If
you were my wife, I'd drink it."

[toc] | [prev] | [next] | [standalone]


#16498

FromArno Welzel <usenet@arnowelzel.de>
Date2016-02-18 15:44 +0100
Message-ID<56C5D8C0.7030703@arnowelzel.de>
In reply to#16494
Tim Streater schrieb am 2016-02-18 um 10:59:

> In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel
> <usenet@arnowelzel.de> wrote:
> 
>> Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33:
>>> James Harris wrote:
>>>
>>>> On 17/02/2016 10:05, Arno Welzel wrote:
>>>>> James Harris schrieb am 2016-02-17 um 10:37:
>> [...]
>>>>>> In addition to the above, could the URL be in Unicode form? I was
>>>>>> thinking to change to index through it with
>>>>>
>>>>> See RFC 1808:
>>>>>
>>>>> <https://tools.ietf.org/html/rfc1808>
>>>>
>>>> I looked through it but could not see anything about Unicode.
>>>
>>> It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains 
>>> something about Unicode, in particular it defines UTF-8 percent-encoding.
>>
>> That's the problem with RFCs - nothing in the old RFC indicates the
>> replacement ;-)
> 
> Yes it does. Look at the top of the old RFC (assuming you didn't
> download it on stone tablets in AD 54) and it will have an "Obsoleted
> by:" line indicating which RFC replaces it.

I stand corrected - of course there is an indication.


-- 
Arno Welzel
http://arnowelzel.de
http://de-rec-fahrrad.de
http://fahrradzukunft.de

[toc] | [prev] | [next] | [standalone]


Page 1 of 5  [1] 2 3 4 5  Next page →

Back to top | Article view | comp.lang.php


csiph-web