Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.php > #16438 > unrolled thread
| Started by | James Harris <james.harris.1@gmail.com> |
|---|---|
| First post | 2016-02-17 09:37 +0000 |
| Last post | 2016-03-22 05:23 -0700 |
| Articles | 20 on this page of 86 — 9 participants |
Back to article view | Back to comp.lang.php
PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 09:37 +0000
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 11:05 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 15:04 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:54 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:09 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 21:12 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:07 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:07 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 20:35 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:40 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 21:02 +0000
Re: PHP processing steps to apply to a URL to make it safe gordonb.lkcah@burditt.org (Gordon Burditt) - 2016-02-29 21:29 -0600
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:33 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:12 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 21:48 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:05 +0100
Re: PHP processing steps to apply to a URL to make it safe Tim Streater <timstreater@greenbee.net> - 2016-02-18 09:59 +0000
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 10:29 +0000
Re: PHP processing steps to apply to a URL to make it safe Tim Streater <timstreater@greenbee.net> - 2016-02-18 11:36 +0000
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:44 +0100
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 21:47 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:02 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:29 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:37 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 21:37 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 15:02 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 15:02 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 08:19 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 14:43 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 15:04 +0100
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 15:17 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 15:47 +0100
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 15:57 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:57 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 17:10 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:11 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 18:51 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:55 -0500
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:06 +0100
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 20:19 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 08:18 +0100
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-18 09:20 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 15:48 +0100
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-18 16:25 +0100
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-18 17:00 +0100
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:26 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-18 10:48 -0500
Re: PHP processing steps to apply to a URL to make it safe "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-17 16:38 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 10:56 -0500
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:08 +0100
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:35 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 09:56 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 16:44 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 12:12 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 19:04 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 09:55 -0500
Re: PHP processing steps to apply to a URL to make it safe "R.Wieser" <address@not.available> - 2016-02-17 16:39 +0100
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 11:00 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-17 15:15 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-17 11:05 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 01:18 +0000
Re: PHP processing steps to apply to a URL to make it safe Arno Welzel <usenet@arnowelzel.de> - 2016-02-17 19:14 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 00:19 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:36 +0100
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:26 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-18 00:36 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-18 22:03 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 19:38 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 21:17 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-20 20:56 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-20 23:48 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 09:37 +0000
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-21 14:42 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 15:03 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:09 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 19:34 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 14:49 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 19:56 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-21 20:29 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-22 07:30 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-22 08:13 -0500
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-22 15:01 +0000
Re: PHP processing steps to apply to a URL to make it safe Jerry Stuckle <jstucklex@attglobal.net> - 2016-02-22 10:31 -0500
Re: PHP processing steps to apply to a URL to make it safe Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-22 00:10 +0100
Re: PHP processing steps to apply to a URL to make it safe James Harris <james.harris.1@gmail.com> - 2016-02-21 23:45 +0000
Re: PHP processing steps to apply to a URL to make it safe pittendrigh <Sandy.Pittendrigh@gmail.com> - 2016-03-22 05:23 -0700
Page 1 of 5 [1] 2 3 4 5 Next page →
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-17 09:37 +0000 |
| Subject | PHP processing steps to apply to a URL to make it safe |
| Message-ID | <na1es3$rsd$1@dont-email.me> |
In line with the principle of a program vetting its input, this is about vetting the supplied URL. In order to process a URL or parts of a URL in PHP what processing ought to be applied to it? I have been using some steps which I will write below but I am becoming increasingly uncertain that I have covered all the bases. The main concern is safety - ensuring that the PHP code is protected against anything that might otherwise trip it up. The second concern is correctness in even corner cases such as odd browsers or extended character sets. I have been working with $_SERVER["REQUEST_URI"], in case that is relevant. Steps so far: urldecode to convert %nn etc trim "/" from the ends in order to normalise check each character is from an allowed set check there are no // parts check there are no .. parts All needed? Any not needed? Latest firefox strips out any .. entries before sending the URL but I am not sure that all earlier browsers would. In addition to the above, could the URL be in Unicode form? I was thinking to change to index through it with $uri[n] but I gather that will index bytes and not characters. Could the URL string that PHP receives be stored as Unicode and should something like mb_substr be used instead? That's all the steps I have come up with so far. It seems a lot. Presumably you guys have steps you use yourselves. Are all the things I have listed necessary? Is there anything else that should be done to process URLs (or components thereof) safely and correctly? James
[toc] | [next] | [standalone]
| From | Arno Welzel <usenet@arnowelzel.de> |
|---|---|
| Date | 2016-02-17 11:05 +0100 |
| Message-ID | <56C445F6.1090008@arnowelzel.de> |
| In reply to | #16438 |
James Harris schrieb am 2016-02-17 um 10:37: > In line with the principle of a program vetting its input, this is about > vetting the supplied URL. > > In order to process a URL or parts of a URL in PHP what processing ought > to be applied to it? I have been using some steps which I will write > below but I am becoming increasingly uncertain that I have covered all > the bases. Well - it depends for which reasons you want to process the URL. > The main concern is safety - ensuring that the PHP code is protected > against anything that might otherwise trip it up. The second concern is > correctness in even corner cases such as odd browsers or extended > character sets. One of the main reasons for unsafe code in PHP is using eval() and SQL statements with values you don't control on your own. You should also never process XML without some sanity checks. Further reading: <http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html> > I have been working with $_SERVER["REQUEST_URI"], in case that is relevant. > > Steps so far: > > urldecode to convert %nn etc > > trim "/" from the ends in order to normalise There is no "normal" URL - in fact the "/" at the end is part of the URL. > check each character is from an allowed set Why? > check there are no // parts Why? > check there are no .. parts Why? > All needed? Any not needed? Latest firefox strips out any .. entries > before sending the URL but I am not sure that all earlier browsers would. And what would happen, if your script get's an URL with ".." and "//" inside? > In addition to the above, could the URL be in Unicode form? I was > thinking to change to index through it with See RFC 1808: <https://tools.ietf.org/html/rfc1808> [...] > That's all the steps I have come up with so far. It seems a lot. > Presumably you guys have steps you use yourselves. Are all the things I > have listed necessary? Is there anything else that should be done to > process URLs (or components thereof) safely and correctly? As I said: It depends on what you want to achieve - do you make sure the URL you get is valid? Do you want to parse the URL in any way? -- Arno Welzel http://arnowelzel.de http://de-rec-fahrrad.de http://fahrradzukunft.de
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-17 15:04 +0000 |
| Message-ID | <na220i$27e$1@dont-email.me> |
| In reply to | #16440 |
On 17/02/2016 10:05, Arno Welzel wrote: > James Harris schrieb am 2016-02-17 um 10:37: > >> In line with the principle of a program vetting its input, this is about >> vetting the supplied URL. >> >> In order to process a URL or parts of a URL in PHP what processing ought >> to be applied to it? I have been using some steps which I will write >> below but I am becoming increasingly uncertain that I have covered all >> the bases. > > Well - it depends for which reasons you want to process the URL. It will go through a series of steps to change it and it will end up identifying something in the file system or something in a database. >> The main concern is safety - ensuring that the PHP code is protected >> against anything that might otherwise trip it up. The second concern is >> correctness in even corner cases such as odd browsers or extended >> character sets. > > One of the main reasons for unsafe code in PHP is using eval() and SQL > statements with values you don't control on your own. You should also > never process XML without some sanity checks. Acknowledged that they are big issues but IMO all inputs should be checked. > Further reading: > <http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html> > >> I have been working with $_SERVER["REQUEST_URI"], in case that is relevant. >> >> Steps so far: >> >> urldecode to convert %nn etc >> >> trim "/" from the ends in order to normalise > > There is no "normal" URL - in fact the "/" at the end is part of the URL. The docs for request_uri (which I'll use as shorthand for the full expression) did not specify whether leading or trailing slashes would be present/maintained. >> check each character is from an allowed set > > Why? My own insecurity, perhaps! >> check there are no // parts > > Why? Because such would designate an empty path component as between "the" and "bad" in http://site.com/this/is/the//bad/part. >> check there are no .. parts > > Why? Because that indicates the parent directory and can be used to move 'up' a level in the path hierarchy. It is thus insecure. >> All needed? Any not needed? Latest firefox strips out any .. entries >> before sending the URL but I am not sure that all earlier browsers would. > > And what would happen, if your script get's an URL with ".." and "//" > inside? At the moment I report them as errors. >> In addition to the above, could the URL be in Unicode form? I was >> thinking to change to index through it with > > See RFC 1808: > > <https://tools.ietf.org/html/rfc1808> I looked through it but could not see anything about Unicode. The info on relative/partial URLs was useful, however. I since found that PHP has functions such as parse_url but that would not help me because while it can extract the piece I am interested in I also want to vet the whole string so extracting a part is not useful. > [...] >> That's all the steps I have come up with so far. It seems a lot. >> Presumably you guys have steps you use yourselves. Are all the things I >> have listed necessary? Is there anything else that should be done to >> process URLs (or components thereof) safely and correctly? > > As I said: It depends on what you want to achieve - do you make sure the > URL you get is valid? Do you want to parse the URL in any way? The URL component (request_uri) will be mapped to a file or a database entry in this case. James
[toc] | [prev] | [next] | [standalone]
| From | Jerry Stuckle <jstucklex@attglobal.net> |
|---|---|
| Date | 2016-02-17 10:54 -0500 |
| Message-ID | <na24tj$nkl$1@jstuckle.eternal-september.org> |
| In reply to | #16451 |
On 2/17/2016 10:04 AM, James Harris wrote: > On 17/02/2016 10:05, Arno Welzel wrote: >> James Harris schrieb am 2016-02-17 um 10:37: >> >>> In line with the principle of a program vetting its input, this is about >>> vetting the supplied URL. >>> >>> In order to process a URL or parts of a URL in PHP what processing ought >>> to be applied to it? I have been using some steps which I will write >>> below but I am becoming increasingly uncertain that I have covered all >>> the bases. >> >> Well - it depends for which reasons you want to process the URL. > > It will go through a series of steps to change it and it will end up > identifying something in the file system or something in a database. > That can be a very complicated process and prone to errors. >>> The main concern is safety - ensuring that the PHP code is protected >>> against anything that might otherwise trip it up. The second concern is >>> correctness in even corner cases such as odd browsers or extended >>> character sets. >> >> One of the main reasons for unsafe code in PHP is using eval() and SQL >> statements with values you don't control on your own. You should also >> never process XML without some sanity checks. > > Acknowledged that they are big issues but IMO all inputs should be checked. > Yes, you need to check all input from the user. However, WHAT you check and HOW you check it depend in a large way as to how the data will be used. >> Further reading: >> <http://phpsecurity.readthedocs.org/en/latest/Injection-Attacks.html> >> >>> I have been working with $_SERVER["REQUEST_URI"], in case that is >>> relevant. >>> >>> Steps so far: >>> >>> urldecode to convert %nn etc >>> >>> trim "/" from the ends in order to normalise >> >> There is no "normal" URL - in fact the "/" at the end is part of the URL. > > The docs for request_uri (which I'll use as shorthand for the full > expression) did not specify whether leading or trailing slashes would be > present/maintained. > It has whatever was passed from the web server. PHP does not change it. >>> check each character is from an allowed set >> >> Why? > > My own insecurity, perhaps! > It can only be in the charset allowed by URLs - or it wouldn't have gotten this far. >>> check there are no // parts >> >> Why? > > Because such would designate an empty path component as between "the" > and "bad" in http://site.com/this/is/the//bad/part. > See above. >>> check there are no .. parts >> >> Why? > > Because that indicates the parent directory and can be used to move 'up' > a level in the path hierarchy. It is thus insecure. > Not past the DOCUMENT_ROOT. The web server will not allow it. >>> All needed? Any not needed? Latest firefox strips out any .. entries >>> before sending the URL but I am not sure that all earlier browsers >>> would. >> >> And what would happen, if your script get's an URL with ".." and "//" >> inside? > > At the moment I report them as errors. > They are perfectly valid. >>> In addition to the above, could the URL be in Unicode form? I was >>> thinking to change to index through it with >> >> See RFC 1808: >> >> <https://tools.ietf.org/html/rfc1808> > > I looked through it but could not see anything about Unicode. The info > on relative/partial URLs was useful, however. > That's because only the ASCII subset of characters (and not all of them) are allowed in a URL. > I since found that PHP has functions such as parse_url but that would > not help me because while it can extract the piece I am interested in I > also want to vet the whole string so extracting a part is not useful. > >> [...] >>> That's all the steps I have come up with so far. It seems a lot. >>> Presumably you guys have steps you use yourselves. Are all the things I >>> have listed necessary? Is there anything else that should be done to >>> process URLs (or components thereof) safely and correctly? >> >> As I said: It depends on what you want to achieve - do you make sure the >> URL you get is valid? Do you want to parse the URL in any way? > > The URL component (request_uri) will be mapped to a file or a database > entry in this case. > > James > I still don't understand exactly what you're trying to do. Are all page requests to your website to be redirected to this script? If so, you'll have to validate each part of the path - not just the entire uri. -- ================== Remove the "x" from my email address Jerry Stuckle jstucklex@attglobal.net ==================
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-18 01:09 +0000 |
| Message-ID | <na35fm$qc4$1@dont-email.me> |
| In reply to | #16456 |
On 17/02/2016 15:54, Jerry Stuckle wrote:
> On 2/17/2016 10:04 AM, James Harris wrote:
>> On 17/02/2016 10:05, Arno Welzel wrote:
>>> James Harris schrieb am 2016-02-17 um 10:37:
...
>>>> I have been working with $_SERVER["REQUEST_URI"], in case that is
>>>> relevant.
>>>>
>>>> Steps so far:
>>>>
>>>> urldecode to convert %nn etc
>>>>
>>>> trim "/" from the ends in order to normalise
>>>
>>> There is no "normal" URL - in fact the "/" at the end is part of the URL.
>>
>> The docs for request_uri (which I'll use as shorthand for the full
>> expression) did not specify whether leading or trailing slashes would be
>> present/maintained.
>>
>
> It has whatever was passed from the web server. PHP does not change it.
It seems right, then, to remove the slash from the front, if there is
one. I will, however, give the potential trailing slash some more
thought. I may be best to leave that in place, if there is one, and
report it as an error. Thanks for pointing that out.
>>>> check each character is from an allowed set
>>>
>>> Why?
>>
>> My own insecurity, perhaps!
>>
>
> It can only be in the charset allowed by URLs - or it wouldn't have
> gotten this far.
My bad. I should have said that I do the charset checking *after* using
urldecode(). The idea is as far as possible to vet the string that the
user sees in his browser's location box. I guess browsers differ but
from tests on recent Firefox and IE they display the glyphs in the
location bar but pass them encoded in the request URI. I need to check
the characters after converting them back to the glyphs.
>>>> check there are no // parts
>>>
>>> Why?
>>
>> Because such would designate an empty path component as between "the"
>> and "bad" in http://site.com/this/is/the//bad/part.
>>
>
> See above.
Sorry, not sure which bit above, but I don't want a URL with a double
slash in the path to work. Perhaps N slashes could be merged into a
single slash but that would make two URLs identify the same resource on
the site. That is not how this is intended to work. It's a bit
complicated but to try to explain
http://site.com/folder/file
http://site.com/folder//file
Those two URLs would identify a virtual folder and a virtual file on the
site. The PHP code maps them to a real resource - which could be in a
database or it could be a file. The URL structure is intended to be
independent of and to outlive any changes to how the files or resources
are stored. I therefore want to make sure that users use the single
correct URL from day 1.
>>>> check there are no .. parts
>>>
>>> Why?
>>
>> Because that indicates the parent directory and can be used to move 'up'
>> a level in the path hierarchy. It is thus insecure.
>>
>
> Not past the DOCUMENT_ROOT. The web server will not allow it.
Nevertheless, for the same reasons as above, the URL structure does not
necessarily mean a part of a filesystem. To extend the earlier example,
http://site.com/folder/file
http://site.com/folder/dummy/../file
If both of those specified a part of the file system then the two URLs
could refer to the same file. But it is not intended to be a file path.
The .. entry, therefore, needs to be reported as an error and not used.
It cannot be used to backspace over the dummy folder name that precedes it.
When I tested this, browsers automatically removed .. entries from their
location bars but I did get at least one browser to send .. by using the
percent code for a dot (%2c, IIRC) twice.
>>>> All needed? Any not needed? Latest firefox strips out any .. entries
>>>> before sending the URL but I am not sure that all earlier browsers
>>>> would.
>>>
>>> And what would happen, if your script get's an URL with ".." and "//"
>>> inside?
>>
>> At the moment I report them as errors.
>>
>
> They are perfectly valid.
I prohibit them for the reasons mentioned above.
...
>>>> That's all the steps I have come up with so far. It seems a lot.
>>>> Presumably you guys have steps you use yourselves. Are all the things I
>>>> have listed necessary? Is there anything else that should be done to
>>>> process URLs (or components thereof) safely and correctly?
>>>
>>> As I said: It depends on what you want to achieve - do you make sure the
>>> URL you get is valid? Do you want to parse the URL in any way?
>>
>> The URL component (request_uri) will be mapped to a file or a database
>> entry in this case.
>>
>> James
>>
>
> I still don't understand exactly what you're trying to do. Are all page
> requests to your website to be redirected to this script? If so, you'll
> have to validate each part of the path - not just the entire uri.
Yes, at least all the relevant ones will go to this script.
What do you mean, "validate each part of the path - not just the entire
uri"? To make an example,
http://site.com:port/folder1/folder2/file
In that, my script will only be activated if http://site.com:port is
correct. So AFAICS, I only need to validate what follows it.
James
[toc] | [prev] | [next] | [standalone]
| From | Jerry Stuckle <jstucklex@attglobal.net> |
|---|---|
| Date | 2016-02-17 21:12 -0500 |
| Message-ID | <na3940$3qd$1@jstuckle.eternal-september.org> |
| In reply to | #16485 |
On 2/17/2016 8:09 PM, James Harris wrote: > On 17/02/2016 15:54, Jerry Stuckle wrote: >> On 2/17/2016 10:04 AM, James Harris wrote: >>> On 17/02/2016 10:05, Arno Welzel wrote: >>>> James Harris schrieb am 2016-02-17 um 10:37: > > ... > >>>>> I have been working with $_SERVER["REQUEST_URI"], in case that is >>>>> relevant. >>>>> >>>>> Steps so far: >>>>> >>>>> urldecode to convert %nn etc >>>>> >>>>> trim "/" from the ends in order to normalise >>>> >>>> There is no "normal" URL - in fact the "/" at the end is part of the >>>> URL. >>> >>> The docs for request_uri (which I'll use as shorthand for the full >>> expression) did not specify whether leading or trailing slashes would be >>> present/maintained. >>> >> >> It has whatever was passed from the web server. PHP does not change it. > > It seems right, then, to remove the slash from the front, if there is > one. I will, however, give the potential trailing slash some more > thought. I may be best to leave that in place, if there is one, and > report it as an error. Thanks for pointing that out. > It is NOT an error to have a trailing slash in a URL. >>>>> check each character is from an allowed set >>>> >>>> Why? >>> >>> My own insecurity, perhaps! >>> >> >> It can only be in the charset allowed by URLs - or it wouldn't have >> gotten this far. > > My bad. I should have said that I do the charset checking *after* using > urldecode(). The idea is as far as possible to vet the string that the > user sees in his browser's location box. I guess browsers differ but > from tests on recent Firefox and IE they display the glyphs in the > location bar but pass them encoded in the request URI. I need to check > the characters after converting them back to the glyphs. > >>>>> check there are no // parts >>>> >>>> Why? >>> >>> Because such would designate an empty path component as between "the" >>> and "bad" in http://site.com/this/is/the//bad/part. >>> >> >> See above. > > Sorry, not sure which bit above, but I don't want a URL with a double > slash in the path to work. Perhaps N slashes could be merged into a > single slash but that would make two URLs identify the same resource on > the site. That is not how this is intended to work. It's a bit > complicated but to try to explain > > http://site.com/folder/file > http://site.com/folder//file > > Those two URLs would identify a virtual folder and a virtual file on the > site. The PHP code maps them to a real resource - which could be in a > database or it could be a file. The URL structure is intended to be > independent of and to outlive any changes to how the files or resources > are stored. I therefore want to make sure that users use the single > correct URL from day 1. > You need to be careful here. You shouldn't be changing the rules as to what is a valid or invalid URL. >>>>> check there are no .. parts >>>> >>>> Why? >>> >>> Because that indicates the parent directory and can be used to move 'up' >>> a level in the path hierarchy. It is thus insecure. >>> >> >> Not past the DOCUMENT_ROOT. The web server will not allow it. > > Nevertheless, for the same reasons as above, the URL structure does not > necessarily mean a part of a filesystem. To extend the earlier example, > > http://site.com/folder/file > http://site.com/folder/dummy/../file > > If both of those specified a part of the file system then the two URLs > could refer to the same file. But it is not intended to be a file path. > The .. entry, therefore, needs to be reported as an error and not used. > It cannot be used to backspace over the dummy folder name that precedes it. > > When I tested this, browsers automatically removed .. entries from their > location bars but I did get at least one browser to send .. by using the > percent code for a dot (%2c, IIRC) twice. > Yes, and you can get it with embedded objects, also, such as an image which resides in a higher level directory (or a subdirectory therein). Remember that embedded objects can be referenced relative to the directory the page is loaded from. >>>>> All needed? Any not needed? Latest firefox strips out any .. entries >>>>> before sending the URL but I am not sure that all earlier browsers >>>>> would. >>>> >>>> And what would happen, if your script get's an URL with ".." and "//" >>>> inside? >>> >>> At the moment I report them as errors. >>> >> >> They are perfectly valid. > > I prohibit them for the reasons mentioned above. > > ... > And you should not, for reasons mentioned above. >>>>> That's all the steps I have come up with so far. It seems a lot. >>>>> Presumably you guys have steps you use yourselves. Are all the >>>>> things I >>>>> have listed necessary? Is there anything else that should be done to >>>>> process URLs (or components thereof) safely and correctly? >>>> >>>> As I said: It depends on what you want to achieve - do you make sure >>>> the >>>> URL you get is valid? Do you want to parse the URL in any way? >>> >>> The URL component (request_uri) will be mapped to a file or a database >>> entry in this case. >>> >>> James >>> >> >> I still don't understand exactly what you're trying to do. Are all page >> requests to your website to be redirected to this script? If so, you'll >> have to validate each part of the path - not just the entire uri. > > Yes, at least all the relevant ones will go to this script. > > What do you mean, "validate each part of the path - not just the entire > uri"? To make an example, > > http://site.com:port/folder1/folder2/file > > In that, my script will only be activated if http://site.com:port is > correct. So AFAICS, I only need to validate what follows it. > > James > Yes, and you need to validate each part following the :port, separately and together. What you are doing is NOT simple, and you need to be very careful. Allow all legitimate URLs while rejecting invalid ones can be very difficult. -- ================== Remove the "x" from my email address Jerry Stuckle jstucklex@attglobal.net ==================
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-20 19:07 +0000 |
| Message-ID | <naadca$bq$1@dont-email.me> |
| In reply to | #16488 |
On 18/02/2016 02:12, Jerry Stuckle wrote: > On 2/17/2016 8:09 PM, James Harris wrote: >> On 17/02/2016 15:54, Jerry Stuckle wrote: >>> On 2/17/2016 10:04 AM, James Harris wrote: >>>> On 17/02/2016 10:05, Arno Welzel wrote: >>>>> James Harris schrieb am 2016-02-17 um 10:37: ... >>>>>> check there are no // parts >>>>> >>>>> Why? >>>> >>>> Because such would designate an empty path component as between "the" >>>> and "bad" in http://site.com/this/is/the//bad/part. >>>> >>> >>> See above. >> >> Sorry, not sure which bit above, but I don't want a URL with a double >> slash in the path to work. Perhaps N slashes could be merged into a >> single slash but that would make two URLs identify the same resource on >> the site. That is not how this is intended to work. It's a bit >> complicated but to try to explain >> >> http://site.com/folder/file >> http://site.com/folder//file >> >> Those two URLs would identify a virtual folder and a virtual file on the >> site. The PHP code maps them to a real resource - which could be in a >> database or it could be a file. The URL structure is intended to be >> independent of and to outlive any changes to how the files or resources >> are stored. I therefore want to make sure that users use the single >> correct URL from day 1. >> > > You need to be careful here. You shouldn't be changing the rules as to > what is a valid or invalid URL. Sorry. I think I have been at fault in not being clear that although I am getting data from a URL my application is really using the URL path as a unique key. What the URL contains after the site name is taken as a hierarchical sequence of key values. In http://site.com/folder/subfolder/file the "folder" is a key within the site; the "subfolder" is a key within the first key, etc. As such, any // element is invalid because it specifies a null key. Also, /../ is invalid. It does not indicate a parent folder but the ".." key. And a trailing slash is (or at least could be seen as) invalid because it falsely indicates that another key is to follow. So I am applying rules on top of those that apply to URLs, not taking or having to vet all URL rules as they stand. James
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-20 21:07 +0100 |
| Message-ID | <20227200.nrqMxn3Nlc@PointedEars.de> |
| In reply to | #16531 |
James Harris wrote: > Sorry. I think I have been at fault in not being clear that although I > am getting data from a URL my application is really using the URL path > as a unique key. What the URL contains after the site name is taken as a > hierarchical sequence of key values. In > > http://site.com/folder/subfolder/file Please use the “example” TLD for examples instead: http://site.example/folder/subfolder/file See also <http://tools.ietf.org/html/rfc2606>. > the "folder" is a key within the site; the "subfolder" is a key within > the first key, etc. > > As such, any // element is invalid because it specifies a null key. > Also, /../ is invalid. It does not indicate a parent folder but the ".." > key. And a trailing slash is (or at least could be seen as) invalid > because it falsely indicates that another key is to follow. > > So I am applying rules on top of those that apply to URLs, not taking or > having to vet all URL rules as they stand. I stand corrected with regard to my earlier statement that it would be guaranteed that $_SERVER['REQUEST_URI'] would always start with a “/”. I have confirmed that it can also be an absolute URI instead: ---------------------------------------------------------------------- $ echo '<?= $_SERVER["REQUEST_URI"] . "\n" ?>' > tmp/server.php $ php -S localhost:1337 -t tmp/ & [1] 25339 $ PHP 5.6.14-0+deb8u1 Development Server started at Sat Feb 20 21:03:50 2016 Listening on http://localhost:1337 Document root is /home/pelinux/tmp Press Ctrl-C to quit. $ telnet localhost 1337 Trying ::1... Connected to localhost. Escape character is '^]'. GET http://localhost:1337/server.php HTTP/1.0 HTTP/1.0 200 OK Connection: close X-Powered-By: PHP/5.6.14-0+deb8u1 Content-type: text/html; charset=UTF-8 http://localhost:1337/server.php Connection closed by foreign host. ---------------------------------------------------------------------- See also <http://tools.ietf.org/html/rfc2616#section-5.1.2>. But what reason do you have to think that someone would make such an HTTP request to your server? Also, what reason do you have to think that someone would make an HTTP request to your server that contains “..”, i.e. that this would not be resolved by the HTTP client already? -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-20 20:35 +0000 |
| Message-ID | <naaii9$kf5$1@dont-email.me> |
| In reply to | #16536 |
On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote: > James Harris wrote: > >> Sorry. I think I have been at fault in not being clear that although I >> am getting data from a URL my application is really using the URL path >> as a unique key. What the URL contains after the site name is taken as a >> hierarchical sequence of key values. In >> >> http://site.com/folder/subfolder/file > > Please use the “example” TLD for examples instead: > > http://site.example/folder/subfolder/file I see there is http://example.com as well. I might use that instead as people could find the .example TLD a bit odd. > See also <http://tools.ietf.org/html/rfc2606>. Thanks. >> the "folder" is a key within the site; the "subfolder" is a key within >> the first key, etc. >> >> As such, any // element is invalid because it specifies a null key. >> Also, /../ is invalid. It does not indicate a parent folder but the ".." >> key. And a trailing slash is (or at least could be seen as) invalid >> because it falsely indicates that another key is to follow. >> >> So I am applying rules on top of those that apply to URLs, not taking or >> having to vet all URL rules as they stand. ... > But what reason do you have to think that someone would make such an HTTP > request to your server? Also, what reason do you have to think that someone > would make an HTTP request to your server that contains “..”, i.e. that this > would not be resolved by the HTTP client already? I don't think they would. I think they might be able to. IMO vetting inputs is not about what sensible people might do but about what those with bad intent could do. And there are so many different versions of browser out there that we cannot test on anything like all of them. James
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-20 21:40 +0100 |
| Message-ID | <2843221.kHc23yZ0jg@PointedEars.de> |
| In reply to | #16541 |
James Harris wrote: > On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote: >> But what reason do you have to think that someone would make [an HTTP >> request with an absolute URI] to your server? Also, what reason do you >> have to think that someone would make an HTTP request to your server that >> contains “..”, i.e. that this would not be resolved by the HTTP client >> already? > > I don't think they would. I think they might be able to. IMO vetting > inputs is not about what sensible people might do but about what those > with bad intent could do. And there are so many different versions of > browser out there that we cannot test on anything like all of them. The “people with bad intent” would literally GET *nothing* in this case. What is the harm in that? -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-20 21:02 +0000 |
| Message-ID | <naak44$q4o$1@dont-email.me> |
| In reply to | #16543 |
On 20/02/2016 20:40, Thomas 'PointedEars' Lahn wrote: > James Harris wrote: > >> On 20/02/2016 20:07, Thomas 'PointedEars' Lahn wrote: >>> But what reason do you have to think that someone would make [an HTTP >>> request with an absolute URI] to your server? Also, what reason do you >>> have to think that someone would make an HTTP request to your server that >>> contains “..”, i.e. that this would not be resolved by the HTTP client >>> already? >> >> I don't think they would. I think they might be able to. IMO vetting >> inputs is not about what sensible people might do but about what those >> with bad intent could do. And there are so many different versions of >> browser out there that we cannot test on anything like all of them. > > The “people with bad intent” would literally GET *nothing* in this case. > What is the harm in that? I am not sure what that means. We may be talking about different things. No problem. James
[toc] | [prev] | [next] | [standalone]
| From | gordonb.lkcah@burditt.org (Gordon Burditt) |
|---|---|
| Date | 2016-02-29 21:29 -0600 |
| Message-ID | <uLydnfP7no0pkUjLnZ2dnUU7-KHNnZ2d@posted.internetamerica> |
| In reply to | #16531 |
A URL is a string. If you don't do anything potentially dangerous with a string, there is no security issue. Contrary to popular opinion, there is no computer equivalent to the "killer joke", which kills people who hear the joke, if the computer doesn't process it. You haven't said whether or not the URL points to YOUR servers or not. Example: the vast majority of the URLs in Google's search engine point to somewhere besides Google. That's the whole point of a worldwide search engine, right? It's also not at all uncommon for ads on a website to have URLs that point to the advertiser's website, not the website you're viewing. A PHP page might be set up to select a random ad, output a link to it, and log the page view, so subsequent runs prefer ads with fewer views. Some potentially dangerous things to do with a string in PHP, especially if it contains user-supplied data, or comes from a database that might have unvetted user-supplied data: (1) Use parts of the string to reference the file system. (bypassing web server permission restrictions) Quoting doesn't really work here - you need to reject attempts to access outside the section of the filesystem. (2) Use parts of the string in a SQL query without proper quoting. (SQL injection) (3) Use the string in content passed to eval() without proper quoting. (Executing arbitrary code) (4) Use the string in content passed to a shell without proper quoting. (Executing arbitrary code, `rm -rf /` being the standard bad example) WARNING: DO NOT USE THIS POST AS A SHELL "HERE" DOCUMENT. (5) Use the string as an email address passed to a mail transport service. (Spamming, injecting extra destinations into email headers.) (6) Output the string to a web page. (Possible XSS attack.) (7) Validate the string using only Javascript, which can be turned off, or with HTML (such as input length limits), which can be bypassed, say, by telnet to the HTTP port, with a URL manually typed in. (Bots tend not to use real browsers anyway.) The string should also be validated on the server side, although letting the user know of a problem BEFORE pressing submit is more user-friendly. (8) Using a non-constant string (subject to variable substitution) in a filesystem (or, *MUCH WORSE*, URL) reference, such as include, require, etc. (9) Using even a constant string that refers to a website you don't control in a reference that executes the output as code, such as include, require, etc. DNS spoofing or playing games with ARP could make even references to *YOUR* sites dangerous. (10) It's generally a bad idea to alter data in response to a HTTP GET request, such as transferring money, ordering merchandise, or deleting records (think about what a webcrawler might do to your database!). Use HTTP POST for that. Logging hits and page view counts is an exception. (11) The data fetched from a URL should be treated as user-supplied. Lots of other stuff I forgot to mention. For example: http://www.google.com/news/../../../../../etc/passwd is not dangerous because: (1) It does not refer to YOUR server, so it's not YOUR security problem. (2) If it *DOES* refer to YOUR server, rejecting the string doesn't improve the situation much if a direct request to your web server processes it anyway. However, you need to avoid introducing new problems of this type by copying a string that is part of a URL reference to a filesystem reference. > Sorry. I think I have been at fault in not being clear that although I > am getting data from a URL my application is really using the URL path > as a unique key. You appear to be using it as part of a *reference into your filesystem*, which is dangerous (but it can be made safe with appropriate checking), and that should suggest what checks are necessary. You also may have to deal with the "unique key" issue of not really being unique if your filesystem is case-insensitive but the URL is case-sensitive. If the URL is supposed to refer to YOUR site, you can put arbitrary restrictions on what is valid. For example, you might allow / (if it's first, and only one of them), and 0, O, o, i, I, L, l, 1, !, and | (Note: look carefully - there are no repeat characters in that list, which intentionally contains characters that are hard to distinguish from each other) (and absolutely *NO* other characters). This might not make a whole lot of sense, but it's allowed. That ends the problem with ../ . It's no longer just a URL, it's a reference into your filesystem. Or database, or whatever. ../ may be a problem in a filesystem reference but not in a database reference. > What the URL contains after the site name is taken as a > hierarchical sequence of key values. In Sequence of key values, or filesystem reference? There's a difference. A sequence of key values doesn't have an issue with ../ other than that it's probably an invalid key. It has other meanings in the filesystem. > http://site.com/folder/subfolder/file > > the "folder" is a key within the site; the "subfolder" is a key within > the first key, etc. > > As such, any // element is invalid because it specifies a null key. > Also, /../ is invalid. It does not indicate a parent folder but the ".." > key. And a trailing slash is (or at least could be seen as) invalid > because it falsely indicates that another key is to follow. > > So I am applying rules on top of those that apply to URLs, not taking or > having to vet all URL rules as they stand.
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-17 23:33 +0100 |
| Message-ID | <3115586.UhRzssAi4L@PointedEars.de> |
| In reply to | #16451 |
James Harris wrote: > On 17/02/2016 10:05, Arno Welzel wrote: >> James Harris schrieb am 2016-02-17 um 10:37: >>> check there are no // parts >> Why? > > Because such would designate an empty path component as between "the" > and "bad" in http://site.com/this/is/the//bad/part. How did you get the idea that this would be a problem? >>> check there are no .. parts >> >> Why? > > Because that indicates the parent directory and can be used to move 'up' > a level in the path hierarchy. It is thus insecure. Why do you use user input verbatim in your code in the first place? >>> In addition to the above, could the URL be in Unicode form? I was >>> thinking to change to index through it with >> >> See RFC 1808: >> >> <https://tools.ietf.org/html/rfc1808> > > I looked through it but could not see anything about Unicode. It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains something about Unicode, in particular it defines UTF-8 percent-encoding. > The info on relative/partial URLs was useful, however. Do not rely on obsolete RFCs. > The URL component (request_uri) will be mapped to a file or a database > entry in this case. BAD. Broken as designed. -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-18 01:12 +0000 |
| Message-ID | <na35l0$qc4$2@dont-email.me> |
| In reply to | #16477 |
On 17/02/2016 22:33, Thomas 'PointedEars' Lahn wrote: ... > Why do you use user input verbatim in your code in the first place? I don't. The whole point of this query is about how to vet user input and I did explain that I use the user input to locate a resource. James
[toc] | [prev] | [next] | [standalone]
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
|---|---|
| Date | 2016-02-18 21:48 +0100 |
| Message-ID | <1460842.BEhbNg33hS@PointedEars.de> |
| In reply to | #16486 |
James Harris wrote: > On 17/02/2016 22:33, Thomas 'PointedEars' Lahn wrote: >> Why do you use user input verbatim in your code in the first place? > > I don't. The whole point of this query is about how to vet user input > and I did explain that I use the user input to locate a resource. But ISTM that you are using and sanitizing the wrong kind of user input. -- PointedEars Zend Certified PHP Engineer <http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2 Please do not cc me. / Bitte keine Kopien per E-Mail.
[toc] | [prev] | [next] | [standalone]
| From | Arno Welzel <usenet@arnowelzel.de> |
|---|---|
| Date | 2016-02-18 08:05 +0100 |
| Message-ID | <56C56D36.2060805@arnowelzel.de> |
| In reply to | #16477 |
Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33: > James Harris wrote: > >> On 17/02/2016 10:05, Arno Welzel wrote: >>> James Harris schrieb am 2016-02-17 um 10:37: [...] >>>> In addition to the above, could the URL be in Unicode form? I was >>>> thinking to change to index through it with >>> >>> See RFC 1808: >>> >>> <https://tools.ietf.org/html/rfc1808> >> >> I looked through it but could not see anything about Unicode. > > It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains > something about Unicode, in particular it defines UTF-8 percent-encoding. That's the problem with RFCs - nothing in the old RFC indicates the replacement ;-) But thanks for the hint. -- Arno Welzel http://arnowelzel.de http://de-rec-fahrrad.de http://fahrradzukunft.de
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2016-02-18 09:59 +0000 |
| Message-ID | <180220160959369722%timstreater@greenbee.net> |
| In reply to | #16490 |
In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel <usenet@arnowelzel.de> wrote: >Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33: >> James Harris wrote: >> >>> On 17/02/2016 10:05, Arno Welzel wrote: >>>> James Harris schrieb am 2016-02-17 um 10:37: >[...] >>>>> In addition to the above, could the URL be in Unicode form? I was >>>>> thinking to change to index through it with >>>> >>>> See RFC 1808: >>>> >>>> <https://tools.ietf.org/html/rfc1808> >>> >>> I looked through it but could not see anything about Unicode. >> >> It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains >> something about Unicode, in particular it defines UTF-8 percent-encoding. > >That's the problem with RFCs - nothing in the old RFC indicates the >replacement ;-) Yes it does. Look at the top of the old RFC (assuming you didn't download it on stone tablets in AD 54) and it will have an "Obsoleted by:" line indicating which RFC replaces it. -- "... you must remember that if you're trying to propagate a creed of poverty, gentleness and tolerance, you need a very rich, powerful, authoritarian organisation to do it." - Vice-Pope Eric
[toc] | [prev] | [next] | [standalone]
| From | James Harris <james.harris.1@gmail.com> |
|---|---|
| Date | 2016-02-18 10:29 +0000 |
| Message-ID | <na468j$b44$1@dont-email.me> |
| In reply to | #16494 |
On 18/02/2016 09:59, Tim Streater wrote: > In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel > <usenet@arnowelzel.de> wrote: ... >> That's the problem with RFCs - nothing in the old RFC indicates the >> replacement ;-) > > Yes it does. Look at the top of the old RFC (assuming you didn't > download it on stone tablets in AD 54) and it will have an "Obsoleted > by:" line indicating which RFC replaces it. AIUI, Arno is right in that RFCs may not be modified once published. I think you mean that the "obsoleted by" header appears in certain presentations of an RFC - e.g. an HTML presentation. Because of the immutable nature of RFCs, such a line cannot be added to the original. Compare these two. https://www.ietf.org/rfc/rfc1808.txt https://tools.ietf.org/html/rfc1808 James
[toc] | [prev] | [next] | [standalone]
| From | Tim Streater <timstreater@greenbee.net> |
|---|---|
| Date | 2016-02-18 11:36 +0000 |
| Message-ID | <180220161136429284%timstreater@greenbee.net> |
| In reply to | #16495 |
In article <na468j$b44$1@dont-email.me>, James Harris <james.harris.1@gmail.com> wrote: >On 18/02/2016 09:59, Tim Streater wrote: >> In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel >> <usenet@arnowelzel.de> wrote: > >... > >>> That's the problem with RFCs - nothing in the old RFC indicates the >>> replacement ;-) >> >> Yes it does. Look at the top of the old RFC (assuming you didn't >> download it on stone tablets in AD 54) and it will have an "Obsoleted >> by:" line indicating which RFC replaces it. > >AIUI, Arno is right in that RFCs may not be modified once published. I >think you mean that the "obsoleted by" header appears in certain >presentations of an RFC - e.g. an HTML presentation. > >Because of the immutable nature of RFCs, such a line cannot be added to >the original. Compare these two. > > https://www.ietf.org/rfc/rfc1808.txt > https://tools.ietf.org/html/rfc1808 Well indeed but why look at the text version of an RFC? Not just for the reason above but also the html versions allow easier navigation within the document. Thass why I have the tools site you cite here on my favourites list. Just a click away. -- Lady Astor: "If you were my husband I'd give you poison." Churchill: "If you were my wife, I'd drink it."
[toc] | [prev] | [next] | [standalone]
| From | Arno Welzel <usenet@arnowelzel.de> |
|---|---|
| Date | 2016-02-18 15:44 +0100 |
| Message-ID | <56C5D8C0.7030703@arnowelzel.de> |
| In reply to | #16494 |
Tim Streater schrieb am 2016-02-18 um 10:59: > In article <56C56D36.2060805@arnowelzel.de>, Arno Welzel > <usenet@arnowelzel.de> wrote: > >> Thomas 'PointedEars' Lahn schrieb am 2016-02-17 um 23:33: >>> James Harris wrote: >>> >>>> On 17/02/2016 10:05, Arno Welzel wrote: >>>>> James Harris schrieb am 2016-02-17 um 10:37: >> [...] >>>>>> In addition to the above, could the URL be in Unicode form? I was >>>>>> thinking to change to index through it with >>>>> >>>>> See RFC 1808: >>>>> >>>>> <https://tools.ietf.org/html/rfc1808> >>>> >>>> I looked through it but could not see anything about Unicode. >>> >>> It has been obsoleted by RFC 3986 eleven(!) years ago now, which contains >>> something about Unicode, in particular it defines UTF-8 percent-encoding. >> >> That's the problem with RFCs - nothing in the old RFC indicates the >> replacement ;-) > > Yes it does. Look at the top of the old RFC (assuming you didn't > download it on stone tablets in AD 54) and it will have an "Obsoleted > by:" line indicating which RFC replaces it. I stand corrected - of course there is an indication. -- Arno Welzel http://arnowelzel.de http://de-rec-fahrrad.de http://fahrradzukunft.de
[toc] | [prev] | [next] | [standalone]
Page 1 of 5 [1] 2 3 4 5 Next page →
Back to top | Article view | comp.lang.php
csiph-web