Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.php > #16475
| Path | csiph.com!news.mixmin.net!weretis.net!feeder1.news.weretis.net!news.solani.org!.POSTED!not-for-mail |
|---|---|
| From | Thomas 'PointedEars' Lahn <PointedEars@web.de> |
| Newsgroups | comp.lang.php |
| Subject | Re: Functions behaving badly: PHP str_split on an empty string |
| Date | Wed, 17 Feb 2016 23:12:54 +0100 |
| Organization | PointedEars Software (PES) |
| Lines | 99 |
| Message-ID | <2399331.b35jDvAURT@PointedEars.de> (permalink) |
| References | <na1g0b$va5$1@dont-email.me> <na1jeb$8hm$1@solani.org> <na25ho$q84$1@dont-email.me> |
| Reply-To | Thomas 'PointedEars' Lahn <php@PointedEars.de> |
| Mime-Version | 1.0 |
| Content-Type | text/plain; charset="UTF-8" |
| Content-Transfer-Encoding | 8Bit |
| X-Trace | solani.org 1455747176 2130 eJwVwoERADEEBMCWCI6Ug6f/EjI/uyZgtCsMavubJCPXU7BTH9/Vkpb2QKzNLfKdSBkEZ+gDG0kREA== (17 Feb 2016 22:12:56 GMT) |
| X-Complaints-To | abuse@news.solani.org |
| NNTP-Posting-Date | Wed, 17 Feb 2016 22:12:56 +0000 (UTC) |
| User-Agent | KNode/4.14.2 |
| X-User-ID | eJwFwQEBBDEIAzBLdNDyL4cx8C/hErqgzhAVXC6Mdmsj77PqyEp//7Ge8Sa014EEVJMs0s/WPmltzrH64QNemBWe |
| Cancel-Lock | sha1:3RM94A8RyyKCXnyBrey6jAYzx3A= |
| X-NNTP-Posting-Host | eJwNycEBwCAIA8CVCpig4xQk+4/Q3vcQNHYugguC2LjnMZm8iv+93RPKDRWyPLxt68yKrLncHyT1EW4= |
| Xref | csiph.com comp.lang.php:16475 |
Show key headers only | View raw
James Harris wrote:
> On 17/02/2016 10:53, Christoph M. Becker wrote:
>> James Harris wrote:
>>> Is there a recommended way to iterate over the characters of a string
>>> when that string could be single-byte ASCII or multi-byte Unicode?
It is a misconception to assume that “Unicode” would be equivalent to
“multi-byte” to begin with.
>> You can use mb_split() or preg_split() for multibyte character
>> encodings. Anyway, you should be aware in which character encoding the
>> string is encoded (mb_detect_encoding() might be helpful).
>
> That was useful in making it possible for me to check the types of
> strings. Given:
>
> $request_uri = $_SERVER["REQUEST_URI"];
> $uri = urldecode($request_uri)
>
> On this server (Apache) $request_uri is ASCII,
No, it is not. It is encoded with ISO-8859-1/Windows-1252, but since it
only contains characters that can be encoded with US-ASCII, "ASCII" is
reported. US-ASCII is a *7*-bit encoding. Keep in mind that the character
encoding of an arbitrary string can only be guessed, not ascertained.
> even if the URL had some Chinese characters in it.
URIs *always* use only ASCII characters (characters that can be encoded
using US-ASCII. Non-ASCII characters have to be UTF-8 percent-encoded,
which is what *raw*urldecode() decodes. See RFC 3986 and the PHP manual.
> As for $uri, if the input is just ASCII characters then $uri (the output
> of urldecode) is plain ASCII. But if the input contains Chinese
> characters
As I indicated, the input can only contain Chinese characters if they are
percent-encoded.
> then $uri will be Unicode.
Unicode is not a character encoding, but a standard for a character set and
several encodings for it. The character encoding of the return value should
be UTF-8 then, where a character with a code point below U+0080 is encoded
with *one* code unit – *one* 8-bit byte (backwards-compatible to US-ASCII).
<http://unicode.org/faq/>
> I presume my choices are essentially either to report non-ASCII
> characters as errors or to do all processing in Unicode.
Non sequitur.
>> However, I suggest you reconsider the approach to iterate over the
>> characters of a string, because it's most likely rather slow in PHP (for
>> one, PHP doesn't have a char type as C does, so a character is stored as
>> string[2]), and there may be better alternatives available (PCRE comes
>> to mind).
>
> There may be a better way to do it. Here is some of my current code
>
> $request_uri = trim(urldecode($_SERVER["REQUEST_URI"]), "/");
> foreach (str_split($request_uri) as $char) {
> if (strpos($path_chars_permitted, $char) === false) {
> $etext .= " Invalid character in URL: \"" .
> htmlspecialchars($char) . "\"<br/>\n";
> }
> }
>
> Suggestions for improvement welcome.
if (preg_match_all('/[^\\x00-\\x7f]/u', $request_uri, $matches))
{
$etext = implode("<br>\n",
array_map(
function ($v) {
return 'Invalid character in URL: "' . htmlspecialchars($v) . '"';
},
$matches[0]
)
);
}
But AISB, if a request URI contains unescaped non-ASCII characters, the UA
making the HTTP request is borked. If it contains percent-encoded
characters, and your PHP program cannot handle that, your detection
algorithm is wrong, and your program is borked from the outset.
> Maybe there is a function to do the lot in one go that would save me
> iterating over the string.
There is.
--
PointedEars
Zend Certified PHP Engineer
<http://www.zend.com/en/yellow-pages/ZEND024953> | Twitter: @PointedEars2
Please do not cc me. / Bitte keine Kopien per E-Mail.
Back to comp.lang.php | Previous | Next — Previous in thread | Find similar | Unroll thread
Functions behaving badly: PHP str_split on an empty string James Harris <james.harris.1@gmail.com> - 2016-02-17 09:57 +0000
Re: Functions behaving badly: PHP str_split on an empty string "Christoph M. Becker" <cmbecker69@arcor.de> - 2016-02-17 11:53 +0100
Re: Functions behaving badly: PHP str_split on an empty string James Harris <james.harris.1@gmail.com> - 2016-02-17 16:04 +0000
Re: Functions behaving badly: PHP str_split on an empty string Thomas 'PointedEars' Lahn <PointedEars@web.de> - 2016-02-17 23:12 +0100
csiph-web