Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #80571 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2021-06-30 07:57 +0000 |
| Last post | 2021-07-01 15:44 +0200 |
| Articles | 20 on this page of 70 — 19 participants |
Back to article view | Back to comp.lang.c++
How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 07:57 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 10:05 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:30 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 12:37 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 09:19 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 11:10 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 11:36 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 16:17 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 10:13 -0700
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 19:27 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 11:21 -0700
Re: How to write wide char string literals? Real Troll <real.troll@trolls.com> - 2021-07-01 18:45 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-02 10:03 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-02 02:31 -0700
Re: How to write wide char string literals? MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk - 2021-07-02 09:38 +0000
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 13:52 +0300
Re: How to write wide char string literals? MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net - 2021-07-02 12:53 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:22 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 17:32 +0300
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:08 +0200
Re: How to write wide char string literals? MrSpud_r5j@ywn9entw2s.org - 2021-06-30 08:39 +0000
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 10:52 +0000
Re: How to write wide char string literals? MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov - 2021-06-30 11:28 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 14:01 +0200
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 10:55 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 11:23 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 06:51 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:09 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:35 -0400
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 14:03 +0300
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:13 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:39 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 13:50 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-06-30 14:19 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-01 13:31 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 04:42 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-01 10:58 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 05:26 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 01:44 -0400
Re: How to write wide char string literals? Bo Persson <bo@bo-persson.se> - 2021-07-02 08:33 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 11:52 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 16:15 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 03:30 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 00:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 13:31 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 08:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 16:28 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 11:48 -0400
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 15:56 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 16:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 17:49 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:11 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 05:52 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 06:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 09:03 -0400
Re: How to write wide char string literals? Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2021-07-03 14:23 +0100
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 17:06 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 14:06 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:13 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:37 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-06-30 16:54 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 17:56 +0300
Re: How to write wide char string literals? Öö Tiib <ootiib@hot.ee> - 2021-06-30 09:22 -0700
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 07:19 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 10:29 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 08:44 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 12:58 +0200
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 14:01 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 14:57 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-07-01 15:44 +0200
Page 3 of 4 — ← Prev page 1 2 [3] 4 Next page →
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-02 11:52 +0000 |
| Message-ID | <sbmul5$1vu5$1@gioia.aioe.org> |
| In reply to | #80617 |
James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
> so you should be able to freely use u prefixed
> string literals and char16_t with code that needs to be portable to
> Windows, but can also be used on other platforms.
But doesn't that have the exact same problem as in my original post?
In other words, in code like this:
const char16_t *str = u"something";
the stuff between the quotation marks in the source code will use
whatever encoding (ostensibly but not assuredly UTF-8), which the
compiler needs to convert to UTF-16 at compile time for writing
it to the output binary.
Or is it guaranteed that the characters between u" and " will
always be interpreted as UTF-8?
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-02 16:15 -0400 |
| Message-ID | <sbns4s$cfd$1@dont-email.me> |
| In reply to | #80623 |
On 7/2/21 7:52 AM, Juha Nieminen wrote:
> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>> so you should be able to freely use u prefixed
>> string literals and char16_t with code that needs to be portable to
>> Windows, but can also be used on other platforms.
>
> But doesn't that have the exact same problem as in my original post?
No, because the original post used wchar_t and the L prefix, for which
the relevant encoding is implementation-defined. That's not the case for
char and the u8 prefix, char16_t and the u prefix, or for char32_t and
the U prefix.
I have a feeling that there's a misunderstanding somewhere in this
conversation, but I'm not sure yet what it is.
> In other words, in code like this:
>
> const char16_t *str = u"something";
>
> the stuff between the quotation marks in the source code will use
> whatever encoding (ostensibly but not assuredly UTF-8), which the
I don't understand why you think that the source code encoding matters.
The only thing that matters is what characters are encoded. As long as
those encodings are for the characters {'u', '""', '\\', 'x', 'C', '2',
'\\', 'x', 'A', '9', '"'}, any fully conforming implementation must give
you the standard defined behavior for u"\xC2\xA9".
> compiler needs to convert to UTF-16 at compile time for writing
> it to the output binary.
When using the u8 prefix, UTF-8 encoding is guaranteed, for which every
codepoint from U+0000 to U+007F is represented by a single character
with a numerical value matching the code point.
When using the u prefix, UTF-16 encoding is guaranteed, for which every
codepoint from U+0000 to U+D7FF, and from U+E000 to U+FFFF, is
represented by a single character with a numerical value matching the
codepoint.
When using the U prefix, UTF-32 encoding is guaranteed, for which every
codepoint from U+0000 to U+D7FF, and from U+E000 to U+10FFFF, is
represented by a single character with a numerical value matching the
codepoint.
Since the meanings of the octal and hexadecimal escape sequences are
defined in terms of the numerical values of the corresponding
characters, if you use any of those prefixes, specifying the value of a
character within the specified range by using an octal or hexadecimal
escape sequence is precisely as portable as using the UCN with the same
numerical value. Using UCNs would be better because they are less
restricted, working just as well with the L prefix and with no prefix.
However, within those ranges, octal and hexadecimal escapes will work
just as well.
Do you know of any implementation of C++ that claims to be fully
conforming, for which that is not the case? If so, how do they justify
that claim?
'\xC2' and '\xA9' are in the ranges for the u and U prefixes. They would
not be in the range for the u8 prefix, but since the context was wide
characters, the u8 prefix is not relevant.
> Or is it guaranteed that the characters between u" and " will
> always be interpreted as UTF-8?
Source code characters between the u" and the " will be interpreted
according to an implementation-defined character encoding. But so long
as they encode {'\\', 'x', 'C', '2', '\\', 'x', 'A', '9'}, you should
get the standard-defined behavior for u"\xC2\xA9".
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-03 03:30 +0200 |
| Message-ID | <sboej7$scd$1@dont-email.me> |
| In reply to | #80628 |
On 2 Jul 2021 22:15, James Kuyper wrote:
> On 7/2/21 7:52 AM, Juha Nieminen wrote:
>> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>>> so you should be able to freely use u prefixed
>>> string literals and char16_t with code that needs to be portable to
>>> Windows, but can also be used on other platforms.
>>
>> But doesn't that have the exact same problem as in my original post?
> No, because the original post used wchar_t and the L prefix, for which
> the relevant encoding is implementation-defined. That's not the case for
> char and the u8 prefix, char16_t and the u prefix, or for char32_t and
> the U prefix.
>
> I have a feeling that there's a misunderstanding somewhere in this
> conversation, but I'm not sure yet what it is.
Juha is concerned about the compiler assuming some other source code
encoding than the actual one.
The implementation defined encoding of `wchar_t`, where in practice the
possibilities as of 2021 are either UTF-16 or UTF-32, doesn't matter.
A correct source code encoding assumption can be guaranteed by simply
statically asserting that the basic execution character set is UTF-8, as
I showed in my original answer in this thread.
>> In other words, in code like this:
>>
>> const char16_t *str = u"something";
>>
>> the stuff between the quotation marks in the source code will use
>> whatever encoding (ostensibly but not assuredly UTF-8), which the
>
> I don't understand why you think that the source code encoding matters.
It matters because if the compiler assumes wrong, and Visual C++
defaults to assuming Windows ANSI when no other indication is present
and it's not forced by options, then one gets incorrect literals.
Which may or may not be caught by unit testing.
> The only thing that matters is what characters are encoded. As long as
> those encodings are for the characters {'u', '""', '\\', 'x', 'C', '2',
> '\\', 'x', 'A', '9', '"'}, any fully conforming implementation must give
> you the standard defined behavior for u"\xC2\xA9".
Consider:
#include <iostream>
using std::cout, std::hex, std::endl;
auto main() -> int
{
const char16_t s16[] = u"\xC2\xA9";
for( const int code: s16 ) {
if( code ) { cout << hex << code << " "; }
}
cout << endl;
}
The output of this program, i.e. the UTF-16 encoding values in `s16`, is
c2 a9
Since Unicode is an extension of Latin-1 the UTF-16 interpretation of
`\xC2` and `xA9` is as Latin-1 characters, respectively "Â" and (not a
coincidence) "©" according to my Windows 10 console in codepage 1252.
Which is not the single "©" that an UTF-8 interpretation gives.
[snip]
- Alf
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-03 00:44 -0400 |
| Message-ID | <sbopvo$qt9$1@dont-email.me> |
| In reply to | #80629 |
On 7/2/21 9:30 PM, Alf P. Steinbach wrote:
> On 2 Jul 2021 22:15, James Kuyper wrote:
>> On 7/2/21 7:52 AM, Juha Nieminen wrote:
>>> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>>>> so you should be able to freely use u prefixed
>>>> string literals and char16_t with code that needs to be portable to
>>>> Windows, but can also be used on other platforms.
>>>
>>> But doesn't that have the exact same problem as in my original post?
>> No, because the original post used wchar_t and the L prefix, for which
>> the relevant encoding is implementation-defined. That's not the case for
>> char and the u8 prefix, char16_t and the u prefix, or for char32_t and
>> the U prefix.
>>
>> I have a feeling that there's a misunderstanding somewhere in this
>> conversation, but I'm not sure yet what it is.
I now have a much better idea what the misunderstanding is. See below.
> Juha is concerned about the compiler assuming some other source code
> encoding than the actual one.
>
> The implementation defined encoding of `wchar_t`, where in practice the
> possibilities as of 2021 are either UTF-16 or UTF-32, doesn't matter.
>
> A correct source code encoding assumption can be guaranteed by simply
> statically asserting that the basic execution character set is UTF-8, as
> I showed in my original answer in this thread.
The encoding of the basic execution character set is irrelevant if the
string literals are prefixed with u8, u, or U, and use only valid escape
sequences to specify members of the extended character set. The encoding
for such literals is explicitly mandated by the standard. Are you (or
he) worrying about a failure to conform to those mandates?
...
>> I don't understand why you think that the source code encoding matters.
>
> It matters because if the compiler assumes wrong, and Visual C++
> defaults to assuming Windows ANSI when no other indication is present
> and it's not forced by options, then one gets incorrect literals.
Even when u8, u or U prefixes are specified?
...
>> The only thing that matters is what characters are encoded. As long as
>> those encodings are for the characters {'u', '""', '\\', 'x', 'C', '2',
>> '\\', 'x', 'A', '9', '"'}, any fully conforming implementation must give
>> you the standard defined behavior for u"\xC2\xA9".
>
> Consider:
>
> #include <iostream>
> using std::cout, std::hex, std::endl;
>
> auto main() -> int
> {
> const char16_t s16[] = u"\xC2\xA9";
> for( const int code: s16 ) {
> if( code ) { cout << hex << code << " "; }
> }
> cout << endl;
> }
>
> The output of this program, i.e. the UTF-16 encoding values in `s16`, is
>
> c2 a9
Yes, that's precisely what the C++ standard mandates, regardless of the
encoding of the source character set. Which is why I mistakenly thought
that's what he was trying to do.
> Since Unicode is an extension of Latin-1 the UTF-16 interpretation of
> `\xC2` and `xA9` is as Latin-1 characters, respectively "Â" and (not a
> coincidence) "©" according to my Windows 10 console in codepage 1252.
>
> Which is not the single "©" that an UTF-8 interpretation gives.
OK - it had not occurred to me that he was trying to encode "©", since
that is not the right way to do so. In a sense, I suppose that's the
point you're making.
My point is that all ten of the following escape sequences should be
perfectly portable ways of specifying that same code point in each of
three Unicode encodings:
UTF-8: u8"\u00A9\U000000A9"
UTF-16: u"\251\xA9\u00A9\U000000A9"
UTF-32: U"\251\xA9\u00A9\U000000A9"
Do you know of any implementation which is non-conforming because it
misinterprets any of those escape sequences?
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-03 13:31 +0200 |
| Message-ID | <sbphpm$vh9$1@dont-email.me> |
| In reply to | #80630 |
On 3 Jul 2021 06:44, James Kuyper wrote:
> On 7/2/21 9:30 PM, Alf P. Steinbach wrote:
>> On 2 Jul 2021 22:15, James Kuyper wrote:
>>> On 7/2/21 7:52 AM, Juha Nieminen wrote:
>>>> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>>>>> so you should be able to freely use u prefixed
>>>>> string literals and char16_t with code that needs to be portable to
>>>>> Windows, but can also be used on other platforms.
>>>>
>>>> But doesn't that have the exact same problem as in my original post?
>>> No, because the original post used wchar_t and the L prefix, for which
>>> the relevant encoding is implementation-defined. That's not the case for
>>> char and the u8 prefix, char16_t and the u prefix, or for char32_t and
>>> the U prefix.
>>>
>>> I have a feeling that there's a misunderstanding somewhere in this
>>> conversation, but I'm not sure yet what it is.
>
> I now have a much better idea what the misunderstanding is. See below.
>> Juha is concerned about the compiler assuming some other source code
>> encoding than the actual one.
>>
>> The implementation defined encoding of `wchar_t`, where in practice the
>> possibilities as of 2021 are either UTF-16 or UTF-32, doesn't matter.
>>
>> A correct source code encoding assumption can be guaranteed by simply
>> statically asserting that the basic execution character set is UTF-8, as
>> I showed in my original answer in this thread.
>
> The encoding of the basic execution character set is irrelevant if the
> string literals are prefixed with u8, u, or U, and use only valid escape
> sequences to specify members of the extended character set.
Having a handy simple way to guarantee a correct source code encoding
assumption doesn't seem irrelevant to me.
On the contrary it's directly a solution to the OP's problem, which to
me appears to be maximally relevant.
> The encoding
> for such literals is explicitly mandated by the standard. Are you (or
> he) worrying about a failure to conform to those mandates?
No, Juha is worrying about the compiler's source code encoding assumption.
> ...
>>> I don't understand why you think that the source code encoding matters.
>>
>> It matters because if the compiler assumes wrong, and Visual C++
>> defaults to assuming Windows ANSI when no other indication is present
>> and it's not forced by options, then one gets incorrect literals.
>
> Even when u8, u or U prefixes are specified?
Yes. As an example, consider
const auto& s = u"Blåbær, Mr. Watson.";
If the source is UTF-8 encoded, without a BOM or other encoding marker,
and if the Visual C++ compiler is not told to assume UTF-8 source code,
then it will incorrectly assume that this is Windows ANSI encoded.
The UTF-8 bytes in the source code will then be interpreted as Windows
ANSI character codes, e.g. as Windows ANSI Western, codepage 1252.
The compiler will then see this source code:
const auto& s = u"Blåbær, Mr. Watson.";
And it will proceeed to encode /that/ string with UTF-8 in the resulting
string value.
> ...
>>> The only thing that matters is what characters are encoded. As long as
>>> those encodings are for the characters {'u', '""', '\\', 'x', 'C', '2',
>>> '\\', 'x', 'A', '9', '"'}, any fully conforming implementation must give
>>> you the standard defined behavior for u"\xC2\xA9".
>>
>> Consider:
>>
>> #include <iostream>
>> using std::cout, std::hex, std::endl;
>>
>> auto main() -> int
>> {
>> const char16_t s16[] = u"\xC2\xA9";
>> for( const int code: s16 ) {
>> if( code ) { cout << hex << code << " "; }
>> }
>> cout << endl;
>> }
>>
>> The output of this program, i.e. the UTF-16 encoding values in `s16`, is
>>
>> c2 a9
>
>
> Yes, that's precisely what the C++ standard mandates, regardless of the
> encoding of the source character set. Which is why I mistakenly thought
> that's what he was trying to do.
>
>> Since Unicode is an extension of Latin-1 the UTF-16 interpretation of
>> `\xC2` and `xA9` is as Latin-1 characters, respectively "Â" and (not a
>> coincidence) "©" according to my Windows 10 console in codepage 1252.
>>
>> Which is not the single "©" that an UTF-8 interpretation gives.
>
> OK - it had not occurred to me that he was trying to encode "©", since
> that is not the right way to do so. In a sense, I suppose that's the
> point you're making.
> My point is that all ten of the following escape sequences should be
> perfectly portable ways of specifying that same code point in each of
> three Unicode encodings:
>
> UTF-8: u8"\u00A9\U000000A9"
> UTF-16: u"\251\xA9\u00A9\U000000A9"
> UTF-32: U"\251\xA9\u00A9\U000000A9"
>
> Do you know of any implementation which is non-conforming because it
> misinterprets any of those escape sequences?
No, they should work. These escapes are an alternative solution to
Juha's problem. However, they lack readability and involve much more
work than necessary, so IMO the thing to do is to assert UTF-8.
- Alf
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-03 08:44 -0400 |
| Message-ID | <sbpm41$thv$1@dont-email.me> |
| In reply to | #80632 |
On 7/3/21 7:31 AM, Alf P. Steinbach wrote: > On 3 Jul 2021 06:44, James Kuyper wrote: >> On 7/2/21 9:30 PM, Alf P. Steinbach wrote: ... >>> It matters because if the compiler assumes wrong, and Visual C++ >>> defaults to assuming Windows ANSI when no other indication is present >>> and it's not forced by options, then one gets incorrect literals. >> >> Even when u8, u or U prefixes are specified? > > Yes. As an example, consider > > const auto& s = u"Blåbær, Mr. Watson."; The comment that led to this sub-thread was specifically about the usability of escape sequences to specify members of the extended character set, and that's the only thing I was talking about. While that string does contain such members, it contains not a single escape sequence. ... >> OK - it had not occurred to me that he was trying to encode "©", since >> that is not the right way to do so. In a sense, I suppose that's the >> point you're making. >> My point is that all ten of the following escape sequences should be >> perfectly portable ways of specifying that same code point in each of >> three Unicode encodings: >> >> UTF-8: u8"\u00A9\U000000A9" >> UTF-16: u"\251\xA9\u00A9\U000000A9" >> UTF-32: U"\251\xA9\u00A9\U000000A9" >> >> Do you know of any implementation which is non-conforming because it >> misinterprets any of those escape sequences? > No, they should work. These escapes are an alternative solution to > Juha's problem. ... They are the only solution that this sub-thread has been about. > ... However, they lack readability and involve much more > work than necessary, so IMO the thing to do is to assert UTF-8. Those are reasonable concerns. That the system's assumptions about the source character set would prevent those escapes from working is not.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-03 16:28 +0200 |
| Message-ID | <sbps5k$b5m$1@dont-email.me> |
| In reply to | #80633 |
On 3 Jul 2021 14:44, James Kuyper wrote: > On 7/3/21 7:31 AM, Alf P. Steinbach wrote: >> On 3 Jul 2021 06:44, James Kuyper wrote: [snippety] >>> >>> UTF-8: u8"\u00A9\U000000A9" >>> UTF-16: u"\251\xA9\u00A9\U000000A9" >>> UTF-32: U"\251\xA9\u00A9\U000000A9" >>> >>> Do you know of any implementation which is non-conforming because it >>> misinterprets any of those escape sequences? >> No, they should work. These escapes are an alternative solution to >> Juha's problem. ... > > They are the only solution that this sub-thread has been about. > >> ... However, they lack readability and involve much more >> work than necessary, so IMO the thing to do is to assert UTF-8. > > Those are reasonable concerns. That the system's assumptions about the > source character set would prevent those escapes from working is not. As far as I know nobody's argued that the source encoding assumption would prevent any escapes from working. If I understand you correctly your “this sub-thread” about escapes and universal character designators -- let's just call them all escapes -- started when you responded to my response Richard Daemon, who had responded to Juha Niemininen, who wrote: [>>] Does that work for wide string literals? Because I don't think it does. In other words: std::wstring s = L"Copyright \xC2\xA9 2001-2020"; [<<] Richard responded to that: [>>] \x works in wide string literal too, and puts in a character with that value. The difference is that if the wide string type isn't unicode encoded then it might get the wrong character in the string. [<<] I responded to Richard: [>>] It gets the wrong characters in the wide string literal, period. [<<] Which it decidedly does. It's trivial to just try it out and see; QED. You responded to that where you snipped Juha's example, indicating some misunderstanding on your part: [>>] The value of a wide character is determined by the current encoding. For wide character literals using the u or U prefixes, that encoding is UTF-16 and UTF-32, respectively, making octal escapes redundant with and less convenient than the use of UCNs. But as he said, they do work for such strings. [<<] So, in your mind this sub-thread may have been about whether escape sequences (including universal character designators) are affected by the source encoding, but to me it has been about whether Juha's example yields the desired string, as he correctly surmised that it didn't. And the outer context from the top thread, is about the source encoding's effect on string literals, which hopefully is now clear. - Alf
[toc] | [prev] | [next] | [standalone]
| From | Richard Damon <Richard@Damon-Family.org> |
|---|---|
| Date | 2021-07-03 11:48 -0400 |
| Message-ID | <YE%DI.10994$835.6256@fx36.iad> |
| In reply to | #80640 |
On 7/3/21 10:28 AM, Alf P. Steinbach wrote: > On 3 Jul 2021 14:44, James Kuyper wrote: >> On 7/3/21 7:31 AM, Alf P. Steinbach wrote: >>> On 3 Jul 2021 06:44, James Kuyper wrote: > [snippety] >>>> >>>> UTF-8: u8"\u00A9\U000000A9" >>>> UTF-16: u"\251\xA9\u00A9\U000000A9" >>>> UTF-32: U"\251\xA9\u00A9\U000000A9" >>>> >>>> Do you know of any implementation which is non-conforming because it >>>> misinterprets any of those escape sequences? >>> No, they should work. These escapes are an alternative solution to >>> Juha's problem. ... >> >> They are the only solution that this sub-thread has been about. >> >>> ... However, they lack readability and involve much more >>> work than necessary, so IMO the thing to do is to assert UTF-8. >> >> Those are reasonable concerns. That the system's assumptions about the >> source character set would prevent those escapes from working is not. > > As far as I know nobody's argued that the source encoding assumption > would prevent any escapes from working. > > If I understand you correctly your “this sub-thread” about escapes and > universal character designators -- let's just call them all escapes -- > started when you responded to my response Richard Daemon, who had > responded to Juha Niemininen, who wrote: > > > [>>] > Does that work for wide string literals? Because I don't think it does. > In other words: > > std::wstring s = L"Copyright \xC2\xA9 2001-2020"; > [<<] > > > Richard responded to that: > > > [>>] > \x works in wide string literal too, and puts in a character with that > value. The difference is that if the wide string type isn't unicode > encoded then it might get the wrong character in the string. > [<<] > > > I responded to Richard: > > > [>>] > It gets the wrong characters in the wide string literal, period. > [<<] > > > Which it decidedly does. It puts into the string exactly the characters that you specified, the character of value 0x00C2 and then the character of value 0x00A9. THAT is what it says to do. If you meant the \x to mean this is a UTF-8 encode string, why are you expecting that? The one issue with \x is it puts in the characters in whatever encoding wide strings use, so you can't just assume unicode values unless you are willing to assume wide string is unicode encoded. > It's trivial to just try it out and see; QED. > > You responded to that where you snipped Juha's example, indicating some > misunderstanding on your part: > > > [>>] > The value of a wide character is determined by the current encoding. For > wide character literals using the u or U prefixes, that encoding is > UTF-16 and UTF-32, respectively, making octal escapes redundant with and > less convenient than the use of UCNs. But as he said, they do work for > such strings. > [<<] > > > So, in your mind this sub-thread may have been about whether escape > sequences (including universal character designators) are affected by > the source encoding, but to me it has been about whether Juha's example > yields the desired string, as he correctly surmised that it didn't. > > And the outer context from the top thread, is about the source > encoding's effect on string literals, which hopefully is now clear. > > > - Alf
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-08 15:56 -0400 |
| Message-ID | <sc7l8m$aao$1@dont-email.me> |
| In reply to | #80640 |
On 7/3/21 10:28 AM, Alf P. Steinbach wrote: > On 3 Jul 2021 14:44, James Kuyper wrote: >> On 7/3/21 7:31 AM, Alf P. Steinbach wrote: >>> On 3 Jul 2021 06:44, James Kuyper wrote: > [snippety] >>>> >>>> UTF-8: u8"\u00A9\U000000A9" >>>> UTF-16: u"\251\xA9\u00A9\U000000A9" >>>> UTF-32: U"\251\xA9\u00A9\U000000A9" >>>> >>>> Do you know of any implementation which is non-conforming because it >>>> misinterprets any of those escape sequences? >>> No, they should work. These escapes are an alternative solution to >>> Juha's problem. ... >> >> They are the only solution that this sub-thread has been about. >> >>> ... However, they lack readability and involve much more >>> work than necessary, so IMO the thing to do is to assert UTF-8. >> >> Those are reasonable concerns. That the system's assumptions about the >> source character set would prevent those escapes from working is not. > > As far as I know nobody's argued that the source encoding assumption > would prevent any escapes from working. You said "It gets the wrong characters in the wide string literal, period.", and other parts of the discussion implicated source encoding assumptions as the reason why. The use of "period" implies no exceptions, and there's a very large set of exceptions: at least two, as as many as four, fully portable working escape sequences for every single Unicode code point. > If I understand you correctly your “this sub-thread” about escapes and > universal character designators -- let's just call them all escapes > -- started when you responded to my response Richard Daemon, who had > responded to Juha Niemininen, who wrote: > > > [>>] > Does that work for wide string literals? Because I don't think it does. > In other words: > > std::wstring s = L"Copyright \xC2\xA9 2001-2020"; > [<<] > > > Richard responded to that: > > > [>>] > \x works in wide string literal too, and puts in a character with that > value. The difference is that if the wide string type isn't unicode > encoded then it might get the wrong character in the string. > [<<] > > > I responded to Richard: > > > [>>] > It gets the wrong characters in the wide string literal, period. > [<<] > > > Which it decidedly does. > > It's trivial to just try it out and see; QED. I did try it: as he said, it can get the wrong character if the string type isn't unicode encoded, and as I pointed out, it can also get the wrong character if the wrong escape sequence is used (which seems trivially obvious). But it's perfectly capable of giving the right characters when the right escape sequence is used with a prefix that mandates a unicode encoding. By saying "... it gets the wrong characters ... period.", you were denying that it's ever possible for it to get the right characters, which is demonstrably false. I've tried out the sequences I specified in the message you quoted above. They all work on my systems, and according to my understanding of the standard, they're required to work on all fully conforming implementations, regardless of source encoding assumptions - if that's not the case, I want to know how the exceptions can be justified. ... > So, in your mind this sub-thread may have been about whether escape > sequences (including universal character designators) are affected by > the source encoding, but to me it has been about whether Juha's example > yields the desired string, as he correctly surmised that it didn't. Yes, but that's because it was the wrong escape sequence, not because there's any inherent problem with using correct escape sequences for that purpose.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-03 16:59 +0000 |
| Message-ID | <sbq522$2v0$1@gioia.aioe.org> |
| In reply to | #80633 |
James Kuyper <jameskuyper@alumni.caltech.edu> wrote: > The comment that led to this sub-thread was specifically about the > usability of escape sequences to specify members of the extended > character set, and that's the only thing I was talking about. While that > string does contain such members, it contains not a single escape sequence. The problem is that the "\xC2\xA9" was presented as a solution to the compiler wrongly assuming some source file encoding other than UTF-8. Those two bytes are the UTF-8 encoding of a non-ascii character. In other words, it's explicitly entering the UTF-8 encoding of that non-ascii character. This works if we are specifying a narrow string literal (and we want it to be UTF-8 encoded). My point is that it doesn't work for a wide string literal. If you say L"\xC2\xA9" you will *not* get that non-ascii character you wanted. Instead, you get two UTF-16 (or UTF-32, depending on how large wchar_t is) characters which are completely different from the one you wanted. You essentially get garbage.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-03 17:49 -0400 |
| Message-ID | <sbqm1r$4qr$1@dont-email.me> |
| In reply to | #80649 |
On 7/3/21 12:59 PM, Juha Nieminen wrote: > James Kuyper <jameskuyper@alumni.caltech.edu> wrote: >> The comment that led to this sub-thread was specifically about the >> usability of escape sequences to specify members of the extended >> character set, and that's the only thing I was talking about. While that >> string does contain such members, it contains not a single escape sequence. > > The problem is that the "\xC2\xA9" was presented as a solution to > the compiler wrongly assuming some source file encoding other than UTF-8. > Those two bytes are the UTF-8 encoding of a non-ascii character. That sequence specifies a character with a value of 0xC2 followed by a character with a value of 0xA9. When the characters in question are wider than 8 bits, that is NOT the UTF-8 encoding of the character you want. Which just means you need to specify the right character. > In other words, it's explicitly entering the UTF-8 encoding of that > non-ascii character. This works if we are specifying a narrow string > literal (and we want it to be UTF-8 encoded). > > My point is that it doesn't work for a wide string literal. If you > say L"\xC2\xA9" you will *not* get that non-ascii character you > wanted. ... That's because you didn't specify what you wanted. You should have used \u00A9 rather than \xC2\xA9. > ... Instead, you get two UTF-16 (or UTF-32, depending on > how large wchar_t is) characters which are completely different > from the one you wanted. You essentially get garbage. You got precisely what you specified - if it's not what you wanted, you need to change your specification.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-08 08:11 +0000 |
| Message-ID | <sc6bvo$1bks$1@gioia.aioe.org> |
| In reply to | #80662 |
James Kuyper <jameskuyper@alumni.caltech.edu> wrote: >> ... Instead, you get two UTF-16 (or UTF-32, depending on >> how large wchar_t is) characters which are completely different >> from the one you wanted. You essentially get garbage. > > You got precisely what you specified - if it's not what you wanted, you > need to change your specification. No, I didn't. I wanted a way to specify wide string literals, and that solution was incorrect.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-08 05:52 -0400 |
| Message-ID | <sc6htf$end$1@dont-email.me> |
| In reply to | #80689 |
On 7/8/21 4:11 AM, Juha Nieminen wrote:
> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>>> ... Instead, you get two UTF-16 (or UTF-32, depending on
>>> how large wchar_t is) characters which are completely different
>>> from the one you wanted. You essentially get garbage.
>>
>> You got precisely what you specified - if it's not what you wanted, you
>> need to change your specification.
>
> No, I didn't. I wanted a way to specify wide string literals, and that
> solution was incorrect.
Paavo Helde's solution of using "\xC2\xA9" was correct for narrow string
literals (on systems with CHAR_BIT==8, a requirement that he didn't
bother mentioning). He was relying upon a UTF-8 => UTF-16 conversion
routine of his own creation to get the corresponding wide string.
You asked whether L"\xC2\xA9" would work, and the answer is "No",
because it specifies two wide characters when only one is desired. You
were aware that it wouldn't work, but seemed to be suggesting that
there's a potentially faulty UTF-8=>UTF-16 conversion involved in it's
failure to be correct. There is no such conversion. L"\xC2\xA9"
specifies directly a wchar_t array of length 3 initialized with {0xC2,
0xA9, 0}, which is not what you wanted.
I initially didn't address that point properly because I hadn't realized
that only one character was desired.
However, u"\xA9" or U"\xA9" would work fine; L"\xA9" should produce the
desired result on systems where wchar_t uses UCS2 or UCS4 (==UTF-32)
encoding.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-03 06:59 +0000 |
| Message-ID | <sbp1td$1sqn$1@gioia.aioe.org> |
| In reply to | #80628 |
James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>> In other words, in code like this:
>>
>> const char16_t *str = u"something";
>>
>> the stuff between the quotation marks in the source code will use
>> whatever encoding (ostensibly but not assuredly UTF-8), which the
>
> I don't understand why you think that the source code encoding matters.
Because the source file will most often be a text file using 8-bit
characters and, in these situations most likely (although not assuredly)
using UTF-8 encoding for non-ascii characters.
However, when you write:
const char16_t *str = u"something";
if that "something" contains non-ascii characters, which in this case will
be (usually) UTF-8 encoded in this source code, the compiler will have to
interpret that UTF-8 string and convert it to UTF-16 for the output
binary.
So the problem is the same as with wchar_t: How does the compiler know
which encoding is being used in this source file? It needs to know that
since it has to generate an UTF-16 string literal into the output binary
from those characters appearing in the source code.
> When using the u8 prefix, UTF-8 encoding is guaranteed, for which every
> codepoint from U+0000 to U+007F is represented by a single character
> with a numerical value matching the code point.
UTF-8 encoding is guaranteed *for the result*, ie. what the compiler writes
to the output binary. Is it guaranteed to *read* the characters in the
source code between the quotation marks and interpret it as UTF-8?
> When using the u prefix, UTF-16 encoding is guaranteed, for which every
> codepoint from U+0000 to U+D7FF, and from U+E000 to U+FFFF, is
> represented by a single character with a numerical value matching the
> codepoint.
Same issue, even more relevantly here.
> Do you know of any implementation of C++ that claims to be fully
> conforming, for which that is not the case? If so, how do they justify
> that claim?
Visual Studio will, by default (ie. with default project settings after
having created a new project) interpret the source files as Windows-1252
(which is very similar to ISO-Latin-1).
This means that when you write L"something" or u"something", if there
are any non-ascii characters between the parentheses, UTF-8 encoded,
then the result will be incorrect. (In order to make Visual Studio do
the correct conversion, you need to specify that the file is UTF-8 encoded
in the project settings).
>> Or is it guaranteed that the characters between u" and " will
>> always be interpreted as UTF-8?
>
> Source code characters between the u" and the " will be interpreted
> according to an implementation-defined character encoding. But so long
> as they encode {'\\', 'x', 'C', '2', '\\', 'x', 'A', '9'}, you should
> get the standard-defined behavior for u"\xC2\xA9".
Yes, but that's not the correct desired character in UTF-16, only in UTF-8.
You'll get garbage as your UTF-16 string literal.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-03 09:03 -0400 |
| Message-ID | <sbpn6b$5ut$1@dont-email.me> |
| In reply to | #80631 |
On 7/3/21 2:59 AM, Juha Nieminen wrote:
> James Kuyper <jameskuyper@alumni.caltech.edu> wrote:
>>> In other words, in code like this:
>>>
>>> const char16_t *str = u"something";
>>>
>>> the stuff between the quotation marks in the source code will use
>>> whatever encoding (ostensibly but not assuredly UTF-8), which the
>>
>> I don't understand why you think that the source code encoding matters.
>
> Because the source file will most often be a text file using 8-bit
> characters and, in these situations most likely (although not assuredly)
> using UTF-8 encoding for non-ascii characters.
Every comment I made on this sub-thread was prefaced on the absence of
any actual members of the extended character set - I was talking only
about the feasibility of using escape sequences to specify such members.
...
>> Do you know of any implementation of C++ that claims to be fully
>> conforming, for which that is not the case? If so, how do they justify
>> that claim?
>
> Visual Studio will, by default (ie. with default project settings after
> having created a new project) interpret the source files as Windows-1252
> (which is very similar to ISO-Latin-1).
So, that shouldn't cause a problem for escape sequences, which, as a
matter of deliberate design, consist entirely of characters from the
basic source character set.
>>> Or is it guaranteed that the characters between u" and " will
>>> always be interpreted as UTF-8?
>>
>> Source code characters between the u" and the " will be interpreted
>> according to an implementation-defined character encoding. But so long
>> as they encode {'\\', 'x', 'C', '2', '\\', 'x', 'A', '9'}, you should
>> get the standard-defined behavior for u"\xC2\xA9".
>
> Yes, but that's not the correct desired character in UTF-16, only in UTF-8.
> You'll get garbage as your UTF-16 string literal.
The 'u' mandates UTF-16, which is the only thing that's relevant to the
interpretation of that string literal. That it is the correct pair of
characters, given that UTF-16 has been mandated. Whether or not it's the
intended character depends upon how well your code expresses your
intentions. Alf says that the character that was intended was U+00A9, so
that code does not correctly express that intention. The correct way to
specify it doesn't depend upon the source character set, it only depends
upon the desired encoding of the string. Each of the following ten
escape sequences is a different portably correct ways of expressing that
intention:
UTF-8: u8"\u00A9\U000000A9"
UTF-16: u"\251\xA9\u00A9\U000000A9"
UTF-32: U"\251\xA9\u00A9\U000000A9"
[toc] | [prev] | [next] | [standalone]
| From | Chris Vine <chris@cvine--nospam--.freeserve.co.uk> |
|---|---|
| Date | 2021-07-03 14:23 +0100 |
| Message-ID | <20210703142356.c218ea02236e886e097c5af9@cvine--nospam--.freeserve.co.uk> |
| In reply to | #80631 |
On Sat, 3 Jul 2021 06:59:59 +0000 (UTC) Juha Nieminen <nospam@thanks.invalid> wrote: [snip] > So the problem is the same as with wchar_t: How does the compiler know > which encoding is being used in this source file? It needs to know that > since it has to generate an UTF-16 string literal into the output binary > from those characters appearing in the source code. For encodings other than the 96 charcters of the basic source character set (which map onto ASCII) that the C++ standard requires, this is implementation defined and the compiler should document it. In the case of gcc, it documents that the source character set is UTF-8 unless a different source file encoding is indicated by the -finput-charset option. With gcc you can also set the narrow execution character set with the -fexec-charset option. Presumably for any one string literal this can be overridden by prefixing it with u8, or it wouldn't be consistent with the standard, but I have never checked whether that is in fact the case. This is what gcc says about character sets, which is somewhat divergent from the C and C++ standards: http://gcc.gnu.org/onlinedocs/cpp/Character-sets.html I doubt this is often relevant. What most multi-lingual programs do is have source strings in English using the ASCII subset of UTF-8 and translate to UTF-8 dynamically by reference to the locale. Gnu's gettext is a quite commonly used implementation of this approach.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-03 17:06 +0000 |
| Message-ID | <sbq5dm$7su$1@gioia.aioe.org> |
| In reply to | #80636 |
Chris Vine <chris@cvine--nospam--.freeserve.co.uk> wrote: > I doubt this is often relevant. What most multi-lingual programs do is > have source strings in English using the ASCII subset of UTF-8 and > translate to UTF-8 dynamically by reference to the locale. Gnu's > gettext is a quite commonly used implementation of this approach. It's quite relevant. For example, if you are writing unit tests for some library dealing with wide strings (or UTF-16 strings), it's quite common to write string literals in your tests, so you need to be aware of this problem: What will work just fine with gcc might not work with Visual Studio, and your unit test will succeed in one but not the other. The solution offered elsewhere in this tread is the correct way to go, ie. using the "\uXXXX" escape codes for such string literals, as they will always be interpreted correctly by the compiler (even if the readability of the source code suffers as a consequence).
[toc] | [prev] | [next] | [standalone]
| From | Richard Damon <Richard@Damon-Family.org> |
|---|---|
| Date | 2021-07-03 14:06 -0400 |
| Message-ID | <wG1EI.8148$SMMb.6116@fx17.iad> |
| In reply to | #80650 |
On 7/3/21 1:06 PM, Juha Nieminen wrote: > Chris Vine <chris@cvine--nospam--.freeserve.co.uk> wrote: >> I doubt this is often relevant. What most multi-lingual programs do is >> have source strings in English using the ASCII subset of UTF-8 and >> translate to UTF-8 dynamically by reference to the locale. Gnu's >> gettext is a quite commonly used implementation of this approach. > > It's quite relevant. For example, if you are writing unit tests for some > library dealing with wide strings (or UTF-16 strings), it's quite common > to write string literals in your tests, so you need to be aware of this > problem: What will work just fine with gcc might not work with Visual > Studio, and your unit test will succeed in one but not the other. > > The solution offered elsewhere in this tread is the correct way to go, > ie. using the "\uXXXX" escape codes for such string literals, as they > will always be interpreted correctly by the compiler (even if the > readability of the source code suffers as a consequence). > Add the solution for the readability is to just write the code as native literals, but NOT as the actual C++ file, and have a filter stage that translates this file into the actual C++ code with the escapes. The language was designed for this sort of functionality.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-08 08:13 +0000 |
| Message-ID | <sc6c2j$1bks$2@gioia.aioe.org> |
| In reply to | #80657 |
Richard Damon <Richard@damon-family.org> wrote: > On 7/3/21 1:06 PM, Juha Nieminen wrote: >> Chris Vine <chris@cvine--nospam--.freeserve.co.uk> wrote: >>> I doubt this is often relevant. What most multi-lingual programs do is >>> have source strings in English using the ASCII subset of UTF-8 and >>> translate to UTF-8 dynamically by reference to the locale. Gnu's >>> gettext is a quite commonly used implementation of this approach. >> >> It's quite relevant. For example, if you are writing unit tests for some >> library dealing with wide strings (or UTF-16 strings), it's quite common >> to write string literals in your tests, so you need to be aware of this >> problem: What will work just fine with gcc might not work with Visual >> Studio, and your unit test will succeed in one but not the other. >> >> The solution offered elsewhere in this tread is the correct way to go, >> ie. using the "\uXXXX" escape codes for such string literals, as they >> will always be interpreted correctly by the compiler (even if the >> readability of the source code suffers as a consequence). >> > > Add the solution for the readability is to just write the code as native > literals, but NOT as the actual C++ file, and have a filter stage that > translates this file into the actual C++ code with the escapes. Clearly you have never written unit tests.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-02 15:37 +0200 |
| Message-ID | <sbn4qd$1cs$1@dont-email.me> |
| In reply to | #80617 |
On 2 Jul 2021 07:44, James Kuyper wrote: > On 7/2/21 1:26 AM, Juha Nieminen wrote: >> James Kuyper <jameskuyper@alumni.caltech.edu> wrote: >>> Why would you use wchar_t if you char about unicode? You should be using >>> string literal using either the u8, u, or U prefixes, and store/access >>> the strings as arrays of char, char16_t, or char32_t, respectively. Such >>> literals are guaranteed to be in UTF-8, UTF-16, or UTF-32 encoding, >>> respectively. >> >> No my choice in this case. > > Recent messages reminded me of Window's strong incentives to use > wchar_t, something I've thankfully had no experience with. As a result, > what I'm about to say may be incorrect - but it seems to me that a > conforming implementation of C++ targeting Windows should have wchar_t > be the same as char16_t, so you should be able to freely use u prefixed > string literals and char16_t with code that needs to be portable to > Windows, but can also be used on other platforms. If your code didn't > need to be portable to other platforms, you could, by definition, rely > upon Window's own guarantees about wchar_t. I agree, but unfortunately both the C and C++ standards require `wchar_t` to be able represent all code points in the largest supported extended character set, and while that requirement worked nicely with original Unicode, already in 1992 or thereabouts (not quite sure) it was in conflict with extremely firmly established practice in Windows. Bringing the standards in agreement with actual practice should be a goal, but for C++ it's seldom done. It /was/ done, in C++11, for the hopeless idealistic C++98 goal of clean C++-ish `<c...>` headers that didn't pollute the global namespace. But C++17 added stuff to only some `<c...>` headers and not to the corresponding `<... .h>` headers, and that fact was then used quite recently in a proposal to un-deprecate the .h headers but add all new stuff only to `<c...>` headers. Which satisfies the politics but not the in-practice of using code that uses C libraries designed as C++-compatible, without running afoul of qualification issues. And it /was/ done, in C++11, for both `throw` specifications and for the `export` keyword, not to mention the C++11 optional conversion between function pointers and `void*`, in order to support the reality of Posix. But it seems to me, it's very very hard to get such changes through the committee, just as individual people generally don't like to admit that they've been wrong. Instead, C++20 started on a path of /introducing/ more conflicts with reality. In particular for `std::filesystem::path`, where they threw the baby out with the bathwater for purely academic idealism reasons. - Alf
[toc] | [prev] | [next] | [standalone]
Page 3 of 4 — ← Prev page 1 2 [3] 4 Next page →
Back to top | Article view | comp.lang.c++
csiph-web