Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #80571 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2021-06-30 07:57 +0000 |
| Last post | 2021-07-01 15:44 +0200 |
| Articles | 20 on this page of 70 — 19 participants |
Back to article view | Back to comp.lang.c++
How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 07:57 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 10:05 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:30 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 12:37 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 09:19 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 11:10 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 11:36 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 16:17 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 10:13 -0700
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 19:27 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 11:21 -0700
Re: How to write wide char string literals? Real Troll <real.troll@trolls.com> - 2021-07-01 18:45 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-02 10:03 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-02 02:31 -0700
Re: How to write wide char string literals? MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk - 2021-07-02 09:38 +0000
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 13:52 +0300
Re: How to write wide char string literals? MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net - 2021-07-02 12:53 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:22 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 17:32 +0300
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:08 +0200
Re: How to write wide char string literals? MrSpud_r5j@ywn9entw2s.org - 2021-06-30 08:39 +0000
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 10:52 +0000
Re: How to write wide char string literals? MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov - 2021-06-30 11:28 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 14:01 +0200
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 10:55 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 11:23 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 06:51 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:09 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:35 -0400
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 14:03 +0300
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:13 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:39 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 13:50 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-06-30 14:19 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-01 13:31 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 04:42 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-01 10:58 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 05:26 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 01:44 -0400
Re: How to write wide char string literals? Bo Persson <bo@bo-persson.se> - 2021-07-02 08:33 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 11:52 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 16:15 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 03:30 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 00:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 13:31 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 08:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 16:28 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 11:48 -0400
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 15:56 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 16:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 17:49 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:11 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 05:52 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 06:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 09:03 -0400
Re: How to write wide char string literals? Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2021-07-03 14:23 +0100
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 17:06 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 14:06 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:13 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:37 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-06-30 16:54 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 17:56 +0300
Re: How to write wide char string literals? Öö Tiib <ootiib@hot.ee> - 2021-06-30 09:22 -0700
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 07:19 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 10:29 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 08:44 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 12:58 +0200
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 14:01 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 14:57 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-07-01 15:44 +0200
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
| From | MrSpud_r5j@ywn9entw2s.org |
|---|---|
| Date | 2021-06-30 08:39 +0000 |
| Message-ID | <sbhajm$1rj2$1@gioia.aioe.org> |
| In reply to | #80571 |
On Wed, 30 Jun 2021 07:57:14 +0000 (UTC) Juha Nieminen <nospam@thanks.invalid> wrote: >Character encoding was a problem in the 1960's, and it's still a problem >today, no matter how much computers advance. Sheesh. > >Problem is, how to reliably write wide char string literals that contain >non-ascii characters? > >Suppose you write for example this: > > const wchar_t* str = L"???"; > >In the *source code* that string literal may be eg. UTF-8 encoded. However, >the compiler needs to convert it to wide chars. > >Problem is, how does the compiler know which encoding is being used in >that 8-bit string literal in the source code, in order for it to convert >it properly to wide chars? > >Some compilers may assume it's UTF-8 encoded source code. Others may >assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio). >Obviously the end result will be garbage if the wrong assumption is made. > >In most compilers (such as Visual Studio) you can specify which encoding >to assume for source files, but this has to be done at the project >settings level. I don't think there's any way to specify the encoding >in the source code itself. > >What does the C++ standard say? Does it say that source code files are >always UTF-8 encoded, or is it up to the implementation? I assume that if >it's the latter, the standard doesn't provide any mechanism to specify >which encoding is being used. Or does it? Why should it care? To C and C++ strings are just a sequence of bytes, the encoding is irrelevant unless you're using functions specific to a particular encoding, eg: utf8_strlen() or similar.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-06-30 10:52 +0000 |
| Message-ID | <sbhie0$1i69$1@gioia.aioe.org> |
| In reply to | #80575 |
MrSpud_r5j@ywn9entw2s.org wrote: > Why should it care? To C and C++ strings are just a sequence of bytes, the > encoding is irrelevant unless you're using functions specific to a particular > encoding, eg: utf8_strlen() or similar. That would be correct if this were a char string literal. But it's not. It's a wide char string literal. L"something". This means that in the source file the stuff between the quotes is, for example, UTF-8 encoded, but the compiler needs to produce a wide char string into the compiled binary, so the compiler needs to perform at compile time a string encoding conversion from 8-bit UTF-8 to whatever a wchar_t* may be (most usually either UTF-16 or UTF-32).
[toc] | [prev] | [next] | [standalone]
| From | MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov |
|---|---|
| Date | 2021-06-30 11:28 +0000 |
| Message-ID | <sbhkhn$jle$1@gioia.aioe.org> |
| In reply to | #80580 |
On Wed, 30 Jun 2021 10:52:50 +0000 (UTC) Juha Nieminen <nospam@thanks.invalid> wrote: >MrSpud_r5j@ywn9entw2s.org wrote: >> Why should it care? To C and C++ strings are just a sequence of bytes, the >> encoding is irrelevant unless you're using functions specific to a particular > >> encoding, eg: utf8_strlen() or similar. > >That would be correct if this were a char string literal. But it's not. >It's a wide char string literal. L"something". > >This means that in the source file the stuff between the quotes is, >for example, UTF-8 encoded, but the compiler needs to produce a wide >char string into the compiled binary, so the compiler needs to perform >at compile time a string encoding conversion from 8-bit UTF-8 to >whatever a wchar_t* may be (most usually either UTF-16 or UTF-32). I suspect thats something only Windows programmers have to worry about. utf8 has been the de facto standard on *nix for years.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2021-06-30 14:01 +0200 |
| Message-ID | <sbhmfn$amc$1@dont-email.me> |
| In reply to | #80584 |
On 30/06/2021 13:28, MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov wrote: > On Wed, 30 Jun 2021 10:52:50 +0000 (UTC) > Juha Nieminen <nospam@thanks.invalid> wrote: >> MrSpud_r5j@ywn9entw2s.org wrote: >>> Why should it care? To C and C++ strings are just a sequence of bytes, the >>> encoding is irrelevant unless you're using functions specific to a particular >> >>> encoding, eg: utf8_strlen() or similar. >> >> That would be correct if this were a char string literal. But it's not. >> It's a wide char string literal. L"something". >> >> This means that in the source file the stuff between the quotes is, >> for example, UTF-8 encoded, but the compiler needs to produce a wide >> char string into the compiled binary, so the compiler needs to perform >> at compile time a string encoding conversion from 8-bit UTF-8 to >> whatever a wchar_t* may be (most usually either UTF-16 or UTF-32). > > I suspect thats something only Windows programmers have to worry about. utf8 > has been the de facto standard on *nix for years. > UTF-8 has been the standard for most purposes for many years (with UTF-32 used internally sometimes). However, there are a few very important exceptions that use UTF-16, because they started using Unicode in the early days when it looked like UTF-16 (or in fact just UCS-2) would be sufficient. That includes Windows, Java, QT and Javascript. Moving these to UTF-8 takes time.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-06-30 10:55 +0200 |
| Message-ID | <sbhbi5$fm3$1@dont-email.me> |
| In reply to | #80571 |
On 30 Jun 2021 09:57, Juha Nieminen wrote:
> Character encoding was a problem in the 1960's, and it's still a problem
> today, no matter how much computers advance. Sheesh.
>
> Problem is, how to reliably write wide char string literals that contain
> non-ascii characters?
>
> Suppose you write for example this:
>
> const wchar_t* str = L"???";
>
> In the *source code* that string literal may be eg. UTF-8 encoded. However,
> the compiler needs to convert it to wide chars.
>
> Problem is, how does the compiler know which encoding is being used in
> that 8-bit string literal in the source code, in order for it to convert
> it properly to wide chars?
The compiler necessarily assumes some source encoding.
g++ and Visual C++ use different schemes for determining the source code
encoding assumption.
g++ uses a single encoding assumption that you can change via options,
while Visual C++ by default determines the encoding for each individual
file, which is a much more flexible scheme. However, in modern
programming work you don't want to use that flexible Visual C++ scheme
because the base assumption, when no other indication is present, is
that a file is Windows ANSI encoded, while in modern programming work
it's most likely UTF-8 encoded. So it's now a good idea to use the
Visual C++ UTF-8 option, plus some others, e.g.
/nologo /utf-8 /EHsc /GR /permissive- /FI"iso646.h" /std:c++17
/Zc:__cplusplus /Zc:externC- /W4 /wd4459 /D _CRT_SECURE_NO_WARNINGS=1 /D
_STL_SECURE_NO_WARNINGS=1
> Some compilers may assume it's UTF-8 encoded source code. Others may
> assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio).
> Obviously the end result will be garbage if the wrong assumption is made.
Yes. You can to some extent prevent Visual C++ mis-interpretation by
using the UTF-8 BOM as an encoding indicator, and I recommend that.
However, there are costs, in particular that mindless Linux fanbois (all
fanbois are mindless, even C++ fanbois) hung up on supporting archaic
Linux tools that can't handle the BOM, can then brand you as this and
that; and that's not hypothetical, it's direct experience. Also, even
though using a BOM is a very strong convention in Windows the Cmd `type`
command can't handle it, so that one is nudged in the direction of
Powershell, which is a monstrosity that I really hate.
> In most compilers (such as Visual Studio) you can specify which encoding
> to assume for source files, but this has to be done at the project
> settings level.
Uhm, no, you can specify compiler options per file if you want, in each
file's properties.
Visual Studio 2019 screenshot: (https://ibb.co/tJ5jNJC)
> I don't think there's any way to specify the encoding
> in the source code itself.
Not in standard C++. For Visual C++ there is an undocumented (or used to
be undocumented) `#pragma` used e.g. in automatically generated resource
scripts, .rc files. I don't recall the name. Also, there is the UTF-8
BOM. An UTF-8 BOM is a pretty surefire way to force UTF-8 assumption.
> What does the C++ standard say? Does it say that source code files are
> always UTF-8 encoded, or is it up to the implementation?
It's totally up to the implementation.
That wouldn't be so bad if the standard had addressed the issue of a
collection of source files, in particular headers, with different
encodings, e.g. if the standard had /required/ all source files in a
translation unit to have the same encoding.
That's the assumption of g++, but not of Visual C++.
> I assume that if
> it's the latter, the standard doesn't provide any mechanism to specify
> which encoding is being used. Or does it?
Right. It's a mess. :-o :-)
But, practical solutions:
• Use UTF-8 BOM and Just Ignore™ whining from Linux fanbois.
• For good measure also use `/utf-8` option with Visual C++.
• Where it matters you can /statically assert/ UTF-8 encoding.
A `static_assert` depends on both that the the compiler's source file
encoding assumption is correct, whatever it is, and that the basic
execution character set (encoding of literals in the executable) is
UTF-8. These are separate encoding choices and can be specified
separately both with g++ and Visual C++. But assuming they both hold,
constexpr inline auto utf8_is_the_execution_character_set()
-> bool
{
constexpr auto& slashed_o = "ø";
return (sizeof( slashed_o ) == 3 and slashed_o[0] == '\xC3' and
slashed_o[1] == '\xB8');
}
When a `static_assert(utf8_is_the_execution_character_set())` holds you
can be pretty sure that the source encoding assumption is correct.
- Alf
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2021-06-30 11:23 +0200 |
| Message-ID | <sbhd6d$a0b$1@dont-email.me> |
| In reply to | #80576 |
On 30/06/2021 10:55, Alf P. Steinbach wrote:
> Right. It's a mess. :-o :-)
>
> But, practical solutions:
>
> • Use UTF-8 BOM and Just Ignore™ whining from Linux fanbois.
Should you also ignore the recommendations from Unicode people? I say
you should ignore Windows Notepad fanbois, and drop the BOM.
Some programs can't handle a UTF-8 BOM. Some programs can't handle (or
at least, can't automatically recognise) UTF-8 encoding without a BOM.
Some programs add a UTF-8 BOM automatically, some remove it
automatically, some don't care whether it is there or not.
Like it or not, your success with or without a UTF-8 BOM is going to
depend on the programs you use.
If you use a lot of programs that can't work properly without it (such
as Windows Notepad), use a BOM. If you are able to live without it,
perhaps by telling your editor to assume UTF-8 or adding a compiler
switch to your build system, do so.
When you have the option, /always/ choose UTF-8 encoding /without/ a
BOM. It is far and away the most popular format for text, and it is the
format you are already using. It is the format used by all the include
files you have for all your libraries (including the standard library)
on your system, and on every other system. That is because plain ASCII
is also in UTF-8 with no BOM. It is the only unicode encoding that is
fully compatible with the files you have - it is therefore your only
option if your preference is to have a single encoding.
The sooner encodings other than BOM-less UTF-8 die out, the better.
That is the only way out of the mess.
> • For good measure also use `/utf-8` option with Visual C++.
> • Where it matters you can /statically assert/ UTF-8 encoding.
> > A `static_assert` depends on both that the the compiler's source file
> encoding assumption is correct, whatever it is, and that the basic
> execution character set (encoding of literals in the executable) is
> UTF-8. These are separate encoding choices and can be specified
> separately both with g++ and Visual C++. But assuming they both hold,
>
> constexpr inline auto utf8_is_the_execution_character_set()
> -> bool
> {
> constexpr auto& slashed_o = "ø";
> return (sizeof( slashed_o ) == 3 and slashed_o[0] == '\xC3' and
> slashed_o[1] == '\xB8');
> }
>
> When a `static_assert(utf8_is_the_execution_character_set())` holds you
> can be pretty sure that the source encoding assumption is correct.
>
Static assertions are always a good idea. So are pragmas forcing
options, for compilers that support that.
[toc] | [prev] | [next] | [standalone]
| From | Richard Damon <Richard@Damon-Family.org> |
|---|---|
| Date | 2021-06-30 06:51 -0400 |
| Message-ID | <_0YCI.4454$_X3.2330@fx40.iad> |
| In reply to | #80571 |
On 6/30/21 3:57 AM, Juha Nieminen wrote: > Character encoding was a problem in the 1960's, and it's still a problem > today, no matter how much computers advance. Sheesh. > > Problem is, how to reliably write wide char string literals that contain > non-ascii characters? > > Suppose you write for example this: > > const wchar_t* str = L"???"; > > In the *source code* that string literal may be eg. UTF-8 encoded. However, > the compiler needs to convert it to wide chars. > > Problem is, how does the compiler know which encoding is being used in > that 8-bit string literal in the source code, in order for it to convert > it properly to wide chars? > > Some compilers may assume it's UTF-8 encoded source code. Others may > assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio). > Obviously the end result will be garbage if the wrong assumption is made. > > In most compilers (such as Visual Studio) you can specify which encoding > to assume for source files, but this has to be done at the project > settings level. I don't think there's any way to specify the encoding > in the source code itself. > > What does the C++ standard say? Does it say that source code files are > always UTF-8 encoded, or is it up to the implementation? I assume that if > it's the latter, the standard doesn't provide any mechanism to specify > which encoding is being used. Or does it? > You have this backwards, by the Standard, you don't tell the implementation what encoding the source files are, the implementation tells you what encoding it specifies that you should use. The implementation is allowed to give you a way to tell it what to tell you, but this is all implementation details. There is a fundamental issue with trying to define an in-source way to specify this, as we don't even have the ability to assume the ASCII is part of the encoding as it could be EBCDIC. Yes, if we got to throw out everything and start fresh, we might do things differently. As to the question of how to put the characters in a string, that is what escape codes like \u and \U are for.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-06-30 11:09 +0000 |
| Message-ID | <sbhjdo$1np$1@gioia.aioe.org> |
| In reply to | #80579 |
Richard Damon <Richard@damon-family.org> wrote: > As to the question of how to put the characters in a string, that is > what escape codes like \u and \U are for. While perhaps not ideal in terms of code readability (or writability, as one needs to look up the unicode code points of each non-ascii character), I suppose this is the best solution for portable code.
[toc] | [prev] | [next] | [standalone]
| From | Richard Damon <Richard@Damon-Family.org> |
|---|---|
| Date | 2021-06-30 07:35 -0400 |
| Message-ID | <AGYCI.2431$4G4.2424@fx10.iad> |
| In reply to | #80582 |
On 6/30/21 7:09 AM, Juha Nieminen wrote: > Richard Damon <Richard@damon-family.org> wrote: >> As to the question of how to put the characters in a string, that is >> what escape codes like \u and \U are for. > > While perhaps not ideal in terms of code readability (or writability, as > one needs to look up the unicode code points of each non-ascii character), > I suppose this is the best solution for portable code. > And, as somewhat common in the standard, this sort of thing was designed so if you wanted to, you could have written the file is some encoding that the compiler doesn't know, and then run it through a pre-processor that performs this transform to the 'ugly' codes. Thus the original source file (that would be the one to edit to make changes) would be very readable, and the uglyness is hidden in a temporary intermediary file. For example, the file might be in a .cpp16 file that is UTF-16 encoded. The make file will have a recipe for convert .cpp16 to .cpp switching to the compilers 'natural' character set.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <myfirstname@osa.pri.ee> |
|---|---|
| Date | 2021-06-30 14:03 +0300 |
| Message-ID | <sbhj2g$d22$1@dont-email.me> |
| In reply to | #80571 |
30.06.2021 10:57 Juha Nieminen kirjutas: > Character encoding was a problem in the 1960's, and it's still a problem > today, no matter how much computers advance. Sheesh. > > Problem is, how to reliably write wide char string literals that contain > non-ascii characters? > > Suppose you write for example this: > > const wchar_t* str = L"???"; > I want my code to work anywhere with any compiler/framework conventions and settings, and all my internal strings are in UTF-8 anyway, so I can use strict ASCII source files with hardcoded UTF-8 characters, e.g.: std::string s = "Copyright \xC2\xA9 2001-2020"; One can find the UTF-8 codes for such symbols quite easily from pages like https://www.fileformat.info/info/unicode/char/a9/index.htm For converting my strings to wide strings for Windows SDK functions, I have small utility functions like Utf2Win(): ::MessageBoxW(nullptr, Utf2Win(s).c_str(), L"About", MB_OK); This setup means I do not have to worry about source code codepage conventions *at all*. Fortunately I do not have much such texts. YMMV.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-06-30 11:13 +0000 |
| Message-ID | <sbhjlc$5cb$1@gioia.aioe.org> |
| In reply to | #80581 |
Paavo Helde <myfirstname@osa.pri.ee> wrote: > I want my code to work anywhere with any compiler/framework conventions > and settings, and all my internal strings are in UTF-8 anyway, so I can > use strict ASCII source files with hardcoded UTF-8 characters, e.g.: > > std::string s = "Copyright \xC2\xA9 2001-2020"; Does that work for wide string literals? Because I don't think it does. In other words: std::wstring s = L"Copyright \xC2\xA9 2001-2020"; However, as suggested in another reply, using "\uXXXX" instead ought to work just fine (regardless of whether it's a narrow or wide char literal). As long as you don't need the readability, of course.
[toc] | [prev] | [next] | [standalone]
| From | Richard Damon <Richard@Damon-Family.org> |
|---|---|
| Date | 2021-06-30 07:39 -0400 |
| Message-ID | <YJYCI.5782$5J1.4863@fx08.iad> |
| In reply to | #80583 |
On 6/30/21 7:13 AM, Juha Nieminen wrote: > Paavo Helde <myfirstname@osa.pri.ee> wrote: >> I want my code to work anywhere with any compiler/framework conventions >> and settings, and all my internal strings are in UTF-8 anyway, so I can >> use strict ASCII source files with hardcoded UTF-8 characters, e.g.: >> >> std::string s = "Copyright \xC2\xA9 2001-2020"; > > Does that work for wide string literals? Because I don't think it does. > In other words: > > std::wstring s = L"Copyright \xC2\xA9 2001-2020"; > > However, as suggested in another reply, using "\uXXXX" instead ought > to work just fine (regardless of whether it's a narrow or wide char > literal). As long as you don't need the readability, of course. > \x works in wide string literal too, and puts in a character with that value. The difference is that if the wide string type isn't unicode encoded then it might get the wrong character in the string.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-06-30 13:50 +0200 |
| Message-ID | <sbhlpp$r64$1@dont-email.me> |
| In reply to | #80586 |
On 30 Jun 2021 13:39, Richard Damon wrote: > On 6/30/21 7:13 AM, Juha Nieminen wrote: >> Paavo Helde <myfirstname@osa.pri.ee> wrote: >>> I want my code to work anywhere with any compiler/framework conventions >>> and settings, and all my internal strings are in UTF-8 anyway, so I can >>> use strict ASCII source files with hardcoded UTF-8 characters, e.g.: >>> >>> std::string s = "Copyright \xC2\xA9 2001-2020"; >> >> Does that work for wide string literals? Because I don't think it does. >> In other words: >> >> std::wstring s = L"Copyright \xC2\xA9 2001-2020"; >> >> However, as suggested in another reply, using "\uXXXX" instead ought >> to work just fine (regardless of whether it's a narrow or wide char >> literal). As long as you don't need the readability, of course. >> > > \x works in wide string literal too, and puts in a character with that > value. The difference is that if the wide string type isn't unicode > encoded then it might get the wrong character in the string. It gets the wrong characters in the wide string literal, period. - Alf
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-06-30 14:19 -0400 |
| Message-ID | <sbicji$8kd$1@dont-email.me> |
| In reply to | #80587 |
On 6/30/21 7:50 AM, Alf P. Steinbach wrote: > On 30 Jun 2021 13:39, Richard Damon wrote: ... >> \x works in wide string literal too, and puts in a character with that >> value. The difference is that if the wide string type isn't unicode >> encoded then it might get the wrong character in the string. > > It gets the wrong characters in the wide string literal, period. "The escape \ooo consists of the backslash followed by one, two, or three octal digits that are taken to specify the value of the desired character. ... The value of a character-literal is implementation-defined if it falls outside of the implementation-defined range defined for ... wchar_t (for character-literals prefixed by L)." (5.13.3p7) The value of a wide character is determined by the current encoding. For wide character literals using the u or U prefixes, that encoding is UTF-16 and UTF-32, respectively, making octal escapes redundant with and less convenient than the use of UCNs. But as he said, they do work for such strings.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-01 13:31 +0200 |
| Message-ID | <sbk92c$mee$1@dont-email.me> |
| In reply to | #80592 |
On 30 Jun 2021 20:19, James Kuyper wrote: > On 6/30/21 7:50 AM, Alf P. Steinbach wrote: >> On 30 Jun 2021 13:39, Richard Damon wrote: > ... >>> \x works in wide string literal too, and puts in a character with that >>> value. The difference is that if the wide string type isn't unicode >>> encoded then it might get the wrong character in the string. >> >> It gets the wrong characters in the wide string literal, period. > "The escape \ooo consists of the backslash followed by one, two, or > three octal digits that are taken to specify the value of the desired > character. ... The value of a character-literal is > implementation-defined if it falls outside of the implementation-defined > range defined for ... wchar_t (for character-literals prefixed by L)." > (5.13.3p7) > > The value of a wide character is determined by the current encoding. For > wide character literals using the u or U prefixes, that encoding is > UTF-16 and UTF-32, respectively, making octal escapes redundant with and > less convenient than the use of UCNs. But as he said, they do work for > such strings. You snipped some context, the example we're talking about. That does decidedly not work in the sense of producing the intended string. Perhaps I can make you understand this by talking about source code in general. Yes, that example code is valid C++, so a conforming compiler shall compile it with no errors; and yes, that code has a well defined meaning, look, here's the C++ standard, it spells it out, what the meaning is. But no, it doesn't do what you intended. - Alf
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-01 04:42 +0000 |
| Message-ID | <sbjh2o$1bq6$1@gioia.aioe.org> |
| In reply to | #80586 |
Richard Damon <Richard@damon-family.org> wrote: > On 6/30/21 7:13 AM, Juha Nieminen wrote: >> Paavo Helde <myfirstname@osa.pri.ee> wrote: >>> I want my code to work anywhere with any compiler/framework conventions >>> and settings, and all my internal strings are in UTF-8 anyway, so I can >>> use strict ASCII source files with hardcoded UTF-8 characters, e.g.: >>> >>> std::string s = "Copyright \xC2\xA9 2001-2020"; >> >> Does that work for wide string literals? Because I don't think it does. >> In other words: >> >> std::wstring s = L"Copyright \xC2\xA9 2001-2020"; >> >> However, as suggested in another reply, using "\uXXXX" instead ought >> to work just fine (regardless of whether it's a narrow or wide char >> literal). As long as you don't need the readability, of course. >> > > \x works in wide string literal too, and puts in a character with that > value. The difference is that if the wide string type isn't unicode > encoded then it might get the wrong character in the string. The problem is that "\xC2\xA9" in UTF-8 is not the same thing as "\xC2\xA9" in UTF-16 or UTF-32 (whichever wchar_t happens to be). "\uXXXX", however, ought to work regardless because it specifies the actual unicode codepoint you want, rather than its encoding.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-01 10:58 -0400 |
| Message-ID | <sbkl6u$cd5$1@dont-email.me> |
| In reply to | #80593 |
On 7/1/21 12:42 AM, Juha Nieminen wrote: > Richard Damon <Richard@damon-family.org> wrote: >> On 6/30/21 7:13 AM, Juha Nieminen wrote: >>> Paavo Helde <myfirstname@osa.pri.ee> wrote: >>>> I want my code to work anywhere with any compiler/framework conventions >>>> and settings, and all my internal strings are in UTF-8 anyway, so I can >>>> use strict ASCII source files with hardcoded UTF-8 characters, e.g.: >>>> >>>> std::string s = "Copyright \xC2\xA9 2001-2020"; >>> >>> Does that work for wide string literals? Because I don't think it does. >>> In other words: >>> >>> std::wstring s = L"Copyright \xC2\xA9 2001-2020"; >>> >>> However, as suggested in another reply, using "\uXXXX" instead ought >>> to work just fine (regardless of whether it's a narrow or wide char >>> literal). As long as you don't need the readability, of course. >>> >> >> \x works in wide string literal too, and puts in a character with that >> value. The difference is that if the wide string type isn't unicode >> encoded then it might get the wrong character in the string. > > The problem is that "\xC2\xA9" in UTF-8 is not the same thing as > "\xC2\xA9" in UTF-16 or UTF-32 (whichever wchar_t happens to be). Why would you use wchar_t if you char about unicode? You should be using string literal using either the u8, u, or U prefixes, and store/access the strings as arrays of char, char16_t, or char32_t, respectively. Such literals are guaranteed to be in UTF-8, UTF-16, or UTF-32 encoding, respectively.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-07-02 05:26 +0000 |
| Message-ID | <sbm82o$1rhp$1@gioia.aioe.org> |
| In reply to | #80606 |
James Kuyper <jameskuyper@alumni.caltech.edu> wrote: > Why would you use wchar_t if you char about unicode? You should be using > string literal using either the u8, u, or U prefixes, and store/access > the strings as arrays of char, char16_t, or char32_t, respectively. Such > literals are guaranteed to be in UTF-8, UTF-16, or UTF-32 encoding, > respectively. No my choice in this case.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-07-02 01:44 -0400 |
| Message-ID | <sbm94h$fif$1@dont-email.me> |
| In reply to | #80616 |
On 7/2/21 1:26 AM, Juha Nieminen wrote: > James Kuyper <jameskuyper@alumni.caltech.edu> wrote: >> Why would you use wchar_t if you char about unicode? You should be using >> string literal using either the u8, u, or U prefixes, and store/access >> the strings as arrays of char, char16_t, or char32_t, respectively. Such >> literals are guaranteed to be in UTF-8, UTF-16, or UTF-32 encoding, >> respectively. > > No my choice in this case. Recent messages reminded me of Window's strong incentives to use wchar_t, something I've thankfully had no experience with. As a result, what I'm about to say may be incorrect - but it seems to me that a conforming implementation of C++ targeting Windows should have wchar_t be the same as char16_t, so you should be able to freely use u prefixed string literals and char16_t with code that needs to be portable to Windows, but can also be used on other platforms. If your code didn't need to be portable to other platforms, you could, by definition, rely upon Window's own guarantees about wchar_t.
[toc] | [prev] | [next] | [standalone]
| From | Bo Persson <bo@bo-persson.se> |
|---|---|
| Date | 2021-07-02 08:33 +0200 |
| Message-ID | <ik7q8uFp0nnU1@mid.individual.net> |
| In reply to | #80617 |
On 2021-07-02 at 07:44, James Kuyper wrote: > On 7/2/21 1:26 AM, Juha Nieminen wrote: >> James Kuyper <jameskuyper@alumni.caltech.edu> wrote: >>> Why would you use wchar_t if you char about unicode? You should be using >>> string literal using either the u8, u, or U prefixes, and store/access >>> the strings as arrays of char, char16_t, or char32_t, respectively. Such >>> literals are guaranteed to be in UTF-8, UTF-16, or UTF-32 encoding, >>> respectively. >> >> No my choice in this case. > > Recent messages reminded me of Window's strong incentives to use > wchar_t, something I've thankfully had no experience with. As a result, > what I'm about to say may be incorrect - but it seems to me that a > conforming implementation of C++ targeting Windows should have wchar_t > be the same as char16_t, so you should be able to freely use u prefixed > string literals and char16_t with code that needs to be portable to > Windows, but can also be used on other platforms. If your code didn't > need to be portable to other platforms, you could, by definition, rely > upon Window's own guarantees about wchar_t. > The problem is that the language explicitly requires wchar_t to be a distinct type. Once upon a time it was implemented as a typedef, but that option was removed already in C++98. Overload resolution and things...
[toc] | [prev] | [next] | [standalone]
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
Back to top | Article view | comp.lang.c++
csiph-web