Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #82565

Re: Character encoding conversion in wide string literals

From "Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Newsgroups comp.lang.c++
Subject Re: Character encoding conversion in wide string literals
Date 2021-12-07 16:44 +0100
Organization A noiseless patient Spider
Message-ID <sonvh8$vc9$1@dont-email.me> (permalink)
References <sonh15$g9j$1@gioia.aioe.org>

Show all headers | View raw


On 7 Dec 2021 12:37, Juha Nieminen wrote:
> Recently I stumbled across a problem where I had wide string literals
> with non-ascii characters UTF-8 encoded. In other words, I had code like
> this (I'm using non-ascii in the code below, I hope it doesn't get
> mangled up, but even if it does, it should nevertheless be clear what
> I'm trying to express):
> 
>      std::wstring str = L"non-ascii chars: ???";
> 
> The C++ source file itself uses UTF-8 encoding, meaning that that line
> of code is likewise UTF-8 encoded. If it were a narrow string literal
> (being assigned to a std::string) then it works just fine (primarily
> because the compiler doesn't need to do anything to it, it can simply
> take those bytes from the source file as is).
> 
> However, since it's a wide string literal (being assigned to a std::wstring)
> it's not as clear-cut anymore. What does the standard say about this
> situation?
> 
> The thing is that it works just fine in Linux using gcc. The compiler will
> re-encode the UTF-8 encoded characters in the source file inside the
> parentheses into whatever encoding wide char string use, so the correct
> content will end up in the executable binary (and thus in the wstring).
> 
> Apparently it does not work correctly in (some recent version of)
> Visual Studio, where apparently it just takes the byte values from the
> source file within the parentheses as-is, and just assigns those values
> as-is to the wide chars that end up in the binary. (Or something like that.)

The Visual C++ compiler assumes that source code is Windows ANSI encoded 
unless

* you use an encoding option such as `/utf8`, or
* the source is UTF-8 with BOM, or
* the source is UTF-16.

Independently of that Visual C++ assumes that the execution character 
set (the byte-based encoding that should be used for text data in the 
executable) is Windows ANSI, unless it's specified as something else. 
The `/utf8` option specifies also that. It's a combo option that 
specifies both source encoding and execution character set as UTF-8.

Unfortunately as of VS 2022 `/utf8` is not set by default in a VS 
project, and unfortunately there's nothing you can just click to set it. 
You have to to type it in (right click project then) Properties -> C/C++ 
-> Command Line. I usually set "/utf-8 /Zc:__cplusplus".


> Does the standard specify what the compiler should do in this situation?
> If not, then what is the proper way of specifying wide string literals
> that contain non-ascii characters?

I'll let others discuss that, but (1) it does, and (2) just so you're 
aware: the main problem is that the C and C++ standards do not conform 
to reality in their requirement that a `wchar_t` value should suffice to 
encode all possible code points in the wide character set.

In Windows wide text is UTF-16, with 16-bit `wchar_t`. Which means that 
some emojis etc. that appear as a single character and constitute one 
21-bit code point, can become a pair of two `wchar_t` values, an UTF-16 
"surrogate pair".

That's probably not your problem though, but it is a/the problem.


- ALf

Back to comp.lang.c++ | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

Character encoding conversion in wide string literals Juha Nieminen <nospam@thanks.invalid> - 2021-12-07 11:37 +0000
  Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 16:44 +0100
    Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:48 -0500
      Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 18:59 +0100
        Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
        Re: Character encoding conversion in wide string literals Öö Tiib <ootiib@hot.ee> - 2021-12-07 16:39 -0800
        Re: Character encoding conversion in wide string literals David Brown <david.brown@hesbynett.no> - 2021-12-08 09:18 +0100
    Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 09:41 -0800
      Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 19:07 +0100
        Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:21 -0800
      Re: Character encoding conversion in wide string literals Paavo Helde <eesnimi@osa.pri.ee> - 2021-12-07 20:32 +0200
        Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
        Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:26 -0800
  Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:38 -0500
    Re: Character encoding conversion in wide string literals Manfred <noname@add.invalid> - 2021-12-07 19:07 +0100

csiph-web