Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #82571

Re: Character encoding conversion in wide string literals

Path csiph.com!eternal-september.org!reader02.eternal-september.org!.POSTED!not-for-mail
From "Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Newsgroups comp.lang.c++
Subject Re: Character encoding conversion in wide string literals
Date Tue, 7 Dec 2021 19:07:00 +0100
Organization A noiseless patient Spider
Lines 61
Message-ID <soo7s7$uce$1@dont-email.me> (permalink)
References <sonh15$g9j$1@gioia.aioe.org> <sonvh8$vc9$1@dont-email.me> <87r1aob9d3.fsf@nosuchdomain.example.com>
Mime-Version 1.0
Content-Type text/plain; charset=UTF-8; format=flowed
Content-Transfer-Encoding 7bit
Injection-Date Tue, 7 Dec 2021 18:07:03 -0000 (UTC)
Injection-Info reader02.eternal-september.org; posting-host="082f8b1fdbf905e3e36c03435b0eccad"; logging-data="31118"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX19oRJuRW0cBnf2ldFK5Y/nQ"
User-Agent Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.3.2
Cancel-Lock sha1:8/JCOGMeDMC2dwTMe6K+je3L8qs=
In-Reply-To <87r1aob9d3.fsf@nosuchdomain.example.com>
Content-Language en-US
Xref csiph.com comp.lang.c++:82571

Show key headers only | View raw


On 7 Dec 2021 18:41, Keith Thompson wrote:
> "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
>> On 7 Dec 2021 12:37, Juha Nieminen wrote:
>>> Recently I stumbled across a problem where I had wide string literals
>>> with non-ascii characters UTF-8 encoded. In other words, I had code like
>>> this (I'm using non-ascii in the code below, I hope it doesn't get
>>> mangled up, but even if it does, it should nevertheless be clear what
>>> I'm trying to express):
>>>       std::wstring str = L"non-ascii chars: ???";
>>> The C++ source file itself uses UTF-8 encoding, meaning that that
>>> line
>>> of code is likewise UTF-8 encoded. If it were a narrow string literal
>>> (being assigned to a std::string) then it works just fine (primarily
>>> because the compiler doesn't need to do anything to it, it can simply
>>> take those bytes from the source file as is).
>>> However, since it's a wide string literal (being assigned to a
>>> std::wstring)
>>> it's not as clear-cut anymore. What does the standard say about this
>>> situation?
>>> The thing is that it works just fine in Linux using gcc. The
>>> compiler will
>>> re-encode the UTF-8 encoded characters in the source file inside the
>>> parentheses into whatever encoding wide char string use, so the correct
>>> content will end up in the executable binary (and thus in the wstring).
>>> Apparently it does not work correctly in (some recent version of)
>>> Visual Studio, where apparently it just takes the byte values from the
>>> source file within the parentheses as-is, and just assigns those values
>>> as-is to the wide chars that end up in the binary. (Or something like that.)
>>
>> The Visual C++ compiler assumes that source code is Windows ANSI
>> encoded unless
>>
>> * you use an encoding option such as `/utf8`, or
>> * the source is UTF-8 with BOM, or
>> * the source is UTF-16.
> 
> What exactly do you mean by "Windows ANSI"?  Windows-1252 or something
> else?  (Microsoft doesn't call it "ANSI", because it isn't.)
> 
> [...]

"Windows ANSI" is the encoding specified by the `GetACP` API function, 
which, but as I recall that's more or less undocumented, just serves up 
the codepage number specified by registry value

Computer\HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage@ACP

This means that "Windows ANSI" is a pretty dynamic thing. Not just 
system-dependent, but at-the-moment-configuration dependent. Though in 
English-speaking countries it's Windows 1252 by default.

And that in turn means that using the defaults with Visual C++, you can 
end up with pretty much any encoding whatsoever of narrow literals.

Which means that it's a good idea to take charge.

Option `/utf8` is one way to take charge.


- Alf

Back to comp.lang.c++ | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

Character encoding conversion in wide string literals Juha Nieminen <nospam@thanks.invalid> - 2021-12-07 11:37 +0000
  Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 16:44 +0100
    Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:48 -0500
      Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 18:59 +0100
        Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
        Re: Character encoding conversion in wide string literals Öö Tiib <ootiib@hot.ee> - 2021-12-07 16:39 -0800
        Re: Character encoding conversion in wide string literals David Brown <david.brown@hesbynett.no> - 2021-12-08 09:18 +0100
    Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 09:41 -0800
      Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 19:07 +0100
        Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:21 -0800
      Re: Character encoding conversion in wide string literals Paavo Helde <eesnimi@osa.pri.ee> - 2021-12-07 20:32 +0200
        Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
        Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:26 -0800
  Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:38 -0500
    Re: Character encoding conversion in wide string literals Manfred <noname@add.invalid> - 2021-12-07 19:07 +0100

csiph-web