Path: csiph.com!eternal-september.org!reader02.eternal-september.org!.POSTED!not-for-mail From: "Alf P. Steinbach" Newsgroups: comp.lang.c++ Subject: Re: Character encoding conversion in wide string literals Date: Tue, 7 Dec 2021 19:07:00 +0100 Organization: A noiseless patient Spider Lines: 61 Message-ID: References: <87r1aob9d3.fsf@nosuchdomain.example.com> Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit Injection-Date: Tue, 7 Dec 2021 18:07:03 -0000 (UTC) Injection-Info: reader02.eternal-september.org; posting-host="082f8b1fdbf905e3e36c03435b0eccad"; logging-data="31118"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX19oRJuRW0cBnf2ldFK5Y/nQ" User-Agent: Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.3.2 Cancel-Lock: sha1:8/JCOGMeDMC2dwTMe6K+je3L8qs= In-Reply-To: <87r1aob9d3.fsf@nosuchdomain.example.com> Content-Language: en-US Xref: csiph.com comp.lang.c++:82571 On 7 Dec 2021 18:41, Keith Thompson wrote: > "Alf P. Steinbach" writes: >> On 7 Dec 2021 12:37, Juha Nieminen wrote: >>> Recently I stumbled across a problem where I had wide string literals >>> with non-ascii characters UTF-8 encoded. In other words, I had code like >>> this (I'm using non-ascii in the code below, I hope it doesn't get >>> mangled up, but even if it does, it should nevertheless be clear what >>> I'm trying to express): >>> std::wstring str = L"non-ascii chars: ???"; >>> The C++ source file itself uses UTF-8 encoding, meaning that that >>> line >>> of code is likewise UTF-8 encoded. If it were a narrow string literal >>> (being assigned to a std::string) then it works just fine (primarily >>> because the compiler doesn't need to do anything to it, it can simply >>> take those bytes from the source file as is). >>> However, since it's a wide string literal (being assigned to a >>> std::wstring) >>> it's not as clear-cut anymore. What does the standard say about this >>> situation? >>> The thing is that it works just fine in Linux using gcc. The >>> compiler will >>> re-encode the UTF-8 encoded characters in the source file inside the >>> parentheses into whatever encoding wide char string use, so the correct >>> content will end up in the executable binary (and thus in the wstring). >>> Apparently it does not work correctly in (some recent version of) >>> Visual Studio, where apparently it just takes the byte values from the >>> source file within the parentheses as-is, and just assigns those values >>> as-is to the wide chars that end up in the binary. (Or something like that.) >> >> The Visual C++ compiler assumes that source code is Windows ANSI >> encoded unless >> >> * you use an encoding option such as `/utf8`, or >> * the source is UTF-8 with BOM, or >> * the source is UTF-16. > > What exactly do you mean by "Windows ANSI"? Windows-1252 or something > else? (Microsoft doesn't call it "ANSI", because it isn't.) > > [...] "Windows ANSI" is the encoding specified by the `GetACP` API function, which, but as I recall that's more or less undocumented, just serves up the codepage number specified by registry value Computer\HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage@ACP This means that "Windows ANSI" is a pretty dynamic thing. Not just system-dependent, but at-the-moment-configuration dependent. Though in English-speaking countries it's Windows 1252 by default. And that in turn means that using the defaults with Visual C++, you can end up with pretty much any encoding whatsoever of narrow literals. Which means that it's a good idea to take charge. Option `/utf8` is one way to take charge. - Alf