Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #82571
| Path | csiph.com!eternal-september.org!reader02.eternal-september.org!.POSTED!not-for-mail |
|---|---|
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
| Newsgroups | comp.lang.c++ |
| Subject | Re: Character encoding conversion in wide string literals |
| Date | Tue, 7 Dec 2021 19:07:00 +0100 |
| Organization | A noiseless patient Spider |
| Lines | 61 |
| Message-ID | <soo7s7$uce$1@dont-email.me> (permalink) |
| References | <sonh15$g9j$1@gioia.aioe.org> <sonvh8$vc9$1@dont-email.me> <87r1aob9d3.fsf@nosuchdomain.example.com> |
| Mime-Version | 1.0 |
| Content-Type | text/plain; charset=UTF-8; format=flowed |
| Content-Transfer-Encoding | 7bit |
| Injection-Date | Tue, 7 Dec 2021 18:07:03 -0000 (UTC) |
| Injection-Info | reader02.eternal-september.org; posting-host="082f8b1fdbf905e3e36c03435b0eccad"; logging-data="31118"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX19oRJuRW0cBnf2ldFK5Y/nQ" |
| User-Agent | Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:91.0) Gecko/20100101 Thunderbird/91.3.2 |
| Cancel-Lock | sha1:8/JCOGMeDMC2dwTMe6K+je3L8qs= |
| In-Reply-To | <87r1aob9d3.fsf@nosuchdomain.example.com> |
| Content-Language | en-US |
| Xref | csiph.com comp.lang.c++:82571 |
Show key headers only | View raw
On 7 Dec 2021 18:41, Keith Thompson wrote: > "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes: >> On 7 Dec 2021 12:37, Juha Nieminen wrote: >>> Recently I stumbled across a problem where I had wide string literals >>> with non-ascii characters UTF-8 encoded. In other words, I had code like >>> this (I'm using non-ascii in the code below, I hope it doesn't get >>> mangled up, but even if it does, it should nevertheless be clear what >>> I'm trying to express): >>> std::wstring str = L"non-ascii chars: ???"; >>> The C++ source file itself uses UTF-8 encoding, meaning that that >>> line >>> of code is likewise UTF-8 encoded. If it were a narrow string literal >>> (being assigned to a std::string) then it works just fine (primarily >>> because the compiler doesn't need to do anything to it, it can simply >>> take those bytes from the source file as is). >>> However, since it's a wide string literal (being assigned to a >>> std::wstring) >>> it's not as clear-cut anymore. What does the standard say about this >>> situation? >>> The thing is that it works just fine in Linux using gcc. The >>> compiler will >>> re-encode the UTF-8 encoded characters in the source file inside the >>> parentheses into whatever encoding wide char string use, so the correct >>> content will end up in the executable binary (and thus in the wstring). >>> Apparently it does not work correctly in (some recent version of) >>> Visual Studio, where apparently it just takes the byte values from the >>> source file within the parentheses as-is, and just assigns those values >>> as-is to the wide chars that end up in the binary. (Or something like that.) >> >> The Visual C++ compiler assumes that source code is Windows ANSI >> encoded unless >> >> * you use an encoding option such as `/utf8`, or >> * the source is UTF-8 with BOM, or >> * the source is UTF-16. > > What exactly do you mean by "Windows ANSI"? Windows-1252 or something > else? (Microsoft doesn't call it "ANSI", because it isn't.) > > [...] "Windows ANSI" is the encoding specified by the `GetACP` API function, which, but as I recall that's more or less undocumented, just serves up the codepage number specified by registry value Computer\HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage@ACP This means that "Windows ANSI" is a pretty dynamic thing. Not just system-dependent, but at-the-moment-configuration dependent. Though in English-speaking countries it's Windows 1252 by default. And that in turn means that using the defaults with Visual C++, you can end up with pretty much any encoding whatsoever of narrow literals. Which means that it's a good idea to take charge. Option `/utf8` is one way to take charge. - Alf
Back to comp.lang.c++ | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Character encoding conversion in wide string literals Juha Nieminen <nospam@thanks.invalid> - 2021-12-07 11:37 +0000
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 16:44 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:48 -0500
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 18:59 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Öö Tiib <ootiib@hot.ee> - 2021-12-07 16:39 -0800
Re: Character encoding conversion in wide string literals David Brown <david.brown@hesbynett.no> - 2021-12-08 09:18 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 09:41 -0800
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 19:07 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:21 -0800
Re: Character encoding conversion in wide string literals Paavo Helde <eesnimi@osa.pri.ee> - 2021-12-07 20:32 +0200
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:26 -0800
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:38 -0500
Re: Character encoding conversion in wide string literals Manfred <noname@add.invalid> - 2021-12-07 19:07 +0100
csiph-web