Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #82571
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Newsgroups | comp.lang.c++ |
| Subject | Re: Character encoding conversion in wide string literals |
| Date | 2021-12-07 19:07 +0100 |
| Organization | A noiseless patient Spider |
| Message-ID | <soo7s7$uce$1@dont-email.me> (permalink) |
| References | <sonh15$g9j$1@gioia.aioe.org> <sonvh8$vc9$1@dont-email.me> <87r1aob9d3.fsf@nosuchdomain.example.com> |
On 7 Dec 2021 18:41, Keith Thompson wrote: > "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes: >> On 7 Dec 2021 12:37, Juha Nieminen wrote: >>> Recently I stumbled across a problem where I had wide string literals >>> with non-ascii characters UTF-8 encoded. In other words, I had code like >>> this (I'm using non-ascii in the code below, I hope it doesn't get >>> mangled up, but even if it does, it should nevertheless be clear what >>> I'm trying to express): >>> std::wstring str = L"non-ascii chars: ???"; >>> The C++ source file itself uses UTF-8 encoding, meaning that that >>> line >>> of code is likewise UTF-8 encoded. If it were a narrow string literal >>> (being assigned to a std::string) then it works just fine (primarily >>> because the compiler doesn't need to do anything to it, it can simply >>> take those bytes from the source file as is). >>> However, since it's a wide string literal (being assigned to a >>> std::wstring) >>> it's not as clear-cut anymore. What does the standard say about this >>> situation? >>> The thing is that it works just fine in Linux using gcc. The >>> compiler will >>> re-encode the UTF-8 encoded characters in the source file inside the >>> parentheses into whatever encoding wide char string use, so the correct >>> content will end up in the executable binary (and thus in the wstring). >>> Apparently it does not work correctly in (some recent version of) >>> Visual Studio, where apparently it just takes the byte values from the >>> source file within the parentheses as-is, and just assigns those values >>> as-is to the wide chars that end up in the binary. (Or something like that.) >> >> The Visual C++ compiler assumes that source code is Windows ANSI >> encoded unless >> >> * you use an encoding option such as `/utf8`, or >> * the source is UTF-8 with BOM, or >> * the source is UTF-16. > > What exactly do you mean by "Windows ANSI"? Windows-1252 or something > else? (Microsoft doesn't call it "ANSI", because it isn't.) > > [...] "Windows ANSI" is the encoding specified by the `GetACP` API function, which, but as I recall that's more or less undocumented, just serves up the codepage number specified by registry value Computer\HKEY_LOCAL_MACHINE\SYSTEM\CurrentControlSet\Control\Nls\CodePage@ACP This means that "Windows ANSI" is a pretty dynamic thing. Not just system-dependent, but at-the-moment-configuration dependent. Though in English-speaking countries it's Windows 1252 by default. And that in turn means that using the defaults with Visual C++, you can end up with pretty much any encoding whatsoever of narrow literals. Which means that it's a good idea to take charge. Option `/utf8` is one way to take charge. - Alf
Back to comp.lang.c++ | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Character encoding conversion in wide string literals Juha Nieminen <nospam@thanks.invalid> - 2021-12-07 11:37 +0000
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 16:44 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:48 -0500
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 18:59 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Öö Tiib <ootiib@hot.ee> - 2021-12-07 16:39 -0800
Re: Character encoding conversion in wide string literals David Brown <david.brown@hesbynett.no> - 2021-12-08 09:18 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 09:41 -0800
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 19:07 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:21 -0800
Re: Character encoding conversion in wide string literals Paavo Helde <eesnimi@osa.pri.ee> - 2021-12-07 20:32 +0200
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:26 -0800
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:38 -0500
Re: Character encoding conversion in wide string literals Manfred <noname@add.invalid> - 2021-12-07 19:07 +0100
csiph-web