Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #82565
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Newsgroups | comp.lang.c++ |
| Subject | Re: Character encoding conversion in wide string literals |
| Date | 2021-12-07 16:44 +0100 |
| Organization | A noiseless patient Spider |
| Message-ID | <sonvh8$vc9$1@dont-email.me> (permalink) |
| References | <sonh15$g9j$1@gioia.aioe.org> |
On 7 Dec 2021 12:37, Juha Nieminen wrote: > Recently I stumbled across a problem where I had wide string literals > with non-ascii characters UTF-8 encoded. In other words, I had code like > this (I'm using non-ascii in the code below, I hope it doesn't get > mangled up, but even if it does, it should nevertheless be clear what > I'm trying to express): > > std::wstring str = L"non-ascii chars: ???"; > > The C++ source file itself uses UTF-8 encoding, meaning that that line > of code is likewise UTF-8 encoded. If it were a narrow string literal > (being assigned to a std::string) then it works just fine (primarily > because the compiler doesn't need to do anything to it, it can simply > take those bytes from the source file as is). > > However, since it's a wide string literal (being assigned to a std::wstring) > it's not as clear-cut anymore. What does the standard say about this > situation? > > The thing is that it works just fine in Linux using gcc. The compiler will > re-encode the UTF-8 encoded characters in the source file inside the > parentheses into whatever encoding wide char string use, so the correct > content will end up in the executable binary (and thus in the wstring). > > Apparently it does not work correctly in (some recent version of) > Visual Studio, where apparently it just takes the byte values from the > source file within the parentheses as-is, and just assigns those values > as-is to the wide chars that end up in the binary. (Or something like that.) The Visual C++ compiler assumes that source code is Windows ANSI encoded unless * you use an encoding option such as `/utf8`, or * the source is UTF-8 with BOM, or * the source is UTF-16. Independently of that Visual C++ assumes that the execution character set (the byte-based encoding that should be used for text data in the executable) is Windows ANSI, unless it's specified as something else. The `/utf8` option specifies also that. It's a combo option that specifies both source encoding and execution character set as UTF-8. Unfortunately as of VS 2022 `/utf8` is not set by default in a VS project, and unfortunately there's nothing you can just click to set it. You have to to type it in (right click project then) Properties -> C/C++ -> Command Line. I usually set "/utf-8 /Zc:__cplusplus". > Does the standard specify what the compiler should do in this situation? > If not, then what is the proper way of specifying wide string literals > that contain non-ascii characters? I'll let others discuss that, but (1) it does, and (2) just so you're aware: the main problem is that the C and C++ standards do not conform to reality in their requirement that a `wchar_t` value should suffice to encode all possible code points in the wide character set. In Windows wide text is UTF-16, with 16-bit `wchar_t`. Which means that some emojis etc. that appear as a single character and constitute one 21-bit code point, can become a pair of two `wchar_t` values, an UTF-16 "surrogate pair". That's probably not your problem though, but it is a/the problem. - ALf
Back to comp.lang.c++ | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
Character encoding conversion in wide string literals Juha Nieminen <nospam@thanks.invalid> - 2021-12-07 11:37 +0000
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 16:44 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:48 -0500
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 18:59 +0100
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Öö Tiib <ootiib@hot.ee> - 2021-12-07 16:39 -0800
Re: Character encoding conversion in wide string literals David Brown <david.brown@hesbynett.no> - 2021-12-08 09:18 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 09:41 -0800
Re: Character encoding conversion in wide string literals "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-12-07 19:07 +0100
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:21 -0800
Re: Character encoding conversion in wide string literals Paavo Helde <eesnimi@osa.pri.ee> - 2021-12-07 20:32 +0200
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 13:55 -0500
Re: Character encoding conversion in wide string literals Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-07 12:26 -0800
Re: Character encoding conversion in wide string literals James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-07 11:38 -0500
Re: Character encoding conversion in wide string literals Manfred <noname@add.invalid> - 2021-12-07 19:07 +0100
csiph-web