Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #80571 > unrolled thread
| Started by | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| First post | 2021-06-30 07:57 +0000 |
| Last post | 2021-07-01 15:44 +0200 |
| Articles | 20 on this page of 70 — 19 participants |
Back to article view | Back to comp.lang.c++
How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 07:57 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 10:05 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:30 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 12:37 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 09:19 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 11:10 +0200
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 11:36 +0200
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 16:17 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 10:13 -0700
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 19:27 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 11:21 -0700
Re: How to write wide char string literals? Real Troll <real.troll@trolls.com> - 2021-07-01 18:45 +0000
Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-02 10:03 +0200
Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-02 02:31 -0700
Re: How to write wide char string literals? MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk - 2021-07-02 09:38 +0000
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 13:52 +0300
Re: How to write wide char string literals? MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net - 2021-07-02 12:53 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:22 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 17:32 +0300
Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:08 +0200
Re: How to write wide char string literals? MrSpud_r5j@ywn9entw2s.org - 2021-06-30 08:39 +0000
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 10:52 +0000
Re: How to write wide char string literals? MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov - 2021-06-30 11:28 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 14:01 +0200
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 10:55 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 11:23 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 06:51 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:09 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:35 -0400
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 14:03 +0300
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:13 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:39 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 13:50 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-06-30 14:19 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-01 13:31 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 04:42 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-01 10:58 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 05:26 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 01:44 -0400
Re: How to write wide char string literals? Bo Persson <bo@bo-persson.se> - 2021-07-02 08:33 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 11:52 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 16:15 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 03:30 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 00:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 13:31 +0200
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 08:44 -0400
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 16:28 +0200
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 11:48 -0400
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 15:56 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 16:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 17:49 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:11 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 05:52 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 06:59 +0000
Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 09:03 -0400
Re: How to write wide char string literals? Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2021-07-03 14:23 +0100
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 17:06 +0000
Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 14:06 -0400
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:13 +0000
Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:37 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-06-30 16:54 +0200
Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 17:56 +0300
Re: How to write wide char string literals? Öö Tiib <ootiib@hot.ee> - 2021-06-30 09:22 -0700
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 07:19 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 10:29 +0200
Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 08:44 +0000
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 12:58 +0200
Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 14:01 +0200
Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 14:57 +0200
Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-07-01 15:44 +0200
Page 1 of 4 [1] 2 3 4 Next page →
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2021-06-30 07:57 +0000 |
| Subject | How to write wide char string literals? |
| Message-ID | <sbh84o$m8t$1@gioia.aioe.org> |
Character encoding was a problem in the 1960's, and it's still a problem
today, no matter how much computers advance. Sheesh.
Problem is, how to reliably write wide char string literals that contain
non-ascii characters?
Suppose you write for example this:
const wchar_t* str = L"???";
In the *source code* that string literal may be eg. UTF-8 encoded. However,
the compiler needs to convert it to wide chars.
Problem is, how does the compiler know which encoding is being used in
that 8-bit string literal in the source code, in order for it to convert
it properly to wide chars?
Some compilers may assume it's UTF-8 encoded source code. Others may
assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio).
Obviously the end result will be garbage if the wrong assumption is made.
In most compilers (such as Visual Studio) you can specify which encoding
to assume for source files, but this has to be done at the project
settings level. I don't think there's any way to specify the encoding
in the source code itself.
What does the C++ standard say? Does it say that source code files are
always UTF-8 encoded, or is it up to the implementation? I assume that if
it's the latter, the standard doesn't provide any mechanism to specify
which encoding is being used. Or does it?
[toc] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-06-30 10:05 +0200 |
| Message-ID | <sbh8kj$ucq$1@gioia.aioe.org> |
| In reply to | #80571 |
Am 30.06.2021 um 09:57 schrieb Juha Nieminen: > Character encoding was a problem in the 1960's, and it's still a problem > today, no matter how much computers advance. Sheesh. > > Problem is, how to reliably write wide char string literals that contain > non-ascii characters? > > Suppose you write for example this: > > const wchar_t* str = L"???"; > > In the *source code* that string literal may be eg. UTF-8 encoded. However, > the compiler needs to convert it to wide chars. > > Problem is, how does the compiler know which encoding is being used in > that 8-bit string literal in the source code, in order for it to convert > it properly to wide chars? > > Some compilers may assume it's UTF-8 encoded source code. Others may > assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio). > Obviously the end result will be garbage if the wrong assumption is made. > > In most compilers (such as Visual Studio) you can specify which encoding > to assume for source files, but this has to be done at the project > settings level. I don't think there's any way to specify the encoding > in the source code itself. > > What does the C++ standard say? Does it say that source code files are > always UTF-8 encoded, or is it up to the implementation? I assume that if > it's the latter, the standard doesn't provide any mechanism to specify > which encoding is being used. Or does it? Use UTF-16 sourcefiles.
[toc] | [prev] | [next] | [standalone]
| From | Ralf Goertz <me@myprovider.invalid> |
|---|---|
| Date | 2021-06-30 10:30 +0200 |
| Message-ID | <sbha3o$nhi$2@dont-email.me> |
| In reply to | #80572 |
Am Wed, 30 Jun 2021 10:05:40 +0200 schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > Am 30.06.2021 um 09:57 schrieb Juha Nieminen: > > What does the C++ standard say? Does it say that source code files > > are always UTF-8 encoded, or is it up to the implementation? I > > assume that if it's the latter, the standard doesn't provide any > > mechanism to specify which encoding is being used. Or does it? > > Use UTF-16 sourcefiles. That doesn't help with gcc. Even if you specify the encoding on the command line with -finput-charset=utf16be you run into trouble since then gcc assumes the include files (even those included implicitely) are assumed to be utf16be.
[toc] | [prev] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-06-30 12:37 +0200 |
| Message-ID | <sbhhi2$14hc$1@gioia.aioe.org> |
| In reply to | #80574 |
Am 30.06.2021 um 10:30 schrieb Ralf Goertz: > Am Wed, 30 Jun 2021 10:05:40 +0200 > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > >> Am 30.06.2021 um 09:57 schrieb Juha Nieminen: >>> What does the C++ standard say? Does it say that source code files >>> are always UTF-8 encoded, or is it up to the implementation? I >>> assume that if it's the latter, the standard doesn't provide any >>> mechanism to specify which encoding is being used. Or does it? >> >> Use UTF-16 sourcefiles. > > That doesn't help with gcc. Even if you specify the encoding on the > command line with -finput-charset=utf16be you run into trouble since > then gcc assumes the include files (even those included implicitely) > are assumed to be utf16be. UTF-16-files have a byte-header which helps the compiler to distinguish ASCII-files and UTF-16-files.
[toc] | [prev] | [next] | [standalone]
| From | Ralf Goertz <me@myprovider.invalid> |
|---|---|
| Date | 2021-07-01 09:19 +0200 |
| Message-ID | <sbjqac$jc7$1@dont-email.me> |
| In reply to | #80578 |
Am Wed, 30 Jun 2021 12:37:55 +0200 schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > Am 30.06.2021 um 10:30 schrieb Ralf Goertz: > > Am Wed, 30 Jun 2021 10:05:40 +0200 > > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > > > >> Am 30.06.2021 um 09:57 schrieb Juha Nieminen: > >>> What does the C++ standard say? Does it say that source code files > >>> are always UTF-8 encoded, or is it up to the implementation? I > >>> assume that if it's the latter, the standard doesn't provide any > >>> mechanism to specify which encoding is being used. Or does it? > >> > >> Use UTF-16 sourcefiles. > > > > That doesn't help with gcc. Even if you specify the encoding on the > > command line with -finput-charset=utf16be you run into trouble since > > then gcc assumes the include files (even those included implicitely) > > are assumed to be utf16be. > > UTF-16-files have a byte-header which helps the compiler to > distinguish ASCII-files and UTF-16-files. I know that. It's called a byte order mark. And gcc ignores it.
[toc] | [prev] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-07-01 11:10 +0200 |
| Message-ID | <sbk0qv$1s4h$1@gioia.aioe.org> |
| In reply to | #80595 |
Am 01.07.2021 um 09:19 schrieb Ralf Goertz: > Am Wed, 30 Jun 2021 12:37:55 +0200 > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > >> Am 30.06.2021 um 10:30 schrieb Ralf Goertz: >>> Am Wed, 30 Jun 2021 10:05:40 +0200 >>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: >>> >>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen: >>>>> What does the C++ standard say? Does it say that source code files >>>>> are always UTF-8 encoded, or is it up to the implementation? I >>>>> assume that if it's the latter, the standard doesn't provide any >>>>> mechanism to specify which encoding is being used. Or does it? >>>> >>>> Use UTF-16 sourcefiles. >>> >>> That doesn't help with gcc. Even if you specify the encoding on the >>> command line with -finput-charset=utf16be you run into trouble since >>> then gcc assumes the include files (even those included implicitely) >>> are assumed to be utf16be. >> >> UTF-16-files have a byte-header which helps the compiler to >> distinguish ASCII-files and UTF-16-files. > > I know that. It's called a byte order mark. And gcc ignores it. No, wrong - gcc honors it since vesion 1.01.
[toc] | [prev] | [next] | [standalone]
| From | Ralf Goertz <me@myprovider.invalid> |
|---|---|
| Date | 2021-07-01 11:36 +0200 |
| Message-ID | <sbk2au$7pf$1@dont-email.me> |
| In reply to | #80598 |
Am Thu, 1 Jul 2021 11:10:55 +0200
schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> Am 01.07.2021 um 09:19 schrieb Ralf Goertz:
> > Am Wed, 30 Jun 2021 12:37:55 +0200
> > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> >
> >> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
> >>> Am Wed, 30 Jun 2021 10:05:40 +0200
> >>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> >>>
> >>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
> >>>>> What does the C++ standard say? Does it say that source code
> >>>>> files are always UTF-8 encoded, or is it up to the
> >>>>> implementation? I assume that if it's the latter, the standard
> >>>>> doesn't provide any mechanism to specify which encoding is
> >>>>> being used. Or does it?
> >>>>
> >>>> Use UTF-16 sourcefiles.
> >>>
> >>> That doesn't help with gcc. Even if you specify the encoding on
> >>> the command line with -finput-charset=utf16be you run into
> >>> trouble since then gcc assumes the include files (even those
> >>> included implicitely) are assumed to be utf16be.
> >>
> >> UTF-16-files have a byte-header which helps the compiler to
> >> distinguish ASCII-files and UTF-16-files.
> >
> > I know that. It's called a byte order mark. And gcc ignores it.
>
> No, wrong - gcc honors it since vesion 1.01.
I created this file b.cc:
int main() {
return 0;
}
using vi with
:set fileencoding=utf16
:set bomb
Then
~/c> file b.cc
b.cc: C source, Unicode text, UTF-16, big-endian text
or
~/c> od -h b.cc
0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
0000040 2000 2000 7200 6500 7400 7500 7200 6e00
0000060 2000 3000 3b00 0a00 7d00 0a00
0000074
where you can see the BOM fffe. Feeding this to gcc (or g++) you get:
~/c> gcc b.cc
b.cc:1:1: error: stray ‘\376’ in program
1 | �� i n t m a i n ( ) {
| ^
b.cc:1:2: error: stray ‘\377’ in program
1 | �� i n t m a i n ( ) {
| ^
b.cc:1:3: warning: null character(s) ignored
1 | �� i n t m a i n ( ) {
| ^
b.cc:1:5: warning: null character(s) ignored
etc.
How does that qualify as “gcc honoring the BOM”?
[toc] | [prev] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-07-01 16:17 +0200 |
| Message-ID | <sbkipq$ifi$1@gioia.aioe.org> |
| In reply to | #80599 |
Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
> Am Thu, 1 Jul 2021 11:10:55 +0200
> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>
>> Am 01.07.2021 um 09:19 schrieb Ralf Goertz:
>>> Am Wed, 30 Jun 2021 12:37:55 +0200
>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>>
>>>> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
>>>>> Am Wed, 30 Jun 2021 10:05:40 +0200
>>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>>>>
>>>>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
>>>>>>> What does the C++ standard say? Does it say that source code
>>>>>>> files are always UTF-8 encoded, or is it up to the
>>>>>>> implementation? I assume that if it's the latter, the standard
>>>>>>> doesn't provide any mechanism to specify which encoding is
>>>>>>> being used. Or does it?
>>>>>>
>>>>>> Use UTF-16 sourcefiles.
>>>>>
>>>>> That doesn't help with gcc. Even if you specify the encoding on
>>>>> the command line with -finput-charset=utf16be you run into
>>>>> trouble since then gcc assumes the include files (even those
>>>>> included implicitely) are assumed to be utf16be.
>>>>
>>>> UTF-16-files have a byte-header which helps the compiler to
>>>> distinguish ASCII-files and UTF-16-files.
>>>
>>> I know that. It's called a byte order mark. And gcc ignores it.
>>
>> No, wrong - gcc honors it since vesion 1.01.
>
> I created this file b.cc:
>
> int main() {
> return 0;
> }
>
> using vi with
>
> :set fileencoding=utf16
> :set bomb
>
> Then
>
> ~/c> file b.cc
> b.cc: C source, Unicode text, UTF-16, big-endian text
>
> or
>
> ~/c> od -h b.cc
> 0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
> 0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
> 0000040 2000 2000 7200 6500 7400 7500 7200 6e00
> 0000060 2000 3000 3b00 0a00 7d00 0a00
> 0000074
>
> where you can see the BOM fffe. Feeding this to gcc (or g++) you get:
>
> ~/c> gcc b.cc
> b.cc:1:1: error: stray ‘\376’ in program
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:2: error: stray ‘\377’ in program
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:3: warning: null character(s) ignored
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:5: warning: null character(s) ignored
>
> etc.
>
> How does that qualify as “gcc honoring the BOM”?
>
Use v1.01.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2021-07-01 10:13 -0700 |
| Message-ID | <871r8i9cu4.fsf@nosuchdomain.example.com> |
| In reply to | #80605 |
Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>> Am Thu, 1 Jul 2021 11:10:55 +0200
>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
[...]
>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>
>>> No, wrong - gcc honors it since vesion 1.01.
>> I created this file b.cc:
[...]
>> How does that qualify as “gcc honoring the BOM”?
>
> Use v1.01.
You think you're funny. You're not.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-07-01 19:27 +0200 |
| Message-ID | <sbktt9$1umt$1@gioia.aioe.org> |
| In reply to | #80607 |
Am 01.07.2021 um 19:13 schrieb Keith Thompson: > Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes: >> Am 01.07.2021 um 11:36 schrieb Ralf Goertz: >>> Am Thu, 1 Jul 2021 11:10:55 +0200 >>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: > [...] >>>>> I know that. It's called a byte order mark. And gcc ignores it. >>>> >>>> No, wrong - gcc honors it since vesion 1.01. >>> I created this file b.cc: > [...] >>> How does that qualify as “gcc honoring the BOM”? >> >> Use v1.01. > > You think you're funny. You're not. v1.01 does honor the BOM.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2021-07-01 11:21 -0700 |
| Message-ID | <87wnq999ok.fsf@nosuchdomain.example.com> |
| In reply to | #80608 |
Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
> Am 01.07.2021 um 19:13 schrieb Keith Thompson:
>> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
>>> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>>>> Am Thu, 1 Jul 2021 11:10:55 +0200
>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>> [...]
>>>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>>>
>>>>> No, wrong - gcc honors it since vesion 1.01.
>>>> I created this file b.cc:
>> [...]
>>>> How does that qualify as “gcc honoring the BOM”?
>>>
>>> Use v1.01.
>> You think you're funny. You're not.
>
> v1.01 does honor the BOM.
At the risk of giving the impression I'm taking you seriously, the
oldest version of gcc available from gnu.org is 1.42, released in 1992.
I've see no evidence that gcc v1.01 would have honored the BOM, but it
doesn't matter, since that version is obsolete and unavailable.
I conclude that you are a troll.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Real Troll <real.troll@trolls.com> |
|---|---|
| Date | 2021-07-01 18:45 +0000 |
| Message-ID | <sbl2k3$386$1@gioia.aioe.org> |
| In reply to | #80609 |
On 01/07/2021 19:21, Keith Thompson wrote: > > I conclude that you are a troll. > It takes one troll to know another. This is an expert opinion of a Real Troll!
[toc] | [prev] | [next] | [standalone]
| From | Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> |
|---|---|
| Date | 2021-07-02 10:03 +0200 |
| Message-ID | <sbmh8m$1lbb$1@gioia.aioe.org> |
| In reply to | #80609 |
Am 01.07.2021 um 20:21 schrieb Keith Thompson: > Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes: >> Am 01.07.2021 um 19:13 schrieb Keith Thompson: >>> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes: >>>> Am 01.07.2021 um 11:36 schrieb Ralf Goertz: >>>>> Am Thu, 1 Jul 2021 11:10:55 +0200 >>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>: >>> [...] >>>>>>> I know that. It's called a byte order mark. And gcc ignores it. >>>>>> >>>>>> No, wrong - gcc honors it since vesion 1.01. >>>>> I created this file b.cc: >>> [...] >>>>> How does that qualify as “gcc honoring the BOM”? >>>> >>>> Use v1.01. >>> You think you're funny. You're not. >> >> v1.01 does honor the BOM. > > At the risk of giving the impression I'm taking you seriously, the > oldest version of gcc available from gnu.org is 1.42, released in 1992. No, the oldest gcc-version is v0.9 from March 22, 1987. > > I've see no evidence that gcc v1.01 would have honored the BOM, but it > doesn't matter, since that version is obsolete and unavailable. > > I conclude that you are a troll. >
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2021-07-02 02:31 -0700 |
| Message-ID | <87sg0x83kj.fsf@nosuchdomain.example.com> |
| In reply to | #80599 |
Ralf Goertz <me@myprovider.invalid> writes:
[...]
> I created this file b.cc:
>
> int main() {
> return 0;
> }
>
> using vi with
>
> :set fileencoding=utf16
> :set bomb
>
> Then
>
> ~/c> file b.cc
> b.cc: C source, Unicode text, UTF-16, big-endian text
>
> or
>
> ~/c> od -h b.cc
> 0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
> 0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
> 0000040 2000 2000 7200 6500 7400 7500 7200 6e00
> 0000060 2000 3000 3b00 0a00 7d00 0a00
> 0000074
>
> where you can see the BOM fffe. Feeding this to gcc (or g++) you get:
>
> ~/c> gcc b.cc
> b.cc:1:1: error: stray ‘\376’ in program
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:2: error: stray ‘\377’ in program
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:3: warning: null character(s) ignored
> 1 | �� i n t m a i n ( ) {
> | ^
> b.cc:1:5: warning: null character(s) ignored
>
> etc.
>
> How does that qualify as “gcc honoring the BOM”?
On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
(I don't know whether there's a way to configure it to do so.)
It does handle UTF-8 with or without a BOM.
$ file b.cpp
b.cpp: C source, UTF-8 Unicode (with BOM) text
$ cat b.cpp
int main() { }
$ hd b.cpp
00000000 ef bb bf 69 6e 74 20 6d 61 69 6e 28 29 20 7b 20 |...int main() { |
00000010 7d 0a |}.|
00000012
$ gcc -c b.cpp
$
gcc 9.3.0 on Ubuntu 20.04. (There is, of course, no point in going back
to ancient versions of gcc.)
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk |
|---|---|
| Date | 2021-07-02 09:38 +0000 |
| Message-ID | <sbmmqi$akd$1@gioia.aioe.org> |
| In reply to | #80620 |
On Fri, 02 Jul 2021 02:31:24 -0700 Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >Ralf Goertz <me@myprovider.invalid> writes: >[...] > >On my system, gcc doesn't handle UTF-16 at all, with or without a BOM. >(I don't know whether there's a way to configure it to do so.) Just out of interest, what byte order is the BOM in? Catch 22?
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <myfirstname@osa.pri.ee> |
|---|---|
| Date | 2021-07-02 13:52 +0300 |
| Message-ID | <sbmr69$tt1$1@dont-email.me> |
| In reply to | #80621 |
02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas: > On Fri, 02 Jul 2021 02:31:24 -0700 > Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >> Ralf Goertz <me@myprovider.invalid> writes: >> [...] >> >> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM. >> (I don't know whether there's a way to configure it to do so.) > > Just out of interest, what byte order is the BOM in? Catch 22? > This is probably a troll question, but answering anyway: the BOM marker U+FEFF is in the correct byte order, in both little-endian and big-endian UTF-16. That's how you tell them apart. The trick is in that the reverse value U+FFFE is not valid Unicode character, so there is no possibility of mixup. Neither sequence is also valid UTF-8, to avoid mixup with that.
[toc] | [prev] | [next] | [standalone]
| From | MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net |
|---|---|
| Date | 2021-07-02 12:53 +0000 |
| Message-ID | <sbn27o$1j5a$1@gioia.aioe.org> |
| In reply to | #80622 |
On Fri, 2 Jul 2021 13:52:53 +0300 Paavo Helde <myfirstname@osa.pri.ee> wrote: >02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas: >> On Fri, 02 Jul 2021 02:31:24 -0700 >> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >>> Ralf Goertz <me@myprovider.invalid> writes: >>> [...] >>> >>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM. >>> (I don't know whether there's a way to configure it to do so.) >> >> Just out of interest, what byte order is the BOM in? Catch 22? >> > >This is probably a troll question, but answering anyway: the BOM marker No, not a troll. >U+FEFF is in the correct byte order, in both little-endian and >big-endian UTF-16. That's how you tell them apart. So its just 2 bytes in sequence, not a 16 bit value?
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2021-07-02 15:22 +0200 |
| Message-ID | <sbn3u7$qti$1@dont-email.me> |
| In reply to | #80624 |
On 2 Jul 2021 14:53, MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net wrote:
> On Fri, 2 Jul 2021 13:52:53 +0300
> Paavo Helde <myfirstname@osa.pri.ee> wrote:
>> 02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas:
>>> On Fri, 02 Jul 2021 02:31:24 -0700
>>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>>> Ralf Goertz <me@myprovider.invalid> writes:
>>>> [...]
>>>>
>>>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>>>> (I don't know whether there's a way to configure it to do so.)
>>>
>>> Just out of interest, what byte order is the BOM in? Catch 22?
>>>
>>
>> This is probably a troll question, but answering anyway: the BOM marker
>
> No, not a troll.
>
>> U+FEFF is in the correct byte order, in both little-endian and
>> big-endian UTF-16. That's how you tell them apart.
>
> So its just 2 bytes in sequence, not a 16 bit value?
The BOM is a Unicode code point, U+FEFF as Paavo mentioned, originally
standing for an invisible zero-width hard space. It's encoded with
either little endian UTF-16 (then as two bytes), or as big endian UTF-16
(then as two bytes), or as endianness agnostic UTF-8 (then as three
bytes). The encoded BOM yields a reliable encoding indicator, though
pedantic people might argue that it's just statistical -- after all, one
just might happen to have a Windows 1252 encoded file with three
characters at the start with the same byte values as the UTF-8 BOM.
In the same vein, one just might happen to have a `.txt` file in Windows
with the letters "MZ" at the very start, like
MZ, Mishtara Zva'it, is the Military Police Corps of Israel. blah
and if you then try to open the file in your default text editor by just
typing the file name in old Cmd, those letters will be misinterpreted as
the initials of Mark Zbikowski, marking the file as an executable...
Since the chance of that happening isn't absolutely 0 one should never
use text file names as commands, or the UTF-8 BOM as an encoding marker.
- Alf
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <myfirstname@osa.pri.ee> |
|---|---|
| Date | 2021-07-02 17:32 +0300 |
| Message-ID | <sbn81a$o77$1@dont-email.me> |
| In reply to | #80624 |
02.07.2021 15:53 MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net kirjutas: > On Fri, 2 Jul 2021 13:52:53 +0300 > Paavo Helde <myfirstname@osa.pri.ee> wrote: >> 02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas: >>> On Fri, 02 Jul 2021 02:31:24 -0700 >>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >>>> Ralf Goertz <me@myprovider.invalid> writes: >>>> [...] >>>> >>>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM. >>>> (I don't know whether there's a way to configure it to do so.) >>> >>> Just out of interest, what byte order is the BOM in? Catch 22? >>> >> >> This is probably a troll question, but answering anyway: the BOM marker > > No, not a troll. > >> U+FEFF is in the correct byte order, in both little-endian and >> big-endian UTF-16. That's how you tell them apart. > > So its just 2 bytes in sequence, not a 16 bit value? The file contains bytes, it's up to the reading code how to interpret the bytes. It can interpret the bytes as uint16_t, i.e. cast the file buffer as 'const uint16_t*' and read the first 2-byte value. If it is 0xFEFF, then it knows this is an UTF-16 file in a matching byte order. If it is 0xFFFE, then it knows it's an UTF-16 file in an opposite byte order, and the rest of the file needs to be byte-swapped. It can also interpret the buffer as containing uint8_t bytes, but then the logic is a bit more complex, it must then know if it itself is running on a big-endian or little-endian machine, and behave accordingly.
[toc] | [prev] | [next] | [standalone]
| From | Ralf Goertz <me@myprovider.invalid> |
|---|---|
| Date | 2021-06-30 10:08 +0200 |
| Message-ID | <sbh8q7$nhi$1@dont-email.me> |
| In reply to | #80571 |
Am Wed, 30 Jun 2021 07:57:14 +0000 (UTC) schrieb Juha Nieminen <nospam@thanks.invalid>: > In most compilers (such as Visual Studio) you can specify which > encoding to assume for source files, but this has to be done at the > project settings level. I don't think there's any way to specify the > encoding in the source code itself. When using UTF-encoding there is always A BOM you could use. Doesn't help much with iso encoding, though. And I also just found out that gcc doesn't notice that a source file with an appropriate byte order mark is encoded in utf32 BE. That's a bit disappointing.
[toc] | [prev] | [next] | [standalone]
Page 1 of 4 [1] 2 3 4 Next page →
Back to top | Article view | comp.lang.c++
csiph-web