Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #80571 > unrolled thread

How to write wide char string literals?

Started byJuha Nieminen <nospam@thanks.invalid>
First post2021-06-30 07:57 +0000
Last post2021-07-01 15:44 +0200
Articles 20 on this page of 70 — 19 participants

Back to article view | Back to comp.lang.c++


Contents

  How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 07:57 +0000
    Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 10:05 +0200
      Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:30 +0200
        Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-06-30 12:37 +0200
          Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 09:19 +0200
            Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 11:10 +0200
              Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-07-01 11:36 +0200
                Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 16:17 +0200
                  Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 10:13 -0700
                    Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-01 19:27 +0200
                      Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-01 11:21 -0700
                        Re: How to write wide char string literals? Real Troll <real.troll@trolls.com> - 2021-07-01 18:45 +0000
                        Re: How to write wide char string literals? Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> - 2021-07-02 10:03 +0200
                Re: How to write wide char string literals? Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-07-02 02:31 -0700
                  Re: How to write wide char string literals? MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk - 2021-07-02 09:38 +0000
                    Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 13:52 +0300
                      Re: How to write wide char string literals? MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net - 2021-07-02 12:53 +0000
                        Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:22 +0200
                        Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-07-02 17:32 +0300
    Re: How to write wide char string literals? Ralf Goertz <me@myprovider.invalid> - 2021-06-30 10:08 +0200
    Re: How to write wide char string literals? MrSpud_r5j@ywn9entw2s.org - 2021-06-30 08:39 +0000
      Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 10:52 +0000
        Re: How to write wide char string literals? MrSpud_3u59h8@0c9tv3nddl090w2ynhm.gov - 2021-06-30 11:28 +0000
          Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 14:01 +0200
    Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 10:55 +0200
      Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-06-30 11:23 +0200
    Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 06:51 -0400
      Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:09 +0000
        Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:35 -0400
    Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 14:03 +0300
      Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-06-30 11:13 +0000
        Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-06-30 07:39 -0400
          Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-06-30 13:50 +0200
            Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-06-30 14:19 -0400
              Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-01 13:31 +0200
          Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 04:42 +0000
            Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-01 10:58 -0400
              Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 05:26 +0000
                Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 01:44 -0400
                  Re: How to write wide char string literals? Bo Persson <bo@bo-persson.se> - 2021-07-02 08:33 +0200
                  Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-02 11:52 +0000
                    Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-02 16:15 -0400
                      Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 03:30 +0200
                        Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 00:44 -0400
                          Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 13:31 +0200
                            Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 08:44 -0400
                              Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-03 16:28 +0200
                                Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 11:48 -0400
                                Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 15:56 -0400
                              Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 16:59 +0000
                                Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 17:49 -0400
                                  Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:11 +0000
                                    Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-08 05:52 -0400
                      Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 06:59 +0000
                        Re: How to write wide char string literals? James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-07-03 09:03 -0400
                        Re: How to write wide char string literals? Chris Vine <chris@cvine--nospam--.freeserve.co.uk> - 2021-07-03 14:23 +0100
                          Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-03 17:06 +0000
                            Re: How to write wide char string literals? Richard Damon <Richard@Damon-Family.org> - 2021-07-03 14:06 -0400
                              Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-08 08:13 +0000
                  Re: How to write wide char string literals? "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2021-07-02 15:37 +0200
        Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-06-30 16:54 +0200
        Re: How to write wide char string literals? Paavo Helde <myfirstname@osa.pri.ee> - 2021-06-30 17:56 +0300
    Re: How to write wide char string literals? Öö Tiib <ootiib@hot.ee> - 2021-06-30 09:22 -0700
    Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 07:19 +0200
      Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 10:29 +0200
        Re: How to write wide char string literals? Juha Nieminen <nospam@thanks.invalid> - 2021-07-01 08:44 +0000
          Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 12:58 +0200
        Re: How to write wide char string literals? Christian Gollwitzer <auriocus@gmx.de> - 2021-07-01 14:01 +0200
          Re: How to write wide char string literals? David Brown <david.brown@hesbynett.no> - 2021-07-01 14:57 +0200
      Re: How to write wide char string literals? Manfred <noname@add.invalid> - 2021-07-01 15:44 +0200

Page 1 of 4  [1] 2 3 4  Next page →


#80571 — How to write wide char string literals?

FromJuha Nieminen <nospam@thanks.invalid>
Date2021-06-30 07:57 +0000
SubjectHow to write wide char string literals?
Message-ID<sbh84o$m8t$1@gioia.aioe.org>
Character encoding was a problem in the 1960's, and it's still a problem
today, no matter how much computers advance. Sheesh.

Problem is, how to reliably write wide char string literals that contain
non-ascii characters?

Suppose you write for example this:

    const wchar_t* str = L"???";

In the *source code* that string literal may be eg. UTF-8 encoded. However,
the compiler needs to convert it to wide chars.

Problem is, how does the compiler know which encoding is being used in
that 8-bit string literal in the source code, in order for it to convert
it properly to wide chars?

Some compilers may assume it's UTF-8 encoded source code. Others may
assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio).
Obviously the end result will be garbage if the wrong assumption is made.

In most compilers (such as Visual Studio) you can specify which encoding
to assume for source files, but this has to be done at the project
settings level. I don't think there's any way to specify the encoding
in the source code itself.

What does the C++ standard say? Does it say that source code files are
always UTF-8 encoded, or is it up to the implementation? I assume that if
it's the latter, the standard doesn't provide any mechanism to specify
which encoding is being used. Or does it?

[toc] | [next] | [standalone]


#80572

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-06-30 10:05 +0200
Message-ID<sbh8kj$ucq$1@gioia.aioe.org>
In reply to#80571
Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
> Character encoding was a problem in the 1960's, and it's still a problem
> today, no matter how much computers advance. Sheesh.
> 
> Problem is, how to reliably write wide char string literals that contain
> non-ascii characters?
> 
> Suppose you write for example this:
> 
>      const wchar_t* str = L"???";
> 
> In the *source code* that string literal may be eg. UTF-8 encoded. However,
> the compiler needs to convert it to wide chars.
> 
> Problem is, how does the compiler know which encoding is being used in
> that 8-bit string literal in the source code, in order for it to convert
> it properly to wide chars?
> 
> Some compilers may assume it's UTF-8 encoded source code. Others may
> assume it's ISO-Latin-1 encoded (I'm looking at you, Visual Studio).
> Obviously the end result will be garbage if the wrong assumption is made.
> 
> In most compilers (such as Visual Studio) you can specify which encoding
> to assume for source files, but this has to be done at the project
> settings level. I don't think there's any way to specify the encoding
> in the source code itself.
> 
> What does the C++ standard say? Does it say that source code files are
> always UTF-8 encoded, or is it up to the implementation? I assume that if
> it's the latter, the standard doesn't provide any mechanism to specify
> which encoding is being used. Or does it?

Use UTF-16 sourcefiles.

[toc] | [prev] | [next] | [standalone]


#80574

FromRalf Goertz <me@myprovider.invalid>
Date2021-06-30 10:30 +0200
Message-ID<sbha3o$nhi$2@dont-email.me>
In reply to#80572
Am Wed, 30 Jun 2021 10:05:40 +0200
schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:

> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
> > What does the C++ standard say? Does it say that source code files
> > are always UTF-8 encoded, or is it up to the implementation? I
> > assume that if it's the latter, the standard doesn't provide any
> > mechanism to specify which encoding is being used. Or does it?  
> 
> Use UTF-16 sourcefiles.

That doesn't help with gcc. Even if you specify the encoding on the
command line with -finput-charset=utf16be you run into trouble since
then gcc assumes the include files (even those included implicitely) are
assumed to be utf16be.

[toc] | [prev] | [next] | [standalone]


#80578

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-06-30 12:37 +0200
Message-ID<sbhhi2$14hc$1@gioia.aioe.org>
In reply to#80574
Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
> Am Wed, 30 Jun 2021 10:05:40 +0200
> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> 
>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
>>> What does the C++ standard say? Does it say that source code files
>>> are always UTF-8 encoded, or is it up to the implementation? I
>>> assume that if it's the latter, the standard doesn't provide any
>>> mechanism to specify which encoding is being used. Or does it?
>>
>> Use UTF-16 sourcefiles.
> 
> That doesn't help with gcc. Even if you specify the encoding on the
> command line with -finput-charset=utf16be you run into trouble since
> then gcc assumes the include files (even those included implicitely)
> are assumed to be utf16be.

UTF-16-files have a byte-header which helps the compiler to distinguish
ASCII-files and UTF-16-files.

[toc] | [prev] | [next] | [standalone]


#80595

FromRalf Goertz <me@myprovider.invalid>
Date2021-07-01 09:19 +0200
Message-ID<sbjqac$jc7$1@dont-email.me>
In reply to#80578
Am Wed, 30 Jun 2021 12:37:55 +0200
schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:

> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
> > Am Wed, 30 Jun 2021 10:05:40 +0200
> > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> >   
> >> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:  
> >>> What does the C++ standard say? Does it say that source code files
> >>> are always UTF-8 encoded, or is it up to the implementation? I
> >>> assume that if it's the latter, the standard doesn't provide any
> >>> mechanism to specify which encoding is being used. Or does it?  
> >>
> >> Use UTF-16 sourcefiles.  
> > 
> > That doesn't help with gcc. Even if you specify the encoding on the
> > command line with -finput-charset=utf16be you run into trouble since
> > then gcc assumes the include files (even those included implicitely)
> > are assumed to be utf16be.  
> 
> UTF-16-files have a byte-header which helps the compiler to
> distinguish ASCII-files and UTF-16-files.

I know that. It's called a byte order mark. And gcc ignores it.

[toc] | [prev] | [next] | [standalone]


#80598

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-07-01 11:10 +0200
Message-ID<sbk0qv$1s4h$1@gioia.aioe.org>
In reply to#80595
Am 01.07.2021 um 09:19 schrieb Ralf Goertz:
> Am Wed, 30 Jun 2021 12:37:55 +0200
> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> 
>> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
>>> Am Wed, 30 Jun 2021 10:05:40 +0200
>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>>    
>>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
>>>>> What does the C++ standard say? Does it say that source code files
>>>>> are always UTF-8 encoded, or is it up to the implementation? I
>>>>> assume that if it's the latter, the standard doesn't provide any
>>>>> mechanism to specify which encoding is being used. Or does it?
>>>>
>>>> Use UTF-16 sourcefiles.
>>>
>>> That doesn't help with gcc. Even if you specify the encoding on the
>>> command line with -finput-charset=utf16be you run into trouble since
>>> then gcc assumes the include files (even those included implicitely)
>>> are assumed to be utf16be.
>>
>> UTF-16-files have a byte-header which helps the compiler to
>> distinguish ASCII-files and UTF-16-files.
> 
> I know that. It's called a byte order mark. And gcc ignores it.

No, wrong - gcc honors it since vesion 1.01.

[toc] | [prev] | [next] | [standalone]


#80599

FromRalf Goertz <me@myprovider.invalid>
Date2021-07-01 11:36 +0200
Message-ID<sbk2au$7pf$1@dont-email.me>
In reply to#80598
Am Thu, 1 Jul 2021 11:10:55 +0200
schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:

> Am 01.07.2021 um 09:19 schrieb Ralf Goertz:
> > Am Wed, 30 Jun 2021 12:37:55 +0200
> > schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> >   
> >> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:  
> >>> Am Wed, 30 Jun 2021 10:05:40 +0200
> >>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> >>>      
> >>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:  
> >>>>> What does the C++ standard say? Does it say that source code
> >>>>> files are always UTF-8 encoded, or is it up to the
> >>>>> implementation? I assume that if it's the latter, the standard
> >>>>> doesn't provide any mechanism to specify which encoding is
> >>>>> being used. Or does it?  
> >>>>
> >>>> Use UTF-16 sourcefiles.  
> >>>
> >>> That doesn't help with gcc. Even if you specify the encoding on
> >>> the command line with -finput-charset=utf16be you run into
> >>> trouble since then gcc assumes the include files (even those
> >>> included implicitely) are assumed to be utf16be.  
> >>
> >> UTF-16-files have a byte-header which helps the compiler to
> >> distinguish ASCII-files and UTF-16-files.  
> > 
> > I know that. It's called a byte order mark. And gcc ignores it.  
> 
> No, wrong - gcc honors it since vesion 1.01.

I created this file b.cc:

int main() {
    return 0;
}

using vi with

:set fileencoding=utf16
:set bomb

Then

~/c> file b.cc 
b.cc: C source, Unicode text, UTF-16, big-endian text

or

~/c> od -h b.cc 
0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
0000040 2000 2000 7200 6500 7400 7500 7200 6e00
0000060 2000 3000 3b00 0a00 7d00 0a00
0000074

where you can see the BOM fffe. Feeding this to gcc (or g++) you get:

~/c> gcc b.cc  
b.cc:1:1: error: stray ‘\376’ in program
    1 | �� i n t   m a i n ( )   { 
      | ^
b.cc:1:2: error: stray ‘\377’ in program
    1 | �� i n t   m a i n ( )   { 
      |  ^
b.cc:1:3: warning: null character(s) ignored
    1 | �� i n t   m a i n ( )   { 
      |   ^
b.cc:1:5: warning: null character(s) ignored

etc.

How does that qualify as “gcc honoring the BOM”?

[toc] | [prev] | [next] | [standalone]


#80605

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-07-01 16:17 +0200
Message-ID<sbkipq$ifi$1@gioia.aioe.org>
In reply to#80599
Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
> Am Thu, 1 Jul 2021 11:10:55 +0200
> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> 
>> Am 01.07.2021 um 09:19 schrieb Ralf Goertz:
>>> Am Wed, 30 Jun 2021 12:37:55 +0200
>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>>    
>>>> Am 30.06.2021 um 10:30 schrieb Ralf Goertz:
>>>>> Am Wed, 30 Jun 2021 10:05:40 +0200
>>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>>>>       
>>>>>> Am 30.06.2021 um 09:57 schrieb Juha Nieminen:
>>>>>>> What does the C++ standard say? Does it say that source code
>>>>>>> files are always UTF-8 encoded, or is it up to the
>>>>>>> implementation? I assume that if it's the latter, the standard
>>>>>>> doesn't provide any mechanism to specify which encoding is
>>>>>>> being used. Or does it?
>>>>>>
>>>>>> Use UTF-16 sourcefiles.
>>>>>
>>>>> That doesn't help with gcc. Even if you specify the encoding on
>>>>> the command line with -finput-charset=utf16be you run into
>>>>> trouble since then gcc assumes the include files (even those
>>>>> included implicitely) are assumed to be utf16be.
>>>>
>>>> UTF-16-files have a byte-header which helps the compiler to
>>>> distinguish ASCII-files and UTF-16-files.
>>>
>>> I know that. It's called a byte order mark. And gcc ignores it.
>>
>> No, wrong - gcc honors it since vesion 1.01.
> 
> I created this file b.cc:
> 
> int main() {
>      return 0;
> }
> 
> using vi with
> 
> :set fileencoding=utf16
> :set bomb
> 
> Then
> 
> ~/c> file b.cc
> b.cc: C source, Unicode text, UTF-16, big-endian text
> 
> or
> 
> ~/c> od -h b.cc
> 0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
> 0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
> 0000040 2000 2000 7200 6500 7400 7500 7200 6e00
> 0000060 2000 3000 3b00 0a00 7d00 0a00
> 0000074
> 
> where you can see the BOM fffe. Feeding this to gcc (or g++) you get:
> 
> ~/c> gcc b.cc
> b.cc:1:1: error: stray ‘\376’ in program
>      1 | �� i n t   m a i n ( )   {
>        | ^
> b.cc:1:2: error: stray ‘\377’ in program
>      1 | �� i n t   m a i n ( )   {
>        |  ^
> b.cc:1:3: warning: null character(s) ignored
>      1 | �� i n t   m a i n ( )   {
>        |   ^
> b.cc:1:5: warning: null character(s) ignored
> 
> etc.
> 
> How does that qualify as “gcc honoring the BOM”?
> 

Use v1.01.

[toc] | [prev] | [next] | [standalone]


#80607

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2021-07-01 10:13 -0700
Message-ID<871r8i9cu4.fsf@nosuchdomain.example.com>
In reply to#80605
Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>> Am Thu, 1 Jul 2021 11:10:55 +0200
>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
[...]
>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>
>>> No, wrong - gcc honors it since vesion 1.01.
>> I created this file b.cc:
[...]
>> How does that qualify as “gcc honoring the BOM”?
>
> Use v1.01.

You think you're funny.  You're not.

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#80608

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-07-01 19:27 +0200
Message-ID<sbktt9$1umt$1@gioia.aioe.org>
In reply to#80607
Am 01.07.2021 um 19:13 schrieb Keith Thompson:
> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
>> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>>> Am Thu, 1 Jul 2021 11:10:55 +0200
>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
> [...]
>>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>>
>>>> No, wrong - gcc honors it since vesion 1.01.
>>> I created this file b.cc:
> [...]
>>> How does that qualify as “gcc honoring the BOM”?
>>
>> Use v1.01.
> 
> You think you're funny.  You're not.

v1.01 does honor the BOM.

[toc] | [prev] | [next] | [standalone]


#80609

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2021-07-01 11:21 -0700
Message-ID<87wnq999ok.fsf@nosuchdomain.example.com>
In reply to#80608
Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
> Am 01.07.2021 um 19:13 schrieb Keith Thompson:
>> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
>>> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>>>> Am Thu, 1 Jul 2021 11:10:55 +0200
>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>> [...]
>>>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>>>
>>>>> No, wrong - gcc honors it since vesion 1.01.
>>>> I created this file b.cc:
>> [...]
>>>> How does that qualify as “gcc honoring the BOM”?
>>>
>>> Use v1.01.
>> You think you're funny.  You're not.
>
> v1.01 does honor the BOM.

At the risk of giving the impression I'm taking you seriously, the
oldest version of gcc available from gnu.org is 1.42, released in 1992.

I've see no evidence that gcc v1.01 would have honored the BOM, but it
doesn't matter, since that version is obsolete and unavailable.

I conclude that you are a troll.

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#80610

FromReal Troll <real.troll@trolls.com>
Date2021-07-01 18:45 +0000
Message-ID<sbl2k3$386$1@gioia.aioe.org>
In reply to#80609
On 01/07/2021 19:21, Keith Thompson wrote:
>
> I conclude that you are a troll.
>
It takes one troll to know another. This is an expert opinion of a Real Troll!


[toc] | [prev] | [next] | [standalone]


#80619

FromKli-Kla-Klawitter <kliklaklawitter69@gmail.com>
Date2021-07-02 10:03 +0200
Message-ID<sbmh8m$1lbb$1@gioia.aioe.org>
In reply to#80609
Am 01.07.2021 um 20:21 schrieb Keith Thompson:
> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
>> Am 01.07.2021 um 19:13 schrieb Keith Thompson:
>>> Kli-Kla-Klawitter <kliklaklawitter69@gmail.com> writes:
>>>> Am 01.07.2021 um 11:36 schrieb Ralf Goertz:
>>>>> Am Thu, 1 Jul 2021 11:10:55 +0200
>>>>> schrieb Kli-Kla-Klawitter <kliklaklawitter69@gmail.com>:
>>> [...]
>>>>>>> I know that. It's called a byte order mark. And gcc ignores it.
>>>>>>
>>>>>> No, wrong - gcc honors it since vesion 1.01.
>>>>> I created this file b.cc:
>>> [...]
>>>>> How does that qualify as “gcc honoring the BOM”?
>>>>
>>>> Use v1.01.
>>> You think you're funny.  You're not.
>>
>> v1.01 does honor the BOM.
> 
> At the risk of giving the impression I'm taking you seriously, the
> oldest version of gcc available from gnu.org is 1.42, released in 1992.

No, the oldest gcc-version is v0.9 from March 22, 1987.
> 
> I've see no evidence that gcc v1.01 would have honored the BOM, but it
> doesn't matter, since that version is obsolete and unavailable.
> 
> I conclude that you are a troll.
> 

[toc] | [prev] | [next] | [standalone]


#80620

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2021-07-02 02:31 -0700
Message-ID<87sg0x83kj.fsf@nosuchdomain.example.com>
In reply to#80599
Ralf Goertz <me@myprovider.invalid> writes:
[...]
> I created this file b.cc:
>
> int main() {
>     return 0;
> }
>
> using vi with
>
> :set fileencoding=utf16
> :set bomb
>
> Then
>
> ~/c> file b.cc 
> b.cc: C source, Unicode text, UTF-16, big-endian text
>
> or
>
> ~/c> od -h b.cc 
> 0000000 fffe 6900 6e00 7400 2000 6d00 6100 6900
> 0000020 6e00 2800 2900 2000 7b00 0a00 2000 2000
> 0000040 2000 2000 7200 6500 7400 7500 7200 6e00
> 0000060 2000 3000 3b00 0a00 7d00 0a00
> 0000074
>
> where you can see the BOM fffe. Feeding this to gcc (or g++) you get:
>
> ~/c> gcc b.cc  
> b.cc:1:1: error: stray ‘\376’ in program
>     1 | �� i n t   m a i n ( )   { 
>       | ^
> b.cc:1:2: error: stray ‘\377’ in program
>     1 | �� i n t   m a i n ( )   { 
>       |  ^
> b.cc:1:3: warning: null character(s) ignored
>     1 | �� i n t   m a i n ( )   { 
>       |   ^
> b.cc:1:5: warning: null character(s) ignored
>
> etc.
>
> How does that qualify as “gcc honoring the BOM”?

On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
(I don't know whether there's a way to configure it to do so.)

It does handle UTF-8 with or without a BOM.

$ file b.cpp
b.cpp: C source, UTF-8 Unicode (with BOM) text
$ cat b.cpp
int main() { }
$ hd b.cpp
00000000  ef bb bf 69 6e 74 20 6d  61 69 6e 28 29 20 7b 20  |...int main() { |
00000010  7d 0a                                             |}.|
00000012
$ gcc -c b.cpp
$

gcc 9.3.0 on Ubuntu 20.04.  (There is, of course, no point in going back
to ancient versions of gcc.)

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#80621

FromMrSpud_oyCn@92wvlb1hltq4dhc.gov.uk
Date2021-07-02 09:38 +0000
Message-ID<sbmmqi$akd$1@gioia.aioe.org>
In reply to#80620
On Fri, 02 Jul 2021 02:31:24 -0700
Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>Ralf Goertz <me@myprovider.invalid> writes:
>[...]
>
>On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>(I don't know whether there's a way to configure it to do so.)

Just out of interest, what byte order is the BOM in? Catch 22?

[toc] | [prev] | [next] | [standalone]


#80622

FromPaavo Helde <myfirstname@osa.pri.ee>
Date2021-07-02 13:52 +0300
Message-ID<sbmr69$tt1$1@dont-email.me>
In reply to#80621
02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas:
> On Fri, 02 Jul 2021 02:31:24 -0700
> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>> Ralf Goertz <me@myprovider.invalid> writes:
>> [...]
>>
>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>> (I don't know whether there's a way to configure it to do so.)
> 
> Just out of interest, what byte order is the BOM in? Catch 22?
> 

This is probably a troll question, but answering anyway: the BOM marker 
U+FEFF is in the correct byte order, in both little-endian and 
big-endian UTF-16. That's how you tell them apart.

The trick is in that the reverse value U+FFFE is not valid Unicode 
character, so there is no possibility of mixup. Neither sequence is also 
valid UTF-8, to avoid mixup with that.

[toc] | [prev] | [next] | [standalone]


#80624

FromMrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net
Date2021-07-02 12:53 +0000
Message-ID<sbn27o$1j5a$1@gioia.aioe.org>
In reply to#80622
On Fri, 2 Jul 2021 13:52:53 +0300
Paavo Helde <myfirstname@osa.pri.ee> wrote:
>02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas:
>> On Fri, 02 Jul 2021 02:31:24 -0700
>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>> Ralf Goertz <me@myprovider.invalid> writes:
>>> [...]
>>>
>>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>>> (I don't know whether there's a way to configure it to do so.)
>> 
>> Just out of interest, what byte order is the BOM in? Catch 22?
>> 
>
>This is probably a troll question, but answering anyway: the BOM marker 

No, not a troll.

>U+FEFF is in the correct byte order, in both little-endian and 
>big-endian UTF-16. That's how you tell them apart.

So its just 2 bytes in sequence, not a 16 bit value?

[toc] | [prev] | [next] | [standalone]


#80625

From"Alf P. Steinbach" <alf.p.steinbach@gmail.com>
Date2021-07-02 15:22 +0200
Message-ID<sbn3u7$qti$1@dont-email.me>
In reply to#80624
On 2 Jul 2021 14:53, MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net wrote:
> On Fri, 2 Jul 2021 13:52:53 +0300
> Paavo Helde <myfirstname@osa.pri.ee> wrote:
>> 02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas:
>>> On Fri, 02 Jul 2021 02:31:24 -0700
>>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>>> Ralf Goertz <me@myprovider.invalid> writes:
>>>> [...]
>>>>
>>>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>>>> (I don't know whether there's a way to configure it to do so.)
>>>
>>> Just out of interest, what byte order is the BOM in? Catch 22?
>>>
>>
>> This is probably a troll question, but answering anyway: the BOM marker
> 
> No, not a troll.
> 
>> U+FEFF is in the correct byte order, in both little-endian and
>> big-endian UTF-16. That's how you tell them apart.
> 
> So its just 2 bytes in sequence, not a 16 bit value?

The BOM is a Unicode code point, U+FEFF as Paavo mentioned, originally 
standing for an invisible zero-width hard space. It's encoded with 
either little endian UTF-16 (then as two bytes), or as big endian UTF-16 
(then as two bytes), or as endianness agnostic UTF-8 (then as three 
bytes). The encoded BOM yields a reliable encoding indicator, though 
pedantic people might argue that it's just statistical -- after all, one 
just might happen to have a Windows 1252 encoded file with three 
characters at the start with the same byte values as the UTF-8 BOM.

In the same vein, one just might happen to have a `.txt` file in Windows 
with the letters "MZ" at the very start, like

     MZ, Mishtara Zva'it, is the Military Police Corps of Israel. blah

and if you then try to open the file in your default text editor by just 
typing the file name in old Cmd, those letters will be misinterpreted as 
the initials of Mark Zbikowski, marking the file as an executable...

Since the chance of that happening isn't absolutely 0 one should never 
use text file names as commands, or the UTF-8 BOM as an encoding marker.


- Alf

[toc] | [prev] | [next] | [standalone]


#80627

FromPaavo Helde <myfirstname@osa.pri.ee>
Date2021-07-02 17:32 +0300
Message-ID<sbn81a$o77$1@dont-email.me>
In reply to#80624
02.07.2021 15:53 MrSpud_zg8yg@nx8574z6ey2rc0y563k5i2.net kirjutas:
> On Fri, 2 Jul 2021 13:52:53 +0300
> Paavo Helde <myfirstname@osa.pri.ee> wrote:
>> 02.07.2021 12:38 MrSpud_oyCn@92wvlb1hltq4dhc.gov.uk kirjutas:
>>> On Fri, 02 Jul 2021 02:31:24 -0700
>>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>>> Ralf Goertz <me@myprovider.invalid> writes:
>>>> [...]
>>>>
>>>> On my system, gcc doesn't handle UTF-16 at all, with or without a BOM.
>>>> (I don't know whether there's a way to configure it to do so.)
>>>
>>> Just out of interest, what byte order is the BOM in? Catch 22?
>>>
>>
>> This is probably a troll question, but answering anyway: the BOM marker
> 
> No, not a troll.
> 
>> U+FEFF is in the correct byte order, in both little-endian and
>> big-endian UTF-16. That's how you tell them apart.
> 
> So its just 2 bytes in sequence, not a 16 bit value?

The file contains bytes, it's up to the reading code how to interpret 
the bytes. It can interpret the bytes as uint16_t, i.e. cast the file 
buffer as 'const uint16_t*' and read the first 2-byte value. If it is 
0xFEFF, then it knows this is an UTF-16 file in a matching byte order. 
If it is  0xFFFE, then it knows it's an UTF-16 file in an opposite byte 
order, and the rest of the file needs to be byte-swapped.

It can also interpret the buffer as containing uint8_t bytes, but then 
the logic is a bit more complex, it must then know if it itself is 
running on a big-endian or little-endian machine, and behave accordingly.

[toc] | [prev] | [next] | [standalone]


#80573

FromRalf Goertz <me@myprovider.invalid>
Date2021-06-30 10:08 +0200
Message-ID<sbh8q7$nhi$1@dont-email.me>
In reply to#80571
Am Wed, 30 Jun 2021 07:57:14 +0000 (UTC)
schrieb Juha Nieminen <nospam@thanks.invalid>:

> In most compilers (such as Visual Studio) you can specify which
> encoding to assume for source files, but this has to be done at the
> project settings level. I don't think there's any way to specify the
> encoding in the source code itself.

When using UTF-encoding there is always A BOM you could use. Doesn't
help much with iso encoding, though. And I also just found out that gcc
doesn't notice that a source file with an appropriate byte order mark is
encoded in utf32 BE. That's a bit disappointing.

[toc] | [prev] | [next] | [standalone]


Page 1 of 4  [1] 2 3 4  Next page →

Back to top | Article view | comp.lang.c++


csiph-web