Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c++ > #83684 > unrolled thread

getline() problem

Started byJivanmukta <jivanmukta@poczta.onet.pl>
First post2022-04-23 19:07 +0200
Last post2022-04-26 11:15 +0300
Articles 20 on this page of 57 — 16 participants

Back to article view | Back to comp.lang.c++


Contents

  getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:07 +0200
    Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 10:17 -0700
      Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:40 +0200
        Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:59 +0200
          Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-23 21:47 +0300
            Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-25 06:00 +0000
              Re: getline() problem Öö Tiib <ootiib@hot.ee> - 2022-04-25 00:12 -0700
              Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 11:40 +0300
              Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 11:37 -0400
                Re: getline() problem Vir Campestris <vir.campestris@invalid.invalid> - 2022-04-25 16:54 +0100
                  Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 12:29 -0400
                    Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 17:29 +0000
                    Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:17 +0200
                  Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-25 10:19 -0700
                  Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:41 -0700
                  Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 21:58 +0300
                Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 15:56 +0000
                  Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 13:12 -0400
                  Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:03 -0700
                    Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 19:05 +0000
                      Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 12:48 -0700
                      Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 21:46 +0100
                    Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 05:43 +0000
                      Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 09:36 +0200
                        Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 09:18 +0000
                          Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-26 02:47 -0700
                            Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 11:06 +0000
                          Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 15:52 +0200
                          Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-26 07:04 -0700
                          Re: getline() problem Manfred <noname@add.invalid> - 2022-04-26 16:49 +0200
                          Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 12:01 -0400
                      Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 11:49 -0400
                  Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:23 +0200
              Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:33 -0700
                Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 10:52 -0700
                  Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:08 -0700
                    Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 21:18 -0700
                      Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 22:04 -0700
                Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:06 +0300
                  Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 13:02 -0700
          Re: getline() problem Barry Schwarz <schwarzb@delq.com> - 2022-04-23 12:54 -0700
          Re: getline() problem Montmorency <none@none.com> - 2022-04-23 14:45 -0700
            Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-23 16:49 -0700
              Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 01:04 +0100
                Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-23 18:24 -0700
                  Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 03:16 +0100
                    Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-24 22:51 -0700
                      Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 11:07 +0100
              Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 17:15 -0700
                Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 18:23 -0700
            Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-24 05:28 +0200
    Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 12:41 +0200
      Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 14:33 +0300
      Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 18:53 +0200
        Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:08 +0300
          Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-26 09:01 +0200
            Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-26 11:15 +0300

Page 1 of 3  [1] 2 3  Next page →


#83684 — getline() problem

FromJivanmukta <jivanmukta@poczta.onet.pl>
Date2022-04-23 19:07 +0200
Subjectgetline() problem
Message-ID<t41bp7$35ldb$1@portraits.wsisiz.edu.pl>
I have a text file (with PHP source code) containing such two lines:

SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
EOL;

I need to read the line into variable: string line. I wrote:

getline(in_file, line);

The problem is that in variable line I receive:

...\Diff\Line:"LineEOL;

I think problem is with \00 characters, probably getline() does not read 
it correctly.

How to read it correctly in my C++ program?

[toc] | [next] | [standalone]


#83685

FromAndrey Tarasevich <andreytarasevich@hotmail.com>
Date2022-04-23 10:17 -0700
Message-ID<t41cbb$qfn$1@dont-email.me>
In reply to#83684
On 4/23/2022 10:07 AM, Jivanmukta wrote:
> I have a text file (with PHP source code) containing such two lines:
> 
> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
> EOL;

What is this supposed to mean? Do you actually have a binary 0 byte in 
the file right after "Line"? Or is it plain text consisting of "\" 
character followed by "00" characters?

I presume it's the latter.

> I need to read the line into variable: string line. I wrote:
> 
> getline(in_file, line);
> 
> The problem is that in variable line I receive:
> 
> ...\Diff\Line:"LineEOL;

> I think problem is with \00 characters, probably getline() does not read 
> it correctly.

No, that's most certainly not the case.

> How to read it correctly in my C++ program?

Most likely you are making something up or misinterpreting something. 
Everything is already read correctly. You are just inspecting the result 
in some misleading way and subsequently convince yourself that something 
is wrong.

-- 
Best regards,
Andrey Tarasevich

[toc] | [prev] | [next] | [standalone]


#83686

FromJivanmukta <jivanmukta@poczta.onet.pl>
Date2022-04-23 19:40 +0200
Message-ID<t41dmo$35p5k$1@portraits.wsisiz.edu.pl>
In reply to#83685
W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze:
> On 4/23/2022 10:07 AM, Jivanmukta wrote:
>> I have a text file (with PHP source code) containing such two lines:
>>
>> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
>> EOL;
> 
> What is this supposed to mean? Do you actually have a binary 0 byte in 
> the file right after "Line"? Or is it plain text consisting of "\" 
> character followed by "00" characters?
> 
> I presume it's the latter.

plain text consisting of "\" character followed by "00" characters

> 
>> I need to read the line into variable: string line. I wrote:
>>
>> getline(in_file, line);
>>
>> The problem is that in variable line I receive:
>>
>> ...\Diff\Line:"LineEOL;
> 
>> I think problem is with \00 characters, probably getline() does not 
>> read it correctly.
> 
> No, that's most certainly not the case.
> 
>> How to read it correctly in my C++ program?
> 
> Most likely you are making something up or misinterpreting something. 
> Everything is already read correctly. You are just inspecting the result 
> in some misleading way and subsequently convince yourself that something 
> is wrong.
> 

Why I have 2 lines concatenated (I am sure of that because I write 
contents of variable line into log file):
	...\Diff\Line:"LineEOL;
There's newline character before EOL; string in the input file!

[toc] | [prev] | [next] | [standalone]


#83687

FromJivanmukta <jivanmukta@poczta.onet.pl>
Date2022-04-23 19:59 +0200
Message-ID<t41er4$35r34$1@portraits.wsisiz.edu.pl>
In reply to#83686
W dniu 23.04.2022 o 19:40, Jivanmukta pisze:
> W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze:
>> On 4/23/2022 10:07 AM, Jivanmukta wrote:
>>> I have a text file (with PHP source code) containing such two lines:
>>>
>>> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
>>> EOL;
>>
>> What is this supposed to mean? Do you actually have a binary 0 byte in 
>> the file right after "Line"? Or is it plain text consisting of "\" 
>> character followed by "00" characters?
>>
>> I presume it's the latter.
> 
> plain text consisting of "\" character followed by "00" characters

Sorry my mistake: I opened the file in VSCode and now I see that the 
file contains NUL bytes.
How to read in C++ line of text finished with newline and containing NUL 
bytes?

[toc] | [prev] | [next] | [standalone]


#83688

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-04-23 21:47 +0300
Message-ID<t41hkn$6p9$1@dont-email.me>
In reply to#83687
23.04.2022 20:59 Jivanmukta kirjutas:
> W dniu 23.04.2022 o 19:40, Jivanmukta pisze:
>> W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze:
>>> On 4/23/2022 10:07 AM, Jivanmukta wrote:
>>>> I have a text file (with PHP source code) containing such two lines:
>>>>
>>>> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
>>>> EOL;
>>>
>>> What is this supposed to mean? Do you actually have a binary 0 byte 
>>> in the file right after "Line"? Or is it plain text consisting of "\" 
>>> character followed by "00" characters?
>>>
>>> I presume it's the latter.
>>
>> plain text consisting of "\" character followed by "00" characters
> 
> Sorry my mistake: I opened the file in VSCode and now I see that the 
> file contains NUL bytes.
> How to read in C++ line of text finished with newline and containing NUL 
> bytes?

There is no problem with reading NUL bytes in C++, they are not handled 
specially in any way, unlike in C.

What you need to keep in mind is that most of C is also a subset of C++, 
so when dealing with strings containing NUL bytes you need to take care 
to NOT use any C-style functionality from that subset which assumes 
zero-terminated strings. Basically it means not to call c_str() on the 
input line string, not to use C functions like strchr(), etc.

Also, the debuggers often only show the string up to the first NUL byte, 
so they might give you a wrong impression what you have really got. Call 
string.length() when in doubt.

Example: reading a file starting with line "ABC\00DEF\n" :

#include <iostream>
#include <fstream>
#include <string>

int main() {

	std::ifstream is("c:/tmp/zero.txt");
	std::string line;
	std::getline(is, line);
	for (char c : line) {
		std::cout << int(c) << " ";
	}
}


Output:
65 66 67 0 68 69 70

You see the NUL byte is prominently present here in the middle of the 
string.

[toc] | [prev] | [next] | [standalone]


#83705

FromJuha Nieminen <nospam@thanks.invalid>
Date2022-04-25 06:00 +0000
Message-ID<t45de2$1sln$2@gioia.aioe.org>
In reply to#83688
Paavo Helde <eesnimi@osa.pri.ee> wrote:
> There is no problem with reading NUL bytes in C++, they are not handled 
> specially in any way, unlike in C.

Still, I wonder if the file stream should be opened in binary mode
(instead of the default). I have experiences of this making a difference
with the Microsoft compilers in Windows. It can do funky things to
binary input if the file stream isn't opened in binary mode.

[toc] | [prev] | [next] | [standalone]


#83707

FromÖö Tiib <ootiib@hot.ee>
Date2022-04-25 00:12 -0700
Message-ID<2fe2e923-2b17-492a-905a-fb76b2e4d420n@googlegroups.com>
In reply to#83705
On Monday, 25 April 2022 at 09:00:53 UTC+3, Juha Nieminen wrote:
> Paavo Helde <ees...@osa.pri.ee> wrote: 
> > There is no problem with reading NUL bytes in C++, they are not handled 
> > specially in any way, unlike in C.
> 
> Still, I wonder if the file stream should be opened in binary mode 
> (instead of the default). I have experiences of this making a difference 
> with the Microsoft compilers in Windows. It can do funky things to 
> binary input if the file stream isn't opened in binary mode.

Yes, always open in binary mode. In Windows text stream a 
character '\0x1A' will end the stream preliminarily ... as example of just
one of horrible bear traps you will face. So simply dump it.

std::ifstream in(filename, ios_base::binary);

The implementation is only allowed to add indeterminate amount of null
characters to end of binary stream, otherwise it must be precise by
byte. And everybody seem to comply.

[toc] | [prev] | [next] | [standalone]


#83709

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-04-25 11:40 +0300
Message-ID<t45mph$is4$1@dont-email.me>
In reply to#83705
25.04.2022 09:00 Juha Nieminen kirjutas:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> There is no problem with reading NUL bytes in C++, they are not handled
>> specially in any way, unlike in C.
> 
> Still, I wonder if the file stream should be opened in binary mode
> (instead of the default). I have experiences of this making a difference
> with the Microsoft compilers in Windows. It can do funky things to
> binary input if the file stream isn't opened in binary mode.

It does not make a difference regarding NUL bytes, at least not with 
recent MSVC on Windows. It does make difference for byte value 26 
though, as commented by Night Wing, so it is indeed a good idea to use 
binary mode for such files which can contain NUL bytes and who knows 
what else.

In general it makes a lot of sense to open all files in binary mode, 
even if they are known to be text files, to avoid any potential platform 
or implementation specific "features" like truncating the file at byte 
26, or producing/expecting platform-specific line endings.

Curiously enough, the concept of file text mode which was introduced to 
make the software more portable has turned into its opposite and is now 
making things less portable. Locales are somewhat similar.

[toc] | [prev] | [next] | [standalone]


#83728

FromJames Kuyper <jameskuyper@alumni.caltech.edu>
Date2022-04-25 11:37 -0400
Message-ID<t46f86$or8$1@dont-email.me>
In reply to#83705
On 4/25/22 02:00, Juha Nieminen wrote:
> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>> There is no problem with reading NUL bytes in C++, they are not handled 
>> specially in any way, unlike in C.
> 
> Still, I wonder if the file stream should be opened in binary mode
> (instead of the default). I have experiences of this making a difference
> with the Microsoft compilers in Windows. It can do funky things to
> binary input if the file stream isn't opened in binary mode.

The C++ standard says very little about the differences between text
mode and binary mode, cross-referencing the C standard for such issues.
The C standard says something important to this issue:

"... Data read in from a text stream will necessarily compare equal to
the data that were earlier written out to that stream only if: the data
consist only of printing characters and the control characters
horizontal tab and new-line; no new-line character is immediately
preceded by space characters; and the last character is a new-line
character. ..." (7.21.2p2).

As a general rule, you should not write data to a text stream if it
contains any feature invalidating that guarantee, and you should not
read data from a text stream if it contains any of those things.

[toc] | [prev] | [next] | [standalone]


#83733

FromVir Campestris <vir.campestris@invalid.invalid>
Date2022-04-25 16:54 +0100
Message-ID<t46g6s$3e7$1@dont-email.me>
In reply to#83728
On 25/04/2022 16:37, James Kuyper wrote:
> On 4/25/22 02:00, Juha Nieminen wrote:
>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>> There is no problem with reading NUL bytes in C++, they are not handled
>>> specially in any way, unlike in C.
>>
>> Still, I wonder if the file stream should be opened in binary mode
>> (instead of the default). I have experiences of this making a difference
>> with the Microsoft compilers in Windows. It can do funky things to
>> binary input if the file stream isn't opened in binary mode.
> 
> The C++ standard says very little about the differences between text
> mode and binary mode, cross-referencing the C standard for such issues.
> The C standard says something important to this issue:
> 
> "... Data read in from a text stream will necessarily compare equal to
> the data that were earlier written out to that stream only if: the data
> consist only of printing characters and the control characters
> horizontal tab and new-line; no new-line character is immediately
> preceded by space characters; and the last character is a new-line
> character. ..." (7.21.2p2).
> 
> As a general rule, you should not write data to a text stream if it
> contains any feature invalidating that guarantee, and you should not
> read data from a text stream if it contains any of those things.

Windows text files have 13,10 as a newline sequence, not just 10 as in 
Linux etc.

This dates back to at least CP/M where the byte stream in the file 
exactly matched the one you would send to your CRT or printer, and a 
line feed really did move the roller but not the print head.

I have a vague feeling that Apple might use 10,13 - certainly I've met 
it somewhere in the past.

The point of text mode is to allow a program to ignore these differences 
by allowing the runtime to handle them.

It's not 100%.

Andy

[toc] | [prev] | [next] | [standalone]


#83738

FromJames Kuyper <jameskuyper@alumni.caltech.edu>
Date2022-04-25 12:29 -0400
Message-ID<t46i9f$mnk$1@dont-email.me>
In reply to#83733
On 4/25/22 11:54, Vir Campestris wrote:
> On 25/04/2022 16:37, James Kuyper wrote:
...
>> The C++ standard says very little about the differences between text
>> mode and binary mode, cross-referencing the C standard for such issues.
>> The C standard says something important to this issue:
>>
>> "... Data read in from a text stream will necessarily compare equal to
>> the data that were earlier written out to that stream only if: the data
>> consist only of printing characters and the control characters
>> horizontal tab and new-line; no new-line character is immediately
>> preceded by space characters; and the last character is a new-line
>> character. ..." (7.21.2p2).
>>
>> As a general rule, you should not write data to a text stream if it
>> contains any feature invalidating that guarantee, and you should not
>> read data from a text stream if it contains any of those things.
> 
> Windows text files have 13,10 as a newline sequence, not just 10 as in 
> Linux etc.
> 
> This dates back to at least CP/M where the byte stream in the file 
> exactly matched the one you would send to your CRT or printer, and a 
> line feed really did move the roller but not the print head.
> 
> I have a vague feeling that Apple might use 10,13 - certainly I've met 
> it somewhere in the past.
> 
> The point of text mode is to allow a program to ignore these differences 
> by allowing the runtime to handle them.
> 
> It's not 100%.

Those requirements explicitly allow new-line characters be writable to
text streams, and that when they are read in again, the result must
compare equal to '\n'.  This means that on platforms where a new-line is
represented in the underlying file by "\n\r" or by "\r\n", then standard
I/O routines must convert '\n' to that sequence on write, and convert
that sequence back to '\n' when reading from the same stream.
If the "not 100%' you're talking about violates that guarantee, then
what you're talking about isn't a conforming implementation of C++.

[toc] | [prev] | [next] | [standalone]


#83744

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-04-25 17:29 +0000
Message-ID<2UA9K.5190$Awz.3215@fx03.iad>
In reply to#83738
James Kuyper <jameskuyper@alumni.caltech.edu> writes:
>On 4/25/22 11:54, Vir Campestris wrote:
>> On 25/04/2022 16:37, James Kuyper wrote:
>...
>>> The C++ standard says very little about the differences between text
>>> mode and binary mode, cross-referencing the C standard for such issues.
>>> The C standard says something important to this issue:
>>>
>>> "... Data read in from a text stream will necessarily compare equal to
>>> the data that were earlier written out to that stream only if: the data
>>> consist only of printing characters and the control characters
>>> horizontal tab and new-line; no new-line character is immediately
>>> preceded by space characters; and the last character is a new-line
>>> character. ..." (7.21.2p2).
>>>

>> 
>> Windows text files have 13,10 as a newline sequence, not just 10 as in 
>> Linux etc.
>> 
>> This dates back to at least CP/M where the byte stream in the file 

It was used by DEC systems, IBM and Burroughs systems (when communicating
with teletypes and other non-line-printer output devices) even before 1974
when CP/M was developed, although generally not part of the file for the
non-DEC systems, rather added by the operating system line discipline functions.

[toc] | [prev] | [next] | [standalone]


#83751

FromManfred <noname@add.invalid>
Date2022-04-25 20:17 +0200
Message-ID<t46ojk$1nv4$1@gioia.aioe.org>
In reply to#83738
On 4/25/2022 6:29 PM, James Kuyper wrote:
> On 4/25/22 11:54, Vir Campestris wrote:
>> On 25/04/2022 16:37, James Kuyper wrote:
> ...
>>> The C++ standard says very little about the differences between text
>>> mode and binary mode, cross-referencing the C standard for such issues.
>>> The C standard says something important to this issue:
>>>
>>> "... Data read in from a text stream will necessarily compare equal to
>>> the data that were earlier written out to that stream only if: the data
>>> consist only of printing characters and the control characters
>>> horizontal tab and new-line; no new-line character is immediately
>>> preceded by space characters; and the last character is a new-line
>>> character. ..." (7.21.2p2).
>>>
>>> As a general rule, you should not write data to a text stream if it
>>> contains any feature invalidating that guarantee, and you should not
>>> read data from a text stream if it contains any of those things.
>>
>> Windows text files have 13,10 as a newline sequence, not just 10 as in
>> Linux etc.
>>
>> This dates back to at least CP/M where the byte stream in the file
>> exactly matched the one you would send to your CRT or printer, and a
>> line feed really did move the roller but not the print head.
>>
>> I have a vague feeling that Apple might use 10,13 - certainly I've met
>> it somewhere in the past.
>>
>> The point of text mode is to allow a program to ignore these differences
>> by allowing the runtime to handle them.
>>
>> It's not 100%.
> 
> Those requirements explicitly allow new-line characters be writable to
> text streams, and that when they are read in again, the result must
> compare equal to '\n'.  This means that on platforms where a new-line is
> represented in the underlying file by "\n\r" or by "\r\n", then standard
> I/O routines must convert '\n' to that sequence on write, and convert
> that sequence back to '\n' when reading from the same stream.
> If the "not 100%' you're talking about violates that guarantee, then
> what you're talking about isn't a conforming implementation of C++.
> 
> 

MS-Windows (and DOS before that) behave as you describe ('\n' is stored 
as '\r\n' on disk and converted back to '\n' upon read) - or, more 
properly, the MS C runtime does that.
So, according to your quote above from the C standard that allows for 
only LF and TAB for round trip consistency of text files, MS 
implementations have been standard compliant since the beginning of time.
That said, interoperability when dealing with text streams between 
Windows and other systems is still painful.

[toc] | [prev] | [next] | [standalone]


#83743

FromKeith Thompson <Keith.S.Thompson+u@gmail.com>
Date2022-04-25 10:19 -0700
Message-ID<87o80paxnb.fsf@nosuchdomain.example.com>
In reply to#83733
Vir Campestris <vir.campestris@invalid.invalid> writes:
[...]
> Windows text files have 13,10 as a newline sequence, not just 10 as in
> Linux etc.
>
> This dates back to at least CP/M where the byte stream in the file
> exactly matched the one you would send to your CRT or printer, and a 
> line feed really did move the roller but not the print head.
>
> I have a vague feeling that Apple might use 10,13 - certainly I've met
> it somewhere in the past.

Versions of MacOS before version X (10) used CR as a line terminator.
OS X, since it's Unix-based, uses LF.

> The point of text mode is to allow a program to ignore these
> differences by allowing the runtime to handle them.
>
> It's not 100%.

-- 
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */

[toc] | [prev] | [next] | [standalone]


#83746

FromAndrey Tarasevich <andreytarasevich@hotmail.com>
Date2022-04-25 10:41 -0700
Message-ID<t46mfm$qkp$1@dont-email.me>
In reply to#83733
On 4/25/2022 8:54 AM, Vir Campestris wrote:
> 
> The point of text mode is to allow a program to ignore these differences 
> by allowing the runtime to handle them.
> 

"Ignoring these differences" comes with a price, like restrictions on 
absolute positioning. E.g. one's only allowed to `fseek` to 0 or to a 
position previously obtained from `ftell`. Arbitrary relative/absolute 
positioning is not supported in text streams. These limitations, if I 
remember correctly, transfer to C++ I/O library unchanged since it 
refers to C's `fseek/ftell' specification where it matters.

In other words, you can't just completely and transparently "ignore" all 
the differences.

-- 
Best regards,
Andrey Tarasevich

[toc] | [prev] | [next] | [standalone]


#83753

FromPaavo Helde <eesnimi@osa.pri.ee>
Date2022-04-25 21:58 +0300
Message-ID<t46r01$1nq$1@dont-email.me>
In reply to#83733
25.04.2022 18:54 Vir Campestris kirjutas:
> 
> Windows text files have 13,10 as a newline sequence, not just 10 as in 
> Linux etc.
> 
> This dates back to at least CP/M where the byte stream in the file 
> exactly matched the one you would send to your CRT or printer, and a 
> line feed really did move the roller but not the print head.
> 
> I have a vague feeling that Apple might use 10,13 - certainly I've met 
> it somewhere in the past.

It was 13 in old Mac, and 10 since Mac OS X, as it is a genuine Unix.

> 
> The point of text mode is to allow a program to ignore these differences 
> by allowing the runtime to handle them.

This only works if the text file has been produced on a platform with 
the same conventions as the currently running program. Nowadays files 
are routinely copied and accessed all over the world and almost always 
one needs to support reading text files produced by "wrong" conventions. 
And there is no guarantee the runtime handles them correctly in text 
mode, rather the opposite.

Even Notepad can read unix text files nowadays, so it's obviously not 
that hard ;-)

Once upon a time there was also FTP text mode which was meant to cope 
with these platform-specific differences, but this approach caused more 
problems than solved.

[toc] | [prev] | [next] | [standalone]


#83735

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-04-25 15:56 +0000
Message-ID<Kwz9K.23424$x9Ea.11744@fx45.iad>
In reply to#83728
James Kuyper <jameskuyper@alumni.caltech.edu> writes:
>On 4/25/22 02:00, Juha Nieminen wrote:
>> Paavo Helde <eesnimi@osa.pri.ee> wrote:
>>> There is no problem with reading NUL bytes in C++, they are not handled 
>>> specially in any way, unlike in C.
>> 
>> Still, I wonder if the file stream should be opened in binary mode
>> (instead of the default). I have experiences of this making a difference
>> with the Microsoft compilers in Windows. It can do funky things to
>> binary input if the file stream isn't opened in binary mode.
>
>The C++ standard says very little about the differences between text
>mode and binary mode, cross-referencing the C standard for such issues.
>The C standard says something important to this issue:
>
>"... Data read in from a text stream will necessarily compare equal to
>the data that were earlier written out to that stream only if: the data
>consist only of printing characters and the control characters
>horizontal tab and new-line; no new-line character is immediately
>preceded by space characters; and the last character is a new-line
>character. ..." (7.21.2p2).
>
>As a general rule, you should not write data to a text stream if it
>contains any feature invalidating that guarantee, and you should not
>read data from a text stream if it contains any of those things.

As a general rule, your general rule only applies to Windows.  On Unix
and Unix-like systems, there is no functional difference between text
and binary.

[toc] | [prev] | [next] | [standalone]


#83742

FromJames Kuyper <jameskuyper@alumni.caltech.edu>
Date2022-04-25 13:12 -0400
Message-ID<t46kq1$c2j$2@dont-email.me>
In reply to#83735
On 4/25/22 11:56, Scott Lurndal wrote:
> James Kuyper <jameskuyper@alumni.caltech.edu> writes:
...
>> The C++ standard says very little about the differences between text
>> mode and binary mode, cross-referencing the C standard for such issues.
>> The C standard says something important to this issue:
>>
>> "... Data read in from a text stream will necessarily compare equal to
>> the data that were earlier written out to that stream only if: the data
>> consist only of printing characters and the control characters
>> horizontal tab and new-line; no new-line character is immediately
>> preceded by space characters; and the last character is a new-line
>> character. ..." (7.21.2p2).
>>
>> As a general rule, you should not write data to a text stream if it
>> contains any feature invalidating that guarantee, and you should not
>> read data from a text stream if it contains any of those things.
>
> As a general rule, your general rule only applies to Windows. On Unix
> and Unix-like systems, there is no functional difference between text
> and binary.

True - but if you want your code to be portable to other systems, you
should follow that rule. It's not exactly an onerous requirement.

[toc] | [prev] | [next] | [standalone]


#83749

FromAndrey Tarasevich <andreytarasevich@hotmail.com>
Date2022-04-25 11:03 -0700
Message-ID<t46np1$5uj$1@dont-email.me>
In reply to#83735
On 4/25/2022 8:56 AM, Scott Lurndal wrote:
>>
>> As a general rule, you should not write data to a text stream if it
>> contains any feature invalidating that guarantee, and you should not
>> read data from a text stream if it contains any of those things.
> 
> As a general rule, your general rule only applies to Windows.  On Unix
> and Unix-like systems, there is no functional difference between text
> and binary.

That's catastrophically incorrect on more than one level.

This is the same fallacy that is often used by numerous "UB deniers", 
claiming that on their hardware/OS platform some forms of "undefined 
behavior" are supposedly "perfectly defined" because they know how their 
platform/hardware/OS will react in some specific situation.

They just can't grasp this little bit of understanding that the primary 
source of "undefined" in undefined behavior in _not_ their hardware/OS. 
It is the abstract language semantics implemented by the compiler. The 
compiler that actively uses UB as an optimization opportunity at a 
rather abstract level. It doesn't even get to the hardware/OS level. 
Hardware/OS matter very little.

The same reasoning can be applied to the matter of binary vs. text streams.

For example, the very moment the abstract specification of text stream 
will provide optimization opportunities (or any other benefits) to the 
implementation of C or C++ standard library, this implementation 
might/will take advantage of these opportunities to make text streams 
work more efficiently than binary streams (within the spec, of course). 
It will break text vs. binary compatibility and it will do this with 
complete disregard to the fact that the underlying OS makes no 
difference between text and binary.

I'm exaggerating, of course, because it something like this ever happens 
to Unix-like I/O, the raging uproar the affected parties will be 
deafening. (We've heard it before with strict aliasing, strict overflow, 
`memcpy` that copies backwards etc.) Nobody will dare to break 
binary-text equivalence of Unix-like systems (and POSIX won't allow 
this). But in essence... this is just a wall of entrenched incompetence 
too high to climb over.

-- 
Best regards,
Andrey Tarasevich

[toc] | [prev] | [next] | [standalone]


#83754

Fromscott@slp53.sl.home (Scott Lurndal)
Date2022-04-25 19:05 +0000
Message-ID<0iC9K.123255$Kdf.44786@fx96.iad>
In reply to#83749
Andrey Tarasevich <andreytarasevich@hotmail.com> writes:
>On 4/25/2022 8:56 AM, Scott Lurndal wrote:
>>>
>>> As a general rule, you should not write data to a text stream if it
>>> contains any feature invalidating that guarantee, and you should not
>>> read data from a text stream if it contains any of those things.
>> 
>> As a general rule, your general rule only applies to Windows.  On Unix
>> and Unix-like systems, there is no functional difference between text
>> and binary.
>
>That's catastrophically incorrect on more than one level.

Actually, that's POSIX defined behavior.

Rest of screed elided.

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | comp.lang.c++


csiph-web