Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #83684 > unrolled thread
| Started by | Jivanmukta <jivanmukta@poczta.onet.pl> |
|---|---|
| First post | 2022-04-23 19:07 +0200 |
| Last post | 2022-04-26 11:15 +0300 |
| Articles | 20 on this page of 57 — 16 participants |
Back to article view | Back to comp.lang.c++
getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:07 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 10:17 -0700
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:40 +0200
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:59 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-23 21:47 +0300
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-25 06:00 +0000
Re: getline() problem Öö Tiib <ootiib@hot.ee> - 2022-04-25 00:12 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 11:40 +0300
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 11:37 -0400
Re: getline() problem Vir Campestris <vir.campestris@invalid.invalid> - 2022-04-25 16:54 +0100
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 12:29 -0400
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 17:29 +0000
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:17 +0200
Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-25 10:19 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:41 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 21:58 +0300
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 15:56 +0000
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 13:12 -0400
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:03 -0700
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 19:05 +0000
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 12:48 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 21:46 +0100
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 05:43 +0000
Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 09:36 +0200
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 09:18 +0000
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-26 02:47 -0700
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 11:06 +0000
Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 15:52 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-26 07:04 -0700
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-26 16:49 +0200
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 12:01 -0400
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 11:49 -0400
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:23 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:33 -0700
Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 10:52 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:08 -0700
Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 21:18 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 22:04 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:06 +0300
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 13:02 -0700
Re: getline() problem Barry Schwarz <schwarzb@delq.com> - 2022-04-23 12:54 -0700
Re: getline() problem Montmorency <none@none.com> - 2022-04-23 14:45 -0700
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-23 16:49 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 01:04 +0100
Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-23 18:24 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 03:16 +0100
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-24 22:51 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 11:07 +0100
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 17:15 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 18:23 -0700
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-24 05:28 +0200
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 12:41 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 14:33 +0300
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 18:53 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:08 +0300
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-26 09:01 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-26 11:15 +0300
Page 1 of 3 [1] 2 3 Next page →
| From | Jivanmukta <jivanmukta@poczta.onet.pl> |
|---|---|
| Date | 2022-04-23 19:07 +0200 |
| Subject | getline() problem |
| Message-ID | <t41bp7$35ldb$1@portraits.wsisiz.edu.pl> |
I have a text file (with PHP source code) containing such two lines: SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}} EOL; I need to read the line into variable: string line. I wrote: getline(in_file, line); The problem is that in variable line I receive: ...\Diff\Line:"LineEOL; I think problem is with \00 characters, probably getline() does not read it correctly. How to read it correctly in my C++ program?
[toc] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-23 10:17 -0700 |
| Message-ID | <t41cbb$qfn$1@dont-email.me> |
| In reply to | #83684 |
On 4/23/2022 10:07 AM, Jivanmukta wrote: > I have a text file (with PHP source code) containing such two lines: > > SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}} > EOL; What is this supposed to mean? Do you actually have a binary 0 byte in the file right after "Line"? Or is it plain text consisting of "\" character followed by "00" characters? I presume it's the latter. > I need to read the line into variable: string line. I wrote: > > getline(in_file, line); > > The problem is that in variable line I receive: > > ...\Diff\Line:"LineEOL; > I think problem is with \00 characters, probably getline() does not read > it correctly. No, that's most certainly not the case. > How to read it correctly in my C++ program? Most likely you are making something up or misinterpreting something. Everything is already read correctly. You are just inspecting the result in some misleading way and subsequently convince yourself that something is wrong. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | Jivanmukta <jivanmukta@poczta.onet.pl> |
|---|---|
| Date | 2022-04-23 19:40 +0200 |
| Message-ID | <t41dmo$35p5k$1@portraits.wsisiz.edu.pl> |
| In reply to | #83685 |
W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze: > On 4/23/2022 10:07 AM, Jivanmukta wrote: >> I have a text file (with PHP source code) containing such two lines: >> >> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}} >> EOL; > > What is this supposed to mean? Do you actually have a binary 0 byte in > the file right after "Line"? Or is it plain text consisting of "\" > character followed by "00" characters? > > I presume it's the latter. plain text consisting of "\" character followed by "00" characters > >> I need to read the line into variable: string line. I wrote: >> >> getline(in_file, line); >> >> The problem is that in variable line I receive: >> >> ...\Diff\Line:"LineEOL; > >> I think problem is with \00 characters, probably getline() does not >> read it correctly. > > No, that's most certainly not the case. > >> How to read it correctly in my C++ program? > > Most likely you are making something up or misinterpreting something. > Everything is already read correctly. You are just inspecting the result > in some misleading way and subsequently convince yourself that something > is wrong. > Why I have 2 lines concatenated (I am sure of that because I write contents of variable line into log file): ...\Diff\Line:"LineEOL; There's newline character before EOL; string in the input file!
[toc] | [prev] | [next] | [standalone]
| From | Jivanmukta <jivanmukta@poczta.onet.pl> |
|---|---|
| Date | 2022-04-23 19:59 +0200 |
| Message-ID | <t41er4$35r34$1@portraits.wsisiz.edu.pl> |
| In reply to | #83686 |
W dniu 23.04.2022 o 19:40, Jivanmukta pisze: > W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze: >> On 4/23/2022 10:07 AM, Jivanmukta wrote: >>> I have a text file (with PHP source code) containing such two lines: >>> >>> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}} >>> EOL; >> >> What is this supposed to mean? Do you actually have a binary 0 byte in >> the file right after "Line"? Or is it plain text consisting of "\" >> character followed by "00" characters? >> >> I presume it's the latter. > > plain text consisting of "\" character followed by "00" characters Sorry my mistake: I opened the file in VSCode and now I see that the file contains NUL bytes. How to read in C++ line of text finished with newline and containing NUL bytes?
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-04-23 21:47 +0300 |
| Message-ID | <t41hkn$6p9$1@dont-email.me> |
| In reply to | #83687 |
23.04.2022 20:59 Jivanmukta kirjutas:
> W dniu 23.04.2022 o 19:40, Jivanmukta pisze:
>> W dniu 23.04.2022 o 19:17, Andrey Tarasevich pisze:
>>> On 4/23/2022 10:07 AM, Jivanmukta wrote:
>>>> I have a text file (with PHP source code) containing such two lines:
>>>>
>>>> SebastianBergmann\Diff\Line\00content";s:7:"2222222";}}}}}}
>>>> EOL;
>>>
>>> What is this supposed to mean? Do you actually have a binary 0 byte
>>> in the file right after "Line"? Or is it plain text consisting of "\"
>>> character followed by "00" characters?
>>>
>>> I presume it's the latter.
>>
>> plain text consisting of "\" character followed by "00" characters
>
> Sorry my mistake: I opened the file in VSCode and now I see that the
> file contains NUL bytes.
> How to read in C++ line of text finished with newline and containing NUL
> bytes?
There is no problem with reading NUL bytes in C++, they are not handled
specially in any way, unlike in C.
What you need to keep in mind is that most of C is also a subset of C++,
so when dealing with strings containing NUL bytes you need to take care
to NOT use any C-style functionality from that subset which assumes
zero-terminated strings. Basically it means not to call c_str() on the
input line string, not to use C functions like strchr(), etc.
Also, the debuggers often only show the string up to the first NUL byte,
so they might give you a wrong impression what you have really got. Call
string.length() when in doubt.
Example: reading a file starting with line "ABC\00DEF\n" :
#include <iostream>
#include <fstream>
#include <string>
int main() {
std::ifstream is("c:/tmp/zero.txt");
std::string line;
std::getline(is, line);
for (char c : line) {
std::cout << int(c) << " ";
}
}
Output:
65 66 67 0 68 69 70
You see the NUL byte is prominently present here in the middle of the
string.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-04-25 06:00 +0000 |
| Message-ID | <t45de2$1sln$2@gioia.aioe.org> |
| In reply to | #83688 |
Paavo Helde <eesnimi@osa.pri.ee> wrote: > There is no problem with reading NUL bytes in C++, they are not handled > specially in any way, unlike in C. Still, I wonder if the file stream should be opened in binary mode (instead of the default). I have experiences of this making a difference with the Microsoft compilers in Windows. It can do funky things to binary input if the file stream isn't opened in binary mode.
[toc] | [prev] | [next] | [standalone]
| From | Öö Tiib <ootiib@hot.ee> |
|---|---|
| Date | 2022-04-25 00:12 -0700 |
| Message-ID | <2fe2e923-2b17-492a-905a-fb76b2e4d420n@googlegroups.com> |
| In reply to | #83705 |
On Monday, 25 April 2022 at 09:00:53 UTC+3, Juha Nieminen wrote: > Paavo Helde <ees...@osa.pri.ee> wrote: > > There is no problem with reading NUL bytes in C++, they are not handled > > specially in any way, unlike in C. > > Still, I wonder if the file stream should be opened in binary mode > (instead of the default). I have experiences of this making a difference > with the Microsoft compilers in Windows. It can do funky things to > binary input if the file stream isn't opened in binary mode. Yes, always open in binary mode. In Windows text stream a character '\0x1A' will end the stream preliminarily ... as example of just one of horrible bear traps you will face. So simply dump it. std::ifstream in(filename, ios_base::binary); The implementation is only allowed to add indeterminate amount of null characters to end of binary stream, otherwise it must be precise by byte. And everybody seem to comply.
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-04-25 11:40 +0300 |
| Message-ID | <t45mph$is4$1@dont-email.me> |
| In reply to | #83705 |
25.04.2022 09:00 Juha Nieminen kirjutas: > Paavo Helde <eesnimi@osa.pri.ee> wrote: >> There is no problem with reading NUL bytes in C++, they are not handled >> specially in any way, unlike in C. > > Still, I wonder if the file stream should be opened in binary mode > (instead of the default). I have experiences of this making a difference > with the Microsoft compilers in Windows. It can do funky things to > binary input if the file stream isn't opened in binary mode. It does not make a difference regarding NUL bytes, at least not with recent MSVC on Windows. It does make difference for byte value 26 though, as commented by Night Wing, so it is indeed a good idea to use binary mode for such files which can contain NUL bytes and who knows what else. In general it makes a lot of sense to open all files in binary mode, even if they are known to be text files, to avoid any potential platform or implementation specific "features" like truncating the file at byte 26, or producing/expecting platform-specific line endings. Curiously enough, the concept of file text mode which was introduced to make the software more portable has turned into its opposite and is now making things less portable. Locales are somewhat similar.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-25 11:37 -0400 |
| Message-ID | <t46f86$or8$1@dont-email.me> |
| In reply to | #83705 |
On 4/25/22 02:00, Juha Nieminen wrote: > Paavo Helde <eesnimi@osa.pri.ee> wrote: >> There is no problem with reading NUL bytes in C++, they are not handled >> specially in any way, unlike in C. > > Still, I wonder if the file stream should be opened in binary mode > (instead of the default). I have experiences of this making a difference > with the Microsoft compilers in Windows. It can do funky things to > binary input if the file stream isn't opened in binary mode. The C++ standard says very little about the differences between text mode and binary mode, cross-referencing the C standard for such issues. The C standard says something important to this issue: "... Data read in from a text stream will necessarily compare equal to the data that were earlier written out to that stream only if: the data consist only of printing characters and the control characters horizontal tab and new-line; no new-line character is immediately preceded by space characters; and the last character is a new-line character. ..." (7.21.2p2). As a general rule, you should not write data to a text stream if it contains any feature invalidating that guarantee, and you should not read data from a text stream if it contains any of those things.
[toc] | [prev] | [next] | [standalone]
| From | Vir Campestris <vir.campestris@invalid.invalid> |
|---|---|
| Date | 2022-04-25 16:54 +0100 |
| Message-ID | <t46g6s$3e7$1@dont-email.me> |
| In reply to | #83728 |
On 25/04/2022 16:37, James Kuyper wrote: > On 4/25/22 02:00, Juha Nieminen wrote: >> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>> There is no problem with reading NUL bytes in C++, they are not handled >>> specially in any way, unlike in C. >> >> Still, I wonder if the file stream should be opened in binary mode >> (instead of the default). I have experiences of this making a difference >> with the Microsoft compilers in Windows. It can do funky things to >> binary input if the file stream isn't opened in binary mode. > > The C++ standard says very little about the differences between text > mode and binary mode, cross-referencing the C standard for such issues. > The C standard says something important to this issue: > > "... Data read in from a text stream will necessarily compare equal to > the data that were earlier written out to that stream only if: the data > consist only of printing characters and the control characters > horizontal tab and new-line; no new-line character is immediately > preceded by space characters; and the last character is a new-line > character. ..." (7.21.2p2). > > As a general rule, you should not write data to a text stream if it > contains any feature invalidating that guarantee, and you should not > read data from a text stream if it contains any of those things. Windows text files have 13,10 as a newline sequence, not just 10 as in Linux etc. This dates back to at least CP/M where the byte stream in the file exactly matched the one you would send to your CRT or printer, and a line feed really did move the roller but not the print head. I have a vague feeling that Apple might use 10,13 - certainly I've met it somewhere in the past. The point of text mode is to allow a program to ignore these differences by allowing the runtime to handle them. It's not 100%. Andy
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-25 12:29 -0400 |
| Message-ID | <t46i9f$mnk$1@dont-email.me> |
| In reply to | #83733 |
On 4/25/22 11:54, Vir Campestris wrote: > On 25/04/2022 16:37, James Kuyper wrote: ... >> The C++ standard says very little about the differences between text >> mode and binary mode, cross-referencing the C standard for such issues. >> The C standard says something important to this issue: >> >> "... Data read in from a text stream will necessarily compare equal to >> the data that were earlier written out to that stream only if: the data >> consist only of printing characters and the control characters >> horizontal tab and new-line; no new-line character is immediately >> preceded by space characters; and the last character is a new-line >> character. ..." (7.21.2p2). >> >> As a general rule, you should not write data to a text stream if it >> contains any feature invalidating that guarantee, and you should not >> read data from a text stream if it contains any of those things. > > Windows text files have 13,10 as a newline sequence, not just 10 as in > Linux etc. > > This dates back to at least CP/M where the byte stream in the file > exactly matched the one you would send to your CRT or printer, and a > line feed really did move the roller but not the print head. > > I have a vague feeling that Apple might use 10,13 - certainly I've met > it somewhere in the past. > > The point of text mode is to allow a program to ignore these differences > by allowing the runtime to handle them. > > It's not 100%. Those requirements explicitly allow new-line characters be writable to text streams, and that when they are read in again, the result must compare equal to '\n'. This means that on platforms where a new-line is represented in the underlying file by "\n\r" or by "\r\n", then standard I/O routines must convert '\n' to that sequence on write, and convert that sequence back to '\n' when reading from the same stream. If the "not 100%' you're talking about violates that guarantee, then what you're talking about isn't a conforming implementation of C++.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-04-25 17:29 +0000 |
| Message-ID | <2UA9K.5190$Awz.3215@fx03.iad> |
| In reply to | #83738 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes: >On 4/25/22 11:54, Vir Campestris wrote: >> On 25/04/2022 16:37, James Kuyper wrote: >... >>> The C++ standard says very little about the differences between text >>> mode and binary mode, cross-referencing the C standard for such issues. >>> The C standard says something important to this issue: >>> >>> "... Data read in from a text stream will necessarily compare equal to >>> the data that were earlier written out to that stream only if: the data >>> consist only of printing characters and the control characters >>> horizontal tab and new-line; no new-line character is immediately >>> preceded by space characters; and the last character is a new-line >>> character. ..." (7.21.2p2). >>> >> >> Windows text files have 13,10 as a newline sequence, not just 10 as in >> Linux etc. >> >> This dates back to at least CP/M where the byte stream in the file It was used by DEC systems, IBM and Burroughs systems (when communicating with teletypes and other non-line-printer output devices) even before 1974 when CP/M was developed, although generally not part of the file for the non-DEC systems, rather added by the operating system line discipline functions.
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2022-04-25 20:17 +0200 |
| Message-ID | <t46ojk$1nv4$1@gioia.aioe.org> |
| In reply to | #83738 |
On 4/25/2022 6:29 PM, James Kuyper wrote:
> On 4/25/22 11:54, Vir Campestris wrote:
>> On 25/04/2022 16:37, James Kuyper wrote:
> ...
>>> The C++ standard says very little about the differences between text
>>> mode and binary mode, cross-referencing the C standard for such issues.
>>> The C standard says something important to this issue:
>>>
>>> "... Data read in from a text stream will necessarily compare equal to
>>> the data that were earlier written out to that stream only if: the data
>>> consist only of printing characters and the control characters
>>> horizontal tab and new-line; no new-line character is immediately
>>> preceded by space characters; and the last character is a new-line
>>> character. ..." (7.21.2p2).
>>>
>>> As a general rule, you should not write data to a text stream if it
>>> contains any feature invalidating that guarantee, and you should not
>>> read data from a text stream if it contains any of those things.
>>
>> Windows text files have 13,10 as a newline sequence, not just 10 as in
>> Linux etc.
>>
>> This dates back to at least CP/M where the byte stream in the file
>> exactly matched the one you would send to your CRT or printer, and a
>> line feed really did move the roller but not the print head.
>>
>> I have a vague feeling that Apple might use 10,13 - certainly I've met
>> it somewhere in the past.
>>
>> The point of text mode is to allow a program to ignore these differences
>> by allowing the runtime to handle them.
>>
>> It's not 100%.
>
> Those requirements explicitly allow new-line characters be writable to
> text streams, and that when they are read in again, the result must
> compare equal to '\n'. This means that on platforms where a new-line is
> represented in the underlying file by "\n\r" or by "\r\n", then standard
> I/O routines must convert '\n' to that sequence on write, and convert
> that sequence back to '\n' when reading from the same stream.
> If the "not 100%' you're talking about violates that guarantee, then
> what you're talking about isn't a conforming implementation of C++.
>
>
MS-Windows (and DOS before that) behave as you describe ('\n' is stored
as '\r\n' on disk and converted back to '\n' upon read) - or, more
properly, the MS C runtime does that.
So, according to your quote above from the C standard that allows for
only LF and TAB for round trip consistency of text files, MS
implementations have been standard compliant since the beginning of time.
That said, interoperability when dealing with text streams between
Windows and other systems is still painful.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2022-04-25 10:19 -0700 |
| Message-ID | <87o80paxnb.fsf@nosuchdomain.example.com> |
| In reply to | #83733 |
Vir Campestris <vir.campestris@invalid.invalid> writes:
[...]
> Windows text files have 13,10 as a newline sequence, not just 10 as in
> Linux etc.
>
> This dates back to at least CP/M where the byte stream in the file
> exactly matched the one you would send to your CRT or printer, and a
> line feed really did move the roller but not the print head.
>
> I have a vague feeling that Apple might use 10,13 - certainly I've met
> it somewhere in the past.
Versions of MacOS before version X (10) used CR as a line terminator.
OS X, since it's Unix-based, uses LF.
> The point of text mode is to allow a program to ignore these
> differences by allowing the runtime to handle them.
>
> It's not 100%.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
Working, but not speaking, for Philips
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 10:41 -0700 |
| Message-ID | <t46mfm$qkp$1@dont-email.me> |
| In reply to | #83733 |
On 4/25/2022 8:54 AM, Vir Campestris wrote: > > The point of text mode is to allow a program to ignore these differences > by allowing the runtime to handle them. > "Ignoring these differences" comes with a price, like restrictions on absolute positioning. E.g. one's only allowed to `fseek` to 0 or to a position previously obtained from `ftell`. Arbitrary relative/absolute positioning is not supported in text streams. These limitations, if I remember correctly, transfer to C++ I/O library unchanged since it refers to C's `fseek/ftell' specification where it matters. In other words, you can't just completely and transparently "ignore" all the differences. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-04-25 21:58 +0300 |
| Message-ID | <t46r01$1nq$1@dont-email.me> |
| In reply to | #83733 |
25.04.2022 18:54 Vir Campestris kirjutas: > > Windows text files have 13,10 as a newline sequence, not just 10 as in > Linux etc. > > This dates back to at least CP/M where the byte stream in the file > exactly matched the one you would send to your CRT or printer, and a > line feed really did move the roller but not the print head. > > I have a vague feeling that Apple might use 10,13 - certainly I've met > it somewhere in the past. It was 13 in old Mac, and 10 since Mac OS X, as it is a genuine Unix. > > The point of text mode is to allow a program to ignore these differences > by allowing the runtime to handle them. This only works if the text file has been produced on a platform with the same conventions as the currently running program. Nowadays files are routinely copied and accessed all over the world and almost always one needs to support reading text files produced by "wrong" conventions. And there is no guarantee the runtime handles them correctly in text mode, rather the opposite. Even Notepad can read unix text files nowadays, so it's obviously not that hard ;-) Once upon a time there was also FTP text mode which was meant to cope with these platform-specific differences, but this approach caused more problems than solved.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-04-25 15:56 +0000 |
| Message-ID | <Kwz9K.23424$x9Ea.11744@fx45.iad> |
| In reply to | #83728 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes: >On 4/25/22 02:00, Juha Nieminen wrote: >> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>> There is no problem with reading NUL bytes in C++, they are not handled >>> specially in any way, unlike in C. >> >> Still, I wonder if the file stream should be opened in binary mode >> (instead of the default). I have experiences of this making a difference >> with the Microsoft compilers in Windows. It can do funky things to >> binary input if the file stream isn't opened in binary mode. > >The C++ standard says very little about the differences between text >mode and binary mode, cross-referencing the C standard for such issues. >The C standard says something important to this issue: > >"... Data read in from a text stream will necessarily compare equal to >the data that were earlier written out to that stream only if: the data >consist only of printing characters and the control characters >horizontal tab and new-line; no new-line character is immediately >preceded by space characters; and the last character is a new-line >character. ..." (7.21.2p2). > >As a general rule, you should not write data to a text stream if it >contains any feature invalidating that guarantee, and you should not >read data from a text stream if it contains any of those things. As a general rule, your general rule only applies to Windows. On Unix and Unix-like systems, there is no functional difference between text and binary.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-25 13:12 -0400 |
| Message-ID | <t46kq1$c2j$2@dont-email.me> |
| In reply to | #83735 |
On 4/25/22 11:56, Scott Lurndal wrote: > James Kuyper <jameskuyper@alumni.caltech.edu> writes: ... >> The C++ standard says very little about the differences between text >> mode and binary mode, cross-referencing the C standard for such issues. >> The C standard says something important to this issue: >> >> "... Data read in from a text stream will necessarily compare equal to >> the data that were earlier written out to that stream only if: the data >> consist only of printing characters and the control characters >> horizontal tab and new-line; no new-line character is immediately >> preceded by space characters; and the last character is a new-line >> character. ..." (7.21.2p2). >> >> As a general rule, you should not write data to a text stream if it >> contains any feature invalidating that guarantee, and you should not >> read data from a text stream if it contains any of those things. > > As a general rule, your general rule only applies to Windows. On Unix > and Unix-like systems, there is no functional difference between text > and binary. True - but if you want your code to be portable to other systems, you should follow that rule. It's not exactly an onerous requirement.
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 11:03 -0700 |
| Message-ID | <t46np1$5uj$1@dont-email.me> |
| In reply to | #83735 |
On 4/25/2022 8:56 AM, Scott Lurndal wrote: >> >> As a general rule, you should not write data to a text stream if it >> contains any feature invalidating that guarantee, and you should not >> read data from a text stream if it contains any of those things. > > As a general rule, your general rule only applies to Windows. On Unix > and Unix-like systems, there is no functional difference between text > and binary. That's catastrophically incorrect on more than one level. This is the same fallacy that is often used by numerous "UB deniers", claiming that on their hardware/OS platform some forms of "undefined behavior" are supposedly "perfectly defined" because they know how their platform/hardware/OS will react in some specific situation. They just can't grasp this little bit of understanding that the primary source of "undefined" in undefined behavior in _not_ their hardware/OS. It is the abstract language semantics implemented by the compiler. The compiler that actively uses UB as an optimization opportunity at a rather abstract level. It doesn't even get to the hardware/OS level. Hardware/OS matter very little. The same reasoning can be applied to the matter of binary vs. text streams. For example, the very moment the abstract specification of text stream will provide optimization opportunities (or any other benefits) to the implementation of C or C++ standard library, this implementation might/will take advantage of these opportunities to make text streams work more efficiently than binary streams (within the spec, of course). It will break text vs. binary compatibility and it will do this with complete disregard to the fact that the underlying OS makes no difference between text and binary. I'm exaggerating, of course, because it something like this ever happens to Unix-like I/O, the raging uproar the affected parties will be deafening. (We've heard it before with strict aliasing, strict overflow, `memcpy` that copies backwards etc.) Nobody will dare to break binary-text equivalence of Unix-like systems (and POSIX won't allow this). But in essence... this is just a wall of entrenched incompetence too high to climb over. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-04-25 19:05 +0000 |
| Message-ID | <0iC9K.123255$Kdf.44786@fx96.iad> |
| In reply to | #83749 |
Andrey Tarasevich <andreytarasevich@hotmail.com> writes: >On 4/25/2022 8:56 AM, Scott Lurndal wrote: >>> >>> As a general rule, you should not write data to a text stream if it >>> contains any feature invalidating that guarantee, and you should not >>> read data from a text stream if it contains any of those things. >> >> As a general rule, your general rule only applies to Windows. On Unix >> and Unix-like systems, there is no functional difference between text >> and binary. > >That's catastrophically incorrect on more than one level. Actually, that's POSIX defined behavior. Rest of screed elided.
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | comp.lang.c++
csiph-web