Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #83684 > unrolled thread
| Started by | Jivanmukta <jivanmukta@poczta.onet.pl> |
|---|---|
| First post | 2022-04-23 19:07 +0200 |
| Last post | 2022-04-26 11:15 +0300 |
| Articles | 20 on this page of 57 — 16 participants |
Back to article view | Back to comp.lang.c++
getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:07 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 10:17 -0700
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:40 +0200
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-23 19:59 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-23 21:47 +0300
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-25 06:00 +0000
Re: getline() problem Öö Tiib <ootiib@hot.ee> - 2022-04-25 00:12 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 11:40 +0300
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 11:37 -0400
Re: getline() problem Vir Campestris <vir.campestris@invalid.invalid> - 2022-04-25 16:54 +0100
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 12:29 -0400
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 17:29 +0000
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:17 +0200
Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-25 10:19 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:41 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 21:58 +0300
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 15:56 +0000
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-25 13:12 -0400
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:03 -0700
Re: getline() problem scott@slp53.sl.home (Scott Lurndal) - 2022-04-25 19:05 +0000
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 12:48 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 21:46 +0100
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 05:43 +0000
Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 09:36 +0200
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 09:18 +0000
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-26 02:47 -0700
Re: getline() problem Juha Nieminen <nospam@thanks.invalid> - 2022-04-26 11:06 +0000
Re: getline() problem David Brown <david.brown@hesbynett.no> - 2022-04-26 15:52 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-26 07:04 -0700
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-26 16:49 +0200
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 12:01 -0400
Re: getline() problem James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-04-26 11:49 -0400
Re: getline() problem Manfred <noname@add.invalid> - 2022-04-25 20:23 +0200
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 10:33 -0700
Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 10:52 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 11:08 -0700
Re: getline() problem "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-04-25 21:18 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 22:04 -0700
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:06 +0300
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-25 13:02 -0700
Re: getline() problem Barry Schwarz <schwarzb@delq.com> - 2022-04-23 12:54 -0700
Re: getline() problem Montmorency <none@none.com> - 2022-04-23 14:45 -0700
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-23 16:49 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 01:04 +0100
Re: getline() problem Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2022-04-23 18:24 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-24 03:16 +0100
Re: getline() problem Malcolm McLean <malcolm.arthur.mclean@gmail.com> - 2022-04-24 22:51 -0700
Re: getline() problem Ben <ben.usenet@bsb.me.uk> - 2022-04-25 11:07 +0100
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 17:15 -0700
Re: getline() problem Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-04-23 18:23 -0700
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-24 05:28 +0200
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 12:41 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 14:33 +0300
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-25 18:53 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-25 22:08 +0300
Re: getline() problem Jivanmukta <jivanmukta@poczta.onet.pl> - 2022-04-26 09:01 +0200
Re: getline() problem Paavo Helde <eesnimi@osa.pri.ee> - 2022-04-26 11:15 +0300
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 12:48 -0700 |
| Message-ID | <t46ttv$pr2$1@dont-email.me> |
| In reply to | #83754 |
On 4/25/2022 12:05 PM, Scott Lurndal wrote: > Andrey Tarasevich <andreytarasevich@hotmail.com> writes: >> On 4/25/2022 8:56 AM, Scott Lurndal wrote: >>>> >>>> As a general rule, you should not write data to a text stream if it >>>> contains any feature invalidating that guarantee, and you should not >>>> read data from a text stream if it contains any of those things. >>> >>> As a general rule, your general rule only applies to Windows. On Unix >>> and Unix-like systems, there is no functional difference between text >>> and binary. >> >> That's catastrophically incorrect on more than one level. > > Actually, that's POSIX defined behavior. > Rest of screed elided. Firstly, you need to work on your reading-&-comprehension-101, since the matter of POSIX is covered in the "elided screed". Secondly, "POSIX defined behavior" here matters as much at was covered by me, i.e. not much, if at all. So, I re-quote the "elided screed" (you're welcome), since it is mandatory reading and mandatory learning for everyone. One day I might decide to hold a quiz here, and I assure you that question about this will be part of it > This is the same fallacy that is often used by numerous "UB deniers", claiming that on their hardware/OS platform some forms of "undefined behavior" are supposedly "perfectly defined" because they know how their platform/hardware/OS will react in some specific situation. > > They just can't grasp this little bit of understanding that the primary source of "undefined" in undefined behavior in _not_ their hardware/OS. It is the abstract language semantics implemented by the compiler. The compiler that actively uses UB as an optimization opportunity at a rather abstract level. It doesn't even get to the hardware/OS level. Hardware/OS matter very little. > > The same reasoning can be applied to the matter of binary vs. text streams. > > For example, the very moment the abstract specification of text stream will provide optimization opportunities (or any other benefits) to the implementation of C or C++ standard library, this implementation might/will take advantage of these opportunities to make text streams work more efficiently than binary streams (within the spec, of course). It will break text vs. binary compatibility and it will do this with complete disregard to the fact that the underlying OS makes no difference between text and binary. > > I'm exaggerating, of course, because it something like this ever happens to Unix-like I/O, the raging uproar the affected parties will be deafening. (We've heard it before with strict aliasing, strict overflow, `memcpy` that copies backwards etc.) Nobody will dare to break binary-text equivalence of Unix-like systems (and POSIX won't allow this). But in essence... this is just a wall of entrenched incompetence too high to climb over. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | Ben <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-04-25 21:46 +0100 |
| Message-ID | <8735i0uc0c.fsf@bsb.me.uk> |
| In reply to | #83754 |
scott@slp53.sl.home (Scott Lurndal) writes: > Andrey Tarasevich <andreytarasevich@hotmail.com> writes: >>On 4/25/2022 8:56 AM, Scott Lurndal wrote: >>>> >>>> As a general rule, you should not write data to a text stream if it >>>> contains any feature invalidating that guarantee, and you should not >>>> read data from a text stream if it contains any of those things. >>> >>> As a general rule, your general rule only applies to Windows. On Unix >>> and Unix-like systems, there is no functional difference between text >>> and binary. >> >>That's catastrophically incorrect on more than one level. > > Actually, that's POSIX defined behavior. ACK. I was going to say the same in way too many more words. AT seems to think that writing code targeted exclusively at POSIX systems is absurd. > Rest of screed elided. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-04-26 05:43 +0000 |
| Message-ID | <t480p6$1gdc$1@gioia.aioe.org> |
| In reply to | #83749 |
Andrey Tarasevich <andreytarasevich@hotmail.com> wrote: > It is the abstract language semantics implemented by the compiler. The > compiler that actively uses UB as an optimization opportunity at a > rather abstract level. By the way, doesn't this break the rule that optimizations must not change the observable behavior of the program (other than by its execution time)? (Granted, I haven't actually checked that the standard actually mandates that "optimizations must not change observable behavior". I'm not sure where I have got that idea from. I would assume that the standard does require this, in one form of another. Or does UB "take precedence" over that requirement, making it void?)
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-04-26 09:36 +0200 |
| Message-ID | <t487du$7ps$1@dont-email.me> |
| In reply to | #83765 |
On 26/04/2022 07:43, Juha Nieminen wrote:
> Andrey Tarasevich <andreytarasevich@hotmail.com> wrote:
>> It is the abstract language semantics implemented by the compiler. The
>> compiler that actively uses UB as an optimization opportunity at a
>> rather abstract level.
>
> By the way, doesn't this break the rule that optimizations must not change
> the observable behavior of the program (other than by its execution time)?
>
> (Granted, I haven't actually checked that the standard actually mandates
> that "optimizations must not change observable behavior". I'm not sure
> where I have got that idea from. I would assume that the standard does
> require this, in one form of another. Or does UB "take precedence" over
> that requirement, making it void?)
UB "takes precedence" over /everything/ - it basically says there are no
guarantees of anything. There was an attempt to distinguish between
"bounded" and "unbounded" undefined behaviour in C (Annex L in the C11
standards), but AFAIK no one implemented any of it. The trouble is,
even "bounded" undefined behaviour can quickly cause unbounded knock-on
effects. A relatively innocent-seeming arithmetic overflow might result
in an incorrect pointer, which then results in a write to a completely
unexpected part of the memory, and running unexpected code (imagine that
a function pointer got stomped on). You can easily see how this could
affect the observable behaviour - perhaps the address written was a
volatile variable. You can see how it could affect even previous
observable behaviour - your code may have written to a file, then run
riot after UB and entered code that deleted the file, affecting
historical observable behaviour before it even hit the disk surface.
Obviously one can hope that the worst consequences of UB are the least
likely to happen, but the implementation can't rule out anything (unless
it specifically adds features to turn UB into defined behaviour - such
as gcc -fwrapv, or sanitizers, etc.) The standards basically refuse to
accept any responsibility for what might happen.
Trying to put strict limits on UB, or even to guarantee that past
observable behaviour is still valid after the UB is executed, requires
locking down the run-time environment significantly. You need a virtual
machine with strict limitations, or dedicated hardware (clearly things
like memory protection units, fault traps, etc., go a long way towards
this in practice - but not all C and C++ targets have these). If you
want a managed language or managed runtime environment, there are many
available - C and C++ are aimed for maximal efficiency, not maximal
safety. ("With great power comes great responsibility", for the
superhero fans.)
That doesn't mean that an implementation should not try to limit the
consequences of executing UB, but it /does/ mean that it is not obliged
to do so by the standards and can pick a balance between improving the
efficiency of correct code verses being as helpful as possible to the
developer in the face of incorrect code.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-04-26 09:18 +0000 |
| Message-ID | <t48de1$11f9$1@gioia.aioe.org> |
| In reply to | #83768 |
David Brown <david.brown@hesbynett.no> wrote:
> UB "takes precedence" over /everything/ - it basically says there are no
> guarantees of anything.
The more I think about it, the more ridiculous that becomes.
If the compiler is allowed to do *anything* it wants when it encounters
Undefined Behavior, doesn't that mean that it can, for example, compile
the entire program into a single "main() { return 0; }"?
Imagine a million-line program where the compiler finds one single
extremely obscure and mostly innocuous case of UB (eg. overflowing
of a signed integer), and that allows it to optimize the *entire program*
away, into a literal on-liner that does absolutely nothing?
After all, if there are no guarantees of *anything*, and the compiler can
do whatever it wants, optimizing the entire program away is "whatever the
compiler wants", and part of "no guarantees of anything".
Nobody would use such a compiler. There has to be *some* limits to what
the compiler is allowed to do even if it encounters some very minor and
innocuous "technically speaking, if we take the absolute strict letter
of the standard, this is undefined behavior".
(Even if there's some kind of caveat like "the undefined behavior can
only happen if that particular piece of code is executed", it's perfectly
possible for the minor UB (eg. signed integer overflow) to happen at
the beginning of main(), after which "nothing is guaranteed". Meaning
that the compiler could just optimize the entire rest of the million-line
program into a single "return 0;" Nobody would use such a compiler.)
[toc] | [prev] | [next] | [standalone]
| From | Malcolm McLean <malcolm.arthur.mclean@gmail.com> |
|---|---|
| Date | 2022-04-26 02:47 -0700 |
| Message-ID | <f65d5129-0b35-4732-b82c-f284726cd66an@googlegroups.com> |
| In reply to | #83772 |
On Tuesday, 26 April 2022 at 10:19:14 UTC+1, Juha Nieminen wrote:
> David Brown <david...@hesbynett.no> wrote:
> > UB "takes precedence" over /everything/ - it basically says there are no
> > guarantees of anything.
> The more I think about it, the more ridiculous that becomes.
>
> If the compiler is allowed to do *anything* it wants when it encounters
> Undefined Behavior, doesn't that mean that it can, for example, compile
> the entire program into a single "main() { return 0; }"?
>
> Imagine a million-line program where the compiler finds one single
> extremely obscure and mostly innocuous case of UB (eg. overflowing
> of a signed integer), and that allows it to optimize the *entire program*
> away, into a literal on-liner that does absolutely nothing?
>
> After all, if there are no guarantees of *anything*, and the compiler can
> do whatever it wants, optimizing the entire program away is "whatever the
> compiler wants", and part of "no guarantees of anything".
>
> Nobody would use such a compiler. There has to be *some* limits to what
> the compiler is allowed to do even if it encounters some very minor and
> innocuous "technically speaking, if we take the absolute strict letter
> of the standard, this is undefined behavior".
>
> (Even if there's some kind of caveat like "the undefined behavior can
> only happen if that particular piece of code is executed", it's perfectly
> possible for the minor UB (eg. signed integer overflow) to happen at
> the beginning of main(), after which "nothing is guaranteed". Meaning
> that the compiler could just optimize the entire rest of the million-line
> program into a single "return 0;" Nobody would use such a compiler.)
>
Basically you want a severe warning, because the programmer might have
deliberately caused the undefined behaviour because he knows what it
will do on his platform. He might want to set the carry flag before
jumping into an assembler routine, for example.
But normally it's best if UB does something dramatic and fatal to normal
execution, because no results are usually better than wrong results, and
no executable is usually better than a buggy executable.
[toc] | [prev] | [next] | [standalone]
| From | Juha Nieminen <nospam@thanks.invalid> |
|---|---|
| Date | 2022-04-26 11:06 +0000 |
| Message-ID | <t48jo1$32b$1@gioia.aioe.org> |
| In reply to | #83775 |
Malcolm McLean <malcolm.arthur.mclean@gmail.com> wrote:
> Basically you want a severe warning, because the programmer might have
> deliberately caused the undefined behaviour because he knows what it
> will do on his platform. He might want to set the carry flag before
> jumping into an assembler routine, for example.
>
> But normally it's best if UB does something dramatic and fatal to normal
> execution, because no results are usually better than wrong results, and
> no executable is usually better than a buggy executable.
One rather interesting example of UB doing something quite unexpected
is this code, when compiled with clang 13.0.0:
//--------------------------------------------
#include <stdio.h>
void f1(void) {
for(int i = 0; i >= 0; i++) {
// Undefined behavior
}
}
void f2(void) {
puts("Formatting /dev/sda1...");
// system("mkfs -t btrfs -f /dev/sda1");
}
void (*volatile p1)(void) = f1;
void (*volatile p2)(void) = f2;
int main(void) {
puts(__VERSION__);
p1();
return 0;
}
//------------------------------------------
That "p1()" might look like it's calling the f1() function, but ends up
executing the f2() function instead (because for some reason this version
of clang compiles f1() into nothing. Not even a 'ret'. And f2() immediately
follows it, so...)
Source: https://bugs.llvm.org/show_bug.cgi?id=49599
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2022-04-26 15:52 +0200 |
| Message-ID | <t48tea$e4a$1@dont-email.me> |
| In reply to | #83772 |
On 26/04/2022 11:18, Juha Nieminen wrote:
> David Brown <david.brown@hesbynett.no> wrote:
>> UB "takes precedence" over /everything/ - it basically says there are no
>> guarantees of anything.
>
> The more I think about it, the more ridiculous that becomes.
>
> If the compiler is allowed to do *anything* it wants when it encounters
> Undefined Behavior, doesn't that mean that it can, for example, compile
> the entire program into a single "main() { return 0; }"?
Yes.
But an implementation that does that is unlikely to be very popular!
>
> Imagine a million-line program where the compiler finds one single
> extremely obscure and mostly innocuous case of UB (eg. overflowing
> of a signed integer), and that allows it to optimize the *entire program*
> away, into a literal on-liner that does absolutely nothing?
>
> After all, if there are no guarantees of *anything*, and the compiler can
> do whatever it wants, optimizing the entire program away is "whatever the
> compiler wants", and part of "no guarantees of anything".
>
> Nobody would use such a compiler. There has to be *some* limits to what
> the compiler is allowed to do even if it encounters some very minor and
> innocuous "technically speaking, if we take the absolute strict letter
> of the standard, this is undefined behavior".
>
> (Even if there's some kind of caveat like "the undefined behavior can
> only happen if that particular piece of code is executed", it's perfectly
> possible for the minor UB (eg. signed integer overflow) to happen at
> the beginning of main(), after which "nothing is guaranteed". Meaning
> that the compiler could just optimize the entire rest of the million-line
> program into a single "return 0;" Nobody would use such a compiler.)
(You can have code in your program that, if executed, would have UB.
The problem only comes when trying to execute it.)
I think you are - as many people seem to do when thinking about UB -
taking your concerns to the extremes.
In the programming world, there is the phrase "garbage in, garbage out"
- that's all this is, nothing more. It's not new or special (Charles
Babbage wrote about the concept in regard to his mechanical difference
engine). Programming is about taking a specification, and writing code
that implements the specification.
If I say "write a function "int_square_root" that takes an int between 0
and 1000 and returns its square root, rounded down", then you can write
code that does that. What should I expect of the code if I call it with
an argument of 1001? Or -1? I have /no/ expectations, and you give me
/no/ guarantees. Maybe the function won't return. Maybe it will crash.
Maybe it will give a "reasonable" result. Who knows?
A compiler is a program with the specification "given an input that is
valid source code conforming to the standards, generate object code that
implements the input program". If your code tries to execute undefined
behaviour, you are asking the compiler for something outside its
specifications. You are asking for the integer that is the square root
of -1. You are asking for the object code for source code that does not
make sense according to the language definitions. The compiler can't
give you an answer, because there is no answer.
Maybe the compiler can make some guesses, or give you helpful feedback,
or skip that bit and carry on. But there is no way the compiler can
give you a /correct/ answer, and the standards abdicate any guarantees
on the matter.
A compiler that spots inevitable UB and generates an empty program might
be correct, but it is unlikely to be popular. One that gives an error
message during compilation would more helpful to most people. Because
the standards give no definition, compilers are free to do as they want
here - such as generate run-time errors to aid debugging. (Contrast
this with languages like Java that attempt to eliminate UB - common
errors such as integer overflow are defined behaviour and the compiler
must do what the language standards require, rather than helping the
programmer find the bugs.) Equally, the compiler can assume states that
lead to UB cannot occur if that aids optimisation, and at no point do
you get any guarantee about what might happen because the compiler
cannot promise to give you the result you expect from what is
essentially meaningless code.
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-26 07:04 -0700 |
| Message-ID | <t48u5u$lgv$1@dont-email.me> |
| In reply to | #83772 |
On 4/26/2022 2:18 AM, Juha Nieminen wrote:
> David Brown <david.brown@hesbynett.no> wrote:
>> UB "takes precedence" over /everything/ - it basically says there are no
>> guarantees of anything.
>
> The more I think about it, the more ridiculous that becomes.
>
> If the compiler is allowed to do *anything* it wants when it encounters
> Undefined Behavior, doesn't that mean that it can, for example, compile
> the entire program into a single "main() { return 0; }"?
>
Absolutely! Or it can refuse to compile the program entirely.
Still there are some ambiguities in the way it is described in the
standard.
On the one hand, it is clear that localized undefined behavior is
permitted to "poison" the whole program.
On the other hand, it appears that in accordance to the intent undefined
code triggers the "actual" undefined behavior only when/if control
passes through the problematic construct. The standard text seems to
imply that undefined code has no harmful effects as long as it is never
executed. And, obviously, whether some portion of the code is executed
or not is generally a run-time matter. Does a piece of _potentially_
executed undefined code unconditionally make the whole program
undefined, i.e. allows transformation into `int main() { return 0; }` at
compile time?
Taking the above into account, a real-life compiler that takes such
"uncompromising" stance towards undefined behavior (e.g. discards the
code) will likely be forced to carefully determine a range of code that
is "poisoned" by each specific undefined construct it managed to
definitively detect. The compiler will have to "localize" the
corresponding transformations of the code.
P.S. Obviously, issuing a diagnostic and refusing to compile the program
in response to detected UB is a non-localized outcome. One can argue
that such outcome contradicts the above intent if the undefined coide is
never executed. But I think that producing a diagnostic and refusing to
compile as a consequence of UB stands apart from other manifestations of
UB. From purely practical point of view most users will probably accept
such outcome favorably, especially if there's is a workaround for
special cases.
--
Best regards,
Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2022-04-26 16:49 +0200 |
| Message-ID | <t490qj$g1a$1@gioia.aioe.org> |
| In reply to | #83772 |
On 4/26/2022 11:18 AM, Juha Nieminen wrote:
> David Brown <david.brown@hesbynett.no> wrote:
>> UB "takes precedence" over /everything/ - it basically says there are no
>> guarantees of anything.
>
> The more I think about it, the more ridiculous that becomes.
>
> If the compiler is allowed to do *anything* it wants when it encounters
> Undefined Behavior, doesn't that mean that it can, for example, compile
> the entire program into a single "main() { return 0; }"?
Technically yes, it can. Others have already said that.
That is why the very concept of Undefined Behaviour keeps raising questions.
>
> Imagine a million-line program where the compiler finds one single
> extremely obscure and mostly innocuous case of UB (eg. overflowing
> of a signed integer), and that allows it to optimize the *entire program*
> away, into a literal on-liner that does absolutely nothing?
>
> After all, if there are no guarantees of *anything*, and the compiler can
> do whatever it wants, optimizing the entire program away is "whatever the
> compiler wants", and part of "no guarantees of anything".
>
> Nobody would use such a compiler. There has to be *some* limits to what
> the compiler is allowed to do even if it encounters some very minor and
> innocuous "technically speaking, if we take the absolute strict letter
> of the standard, this is undefined behavior".
"Minor" is a qualitative and subjective attribute. Humans know that,
machines don't. Compiler writers might make choices about how to handle
some cases of UB, but then all you get is implementation defined
behavior, and non portable code.
My point is that the real burden of dealing with UB is not (and it
cannot be) on the compiler. Rather, it is on the people who write the
standard, which actually get to decide what is UB and what is not.
>
> (Even if there's some kind of caveat like "the undefined behavior can
> only happen if that particular piece of code is executed", it's perfectly
> possible for the minor UB (eg. signed integer overflow) to happen at
> the beginning of main(), after which "nothing is guaranteed". Meaning
> that the compiler could just optimize the entire rest of the million-line
> program into a single "return 0;" Nobody would use such a compiler.)
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-26 12:01 -0400 |
| Message-ID | <t4951n$gdm$1@dont-email.me> |
| In reply to | #83772 |
On 4/26/22 05:18, Juha Nieminen wrote:
> David Brown <david.brown@hesbynett.no> wrote:
>> UB "takes precedence" over /everything/ - it basically says there are no
>> guarantees of anything.
>
> The more I think about it, the more ridiculous that becomes.
>
> If the compiler is allowed to do *anything* it wants when it encounters
> Undefined Behavior, doesn't that mean that it can, for example, compile
> the entire program into a single "main() { return 0; }"?
If it's compile-time undefined behavior, or run-time undefined behavior
that will unavoidably occur some time after program start, yes, it can.
> Imagine a million-line program where the compiler finds one single
> extremely obscure and mostly innocuous case of UB (eg. overflowing
> of a signed integer), and that allows it to optimize the *entire program*
> away, into a literal on-liner that does absolutely nothing?
In principle, yes. In practice, in many situations the effect will be
must more limited.
> Nobody would use such a compiler. There has to be *some* limits to what
> the compiler is allowed to do even if it encounters some very minor and
> innocuous "technically speaking, if we take the absolute strict letter
> of the standard, this is undefined behavior".
Everybody uses such a compiler. As a general rule, the behavior is not
undefined unless there's at least one platform (and possibly many) where
it would be unacceptably difficult to diagnose violations of the rules.
The actually behavior that occurs is usually just whatever goes wrong as
a result of the implementation generating code that assumes that the
case with undefined behavior won't come up. It's not uncommon for that
behavior to be catastrophic.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-26 11:49 -0400 |
| Message-ID | <t494b6$a1v$1@dont-email.me> |
| In reply to | #83765 |
On 4/26/22 01:43, Juha Nieminen wrote: > Andrey Tarasevich <andreytarasevich@hotmail.com> wrote: >> It is the abstract language semantics implemented by the compiler. The >> compiler that actively uses UB as an optimization opportunity at a >> rather abstract level. > > By the way, doesn't this break the rule that optimizations must not change > the observable behavior of the program (other than by its execution time)? > > (Granted, I haven't actually checked that the standard actually mandates > that "optimizations must not change observable behavior". I'm not sure > where I have got that idea from. I would assume that the standard does > require this, in one form of another. ... You're confusing two different kinds of optimization. The first kind are the "as-if" optimizations. It's not true that such optimizations can't change the observable behavior, but that's close to being the correct description. For any given piece of code, because of the many things that the standard leaves unspecified, there's often multiple (sometimes, infinitely many) different sequences of observable behavior that are allowed. When a program is well-formed, the implementation is required generate one of those sequences. An optimization can change which of those sequences occurs, but it cannot allow a sequence not on that list to occur. The other kinds of optimizations are enabled by the fact that certain kinds of code can have undefined behavior under some circumstances. An implementation is free to ignore the possibility that those circumstances might occur, which sometimes enables optimizations for code that doesn't have undefined behavior.. For example, `b+a-b` for signed integer types can be optimized to a, because the overflow that might occur when b+a is evaluated has undefined behavior, so the compiler doesn't have to worry about that possibility. If the behavior were defined, the compiler might (depending upon the target platform) have to complicate the code in order to ensure that the defined behavior actually occurs. > ... Or does UB "take precedence" over > that requirement, making it void?) Undefined behavior releases an implementation from ALL requirements imposed by the standard, even than one.
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2022-04-25 20:23 +0200 |
| Message-ID | <t46ova$1tav$1@gioia.aioe.org> |
| In reply to | #83735 |
On 4/25/2022 5:56 PM, Scott Lurndal wrote: > James Kuyper <jameskuyper@alumni.caltech.edu> writes: >> On 4/25/22 02:00, Juha Nieminen wrote: >>> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>>> There is no problem with reading NUL bytes in C++, they are not handled >>>> specially in any way, unlike in C. >>> >>> Still, I wonder if the file stream should be opened in binary mode >>> (instead of the default). I have experiences of this making a difference >>> with the Microsoft compilers in Windows. It can do funky things to >>> binary input if the file stream isn't opened in binary mode. >> >> The C++ standard says very little about the differences between text >> mode and binary mode, cross-referencing the C standard for such issues. >> The C standard says something important to this issue: >> >> "... Data read in from a text stream will necessarily compare equal to >> the data that were earlier written out to that stream only if: the data >> consist only of printing characters and the control characters >> horizontal tab and new-line; no new-line character is immediately >> preceded by space characters; and the last character is a new-line >> character. ..." (7.21.2p2). >> >> As a general rule, you should not write data to a text stream if it >> contains any feature invalidating that guarantee, and you should not >> read data from a text stream if it contains any of those things. > > As a general rule, your general rule only applies to Windows. On Unix > and Unix-like systems, there is no functional difference between text > and binary. Well, the general rule is correctly a general rule as far as the C and C++ standard guarantees go. On Unix and Unix-like systems you may well have other guarantees in action, which are not as general as the C standard.
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 10:33 -0700 |
| Message-ID | <t46m1i$kje$1@dont-email.me> |
| In reply to | #83705 |
On 4/24/2022 11:00 PM, Juha Nieminen wrote: > Paavo Helde <eesnimi@osa.pri.ee> wrote: >> There is no problem with reading NUL bytes in C++, they are not handled >> specially in any way, unlike in C. > > Still, I wonder if the file stream should be opened in binary mode > (instead of the default). I have experiences of this making a difference > with the Microsoft compilers in Windows. It can do funky things to > binary input if the file stream isn't opened in binary mode. No, it can't. All it does on the input side is translate the line endings to '\n', which it properly does. "Funky things" usually come as a consequence of unfounded expectations, like expecting to be able to use absolute positioning in text streams, even though the language standard clearly says that this is not supported (for obvious reasons). -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-25 10:52 -0700 |
| Message-ID | <2b5dfef5-a18c-4a13-acc1-da3466c0cb41n@googlegroups.com> |
| In reply to | #83745 |
On Monday, April 25, 2022 at 1:33:54 PM UTC-4, Andrey Tarasevich wrote: > On 4/24/2022 11:00 PM, Juha Nieminen wrote: > > Paavo Helde <ees...@osa.pri.ee> wrote: > >> There is no problem with reading NUL bytes in C++, they are not handled > >> specially in any way, unlike in C. > > > > Still, I wonder if the file stream should be opened in binary mode > > (instead of the default). I have experiences of this making a difference > > with the Microsoft compilers in Windows. It can do funky things to > > binary input if the file stream isn't opened in binary mode. > No, it can't. All it does on the input side is translate the line > endings to '\n', which it properly does. That might be true of a particular implementation, but in general, data read from a text stream might not compare equal to data previously written to that stream, unless "... the data consist only of printing characters and the control characters horizontal tab and new-line; no new-line character is immediately preceded by space characters; and the last character is a new-line character. ..." (C standard 7.21.2p2, which is incorporated by reference into the C++ standard). There's two ways in which failures to meet those requirements might result in a failure to compare equal: something might be lost, changed, or added when writing to the text stream, or something might be lost, changed, or added when reading in from a text stream. 7.21.2p2 permits both possibilities, by failing to say anything preventing either one. You ignore those possibilities at your own risk (which might be entirely reasonable if you know that the target implementation doesn't do anything of the sort).
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 11:08 -0700 |
| Message-ID | <t46o2r$8o9$1@dont-email.me> |
| In reply to | #83747 |
On 4/25/2022 10:52 AM, james...@alumni.caltech.edu wrote: > On Monday, April 25, 2022 at 1:33:54 PM UTC-4, Andrey Tarasevich wrote: >> On 4/24/2022 11:00 PM, Juha Nieminen wrote: >>> Paavo Helde <ees...@osa.pri.ee> wrote: >>>> There is no problem with reading NUL bytes in C++, they are not handled >>>> specially in any way, unlike in C. >>> >>> Still, I wonder if the file stream should be opened in binary mode >>> (instead of the default). I have experiences of this making a difference >>> with the Microsoft compilers in Windows. It can do funky things to >>> binary input if the file stream isn't opened in binary mode. >> No, it can't. All it does on the input side is translate the line >> endings to '\n', which it properly does. > > That might be true of a particular implementation, but in general, data read > from a text stream might not compare equal to data previously written to that > stream, unless > Yes, and as I said previously, these transformations are allowed on the output side. It is true that a character buffer being written into a text stream might not compare equal to a character buffer later read from the same test stream. This is true because of the transformations it can be subjected to by a standard _output_ function. However, this topic was originally dedicated to the _input_ behavior only. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-04-25 21:18 -0700 |
| Message-ID | <8fe7f117-270d-4595-a3d8-dfc0d449b0fbn@googlegroups.com> |
| In reply to | #83750 |
On Monday, April 25, 2022 at 2:08:42 PM UTC-4, Andrey Tarasevich wrote: > On 4/25/2022 10:52 AM, james...@alumni.caltech.edu wrote: > > On Monday, April 25, 2022 at 1:33:54 PM UTC-4, Andrey Tarasevich wrote: ... > >> No, it can't. All it does on the input side is translate the line > >> endings to '\n', which it properly does. > > > > That might be true of a particular implementation, but in general, data read > > from a text stream might not compare equal to data previously written to that > > stream, unless > > > Yes, and as I said previously, these transformations are allowed on the > output side Please cite the text that restricts the transforms to the output functions. ... > ... It is true that a character buffer being written into a > text stream might not compare equal to a character buffer later read > from the same test stream. Notice that "write" and "read" play equal roles in that specification. You can't assign responsibility for the failure to compare equal uniquely on the output functions; an implementation where it was due to the input functions, or to both, wouldn't violate any requirement specified by the standard. > ... This is true because of the transformations > it can be subjected to by a standard _output_ function. Or by a standard input fuction. > However, this topic was originally dedicated to the _input_ behavior only. Which is therefore sufficient for the problem to occur.
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 22:04 -0700 |
| Message-ID | <t47uhl$hsn$1@dont-email.me> |
| In reply to | #83761 |
On 4/25/2022 9:18 PM, james...@alumni.caltech.edu wrote: > On Monday, April 25, 2022 at 2:08:42 PM UTC-4, Andrey Tarasevich wrote: >> On 4/25/2022 10:52 AM, james...@alumni.caltech.edu wrote: >>> On Monday, April 25, 2022 at 1:33:54 PM UTC-4, Andrey Tarasevich wrote: > ... >>>> No, it can't. All it does on the input side is translate the line >>>> endings to '\n', which it properly does. >>> >>> That might be true of a particular implementation, but in general, data read >>> from a text stream might not compare equal to data previously written to that >>> stream, unless >>> >> Yes, and as I said previously, these transformations are allowed on the >> output side > > Please cite the text that restricts the transforms to the output functions. > ... Look like I was wrong, and I take it back. Now, after rereading the corresponding portion of the specification I agree that it is fairly symmetrical with regard to text input and output. The transformations, like "optimizing out" spaces before the newline, can occur on input as well. -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
| From | Paavo Helde <eesnimi@osa.pri.ee> |
|---|---|
| Date | 2022-04-25 22:06 +0300 |
| Message-ID | <t46rfd$622$1@dont-email.me> |
| In reply to | #83745 |
25.04.2022 20:33 Andrey Tarasevich kirjutas: > On 4/24/2022 11:00 PM, Juha Nieminen wrote: >> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>> There is no problem with reading NUL bytes in C++, they are not handled >>> specially in any way, unlike in C. >> >> Still, I wonder if the file stream should be opened in binary mode >> (instead of the default). I have experiences of this making a difference >> with the Microsoft compilers in Windows. It can do funky things to >> binary input if the file stream isn't opened in binary mode. > > No, it can't. All it does on the input side is translate the line > endings to '\n', which it properly does. It also translates 26 to EOF, leaving the rest of the file unread. This behavior comes as a surprise to many people and can be indeed called "funky".
[toc] | [prev] | [next] | [standalone]
| From | Andrey Tarasevich <andreytarasevich@hotmail.com> |
|---|---|
| Date | 2022-04-25 13:02 -0700 |
| Message-ID | <t46uol$c6$1@dont-email.me> |
| In reply to | #83755 |
On 4/25/2022 12:06 PM, Paavo Helde wrote: > 25.04.2022 20:33 Andrey Tarasevich kirjutas: >> On 4/24/2022 11:00 PM, Juha Nieminen wrote: >>> Paavo Helde <eesnimi@osa.pri.ee> wrote: >>>> There is no problem with reading NUL bytes in C++, they are not handled >>>> specially in any way, unlike in C. >>> >>> Still, I wonder if the file stream should be opened in binary mode >>> (instead of the default). I have experiences of this making a difference >>> with the Microsoft compilers in Windows. It can do funky things to >>> binary input if the file stream isn't opened in binary mode. >> >> No, it can't. All it does on the input side is translate the line >> endings to '\n', which it properly does. > > It also translates 26 to EOF, leaving the rest of the file unread. This > behavior comes as a surprise to many people and can be indeed called > "funky". True. I have to admit that I managed to forget about this treatment for `1A` character. However, practice shows that a lot more people run into "funky" unexpected EOF conditions when they attempt to `getchar()` an `FF` byte into a `signed char` variable, instead of using the proper `int` as a recipient :) -- Best regards, Andrey Tarasevich
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.lang.c++
csiph-web