Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c++ > #82895 > unrolled thread
| Started by | Bonita Montero <Bonita.Montero@gmail.com> |
|---|---|
| First post | 2022-02-04 13:54 +0100 |
| Last post | 2022-02-07 18:30 +0100 |
| Articles | 20 on this page of 103 — 18 participants |
Back to article view | Back to comp.lang.c++
C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-04 13:54 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-04 14:25 +0000
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-04 15:31 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-04 14:54 +0000
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-04 17:41 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-05 11:31 +0000
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-05 12:49 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-05 12:20 +0000
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-04 16:15 +0000
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-04 17:43 +0100
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-05 01:51 +0100
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-05 01:59 +0000
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-04 23:05 -0500
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-05 11:56 +0100
Re: C++20 concepts rocks David Brown <david.brown@hesbynett.no> - 2022-02-05 13:53 +0100
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-05 14:12 +0100
Re: C++20 concepts rocks Anand Hariharan <mailto.anand.hariharan@gmail.com> - 2022-02-12 11:13 -0800
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-12 22:18 -0500
Re: C++20 concepts rocks Richard Damon <Richard@Damon-Family.org> - 2022-02-13 12:51 -0500
Re: C++20 concepts rocks Paavo Helde <eesnimi@osa.pri.ee> - 2022-02-13 20:16 +0200
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-13 14:12 -0500
Re: C++20 concepts rocks Richard Damon <Richard@Damon-Family.org> - 2022-02-13 16:22 -0500
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-13 17:18 -0500
Re: C++20 concepts rocks Öö Tiib <ootiib@hot.ee> - 2022-02-14 00:51 -0800
Re: C++20 concepts rocks "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-02-14 08:29 -0800
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-04 23:37 -0800
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-05 11:58 +0100
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 02:16 -0800
Re: C++20 concepts rocks "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2022-02-05 12:35 -0800
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 05:54 -0800
Re: C++20 concepts rocks Manfred <noname@add.invalid> - 2022-02-06 19:43 +0100
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-04 23:53 -0800
Re: C++20 concepts rocks red floyd <no.spam.here@its.invalid> - 2022-02-05 10:10 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-05 19:51 +0100
Re: C++20 concepts rocks red floyd <no.spam.here@its.invalid> - 2022-02-05 16:42 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-06 01:39 +0000
Re: C++20 concepts rocks red floyd <no.spam.here@its.invalid> - 2022-02-05 18:02 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-06 09:57 +0100
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-06 10:57 +0100
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-06 12:40 +0100
Re: C++20 concepts rocks Öö Tiib <ootiib@hot.ee> - 2022-02-06 03:53 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-06 14:35 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 11:20 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-06 21:53 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 19:03 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-07 11:50 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-07 06:26 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-07 15:48 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-07 12:16 -0800
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-07 22:27 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-09 21:43 +0000
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-09 21:07 -0500
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-09 21:25 -0800
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-10 13:05 +0000
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-10 11:19 -0500
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-11 05:41 -0800
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-11 17:06 +0100
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-11 11:50 -0500
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-11 21:13 +0100
Re: C++20 concepts rocks scott@slp53.sl.home (Scott Lurndal) - 2022-02-11 17:06 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-15 00:10 -0800
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-15 20:58 +0100
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-16 06:40 +0000
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-16 09:54 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-16 09:00 +0000
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-16 20:25 +0100
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-17 08:53 +0000
Re: C++20 concepts rocks Paavo Helde <eesnimi@osa.pri.ee> - 2022-02-17 16:11 +0200
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-17 10:58 -0500
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-17 07:32 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-04-19 02:21 -0700
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-04-19 11:56 +0200
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-04-25 03:02 -0700
Re: C++20 concepts rocks Ben Bacarisse <ben.usenet@bsb.me.uk> - 2022-02-12 00:01 +0000
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-14 23:55 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-07 18:31 +0100
Re: C++20 concepts rocks scott@slp53.sl.home (Scott Lurndal) - 2022-02-07 17:43 +0000
Re: C++20 concepts rocks James Kuyper <jameskuyper@alumni.caltech.edu> - 2022-02-06 14:25 -0500
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 02:01 -0800
Re: C++20 concepts rocks red floyd <no.spam.here@its.invalid> - 2022-02-06 19:37 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-07 10:18 +0100
Re: C++20 concepts rocks Paavo Helde <eesnimi@osa.pri.ee> - 2022-02-07 11:32 +0200
Re: C++20 concepts rocks red floyd <no.spam.here@its.invalid> - 2022-02-07 07:55 -0800
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 02:11 -0800
Re: C++20 concepts rocks "Alf P. Steinbach" <alf.p.steinbach@gmail.com> - 2022-02-06 13:13 +0100
Re: C++20 concepts rocks Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-06 04:58 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-06 14:08 +0100
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-05 12:48 +0100
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-04 20:15 +0000
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-05 11:32 +0000
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-06 08:28 +0000
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-07 09:52 +0000
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-08 07:05 +0000
Re: C++20 concepts rocks Muttley@dastardlyhq.com - 2022-02-08 09:25 +0000
Re: C++20 concepts rocks Juha Nieminen <nospam@thanks.invalid> - 2022-02-08 10:09 +0000
Re: C++20 concepts rocks Andrey Tarasevich <andreytarasevich@hotmail.com> - 2022-02-04 12:10 -0800
Re: C++20 concepts rocks Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-05 13:41 +0100
Re: C++20 concepts rocks "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2022-02-05 19:28 -0800
A little benchmark: Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-05 13:22 +0100
A better benchmark Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-06 14:56 +0100
A improved routine Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-07 15:19 +0100
Re: A better benchmark Tim Rentsch <tr.17687@z991.linuxsc.com> - 2022-02-07 06:57 -0800
Re: A better benchmark Bonita Montero <Bonita.Montero@gmail.com> - 2022-02-07 18:30 +0100
Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6 Next page →
| From | Öö Tiib <ootiib@hot.ee> |
|---|---|
| Date | 2022-02-06 03:53 -0800 |
| Message-ID | <7770625c-6446-4f1e-8d8f-7d86aa90e6cbn@googlegroups.com> |
| In reply to | #82929 |
On Sunday, 6 February 2022 at 11:58:18 UTC+2, alf.p.s...@gmail.com wrote: > On 6 Feb 2022 09:57, Bonita Montero wrote: > > Am 06.02.2022 um 01:42 schrieb red floyd: > >> On 2/5/2022 10:51 AM, Bonita Montero wrote: > >>>> On 2/4/2022 11:53 PM, Tim Rentsch wrote: > >>>>> char *ep = (&result)[1]; > >>> > >>> > Am 05.02.2022 um 19:10 schrieb red floyd: > >>>> Why the oddly unreadable initialization of ep? Why not just > >>>> > >>>> char *ep = result + sizeof(result)? > >>>> > >>> > >>> I also consider (&result)[1] as the more elegant way. > >> > >> To be honest, I don't give a damn about your opinion. > > > > Your solution works only with char-arrays. > > Tims solution works with all arrays. > > Therefore it's preferrable. > Red Floyd used `sizeof` because it was a `char`-array. That avoided an > otherwise needless include directive or DIY function definition. The > standard library offers `std::ssize` (plus the unsigned `std::size` in > C++17 and earlier) for the general case. > > That's portable and clear code. > > The code using `(&result)[1]` involves a couple of extra conversions and > indirection and is thus needlessly complex, plus it's formally UB, > though I guess to most of us that formal UB doesn't matter except as an > example of the improvement potential of the standardization process. It matters to most. Basically (&result)[1] is required to be more elegant way to write *((&result)+1) and that * there looks like dereference that is not allowed.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-06 14:35 +0000 |
| Message-ID | <87mtj4qd0h.fsf@bsb.me.uk> |
| In reply to | #82929 |
"Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
> The code using `(&result)[1]` involves a couple of extra conversions
Eh? char *ep = r + sizeof(r); involves one array-to-pointer conversion
just as char *ep = (&r)[1]; does.
> and indirections and is thus needlessly complex, plus it's formally
> UB,
Can you point to where this is made UB in the C++ standard?
There is an ambiguous phrase in the C standard ("undefined behaviour if
the * is evaluated") but that phrase is entirely missing in the C++
text. And I got lost in the forest of rvalues, lvalues, prvalues,
xvalues and glvalues without finding anything undefined. Can you find
it?
C++ may well make it explicitly UB, but then C++ has a stronger reason
not to, since taking a reference is (surely?) not intended to outlawed:
int a[10];
int (&r)[10] = (&a)[1];
C++ has special wording to permit taking a reference to the lvalue that
results from applying * to some pointers with incomplete types -- an
operation this unequivocally undefined by the C standard -- so applying
* in some otherwise undefined situations is an area with some recognised
special cases.
> though I guess to most of us that formal UB doesn't matter except as
> an example of the improvement potential of the standardization
> process.
There was talk about filing a DR (for WG14) over in comp.lang.c but I
don't think anyone did.
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-06 11:20 -0800 |
| Message-ID | <86fsovkdjw.fsf@linuxsc.com> |
| In reply to | #82940 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
>
>> The code using `(&result)[1]` involves a couple of extra conversions
>
> Eh? char *ep = r + sizeof(r); involves one array-to-pointer conversion
> just as char *ep = (&r)[1]; does.
>
>> and indirections and is thus needlessly complex, plus it's formally
>> UB,
>
> Can you point to where this is made UB in the C++ standard?
>
> There is an ambiguous phrase in the C standard ("undefined behaviour if
> the * is evaluated") but that phrase is entirely missing in the C++
> text. And I got lost in the forest of rvalues, lvalues, prvalues,
> xvalues and glvalues without finding anything undefined. Can you find
> it? [...]
Let me take a stab at explaining. There is undefined behavior in
the code that Alf originally responded to, but where the UB is
and why it is UB has not been explained very well. Two points to
start: there is nothing wrong with the original initializing
declaration; and the aspect of dereferencing is a red herring.
To begin suppose we have this code:
char foo[2][20];
char *p = foo[1];
There is nothing wrong with this code. After these declarations
the variable p points at the first element of foo[1];
But a problem happens if we want to use p to create pointer
values that point to elements of foo[0], as for example
char *p_prime = p-1;
The rules for pointer arithmetic don't allow this. The reason
is p points an element of foo[1], but p-1 doesn't. In both C
and C++ the semantic description for adding an integer to a
pointer defines the result only if P and P+N are both elements
of the same array (or one past the last element). But that
isn't true here. The variable p unambiguously points at an
element of foo[1], but p-1 does not. If we have two pointers
p and q
char foo[2][20];
char *p = foo[1];
char *q = foo[0] + 20;
the pointers p and q will (and must) compare equal, but they are
not interchangeable, because they point into different subarrays
of foo.
Going back to the earlier code, we have
char *ep = (&result)[1];
There is nothing wrong with this initialization. But by virtue
of taking the address of result, which is an array, we have in
effect created a two-dimensional array, and have initialized ep
with a pointer that points into the "second subarray" of that
two-dimensional array. The created pointer value is not allowed
to participate in arithmetic that would take it into the first
subarray. So later in the earlier code when we use this
expression
ep - 4
the code is asking for a pointer to an object in (&result)[1].
But there is no such object. The expressions
ep - 0
and
ep + 0
would both be okay, but only because what is being added or
subtracted is zero; if any non-zero value were used there
is undefined behavior, because the resulting pointer value
does not point to an object in the same subarray. (In effect
there is one and only one object in the second subarray, but
trying to access that mythical object also falls into the
realm of undefined behavior.)
References: section 6.5.6 paragraph 8 for C (n1570);
section 7.6.6 paragraph 4.2 for C++ (N4860).
For the record, IMO neither the C standard nor the C++ standard
describes plainly enough what the rules are for multi-dimensional
arrays. I understand why the rules are the way they are, but I
would like to see one or both standards state how things work in
a more lucid fashion.
Does this help clear things up?
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-06 21:53 +0000 |
| Message-ID | <877da7r7ap.fsf@bsb.me.uk> |
| In reply to | #82942 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
>>
>>> The code using `(&result)[1]` involves a couple of extra conversions
>>
>> Eh? char *ep = r + sizeof(r); involves one array-to-pointer conversion
>> just as char *ep = (&r)[1]; does.
>>
>>> and indirections and is thus needlessly complex, plus it's formally
>>> UB,
>>
>> Can you point to where this is made UB in the C++ standard?
>>
>> There is an ambiguous phrase in the C standard ("undefined behaviour if
>> the * is evaluated") but that phrase is entirely missing in the C++
>> text. And I got lost in the forest of rvalues, lvalues, prvalues,
>> xvalues and glvalues without finding anything undefined. Can you find
>> it? [...]
>
> Let me take a stab at explaining. There is undefined behavior in
> the code that Alf originally responded to, but where the UB is
> and why it is UB has not been explained very well. Two points to
> start: there is nothing wrong with the original initializing
> declaration; and the aspect of dereferencing is a red herring.
>
> To begin suppose we have this code:
>
> char foo[2][20];
> char *p = foo[1];
>
> There is nothing wrong with this code. After these declarations
> the variable p points at the first element of foo[1];
>
> But a problem happens if we want to use p to create pointer
> values that point to elements of foo[0], as for example
>
> char *p_prime = p-1;
>
> The rules for pointer arithmetic don't allow this. The reason
> is p points an element of foo[1], but p-1 doesn't. In both C
> and C++ the semantic description for adding an integer to a
> pointer defines the result only if P and P+N are both elements
> of the same array (or one past the last element). But that
> isn't true here. The variable p unambiguously points at an
> element of foo[1], but p-1 does not. If we have two pointers
> p and q
>
> char foo[2][20];
> char *p = foo[1];
> char *q = foo[0] + 20;
>
> the pointers p and q will (and must) compare equal, but they are
> not interchangeable, because they point into different subarrays
> of foo.
>
> Going back to the earlier code, we have
>
> char *ep = (&result)[1];
>
> There is nothing wrong with this initialization. But by virtue
> of taking the address of result, which is an array, we have in
> effect created a two-dimensional array, and have initialized ep
> with a pointer that points into the "second subarray" of that
> two-dimensional array. The created pointer value is not allowed
> to participate in arithmetic that would take it into the first
> subarray. So later in the earlier code when we use this
> expression
>
> ep - 4
>
> the code is asking for a pointer to an object in (&result)[1].
> But there is no such object. The expressions
>
> ep - 0
>
> and
> ep + 0
>
> would both be okay, but only because what is being added or
> subtracted is zero; if any non-zero value were used there
> is undefined behavior, because the resulting pointer value
> does not point to an object in the same subarray. (In effect
> there is one and only one object in the second subarray, but
> trying to access that mythical object also falls into the
> realm of undefined behavior.)
>
> References: section 6.5.6 paragraph 8 for C (n1570);
> section 7.6.6 paragraph 4.2 for C++ (N4860).
So given
int i;
char *cp = (void *)(&i + 1);
accessing the bytes of i from cp is also undefined. Whilst I don't
think I've ever used the (&array)[1] pointer before now as anything
other than an "end marker" for comparison (which is defined), I'm pretty
sure I've written something like the above more than once. It just
seems simpler than adding sizeof (i) to (char *)&i.
> For the record, IMO neither the C standard nor the C++ standard
> describes plainly enough what the rules are for multi-dimensional
> arrays. I understand why the rules are the way they are, but I
> would like to see one or both standards state how things work in
> a more lucid fashion.
>
> Does this help clear things up?
It explains something else, but not what I was asking about. I did not
consider the problem of "going backwards", so thank you for that. (I
originally asked for citations to support what Alf thought was UB about
the (&array)[1] construct because C++ and C are worded differently
there.)
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-06 19:03 -0800 |
| Message-ID | <86bkzjjs3c.fsf@linuxsc.com> |
| In reply to | #82944 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>
>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>
>>> "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
>>>
>>>> The code using `(&result)[1]` involves a couple of extra conversions
>>>
>>> Eh? char *ep = r + sizeof(r); involves one array-to-pointer conversion
>>> just as char *ep = (&r)[1]; does.
>>>
>>>> and indirections and is thus needlessly complex, plus it's formally
>>>> UB,
>>>
>>> Can you point to where this is made UB in the C++ standard?
>>>
>>> There is an ambiguous phrase in the C standard ("undefined behaviour if
>>> the * is evaluated") but that phrase is entirely missing in the C++
>>> text. And I got lost in the forest of rvalues, lvalues, prvalues,
>>> xvalues and glvalues without finding anything undefined. Can you find
>>> it? [...]
>>
>> Let me take a stab at explaining. There is undefined behavior in
>> the code that Alf originally responded to, but where the UB is
>> and why it is UB has not been explained very well. Two points to
>> start: there is nothing wrong with the original initializing
>> declaration; and the aspect of dereferencing is a red herring.
>>
>> To begin suppose we have this code:
>>
>> char foo[2][20];
>> char *p = foo[1];
>>
>> There is nothing wrong with this code. After these declarations
>> the variable p points at the first element of foo[1];
>>
>> But a problem happens if we want to use p to create pointer
>> values that point to elements of foo[0], as for example
>>
>> char *p_prime = p-1;
>>
>> The rules for pointer arithmetic don't allow this. The reason
>> is p points an element of foo[1], but p-1 doesn't. In both C
>> and C++ the semantic description for adding an integer to a
>> pointer defines the result only if P and P+N are both elements
>> of the same array (or one past the last element). But that
>> isn't true here. The variable p unambiguously points at an
>> element of foo[1], but p-1 does not. If we have two pointers
>> p and q
>>
>> char foo[2][20];
>> char *p = foo[1];
>> char *q = foo[0] + 20;
>>
>> the pointers p and q will (and must) compare equal, but they are
>> not interchangeable, because they point into different subarrays
>> of foo.
>>
>> Going back to the earlier code, we have
>>
>> char *ep = (&result)[1];
>>
>> There is nothing wrong with this initialization. But by virtue
>> of taking the address of result, which is an array, we have in
>> effect created a two-dimensional array, and have initialized ep
>> with a pointer that points into the "second subarray" of that
>> two-dimensional array. The created pointer value is not allowed
>> to participate in arithmetic that would take it into the first
>> subarray. So later in the earlier code when we use this
>> expression
>>
>> ep - 4
>>
>> the code is asking for a pointer to an object in (&result)[1].
>> But there is no such object. The expressions
>>
>> ep - 0
>>
>> and
>> ep + 0
>>
>> would both be okay, but only because what is being added or
>> subtracted is zero; if any non-zero value were used there
>> is undefined behavior, because the resulting pointer value
>> does not point to an object in the same subarray. (In effect
>> there is one and only one object in the second subarray, but
>> trying to access that mythical object also falls into the
>> realm of undefined behavior.)
>>
>> References: section 6.5.6 paragraph 8 for C (n1570);
>> section 7.6.6 paragraph 4.2 for C++ (N4860).
>
> So given
>
> int i;
> char *cp = (void *)(&i + 1);
>
> accessing the bytes of i from cp is also undefined.
No, accessing cp[-1], etc, is defined behavior. The two situations
are not analogous. The reason is that in this case there is only
one array, without any subarrays. The expression (&i+1) has type
pointer to int; it points to one past the last element of the
implicit array that surrounds a non-array object. Because that
expression points to (or just past) an element of the one array, C
allows access to other elements of that array, namely the bytes of
(&i)[0].
By analogy, if we were to do this
char result[20];
char *ep = (void*)(&result + 1);
then ep can be used to access the bytes in result, because the
pointer being converted points to an element of &result, rather
than to an element of (&result)[1]: in the first case the value
points to an element of the first level array, whereas in the
second case the value points to an element of the second level
array (which is a subarray, so the pointer value points into
the subarray rather than to the subarray as a whole). Similarly,
char result[20];
char *ep = (void*) &(&result)[1];
allows ep to be used to access the bytes in result, because
the pointer being converted points to an element of &result.
In the original case
char result[20];
char *ep = (&result)[1];
what kills us is the implicit conversion from (char[20]) to
(char *). That conversion gives a pointer that points /into/
(&result)[1] rather than a pointer /to/ an element of &result.
Taking the address of (&result)[1] -- ie, &(&result)[1] --
solves the undefined behavior aspect, but of course then the
expression does not have type (char *), so some kind of casting
would be needed.
> [...] While I don't think I've ever used the (&array)[1] pointer
> before now as anything other than an "end marker" for comparison
> (which is defined), I'm pretty sure I've written something like the
> above more than once. It just seems simpler than adding sizeof (i)
> to (char *)&i.
If you don't mind using a cast or compound literal, you can use
the &x+1 form regardless of whether x is a scalar or an array:
char result[20];
char *ep = (void*){ &result + 1 };
and sidestep the problem with undefined behavior. (Unfortunately
the compound literal form works only with function locals, and not
at the top level.)
For the record, I think your idiom is very nice, and I am quite
disappointed that it inadvertently wanders into the land of
undefined behavior.
>> For the record, IMO neither the C standard nor the C++ standard
>> describes plainly enough what the rules are for multi-dimensional
>> arrays. I understand why the rules are the way they are, but I
>> would like to see one or both standards state how things work in
>> a more lucid fashion.
>>
>> Does this help clear things up?
>
> It explains something else, but not what I was asking about. I did not
> consider the problem of "going backwards", so thank you for that. (I
> originally asked for citations to support what Alf thought was UB about
> the (&array)[1] construct because C++ and C are worded differently
> there.)
I must confess I find some of Alf's statements hard to understand,
so unfortunately I can't help you there. AFAICT the construct
(&array)[1] does not by itself have undefined behavior in either
C or C++.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-07 11:50 +0000 |
| Message-ID | <87k0e6q4j8.fsf@bsb.me.uk> |
| In reply to | #82946 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>
>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>>
>>>> "Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes:
>>>>
>>>>> The code using `(&result)[1]` involves a couple of extra conversions
>>>>
>>>> Eh? char *ep = r + sizeof(r); involves one array-to-pointer conversion
>>>> just as char *ep = (&r)[1]; does.
>>>>
>>>>> and indirections and is thus needlessly complex, plus it's formally
>>>>> UB,
>>>>
>>>> Can you point to where this is made UB in the C++ standard?
>>>>
>>>> There is an ambiguous phrase in the C standard ("undefined behaviour if
>>>> the * is evaluated") but that phrase is entirely missing in the C++
>>>> text. And I got lost in the forest of rvalues, lvalues, prvalues,
>>>> xvalues and glvalues without finding anything undefined. Can you find
>>>> it? [...]
>>>
>>> Let me take a stab at explaining. There is undefined behavior in
>>> the code that Alf originally responded to, but where the UB is
>>> and why it is UB has not been explained very well. Two points to
>>> start: there is nothing wrong with the original initializing
>>> declaration; and the aspect of dereferencing is a red herring.
>>>
>>> To begin suppose we have this code:
>>>
>>> char foo[2][20];
>>> char *p = foo[1];
>>>
>>> There is nothing wrong with this code. After these declarations
>>> the variable p points at the first element of foo[1];
>>>
>>> But a problem happens if we want to use p to create pointer
>>> values that point to elements of foo[0], as for example
>>>
>>> char *p_prime = p-1;
>>>
>>> The rules for pointer arithmetic don't allow this. The reason
>>> is p points an element of foo[1], but p-1 doesn't. In both C
>>> and C++ the semantic description for adding an integer to a
>>> pointer defines the result only if P and P+N are both elements
>>> of the same array (or one past the last element). But that
>>> isn't true here. The variable p unambiguously points at an
>>> element of foo[1], but p-1 does not. If we have two pointers
>>> p and q
>>>
>>> char foo[2][20];
>>> char *p = foo[1];
>>> char *q = foo[0] + 20;
>>>
>>> the pointers p and q will (and must) compare equal, but they are
>>> not interchangeable, because they point into different subarrays
>>> of foo.
>>>
>>> Going back to the earlier code, we have
>>>
>>> char *ep = (&result)[1];
>>>
>>> There is nothing wrong with this initialization. But by virtue
>>> of taking the address of result, which is an array, we have in
>>> effect created a two-dimensional array, and have initialized ep
>>> with a pointer that points into the "second subarray" of that
>>> two-dimensional array. The created pointer value is not allowed
>>> to participate in arithmetic that would take it into the first
>>> subarray. So later in the earlier code when we use this
>>> expression
>>>
>>> ep - 4
>>>
>>> the code is asking for a pointer to an object in (&result)[1].
>>> But there is no such object. The expressions
>>>
>>> ep - 0
>>>
>>> and
>>> ep + 0
>>>
>>> would both be okay, but only because what is being added or
>>> subtracted is zero; if any non-zero value were used there
>>> is undefined behavior, because the resulting pointer value
>>> does not point to an object in the same subarray. (In effect
>>> there is one and only one object in the second subarray, but
>>> trying to access that mythical object also falls into the
>>> realm of undefined behavior.)
>>>
>>> References: section 6.5.6 paragraph 8 for C (n1570);
>>> section 7.6.6 paragraph 4.2 for C++ (N4860).
>>
>> So given
>>
>> int i;
>> char *cp = (void *)(&i + 1);
>>
>> accessing the bytes of i from cp is also undefined.
>
> No, accessing cp[-1], etc, is defined behavior. The two situations
> are not analogous. The reason is that in this case there is only
> one array, without any subarrays.
Does that not depend how literally one takes the "an object is an array
of length one" rule?
(I'm not being 100% serious here. It's obviously intended to mean, "an
object is the sole element in an array of length one".)
I've cut the rest of your very helpful explanation except for a detail:
> If you don't mind using a cast or compound literal, you can use
> the &x+1 form regardless of whether x is a scalar or an array:
>
> char result[20];
> char *ep = (void*){ &result + 1 };
Were are talking about C++ here. Maybe there is some default
constructor equivalent of this.
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-07 06:26 -0800 |
| Message-ID | <8635kukb17.fsf@linuxsc.com> |
| In reply to | #82951 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>
>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
[.. considering the idiom (&x)[1], where x is an array ..]
>>> So given
>>>
>>> int i;
>>> char *cp = (void *)(&i + 1);
>>>
>>> accessing the bytes of i from cp is also undefined.
>>
>> No, accessing cp[-1], etc, is defined behavior. The two
>> situations are not analogous. The reason is that in this case
>> there is only one array, without any subarrays.
>
> Does that not depend how literally one takes the "an object is an
> array of length one" rule?
>
> (I'm not being 100% serious here. It's obviously intended to
> mean, "an object is the sole element in an array of length one".)
I believe the rule about treating objects as an array of length
one does not enter into the question; all that matters is the
types involved. The type of &i is pointer to int. However, if
we did this
int i;
char *cp = ( (char (*)[sizeof i]) &i )[1];
then there would again be undefined behavior, because now the
type of the expression before the [1] is a pointer-to-array type,
and so that expression ultimately gets converted to a pointer to
an element of the second subarray.
> I've cut the rest of your very helpful explanation except for a
> detail:
>
>> If you don't mind using a cast or compound literal, you can use
>> the &x+1 form regardless of whether x is a scalar or an array:
>>
>> char result[20];
>> char *ep = (void*){ &result + 1 };
>
> We're are talking about C++ here. Maybe there is some default
> constructor equivalent of this.
Sorry, yes. In some cases in C, and in C++, using a cast may be
the only option:
char result[20];
char *ep = (char*)( &result + 1 );
Probably it's true that in C++ something could be done to avoid
having to use a cast, but unfortunately I find the rules for
automatic conversions ("coercions") in C++ too complicated to
offer a reliable answer to the question.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-07 15:48 +0000 |
| Message-ID | <878rumptie.fsf@bsb.me.uk> |
| In reply to | #82955 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>
>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
> [.. considering the idiom (&x)[1], where x is an array ..]
>
>>>> So given
>>>>
>>>> int i;
>>>> char *cp = (void *)(&i + 1);
>>>>
>>>> accessing the bytes of i from cp is also undefined.
>>>
>>> No, accessing cp[-1], etc, is defined behavior. The two
>>> situations are not analogous. The reason is that in this case
>>> there is only one array, without any subarrays.
>>
>> Does that not depend how literally one takes the "an object is an
>> array of length one" rule?
>>
>> (I'm not being 100% serious here. It's obviously intended to
>> mean, "an object is the sole element in an array of length one".)
>
> I believe the rule about treating objects as an array of length
> one does not enter into the question;
I thought it comes into play every time arithmetic is done on a pointer
to an object that is not an element of an array. There appears to be no
other explicit justification for the "one past the end" pointer &i + 1.
Mind you, footnote 76 makes it clear: "an object that is not an array
element is considered to belong to a single-element array for this
purpose" and C has similar wording. I misremembered the rule: it's not
that the object is considered to /be/ an array but to be /in/ an array
or length one. I think the rule of often misquoted.
> all that matters is the
> types involved. The type of &i is pointer to int. However, if
> we did this
>
> int i;
> char *cp = ( (char (*)[sizeof i]) &i )[1];
>
> then there would again be undefined behavior, because now the
> type of the expression before the [1] is a pointer-to-array type,
> and so that expression ultimately gets converted to a pointer to
> an element of the second subarray.
>
>> I've cut the rest of your very helpful explanation except for a
>> detail:
>>
>>> If you don't mind using a cast or compound literal, you can use
>>> the &x+1 form regardless of whether x is a scalar or an array:
>>>
>>> char result[20];
>>> char *ep = (void*){ &result + 1 };
>>
>> We're are talking about C++ here. Maybe there is some default
>> constructor equivalent of this.
>
> Sorry, yes. In some cases in C, and in C++, using a cast may be
> the only option:
>
> char result[20];
> char *ep = (char*)( &result + 1 );
Thinking about this a bit more I not 100% convinced; or perhaps I am
just a little uneasy about how easy it would be to get this wrong.
Switching to int (so that any special rules about accessing an object's
byte representation using a pointer to char don't come into play) this
int result[20];
int *ip = (int *)(&result + 1);
allows accesses like ip[-1], so (unless I'm missing some other rules)
this
int result2[2][20];
int *ip = (int *)(&result2[0] + 1);
allows access to result2[0] via negative indexes, but not to result2[1]
using positive ones, despite the fact that the pointer being converted
is a pointer to the second subarray of 'result2'. Is that how you see
things?
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-07 12:16 -0800 |
| Message-ID | <86leymiga7.fsf@linuxsc.com> |
| In reply to | #82957 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>
>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>
>>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>>
>>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>
>> [.. considering the idiom (&x)[1], where x is an array ..]
>>
>>>>> So given
>>>>>
>>>>> int i;
>>>>> char *cp = (void *)(&i + 1);
>>>>>
>>>>> accessing the bytes of i from cp is also undefined.
>>>>
>>>> No, accessing cp[-1], etc, is defined behavior. The two
>>>> situations are not analogous. The reason is that in this case
>>>> there is only one array, without any subarrays.
>>>
>>> Does that not depend how literally one takes the "an object is an
>>> array of length one" rule?
>>>
>>> (I'm not being 100% serious here. It's obviously intended to
>>> mean, "an object is the sole element in an array of length one".)
>>
>> I believe the rule about treating objects as an array of length
>> one does not enter into the question;
>
> I thought it comes into play every time arithmetic is done on a
> pointer to an object that is not an element of an array. There
> appears to be no other explicit justification for the "one past the
> end" pointer &i + 1.
Yes, it is important that the expression &i+1 obeys the rule
about treating pointers to objects not in an array the same as if
the object were an element in an array of length 1. What I meant
was that this rule doesn't matter for whether there is undefined
behavior -- the expression &x + 1 always works, no matter what x
is (assuming the expression x doesn't refer to a bitfield or
something like that).
> Mind you, footnote 76 makes it clear: "an object that is not an
> array element is considered to belong to a single-element array for
> this purpose" and C has similar wording. I misremembered the rule:
> it's not that the object is considered to /be/ an array but to be
> /in/ an array or length one. I think the rule of often misquoted.
Looking up the relevant passages in the two standards, I find
they are slightly different. In C++ the rule is primarily about
objects, whereas in C the rule centers around pointer values.
Probably the two rules have the same consequences, but I haven't
tried to verify that.
>> all that matters is the
>> types involved. The type of &i is pointer to int. However, if
>> we did this
>>
>> int i;
>> char *cp = ( (char (*)[sizeof i]) &i )[1];
>>
>> then there would again be undefined behavior, because now the
>> type of the expression before the [1] is a pointer-to-array type,
>> and so that expression ultimately gets converted to a pointer to
>> an element of the second subarray.
>>
>>> I've cut the rest of your very helpful explanation except for a
>>> detail:
>>>
>>>> If you don't mind using a cast or compound literal, you can use
>>>> the &x+1 form regardless of whether x is a scalar or an array:
>>>>
>>>> char result[20];
>>>> char *ep = (void*){ &result + 1 };
>>>
>>> We're are talking about C++ here. Maybe there is some default
>>> constructor equivalent of this.
>>
>> Sorry, yes. In some cases in C, and in C++, using a cast may be
>> the only option:
>>
>> char result[20];
>> char *ep = (char*)( &result + 1 );
>
> Thinking about this a bit more I not 100% convinced; or perhaps I am
> just a little uneasy about how easy it would be to get this wrong.
>
> Switching to int (so that any special rules about accessing an object's
> byte representation using a pointer to char don't come into play) this
>
> int result[20];
> int *ip = (int *)(&result + 1);
>
> allows accesses like ip[-1], so (unless I'm missing some other rules)
> this
>
> int result2[2][20];
> int *ip = (int *)(&result2[0] + 1);
>
> allows access to result2[0] via negative indexes, but not to result2[1]
> using positive ones, despite the fact that the pointer being converted
> is a pointer to the second subarray of 'result2'. Is that how you see
> things?
If you think you're going to confuse with these questions, well
it is quite possible that you will. :)
Having said that, I press on.
Certainly the expression 'result2' is able to access any object
within the confines of the two-dimensional array result2.
Similarly the expression 'result2 + 0' is able to access any
object with in the confines of the two-dimensional array result2,
because 'result2 + 0' is exactly the same pointer as 'result2'.
The expression '&result2[0]' is, by definition, the same as the
expression '& * (result2 + 0)'. (Sidebar: in C that is always
true, whereas in C++ things may be different because of operator
overloading or something like that. I am assuming that these
possibilities don't come into play here, and '&result2[0]' acts
the same way in C++ as it does in C. (End sidebar.))
In C there is a rule that &*(E) is the same as (E) as long as the
constraints for the two operators are met. So in C '&result2[0]'
gives the same results as just 'result2', which is to say it is able
to access any object within the confines of the two-dimensional
array result2.
Incidentally this observation agrees with my usual practice when
calling functions taking a pointer parameter. If at a call site
I want to indicate that the argument is a pointer to a single
element I use an array indexing form, eg, &p[k]. Conversely if
I want to indicate that the argument is a pointer to an element
within a larger array I use an addition form, eg, p+k. I know
that both of these argument expressions allow access to the
entire array, and so are fully interchangeable as far as the
semantics goes; however I think it is useful to distinguish the
two intended meanings, so I normally use different argument forms
in the two cases.
Returning to your question, the expression '&result2[0]+1' is, in
C, the same as 'result2+1', which allows access to all of result2.
In C++, there is AFAICT no rule paralleling the '&*' rule in C.
However, looking at paragraph 3.2 in section 7.6.2.1 (in the C++
document N4860), I believe the same result obtains. The question
is more complicated in C++ because there might be overloading of
operator*(), and perhaps also of operator&(), but presumably those
possibilities aren't in play here. So, I believe that C++ gives
the same result that C does, allowing access to all of result2,
with the disclaimer that my knowledge and understanding of C++
is not what I would call authoritative or necessarily reliable.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-07 22:27 -0800 |
| Message-ID | <86h799j2jv.fsf@linuxsc.com> |
| In reply to | #82962 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes: [...] > If you think you're going to confuse with these questions, well > it is quite possible that you will. :) That should have been, If you think you're going to confuse /me/ with these questions, well it is quite possible that you will. A rather ironic omission.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-09 21:43 +0000 |
| Message-ID | <877da3n2au.fsf@bsb.me.uk> |
| In reply to | #82962 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>
>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>>
>>>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>>>
>>>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>>
>>> [.. considering the idiom (&x)[1], where x is an array ..]
>>>
>>>>>> So given
>>>>>>
>>>>>> int i;
>>>>>> char *cp = (void *)(&i + 1);
>>>>>>
>>>>>> accessing the bytes of i from cp is also undefined.
>>>>>
>>>>> No, accessing cp[-1], etc, is defined behavior. The two
>>>>> situations are not analogous. The reason is that in this case
>>>>> there is only one array, without any subarrays.
>>>>
>>>> Does that not depend how literally one takes the "an object is an
>>>> array of length one" rule?
>>>>
>>>> (I'm not being 100% serious here. It's obviously intended to
>>>> mean, "an object is the sole element in an array of length one".)
>>>
>>> I believe the rule about treating objects as an array of length
>>> one does not enter into the question;
>>
>> I thought it comes into play every time arithmetic is done on a
>> pointer to an object that is not an element of an array. There
>> appears to be no other explicit justification for the "one past the
>> end" pointer &i + 1.
>
> Yes, it is important that the expression &i+1 obeys the rule
> about treating pointers to objects not in an array the same as if
> the object were an element in an array of length 1. What I meant
> was that this rule doesn't matter for whether there is undefined
> behavior -- the expression &x + 1 always works, no matter what x
> is (assuming the expression x doesn't refer to a bitfield or
> something like that).
>
>> Mind you, footnote 76 makes it clear: "an object that is not an
>> array element is considered to belong to a single-element array for
>> this purpose" and C has similar wording. I misremembered the rule:
>> it's not that the object is considered to /be/ an array but to be
>> /in/ an array or length one. I think the rule of often misquoted.
>
> Looking up the relevant passages in the two standards, I find
> they are slightly different. In C++ the rule is primarily about
> objects, whereas in C the rule centers around pointer values.
> Probably the two rules have the same consequences, but I haven't
> tried to verify that.
>
>>> all that matters is the
>>> types involved. The type of &i is pointer to int. However, if
>>> we did this
>>>
>>> int i;
>>> char *cp = ( (char (*)[sizeof i]) &i )[1];
>>>
>>> then there would again be undefined behavior, because now the
>>> type of the expression before the [1] is a pointer-to-array type,
>>> and so that expression ultimately gets converted to a pointer to
>>> an element of the second subarray.
>>>
>>>> I've cut the rest of your very helpful explanation except for a
>>>> detail:
>>>>
>>>>> If you don't mind using a cast or compound literal, you can use
>>>>> the &x+1 form regardless of whether x is a scalar or an array:
>>>>>
>>>>> char result[20];
>>>>> char *ep = (void*){ &result + 1 };
>>>>
>>>> We're are talking about C++ here. Maybe there is some default
>>>> constructor equivalent of this.
>>>
>>> Sorry, yes. In some cases in C, and in C++, using a cast may be
>>> the only option:
>>>
>>> char result[20];
>>> char *ep = (char*)( &result + 1 );
>>
>> Thinking about this a bit more I not 100% convinced; or perhaps I am
>> just a little uneasy about how easy it would be to get this wrong.
>>
>> Switching to int (so that any special rules about accessing an object's
>> byte representation using a pointer to char don't come into play) this
>>
>> int result[20];
>> int *ip = (int *)(&result + 1);
>>
>> allows accesses like ip[-1], so (unless I'm missing some other rules)
>> this
>>
>> int result2[2][20];
>> int *ip = (int *)(&result2[0] + 1);
>>
>> allows access to result2[0] via negative indexes, but not to result2[1]
>> using positive ones, despite the fact that the pointer being converted
>> is a pointer to the second subarray of 'result2'. Is that how you see
>> things?
>
> If you think you're going to confuse with these questions, well
> it is quite possible that you will. :)
>
> Having said that, I press on.
>
> Certainly the expression 'result2' is able to access any object
> within the confines of the two-dimensional array result2.
>
> Similarly the expression 'result2 + 0' is able to access any
> object with in the confines of the two-dimensional array result2,
> because 'result2 + 0' is exactly the same pointer as 'result2'.
>
> The expression '&result2[0]' is, by definition, the same as the
> expression '& * (result2 + 0)'. (Sidebar: in C that is always
> true, whereas in C++ things may be different because of operator
> overloading or something like that. I am assuming that these
> possibilities don't come into play here, and '&result2[0]' acts
> the same way in C++ as it does in C. (End sidebar.))
>
> In C there is a rule that &*(E) is the same as (E) as long as the
> constraints for the two operators are met. So in C '&result2[0]'
> gives the same results as just 'result2', which is to say it is able
> to access any object within the confines of the two-dimensional
> array result2.
>
> Incidentally this observation agrees with my usual practice when
> calling functions taking a pointer parameter. If at a call site
> I want to indicate that the argument is a pointer to a single
> element I use an array indexing form, eg, &p[k]. Conversely if
> I want to indicate that the argument is a pointer to an element
> within a larger array I use an addition form, eg, p+k. I know
> that both of these argument expressions allow access to the
> entire array, and so are fully interchangeable as far as the
> semantics goes; however I think it is useful to distinguish the
> two intended meanings, so I normally use different argument forms
> in the two cases.
>
> Returning to your question, the expression '&result2[0]+1' is, in
> C, the same as 'result2+1', which allows access to all of result2.
Sorry for the delay. Life got in the way. Your arguments are
compelling, but I am left unsure if, for any given pointer expression, I
could work out, using the C++ standard text, the n and i that would
allow me to know what values of j are permitted. (I'm referring to the
0 <= i+j <= n and the 0 <= i-j <= n in 7.6.6 p4.2.
Anyway, to paraphrase F E Smith, I find myself none the wiser but much
better informed.
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-02-09 21:07 -0500 |
| Message-ID | <su1s17$mk0$1@dont-email.me> |
| In reply to | #82976 |
On 2/9/22 16:43, Ben Bacarisse wrote: > Tim Rentsch <tr.17687@z991.linuxsc.com> writes: > >> Ben Bacarisse <ben.usenet@bsb.me.uk> writes: ... >>> int result[20]; >>> int *ip = (int *)(&result + 1); >>> >>> allows accesses like ip[-1], so (unless I'm missing some other rules) >>> this >>> >>> int result2[2][20]; >>> int *ip = (int *)(&result2[0] + 1); ... > Sorry for the delay. Life got in the way. Your arguments are > compelling, but I am left unsure if, for any given pointer expression, I > could work out, using the C++ standard text, the n and i that would > allow me to know what values of j are permitted. (I'm referring to the > 0 <= i+j <= n and the 0 <= i-j <= n in 7.6.6 p4.2. &result: i=0, n=1 &result+1: i=1, n=1 result: i=0, n=20 &result2: i=0, n=1 result2: i=0, n=2 result2 + 1 (== &result2[0] + 1): i=1, n=2 result2[0]: i=0, n=20 If any of those cases seem less than perfectly obvious, let me know, and I'll try to explain my reasoning.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-09 21:25 -0800 |
| Message-ID | <86czjvi97y.fsf@linuxsc.com> |
| In reply to | #82976 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>
>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>
>>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>>
>>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>>>
>>>>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>>>>
>>>>>> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>>>>
>>>> [.. considering the idiom (&x)[1], where x is an array ..]
[...]
>>>> Sorry, yes. In some cases in C, and in C++, using a cast may be
>>>> the only option:
>>>>
>>>> char result[20];
>>>> char *ep = (char*)( &result + 1 );
>>>
>>> Thinking about this a bit more I not 100% convinced; or perhaps
>>> I am just a little uneasy about how easy it would be to get this
>>> wrong.
>>>
>>> Switching to int (so that any special rules about accessing an
>>> object's byte representation using a pointer to char don't come
>>> into play) this
>>>
>>> int result[20];
>>> int *ip = (int *)(&result + 1);
>>>
>>> allows accesses like ip[-1], so (unless I'm missing some other
>>> rules) this
>>>
>>> int result2[2][20];
>>> int *ip = (int *)(&result2[0] + 1);
>>>
>>> allows access to result2[0] via negative indexes, but not to
>>> result2[1] using positive ones, despite the fact that the
>>> pointer being converted is a pointer to the second subarray of
>>> 'result2'. Is that how you see things?
>>
>> If you think you're going to confuse [me] with these questions,
>> well it is quite possible that you will. :)
>>
>> Having said that, I press on.
>>
>> Certainly the expression 'result2' is able to access any object
>> within the confines of the two-dimensional array result2.
>>
>> Similarly the expression 'result2 + 0' is able to access any
>> object with in the confines of the two-dimensional array result2,
>> because 'result2 + 0' is exactly the same pointer as 'result2'.
>>
>> The expression '&result2[0]' is, by definition, the same as the
>> expression '& * (result2 + 0)'. (Sidebar: in C that is always
>> true, whereas in C++ things may be different because of operator
>> overloading or something like that. I am assuming that these
>> possibilities don't come into play here, and '&result2[0]' acts
>> the same way in C++ as it does in C. (End sidebar.))
>>
>> In C there is a rule that &*(E) is the same as (E) as long as the
>> constraints for the two operators are met. So in C '&result2[0]'
>> gives the same results as just 'result2', which is to say it is able
>> to access any object within the confines of the two-dimensional
>> array result2.
>>
>> [...]
>>
>> Returning to your question, the expression '&result2[0]+1' is, in
>> C, the same as 'result2+1', which allows access to all of result2.
>
> Sorry for the delay. Life got in the way. Your arguments are
> compelling, but I am left unsure if, for any given pointer
> expression, I could work out, using the C++ standard text, the n
> and i that would allow me to know what values of j are permitted.
> (I'm referring to the 0 <= i+j <= n and the 0 <= i-j <= n in 7.6.6
> p4.2.
Rather than thinking about pointer values, it might help to start
with objects. If we have this declaration
int a[2][3];
what are the objects? They are
a ( == (&a)[0] )
a[0] a[1]
a[0][0] a[0][1] a[0][2]
a[1][0] a[1][1] a[1][2]
along with the "hypothetical" objects
(&a)[1]
(&a)[1][0]
(&a)[1][0][0]
a[2]
a[2][0]
a[0][3]
a[1][3]
Keep in mind that these forms are not meant to be read as
expressions but just as a way of naming objects.
If we put an & in front of any of the above, now considered as an
expression, we get a pointer value that is allowed to be adjusted
to (ie, using pointer arithmetic) any other form that is the same
except for the last dimension. So if for example we have
&a[0][1]
we can use pointer arithmetic to obtain a pointer to any of
a[0][0] a[0][1] a[0][2] a[0][3]
(with the last item naming a "hypothetical" object), but not any
of the other objects.
Of course if we have or can get a pointer to a containing object
we can obtain a pointer to any of its contained objects. To say
that another way, if we have a pointer to an array, we can obtain
a pointer to any of the elements in the array (and that holds
recursively).
Casting a pointer to a different type doesn't change what objects
it can access. At least, this rule holds in C, and I believe it
also holds in C++. Disclaimer: it is possible C++ has some sort
of magic library functions that do allow something along these
lines, but if it does I don't know about it.
Forgive me if any of the above seems obvious. I am unsure about
what kind of situations you might not be sure of.
Does this explanation help with your uncertainties?
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2022-02-10 13:05 +0000 |
| Message-ID | <87czjulvm9.fsf@bsb.me.uk> |
| In reply to | #82979 |
Tim Rentsch <tr.17687@z991.linuxsc.com> writes: > Rather than thinking about pointer values, it might help to start > with objects. If we have this declaration > > int a[2][3]; > > what are the objects? They are > > a ( == (&a)[0] ) > > a[0] a[1] > > a[0][0] a[0][1] a[0][2] > a[1][0] a[1][1] a[1][2] > > along with the "hypothetical" objects > > (&a)[1] > (&a)[1][0] > (&a)[1][0][0] > > a[2] > a[2][0] > > a[0][3] > a[1][3] > > Keep in mind that these forms are not meant to be read as > expressions but just as a way of naming objects. > > If we put an & in front of any of the above, now considered as an > expression, we get a pointer value that is allowed to be adjusted > to (ie, using pointer arithmetic) any other form that is the same > except for the last dimension. So if for example we have > > &a[0][1] > > we can use pointer arithmetic to obtain a pointer to any of > > a[0][0] a[0][1] a[0][2] a[0][3] > > (with the last item naming a "hypothetical" object), but not any > of the other objects. > > Of course if we have or can get a pointer to a containing object > we can obtain a pointer to any of its contained objects. To say > that another way, if we have a pointer to an array, we can obtain > a pointer to any of the elements in the array (and that holds > recursively). > > Casting a pointer to a different type doesn't change what objects > it can access. At least, this rule holds in C, and I believe it > also holds in C++. Disclaimer: it is possible C++ has some sort > of magic library functions that do allow something along these > lines, but if it does I don't know about it. > > Forgive me if any of the above seems obvious. I am unsure about > what kind of situations you might not be sure of. > > Does this explanation help with your uncertainties? Really the only issue I had is resolved by the rule that "the array" is always the smallest enclosing array that contains the thing pointed to (before any conversions of course). In the back of my mind I had thought that (at least for C) the standard had been written to permit huge arrays to have some or all rows in separate segments, whilst only requiring full pointer arithmetic for the case of a char * accessing the object's representation. But that's not the case, and (T *)&A can access all elements of any array A with base type T. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-02-10 11:19 -0500 |
| Message-ID | <su3dud$kt9$1@dont-email.me> |
| In reply to | #82980 |
On 2/10/22 08:05, Ben Bacarisse wrote: ... > Really the only issue I had is resolved by the rule that "the array" is > always the smallest enclosing array that contains the thing pointed to > (before any conversions of course). While that's correct, it's more directly to the point that the array has an element type that is the same as the pointed-at type.
[toc] | [prev] | [next] | [standalone]
| From | Tim Rentsch <tr.17687@z991.linuxsc.com> |
|---|---|
| Date | 2022-02-11 05:41 -0800 |
| Message-ID | <868ruhikqf.fsf@linuxsc.com> |
| In reply to | #82980 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>
>> Rather than thinking about pointer values, it might help to start
>> with objects. [..illustrative examples..]
>>
>> Casting a pointer to a different type doesn't change what objects
>> it can access. At least, this rule holds in C, [...]
> Really the only issue I had is resolved by the rule that "the array" is
> always the smallest enclosing array that contains the thing pointed to
> (before any conversions of course).
Interesting. I hadn't thought of it that way before. Seems right.
I should make a clarifying statement about casting not changing
what objects can be accessed. If we have this code fragment
int foo[10][20];
extern void set_elements( int *, size_t, int )
set_elements( (int*) &foo, 10*20, -1 );
an argument could be made that set_elements() cannot use pointer
arithmetic (including that implied by use of []) on its first
argument other than to access between foo[0][0] and foo[0][19] (or
to construct a pointer to foo[0][20]). The reasoning would be that
there is no array of 200 elements, so the conditions for address
arithmetic would not be met for index values other than between 0 and
20, and so would technically be undefined behavior. It would be
surprising if that putative undefined behavior would result in the
code "doing the wrong thing", but I feel obliged to point out the
possible alternative reading nonetheless.
Even if the reading proposed above holds, set_elements() can be
written in a way that avoids the putative undefined behavior:
void
set_elements( int *v, size_t n, int value ){
while( n-- > 0 ){
*(int*)( (char*)v + n * sizeof *v ) = value;
}
}
This code has to work for the call that passes '(int*)&foo',
because all objects have an implied character array that overlays
the entire object, which is all of foo in this case.
> In the back of my mind I had thought that (at least for C) the standard
> had been written to permit huge arrays to have some or all rows in
> separate segments, while only requiring full pointer arithmetic for the
> case of a char * accessing the object's representation. But that's not
> the case, and (T *)&A can access all elements of any array A with base
> type T.
My understanding is that the underlying reasons have to do with
possible program optimization. For example, if the array foo and
the function set_elements() have been declared as shown above, and
there is a call
set_elements( foo[0], 20, 10 );
we would like a compiler to be able to assume that foo[1], foo[2],
... , foo[9], are neither changed nor referenced by this call to
set_elements().
AFAIAA all of the above statements apply to C++ as well as C
(assuming of course there is no operator overloading or anything
else along those lines). If anyone knows of any indication to the
contrary in the C++ standard I would be interested to hear about
that.
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2022-02-11 17:06 +0100 |
| Message-ID | <su61id$an$1@dont-email.me> |
| In reply to | #82982 |
On 11 Feb 2022 14:41, Tim Rentsch wrote:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> Tim Rentsch <tr.17687@z991.linuxsc.com> writes:
>>
>>> Rather than thinking about pointer values, it might help to start
>>> with objects. [..illustrative examples..]
>>>
>>> Casting a pointer to a different type doesn't change what objects
>>> it can access. At least, this rule holds in C, [...]
>
>> Really the only issue I had is resolved by the rule that "the array" is
>> always the smallest enclosing array that contains the thing pointed to
>> (before any conversions of course).
>
> Interesting. I hadn't thought of it that way before. Seems right.
>
> I should make a clarifying statement about casting not changing
> what objects can be accessed. If we have this code fragment
>
> int foo[10][20];
> extern void set_elements( int *, size_t, int )
>
> set_elements( (int*) &foo, 10*20, -1 );
>
> an argument could be made that set_elements() cannot use pointer
> arithmetic (including that implied by use of []) on its first
> argument other than to access between foo[0][0] and foo[0][19] (or
> to construct a pointer to foo[0][20]). The reasoning would be that
> there is no array of 200 elements, so the conditions for address
> arithmetic would not be met for index values other than between 0 and
> 20, and so would technically be undefined behavior.
First, for C++20 the above `reinterpret_cast`-expressed-as-C-cast
suffers from not involving "interconvertible" pointers, as noted in
§6.8.2/4.4:
"An array object and its first element are not pointer-interconvertible,
even though they have the same address."
Formally notes are not part of the formal language specification,
they're not "normative". But it's my impression that the C++ committee
has made their own more relaxed rules for ISO standards since (and maybe
including) C++11, and anyway the note surely reflects the committee's
intent.
Secondly, the `foo` object spans 200*sizeof(int) bytes, and an
implementation with range-checked pointers will necessarily allow that
byte range to be addressed. Even after a cast.
And as I see it that in-hypothetical-practice fat pointer view is a good
way to understand the otherwise baffling restrictions for this in the
standard. Fat pointer (slur: "phat pointer"): a pointer outfitted with
begin and end addresses for range checking. It's a good interpretation
because it's backed by a simple hypothetical implementation where one
can more easily reason with concrete examples about what goes on.
> It would be
> surprising if that putative undefined behavior would result in the
> code "doing the wrong thing", but I feel obliged to point out the
> possible alternative reading nonetheless.
>
> Even if the reading proposed above holds, set_elements() can be
> written in a way that avoids the putative undefined behavior:
>
> void
> set_elements( int *v, size_t n, int value ){
> while( n-- > 0 ){
> *(int*)( (char*)v + n * sizeof *v ) = value;
> }
> }
>
> This code has to work for the call that passes '(int*)&foo',
> because all objects have an implied character array that overlays
> the entire object, which is all of foo in this case.
Again, for C++ (though I understand you're discussing the C case) you're
colliding with a formal brick wall.
The general consensus is that only `memcpy` plus one more mechanism I
don't recall now (citing lack of coffee plus excessive blood sugar etc.)
is sufficiently formally supported to avoid formal UB for copying bytes
of objects.
In particular a `reinterpret_cast`, in the above code expressed as a C
style cast, is a good way to let loose the strict aliasing demons of the
g++ compiler. "Oh, formal UB? Well I'll optimize away all code paths
that lead to that, because surely they will never actually be executed,
and if they would it would be a bug. Don't thank me. I'm just doing this
in your best interest, giving you a really fastish program."
>> In the back of my mind I had thought that (at least for C) the standard
>> had been written to permit huge arrays to have some or all rows in
>> separate segments, while only requiring full pointer arithmetic for the
>> case of a char * accessing the object's representation. But that's not
>> the case, and (T *)&A can access all elements of any array A with base
>> type T.
>
> My understanding is that the underlying reasons have to do with
> possible program optimization. For example, if the array foo and
> the function set_elements() have been declared as shown above, and
> there is a call
>
> set_elements( foo[0], 20, 10 );
>
> we would like a compiler to be able to assume that foo[1], foo[2],
> ... , foo[9], are neither changed nor referenced by this call to
> set_elements().
Someone else mentioned (I believe in this thread) the possibility of
having a really large array with sub-arrays placed in two or more
distinct segments of memory, and pointer arithmetic on segment
selector+offset can't in general take you from one segment to another.
I hadn't thought of that; I always thought of this in terms fat pointers.
But as a rationale it makes sense. Just that it's now mostly irrelevant.
I can't imagine an embedded system with segmented memory, though I
assume that as everything that /can/ exist, some such gnarled beast must
exist. But surely C++ doesn't need to explicitly support it. Those who
do program it can surely note that any problem in their unit tests.
> AFAIAA all of the above statements apply to C++ as well as C
> (assuming of course there is no operator overloading or anything
> else along those lines). If anyone knows of any indication to the
> contrary in the C++ standard I would be interested to hear about
> that.
See above; I don't know if my C++ comments apply to C, but I suspect
that they do not, that C is now a far more practical language for stuff
like this.
So maybe one needs to combine C and C++ in order to do formally well
defined programaming.
C for union-based type conversions, multi-dimensional array traversing,
etc.; C++ for higher level control and the bulk of the application code;
C# for user interface (wait, wait, I'm joking, I'm just joking).
- Alf
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2022-02-11 11:50 -0500 |
| Message-ID | <su6452$jv9$1@dont-email.me> |
| In reply to | #82983 |
On 2/11/22 11:06, Alf P. Steinbach wrote:
> On 11 Feb 2022 14:41, Tim Rentsch wrote:
...
>> This code has to work for the call that passes '(int*)&foo',
>> because all objects have an implied character array that overlays
>> the entire object, which is all of foo in this case.
>
> Again, for C++ (though I understand you're discussing the C case) you're
> colliding with a formal brick wall.
>
> The general consensus is that only `memcpy` plus one more mechanism I
> don't recall now (citing lack of coffee plus excessive blood sugar etc.)
> is sufficiently formally supported to avoid formal UB for copying bytes
> of objects.
"For any object (other than a potentially-overlapping subobject) of
trivially copyable type T, whether or not the object holds a valid value
of type T, the underlying bytes (6.7.1) making up the object can be
copied into an array of char, unsigned char, or std::byte (17.2.1). 36
If the content of that array is copied back into the object, the object
shall subsequently hold its original value." {C++ 6.8p2).
The corresponding clause from the C standard is:
"Values stored in non-bit-field objects of any other object type consist
of n × CHAR_BIT bits, where n is the size of an object of that type, in
bytes. The value may be copied into an object of type unsigned char [ n
] (e.g., by memcpy ); the resulting set of bytes is called the object
representation of the value." (C 6.2.6.1p4)
In C++, the guarantee applies to any trivially copyable type. For C, it
applies to any non-bit-field object. The C++ standard uses the concept
of underlying bytes, and requires that the bytes be copied back to their
original location in order for the result to be guaranteed to have the
same value as the original. The C standard gives memcpy() as an example
of how the object could be copied into an array of unsigned char. The
C++ standard leaves the usability of memcpy() as something for the
reader to derive from the definition of that function's behavior.
"If a program attempts to access (3.1) the stored value of an object
through a glvalue whose type is not similar (7.3.5) to one of the
following types the behavior is undefined:
...
— a char, unsigned char, or std::byte type." (C++ 7.2.1p11)
...
> Someone else mentioned (I believe in this thread) the possibility of
> having a really large array with sub-arrays placed in two or more
> distinct segments of memory, and pointer arithmetic on segment
> selector+offset can't in general take you from one segment to another.
Are you talking about a pointers to one of the sub-arrays, or a pointer
to one of the elements of a sub-array? If the former, I don't see how
such an implementation could be conforming. The C++ rules for adding
integers to pointers (7.6.6p4) don't have any exception that I'm aware
of that would allow such an addition to fail.
Even if you're only talking about a pointer to an element of one of the
sub-arrays, it still doesn't sound reasonable. Why would crossing a
segment boundary be a problem for such a pointer, without also being a
problem for a pointer to the sub-array itself?
[toc] | [prev] | [next] | [standalone]
| From | "Alf P. Steinbach" <alf.p.steinbach@gmail.com> |
|---|---|
| Date | 2022-02-11 21:13 +0100 |
| Message-ID | <su6g0g$gpv$1@dont-email.me> |
| In reply to | #82984 |
On 11 Feb 2022 17:50, James Kuyper wrote:
> On 2/11/22 11:06, Alf P. Steinbach wrote:
>> On 11 Feb 2022 14:41, Tim Rentsch wrote:
> ...
>>> This code has to work for the call that passes '(int*)&foo',
>>> because all objects have an implied character array that overlays
>>> the entire object, which is all of foo in this case.
>>
>> Again, for C++ (though I understand you're discussing the C case) you're
>> colliding with a formal brick wall.
>>
>> The general consensus is that only `memcpy` plus one more mechanism I
>> don't recall now (citing lack of coffee plus excessive blood sugar etc.)
>> is sufficiently formally supported to avoid formal UB for copying bytes
>> of objects.
>
> "For any object (other than a potentially-overlapping subobject) of
> trivially copyable type T, whether or not the object holds a valid value
> of type T, the underlying bytes (6.7.1) making up the object can be
> copied into an array of char, unsigned char, or std::byte (17.2.1). 36
> If the content of that array is copied back into the object, the object
> shall subsequently hold its original value." {C++ 6.8p2).
Yes. I commented on how to do that copying. `reinterpret_cast` for that
is ungood.
[snip]
> sub-arrays, it still doesn't sound reasonable. Why would crossing a
> segment boundary be a problem for such a pointer, without also being a
> problem for a pointer to the sub-array itself?
Crossing a segment boundary is a problem because it's inefficient to
check for segment crossing (and find next segment's selector) at every
pointer increment.
With such a hypothetical implementation that possibly the C committee
used as rationale for their nonsense decision, a pointer to
array-residing-in-separate-segment would be a different kind of beast
with much less efficient increment.
A `char*` or `void*` pointer would have to be quite large, and the
standards allow these two pointer types to be larger.
- Alf
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2022-02-11 17:06 +0000 |
| Message-ID | <SIwNJ.3717$9%C9.109@fx21.iad> |
| In reply to | #82983 |
"Alf P. Steinbach" <alf.p.steinbach@gmail.com> writes: >On 11 Feb 2022 14:41, Tim Rentsch wrote: > >In particular a `reinterpret_cast`, in the above code expressed as a C >style cast, is a good way to let loose the strict aliasing demons of the >g++ compiler. "Oh, formal UB? Well I'll optimize away all code paths >that lead to that, because surely they will never actually be executed, >and if they would it would be a bug. Don't thank me. I'm just doing this >in your best interest, giving you a really fastish program." Fortunately, g++ supports the -fno-strict-aliasing option.
[toc] | [prev] | [next] | [standalone]
Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6 Next page →
Back to top | Article view | comp.lang.c++
csiph-web