Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c > #163857 > unrolled thread
| Started by | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| First post | 2021-12-16 11:13 -0300 |
| Last post | 2021-12-18 12:11 -0300 |
| Articles | 20 on this page of 44 — 9 participants |
Back to article view | Back to comp.lang.c
on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-16 11:13 -0300
Re: on understanding & and pointer arithmetic Bart <bc@freeuk.com> - 2021-12-16 14:23 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-16 11:57 -0300
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 17:42 +0000
Re: on understanding & and pointer arithmetic Bart <bc@freeuk.com> - 2021-12-16 18:14 +0000
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 21:14 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-16 16:42 -0300
Re: on understanding & and pointer arithmetic Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-16 12:44 -0800
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:06 -0300
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 21:34 +0000
Re: on understanding & and pointer arithmetic Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-16 14:14 -0800
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 23:44 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:09 -0300
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-16 16:23 +0100
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-16 16:25 -0300
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-16 21:08 +0100
Re: on understanding & and pointer arithmetic Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2021-12-16 12:54 -0800
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-17 10:53 +0100
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 11:39 -0300
Re: on understanding & and pointer arithmetic scott@slp53.sl.home (Scott Lurndal) - 2021-12-18 16:31 +0000
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-18 18:56 +0100
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-25 21:29 -0300
Re: on understanding & and pointer arithmetic Bart <bc@freeuk.com> - 2021-12-18 18:29 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 11:32 -0300
Re: on understanding & and pointer arithmetic Manfred <noname@add.invalid> - 2021-12-16 21:34 +0100
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 22:20 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:02 -0300
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:34 -0300
Re: on understanding & and pointer arithmetic Manfred <noname@add.invalid> - 2021-12-18 18:36 +0100
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-18 21:17 +0000
Re: on understanding & and pointer arithmetic James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-16 18:59 -0500
Re: on understanding & and pointer arithmetic James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-16 18:38 -0500
Re: on understanding & and pointer arithmetic Bart <bc@freeuk.com> - 2021-12-16 23:50 +0000
Re: on understanding & and pointer arithmetic "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> - 2021-12-17 07:34 -0800
Re: on understanding & and pointer arithmetic Bart <bc@freeuk.com> - 2021-12-17 17:48 +0000
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-16 23:56 +0000
Re: on understanding & and pointer arithmetic James Kuyper <jameskuyper@alumni.caltech.edu> - 2021-12-16 19:06 -0500
Re: on understanding & and pointer arithmetic Ben Bacarisse <ben.usenet@bsb.me.uk> - 2021-12-17 00:10 +0000
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:32 -0300
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:30 -0300
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-17 11:08 +0100
Re: on understanding & and pointer arithmetic Manfred <noname@add.invalid> - 2021-12-17 16:24 +0100
Re: on understanding & and pointer arithmetic David Brown <david.brown@hesbynett.no> - 2021-12-17 18:17 +0100
Re: on understanding & and pointer arithmetic Meredith Montgomery <mmontgomery@levado.to> - 2021-12-18 12:11 -0300
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2021-12-18 18:56 +0100 |
| Message-ID | <spl7cj$3kl$1@dont-email.me> |
| In reply to | #163947 |
On 18/12/2021 15:39, Meredith Montgomery wrote: > > It's also nice to see that the situation is even deeper than I would > have thought, though. For instance, right now my intuition is that the > size of a pointer on a 64-bit machine is always 8 bytes. Can anyone > show an easy example of when this isn't true? > There is the x32 ABI, which is an alternative ABI for 64-bit x86 systems that uses 32-bit pointers but 64-bit integer registers and the additional features and registers that x86-64 provides beyond x86-32. The result was more efficient for some kinds of code, but it never really took off. Some Linux distributions provide support for it (such as libraries) alongside their main 64-bit stuff. There are also a few systems, mostly historic and/or academic, that have wider pointers containing security or protection information in addition to addresses. It is usually in smaller systems and nice areas (like DSPs) that you have unusual sizes. For example, gcc for the the 8-bit AVR microcontroller family has 16-bit int, 16-bit pointers for data and code that are separate address spaces (i.e., a data pointer and a function pointer could contain the same "value" if you were to cast them to a uintptr_t, yet, point to completely different memory locations). It also has 24-bit code pointers, and 24-bit generic pointers as well. This is the best example I know of for strangely sized and incompatible pointers, in devices that are popular and modern.
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-25 21:29 -0300 |
| Message-ID | <8635mg2opw.fsf@levado.to> |
| In reply to | #163969 |
David Brown <david.brown@hesbynett.no> writes: > On 18/12/2021 15:39, Meredith Montgomery wrote: > >> >> It's also nice to see that the situation is even deeper than I would >> have thought, though. For instance, right now my intuition is that the >> size of a pointer on a 64-bit machine is always 8 bytes. Can anyone >> show an easy example of when this isn't true? >> > > There is the x32 ABI, which is an alternative ABI for 64-bit x86 systems > that uses 32-bit pointers but 64-bit integer registers and the > additional features and registers that x86-64 provides beyond x86-32. > The result was more efficient for some kinds of code, but it never > really took off. Some Linux distributions provide support for it (such > as libraries) alongside their main 64-bit stuff. > > There are also a few systems, mostly historic and/or academic, that have > wider pointers containing security or protection information in addition > to addresses. > > It is usually in smaller systems and nice areas (like DSPs) that you > have unusual sizes. For example, gcc for the the 8-bit AVR > microcontroller family has 16-bit int, 16-bit pointers for data and code > that are separate address spaces (i.e., a data pointer and a function > pointer could contain the same "value" if you were to cast them to a > uintptr_t, yet, point to completely different memory locations). It > also has 24-bit code pointers, and 24-bit generic pointers as well. > This is the best example I know of for strangely sized and incompatible > pointers, in devices that are popular and modern. Nice to know. Thank you and everyone else that also contributed.
[toc] | [prev] | [next] | [standalone]
| From | Bart <bc@freeuk.com> |
|---|---|
| Date | 2021-12-18 18:29 +0000 |
| Message-ID | <spl99d$gne$1@dont-email.me> |
| In reply to | #163947 |
On 18/12/2021 14:39, Meredith Montgomery wrote: > David Brown <david.brown@hesbynett.no> writes: >> I know the real picture can be more complicated - different pointer >> types can have different sizes, pointers can contain more than just an >> address or memory location, two pointers of different types might >> contain the same raw value but refer to different address spaces, they >> might contain different raw values but refer to the same location and >> may or may not compare equal, there can be trap representations, etc. >> (That list is not complete.) So what I am writing is "lies to >> children", rather than trying to be complete and precise (partly because >> others here, such as yourself, are significantly better at that kind of >> answer). If you are not familiar with the specific phrase "lies to >> children", ask Santa for "The Science of the Discworld" :-) > > For the record, I enjoyed the simplicity of this (sub)thread. I was > trying to understand why my expectation was wrong. So, I was already in > some trouble; bringing in more details at that point would not have > helped more than not bringing them in. > > It's also nice to see that the situation is even deeper than I would > have thought, though. For instance, right now my intuition is that the > size of a pointer on a 64-bit machine is always 8 bytes. Can anyone > show an easy example of when this isn't true? C running on a typical 64-bit desktop PC will have sizeof(int) as 4 bytes, and sizeof(void*) as 8 bytes. Unless you tell it to use a 32-bit target (eg. using -m32 option), when pointers will be 4 bytes. There might be a few odd implementations where it's different (for example I've implemented a language with 32-bit pointers even in 64-bit mode), but I wouldn't worry about that.
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-18 11:32 -0300 |
| Message-ID | <86v8zm7zmd.fsf@levado.to> |
| In reply to | #163878 |
David Brown <david.brown@hesbynett.no> writes:
> On 16/12/2021 20:25, Meredith Montgomery wrote:
>> David Brown <david.brown@hesbynett.no> writes:
>>
>>> On 16/12/2021 15:13, Meredith Montgomery wrote:
>>>> I'm investigating the syntax array[index] and I'm getting surprised at
>>>> some places. My intuition says that a[3] is the same as a + 3. I also
>>>> know that /&a/ is the same as /a/. I first expected the following
>>>> program to print ``lo, world\n'' three times, but it does not.
>>>
>>> I can see why you are confused - you are close, but a bit mixed up.
>>>
>>> "a[3]" is the same as "*(a + 3)". The dereferencing is crucial.
>>
>> Thank you!
>>
>>> In the code below, "a" is an array of 13 char - and that is its type.
>>> When you take its address, "&a", you have a pointer to an array of 13
>>> char. This will have the same /value/ as &a[0], which is a pointer to
>>> the first element of the array - a pointer to a char.
>>
>> Yes, so that corrects me at least once. I didn't think &a would be of a
>> different type. So the type of &a is array of 13 char. I get that.
>
> /No/. "a" is of type "array of 13 char". "&a" is of type "pointer to
> array of 13 char". These are different.
Yes --- thanks. That's what I meant, but didn't write.
>> The value of &a is the same /value/ as &a[0]. I also get that. I'm
>> good here.
>
> Correct. As Ben said (and Ben is very good at explaining things
> accurately), a pointer has a type and has an address as it's value. So
> the values of pointers here are the same, but the pointers themselves
> are different because they have different types.
Yes, I think I will never make that mistake again because that corrected
my intuition now. I was thinking of a pointer as just a variable
holding an integer, but there is a type associated to it which totally
changes arithmetic with it. That's my new intuition. Thanks for
helping me get there --- and for checking my work in your previous post.
[...]
> Once you feel you have got the hang of this, repeat the whole thing with
> an array of "int" rather than an array of "char". That will help you
> appreciate where the scalings come in. (Note that while an "int" is 4
> bytes, or 4 chars, on your platform, it is not necessarily the case on
> other C implementations, especially for very small devices.)
I got the hang of it. Let's see. I'll write the program with my
predictions in comments --- and run it.
--8<---------------cut here---------------start------------->8---
#include <stdio.h>
int main(void) {
int a[] = {1, 2, 3};
printf("%lu\n", (unsigned long) a); // some address
printf("%lu\n", (unsigned long) &a); // same address
printf("%lu bytes\n", (unsigned long) sizeof a); // 12 bytes
printf("%lu bytes\n", (unsigned long) sizeof &a); // 8 bytes
// this will jump 12 bytes to the right
printf("%lu\n", (unsigned long) (&a + 1));
printf("Jumped %lu bytes to the right\n",
(unsigned long) (&a + 1) - (unsigned long) &a);
// this will hump 36 bytes to the right
printf("%lu\n", (unsigned long) (&a + 3));
printf("Jumped %lu bytes to the right\n",
(unsigned long) (&a + 3) - (unsigned long) &a);
}
--8<---------------cut here---------------end--------------->8---
%make arrint
cc -Wall -x c -g -std=c99 -pedantic-errors arrint.c -o arrint
%./arrint.exe
4294954036
4294954036
12 bytes
8 bytes
4294954048
Jumped 12 bytes to the right
4294954072
Jumped 36 bytes to the right
%
That looks good! Thanks very much!
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2021-12-16 21:34 +0100 |
| Message-ID | <spg7sd$10ue$1@gioia.aioe.org> |
| In reply to | #163862 |
On 12/16/2021 4:23 PM, David Brown wrote:
> On 16/12/2021 15:13, Meredith Montgomery wrote:
[...]
>
>> printf("a: %s\n", &a + 3);
>
> "a" here is the full array, so "&a" is a pointer to an array of 13 char.
> Adding 3 gives a new pointer to an array of 13 char, at the address 39
> bytes higher than address of "a". Basically, you are pretending that
> there is an array of 4 elements, each of which is itself an array of 13
> chars - with "a" being the first of those 4 elements. Now you are
> asking to print the 4th element here, which is just whatever happens to
> be in the memory at that address.
>
It might be worth mentioning that accessing memory at that address (&a +
3) technically yields undefined behavior, because it is an attempt to
access a location outside the array (a).
This is to say that attempting to access a memory location that has not
been explicitly allocated or mapped by the program is invalid in C.
Common practical results of many implementations might range from
printing garbage (as you say, whatever happens to be there), run into a
memory access violation (aka segmentation fault), or more unexpected
behaviors.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2021-12-16 22:20 +0000 |
| Message-ID | <87fsqsur88.fsf@bsb.me.uk> |
| In reply to | #163879 |
Manfred <noname@add.invalid> writes:
> On 12/16/2021 4:23 PM, David Brown wrote:
>> On 16/12/2021 15:13, Meredith Montgomery wrote:
> [...]
>>
>>> printf("a: %s\n", &a + 3);
>> "a" here is the full array, so "&a" is a pointer to an array of 13 char.
>> Adding 3 gives a new pointer to an array of 13 char, at the address 39
>> bytes higher than address of "a". Basically, you are pretending that
>> there is an array of 4 elements, each of which is itself an array of 13
>> chars - with "a" being the first of those 4 elements. Now you are
>> asking to print the 4th element here, which is just whatever happens to
>> be in the memory at that address.
>>
>
> It might be worth mentioning that accessing memory at that address (&a
> + 3) technically yields undefined behavior, because it is an attempt
> to access a location outside the array (a).
That's a good point to make. It reminds me of a construct that I rather
like:
char buffer[some-complex-size];
char *end_of_buffer = (&buffer)[1];
(&buffer)[1] means *(&buffer + 1). The constructed pointer &buffer + 1
is valid, and the * does not attempt to dereference it as *(&buffer + 1)
is an array-valued expression. It is, instead, converted to a pointer
to the first element of this non-existent array object -- a pointer one
past the end of 'buffer'.
--
Ben.
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-18 12:02 -0300 |
| Message-ID | <86o85e6jo0.fsf@levado.to> |
| In reply to | #163889 |
Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
> Manfred <noname@add.invalid> writes:
>
>> On 12/16/2021 4:23 PM, David Brown wrote:
>>> On 16/12/2021 15:13, Meredith Montgomery wrote:
>> [...]
>>>
>>>> printf("a: %s\n", &a + 3);
>>> "a" here is the full array, so "&a" is a pointer to an array of 13 char.
>>> Adding 3 gives a new pointer to an array of 13 char, at the address 39
>>> bytes higher than address of "a". Basically, you are pretending that
>>> there is an array of 4 elements, each of which is itself an array of 13
>>> chars - with "a" being the first of those 4 elements. Now you are
>>> asking to print the 4th element here, which is just whatever happens to
>>> be in the memory at that address.
>>>
>>
>> It might be worth mentioning that accessing memory at that address (&a
>> + 3) technically yields undefined behavior, because it is an attempt
>> to access a location outside the array (a).
>>
>
> That's a good point to make.
Indeed. Thanks. (I was aware of that. In fact, I thought that my OP
would produce a lot of --- that's undefined behavior and, so, C has
nothing to do with it. Lol. But, thankfully, you guys went straight
into helping me clear up my troubles.)
> It reminds me of a construct that I rather like:
>
> char buffer[some-complex-size];
> char *end_of_buffer = (&buffer)[1];
>
> (&buffer)[1] means *(&buffer + 1). The constructed pointer &buffer + 1
> is valid, and the * does not attempt to dereference it as *(&buffer + 1)
> is an array-valued expression. It is, instead, converted to a pointer
> to the first element of this non-existent array object -- a pointer one
> past the end of 'buffer'.
Wow, that's a cool application. Thanks for sharing.
I needed to redefine my definition of ``end''. I tend to think the end
of an array is its last byte, but agreed --- that's not quite its end
yet. It seems hard to define the exact end of something.
The end of my property should be on my property or should it be outside
of it? If *it* is outside, then because because *it* is relative to
*my* property, this *it* must be mine, so it belongs to my property. So
it's not outside of it. :-)
I would have done
char *end_of_buffer = (&buffer)[1] - 1;
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-18 12:34 -0300 |
| Message-ID | <867dc253mb.fsf@levado.to> |
| In reply to | #163950 |
ram@zedat.fu-berlin.de (Stefan Ram) writes:
> Meredith Montgomery <mmontgomery@levado.to> writes:
>>of an array is its last byte, but agreed --- that's not quite its end
>>yet. It seems hard to define the exact end of something.
>
> I have developed my personal terminology for integer ranges:
>
> For example, take the range { 3, 4, 5, 6, 7 }.
>
> The "begin of the range" is the first number of the range,
> i.e., 3.
>
> The "bottom of the range" is the number before the first
> number of the range, i.e., 2.
>
> The "end of the range" is the last number of the range,
> i.e., 7.
>
> The "top of the range" is the number after the last number
> of the range, i.e., 8.
>
> In the standard library of C++, they call "end" what I
> call "top".
That's interesting. In technical context, it's nice to have a language
that reflects distinctions where there is.
[toc] | [prev] | [next] | [standalone]
| From | Manfred <noname@add.invalid> |
|---|---|
| Date | 2021-12-18 18:36 +0100 |
| Message-ID | <spl66f$1t5v$1@gioia.aioe.org> |
| In reply to | #163950 |
On 12/18/2021 4:02 PM, Meredith Montgomery wrote:
> Ben Bacarisse <ben.usenet@bsb.me.uk> writes:
>
>> Manfred <noname@add.invalid> writes:
>>
>>> On 12/16/2021 4:23 PM, David Brown wrote:
>>>> On 16/12/2021 15:13, Meredith Montgomery wrote:
>>> [...]
>>>>
>>>>> printf("a: %s\n", &a + 3);
>>>> "a" here is the full array, so "&a" is a pointer to an array of 13 char.
>>>> Adding 3 gives a new pointer to an array of 13 char, at the address 39
>>>> bytes higher than address of "a". Basically, you are pretending that
>>>> there is an array of 4 elements, each of which is itself an array of 13
>>>> chars - with "a" being the first of those 4 elements. Now you are
>>>> asking to print the 4th element here, which is just whatever happens to
>>>> be in the memory at that address.
>>>>
>>>
>>> It might be worth mentioning that accessing memory at that address (&a
>>> + 3) technically yields undefined behavior, because it is an attempt
>>> to access a location outside the array (a).
>>>
>>
>> That's a good point to make.
>
> Indeed. Thanks. (I was aware of that. In fact, I thought that my OP
> would produce a lot of --- that's undefined behavior and, so, C has
> nothing to do with it. Lol. But, thankfully, you guys went straight
> into helping me clear up my troubles.)
>
>> It reminds me of a construct that I rather like:
>>
>> char buffer[some-complex-size];
>> char *end_of_buffer = (&buffer)[1];
>>
>> (&buffer)[1] means *(&buffer + 1). The constructed pointer &buffer + 1
>> is valid, and the * does not attempt to dereference it as *(&buffer + 1)
>> is an array-valued expression. It is, instead, converted to a pointer
>> to the first element of this non-existent array object -- a pointer one
>> past the end of 'buffer'.
>
> Wow, that's a cool application. Thanks for sharing.
>
> I needed to redefine my definition of ``end''. I tend to think the end
> of an array is its last byte, but agreed --- that's not quite its end
> yet. It seems hard to define the exact end of something.
>
> The end of my property should be on my property or should it be outside
> of it? If *it* is outside, then because because *it* is relative to
> *my* property, this *it* must be mine, so it belongs to my property. So
> it's not outside of it. :-)
>
> I would have done
>
> char *end_of_buffer = (&buffer)[1] - 1;
>
It's more about habits in the context of C and how this "end" is used.
C uses the convention that indexes start at 0, which implies that a
collection of N elements has length N and is indexed from 0 to N-1.
Following the same reasoning, the following is customary in C:
#define N 42
int arr[N];
int* start = arr;
int* first = arr;
int* last = arr+N-1;
int* end = arr+N;
"end" defined this way is never dereferenced, but is typically used
(a.o.) as:
for (int* p = start; p != end; ++p)
{
*p = whatever;
}
Writing the same loop using "last" is also possible, but requires using
"<=" with pointers, which involves additional ordering requirements.
The standard follows the same habit, but is more accurate: it usually
refers to the boundary of a collection as "one past the end" of it.
However, it is this "one past the end" that is more often used rather
than the "last element" of the collection.
In fact, the "one past the end" pointer gets special attention from the
standard for the very purpose of making code like the above legal.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2021-12-18 21:17 +0000 |
| Message-ID | <87a6gxr4sf.fsf@bsb.me.uk> |
| In reply to | #163950 |
Meredith Montgomery <mmontgomery@levado.to> writes: > I needed to redefine my definition of ``end''. I tend to think the end > of an array is its last byte, but agreed --- that's not quite its end > yet. It seems hard to define the exact end of something. C permits, for largely historical reasons, the construction of a pointer "just past" the end of an object. As a result, you can process an array (or some part of an array) with a conventional idiom: for (char *ptr = start; ptr < end; ptr++) ... which parallels the method using indexes: for (int i = 0; i < n; i++) ... You can get the index from the pointer (start - ptr) and the pointer form the index (start + i) and everything works provided you don't access the data at the "just past" address. > The end of my property should be on my property or should it be outside > of it? If *it* is outside, then because because *it* is relative to > *my* property, this *it* must be mine, so it belongs to my property. So > it's not outside of it. :-) It's handy that the start and the end are not the same pointer unless the range is zero. When you get into the habit of using the "just past" pointer you will find all sorts of little things drop into place. -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-12-16 18:59 -0500 |
| Message-ID | <spgjsg$jml$1@dont-email.me> |
| In reply to | #163879 |
On 12/16/21 3:34 PM, Manfred wrote: [re: char a[] = "hello, world!"; ,,, > It might be worth mentioning that accessing memory at that address (&a + > 3) technically yields undefined behavior, because it is an attempt to > access a location outside the array (a). There's two problems with that statement. First of all, the expression &a+3 has undefined behavior all by itself (6.5.6p9), even if the code made no attempt to access the memory pointed at by the result of that expression. Secondly, such code simply has undefined behavior. There's no need to qualify it with "technically". Undefined behavior simply means that the standard doesn't impose any requirements on the resulting behavior - and it doesn't impose any requirements on this code. Keep in mind that undefined behavior allows, among infinitely many other things, having the code behave in precisely the manner you incorrectly thought it was required to behave. That case comes up pretty often, which isn't a coincidence. The false idea that there is some particular way that such code is required to behave doesn't just come out of nowhere - it develops in part because there are real implementations where it happens to have that behavior.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-12-16 18:38 -0500 |
| Message-ID | <spgimb$tct$1@dont-email.me> |
| In reply to | #163857 |
On 12/16/21 9:13 AM, Meredith Montgomery wrote:
> I'm investigating the syntax array[index] and I'm getting surprised at
> some places. My intuition says that a[3] is the same as a + 3. I also
That is correct. Interestingly, since a + 3 == 3 + a, it is also the
same as 3[a].
> know that /&a/ is the same as /a/. I first expected the following
That is incorrect. The relevant rule is
"Except when it is the operand of the sizeof operator, or the unary &
operator, or is a string literal used to initialize an array, an
expression that has type "array of type" is converted to an expression
with type "pointer to type" that points to the initial element of the
array object and is not an lvalue." (6.3.2.1p3).
Both the main rule, and two of the exceptions, comes into play in your
code below.
> program to print ``lo, world\n'' three times, but it does not.
>
> --8<---------------cut here---------------start------------->8---
> #include <stdio.h>
> int main() {
> char a[] = "hello, world";
"hello, world" is an expression of type char[12], and according to
6.3.2.1p3, would normally convert into a pointer to the 'h' at the start
of that array. In this context, that would be a constraint violation - a
pointer cannot be used to initialize an array. However, as explained in
6.3.2.1p3, that conversion does not occur when the string literal is
being used as an initializer for an array.
> printf("a: %s\n", &a[3]);
a[3] parses as a post-fix expression, and therefore as a
unary-expression, and the right operand of unary & is required to be a
unary expression.
Therefore, &a[3] gets parsed as &(a[3]), so 'a' does get converted into
a char* pointing at it's first element, so a[3] is equivalent to *(a+3).
The '&' doesn't apply to 'a' itself (which would prevent that conversion
from occurring), but to a[3], which is not an expression of array type,
so 6.3.2.1p3 doesn't come into play. Thus, we have the equivalent of
&*(a+3). The & and * cancel each other out, so it's simply a+3.
> printf("a: %s\n", a + 3);
'a' also converts into a pointer to it's first element here.
> printf("a: %s\n", &a + 3);
However, in this case a + 3 would be an additive expression. Unary &
cannot take an additive expression as it's right operand. Therefore,
this parses as (&a)+3. &a is one of the exceptions to the rule given in
6.3.2.1p3. Therefore, 'a' does NOT get converted into a value of pointer
type, which is good, because that value is explicitly not an lvalue, and
you can't take the address of something that isn't an lvalue.
'a' itself is an lvalue, and &a therefore results in a pointer of type
char(*)[12], a pointer to an entire array of 12 chars, not just one
char. And that's the fundamental problem.
The result of adding an integer to a pointer is defined in terms of
positions in an array of the pointed-at type. The pointed-at type in
this case is char[12]. There's only one such object in that location,
which is a itself. A single object of the pointed-at type is treated,
for the purposes of this rule, as if it were the first and only object
in a single-element array of the pointed-at type, or in other words, as
if it were char[1][12]. The problem is that you are asking it to move
the pointer 3 elements ahead in that array, and that array has only one
element. Such an expression has undefined behavior.
As a practical matter, what would happen on many systems is that they
would treats 'a' as if it were the first element in array[n][12], where
n is >= 4. Thus, (&a)+3 would result in a pointer to the 4th element of
that array. However, since there is not actually any such array, it
merely refers to whatever is actually located at the place where that
fourth element would have been, if it had existed. There's no guarantee
what's in that location, but it apparently is not another copy of the
"hello, world!" array.
[toc] | [prev] | [next] | [standalone]
| From | Bart <bc@freeuk.com> |
|---|---|
| Date | 2021-12-16 23:50 +0000 |
| Message-ID | <spgjc7$84b$1@dont-email.me> |
| In reply to | #163891 |
On 16/12/2021 23:38, James Kuyper wrote: > On 12/16/21 9:13 AM, Meredith Montgomery wrote: >> I'm investigating the syntax array[index] and I'm getting surprised at >> some places. My intuition says that a[3] is the same as a + 3. I also > > That is correct. Really? It's only equivalent in certain cases.
[toc] | [prev] | [next] | [standalone]
| From | "james...@alumni.caltech.edu" <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-12-17 07:34 -0800 |
| Message-ID | <1b0bd00a-a998-426b-9fbf-4a0461e60908n@googlegroups.com> |
| In reply to | #163893 |
On Thursday, December 16, 2021 at 6:50:41 PM UTC-5, Bart wrote: > On 16/12/2021 23:38, James Kuyper wrote: > > On 12/16/21 9:13 AM, Meredith Montgomery wrote: > >> I'm investigating the syntax array[index] and I'm getting surprised at > >> some places. My intuition says that a[3] is the same as a + 3. I also > > > > That is correct. > Really? It's only equivalent in certain cases. Meredith and I both made mistakes. He should have said `(*(a+3))`. Since he didn't say that, I should have pointed out his mistake. However, your comment surprises me. Were you referring to his statement as he originally wrote it? If so, what are the "certain cases" you're referring to? `a[3]` has the type char, while `a+3` has the type char* - it's hard for me to see how they could ever be equivalent. Or did you misread his statement the same way I did? If so, under what circumstances would `a[3]` not be equivalent to `(*(a+3))`? Section 6.5.2.1p2 says: > ... E1[E2] is identical to (*((E1)+(E2))).
[toc] | [prev] | [next] | [standalone]
| From | Bart <bc@freeuk.com> |
|---|---|
| Date | 2021-12-17 17:48 +0000 |
| Message-ID | <spiih9$vsu$1@dont-email.me> |
| In reply to | #163918 |
On 17/12/2021 15:34, james...@alumni.caltech.edu wrote:
> On Thursday, December 16, 2021 at 6:50:41 PM UTC-5, Bart wrote:
>> On 16/12/2021 23:38, James Kuyper wrote:
>>> On 12/16/21 9:13 AM, Meredith Montgomery wrote:
>>>> I'm investigating the syntax array[index] and I'm getting surprised at
>>>> some places. My intuition says that a[3] is the same as a + 3. I also
>>>
>>> That is correct.
>> Really? It's only equivalent in certain cases.
>
> Meredith and I both made mistakes.
Well, the OP is a newbie and that is a genuine misunderstanding.
> He should have said `(*(a+3))`. Since
> he didn't say that, I should have pointed out his mistake. However, your
> comment surprises me.
>
> Were you referring to his statement as he originally wrote it? If so, what
> are the "certain cases" you're referring to? `a[3]`
I'm refering partly to his first paragraph (I'd only glanced at the rest
of the post), and partly to my original reply to that, which I will
repeat here:
-------------------------------------------------------
It's the same as *(a+3).
However if a is an array of arrays, then the array element at *(a+3)
might decay to a pointer of that array, so it could end up as the same
value as a+3, if different type:
int a[5][4];
printf("%p\n", a[3]);
printf("%p\n", *(a+3));
printf("%p\n", a+3);
These all show the same address. But change 'a' to 'int a[5]', and the
first two lines show the value, the element a[3] (some random value if
not initialised), and the last shows the address of that element.
-------------------------------------------------------
So I say that sometimes, *(a+3) can have the same vaue as a+3.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2021-12-16 23:56 +0000 |
| Message-ID | <87k0g4t88f.fsf@bsb.me.uk> |
| In reply to | #163891 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes: > On 12/16/21 9:13 AM, Meredith Montgomery wrote: >> I'm investigating the syntax array[index] and I'm getting surprised at >> some places. My intuition says that a[3] is the same as a + 3. I also > > That is correct. Interestingly, since a + 3 == 3 + a, it is also the > same as 3[a]. I think you have not looked at what the OP wrote carefully enough. a[3] == *(a + 3) == *(3 + a) == 3[a] but not without the *(...) part! -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2021-12-16 19:06 -0500 |
| Message-ID | <spgkam$s7k$1@dont-email.me> |
| In reply to | #163894 |
On 12/16/21 6:56 PM, Ben Bacarisse wrote: > James Kuyper <jameskuyper@alumni.caltech.edu> writes: > >> On 12/16/21 9:13 AM, Meredith Montgomery wrote: >>> I'm investigating the syntax array[index] and I'm getting surprised at >>> some places. My intuition says that a[3] is the same as a + 3. I also >> >> That is correct. Interestingly, since a + 3 == 3 + a, it is also the >> same as 3[a]. > > I think you have not looked at what the OP wrote carefully enough. I did look closely enough. > a[3] == *(a + 3) == *(3 + a) == 3[a] > > but not without the *(...) part! I just forgot to put in the '*'. I also miscounted the length of the array, which I consider to be even more embarrassing.
[toc] | [prev] | [next] | [standalone]
| From | Ben Bacarisse <ben.usenet@bsb.me.uk> |
|---|---|
| Date | 2021-12-17 00:10 +0000 |
| Message-ID | <87ee6ct7jq.fsf@bsb.me.uk> |
| In reply to | #163897 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes: > On 12/16/21 6:56 PM, Ben Bacarisse wrote: >> James Kuyper <jameskuyper@alumni.caltech.edu> writes: >> >>> On 12/16/21 9:13 AM, Meredith Montgomery wrote: >>>> I'm investigating the syntax array[index] and I'm getting surprised at >>>> some places. My intuition says that a[3] is the same as a + 3. I also >>> >>> That is correct. Interestingly, since a + 3 == 3 + a, it is also the >>> same as 3[a]. >> >> I think you have not looked at what the OP wrote carefully enough. > > I did look closely enough. The OP wrote "a[3] is the same as a + 3" and you said "That is correct". -- Ben.
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-18 12:32 -0300 |
| Message-ID | <86fsqq53p9.fsf@levado.to> |
| In reply to | #163897 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes: > On 12/16/21 6:56 PM, Ben Bacarisse wrote: >> James Kuyper <jameskuyper@alumni.caltech.edu> writes: >> >>> On 12/16/21 9:13 AM, Meredith Montgomery wrote: >>>> I'm investigating the syntax array[index] and I'm getting surprised at >>>> some places. My intuition says that a[3] is the same as a + 3. I also >>> >>> That is correct. Interestingly, since a + 3 == 3 + a, it is also the >>> same as 3[a]. >> >> I think you have not looked at what the OP wrote carefully enough. > > I did look closely enough. > >> a[3] == *(a + 3) == *(3 + a) == 3[a] >> >> but not without the *(...) part! > > I just forgot to put in the '*'. I also miscounted the length of the > array, which I consider to be even more embarrassing. I assumed you just skipped a letter, which could happen to anyone.
[toc] | [prev] | [next] | [standalone]
| From | Meredith Montgomery <mmontgomery@levado.to> |
|---|---|
| Date | 2021-12-18 12:30 -0300 |
| Message-ID | <86o85e53sz.fsf@levado.to> |
| In reply to | #163891 |
James Kuyper <jameskuyper@alumni.caltech.edu> writes:
> On 12/16/21 9:13 AM, Meredith Montgomery wrote:
>> I'm investigating the syntax array[index] and I'm getting surprised at
>> some places. My intuition says that a[3] is the same as a + 3. I also
>
> That is correct. Interestingly, since a + 3 == 3 + a, it is also the
> same as 3[a].
>
>> know that /&a/ is the same as /a/. I first expected the following
>
> That is incorrect. The relevant rule is
>
> "Except when it is the operand of the sizeof operator, or the unary &
> operator, or is a string literal used to initialize an array, an
> expression that has type "array of type" is converted to an expression
> with type "pointer to type" that points to the initial element of the
> array object and is not an lvalue." (6.3.2.1p3).
Thanks for the citation and the technical analysis. Much appreciated.
[...]
>> program to print ``lo, world\n'' three times, but it does not.
>>
>> --8<---------------cut here---------------start------------->8---
>> #include <stdio.h>
>> int main() {
>> char a[] = "hello, world";
>
> "hello, world" is an expression of type char[12], [...]
You must have skipped a letter while counting. With '\0' at the end, we
get char[13], but that did not complicated the reading in any way.
[...]
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | comp.lang.c
csiph-web