Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.c > #41434 > unrolled thread

Padding involved

Started byanish singh <anish198519851985@gmail.com>
First post2014-03-07 13:13 -0800
Last post2014-03-08 09:01 -0800
Articles 19 on this page of 79 — 17 participants

Back to article view | Back to comp.lang.c


Contents

  Padding involved anish singh <anish198519851985@gmail.com> - 2014-03-07 13:13 -0800
    Re: Padding involved jacob navia <jacob@spamsink.net> - 2014-03-07 22:17 +0100
    Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-07 16:26 -0500
      Re: Padding involved anish singh <anish198519851985@gmail.com> - 2014-03-07 13:49 -0800
        Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-07 18:11 -0500
          Re: Padding involved anish kumar <yesanishhere@gmail.com> - 2014-03-07 15:19 -0800
            Re: Padding involved Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2014-03-07 16:26 -0700
              Re: Padding involved anish kumar <yesanishhere@gmail.com> - 2014-03-07 15:41 -0800
                Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-07 18:50 -0500
              Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-08 00:10 +0000
                Re: Padding involved Joe Pfeiffer <pfeiffer@cs.nmsu.edu> - 2014-03-08 09:27 -0700
            Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-07 16:24 -0800
              Re: Padding involved anish kumar <yesanishhere@gmail.com> - 2014-03-07 16:45 -0800
                Re: Padding involved Kaz Kylheku <kaz@kylheku.com> - 2014-03-08 01:03 +0000
                  Re: Padding involved David Thompson <dave.thompson2@verizon.net> - 2014-03-28 06:11 -0400
                    Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-28 18:45 +0000
                      Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-28 15:23 -0400
                        Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-28 12:40 -0700
                        Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-28 21:19 +0000
                          Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-28 17:07 -0500
                            Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-28 18:25 -0400
                              Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-28 21:15 -0500
                                Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-29 01:03 -0700
                                  Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-03-29 14:23 -0700
                                    Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-30 01:43 +0000
                            Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-28 21:32 -0400
                              Re: Padding involved Richard Damon <Richard@Damon-Family.org> - 2014-03-28 21:59 -0400
                              Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-29 00:46 -0500
                                Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-29 08:51 -0400
                                  Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-29 08:57 -0400
                                    Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-30 18:53 -0500
                                      Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-30 20:25 -0400
                                        Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-31 11:34 -0500
                                          Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-31 11:27 -0700
                                            Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-31 11:36 -0700
                                              Re: Padding involved Kaz Kylheku <kaz@kylheku.com> - 2014-03-31 19:14 +0000
                                                Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-31 15:43 -0500
                                                  Re: Padding involved Phil Carmody <thefatphil_demunged@yahoo.co.uk> - 2014-04-01 03:12 +0300
                                                    Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-04-03 00:45 -0500
                                  Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-29 13:02 -0500
                                    Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-29 14:44 -0400
                                    Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-29 19:49 +0000
                                    Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-30 19:12 -0400
                                      Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-31 16:20 -0500
                                        Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-31 17:53 -0400
                                          Re: Padding involved Phil Carmody <thefatphil_demunged@yahoo.co.uk> - 2014-04-01 03:22 +0300
                                          Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-04-04 14:36 -0500
                                            Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-04-04 17:33 -0400
                                            Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-04-14 13:46 -0700
                                              Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-04-14 22:00 +0000
                                                Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-04-14 15:37 -0700
                                            Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-04-17 09:11 -0700
                                      Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-04-14 13:25 -0700
                              Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-29 00:57 -0700
                                Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-29 08:57 -0400
                                  Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-29 14:18 -0700
                                    Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-03-30 12:02 -0700
                                      Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-30 19:35 -0700
                                        Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-04-14 12:54 -0700
                                  Re: Padding involved Tim Rentsch <txr@alumni.caltech.edu> - 2014-03-29 14:44 -0700
                            Re: Padding involved glen herrmannsfeldt <gah@ugcs.caltech.edu> - 2014-03-29 03:50 +0000
                              Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-29 09:02 -0400
                              Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-29 13:19 -0500
                          Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-28 21:21 -0400
                      Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-28 15:27 -0400
                      Re: Padding involved Kaz Kylheku <kaz@kylheku.com> - 2014-03-28 19:54 +0000
                      Re: Padding involved Stephen Sprunk <stephen@sprunk.org> - 2014-03-28 15:02 -0500
            Re: Padding involved Kaz Kylheku <kaz@kylheku.com> - 2014-03-08 00:43 +0000
            Re: Padding involved Eric Sosman <esosman@comcast-dot-net.invalid> - 2014-03-07 20:20 -0500
      Re: Padding involved anish singh <anish198519851985@gmail.com> - 2014-03-07 13:49 -0800
    Re: Padding involved James Kuyper <jameskuyper@verizon.net> - 2014-03-07 16:46 -0500
    Re: Padding involved "BartC" <bc@freeuk.com> - 2014-03-07 21:59 +0000
      Re: Padding involved anish kumar <yesanishhere@gmail.com> - 2014-03-07 14:34 -0800
        Re: Padding involved "BartC" <bc@freeuk.com> - 2014-03-07 23:59 +0000
      Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-07 15:15 -0800
        Re: Padding involved "BartC" <bc@freeuk.com> - 2014-03-08 10:08 +0000
          Re: Padding involved Ben Bacarisse <ben.usenet@bsb.me.uk> - 2014-03-08 12:39 +0000
          Re: Padding involved Keith Thompson <kst-u@mib.org> - 2014-03-08 14:03 -0800
    Re: Padding involved Malcolm McLean <malcolm.mclean5@btinternet.com> - 2014-03-08 09:01 -0800

Page 4 of 4 — ← Prev page 1 2 3 [4]


#42273

Fromglen herrmannsfeldt <gah@ugcs.caltech.edu>
Date2014-03-29 03:50 +0000
Message-ID<lh5ftp$ehe$1@speranza.aioe.org>
In reply to#42266
Stephen Sprunk <stephen@sprunk.org> wrote:
(snip)
>>> On the platform you describe, must every double be aligned on a 16
>>> byte address, so the SSE instructions can always be used?

(then I wrote)
>> Pairs of doubles are aligned on 16 byte boundaries.
 
> Standard C has no type "pair of doubles".
 
> Standard C guarantees that _Alignof(double) <= sizeof(double).  
> On x86, we know that sizeof(double) == 8, so _Alignof(double) == 16 
> is not allowed.

Standard C doesn't care about speed at all, but users often do.

Note that x86, back to the 8086, doesn't require alignment, but it
is often faster if properly aligned. When the 80486 was popular,
and four byte alignment of double was all that was needed.
(A 32 bit system, with a 32 bit data bus.)  C compilers, and more
important most of the time, malloc() would generate four byte
alignment.

For way too long after the pentium became popular, C was still
generating four byte alignment.

>> If you have an array of even length, you could process them two at a
>> time if appropriately aligned. You can then, for example, add a pair of 
>> doubles to another pair in one operation.
 
> Standard C only guarantees that your array of doubles will have 
> the same alignment as one double, i.e. 8 bytes on x86.

Note as above, x86 doesn't require 8 byte alignment for doubles,
so C might as well not do any padding, and malloc() might just
as well return odd addresses. 
 
> If you want a guarantee that your array has 16-byte alignment, then you
> must either use/create another type with 16-byte alignment as your array
> element or use an extension to tell the compiler you want stricter
> alignment for a double (or array of doubles) than Standard C requires.

Or give up C and move onto other languages?
 
> Note that Standard C doesn't guarantee the existence of _any_ type with
> 16-byte alignment or the ability to create such, so the former may not
> be possible, and the latter is inherently outside the Standard.
 
(snip)

> If the subroutine is written in assembler, then obviously Standard C
> says nothing about what it can or can't do, nor does Standard C
> guarantee that a pointer-to-double you pass to it will be aligned as
> expected.
 
> If the subroutine were in Standard C, the compiler must properly handle
> the 8-byte aligned case.  However, there is nothing stopping it from
> _also_ detecting the 16-byte aligned case and then using more efficient
> vector instructions.

But if you can't reliably generate them, that doesn't help much.
 
>> In the struct case, one might have a struct with a pair of doubles 
>> (or one complex double) along with some other types, and want the 
>> pair of doubles appropriately aligned, even in an array of such.
 
> Assuming your struct only contains doubles, then the alignment will be
> the same as for one double, like in the array case above.

And if it doesn't?

-- glen

[toc] | [prev] | [next] | [standalone]


#42284

FromJames Kuyper <jameskuyper@verizon.net>
Date2014-03-29 09:02 -0400
Message-ID<lh6ga4$qje$1@dont-email.me>
In reply to#42273
On 03/28/2014 11:50 PM, glen herrmannsfeldt wrote:
> Stephen Sprunk <stephen@sprunk.org> wrote:
...
>> If you want a guarantee that your array has 16-byte alignment, then you
>> must either use/create another type with 16-byte alignment as your array
>> element or use an extension to tell the compiler you want stricter
>> alignment for a double (or array of doubles) than Standard C requires.
> 
> Or give up C and move onto other languages?

That seems like an over-reaction. Actual C compilers are far less
user-unfriendly than the C standard allows them to be. Also, C2011 added
_Alignas(), which would seem to address the issue you're raising.
-- 
James Kuyper

[toc] | [prev] | [next] | [standalone]


#42299

FromStephen Sprunk <stephen@sprunk.org>
Date2014-03-29 13:19 -0500
Message-ID<lh72s7$4ki$1@dont-email.me>
In reply to#42273
On 28-Mar-14 22:50, glen herrmannsfeldt wrote:
> Stephen Sprunk <stephen@sprunk.org> wrote: (snip)
>>
>>> Pairs of doubles are aligned on 16 byte boundaries.
>> 
>> Standard C has no type "pair of doubles".
>> 
>> Standard C guarantees that _Alignof(double) <= sizeof(double). On
>> x86, we know that sizeof(double) == 8, so _Alignof(double) == 16 is
>> not allowed.
> 
> Standard C doesn't care about speed at all, but users often do.
> 
> Note that x86, back to the 8086, doesn't require alignment, but it is
> often faster if properly aligned. When the 80486 was popular, and
> four byte alignment of double was all that was needed. (A 32 bit
> system, with a 32 bit data bus.)  C compilers, and more important
> most of the time, malloc() would generate four byte alignment.
> 
> For way too long after the pentium became popular, C was still 
> generating four byte alignment.

Changing the alignment could break binary compatibility, which is a
bigger issue on some platforms than others.

My GCC targets "i486", and it offers 8-byte alignment for double.  There
is little harm in requiring 8-byte alignment when its output runs on a
486, but there is significant benefit when it runs on a Pentium or later.

>>> If you have an array of even length, you could process them two
>>> at a time if appropriately aligned. You can then, for example,
>>> add a pair of doubles to another pair in one operation.
>> 
>> Standard C only guarantees that your array of doubles will have the
>> same alignment as one double, i.e. 8 bytes on x86.
> 
> Note as above, x86 doesn't require 8 byte alignment for doubles, so C
> might as well not do any padding, and malloc() might just as well
> return odd addresses.

True, but we know that unaligned accesses have a performance penalty, so
implementations _should_ align when possible as a QoI matter.  And C has
to work on systems that don't support unaligned access at all.

>> If you want a guarantee that your array has 16-byte alignment, then
>> you must either use/create another type with 16-byte alignment as
>> your array element or use an extension to tell the compiler you
>> want stricter alignment for a double (or array of doubles) than
>> Standard C requires.
> 
> Or give up C and move onto other languages?

*shrug* That's obviously OT here.

>> If the subroutine were in Standard C, the compiler must properly
>> handle the 8-byte aligned case.  However, there is nothing stopping
>> it from _also_ detecting the 16-byte aligned case and then using
>> more efficient vector instructions.
> 
> But if you can't reliably generate them, that doesn't help much.

An implementation _can_ fairly reliably generate them as a QoI matter,
even if it isn't required to do so.  But since it's not required, it
still has to be able to handle the unaligned case, which isn't
completely avoidable.

>>> In the struct case, one might have a struct with a pair of
>>> doubles (or one complex double) along with some other types, and
>>> want the pair of doubles appropriately aligned, even in an array
>>> of such.
>> 
>> Assuming your struct only contains doubles, then the alignment will
>> be the same as for one double, like in the array case above.
> 
> And if it doesn't?

If it doesn't contain only doubles?  Then we'd need to know what type
the other members are and what their alignment is to be able to
determine the alignment for the whole structure.

S

-- 
Stephen Sprunk         "God does not play dice."  --Albert Einstein
CCIE #3723         "God is an inveterate gambler, and He throws the
K5SSS        dice at every possible opportunity." --Stephen Hawking

[toc] | [prev] | [next] | [standalone]


#42268

FromJames Kuyper <jameskuyper@verizon.net>
Date2014-03-28 21:21 -0400
Message-ID<lh5776$ot6$1@dont-email.me>
In reply to#42265
On 03/28/2014 05:19 PM, glen herrmannsfeldt wrote:
> James Kuyper <jameskuyper@verizon.net> wrote:
...
>> (then I  [glen herrmannsfeldt] wrote)
>>> As I understand it, for some systems an alignment greater than
>>> size is necessary for optimal use. Specifically, some of the SSE
>>> instructions, as I understand it, will process pairs of doubles
>>> aligned to 16 byte boundaries.  I believe other combinations, such
>>> as four floats or four ints.
>  
>> In C, every object of a given type must be allocated at a location which
>> is correctly aligned for it's type. In an array, that means that the
>> first element of the array and the second element must both be correctly
>> aligned - but those two positions are also required to be separated by
>> exactly sizeof(type) bytes. That's not possible unless sizeof(type) is
>> an integer multiple of _Alignof(type).
>  
>> On the platform you describe, must every double be aligned on a 16 byte
>> address, so the SSE instructions can always be used? 
> 
> Pairs of doubles are aligned on 16 byte boundaries.

If two double can be 8 bytes apart, then _Alignof(double)<=8. Since
you'll probably want have at least one double in any group of two or
more doubles to be aligned on a 16 byte boundary, that suggests that an
implementation for that platform should choose _Alignof(double)==8.

> ... If you have an
> array of even length, you could process them two at a time if
> appropriately aligned. You can then, for example, add a pair of 
> doubles to another pair in one operation. 

That is an optimization that a compiler is allowed to take advantage of,
when it can - but if it doesn't prevent the the existence of doubles
starting on addresses that are not multiples of 16, then it does NOT
mean that _Alignof(double) == 16.

> In some cases, the compiler might be able to generate appropriate
> code, for example adding complex data.

That would suggest that there is a strong incentive for _Alignof(double
_Complex) == 16, but that's a different issue.
-- 
James Kuyper

[toc] | [prev] | [next] | [standalone]


#42261

FromEric Sosman <esosman@comcast-dot-net.invalid>
Date2014-03-28 15:27 -0400
Message-ID<lh4ifo$42q$1@dont-email.me>
In reply to#42258
On 3/28/2014 2:45 PM, glen herrmannsfeldt wrote:
> David Thompson <dave.thompson2@verizon.net> wrote:
>
> (snip)
>
>> No, the reverse. The offset must divide the size, or the size must be
>> divisible by the offset. E.g. struct { long a; float b; } if both long
>> and float are 4 bytes (as is common, though not universal and not
>> required by the Standard) then the struct has size at least 8 but is
>> unlikely to have alignment more than 4.
>
> (snip)
>
>> Note that alignment can be less than size, most commonly on systems
>> that can align everything to 1, but I've used a system where int is 4
>> bytes and aligned to 2. Alignment cannot be more than size.
>
> As I understand it, for some systems an alignment greater than
> size is necessary for optimal use. Specifically, some of the SSE
> instructions, as I understand it, will process pairs of doubles
> aligned to 16 byte boundaries.  I believe other combinations, such
> as four floats or four ints.
>
> Also, GPUs might have different alignment requirements than
> traditional processors.

     From C's standpoint, alignment can never exceed size: Arrays
would not work if it did.

     It can still be true -- is true -- that the host system may
require or benefit from alignments that are unknown to C.  For
example, O/S interfaces like Unix' mmap() require alignment on
memory pages.  But "memory page" is not a C type, nor even a C
concept, and there's no direct way for C to control memory page
alignment.  (Even with C11's _Alignas keyword, there's no way C
can discover the memory page size unaided -- and on systems that
support multiple page sizes simultaneously, the situation gets
even thornier.)

-- 
Eric Sosman
esosman@comcast-dot-net.invalid

[toc] | [prev] | [next] | [standalone]


#42263

FromKaz Kylheku <kaz@kylheku.com>
Date2014-03-28 19:54 +0000
Message-ID<20140328124409.987@kylheku.com>
In reply to#42258
On 2014-03-28, glen herrmannsfeldt <gah@ugcs.caltech.edu> wrote:
> David Thompson <dave.thompson2@verizon.net> wrote:
>
> (snip)
>
>> No, the reverse. The offset must divide the size, or the size must be
>> divisible by the offset. E.g. struct { long a; float b; } if both long
>> and float are 4 bytes (as is common, though not universal and not
>> required by the Standard) then the struct has size at least 8 but is
>> unlikely to have alignment more than 4.
>
> (snip)
>
>> Note that alignment can be less than size, most commonly on systems
>> that can align everything to 1, but I've used a system where int is 4
>> bytes and aligned to 2. Alignment cannot be more than size.
>
> As I understand it, for some systems an alignment greater than
> size is necessary for optimal use.

C does not support this. C compilers can support extra alignment for efficient
access in the way local variables are laid out and perhaps struct members.

No such thing will be supported for arrays and pointers.

If a greater alignment than size is required for correctness, then misaligned
access for pointers and arrays must be implemented.

E.g. if a short is two bytes, but must be aligned on a four-byte boundary, then
code generates for array indexing and pointer dereferncing has to somehow
handle the accesses at odd indices. Perhaps by rounding down to an address
divisible by four, loading a four byte word, and then shifting down
the half-word.

The guys who designed C were no strangers to machines that didn't provide
access to certain small types such as characters.

In fact, the B language, predecessor to C, handled strings similarly to C:
characters were packed into arrays of cells, which had to be unpacked and
re-packed by routines.

   http://cm.bell-labs.com/who/dmr/chist.html

  "[B's] character-handling mechanisms, inherited with few changes from BCPL,
   were clumsy: using library procedures to spread packed strings into
   individual cells and then repack, or to access and replace individual
   characters, began to feel awkward, even silly, on a byte-oriented machine. "

So at that point Ritchie went for an addressable character type.
That can basically be seen as the point of departure at which the design
of C shifted toward "every type, down to the character/byte, is accessible at
an address that is no more strictly aligned than a multiple of its size".

[toc] | [prev] | [next] | [standalone]


#42264

FromStephen Sprunk <stephen@sprunk.org>
Date2014-03-28 15:02 -0500
Message-ID<lh4kh9$ke6$1@dont-email.me>
In reply to#42258
On 28-Mar-14 13:45, glen herrmannsfeldt wrote:
> David Thompson <dave.thompson2@verizon.net> wrote:
>> Note that alignment can be less than size, most commonly on
>> systems that can align everything to 1, but I've used a system
>> where int is 4 bytes and aligned to 2. Alignment cannot be more
>> than size.
> 
> As I understand it, for some systems an alignment greater than size
> is necessary for optimal use. Specifically, some of the SSE 
> instructions, as I understand it, will process pairs of doubles 
> aligned to 16 byte boundaries.  I believe other combinations, such as
> four floats or four ints.

Such instructions operate not on a single object but on a group of
objects, and it is the _group_ that must have greater alignment;
however, that is invisible at the C level as long as you're using ints,
floats, etc.  If the compiler wants to auto-vectorize access to an array
of such objects, it is responsible for generating code to handle any
potential alignment issues at the front--and dealing with remainders at
the end.

One alternative is a compiler extension to create vector types, such as
GCC's vector_size attribute; they are similar to (short) arrays but
always have the correct alignment for vector instructions, unlike normal
arrays, and that carries through to arrays of vectors.  For instance:

typedef int v4si __attribute__ ((vector_size (16)));
v4si a = {1,2,3,4};                // always aligned
v4si b[2] = {{1,2,3,4},{5,6,7,8}}; // always aligned
int c[4] = {1,2,3,4};              // maybe unaligned

S

-- 
Stephen Sprunk         "God does not play dice."  --Albert Einstein
CCIE #3723         "God is an inveterate gambler, and He throws the
K5SSS        dice at every possible opportunity." --Stephen Hawking

[toc] | [prev] | [next] | [standalone]


#41454

FromKaz Kylheku <kaz@kylheku.com>
Date2014-03-08 00:43 +0000
Message-ID<20140307162649.67@kylheku.com>
In reply to#41447
On 2014-03-07, anish kumar <yesanishhere@gmail.com> wrote:
> On Friday, March 7, 2014 3:11:20 PM UTC-8, Eric Sosman wrote:
>> On 3/7/2014 4:49 PM, anish singh wrote:
>> 
>>  > [... a question about the size and padding of a struct
>> 
>>  > "defined" by uncompilable code ...]
>> 
>> > I have not given a compilable code. I am just
>> 
>> > asking the size of the struct given.
>> 
>> 
>> 
>>      How big is this array:
>> 
>> 
>> 
>> 	int array<7>;
>> 
>> 
>> 
>> ?  In other words, if the code describing your struct won't even
>> 
>> compile, then you have not "given" a struct at all.  If there is
>> 
>> no struct, it has no size and no padding -- and no existence.
> Understood. How about below:
> int main(void) {
>  	struct test {
> 		char a;
> 		int b;
> 		short c;
> 	};
> 	printf("%d\n", sizeof(struct test));
> 	return 0;
> }
> I completely understand what will be the size of the struct with and
> without mapping but the question is what parameters decides the padding
> involved? Such as size of registers or size of address/data bus of the
> processor?

Padding involved is determined by the "ABI" rules for the given architecture.
It has grave impact for the interoperability of programs, especially ina mixed
environment either with multiple compilers for C or C-related dialects, and
other languages that need to "bind" to C interfaces.

The width of the address or data bus of the processor is a very low-level
implementation detail on the actual silicon die, and is largely irrelevant.
Moreover, there is more than noe bus. Are you talking about the connection
betwen the L1 cache and L2 cache? A processor may read an entire cache line (or
several of them in burst mode) at a time from main memory nowadays; that
doesn't mean we align every structure member to a cache line.

ABI rules also span multiple implementations of an architecture.  If we are
compiling for 32 bit x86, we might tell the compiler to optimize for a 386,
486, Pentium, i7 or whatever, but the layout of the structures should be
interoperable across the family.

I think how it will work on GCC targetting 32 bit Intel is this.

The int will be padded so that it is aligned to an offset divisible by 4,
and the short will be padded so that it is aligned to an offset divisible by 2.

The reason for this is not that the alignment must be there, because processors
in this family support unaligned reads. It's for efficiency of access.
Even processors that can read a word at any byte address stil read it faster
if the address is aligned.

So there is a byte for "a", then three padding bytes. Then "b" is placed,
occupying four bytes, bringing us to offset 8. This is divisible by two, so "c"
is placed there taking two bytes, for a total of ten.

If we add a second "char a2" after "char a", it should go into the padding.

Generally, if a type with a weaker alignment is placed after a type with a
stronger alignment, it shouldn't need alignment. Conversely,
if a type with stronger alignment is placed after one with weaker alignment,
then it may require padding up to an offset that is a multiple of its
alignment.

Furthermore, there may be additional padding at the end of a structure, to
support the notion that structures can be combined together to form an array,
whereby the padding at the end of element [n] establishes the alignment of the
first member of element[n+1].

Thus a structure which is like this { int a; char b; } might have,
depending no architectural details, some bytes of padding after the b, so that
the overall size is divisible by sizeof(int). On Intel x86, I might expect
no padding between a and b, and three bytes after b.

These are very general concepts and the deatils vary quite a lot among
architectures.

For some architectures, the vendors or other organizations who standardize the
architectures, develop a set of documents which specify the ABI. The documents
dicate everything from how structures are laid out, to what stack frames look
like, what registers are used for what, how arguments are passed between
functions and so on. If there is such a body of standards, then generally the
compiler implementor follows that.

[toc] | [prev] | [next] | [standalone]


#41457

FromEric Sosman <esosman@comcast-dot-net.invalid>
Date2014-03-07 20:20 -0500
Message-ID<lfdr9r$7qs$1@dont-email.me>
In reply to#41447
On 3/7/2014 6:19 PM, anish kumar wrote:
> On Friday, March 7, 2014 3:11:20 PM UTC-8, Eric Sosman wrote:
>> On 3/7/2014 4:49 PM, anish singh wrote:
>>   > [... a question about the size and padding of a struct
>>   > "defined" by uncompilable code ...]
>>> I have not given a compilable code. I am just
>>> asking the size of the struct given.
>>
>>       How big is this array:
>>
>> 	int array<7>;
>>
>> ?  In other words, if the code describing your struct won't even
>> compile, then you have not "given" a struct at all.  If there is
>>
>> no struct, it has no size and no padding -- and no existence.
> Understood. How about below:
> int main(void) {
>   	struct test {
> 		char a;
> 		int b;
> 		short c;
> 	};

     *Much* better!

> 	printf("%d\n", sizeof(struct test));
> 	return 0;
> }
> I completely understand what will be the size of the struct with and
> without mapping but the question is what parameters decides the padding
> involved? Such as size of registers or size of address/data bus of the
> processor?

     There are different answers at different levels.

     At the hardware level, alignment and the padding to attain
it are artifacts of the memory subsystem: Not just the busses,
but the address translation circuitry, the various levels of
cache, the inter-processor data-consistency protocols, and so
on.  This collection of components may find it easier or cheaper
or faster to access particular types on a restricted set of
addresses: For example, it might be advantageous to position a
`double' object on an eight-byte boundary or an `int' on a
four-byte boundary.  (I don't think register widths have much
to do with this -- but I'm no hardware designer, so don't treat
my "I don't think" as Gospel.)

     But that's not the whole story.  A host seldom consists only
of the hardware; there's an operating system to think about.  The
O/S usually specifies an "application binary interface" or ABI
that describes how data should be arranged when invoking system
services or when dissecting their results.  For example, even if
the hardware is able to cope with an `int' at an arbitrary address,
the ABI might insist on four-byte alignment (one possible reason
to do so could be to simplify the "Does the caller really have
access to all the bytes implied by this pointer?" test, by not
having to worry about crossing page boundaries).  Then again, an
ABI might choose *not* to cater to all the hardware's whims: For
example, a widely-used ABI calls for four-byte alignment of all
stack-allocated data, even `double' objects that would be
significantly faster if allocated on eight-byte boundaries.

     But even that isn't the whole story.  Eventually, it's the
developers of the compiler itself who decide what policies it will
enforce.  One developer might say "Speed is important: We'll put
every object on an address that makes accesses the very fastest
they can possibly be."  Another might say "Memory bloat should be
avoided: We'll pack the objects as tightly as we can while still
maintaining reasonable (not optimal) speed."  Yet another might
say "This is an embedded machine with only 4KB of RAM, so memory
is an extremely scarce resource and we'll pack everything down to
the absolute minimum."  In the end, that is, it's a human choice.

-- 
Eric Sosman
esosman@comcast-dot-net.invalid

[toc] | [prev] | [next] | [standalone]


#41439

Fromanish singh <anish198519851985@gmail.com>
Date2014-03-07 13:49 -0800
Message-ID<5ef8fc0f-f43f-40c7-9c2b-37a6afa5d2a4@googlegroups.com>
In reply to#41436
I have not given a compilable code. I am just 
asking the size of the struct given.

[toc] | [prev] | [next] | [standalone]


#41437

FromJames Kuyper <jameskuyper@verizon.net>
Date2014-03-07 16:46 -0500
Message-ID<531A3E4B.1090301@verizon.net>
In reply to#41434
On 03/07/2014 04:13 PM, anish singh wrote:
> Struct abbcd{
>    Char c;
>    Int b;
>    Short d;
> };

The keywords "struct", "char", "int" and "short" are all lower case.
This may seem like quibbling, but C is a case sensitive language - if
you plan to use C, you need to learn to be careful about case.

> What will be the size of abbcd ? If padding involved and without padding?Suppose that the processor has only 4 byte registers.
> 
> What will be the size if the particular processor has register for 2 byte, 4 byte and 1 byte?
> 
> Note; size of char is 1,size of int, short is 4 and 2 respectively.

The only answer that works across all implementations is "sizeof(struct
abbcd)".

Without padding, the size will be 7 bytes, though it depends upon the
implementation whether you even have the option of avoiding padding.
With padding, it will be larger than 7, and almost certainly smaller
than SIZE_MAX. The actual value within that range depends upon the
implementation. If you want a more specific answer, the information
you've provided is insufficient to answer it. You need to fully specify
which implementation of C you're using: identify which compiler you're
using, and what compiler options you've chosen, including the target
platform. Once you've specified those things, the easiest way to find
out is to print out the value of sizeof(struct abbcd). That's a lot
quicker than asking us.

[toc] | [prev] | [next] | [standalone]


#41441

From"BartC" <bc@freeuk.com>
Date2014-03-07 21:59 +0000
Message-ID<vjrSu.26037$r53.19153@fx02.am4>
In reply to#41434

"anish singh" <anish198519851985@gmail.com> wrote in message 
news:43b13d99-38eb-4711-824d-5fc40d55edf3@googlegroups.com...
> Struct abbcd{
>   Char c;
>   Int b;
>   Short d;
> };
>
> What will be the size of abbcd ? If padding involved and without 
> padding?Suppose that the processor has only 4 byte registers.
>
> What will be the size if the particular processor has register for 2 byte, 
> 4 byte and 1 byte?
>
> Note; size of char is 1,size of int, short is 4 and 2 respectively.

Do you have access to a C compiler? Then you can do your own experiments 
with programs such as the following. (Note the #pragma line, to turn off 
padding for alignment, will vary between compilers.)

The offsets and padding will depend more on the memory alignments needed for 
the machine, then the sizes of the registers.

#include <stdio.h>
#include <stddef.h>

int main(void) {

struct abbcd {
 char c;
 int b;
 short d;
};

#pragma pack(1)
struct abbcd_packed {
 char c;
 int b;
 short d;
};

printf("Size of char  = %d\n",sizeof(char));
printf("Size of int   = %d\n",sizeof(int));
printf("Size of short = %d\n",sizeof(short));
puts("");

puts("Normal padding:");
printf("Offset of c   = %d\n",offsetof(struct abbcd,c));
printf("Offset of b   = %d\n",offsetof(struct abbcd,b));
printf("Offset of d   = %d\n",offsetof(struct abbcd,d));
printf("Size of abbcd = %d\n",sizeof(struct abbcd));
puts("");

puts("Without padding:");
printf("Offset of c          = %d\n",offsetof(struct abbcd_packed,c));
printf("Offset of b          = %d\n",offsetof(struct abbcd_packed,b));
printf("Offset of d          = %d\n",offsetof(struct abbcd_packed,d));
printf("Size of abbcd_packed = %d\n",sizeof(struct abbcd_packed));

}

-- 
Bartc 

[toc] | [prev] | [next] | [standalone]


#41442

Fromanish kumar <yesanishhere@gmail.com>
Date2014-03-07 14:34 -0800
Message-ID<029a1f3e-c6fe-47a5-b43b-f5050b6d94d0@googlegroups.com>
In reply to#41441
On Friday, March 7, 2014 1:59:11 PM UTC-8, Bart wrote:
> "anish singh" <anish198519851985@gmail.com> wrote in message 
> 
> news:43b13d99-38eb-4711-824d-5fc40d55edf3@googlegroups.com...
> 
> > Struct abbcd{
> 
> >   Char c;
> 
> >   Int b;
> 
> >   Short d;
> 
> > };
> 
> >
> 
> > What will be the size of abbcd ? If padding involved and without 
> 
> > padding?Suppose that the processor has only 4 byte registers.
> 
> >
> 
> > What will be the size if the particular processor has register for 2 byte, 
> 
> > 4 byte and 1 byte?
> 
> >
> 
> > Note; size of char is 1,size of int, short is 4 and 2 respectively.
> 
> 
> 
> Do you have access to a C compiler? Then you can do your own experiments 
> 
> with programs such as the following. (Note the #pragma line, to turn off 
> 
> padding for alignment, will vary between compilers.)
> 
> 
> 
> The offsets and padding will depend more on the memory alignments needed for 
> 
> the machine, then the sizes of the registers.
Are you sure that memory alignments have nothing to do with register size
or the address/data bus size?
> 
> 
> 
> #include <stdio.h>
> 
> #include <stddef.h>
> 
> 
> 
> int main(void) {
> 
> 
> 
> struct abbcd {
> 
>  char c;
> 
>  int b;
> 
>  short d;
> 
> };
> 
> 
> 
> #pragma pack(1)
> 
> struct abbcd_packed {
> 
>  char c;
> 
>  int b;
> 
>  short d;
> 
> };
> 
> 
> 
> printf("Size of char  = %d\n",sizeof(char));
> 
> printf("Size of int   = %d\n",sizeof(int));
> 
> printf("Size of short = %d\n",sizeof(short));
> 
> puts("");
> 
> 
> 
> puts("Normal padding:");
> 
> printf("Offset of c   = %d\n",offsetof(struct abbcd,c));
> 
> printf("Offset of b   = %d\n",offsetof(struct abbcd,b));
> 
> printf("Offset of d   = %d\n",offsetof(struct abbcd,d));
> 
> printf("Size of abbcd = %d\n",sizeof(struct abbcd));
> 
> puts("");
> 
> 
> 
> puts("Without padding:");
> 
> printf("Offset of c          = %d\n",offsetof(struct abbcd_packed,c));
> 
> printf("Offset of b          = %d\n",offsetof(struct abbcd_packed,b));
> 
> printf("Offset of d          = %d\n",offsetof(struct abbcd_packed,d));
> 
> printf("Size of abbcd_packed = %d\n",sizeof(struct abbcd_packed));
> 
> 
> 
> }
> 
> 
> 
> -- 
> 
> Bartc

[toc] | [prev] | [next] | [standalone]


#41451

From"BartC" <bc@freeuk.com>
Date2014-03-07 23:59 +0000
Message-ID<w4tSu.83912$8R3.14785@fx30.am4>
In reply to#41442
"anish kumar" <yesanishhere@gmail.com> wrote in message 
news:029a1f3e-c6fe-47a5-b43b-f5050b6d94d0@googlegroups.com...
> On Friday, March 7, 2014 1:59:11 PM UTC-8, Bart wrote:
>> "anish singh" <anish198519851985@gmail.com> wrote in message

>> The offsets and padding will depend more on the memory alignments needed 
>> for
>>
>> the machine, then the sizes of the registers.
> Are you sure that memory alignments have nothing to do with register size
> or the address/data bus size?

There will be a relationship between memory organisation and register width, 
but the memory layout will be more important.

In your example, there are three kinds of alignment, but there might only 
one width of register.

Or you can have the same register model, but another version of the 
processor might arrange the memory and data bus differently.

Have you been looking at real processors, or made-up ones?

In the example you gave of one-byte registers and 4-byte ints, then such a 
machine could have an 8-bit databus (so alignment is not important), but 
could also have a 16, 32 or 64-bit one, if the processor could use that to 
advantage.

-- 
Bartc 

[toc] | [prev] | [next] | [standalone]


#41446

FromKeith Thompson <kst-u@mib.org>
Date2014-03-07 15:15 -0800
Message-ID<lnvbvpvbvw.fsf@nuthaus.mib.org>
In reply to#41441
"BartC" <bc@freeuk.com> writes:
> "anish singh" <anish198519851985@gmail.com> wrote in message 
> news:43b13d99-38eb-4711-824d-5fc40d55edf3@googlegroups.com...
>> Struct abbcd{
>>   Char c;
>>   Int b;
>>   Short d;
>> };
>>
>> What will be the size of abbcd ? If padding involved and without 
>> padding?Suppose that the processor has only 4 byte registers.
>>
>> What will be the size if the particular processor has register for 2 byte, 
>> 4 byte and 1 byte?
>>
>> Note; size of char is 1,size of int, short is 4 and 2 respectively.

[...]
> #pragma pack(1)
[...]

The OP may not be aware that #pragma pack is non-standard.  It's an
extension implemented by gcc (and probably other C compilers).

-- 
Keith Thompson (The_Other_Keith) kst-u@mib.org  <http://www.ghoti.net/~kst>
Working, but not speaking, for JetHead Development, Inc.
"We must do something.  This is something.  Therefore, we must do this."
    -- Antony Jay and Jonathan Lynn, "Yes Minister"

[toc] | [prev] | [next] | [standalone]


#41459

From"BartC" <bc@freeuk.com>
Date2014-03-08 10:08 +0000
Message-ID<m%BSu.56526$NZ3.32303@fx33.am4>
In reply to#41446
"Keith Thompson" <kst-u@mib.org> wrote in message
news:lnvbvpvbvw.fsf@nuthaus.mib.org...
> "BartC" <bc@freeuk.com> writes:

> [...]
>> #pragma pack(1)
> [...]
>
> The OP may not be aware that #pragma pack is non-standard.  It's an
> extension implemented by gcc (and probably other C compilers).

I mentioned that in my post.

"Keith Thompson" <kst-u@mib.org> wrote in message
news:lnr46dv8ol.fsf@nuthaus.mib.org...

> One more nitpick: sizeof yields a result of type size_t; the "%d"
> format requires an argument of type int.  Use the "%zu" format, or
> convert the sizeof result to int (or to unsigned long and use "%lu").

But you didn't pick on the dozen or so non-uses of "%zu" in the rest of my
post.

However I dislike having to remember and use all these weird and wonderful
format specifiers (a 'zoo' of them almost!). If the right one is that
important, then the compiler should tell me about it (but only lccwin32
seems to do so at default warning levels).

Ideally it should figure it out for itself (a few years ago, I proposed a %?
specifier for that purpose, for use in the 99.9% of cases where the format
string was a constant). Because managing these format strings can be a lot
of work (you change one type from int to long long, then you have to change
hundreds of %d to %lld or %x to %llx).

(FWIW, not using %zu doesn't seem to matter on my machine; when I'm 
compiling for 64-bits and a size_t value occupies 8 bytes, while int is 4 
bytes, then presumably the parameter stack is also 64-bit aligned so use of 
%d seems to have no ill-effects. Tested with 3 x64 compilers.)

-- 
Bartc 

[toc] | [prev] | [next] | [standalone]


#41461

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2014-03-08 12:39 +0000
Message-ID<0.14a256d8294218f1eb5c.20140308123924GMT.87lhwk50g3.fsf@bsb.me.uk>
In reply to#41459
"BartC" <bc@freeuk.com> writes:
<snip>
> "Keith Thompson" <kst-u@mib.org> wrote in message
> news:lnr46dv8ol.fsf@nuthaus.mib.org...
>
>> One more nitpick: sizeof yields a result of type size_t; the "%d"
>> format requires an argument of type int.  Use the "%zu" format, or
>> convert the sizeof result to int (or to unsigned long and use "%lu").
>
> But you didn't pick on the dozen or so non-uses of "%zu" in the rest of my
> post.

You are commenting on a post to someone else.  That post had a single
use of %d.  What else was there to point out?

> However I dislike having to remember and use all these weird and wonderful
> format specifiers (a 'zoo' of them almost!). If the right one is that
> important, then the compiler should tell me about it (but only lccwin32
> seems to do so at default warning levels).

On my machine gcc does too, but in any case it's wise to choose the
warnings you care about.  With gcc, I ask for almost everything a turn
off the couple that I find annoying.

<snip>
-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#41491

FromKeith Thompson <kst-u@mib.org>
Date2014-03-08 14:03 -0800
Message-ID<lna9d0uz52.fsf@nuthaus.mib.org>
In reply to#41459
"BartC" <bc@freeuk.com> writes:
> "Keith Thompson" <kst-u@mib.org> wrote in message
> news:lnvbvpvbvw.fsf@nuthaus.mib.org...
>> "BartC" <bc@freeuk.com> writes:
>
>> [...]
>>> #pragma pack(1)
>> [...]
>>
>> The OP may not be aware that #pragma pack is non-standard.  It's an
>> extension implemented by gcc (and probably other C compilers).
>
> I mentioned that in my post.

Sorry I missed that.  But you wrote that it "will vary between
compilers".  Not all compilers necessarily have a way to specify
packing of structure members.

> "Keith Thompson" <kst-u@mib.org> wrote in message
> news:lnr46dv8ol.fsf@nuthaus.mib.org...
>
>> One more nitpick: sizeof yields a result of type size_t; the "%d"
>> format requires an argument of type int.  Use the "%zu" format, or
>> convert the sizeof result to int (or to unsigned long and use "%lu").
>
> But you didn't pick on the dozen or so non-uses of "%zu" in the rest of my
> post.

I was replying to someone else.  I don't point out every error in every
post.

> However I dislike having to remember and use all these weird and wonderful
> format specifiers (a 'zoo' of them almost!). If the right one is that
> important, then the compiler should tell me about it (but only lccwin32
> seems to do so at default warning levels).

gcc warns about about mismatches between format strings and arguments,
at least in many cases.  But warning about such mismatches in all cases
is not possible.  Format strings are interpreted at run time.  A format
string is commonly a string literal, but needn't be.  You just have to
develop the habit of using the right format yourself if you want to
avoid undefined behavior.

> Ideally it should figure it out for itself (a few years ago, I proposed a %?
> specifier for that purpose, for use in the 99.9% of cases where the format
> string was a constant). Because managing these format strings can be a lot
> of work (you change one type from int to long long, then you have to change
> hundreds of %d to %lld or %x to %llx).

You can always convert the argument to a known type.  For example, if u
is of some unsigned type, but you're not sure which one, you can do:

    printf("%llu\n", (unsigned long long)u);

Or if you happen to know that the value of u is fairly small (say,
because it's the size of a structure that you know is smaller than 32
kbytes), you can just convert to int:

    printf("%d\n", (int)sizeof whatever);

> (FWIW, not using %zu doesn't seem to matter on my machine; when I'm 
> compiling for 64-bits and a size_t value occupies 8 bytes, while int is 4 
> bytes, then presumably the parameter stack is also 64-bit aligned so use of 
> %d seems to have no ill-effects. Tested with 3 x64 compilers.)

I see the same behavior.  I wouldn't be surprised to see it fail on a
big-endian system.  (Actually I just tried it and it "worked"; I'm not
sure why.)

But by using the correct format, perhaps with a cast, I don't have to
worry about it; I know it will work.

-- 
Keith Thompson (The_Other_Keith) kst-u@mib.org  <http://www.ghoti.net/~kst>
Working, but not speaking, for JetHead Development, Inc.
"We must do something.  This is something.  Therefore, we must do this."
    -- Antony Jay and Jonathan Lynn, "Yes Minister"

[toc] | [prev] | [next] | [standalone]


#41477

FromMalcolm McLean <malcolm.mclean5@btinternet.com>
Date2014-03-08 09:01 -0800
Message-ID<52e34051-beb0-4fb2-be38-4f129a1866eb@googlegroups.com>
In reply to#41434
On Friday, March 7, 2014 9:13:02 PM UTC, anish singh wrote:
> Struct abbcd{
> 
>    Char c;
>    Int b;
>    Short d;
> 
> };
> 
> 
> 
> What will be the size of abbcd ? If padding involved and without padding?
> Suppose that the processor has only 4 byte registers.
> 
> 
> Note; size of char is 1,size of int, short is 4 and 2 respectively.
>
The compiler isn't allowed to alter the order of the members. So b must come after c in memory and d must come last. it also must place the first member
right at the top of the structure. So struct abbcd x; char *ptr = (char *)&x; 
must give you the address of c.

But it can insert other padding elements at will. Register size isn't a good
guide, because often processors allow half word access, even have special half
word registers, but make it less efficient than full-word access. 

[toc] | [prev] | [standalone]


Page 4 of 4 — ← Prev page 1 2 3 [4]

Back to top | Article view | comp.lang.c


csiph-web