Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming > #2486 > unrolled thread

little-endian

Started bybob <bob@coolfone.comze.com>
First post2012-11-15 13:45 -0800
Last post2012-11-28 23:45 -0800
Articles 20 on this page of 34 — 12 participants

Back to article view | Back to comp.programming


Contents

  little-endian bob <bob@coolfone.comze.com> - 2012-11-15 13:45 -0800
    Re: little-endian JJ <jaejunks@glegooilma-swapit.com> - 2012-11-15 23:33 +0000
    Re: little-endian bob <bob@coolfone.comze.com> - 2012-11-16 10:38 -0800
      Re: little-endian "Pascal J. Bourguignon" <pjb@informatimago.com> - 2012-11-16 23:32 +0100
        Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-16 21:17 -0600
          Re: little-endian "Pascal J. Bourguignon" <pjb@informatimago.com> - 2012-11-17 10:26 +0100
            Re: little-endian "BartC" <bc@freeuk.com> - 2012-11-17 11:38 +0000
              Re: little-endian "Pascal J. Bourguignon" <pjb@informatimago.com> - 2012-11-17 13:56 +0100
                Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-17 11:58 -0600
              Re: little-endian Ben Bacarisse <ben.usenet@bsb.me.uk> - 2012-11-17 19:35 +0000
                Re: little-endian "BartC" <bc@freeuk.com> - 2012-11-17 19:50 +0000
            Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-17 11:30 -0600
              Re: little-endian Ian Collins <ian-news@hotmail.com> - 2012-11-18 09:28 +1300
                Re: little-endian Robert Wessel <robertwessel2@yahoo.com> - 2012-11-17 14:48 -0600
                  Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-17 16:22 -0600
                    Re: little-endian Ian Collins <ian-news@hotmail.com> - 2012-11-18 15:17 +1300
                      Re: little-endian "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-11-18 09:13 +0100
                        Re: little-endian Ian Collins <ian-news@hotmail.com> - 2012-11-18 21:33 +1300
                          Re: little-endian "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-11-18 10:15 +0100
                        Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-18 03:03 -0600
                          Re: little-endian "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-11-18 10:29 +0100
                            Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-18 12:50 -0600
    Re: little-endian Robin Vowels <robin.vowels@gmail.com> - 2012-11-22 04:49 -0800
      Re: little-endian Jongware <jongware@no-spam.plz> - 2012-11-23 15:22 +0100
        Re: little-endian pacman@kosh.dhis.org (Alan Curry) - 2012-11-25 19:39 +0000
          Re: little-endian Jongware <jongware@no-spam.plz> - 2012-11-26 12:01 +0100
            Re: little-endian "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-11-26 14:02 +0100
              Re: little-endian Jongware <jongware@no-spam.plz> - 2012-11-26 15:34 +0100
                Re: little-endian "Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de> - 2012-11-26 17:03 +0100
                Re: little-endian Robert Wessel <robertwessel2@yahoo.com> - 2012-11-26 17:44 -0600
                  Re: little-endian Robin Vowels <robin.vowels@gmail.com> - 2012-11-28 23:43 -0800
                    Re: little-endian Robert Wessel <robertwessel2@yahoo.com> - 2012-11-29 22:46 -0600
            Re: little-endian BGB <cr88192@hotmail.com> - 2012-11-26 08:30 -0600
          Re: little-endian Robin Vowels <robin.vowels@gmail.com> - 2012-11-28 23:45 -0800

Page 1 of 2  [1] 2  Next page →


#2486 — little-endian

Frombob <bob@coolfone.comze.com>
Date2012-11-15 13:45 -0800
Subjectlittle-endian
Message-ID<2f491c75-8fbc-4dca-9abe-f11e784454a2@googlegroups.com>
Am I the only one who constantly gets confused that little-endian stores the big stuff at the end?

Seems like a misnomer.

[toc] | [next] | [standalone]


#2487

FromJJ <jaejunks@glegooilma-swapit.com>
Date2012-11-15 23:33 +0000
Message-ID<XnsA10D434801618jaejunksglegooilma@0.0.0.9>
In reply to#2486
bob <bob@coolfone.comze.com> wrote:
> Am I the only one who constantly gets confused that little-endian stores 
the big stuff at the end?
> 
> Seems like a misnomer.

It's due to processor architectural difference in comparison with big-
endian processors.

I'm more confused how your message got end up in base64 encoding.

[toc] | [prev] | [next] | [standalone]


#2488

Frombob <bob@coolfone.comze.com>
Date2012-11-16 10:38 -0800
Message-ID<f9e47587-00b6-4292-b12c-bc5b62b9d54c@googlegroups.com>
In reply to#2486
On Thursday, November 15, 2012 4:06:17 PM UTC-6, robert...@yahoo.com wrote:
> On Thu, 15 Nov 2012 13:45:09 -0800 (PST), bob <bob@coolfone.comze.com>
> 
> wrote:
> 
> 
> 
> >Am I the only one who constantly gets confused that little-endian stores the big stuff at the end?
> 
> >
> 
> >Seems like a misnomer.
> 
> 
> 
> 
> 
> I've posted this before but:
> 
> 
> 
> The terms actually come from Jonathan Swift's, "Gulliver�s Travels"
> 
> where there were two rival kingdoms, one where they ate eggs starting
> 
> at the little end (aka the "little endians") and the other in which it
> 
> was correct to eat the egg from the big end. So it's really a
> 
> statement of which end you start the number (or egg) from.

I don't see the explicit mention of 'Little-endian' in the book, but I guess it's implied:

It is computed that eleven thousand persons have at several times suffered death, rather than submit to break their eggs at the smaller end.   Many hundred large volumes have been published upon this controversy: but the books of the Big-endians have been long forbidden, and the whole party rendered incapable by law of holding employments.

Swift, Jonathan (2012-05-12). Gulliver's Travels (Timeless Classics) (Kindle Locations 530-532). Saddleback Educational Publishing. Kindle Edition.

[toc] | [prev] | [next] | [standalone]


#2489

From"Pascal J. Bourguignon" <pjb@informatimago.com>
Date2012-11-16 23:32 +0100
Message-ID<87r4ntp4b1.fsf@informatimago.com>
In reply to#2488
bob <bob@coolfone.comze.com> writes:

> On Thursday, November 15, 2012 4:06:17 PM UTC-6, robert...@yahoo.com wrote:
>> The terms actually come from Jonathan Swift's, "Gulliver�s Travels"
>> 
>> where there were two rival kingdoms, one where they ate eggs starting
>> 
>> at the little end (aka the "little endians") and the other in which it
>> 
>> was correct to eat the egg from the big end. So it's really a
>> 
>> statement of which end you start the number (or egg) from.
>
> I don't see the explicit mention of 'Little-endian' in the book, but I guess it's implied:
>
> It is computed that eleven thousand persons have at several times
> suffered death, rather than submit to break their eggs at the smaller
> end.   Many hundred large volumes have been published upon this
> controversy: but the books of the Big-endians have been long
> forbidden, and the whole party rendered incapable by law of holding
> employments.
>
> Swift, Jonathan (2012-05-12). Gulliver's Travels (Timeless Classics)
> (Kindle Locations 530-532). Saddleback Educational Publishing. Kindle
> Edition.

I'm a little endian egg eater, but a big endian programmer.  I prefer
my numbers stored in memory big endian first :-)


-- 
__Pascal Bourguignon__
http://www.informatimago.com

[toc] | [prev] | [next] | [standalone]


#2490

FromBGB <cr88192@hotmail.com>
Date2012-11-16 21:17 -0600
Message-ID<k86vq4$hl0$1@news.albasani.net>
In reply to#2489
On 11/16/2012 4:32 PM, Pascal J. Bourguignon wrote:
> bob <bob@coolfone.comze.com> writes:
>
>> On Thursday, November 15, 2012 4:06:17 PM UTC-6, robert...@yahoo.com wrote:
>>> The terms actually come from Jonathan Swift's, "Gulliver�s Travels"
>>>
>>> where there were two rival kingdoms, one where they ate eggs starting
>>>
>>> at the little end (aka the "little endians") and the other in which it
>>>
>>> was correct to eat the egg from the big end. So it's really a
>>>
>>> statement of which end you start the number (or egg) from.
>>
>> I don't see the explicit mention of 'Little-endian' in the book, but I guess it's implied:
>>
>> It is computed that eleven thousand persons have at several times
>> suffered death, rather than submit to break their eggs at the smaller
>> end.   Many hundred large volumes have been published upon this
>> controversy: but the books of the Big-endians have been long
>> forbidden, and the whole party rendered incapable by law of holding
>> employments.
>>
>> Swift, Jonathan (2012-05-12). Gulliver's Travels (Timeless Classics)
>> (Kindle Locations 530-532). Saddleback Educational Publishing. Kindle
>> Edition.
>
> I'm a little endian egg eater, but a big endian programmer.  I prefer
> my numbers stored in memory big endian first :-)
>

some of my file-formats end up little-endian, and others big-endian.

then (when dealing with bitstream formats) there is also the whole 
matter of bit-ordering, which may or may not match the byte-ordering.

personally, I prefer little endian, but don't really think it is a big 
deal either.


some of it depends on what the format is based-on, for example, I have a 
network protocol which is more directly based on the Deflate bitstream, 
and so uses little-little ordering. basically, a specialized 
Deflate-like bitstream with the ability to directly encode S-Expression 
data and limited ability to build a context model from these expressions 
(to save bits by avoiding retransmitting repeating values), as well as 
tweaks to make if more efficient for a continuous stream of small 
messages (the design of deflate not being optimally suited for a stream 
of small messages).


I have a graphics format design which is roughly based on JPEG, and 
would use big-big ordering, despite many parts of the bitstream design 
being reused from the network protocol. it differs mostly from JPEG in 
that it would use a lot of space-saving "fine-tuning" adjustments (more 
compact tables, tweaks to the VLC coding, support for 64x64 pixel 
"megablocks", block-based motion compensation, ...).

as-is, I am already using a JPEG-based graphics format (which adds 
layers, an alpha-channel, and lossless coding, and is "mostly backwards 
compatible"), but the difference would be that the new format would not 
be backwards compatible with JPEG, but could potentially compress a 
little better. as-is, the current format for lossless coding is about 
4%-8% worse than HD-Photo / JPEG-XR, IOW: XR would give 89kB and mine 
94kB, for a 512x512 image of a mountain, vs about 270kB for a PNG 
version (encoded with MS's PNG encoder).

it is mostly just debatable if this would be worthwhile.


or such...

[toc] | [prev] | [next] | [standalone]


#2491

From"Pascal J. Bourguignon" <pjb@informatimago.com>
Date2012-11-17 10:26 +0100
Message-ID<87mwygpolg.fsf@informatimago.com>
In reply to#2490
BGB <cr88192@hotmail.com> writes:

> some of my file-formats end up little-endian, and others big-endian.

That's unfortunate.  The internet is in general specified to use
big-endian.  For files you'd be well advised to apply the same rule.

Always use htonl to write to files/network and ntohl to read from
files/network.


> then (when dealing with bitstream formats) there is also the whole
> matter of bit-ordering, which may or may not match the byte-ordering.

Granted, but then bits are usually gathered into bytes by the hardware,
so the issue doesn't concerns us programmers, in general.  That said,
some hardware protocols transmit bytes in little endian bit order, while
others transmit them in big endian bit order too :-)


> personally, I prefer little endian, but don't really think it is a big
> deal either.
> […]
> it is mostly just debatable if this would be worthwhile.

Or if the debate is worthwhile :-)

-- 
__Pascal Bourguignon__
http://www.informatimago.com

[toc] | [prev] | [next] | [standalone]


#2492

From"BartC" <bc@freeuk.com>
Date2012-11-17 11:38 +0000
Message-ID<k87t1r$skv$1@dont-email.me>
In reply to#2491

"Pascal J. Bourguignon" <pjb@informatimago.com> wrote in message 
news:87mwygpolg.fsf@informatimago.com...
> BGB <cr88192@hotmail.com> writes:
>
>> some of my file-formats end up little-endian, and others big-endian.
>
> That's unfortunate.  The internet is in general specified to use
> big-endian.  For files you'd be well advised to apply the same rule.

Or maybe it shouldn't be anything to do with the internet.

Does it also specify whether RGB triples should have R first, or B first? It 
should be a concern of the file format.

Or files can be sent (highly inefficiently) in true text format (not binary 
disguised as text), which is invariably big-endian. Then the only issue is 
which character sequence represents end-of-line...

-- 
Bartc 

[toc] | [prev] | [next] | [standalone]


#2493

From"Pascal J. Bourguignon" <pjb@informatimago.com>
Date2012-11-17 13:56 +0100
Message-ID<87ehjspevu.fsf@informatimago.com>
In reply to#2492
"BartC" <bc@freeuk.com> writes:

> "Pascal J. Bourguignon" <pjb@informatimago.com> wrote in message
> news:87mwygpolg.fsf@informatimago.com...
>> BGB <cr88192@hotmail.com> writes:
>>
>>> some of my file-formats end up little-endian, and others big-endian.
>>
>> That's unfortunate.  The internet is in general specified to use
>> big-endian.  For files you'd be well advised to apply the same rule.
>
> Or maybe it shouldn't be anything to do with the internet.
>
> Does it also specify whether RGB triples should have R first, or B
> first? It should be a concern of the file format.

Definitely.  But it's not the bytesex anymore, it's the colorsex. :-)



> Or files can be sent (highly inefficiently) in true text format (not
> binary disguised as text), which is invariably big-endian. Then the
> only issue is which character sequence represents end-of-line...

Just choose one of the end-of-line characters in the unicode character
set :-)


-- 
__Pascal Bourguignon__
http://www.informatimago.com

[toc] | [prev] | [next] | [standalone]


#2495

FromBGB <cr88192@hotmail.com>
Date2012-11-17 11:58 -0600
Message-ID<k88je9$jos$1@news.albasani.net>
In reply to#2493
On 11/17/2012 6:56 AM, Pascal J. Bourguignon wrote:
> "BartC" <bc@freeuk.com> writes:
>
>> "Pascal J. Bourguignon" <pjb@informatimago.com> wrote in message
>> news:87mwygpolg.fsf@informatimago.com...
>>> BGB <cr88192@hotmail.com> writes:
>>>
>>>> some of my file-formats end up little-endian, and others big-endian.
>>>
>>> That's unfortunate.  The internet is in general specified to use
>>> big-endian.  For files you'd be well advised to apply the same rule.
>>
>> Or maybe it shouldn't be anything to do with the internet.
>>
>> Does it also specify whether RGB triples should have R first, or B
>> first? It should be a concern of the file format.
>
> Definitely.  But it's not the bytesex anymore, it's the colorsex. :-)
>

usually my preference is for RGBA for color bytes.


the downside here is that if loaded as a 32-bit integer on a LE-target 
and displayed in hex, it will be seen in the order:
AABBGGRR

then there is the whole thing that the graphics hardware typically 
internally uses BGRA ordering.

displayed in hex, BGRA looks like:
AARRGGBB

this is more related to efficiency when streaming texture images to/from 
the graphics hardware, as avoiding a flip can make getting/setting the 
images go faster (this can be relevant, say, when streaming video files 
onto textures in a 3D engine).


for related reasons, I tend to nearly always treat FOURCC values as 
big-endian, regardless of whether the format is BE or LE (this means 
that in-memory, the FOURCC values are often transposed, but at least 
they will appear in the correct order if printed as hex).


>
>> Or files can be sent (highly inefficiently) in true text format (not
>> binary disguised as text), which is invariably big-endian. Then the
>> only issue is which character sequence represents end-of-line...
>
> Just choose one of the end-of-line characters in the unicode character
> set :-)
>

I think I read before that CR+LF is the defined network line-ending.

typically it is easier just to make both LF and CR+LF valid line-endings 
when reading files.

[toc] | [prev] | [next] | [standalone]


#2496

FromBen Bacarisse <ben.usenet@bsb.me.uk>
Date2012-11-17 19:35 +0000
Message-ID<0.abf536d710deb4d3bf3e.20121117193507GMT.87k3tkc9as.fsf@bsb.me.uk>
In reply to#2492
"BartC" <bc@freeuk.com> writes:
<snip>
> Or files can be sent (highly inefficiently) in true text format (not
> binary disguised as text), which is invariably big-endian.

Why invariably?  Someone used to Arabic (for example) might well output
a number starting with the least significant digit first.

-- 
Ben.

[toc] | [prev] | [next] | [standalone]


#2497

From"BartC" <bc@freeuk.com>
Date2012-11-17 19:50 +0000
Message-ID<k88psp$4rm$1@dont-email.me>
In reply to#2496

"Ben Bacarisse" <ben.usenet@bsb.me.uk> wrote in message 
news:0.abf536d710deb4d3bf3e.20121117193507GMT.87k3tkc9as.fsf@bsb.me.uk...
> "BartC" <bc@freeuk.com> writes:
> <snip>
>> Or files can be sent (highly inefficiently) in true text format (not
>> binary disguised as text), which is invariably big-endian.
>
> Why invariably?  Someone used to Arabic (for example) might well output
> a number starting with the least significant digit first.

You mean text that might in a Western alphabet look like:

 ABC 12345 DEF

could appear (making no attempt write actual Arabic here...) as:

 FED 12345 BCA

(since all such examples I've seen don't seem to reverse their numbers). 
You're saying that that might be stored in the text itself, if it starts at 
"A", as:

ABC 54321 DEF ?

If so then perhaps you have a point. But there might then be bigger problems 
to deal with than the endianness of numbers...


-- 
Bartc 

[toc] | [prev] | [next] | [standalone]


#2494

FromBGB <cr88192@hotmail.com>
Date2012-11-17 11:30 -0600
Message-ID<k88hoj$fvl$1@news.albasani.net>
In reply to#2491
On 11/17/2012 3:26 AM, Pascal J. Bourguignon wrote:
> BGB <cr88192@hotmail.com> writes:
>
>> some of my file-formats end up little-endian, and others big-endian.
>
> That's unfortunate.  The internet is in general specified to use
> big-endian.  For files you'd be well advised to apply the same rule.
>
> Always use htonl to write to files/network and ntohl to read from
> files/network.
>

granted though, many file formats (in general) are designed in a world 
where the dominant processor architectures (namely x86 and ARM) use 
little-endian.

a lot often comes down to whether the original designer felt like big or 
little endian, or if the format it is based on uses big or little, where 
mostly it is just good to be consistent (as formats which are partly in 
big and partly in little are a little ugly...).



typically, a person may develop and use functions like:
ReadInt32BE
and WriteInt64LE and similar, and use these for most of the encoding.


also, a big downside of using htonl and ntohl as a general mechanism is 
that it makes all of ones' libraries need to depend on winsock.

also, "word flipping" is sort of an ugly trick IMO, better to just 
read/write values with an explicit endianess.



>
>> then (when dealing with bitstream formats) there is also the whole
>> matter of bit-ordering, which may or may not match the byte-ordering.
>
> Granted, but then bits are usually gathered into bytes by the hardware,
> so the issue doesn't concerns us programmers, in general.  That said,
> some hardware protocols transmit bytes in little endian bit order, while
> others transmit them in big endian bit order too :-)
>

a person does have to deal-with / define bit-ordering, if they are 
dealing with something like a Huffman encoder/decoder.

it doesn't necessarily have anything to do with the hardware bit 
ordering, but more, how bits are packed into the bytes for multi-bit values.

example:
start with high-bit, and emit high-bits first (big-big);
start with low-bit, and emit low-bits first (little-little).

other schemes are possible, but are a pain to encode/decode efficiently 
(usually they result from codecs which read/write 1 bit at a time), 
though often a transpose-table is a workable strategy.

conceptually, deflate transposes the Huffman codes (writing the MSB in 
LSB position), but this is partly because Huffman tends to require MSB 
first in-order to decode, and typically people don't really care about 
the numerical value of a Huffman code anyways.


>
>> personally, I prefer little endian, but don't really think it is a big
>> deal either.
>> […]
>> it is mostly just debatable if this would be worthwhile.
>
> Or if the debate is worthwhile :-)
>

could be.

I did experimentally throw together support for 64x64 blocks in my image 
codec, but thus far the results aren't really looking all that 
impressive (except one form which did really good on my test-patterns, 
but produces larger output for more normal images).

[toc] | [prev] | [next] | [standalone]


#2498

FromIan Collins <ian-news@hotmail.com>
Date2012-11-18 09:28 +1300
Message-ID<agqab0F2uveU1@mid.individual.net>
In reply to#2494
On 11/18/12 06:30, BGB wrote:
>
> typically, a person may develop and use functions like:
> ReadInt32BE
> and WriteInt64LE and similar, and use these for most of the encoding.

Typically a programmer will use the utilities their platform provides.

> also, a big downside of using htonl and ntohl as a general mechanism is
> that it makes all of ones' libraries need to depend on winsock.

That's silly on at least three counts:

1) the whole would doesn't revolve around windows.

2) those utility functions are usually implemented in-line (typically as 
macros), so there isn't a library dependency.

3) if you are sending data over IP, you need the socket libraries one 
way or another!

-- 
Ian Collins

[toc] | [prev] | [next] | [standalone]


#2499

FromRobert Wessel <robertwessel2@yahoo.com>
Date2012-11-17 14:48 -0600
Message-ID<7otfa8pfb07sn1ao039nlfc9cr351h2vtr@4ax.com>
In reply to#2498
On Sun, 18 Nov 2012 09:28:16 +1300, Ian Collins <ian-news@hotmail.com>
wrote:

>On 11/18/12 06:30, BGB wrote:
>>
>> typically, a person may develop and use functions like:
>> ReadInt32BE
>> and WriteInt64LE and similar, and use these for most of the encoding.
>
>Typically a programmer will use the utilities their platform provides.
>
>> also, a big downside of using htonl and ntohl as a general mechanism is
>> that it makes all of ones' libraries need to depend on winsock.
>
>That's silly on at least three counts:
>
>1) the whole would doesn't revolve around windows.
>
>2) those utility functions are usually implemented in-line (typically as 
>macros), so there isn't a library dependency.
>
>3) if you are sending data over IP, you need the socket libraries one 
>way or another!


Some more valid criticisms of the hton*/ntoh* functions are that they
only support one byte order on the "network" side, support only a
couple of sizes, and only support unsigned values.

[toc] | [prev] | [next] | [standalone]


#2500

FromBGB <cr88192@hotmail.com>
Date2012-11-17 16:22 -0600
Message-ID<k892t5$p77$1@news.albasani.net>
In reply to#2499
On 11/17/2012 2:48 PM, Robert Wessel wrote:
> On Sun, 18 Nov 2012 09:28:16 +1300, Ian Collins <ian-news@hotmail.com>
> wrote:
>
>> On 11/18/12 06:30, BGB wrote:
>>>
>>> typically, a person may develop and use functions like:
>>> ReadInt32BE
>>> and WriteInt64LE and similar, and use these for most of the encoding.
>>
>> Typically a programmer will use the utilities their platform provides.
>>

typically a programmer will do whatever is most convenient at the moment.

in cases where a person isn't directly using sockets or similar, it is 
usually plenty sufficient just to write a few functions to read/write 
values directly in the intended format (typically using the 
shift-trick). (this makes sense both for the common cases of 
implementing a reader/writer directly to a byte-array, or via 
fgetc/fputc calls or similar).


>>> also, a big downside of using htonl and ntohl as a general mechanism is
>>> that it makes all of ones' libraries need to depend on winsock.
>>
>> That's silly on at least three counts:
>>
>> 1) the whole would doesn't revolve around windows.
>>
>> 2) those utility functions are usually implemented in-line (typically as
>> macros), so there isn't a library dependency.
>>
>> 3) if you are sending data over IP, you need the socket libraries one
>> way or another!
>

you still have to include "winsock.h" or "sys/socket.h" (on Linux), and 
all the relevant #ifdef's, ..., or similar, which means that they don't 
really make sense except for code dealing with sockets.

for most "general" file reader/writer code, it doesn't make sense to 
have such dependencies.


>
> Some more valid criticisms of the hton*/ntoh* functions are that they
> only support one byte order on the "network" side, support only a
> couple of sizes, and only support unsigned values.
>

and they only do direct value -> value mappings...


far more common I think is to implement file-readers more like:
int MyFile_ReadInt32LE(FILE *fd)
{
	int i;
	i=fgetc(fd); i+=fgetc(fd)<<8;
	i+=fgetc(fd)<<16; i+=fgetc(fd)<<24;
	return(i);
}

or, if using byte arrays:
byte *MyFile_ReadInt32LE(byte *cs, int *ri)
{
	int i;
	i=*cs++; i+=(*cs++)<<8;
	i+=(*cs++)<<16; i+=(*cs++)<<24;
	*ri=i;
	return(cs);
}

where somewhere someone has a typedef like:
typedef unsigned char byte;


a lot of this can then be reused as-needed.


or similar...

[toc] | [prev] | [next] | [standalone]


#2501

FromIan Collins <ian-news@hotmail.com>
Date2012-11-18 15:17 +1300
Message-ID<agquqgF2uveU2@mid.individual.net>
In reply to#2500
On 11/18/12 11:22, BGB wrote:
> On 11/17/2012 2:48 PM, Robert Wessel wrote:
>> On Sun, 18 Nov 2012 09:28:16 +1300, Ian Collins<ian-news@hotmail.com>
>> wrote:
>>
>>> On 11/18/12 06:30, BGB wrote:
>>>>
>>>> typically, a person may develop and use functions like:
>>>> ReadInt32BE
>>>> and WriteInt64LE and similar, and use these for most of the encoding.
>>>
>>> Typically a programmer will use the utilities their platform provides.
>>>
>
> typically a programmer will do whatever is most convenient at the moment.

Which tends not to involve reinventing wheels!

<snip>

>>>> also, a big downside of using htonl and ntohl as a general mechanism is
>>>> that it makes all of ones' libraries need to depend on winsock.
>>>
>>> That's silly on at least three counts:
>>>
>>> 1) the whole would doesn't revolve around windows.
>>>
>>> 2) those utility functions are usually implemented in-line (typically as
>>> macros), so there isn't a library dependency.
>>>
>>> 3) if you are sending data over IP, you need the socket libraries one
>>> way or another!
>>
>
> you still have to include "winsock.h" or "sys/socket.h" (on Linux), and
> all the relevant #ifdef's, ..., or similar, which means that they don't
> really make sense except for code dealing with sockets.

The headers take care of the relevant defines.

> for most "general" file reader/writer code, it doesn't make sense to
> have such dependencies.

Unless you want to avoid wheel reinventing...

>> Some more valid criticisms of the hton*/ntoh* functions are that they
>> only support one byte order on the "network" side, support only a
>> couple of sizes, and only support unsigned values.

long long version are a simple extension.  The signed or unsigned nature 
of the type isn't relevant for byte ordering.

> and they only do direct value ->  value mappings...

Which is their purpose.

> far more common I think is to implement file-readers more like:
> int MyFile_ReadInt32LE(FILE *fd)

The reason for using the common byte ordering functions is to avoid 
having to care about the ordering on the wire (or in a file).

-- 
Ian Collins

[toc] | [prev] | [next] | [standalone]


#2502

From"Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de>
Date2012-11-18 09:13 +0100
Message-ID<1d8xuw93lwnxy.14j0ojrkbbe4e$.dlg@40tude.net>
In reply to#2501
On Sun, 18 Nov 2012 15:17:51 +1300, Ian Collins wrote:

> On 11/18/12 11:22, BGB wrote:
>>> Some more valid criticisms of the hton*/ntoh* functions are that they
>>> only support one byte order on the "network" side, support only a
>>> couple of sizes, and only support unsigned values.
> 
> long long version are a simple extension.

Nope. It could be middle-endian.

>The signed or unsigned nature 
> of the type isn't relevant for byte ordering.

How so? Byte ordering is about encoding things into a byte stream. Signed
integers must be encoded too.

>> far more common I think is to implement file-readers more like:
>> int MyFile_ReadInt32LE(FILE *fd)
> 
> The reason for using the common byte ordering functions is to avoid 
> having to care about the ordering on the wire (or in a file).

There are far more types of objects than unsigned integers. When
implementing an application protocol on top of some octet or bit stream
these must be encoded and decoded (serialized/deserialozed) too. There is
nothing special in unsigned integers.

Usage of functions like hton* should be depreciated unless the protocol
specification explicitly states that the given unsigned integer object is
in the "network" format as implemented by hton*.

It is always cleaner to provide a fair implementation of the protocol,
which could be done in a portable way as BGB described. Which is really the
recommended way. A rare exception might be when you have a library that
already implements the protocol layer of interest completely.

-- 
Regards,
Dmitry A. Kazakov
http://www.dmitry-kazakov.de

[toc] | [prev] | [next] | [standalone]


#2503

FromIan Collins <ian-news@hotmail.com>
Date2012-11-18 21:33 +1300
Message-ID<agrkq0F2uveU3@mid.individual.net>
In reply to#2502
On 11/18/12 21:13, Dmitry A. Kazakov wrote:
> On Sun, 18 Nov 2012 15:17:51 +1300, Ian Collins wrote:
>
>> On 11/18/12 11:22, BGB wrote:
>>>> Some more valid criticisms of the hton*/ntoh* functions are that they
>>>> only support one byte order on the "network" side, support only a
>>>> couple of sizes, and only support unsigned values.
>>
>> long long version are a simple extension.
>
> Nope.

Well the machine I'm typing this on has ntohll and htonll.

> It could be middle-endian.

On the wire?

>> The signed or unsigned nature
>> of the type isn't relevant for byte ordering.
>
> How so? Byte ordering is about encoding things into a byte stream. Signed
> integers must be encoded too.

Bytes are bytes, whether they represent a signed or unsigned (or even 
floating point) type is irrelevant.  If it were, there would be a bigger 
common set of byte order reversal functions.

>>> far more common I think is to implement file-readers more like:
>>> int MyFile_ReadInt32LE(FILE *fd)
>>
>> The reason for using the common byte ordering functions is to avoid
>> having to care about the ordering on the wire (or in a file).
>
> There are far more types of objects than unsigned integers. When
> implementing an application protocol on top of some octet or bit stream
> these must be encoded and decoded (serialized/deserialozed) too. There is
> nothing special in unsigned integers.

Did I say there was?

> Usage of functions like hton* should be depreciated unless the protocol
> specification explicitly states that the given unsigned integer object is
> in the "network" format as implemented by hton*.

Which oddly enough, they often are.  The whole point of "network byte 
ordering" is to provide a common standard representation.

> It is always cleaner to provide a fair implementation of the protocol,
> which could be done in a portable way as BGB described. Which is really the
> recommended way. A rare exception might be when you have a library that
> already implements the protocol layer of interest completely.

It's good to know I've been doing things wrong all these years....

-- 
Ian Collins

[toc] | [prev] | [next] | [standalone]


#2505

From"Dmitry A. Kazakov" <mailbox@dmitry-kazakov.de>
Date2012-11-18 10:15 +0100
Message-ID<1elrck7ulw2z3$.bdt6xmotgmne.dlg@40tude.net>
In reply to#2503
On Sun, 18 Nov 2012 21:33:03 +1300, Ian Collins wrote:

> On 11/18/12 21:13, Dmitry A. Kazakov wrote:
>> On Sun, 18 Nov 2012 15:17:51 +1300, Ian Collins wrote:
>>
>>> On 11/18/12 11:22, BGB wrote:
>>>>> Some more valid criticisms of the hton*/ntoh* functions are that they
>>>>> only support one byte order on the "network" side, support only a
>>>>> couple of sizes, and only support unsigned values.
>>>
>>> long long version are a simple extension.
>>
>> Nope.
> 
> Well the machine I'm typing this on has ntohll and htonll.

It is not "a simple extension," provided you meant implementation of DWORD
I/O in terms of a WORD stream. It could again be little or big-endian. You
have all endianness issues for each type of stream element. There is
nothing special in specifically octets, except that for some protocols are
defined as octet streams. Other protocols are bit streams. Some allow you
to define mapping more or less freely (e.g. EtherCAT's FMMUs).

>> It could be middle-endian.
> 
> On the wire?

Yep. Not even contiguous, when you read it out. E.g. some ModBus terminals
have integers stored in a quite arbitrary order of words (they are
word-oriented) with some bit-patterns reserved (ugly mess, in short).

>>> The signed or unsigned nature
>>> of the type isn't relevant for byte ordering.
>>
>> How so? Byte ordering is about encoding things into a byte stream. Signed
>> integers must be encoded too.
> 
> Bytes are bytes, whether they represent a signed or unsigned (or even 
> floating point) type is irrelevant.

As well as their order is. If you are talking about the transport level,
e.g. byte stream, then the order is fixed 1st byte, 2nd byte etc.

If you are talking about an encoding of some entities like integer, float,
employee ID as ordered sequences of bytes, then we return to my point. Any
such encoding must be specified, and big/little-endian is such a
specification for EXCLUSIVELY unsigned integers with the ranges of power of
two. It is meaningless for other types. There is no byte order of a signed
integer. There could be byte order an unsigned integer [used to encode a
signed integer or whatever].

>> Usage of functions like hton* should be depreciated unless the protocol
>> specification explicitly states that the given unsigned integer object is
>> in the "network" format as implemented by hton*.
> 
> Which oddly enough, they often are.  The whole point of "network byte 
> ordering" is to provide a common standard representation.

Standard representation of what?

>> It is always cleaner to provide a fair implementation of the protocol,
>> which could be done in a portable way as BGB described. Which is really the
>> recommended way. A rare exception might be when you have a library that
>> already implements the protocol layer of interest completely.
> 
> It's good to know I've been doing things wrong all these years....

It is never too late... (:-))

-- 
Regards,
Dmitry A. Kazakov
http://www.dmitry-kazakov.de

[toc] | [prev] | [next] | [standalone]


#2504

FromBGB <cr88192@hotmail.com>
Date2012-11-18 03:03 -0600
Message-ID<k8a8fd$epd$1@news.albasani.net>
In reply to#2502
On 11/18/2012 2:13 AM, Dmitry A. Kazakov wrote:
> On Sun, 18 Nov 2012 15:17:51 +1300, Ian Collins wrote:
>
>> On 11/18/12 11:22, BGB wrote:
>>>> Some more valid criticisms of the hton*/ntoh* functions are that they
>>>> only support one byte order on the "network" side, support only a
>>>> couple of sizes, and only support unsigned values.
>>
>> long long version are a simple extension.
>
> Nope. It could be middle-endian.
>
>> The signed or unsigned nature
>> of the type isn't relevant for byte ordering.
>
> How so? Byte ordering is about encoding things into a byte stream. Signed
> integers must be encoded too.
>

yeah. signed vs unsigned storage is an issue all to itself.

>>> far more common I think is to implement file-readers more like:
>>> int MyFile_ReadInt32LE(FILE *fd)
>>
>> The reason for using the common byte ordering functions is to avoid
>> having to care about the ordering on the wire (or in a file).
>
> There are far more types of objects than unsigned integers. When
> implementing an application protocol on top of some octet or bit stream
> these must be encoded and decoded (serialized/deserialozed) too. There is
> nothing special in unsigned integers.
>
> Usage of functions like hton* should be depreciated unless the protocol
> specification explicitly states that the given unsigned integer object is
> in the "network" format as implemented by hton*.
>

yeah.

hton* and ntoh* make sense when dealing with socket-related stuff, but 
don't make nearly as much sense when dealing with file-formats.


typically, the layout of data within a file-format is a fairly important 
part of the file-format, and not really all that consistent between one 
file format and the next, and sometimes you are lucky even if all of the 
members within the same file-format have the same ordering.

worse-still is fileformats where the designer thought about being clever 
and defining the endianess as "whatever was convenient for the writer 
program", which basically amounts to the reader having to keep track of 
a flag for which endianess the file was stored in.


> It is always cleaner to provide a fair implementation of the protocol,
> which could be done in a portable way as BGB described. Which is really the
> recommended way. A rare exception might be when you have a library that
> already implements the protocol layer of interest completely.
>

yeah.

or when dealing with file-formats.

with explicit reader/writer functions, it is also fairly straightforward 
to write functions to deal with pretty much any types the format may 
include, including things like variable-length integers, and values of 
types which don't necessarily start/end on a byte-boundary (common with 
bitstreams).

typically, any structures used are kept purely internal, and the exact 
layouts and types of data present in a file are not tied to the structs 
(and how the particular compiler and target has decided that they should 
be laid out), and it isn't really that much more effort to read/write 
all of the members directly via function calls, than it is to read/write 
structs and swizzle the values afterwards.


other times (depending on how the format is used), it may make sense to 
use structs, but simply define all of the members as byte-arrays, and 
then use functions to read/write the values from the structures (this 
makes things like type-alignment largely a non-issue).

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | comp.programming


csiph-web