Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch.embedded > #31788 > unrolled thread

Embedding a Checksum in an Image File

Started byRick C <gnuarm.deletethisbit@gmail.com>
First post2023-04-19 19:06 -0700
Last post2023-04-27 18:27 +0200
Articles 20 on this page of 85 — 14 participants

Back to article view | Back to comp.arch.embedded


Contents

  Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-19 19:06 -0700
    Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-20 12:14 +0300
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:18 -0700
    Re: Embedding a Checksum in an Image File "Peter Heitzer" <peter.heitzer@rz.uni-regensburg.de> - 2023-04-20 11:30 +0000
    Re: Embedding a Checksum in an Image File dalai lamah <antonio12358@hotmail.com> - 2023-04-20 13:47 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:04 -0700
    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 16:46 +0200
    Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 11:33 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 09:45 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 22:26 +0200
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:36 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:12 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:35 +0200
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:44 -0400
          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-20 22:37 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 12:43 +0200
              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 04:39 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 16:50 +0200
                  Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 17:29 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 16:57 +0200
                      Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-24 00:32 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 16:37 +0200
                          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-05-03 00:15 -0700
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-03 14:48 +0200
                              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:42 +0200
                                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:06 +0200
                                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-10 12:03 +0200
                                    Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:48 -0700
                              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:42 -0700
          Re: Embedding a Checksum in an Image File Stefan Reuther <stefan.news@arcor.de> - 2023-04-21 19:40 +0200
      Re: Embedding a Checksum in an Image File Tauno Voipio <tauno.voipio@notused.fi.invalid> - 2023-04-20 20:17 +0300
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:49 -0400
    Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-20 22:09 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 19:41 -0700
        Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-21 19:30 -0400
    Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 01:53 -0700
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 05:12 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 17:02 +0200
          Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:56 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:01 +0200
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:14 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:13 +0200
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:56 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 19:54 +0200
                  Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-22 20:05 +0000
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 17:37 +0200
                      Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-23 17:37 +0000
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:45 +0200
                          Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-23 18:16 -0400
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:13 +0200
                  Re: Embedding a Checksum in an Image File boB <boB@K7IQ.com> - 2023-04-22 13:41 -0700
                  Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 10:34 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:58 +0200
                      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 15:24 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:17 +0200
                          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-24 01:07 -0700
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:42 +0200
              Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:20 +0200
                Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:44 +0200
        Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:52 -0700
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:23 -0700
            Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-22 07:07 -0700
              Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-22 10:31 -0400
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:54 -0700
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:26 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-27 10:09 -0700
        Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:29 +0300
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:39 +0300
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:44 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:38 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:50 +0200
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:04 +0200
                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:03 +0200
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-30 16:19 +0200
                      Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:34 +0200
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:18 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:33 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:36 +0200
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-28 01:10 +0300
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:54 +0200
      Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:24 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:56 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:09 +0200
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:02 +0200
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:27 +0200

Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →


#31815

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-21 20:14 -0700
Message-ID<ef5ad4e6-57ed-4950-baed-3a6746b9d16en@googlegroups.com>
In reply to#31809
On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
> On 21/04/2023 14:12, Rick C wrote: 
> > 
> > This is simply to be able to say this version is unique, regardless 
> > of what the version number says. Version numbers are set manually 
> > and not always done correctly. I'm looking for something as a backup 
> > so that if the checksums are different, I can be sure the versions 
> > are not the same. 
> > 
> > The less work involved, the better. 
> >
> Run a simple 32-bit crc over the image. The result is a hash of the 
> image. Any change in the image will show up as a change in the crc.

No one is trying to detect changes in the image.  I'm trying to label the image in a way that can be read in operation.  I'm using the checksum simply because that is easy to generate.  I've had problems with version numbering in the past.  It will be used, but I want it supplemented with a number that will change every time the design changes, at least with a high probability, such as 1 in 64k. 

-- 

  Rick C.

  --- Get 1,000 miles of free Supercharging
  --- Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31822

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-22 17:13 +0200
Message-ID<u20tin$3alrj$1@dont-email.me>
In reply to#31815
On 22/04/2023 05:14, Rick C wrote:
> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
>> On 21/04/2023 14:12, Rick C wrote:
>>> 
>>> This is simply to be able to say this version is unique,
>>> regardless of what the version number says. Version numbers are
>>> set manually and not always done correctly. I'm looking for
>>> something as a backup so that if the checksums are different, I
>>> can be sure the versions are not the same.
>>> 
>>> The less work involved, the better.
>>> 
>> Run a simple 32-bit crc over the image. The result is a hash of
>> the image. Any change in the image will show up as a change in the
>> crc.
> 
> No one is trying to detect changes in the image.  I'm trying to label
> the image in a way that can be read in operation.  I'm using the
> checksum simply because that is easy to generate.  I've had problems
> with version numbering in the past.  It will be used, but I want it
> supplemented with a number that will change every time the design
> changes, at least with a high probability, such as 1 in 64k.
> 

Again - use a CRC.  It will give you what you want.

You might want to go for 32-bit CRC rather than a 16-bit CRC, depending 
on the kind of program, how often you build it, and what consequences a 
hash collision could have.  With a 16-bit CRC, you have a 5% chance of a 
collision after 82 builds.  If collisions only matter for releases, and 
you only release a couple of updates, fine - but if they matter during 
development builds, you are getting a more significant risk.  Since a 
32-bit CRC is quick and easy, it's worth using.

[toc] | [prev] | [next] | [standalone]


#31824

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-22 09:56 -0700
Message-ID<3986aac8-aec3-4cde-8bb8-8a61c20084c9n@googlegroups.com>
In reply to#31822
On Saturday, April 22, 2023 at 11:13:32 AM UTC-4, David Brown wrote:
> On 22/04/2023 05:14, Rick C wrote: 
> > On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote: 
> >> On 21/04/2023 14:12, Rick C wrote: 
> >>> 
> >>> This is simply to be able to say this version is unique, 
> >>> regardless of what the version number says. Version numbers are 
> >>> set manually and not always done correctly. I'm looking for 
> >>> something as a backup so that if the checksums are different, I 
> >>> can be sure the versions are not the same. 
> >>> 
> >>> The less work involved, the better. 
> >>> 
> >> Run a simple 32-bit crc over the image. The result is a hash of 
> >> the image. Any change in the image will show up as a change in the 
> >> crc. 
> > 
> > No one is trying to detect changes in the image. I'm trying to label 
> > the image in a way that can be read in operation. I'm using the 
> > checksum simply because that is easy to generate. I've had problems 
> > with version numbering in the past. It will be used, but I want it 
> > supplemented with a number that will change every time the design 
> > changes, at least with a high probability, such as 1 in 64k. 
> >
> Again - use a CRC. It will give you what you want. 

Again - as will a simple addition checksum.


> You might want to go for 32-bit CRC rather than a 16-bit CRC, depending 
> on the kind of program, how often you build it, and what consequences a 
> hash collision could have. With a 16-bit CRC, you have a 5% chance of a 
> collision after 82 builds. If collisions only matter for releases, and 
> you only release a couple of updates, fine - but if they matter during 
> development builds, you are getting a more significant risk. Since a 
> 32-bit CRC is quick and easy, it's worth using.

Or, I might want to go with a simple checksum. 

Thanks for your comments. 

-- 

  Rick C.

  -++ Get 1,000 miles of free Supercharging
  -++ Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31825

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-22 19:54 +0200
Message-ID<u2171f$3cbcu$1@dont-email.me>
In reply to#31824
On 22/04/2023 18:56, Rick C wrote:
> On Saturday, April 22, 2023 at 11:13:32 AM UTC-4, David Brown wrote:
>> On 22/04/2023 05:14, Rick C wrote:
>>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
>>>> On 21/04/2023 14:12, Rick C wrote:
>>>>>
>>>>> This is simply to be able to say this version is unique,
>>>>> regardless of what the version number says. Version numbers are
>>>>> set manually and not always done correctly. I'm looking for
>>>>> something as a backup so that if the checksums are different, I
>>>>> can be sure the versions are not the same.
>>>>>
>>>>> The less work involved, the better.
>>>>>
>>>> Run a simple 32-bit crc over the image. The result is a hash of
>>>> the image. Any change in the image will show up as a change in the
>>>> crc.
>>>
>>> No one is trying to detect changes in the image. I'm trying to label
>>> the image in a way that can be read in operation. I'm using the
>>> checksum simply because that is easy to generate. I've had problems
>>> with version numbering in the past. It will be used, but I want it
>>> supplemented with a number that will change every time the design
>>> changes, at least with a high probability, such as 1 in 64k.
>>>
>> Again - use a CRC. It will give you what you want.
> 
> Again - as will a simple addition checksum.

A simple addition checksum might be okay much of the time, but it 
doesn't have the resolving power of a CRC.  If the source code changes 
"a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum is likely to 
be exactly the same despite the change in the source.  In general, you 
will have much higher chance of collisions, though I think it would be 
very hard to quantify that.

Maybe it will be good enough for you.  Simple checksums were popular 
once, and can still make sense if you are very short on program space. 
But there are good reasons why they fell out of favour in many uses.

> 
> 
>> You might want to go for 32-bit CRC rather than a 16-bit CRC, depending
>> on the kind of program, how often you build it, and what consequences a
>> hash collision could have. With a 16-bit CRC, you have a 5% chance of a
>> collision after 82 builds. If collisions only matter for releases, and
>> you only release a couple of updates, fine - but if they matter during
>> development builds, you are getting a more significant risk. Since a
>> 32-bit CRC is quick and easy, it's worth using.
> 
> Or, I might want to go with a simple checksum.
> 
> Thanks for your comments.
> 


It's your choice (obviously).  I only point out the weaknesses in case 
anyone else is listening in to the thread.

If you like, I can post code for a 32-bit CRC.  It's a table, and a few 
lines of C code.



[toc] | [prev] | [next] | [standalone]


#31826

FromGrant Edwards <invalid@invalid.invalid>
Date2023-04-22 20:05 +0000
Message-ID<u21eln$25c$1@reader2.panix.com>
In reply to#31825
On 2023-04-22, David Brown <david.brown@hesbynett.no> wrote:

> A simple addition checksum might be okay much of the time, but it 
> doesn't have the resolving power of a CRC.  If the source code changes 
> "a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum is likely to 
> be exactly the same despite the change in the source.  In general, you 
> will have much higher chance of collisions, though I think it would be 
> very hard to quantify that.

I remember a long discussion about this a few decades ago. An N bit
additive checksum maps the source data into the same hash space
as a N-bit crc.

Therefore, for two randomly chosen sets of input bits, they both have
a 1 in 2^N chance of a collision.  I think that means that for random
changes to an input set of unspecified properties, they would both
have the same chance that the hash is unchanged.

However... IIRC, somebody (probably at somewhere like Bell labs)
noticed that errors in data transmitted over media like phone lines
and microwave links are _not_ random. Errors tend to be "bursty" and
can be statistically characterized. And it was shown that for the
common error modes for _those_ media, CRCs were better at detecting
real-world failures than additive checksum. And (this is also
important) a CRC is far, far simpler to implement in hardware than an
additive checksum. For the same reasons, CRCs tend to get used for
things like Ethernet frames, disc sectors, etc.

Later people seem to have adopted CRCs for detecting failures in other
very dissimilar media (e.g. EPROMs) where implementing a CRC is _more_
work than an additive checksum. If the failure modes for EPROM are
similar to those studied at <wherever> when CRCs were chosen, then
CRCs are probably also a good choice for EPROMs despite the additional
overhead. If the failure modes for EPROMs are significantly different,
then CRCs might be both sub-optimal and unnecessarily expensive.

I have no hard data either way, but it was never obvious to me that
the arguments people use in favor of CRCs (better at detecting burst
errors on transmission media) necessarily applied to EPROMs.

That said, I do use CRCs rather than additive checksums for things
like EPROM and flash.

--
Grant


[toc] | [prev] | [next] | [standalone]


#31828

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-23 17:37 +0200
Message-ID<u23jc6$3s2qo$1@dont-email.me>
In reply to#31826
On 22/04/2023 22:05, Grant Edwards wrote:
> On 2023-04-22, David Brown <david.brown@hesbynett.no> wrote:
> 
>> A simple addition checksum might be okay much of the time, but it
>> doesn't have the resolving power of a CRC.  If the source code changes
>> "a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum is likely to
>> be exactly the same despite the change in the source.  In general, you
>> will have much higher chance of collisions, though I think it would be
>> very hard to quantify that.
> 
> I remember a long discussion about this a few decades ago. An N bit
> additive checksum maps the source data into the same hash space
> as a N-bit crc.
> 
> Therefore, for two randomly chosen sets of input bits, they both have
> a 1 in 2^N chance of a collision.  I think that means that for random
> changes to an input set of unspecified properties, they would both
> have the same chance that the hash is unchanged.
> 
> However... IIRC, somebody (probably at somewhere like Bell labs)
> noticed that errors in data transmitted over media like phone lines
> and microwave links are _not_ random. Errors tend to be "bursty" and
> can be statistically characterized. And it was shown that for the
> common error modes for _those_ media, CRCs were better at detecting
> real-world failures than additive checksum. And (this is also
> important) a CRC is far, far simpler to implement in hardware than an
> additive checksum. For the same reasons, CRCs tend to get used for
> things like Ethernet frames, disc sectors, etc.
> 
> Later people seem to have adopted CRCs for detecting failures in other
> very dissimilar media (e.g. EPROMs) where implementing a CRC is _more_
> work than an additive checksum. If the failure modes for EPROM are
> similar to those studied at <wherever> when CRCs were chosen, then
> CRCs are probably also a good choice for EPROMs despite the additional
> overhead. If the failure modes for EPROMs are significantly different,
> then CRCs might be both sub-optimal and unnecessarily expensive.
> 
> I have no hard data either way, but it was never obvious to me that
> the arguments people use in favor of CRCs (better at detecting burst
> errors on transmission media) necessarily applied to EPROMs.
> 
> That said, I do use CRCs rather than additive checksums for things
> like EPROM and flash.
> 

That's a lot of good points.  You are absolutely correct that CRC's are 
better for the types of errors that are often seen in transmission 
systems.  The person at Bell Labs that you are thinking about is 
probably Claude Shannon, famous for his quantitive definition of 
information and work on the information capacity of communication 
channels with noise.

Another thing you can look at is the distribution of checksum outputs, 
for random inputs.  For an additive checksum, you can consider your 
input as N independent 0-255 random values, added together.  The result 
will be a normal distribution of the checksum.  If you have, say, a 100 
byte data block and a 16-bit checksum, it's clear that you will never 
get a checksum value greater than 25500, and that you are much more 
likely to get a value close to 12750.  This kind of clustering means 
that the 16-bit checksum contains a lot less than 16 bits of 
information.  Real data - program images, data telegrams, etc., - are 
not fully random and the result is even more clustering and less 
information in the checksum.

Taking the additive checksum over a larger range, then "folding" the 
distribution back by wrapping the checksum to 8-bit or 16-bit will 
greatly reduce the clustering.  That will help a lot if you have a 
program image and use a 16-bit additive checksum, but if you need more 
than "1 in 65536" integrity, it's hard to get.

A particular weakness of purely additive checksums is that they only 
consider the values of the bytes, not their order - re-arranging the 
order of the same data gives the same additive checksum.

CRC's are not as good as more advanced hashes like SHA or MD5.  But 
their distributions are vastly better than additive checksums, and they 
provide integrity checks for a wider variety of possible errors.


Of course, for some uses, an additive checksum might be considered good 
enough.  There's no need to be more complicated than you need to be. 
But since CRC's are usually very simple and efficient to calculate, they 
give an option that is a lot better than an additive checksum for little 
extra cost, while going beyond them to MD5 or SHA involves significantly 
more effort.  (SHA is your first choice if you are protecting against 
malicious changes.)





[toc] | [prev] | [next] | [standalone]


#31830

FromGrant Edwards <invalid@invalid.invalid>
Date2023-04-23 17:37 +0000
Message-ID<u23qc9$g6p$1@reader2.panix.com>
In reply to#31828
On 2023-04-23, David Brown <david.brown@hesbynett.no> wrote:

> Another thing you can look at is the distribution of checksum outputs, 
> for random inputs.  For an additive checksum, you can consider your 
> input as N independent 0-255 random values, added together.  The result 
> will be a normal distribution of the checksum.  If you have, say, a 100 
> byte data block and a 16-bit checksum, it's clear that you will never 
> get a checksum value greater than 25500, and that you are much more 
> likely to get a value close to 12750.

It never occurred to me that for an N-bit checksum, you would sum
something other than N-bit "words" of the input data.

--
Grant

[toc] | [prev] | [next] | [standalone]


#31831

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-23 23:45 +0200
Message-ID<u248tk$3vk28$1@dont-email.me>
In reply to#31830
On 23/04/2023 19:37, Grant Edwards wrote:
> On 2023-04-23, David Brown <david.brown@hesbynett.no> wrote:
> 
>> Another thing you can look at is the distribution of checksum outputs,
>> for random inputs.  For an additive checksum, you can consider your
>> input as N independent 0-255 random values, added together.  The result
>> will be a normal distribution of the checksum.  If you have, say, a 100
>> byte data block and a 16-bit checksum, it's clear that you will never
>> get a checksum value greater than 25500, and that you are much more
>> likely to get a value close to 12750.
> 
> It never occurred to me that for an N-bit checksum, you would sum
> something other than N-bit "words" of the input data.
> 

Usually - in my experience - you sum bytes, using an unsigned integer 
8-bit or 16-bit wide.  Simple additive checksums are often used on small 
8-bit microcontrollers where CRC's are seen (rightly or wrongly) as too 
demanding.  Perhaps other people have different experiences.

You could certainly sum 16-bit words to get your 16-bit additive 
checksum, and that would give a different kind of clustering - maybe 
better, maybe not.

[toc] | [prev] | [next] | [standalone]


#31833

FromRichard Damon <Richard@Damon-Family.org>
Date2023-04-23 18:16 -0400
Message-ID<A6i1M.2358052$iU59.1633184@fx14.iad>
In reply to#31831
On 4/23/23 5:45 PM, David Brown wrote:
> On 23/04/2023 19:37, Grant Edwards wrote:
>> On 2023-04-23, David Brown <david.brown@hesbynett.no> wrote:
>>
>>> Another thing you can look at is the distribution of checksum outputs,
>>> for random inputs.  For an additive checksum, you can consider your
>>> input as N independent 0-255 random values, added together.  The result
>>> will be a normal distribution of the checksum.  If you have, say, a 100
>>> byte data block and a 16-bit checksum, it's clear that you will never
>>> get a checksum value greater than 25500, and that you are much more
>>> likely to get a value close to 12750.
>>
>> It never occurred to me that for an N-bit checksum, you would sum
>> something other than N-bit "words" of the input data.
>>
> 
> Usually - in my experience - you sum bytes, using an unsigned integer 
> 8-bit or 16-bit wide.  Simple additive checksums are often used on small 
> 8-bit microcontrollers where CRC's are seen (rightly or wrongly) as too 
> demanding.  Perhaps other people have different experiences.
> 
> You could certainly sum 16-bit words to get your 16-bit additive 
> checksum, and that would give a different kind of clustering - maybe 
> better, maybe not.
> 
> 

I have seen 16-bit checksums done both ways. Summing 16 bit units does 
eliminate the issue of clustering, and makes adjacent byte swaps 
detectable.

[toc] | [prev] | [next] | [standalone]


#31835

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-24 09:13 +0200
Message-ID<u25a69$87j1$1@dont-email.me>
In reply to#31833
On 24/04/2023 00:16, Richard Damon wrote:
> On 4/23/23 5:45 PM, David Brown wrote:
>> On 23/04/2023 19:37, Grant Edwards wrote:
>>> On 2023-04-23, David Brown <david.brown@hesbynett.no> wrote:
>>>
>>>> Another thing you can look at is the distribution of checksum outputs,
>>>> for random inputs.  For an additive checksum, you can consider your
>>>> input as N independent 0-255 random values, added together.  The result
>>>> will be a normal distribution of the checksum.  If you have, say, a 100
>>>> byte data block and a 16-bit checksum, it's clear that you will never
>>>> get a checksum value greater than 25500, and that you are much more
>>>> likely to get a value close to 12750.
>>>
>>> It never occurred to me that for an N-bit checksum, you would sum
>>> something other than N-bit "words" of the input data.
>>>
>>
>> Usually - in my experience - you sum bytes, using an unsigned integer 
>> 8-bit or 16-bit wide.  Simple additive checksums are often used on 
>> small 8-bit microcontrollers where CRC's are seen (rightly or wrongly) 
>> as too demanding.  Perhaps other people have different experiences.
>>
>> You could certainly sum 16-bit words to get your 16-bit additive 
>> checksum, and that would give a different kind of clustering - maybe 
>> better, maybe not.
>>
>>
> 
> I have seen 16-bit checksums done both ways. Summing 16 bit units does 
> eliminate the issue of clustering, and makes adjacent byte swaps 
> detectable.

Long ago, there used to be a definite risk of mixing up endianness when 
dealing with program images burned to flash or eeprom.  Popular "hex" 
formats like Intel Hex and Motorola SRecord could differ in endianness. 
So byte swaps in the entire image was a real possibility, and good to 
guard against.  But it's hard to imagine how an individual byte swap 
could occur - I see bigger movements and re-arrangements being more 
likely, and using 16-bit units will not help much there.  Still, I think 
there is little doubt that using 16-bit units is better than using 8-bit 
units in many ways (except for efficient implementation on small 8-bit 
devices).

[toc] | [prev] | [next] | [standalone]


#31827

FromboB <boB@K7IQ.com>
Date2023-04-22 13:41 -0700
Message-ID<ich84idn1vub7t4lggn1rkdrq00crmh3rc@4ax.com>
In reply to#31825
On Sat, 22 Apr 2023 19:54:54 +0200, David Brown
<david.brown@hesbynett.no> wrote:

>On 22/04/2023 18:56, Rick C wrote:
>> On Saturday, April 22, 2023 at 11:13:32?AM UTC-4, David Brown wrote:
>>> On 22/04/2023 05:14, Rick C wrote:
>>>> On Friday, April 21, 2023 at 11:02:28?AM UTC-4, David Brown wrote:
>>>>> On 21/04/2023 14:12, Rick C wrote:
>>>>>>
>>>>>> This is simply to be able to say this version is unique,
>>>>>> regardless of what the version number says. Version numbers are
>>>>>> set manually and not always done correctly. I'm looking for
>>>>>> something as a backup so that if the checksums are different, I
>>>>>> can be sure the versions are not the same.
>>>>>>
>>>>>> The less work involved, the better.
>>>>>>
>>>>> Run a simple 32-bit crc over the image. The result is a hash of
>>>>> the image. Any change in the image will show up as a change in the
>>>>> crc.
>>>>
>>>> No one is trying to detect changes in the image. I'm trying to label
>>>> the image in a way that can be read in operation. I'm using the
>>>> checksum simply because that is easy to generate. I've had problems
>>>> with version numbering in the past. It will be used, but I want it
>>>> supplemented with a number that will change every time the design
>>>> changes, at least with a high probability, such as 1 in 64k.
>>>>
>>> Again - use a CRC. It will give you what you want.
>> 
>> Again - as will a simple addition checksum.
>
>A simple addition checksum might be okay much of the time, but it 
>doesn't have the resolving power of a CRC.  If the source code changes 
>"a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum is likely to 
>be exactly the same despite the change in the source.  In general, you 
>will have much higher chance of collisions, though I think it would be 
>very hard to quantify that.
>
>Maybe it will be good enough for you.  Simple checksums were popular 
>once, and can still make sense if you are very short on program space. 
>But there are good reasons why they fell out of favour in many uses.
>
>> 
>> 
>>> You might want to go for 32-bit CRC rather than a 16-bit CRC, depending
>>> on the kind of program, how often you build it, and what consequences a
>>> hash collision could have. With a 16-bit CRC, you have a 5% chance of a
>>> collision after 82 builds. If collisions only matter for releases, and
>>> you only release a couple of updates, fine - but if they matter during
>>> development builds, you are getting a more significant risk. Since a
>>> 32-bit CRC is quick and easy, it's worth using.

Totally agree !  I stopped using simple checksums years ago.
 Many processors these days also have a CRC peripheral that makes it
easy to use.  And I can simply chop that off to 16 bits if I don't
want to transmit all 32 bits.  OR even 24 bits.

boB





>> 
>> Or, I might want to go with a simple checksum.
>> 
>> Thanks for your comments.
>> 
>
>
>It's your choice (obviously).  I only point out the weaknesses in case 
>anyone else is listening in to the thread.
>
>If you like, I can post code for a 32-bit CRC.  It's a table, and a few 
>lines of C code.
>
>
>

[toc] | [prev] | [next] | [standalone]


#31829

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-23 10:34 -0700
Message-ID<99aaf8df-1f0e-4911-9706-0bac770e3d1cn@googlegroups.com>
In reply to#31825
On Saturday, April 22, 2023 at 1:55:01 PM UTC-4, David Brown wrote:
> On 22/04/2023 18:56, Rick C wrote: 
> > On Saturday, April 22, 2023 at 11:13:32 AM UTC-4, David Brown wrote: 
> >> On 22/04/2023 05:14, Rick C wrote: 
> >>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote: 
> >>>> On 21/04/2023 14:12, Rick C wrote: 
> >>>>> 
> >>>>> This is simply to be able to say this version is unique, 
> >>>>> regardless of what the version number says. Version numbers are 
> >>>>> set manually and not always done correctly. I'm looking for 
> >>>>> something as a backup so that if the checksums are different, I 
> >>>>> can be sure the versions are not the same. 
> >>>>> 
> >>>>> The less work involved, the better. 
> >>>>> 
> >>>> Run a simple 32-bit crc over the image. The result is a hash of 
> >>>> the image. Any change in the image will show up as a change in the 
> >>>> crc. 
> >>> 
> >>> No one is trying to detect changes in the image. I'm trying to label 
> >>> the image in a way that can be read in operation. I'm using the 
> >>> checksum simply because that is easy to generate. I've had problems 
> >>> with version numbering in the past. It will be used, but I want it 
> >>> supplemented with a number that will change every time the design 
> >>> changes, at least with a high probability, such as 1 in 64k. 
> >>> 
> >> Again - use a CRC. It will give you what you want. 
> > 
> > Again - as will a simple addition checksum.
> A simple addition checksum might be okay much of the time, but it 
> doesn't have the resolving power of a CRC. If the source code changes 
> "a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum is likely to 
> be exactly the same despite the change in the source. In general, you 
> will have much higher chance of collisions, though I think it would be 
> very hard to quantify that. 
> 
> Maybe it will be good enough for you. Simple checksums were popular 
> once, and can still make sense if you are very short on program space. 
> But there are good reasons why they fell out of favour in many uses.
> > 
> > 
> >> You might want to go for 32-bit CRC rather than a 16-bit CRC, depending 
> >> on the kind of program, how often you build it, and what consequences a 
> >> hash collision could have. With a 16-bit CRC, you have a 5% chance of a 
> >> collision after 82 builds. If collisions only matter for releases, and 
> >> you only release a couple of updates, fine - but if they matter during 
> >> development builds, you are getting a more significant risk. Since a 
> >> 32-bit CRC is quick and easy, it's worth using. 
> > 
> > Or, I might want to go with a simple checksum. 
> > 
> > Thanks for your comments. 
> >
> It's your choice (obviously). I only point out the weaknesses in case 
> anyone else is listening in to the thread. 
> 
> If you like, I can post code for a 32-bit CRC. It's a table, and a few 
> lines of C code.

You know nothing of the project I am working on or those that I typically work on.  But thanks for the advice. 

-- 

  Rick C.

  +-- Get 1,000 miles of free Supercharging
  +-- Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31832

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-23 23:58 +0200
Message-ID<u249ml$3vn9o$1@dont-email.me>
In reply to#31829
On 23/04/2023 19:34, Rick C wrote:
> On Saturday, April 22, 2023 at 1:55:01 PM UTC-4, David Brown wrote:
>> On 22/04/2023 18:56, Rick C wrote:
>>> On Saturday, April 22, 2023 at 11:13:32 AM UTC-4, David Brown
>>> wrote:
>>>> On 22/04/2023 05:14, Rick C wrote:
>>>>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown
>>>>> wrote:
>>>>>> On 21/04/2023 14:12, Rick C wrote:
>>>>>>> 
>>>>>>> This is simply to be able to say this version is unique, 
>>>>>>> regardless of what the version number says. Version
>>>>>>> numbers are set manually and not always done correctly.
>>>>>>> I'm looking for something as a backup so that if the
>>>>>>> checksums are different, I can be sure the versions are
>>>>>>> not the same.
>>>>>>> 
>>>>>>> The less work involved, the better.
>>>>>>> 
>>>>>> Run a simple 32-bit crc over the image. The result is a
>>>>>> hash of the image. Any change in the image will show up as
>>>>>> a change in the crc.
>>>>> 
>>>>> No one is trying to detect changes in the image. I'm trying
>>>>> to label the image in a way that can be read in operation.
>>>>> I'm using the checksum simply because that is easy to
>>>>> generate. I've had problems with version numbering in the
>>>>> past. It will be used, but I want it supplemented with a
>>>>> number that will change every time the design changes, at
>>>>> least with a high probability, such as 1 in 64k.
>>>>> 
>>>> Again - use a CRC. It will give you what you want.
>>> 
>>> Again - as will a simple addition checksum.
>> A simple addition checksum might be okay much of the time, but it 
>> doesn't have the resolving power of a CRC. If the source code
>> changes "a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum
>> is likely to be exactly the same despite the change in the source.
>> In general, you will have much higher chance of collisions, though
>> I think it would be very hard to quantify that.
>> 
>> Maybe it will be good enough for you. Simple checksums were
>> popular once, and can still make sense if you are very short on
>> program space. But there are good reasons why they fell out of
>> favour in many uses.
>>> 
>>> 
>>>> You might want to go for 32-bit CRC rather than a 16-bit CRC,
>>>> depending on the kind of program, how often you build it, and
>>>> what consequences a hash collision could have. With a 16-bit
>>>> CRC, you have a 5% chance of a collision after 82 builds. If
>>>> collisions only matter for releases, and you only release a
>>>> couple of updates, fine - but if they matter during development
>>>> builds, you are getting a more significant risk. Since a 32-bit
>>>> CRC is quick and easy, it's worth using.
>>> 
>>> Or, I might want to go with a simple checksum.
>>> 
>>> Thanks for your comments.
>>> 
>> It's your choice (obviously). I only point out the weaknesses in
>> case anyone else is listening in to the thread.
>> 
>> If you like, I can post code for a 32-bit CRC. It's a table, and a
>> few lines of C code.
> 
> You know nothing of the project I am working on or those that I
> typically work on.  But thanks for the advice.
> 

You haven't given much to go on.  It is still not really clear (to me, 
at least) if you are asking about checksums or how to manipulate binary 
images as part of a build process, or what you are really asking.

When someone wants a checksum on an image file, the appropriate choice 
in most cases is a CRC.  If security is an issue, then a secure hash is 
needed.  For a very limited system, additive checksums might be then 
only realistic choice.

But more often, the reason people pick additive checksums rather than 
CRCs is because they don't realise that CRCs are actually very simple 
and efficient to implement.  People unfamiliar with them might have read 
a little, and think they need to do calculations for each bit (which is 
possible but /slow/), or that they would have to understand the theory 
of binary polynomial division rings (they don't).  They think CRC's are 
complicated and advanced, and shy away from them.

There are a number of people who read this group - maybe some of them 
have learned a little from this thread.


[toc] | [prev] | [next] | [standalone]


#31834

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-23 15:24 -0700
Message-ID<070fc693-6b70-4767-af26-2d3ccb4f7919n@googlegroups.com>
In reply to#31832
On Sunday, April 23, 2023 at 5:58:51 PM UTC-4, David Brown wrote:
> On 23/04/2023 19:34, Rick C wrote: 
> > On Saturday, April 22, 2023 at 1:55:01 PM UTC-4, David Brown wrote: 
> >> On 22/04/2023 18:56, Rick C wrote: 
> >>> On Saturday, April 22, 2023 at 11:13:32 AM UTC-4, David Brown 
> >>> wrote: 
> >>>> On 22/04/2023 05:14, Rick C wrote: 
> >>>>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown 
> >>>>> wrote: 
> >>>>>> On 21/04/2023 14:12, Rick C wrote: 
> >>>>>>> 
> >>>>>>> This is simply to be able to say this version is unique, 
> >>>>>>> regardless of what the version number says. Version 
> >>>>>>> numbers are set manually and not always done correctly. 
> >>>>>>> I'm looking for something as a backup so that if the 
> >>>>>>> checksums are different, I can be sure the versions are 
> >>>>>>> not the same. 
> >>>>>>> 
> >>>>>>> The less work involved, the better. 
> >>>>>>> 
> >>>>>> Run a simple 32-bit crc over the image. The result is a 
> >>>>>> hash of the image. Any change in the image will show up as 
> >>>>>> a change in the crc. 
> >>>>> 
> >>>>> No one is trying to detect changes in the image. I'm trying 
> >>>>> to label the image in a way that can be read in operation. 
> >>>>> I'm using the checksum simply because that is easy to 
> >>>>> generate. I've had problems with version numbering in the 
> >>>>> past. It will be used, but I want it supplemented with a 
> >>>>> number that will change every time the design changes, at 
> >>>>> least with a high probability, such as 1 in 64k. 
> >>>>> 
> >>>> Again - use a CRC. It will give you what you want. 
> >>> 
> >>> Again - as will a simple addition checksum. 
> >> A simple addition checksum might be okay much of the time, but it 
> >> doesn't have the resolving power of a CRC. If the source code 
> >> changes "a = 1; b = 2;" to "a = 2; b = 1;", the addition checksum 
> >> is likely to be exactly the same despite the change in the source. 
> >> In general, you will have much higher chance of collisions, though 
> >> I think it would be very hard to quantify that. 
> >> 
> >> Maybe it will be good enough for you. Simple checksums were 
> >> popular once, and can still make sense if you are very short on 
> >> program space. But there are good reasons why they fell out of 
> >> favour in many uses. 
> >>> 
> >>> 
> >>>> You might want to go for 32-bit CRC rather than a 16-bit CRC, 
> >>>> depending on the kind of program, how often you build it, and 
> >>>> what consequences a hash collision could have. With a 16-bit 
> >>>> CRC, you have a 5% chance of a collision after 82 builds. If 
> >>>> collisions only matter for releases, and you only release a 
> >>>> couple of updates, fine - but if they matter during development 
> >>>> builds, you are getting a more significant risk. Since a 32-bit 
> >>>> CRC is quick and easy, it's worth using. 
> >>> 
> >>> Or, I might want to go with a simple checksum. 
> >>> 
> >>> Thanks for your comments. 
> >>> 
> >> It's your choice (obviously). I only point out the weaknesses in 
> >> case anyone else is listening in to the thread. 
> >> 
> >> If you like, I can post code for a 32-bit CRC. It's a table, and a 
> >> few lines of C code. 
> > 
> > You know nothing of the project I am working on or those that I 
> > typically work on. But thanks for the advice. 
> >
> You haven't given much to go on. It is still not really clear (to me, 
> at least) if you are asking about checksums or how to manipulate binary 
> images as part of a build process, or what you are really asking. 

If you don't understand, you are making this far more complicated than it is.  I don't know what to tell you.  There are no other details that are relevant.  Don't read into this, what is not there. 


> When someone wants a checksum on an image file, the appropriate choice 
> in most cases is a CRC. 

Why?  What makes a CRC an "appropriate" choice.  Normally, when I design something, I establish the requirements.  What requirements are you assuming, that would make the CRC more desireable than a simple checksum? 


> If security is an issue, then a secure hash is 
> needed. For a very limited system, additive checksums might be then 
> only realistic choice. 

What have I said that makes you think security is an issue???  I don't recall ever mentioning anything about security.  Do you recall what I did say? 


> But more often, the reason people pick additive checksums rather than 
> CRCs is because they don't realise that CRCs are actually very simple 
> and efficient to implement. 

The fact that they are "simple and efficient" is not a reason to use them.  I repeat, what are the requirements? 


> People unfamiliar with them might have read 
> a little, and think they need to do calculations for each bit (which is 
> possible but /slow/), or that they would have to understand the theory 
> of binary polynomial division rings (they don't). They think CRC's are 
> complicated and advanced, and shy away from them. 
> 
> There are a number of people who read this group - maybe some of them 
> have learned a little from this thread.

I suppose there is that possibility.  But when people make claims about something being good or "better", without substantiation, there's not much to learn. 

If you think a discussion of CRC calculations would be useful, why don't you open a thread and discuss them, instead of insisting they are the right solution to my problem, when you don't even know what the problem requirements are?  It's all here in the thread.  You only need to read, without projecting your opinions on the problem statement. 

-- 

  Rick C.

  +-+ Get 1,000 miles of free Supercharging
  +-+ Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31836

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-24 09:17 +0200
Message-ID<u25ae7$87j1$2@dont-email.me>
In reply to#31834
On 24/04/2023 00:24, Rick C wrote:
> On Sunday, April 23, 2023 at 5:58:51 PM UTC-4, David Brown wrote:
> 
>> When someone wants a checksum on an image file, the appropriate
>> choice in most cases is a CRC.
> 
> Why?  What makes a CRC an "appropriate" choice.  Normally, when I
> design something, I establish the requirements.  What requirements
> are you assuming, that would make the CRC more desireable than a
> simple checksum?
> 

I've already explained this in quite a lot of detail in this thread (as 
have others).  If you don't like my explanation, or didn't read it, 
that's okay.  You are under no obligation to learn about CRCs.  Or if 
you prefer to look it up in other sources, that's obviously also an option.

> 
>> If security is an issue, then a secure hash is needed. For a very
>> limited system, additive checksums might be then only realistic
>> choice.
> 
> What have I said that makes you think security is an issue???  I
> don't recall ever mentioning anything about security.  Do you recall
> what I did say?
> 
> 
> If you think a discussion of CRC calculations would be useful, why
> don't you open a thread and discuss them, instead of insisting they
> are the right solution to my problem, when you don't even know what
> the problem requirements are?  It's all here in the thread.  You only
> need to read, without projecting your opinions on the problem
> statement.
> 


I've asked you this before - are you /sure/ you understand how Usenet works?

[toc] | [prev] | [next] | [standalone]


#31838

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-24 01:07 -0700
Message-ID<cbfefcb9-2d39-4e75-bb4f-cbda15e1b91en@googlegroups.com>
In reply to#31836
On Monday, April 24, 2023 at 3:17:33 AM UTC-4, David Brown wrote:
> On 24/04/2023 00:24, Rick C wrote: 
> > On Sunday, April 23, 2023 at 5:58:51 PM UTC-4, David Brown wrote: 
> > 
> >> When someone wants a checksum on an image file, the appropriate 
> >> choice in most cases is a CRC. 
> > 
> > Why? What makes a CRC an "appropriate" choice. Normally, when I 
> > design something, I establish the requirements. What requirements 
> > are you assuming, that would make the CRC more desireable than a 
> > simple checksum? 
> >
> I've already explained this in quite a lot of detail in this thread (as 
> have others). If you don't like my explanation, or didn't read it, 
> that's okay. You are under no obligation to learn about CRCs. Or if 
> you prefer to look it up in other sources, that's obviously also an option.

Hmmm...  I ask you a question about why you think CRC is better for my application and you respond oddly.  So you can't explain why the CRC would be better for my application?  OK, thanks anyway. 


> >> If security is an issue, then a secure hash is needed. For a very 
> >> limited system, additive checksums might be then only realistic 
> >> choice. 
> > 
> > What have I said that makes you think security is an issue??? I 
> > don't recall ever mentioning anything about security. Do you recall 
> > what I did say? 
> > 
> >
> > If you think a discussion of CRC calculations would be useful, why 
> > don't you open a thread and discuss them, instead of insisting they 
> > are the right solution to my problem, when you don't even know what 
> > the problem requirements are? It's all here in the thread. You only 
> > need to read, without projecting your opinions on the problem 
> > statement. 
> >
> I've asked you this before - are you /sure/ you understand how Usenet works?

I will say this again, rather than burying your comments on CRC in this thread about checksums, why not open a new thread, and allow the world to read what you have to say, instead of commenting as a side topic in a thread where most people have tuned out long ago?  You can use an appropriate subject line like, "Why CRC is better than checksums for some applications".  

Or you can continue to muddy up the waters here by discussing something that is of no value in this application. 

-- 

  Rick C.

  ++- Get 1,000 miles of free Supercharging
  ++- Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31858

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-04-27 18:42 +0200
Message-ID<u2e8lu$205k5$4@dont-email.me>
In reply to#31815
Den 2023-04-22 kl. 05:14, skrev Rick C:
> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
>> On 21/04/2023 14:12, Rick C wrote:
>>>
>>> This is simply to be able to say this version is unique, regardless
>>> of what the version number says. Version numbers are set manually
>>> and not always done correctly. I'm looking for something as a backup
>>> so that if the checksums are different, I can be sure the versions
>>> are not the same.
>>>
>>> The less work involved, the better.
>>>
>> Run a simple 32-bit crc over the image. The result is a hash of the
>> image. Any change in the image will show up as a change in the crc.
> 
> No one is trying to detect changes in the image.  I'm trying to label the image in a way that can be read in operation.  I'm using the checksum simply because that is easy to generate.  I've had problems with version numbering in the past.  It will be used, but I want it supplemented with a number that will change every time the design changes, at least with a high probability, such as 1 in 64k.
> 

Another thing I added (and was later removed) was a timestamp directive.
A 64 bit integer with the number of seconds since 1970-01-01 00:00.

/Ulf

[toc] | [prev] | [next] | [standalone]


#31872

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-28 09:20 +0200
Message-ID<u2fs41$2bnd0$1@dont-email.me>
In reply to#31858
On 27/04/2023 18:42, Ulf Samuelsson wrote:
> Den 2023-04-22 kl. 05:14, skrev Rick C:
>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
>>> On 21/04/2023 14:12, Rick C wrote:
>>>>
>>>> This is simply to be able to say this version is unique, regardless
>>>> of what the version number says. Version numbers are set manually
>>>> and not always done correctly. I'm looking for something as a backup
>>>> so that if the checksums are different, I can be sure the versions
>>>> are not the same.
>>>>
>>>> The less work involved, the better.
>>>>
>>> Run a simple 32-bit crc over the image. The result is a hash of the
>>> image. Any change in the image will show up as a change in the crc.
>>
>> No one is trying to detect changes in the image.  I'm trying to label 
>> the image in a way that can be read in operation.  I'm using the 
>> checksum simply because that is easy to generate.  I've had problems 
>> with version numbering in the past.  It will be used, but I want it 
>> supplemented with a number that will change every time the design 
>> changes, at least with a high probability, such as 1 in 64k.
>>
> 
> Another thing I added (and was later removed) was a timestamp directive.
> A 64 bit integer with the number of seconds since 1970-01-01 00:00.
> 

Timestamping a build in some way (as part of the "make", using __DATE__ 
or __TIME__ in source code, or some feature of a revision control 
system) is very tempting, and can be helpful for tracking exactly what 
code you have on the system.

However, IMHO having reproducible builds is much more valuable.  I am 
not happy with a project build until I am getting identical binaries 
built on multiple hosts (Windows and Linux).  That's how you can be 
absolutely sure of what code went into a particular binary, even years 
or decades later.

A compromise that can work is to distinguish development builds and 
production builds, and have timestamping in development builds.  That 
also reduces the rate at which your minor version number or build number 
goes up, and avoids endless changes to your "version.h" include file.



[toc] | [prev] | [next] | [standalone]


#31880

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-04-28 10:44 +0200
Message-ID<u2g10k$2cdfn$2@dont-email.me>
In reply to#31872
Den 2023-04-28 kl. 09:20, skrev David Brown:
> On 27/04/2023 18:42, Ulf Samuelsson wrote:
>> Den 2023-04-22 kl. 05:14, skrev Rick C:
>>> On Friday, April 21, 2023 at 11:02:28 AM UTC-4, David Brown wrote:
>>>> On 21/04/2023 14:12, Rick C wrote:
>>>>>
>>>>> This is simply to be able to say this version is unique, regardless
>>>>> of what the version number says. Version numbers are set manually
>>>>> and not always done correctly. I'm looking for something as a backup
>>>>> so that if the checksums are different, I can be sure the versions
>>>>> are not the same.
>>>>>
>>>>> The less work involved, the better.
>>>>>
>>>> Run a simple 32-bit crc over the image. The result is a hash of the
>>>> image. Any change in the image will show up as a change in the crc.
>>>
>>> No one is trying to detect changes in the image.  I'm trying to label 
>>> the image in a way that can be read in operation.  I'm using the 
>>> checksum simply because that is easy to generate.  I've had problems 
>>> with version numbering in the past.  It will be used, but I want it 
>>> supplemented with a number that will change every time the design 
>>> changes, at least with a high probability, such as 1 in 64k.
>>>
>>
>> Another thing I added (and was later removed) was a timestamp directive.
>> A 64 bit integer with the number of seconds since 1970-01-01 00:00.
>>
> 
> Timestamping a build in some way (as part of the "make", using __DATE__ 
> or __TIME__ in source code, or some feature of a revision control 
> system) is very tempting, and can be helpful for tracking exactly what 
> code you have on the system.
> 
> However, IMHO having reproducible builds is much more valuable.  I am 
> not happy with a project build until I am getting identical binaries 
> built on multiple hosts (Windows and Linux).  That's how you can be 
> absolutely sure of what code went into a particular binary, even years 
> or decades later.

With the timestamp located in the header, you can simply compare the 
non-header area.
Make with __DATE__ or __TIME__ will tell you when that module is 
compiled, not when the program is generated.
That is why TIMESTAMP is best generated in the linker.

/Ulf

> 
> A compromise that can work is to distinguish development builds and 
> production builds, and have timestamping in development builds.  That 
> also reduces the rate at which your minor version number or build number 
> goes up, and avoids endless changes to your "version.h" include file.
> 
> 
> 
> 

[toc] | [prev] | [next] | [standalone]


#31812

FromBrian Cockburn <brian.cockburn.1959@gmail.com>
Date2023-04-21 16:52 -0700
Message-ID<1f26bbc6-964c-4081-b9f6-f460a799c9b0n@googlegroups.com>
In reply to#31807
On Friday, April 21, 2023 at 10:12:49 PM UTC+10, Rick C wrote:
> On Friday, April 21, 2023 at 4:53:18 AM UTC-4, Brian Cockburn wrote: 
> > On Thursday, April 20, 2023 at 12:06:36 PM UTC+10, Rick C wrote: 
> > > This is a bit of the chicken and egg thing. If you want a embed a checksum in a code module to report the checksum, is there a way of doing this? It's a bit like being your own grandfather, I think. 
> > > 
> > > I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16. 
> > > 
> > > I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way. Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum. 
> > > 
> > > I'm not thinking of any special checksum generator that excludes the checksum data. That would be too messy. 
> > > 
> > > I keep thinking there is a different way of looking at this to achieve the result I want... 
> > > 
> > > Maybe I can prove it is impossible. Assume the file checksums to X when the checksum data is zero. The goal would then be to include the checksum data value Y in the file, that would change X to Y. Given the properties of the module N checksum, this would appear to be impossible for the general case, unless... Add another data value, called, checksum normalizer. This data value checksums with the original checksum to give the result zero. Then, when the checksum is also added, the resulting checksum is, in fact, the checksum. Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact. 
> > > 
> > > This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think. At least, this could work in software, where data can be included in an image file as itself. In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted. Memory might have data that is stored as itself. I'll need to look into that. 
> > > 
> > > -- 
> > > 
> > > Rick C. 
> > > 
> > > - Get 1,000 miles of free Supercharging 
> > > - Tesla referral code - https://ts.la/richard11209 
> > Rick, What is the purpose of this? Is it (1) to be able to externally identify a binary, as one might a ROM image by computing a checksum? Is it (2) for a run-able binary to be able to check itself? This would of course only be able to detect corruption, not tampering. Is it (3) for the loader (whatever that might be) to be able to say 'this binary has the correct checksum' and only jump to it if it does? Again this would only be able to detect corruption, not tampering. Are you hoping for more than corruption detection?
> This is simply to be able to say this version is unique, regardless of what the version number says. Version numbers are set manually and not always done correctly. I'm looking for something as a backup so that if the checksums are different, I can be sure the versions are not the same. 
> 
> The less work involved, the better. 
> 
> -- 
> 
> Rick C. 
> 
> ++ Get 1,000 miles of free Supercharging 
> ++ Tesla referral code - https://ts.la/richard11209
Rick, so you want the executable to, as part of its execution, print on the console the 'checksum' of itself?  Or do you want to be able to inspect the executable with some other tool to calculate its 'checksum'?  For the latter there are lots of tools to do that (your OS or PROM programmer for instance), for the former you need to embed the calculation code into the executable (along with the length over which to calculate) and run this when asked.  Neither of these involve embedding the 'checksum' value.
And just to be sure I understand what you wrote in a somewhat convoluted way.  When you have two binary executables that report the same version number you want to be able to distinguish them with a 'checksum', right?

[toc] | [prev] | [next] | [standalone]


Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →

Back to top | Article view | comp.arch.embedded


csiph-web