Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch.embedded > #31788 > unrolled thread

Embedding a Checksum in an Image File

Started byRick C <gnuarm.deletethisbit@gmail.com>
First post2023-04-19 19:06 -0700
Last post2023-04-27 18:27 +0200
Articles 20 on this page of 85 — 14 participants

Back to article view | Back to comp.arch.embedded


Contents

  Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-19 19:06 -0700
    Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-20 12:14 +0300
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:18 -0700
    Re: Embedding a Checksum in an Image File "Peter Heitzer" <peter.heitzer@rz.uni-regensburg.de> - 2023-04-20 11:30 +0000
    Re: Embedding a Checksum in an Image File dalai lamah <antonio12358@hotmail.com> - 2023-04-20 13:47 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:04 -0700
    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 16:46 +0200
    Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 11:33 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 09:45 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 22:26 +0200
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:36 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:12 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:35 +0200
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:44 -0400
          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-20 22:37 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 12:43 +0200
              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 04:39 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 16:50 +0200
                  Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 17:29 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 16:57 +0200
                      Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-24 00:32 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 16:37 +0200
                          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-05-03 00:15 -0700
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-03 14:48 +0200
                              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:42 +0200
                                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:06 +0200
                                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-10 12:03 +0200
                                    Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:48 -0700
                              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:42 -0700
          Re: Embedding a Checksum in an Image File Stefan Reuther <stefan.news@arcor.de> - 2023-04-21 19:40 +0200
      Re: Embedding a Checksum in an Image File Tauno Voipio <tauno.voipio@notused.fi.invalid> - 2023-04-20 20:17 +0300
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:49 -0400
    Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-20 22:09 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 19:41 -0700
        Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-21 19:30 -0400
    Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 01:53 -0700
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 05:12 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 17:02 +0200
          Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:56 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:01 +0200
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:14 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:13 +0200
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:56 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 19:54 +0200
                  Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-22 20:05 +0000
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 17:37 +0200
                      Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-23 17:37 +0000
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:45 +0200
                          Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-23 18:16 -0400
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:13 +0200
                  Re: Embedding a Checksum in an Image File boB <boB@K7IQ.com> - 2023-04-22 13:41 -0700
                  Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 10:34 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:58 +0200
                      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 15:24 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:17 +0200
                          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-24 01:07 -0700
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:42 +0200
              Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:20 +0200
                Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:44 +0200
        Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:52 -0700
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:23 -0700
            Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-22 07:07 -0700
              Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-22 10:31 -0400
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:54 -0700
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:26 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-27 10:09 -0700
        Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:29 +0300
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:39 +0300
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:44 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:38 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:50 +0200
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:04 +0200
                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:03 +0200
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-30 16:19 +0200
                      Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:34 +0200
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:18 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:33 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:36 +0200
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-28 01:10 +0300
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:54 +0200
      Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:24 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:56 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:09 +0200
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:02 +0200
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:27 +0200

Page 1 of 5  [1] 2 3 4 5  Next page →


#31788 — Embedding a Checksum in an Image File

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-19 19:06 -0700
SubjectEmbedding a Checksum in an Image File
Message-ID<116ff07e-5e25-469a-90a0-9474108aadd3n@googlegroups.com>
This is a bit of the chicken and egg thing.  If you want a embed a checksum in a code module to report the checksum, is there a way of doing this?  It's a bit like being your own grandfather, I think. 

I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16.  

I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way.  Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum.  

I'm not thinking of any special checksum generator that excludes the checksum data.  That would be too messy.  

I keep thinking there is a different way of looking at this to achieve the result I want... 

Maybe I can prove it is impossible.  Assume the file checksums to X when the checksum data is zero.  The goal would then be to include the checksum data value Y in the file, that would change X to Y.  Given the properties of the module N checksum, this would appear to be impossible for the general case, unless...  Add another data value, called, checksum normalizer.  This data value checksums with the original checksum to give the result zero.  Then, when the checksum is also added, the resulting checksum is, in fact, the checksum.  Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact. 

This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think.  At least, this could work in software, where data can be included in an image file as itself.  In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted.  Memory might have data that is stored as itself.  I'll need to look into that. 

-- 

  Rick C.

  - Get 1,000 miles of free Supercharging
  - Tesla referral code - https://ts.la/richard11209

[toc] | [next] | [standalone]


#31789

FromNiklas Holsti <niklas.holsti@tidorum.invalid>
Date2023-04-20 12:14 +0300
Message-ID<kace4iFke6vU1@mid.individual.net>
In reply to#31788
On 2023-04-20 5:06, Rick C wrote:
> This is a bit of the chicken and egg thing.  If you want a embed a
> checksum in a code module to report the checksum, is there a way of
> doing this?  It's a bit like being your own grandfather, I think.
> 
> I'm not thinking anything too fancy, like a CRC, but rather a simple
> modulo N addition, maybe N being 2^16.

Some decades ago I was involved with a project for an 8052-based device, 
which was required to perform a code-check-sum check at boot.

We decided to use a byte-per-byte xor checksum and make the correct 
check-sum be zero. We had a code module (possibly in assembler, I don't 
remember) that defined a one-byte "adjustment" constant in code memory. 
For each new version of the code, we first set the adjustment constant 
to zero, then ran the program, and it usually reported an error at boot 
because the check-sum was not zero. We then changed the adjustment 
constant to the actual reported checksum, C say, and that zeroed the 
check-sum because C xor C = 0. Bingo. You can use this method to make 
the checksum anything you like, for example hex 55.

With a more advanced order-sensitive check-sum such as a CRC you could 
use the same method if you also ensure (by linker commands) that the 
adjustment value is always the last value that enters in the computed 
check-sum (assuming that the linking order of the other code modules is 
not incidentally changed when the value of the adjustment constant is 
changed).

[toc] | [prev] | [next] | [standalone]


#31793

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-20 06:18 -0700
Message-ID<f4e757c5-31e2-45bd-aaef-34a944a90ebdn@googlegroups.com>
In reply to#31789
On Thursday, April 20, 2023 at 5:15:04 AM UTC-4, Niklas Holsti wrote:
> On 2023-04-20 5:06, Rick C wrote: 
> > This is a bit of the chicken and egg thing. If you want a embed a 
> > checksum in a code module to report the checksum, is there a way of 
> > doing this? It's a bit like being your own grandfather, I think. 
> > 
> > I'm not thinking anything too fancy, like a CRC, but rather a simple 
> > modulo N addition, maybe N being 2^16.
> Some decades ago I was involved with a project for an 8052-based device, 
> which was required to perform a code-check-sum check at boot. 
> 
> We decided to use a byte-per-byte xor checksum and make the correct 
> check-sum be zero. We had a code module (possibly in assembler, I don't 
> remember) that defined a one-byte "adjustment" constant in code memory. 
> For each new version of the code, we first set the adjustment constant 
> to zero, then ran the program, and it usually reported an error at boot 
> because the check-sum was not zero. We then changed the adjustment 
> constant to the actual reported checksum, C say, and that zeroed the 
> check-sum because C xor C = 0. Bingo. You can use this method to make 
> the checksum anything you like, for example hex 55. 
> 
> With a more advanced order-sensitive check-sum such as a CRC you could 
> use the same method if you also ensure (by linker commands) that the 
> adjustment value is always the last value that enters in the computed 
> check-sum (assuming that the linking order of the other code modules is 
> not incidentally changed when the value of the adjustment constant is 
> changed).

Yes, it had occurred to me that a simple checksum could be used with adjustment codes.  But I don't want the checksum to be set to some value, in this way.  I would like to embed the check sum generated from the file.  The way to do this is to embed the checksum in the spot where it can be read for reporting.  Then another value can be embedded elsewhere, that complements the checksum, keeping the file checksum constant.  

Your mention of the XOR checksum makes me realize that if I use addition, rather than XOR, a 16 bit checksum only has a complement if the data used in the calculation are 16 bit quantities.  If the 16 bit checksum is calculated using 8 bit data, there will be a carry out of the lower 8 bits changing the final checksum.  The XOR checksum is really the equivalent of 8 separate bit level checksums.  This has the short coming of one bit detection, but two bit changes in the same bit of two bytes not being detected.  But since I'm not trying to protect against changes, this isn't really a problem.  I'm using this as a verification of the version number.  

-- 

  Rick C.

  -- Get 1,000 miles of free Supercharging
  -- Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31790

From"Peter Heitzer" <peter.heitzer@rz.uni-regensburg.de>
Date2023-04-20 11:30 +0000
Message-ID<kacm2qFlk42U1@mid.individual.net>
In reply to#31788
Rick C <gnuarm.deletethisbit@gmail.com> wrote:
>This is a bit of the chicken and egg thing.  If you want a embed a checksum in a code module to report the checksum, is there a way of doing this?  It's a bit like being your own grandfather, I think. 

>I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16.  

What about putting the following structure at a fixed address at the end of
ROM?:
<startaddr><len><checksum>
Your check function then for example does a 16 bit sum of the bytes from
<startaddr>..<startaddr>+<len>-1 and compares with <checksum>
<startaddr>, <len> an <checksum> can be evaluated at compile time.


-- 
Dipl.-Inform(FH) Peter Heitzer, peter.heitzer@rz.uni-regensburg.de

[toc] | [prev] | [next] | [standalone]


#31791

Fromdalai lamah <antonio12358@hotmail.com>
Date2023-04-20 13:47 +0200
Message-ID<5oqjhwwdrcir.1fwtjz4f5xour.dlg@40tude.net>
In reply to#31788
Un bel giorno Rick C digitò:

> This is a bit of the chicken and egg thing.  If you want a embed a
> checksum in a code module to report the checksum, is there a way of
> doing this?  It's a bit like being your own grandfather, I think. 
> 
> I'm not thinking anything too fancy, like a CRC, but rather a simple
> modulo N addition, maybe N being 2^16. 
> 
> I keep thinking of using a placeholder, but that doesn't seem to work
> out in any useful way.  Even if you try to anticipate the impact of
> adding the checksum, that only gives you a different checksum, that you
> then need to anticipate further... ad infinitum. 

I'm probably not understanding what you mean, but normally the checksum is
stored in a memory section which is not subjected to the checksum
calculation itself.

The actual implementation depends on the tools you are using. Many linkers
support this directly: you specify the memory section(s) subjected to
checksum calculation, the type of checksum (CRC16, CRC32 etc) and the
memory section that will store the checksum.

Here is a technical note for IAR:
https://www.iar.com/knowledge/support/technical-notes/general/checksum-calculation-with-xlink/

A "poor man" solution is to do it manually:

-In the source code, declare your checksum initializing to a known, fixed
value (e.g. 0xDEADBEEF)
-Run the program with a debugger; set a breakpoint when it calculates the
checksum (and fails), and write down the correct checksum
-Using a binary editor, find the fixed value into the executable binary,
and replace it with the correct value.

-- 
Fletto i muscoli e sono nel vuoto.

[toc] | [prev] | [next] | [standalone]


#31792

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-20 06:04 -0700
Message-ID<ab28a285-2a16-4f43-840b-d0b749cdbc6fn@googlegroups.com>
In reply to#31791
On Thursday, April 20, 2023 at 7:47:54 AM UTC-4, dalai lamah wrote:
> Un bel giorno Rick C digitò:
> > This is a bit of the chicken and egg thing. If you want a embed a 
> > checksum in a code module to report the checksum, is there a way of 
> > doing this? It's a bit like being your own grandfather, I think. 
> > 
> > I'm not thinking anything too fancy, like a CRC, but rather a simple 
> > modulo N addition, maybe N being 2^16. 
> > 
> > I keep thinking of using a placeholder, but that doesn't seem to work 
> > out in any useful way. Even if you try to anticipate the impact of 
> > adding the checksum, that only gives you a different checksum, that you 
> > then need to anticipate further... ad infinitum.
> I'm probably not understanding what you mean, but normally the checksum is 
> stored in a memory section which is not subjected to the checksum 
> calculation itself. 

Yes, I didn't explain it clearly.  I am not looking for a way to calculate the checksum from a processor.  That would be trivial.  I want to embed the checksum in the code, so that it can be provided at run time as an ID, a way to validate the version number.  


> The actual implementation depends on the tools you are using. Many linkers 
> support this directly: you specify the memory section(s) subjected to 
> checksum calculation, the type of checksum (CRC16, CRC32 etc) and the 
> memory section that will store the checksum. 

I wish to perform this checksum on the executable file.  


> Here is a technical note for IAR: 
> https://www.iar.com/knowledge/support/technical-notes/general/checksum-calculation-with-xlink/ 
> 
> A "poor man" solution is to do it manually: 
> 
> -In the source code, declare your checksum initializing to a known, fixed 
> value (e.g. 0xDEADBEEF) 
> -Run the program with a debugger; set a breakpoint when it calculates the 
> checksum (and fails), and write down the correct checksum 
> -Using a binary editor, find the fixed value into the executable binary, 
> and replace it with the correct value. 

Yeah, this is not useful, because changing the value stored changes the checksum. It also makes assumptions about the target. 

Maybe this was not the best group to ask the question in.  I thought this was more of a math problem with I started writing the question and the embedded community had already dealt with it.  

-- 

  Rick C.

  + Get 1,000 miles of free Supercharging
  + Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31794

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-20 16:46 +0200
Message-ID<u1rj8b$kpek$1@dont-email.me>
In reply to#31788
On 20/04/2023 04:06, Rick C wrote:
> This is a bit of the chicken and egg thing.  If you want a embed a
> checksum in a code module to report the checksum, is there a way of
> doing this?  It's a bit like being your own grandfather, I think.
> 
> I'm not thinking anything too fancy, like a CRC, but rather a simple
> modulo N addition, maybe N being 2^16.
> 
> I keep thinking of using a placeholder, but that doesn't seem to work
> out in any useful way.  Even if you try to anticipate the impact of
> adding the checksum, that only gives you a different checksum, that
> you then need to anticipate further... ad infinitum.
> 
> I'm not thinking of any special checksum generator that excludes the
> checksum data.  That would be too messy.
> 
> I keep thinking there is a different way of looking at this to
> achieve the result I want...
> 
> Maybe I can prove it is impossible.  Assume the file checksums to X
> when the checksum data is zero.  The goal would then be to include
> the checksum data value Y in the file, that would change X to Y.
> Given the properties of the module N checksum, this would appear to
> be impossible for the general case, unless...  Add another data
> value, called, checksum normalizer.  This data value checksums with
> the original checksum to give the result zero.  Then, when the
> checksum is also added, the resulting checksum is, in fact, the
> checksum.  Another way of looking at this is to add a value that
> combines with the added checksum, to be zero, leaving the original
> checksum intact.
> 
> This might be inordinately hard for a CRC, but a simple checksum
> would not be an issue, I think.  At least, this could work in
> software, where data can be included in an image file as itself.  In
> a device like an FPGA, it might not be included in the bit stream
> file so directly... but that might depend on where in the device it
> is inserted.  Memory might have data that is stored as itself.  I'll
> need to look into that.
> 


I am not sure what your intended use-case is here.  But it is very 
common to add a checksum of some sort to binary image files after 
generating them.  This is done post-link.  You have a struct in your 
read-only data that you link at a known fixed point in the binary.  Your 
post-link patcher can read this struct (for example, to get the program 
version number that is then used to rename the final image file).  It 
can modify the struct (such as inserting the length of the image).  Then 
it calculates a CRC and appends it to the end of the image.

[toc] | [prev] | [next] | [standalone]


#31795

FromGeorge Neuner <gneuner2@comcast.net>
Date2023-04-20 11:33 -0400
Message-ID<pvl24i57aef4vc7bbdk9mvj7sic9dsh64t@4ax.com>
In reply to#31788
On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C
<gnuarm.deletethisbit@gmail.com> wrote:

>This is a bit of the chicken and egg thing.  If you want a embed a
>checksum in a code module to report the checksum, is there a way of
>doing this?  It's a bit like being your own grandfather, I think. 

Take a look at the old xmodem/ymodem CRC.  It was designed such that
when the CRC was sent immediately following the data, a receiver
computing CRC over the whole incoming packet (data and CRC both) would
get a result of zero.

But AFAIK it doesn't work with CCITT equation(s) - you have to use
xmodem/ymodem.


>I'm not thinking anything too fancy, like a CRC, but rather a simple
>modulo N addition, maybe N being 2^16.  

Sorry, I don't know a way to do it with a modular checksum.
YMMV, but I think 16-bit CRC is pretty simple.

George

[toc] | [prev] | [next] | [standalone]


#31796

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-20 09:45 -0700
Message-ID<f0afa198-e735-4da1-a16a-82764af3de4dn@googlegroups.com>
In reply to#31795
On Thursday, April 20, 2023 at 11:33:28 AM UTC-4, George Neuner wrote:
> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
> <gnuarm.del...@gmail.com> wrote: 
> 
> >This is a bit of the chicken and egg thing. If you want a embed a 
> >checksum in a code module to report the checksum, is there a way of 
> >doing this? It's a bit like being your own grandfather, I think.
> Take a look at the old xmodem/ymodem CRC. It was designed such that 
> when the CRC was sent immediately following the data, a receiver 
> computing CRC over the whole incoming packet (data and CRC both) would 
> get a result of zero. 
> 
> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
> xmodem/ymodem.
> >I'm not thinking anything too fancy, like a CRC, but rather a simple 
> >modulo N addition, maybe N being 2^16.
> Sorry, I don't know a way to do it with a modular checksum. 
> YMMV, but I think 16-bit CRC is pretty simple. 
> 
> George

CRC is not complicated, but I would not know how to calculate an inserted value to force the resulting CRC to zero.  How do you do that? 

Even so, I'm not trying to validate the file.  I'm trying to come up with a substitute for a time stamp or version number.  I don't want to have to rely on my consistency in handling the version number correctly.  This would be a backup in case there was more than one version released, even only within the "lab", that were different.  A checksum that could be read by the controlling software would do the job.  

I have run into this before, where the version number was not a 100% indication of the uniqueness of an executable.  The checksum would be a second indicator. 

I should mention that I'm not looking for a solution that relies on any specific details of the tools. 

-- 

  Rick C.

  -+ Get 1,000 miles of free Supercharging
  -+ Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31798

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-20 22:26 +0200
Message-ID<u1s76b$o11h$1@dont-email.me>
In reply to#31796
On 20/04/2023 18:45, Rick C wrote:
> On Thursday, April 20, 2023 at 11:33:28 AM UTC-4, George Neuner
> wrote:
>> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
>> <gnuarm.del...@gmail.com> wrote:
>> 
>>> This is a bit of the chicken and egg thing. If you want a embed
>>> a checksum in a code module to report the checksum, is there a
>>> way of doing this? It's a bit like being your own grandfather, I
>>> think.
>> Take a look at the old xmodem/ymodem CRC. It was designed such
>> that when the CRC was sent immediately following the data, a
>> receiver computing CRC over the whole incoming packet (data and CRC
>> both) would get a result of zero.
>> 
>> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
>> xmodem/ymodem.
>>> I'm not thinking anything too fancy, like a CRC, but rather a
>>> simple modulo N addition, maybe N being 2^16.
>> Sorry, I don't know a way to do it with a modular checksum. YMMV,
>> but I think 16-bit CRC is pretty simple.
>> 
>> George
> 
> CRC is not complicated, but I would not know how to calculate an
> inserted value to force the resulting CRC to zero.  How do you do
> that?

You "insert" the value at the end.  Anything else is insane.

CRC's are quite good hashes, for suitable sized data.  There are perhaps 
some special cases, but basically you'd be doing trial-and-error 
searches to find an inserted value that gives you a zero CRC overall. 
2^16 is not an overwhelming search space, but the whole idea is pointless.

> 
> Even so, I'm not trying to validate the file.  I'm trying to come up
> with a substitute for a time stamp or version number.  I don't want
> to have to rely on my consistency in handling the version number
> correctly.  This would be a backup in case there was more than one
> version released, even only within the "lab", that were different.  A
> checksum that could be read by the controlling software would do the
> job.

A CRC is fine for that.

> 
> I have run into this before, where the version number was not a 100%
> indication of the uniqueness of an executable.  The checksum would be
> a second indicator.
> 
> I should mention that I'm not looking for a solution that relies on
> any specific details of the tools.
> 

A table-based CRC is easy, runs quickly, and can be quickly ported to 
pretty much any language (the C and Python code, for example, is almost 
the same).

[toc] | [prev] | [next] | [standalone]


#31857

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-04-27 18:36 +0200
Message-ID<u2e8ba$205k5$3@dont-email.me>
In reply to#31798
Den 2023-04-20 kl. 22:26, skrev David Brown:
> On 20/04/2023 18:45, Rick C wrote:
>> On Thursday, April 20, 2023 at 11:33:28 AM UTC-4, George Neuner
>> wrote:
>>> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
>>> <gnuarm.del...@gmail.com> wrote:
>>>
>>>> This is a bit of the chicken and egg thing. If you want a embed
>>>> a checksum in a code module to report the checksum, is there a
>>>> way of doing this? It's a bit like being your own grandfather, I
>>>> think.
>>> Take a look at the old xmodem/ymodem CRC. It was designed such
>>> that when the CRC was sent immediately following the data, a
>>> receiver computing CRC over the whole incoming packet (data and CRC
>>> both) would get a result of zero.
>>>
>>> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
>>> xmodem/ymodem.
>>>> I'm not thinking anything too fancy, like a CRC, but rather a
>>>> simple modulo N addition, maybe N being 2^16.
>>> Sorry, I don't know a way to do it with a modular checksum. YMMV,
>>> but I think 16-bit CRC is pretty simple.
>>>
>>> George
>>
>> CRC is not complicated, but I would not know how to calculate an
>> inserted value to force the resulting CRC to zero.  How do you do
>> that?
> 
> You "insert" the value at the end.  Anything else is insane.

In all projects I have been involved with, the application binary starts
with a header looking like this.


MAGIC WORD 1
CRC
Entry Point
Size
other info...
MAGIC WORD 2
APPLICATION_START
...
APPLICATION_END (aligned with flash sector)


The bootloader first checks the two magic words.
It then computes CRC on the header (from Entry Point) to APPLICATION_END

I ported the IAR ielftool (open source) to Linux at
https://github.com/emagii/ielftool

This can insert the CRC in the ELF file, but needs tweaks to work
with an ELF file generated by the GNU tools.

/Ulf


> 
> CRC's are quite good hashes, for suitable sized data.  There are perhaps 
> some special cases, but basically you'd be doing trial-and-error 
> searches to find an inserted value that gives you a zero CRC overall. 
> 2^16 is not an overwhelming search space, but the whole idea is pointless.
> 
>>
>> Even so, I'm not trying to validate the file.  I'm trying to come up
>> with a substitute for a time stamp or version number.  I don't want
>> to have to rely on my consistency in handling the version number
>> correctly.  This would be a backup in case there was more than one
>> version released, even only within the "lab", that were different.  A
>> checksum that could be read by the controlling software would do the
>> job.
> 
> A CRC is fine for that.
> 
>>
>> I have run into this before, where the version number was not a 100%
>> indication of the uniqueness of an executable.  The checksum would be
>> a second indicator.
>>
>> I should mention that I'm not looking for a solution that relies on
>> any specific details of the tools.
>>
> 
> A table-based CRC is easy, runs quickly, and can be quickly ported to 
> pretty much any language (the C and Python code, for example, is almost 
> the same).

[toc] | [prev] | [next] | [standalone]


#31871

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-28 09:12 +0200
Message-ID<u2frll$2bncf$1@dont-email.me>
In reply to#31857
On 27/04/2023 18:36, Ulf Samuelsson wrote:
> Den 2023-04-20 kl. 22:26, skrev David Brown:
>> On 20/04/2023 18:45, Rick C wrote:
>>> On Thursday, April 20, 2023 at 11:33:28 AM UTC-4, George Neuner
>>> wrote:
>>>> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
>>>> <gnuarm.del...@gmail.com> wrote:
>>>>
>>>>> This is a bit of the chicken and egg thing. If you want a embed
>>>>> a checksum in a code module to report the checksum, is there a
>>>>> way of doing this? It's a bit like being your own grandfather, I
>>>>> think.
>>>> Take a look at the old xmodem/ymodem CRC. It was designed such
>>>> that when the CRC was sent immediately following the data, a
>>>> receiver computing CRC over the whole incoming packet (data and CRC
>>>> both) would get a result of zero.
>>>>
>>>> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
>>>> xmodem/ymodem.
>>>>> I'm not thinking anything too fancy, like a CRC, but rather a
>>>>> simple modulo N addition, maybe N being 2^16.
>>>> Sorry, I don't know a way to do it with a modular checksum. YMMV,
>>>> but I think 16-bit CRC is pretty simple.
>>>>
>>>> George
>>>
>>> CRC is not complicated, but I would not know how to calculate an
>>> inserted value to force the resulting CRC to zero.  How do you do
>>> that?
>>
>> You "insert" the value at the end.  Anything else is insane.
> 
> In all projects I have been involved with, the application binary starts
> with a header looking like this.
> 
> 
> MAGIC WORD 1
> CRC
> Entry Point
> Size
> other info...
> MAGIC WORD 2
> APPLICATION_START
> ...
> APPLICATION_END (aligned with flash sector)
> 
> 
> The bootloader first checks the two magic words.
> It then computes CRC on the header (from Entry Point) to APPLICATION_END
> 
> I ported the IAR ielftool (open source) to Linux at
> https://github.com/emagii/ielftool
> 
> This can insert the CRC in the ELF file, but needs tweaks to work
> with an ELF file generated by the GNU tools.
> 
> /Ulf

That can work for some microcontrollers, but is unsuitable for others - 
it depends on how the flash is organised.  For an msp430, for example, 
it would be fine, as the interrupt vectors (including the reset vector) 
are at the end of flash.  But for most ARM Cortex M devices, it would 
not be suitable - they expect the reset vector and initial stack pointer 
at the start of the flash image.  Some devices have a boot ROM, and then 
you have to match their specifics for the header - or you can have your 
own boot program, and make the header how ever you like.

I am absolutely a fan of having some kind of header like this (and 
sometimes even a human-readable copyright notice, identifier and version 
information).  And having it as near the beginning as possible is good. 
But for many microcontrollers, having it at the start is not feasible. 
And if you can't put the CRC at the start like you do, you have to put 
it at the end of the image.


I've never really thought about trying to inject a CRC into an elf file. 
  I use elfs (or should that be "elves" ?) for debugging, not flash 
programming.  And usually the main concern for having a CRC at the end 
of the image is when you have an online update of some kind, to check 
that nothing has gone wrong during the transfer or in-field update.

[toc] | [prev] | [next] | [standalone]


#31878

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-04-28 10:35 +0200
Message-ID<u2g0gc$2cdfn$1@dont-email.me>
In reply to#31871
Den 2023-04-28 kl. 09:12, skrev David Brown:
> On 27/04/2023 18:36, Ulf Samuelsson wrote:
>> Den 2023-04-20 kl. 22:26, skrev David Brown:
>>> On 20/04/2023 18:45, Rick C wrote:
>>>> On Thursday, April 20, 2023 at 11:33:28 AM UTC-4, George Neuner
>>>> wrote:
>>>>> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
>>>>> <gnuarm.del...@gmail.com> wrote:
>>>>>
>>>>>> This is a bit of the chicken and egg thing. If you want a embed
>>>>>> a checksum in a code module to report the checksum, is there a
>>>>>> way of doing this? It's a bit like being your own grandfather, I
>>>>>> think.
>>>>> Take a look at the old xmodem/ymodem CRC. It was designed such
>>>>> that when the CRC was sent immediately following the data, a
>>>>> receiver computing CRC over the whole incoming packet (data and CRC
>>>>> both) would get a result of zero.
>>>>>
>>>>> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
>>>>> xmodem/ymodem.
>>>>>> I'm not thinking anything too fancy, like a CRC, but rather a
>>>>>> simple modulo N addition, maybe N being 2^16.
>>>>> Sorry, I don't know a way to do it with a modular checksum. YMMV,
>>>>> but I think 16-bit CRC is pretty simple.
>>>>>
>>>>> George
>>>>
>>>> CRC is not complicated, but I would not know how to calculate an
>>>> inserted value to force the resulting CRC to zero.  How do you do
>>>> that?
>>>
>>> You "insert" the value at the end.  Anything else is insane.
>>
>> In all projects I have been involved with, the application binary starts
>> with a header looking like this.
>>
>>
>> MAGIC WORD 1
>> CRC
>> Entry Point
>> Size
>> other info...
>> MAGIC WORD 2
>> APPLICATION_START
>> ...
>> APPLICATION_END (aligned with flash sector)
>>
>>
>> The bootloader first checks the two magic words.
>> It then computes CRC on the header (from Entry Point) to APPLICATION_END
>>
>> I ported the IAR ielftool (open source) to Linux at
>> https://github.com/emagii/ielftool
>>
>> This can insert the CRC in the ELF file, but needs tweaks to work
>> with an ELF file generated by the GNU tools.
>>
>> /Ulf
> 
> That can work for some microcontrollers, but is unsuitable for others - 
> it depends on how the flash is organised.  For an msp430, for example, 
> it would be fine, as the interrupt vectors (including the reset vector) 
> are at the end of flash.  But for most ARM Cortex M devices, it would 
> not be suitable - they expect the reset vector and initial stack pointer 
> at the start of the flash image.  Some devices have a boot ROM, and then 
> you have to match their specifics for the header - or you can have your 
> own boot program, and make the header how ever you like.


All projects I am involved with have a custom bootloader.
If there is a problem with the reset vector, then the program will fail 
immediately.
The CRC is right after the initial vector table.
The bootloader application contains a copy of the vector table.

THe first thing the bootloader does is to check the CRC from right after 
the CRC. Then it compares the vector table with the copy.

The header is only for the application.

> 
> I am absolutely a fan of having some kind of header like this (and 
> sometimes even a human-readable copyright notice, identifier and version 
> information).  And having it as near the beginning as possible is good. 
> But for many microcontrollers, having it at the start is not feasible. 
> And if you can't put the CRC at the start like you do, you have to put 
> it at the end of the image.
> 
> 
> I've never really thought about trying to inject a CRC into an elf file. 
>   I use elfs (or should that be "elves" ?) for debugging, not flash 
> programming.  And usually the main concern for having a CRC at the end 
> of the image is when you have an online update of some kind, to check 
> that nothing has gone wrong during the transfer or in-field update.
> 
> 

The last bootloader I wrote download using Y-Modem which has CRC 
checking. Since it had more RAM than internal flash, the whole 
application was downloaded to RAM first, and then when everything is OK,
the flash can be programmed. Finally, the header is analyzed and the 
flash contents checked. There is absolutely no need to have the CRC at 
the end since the CRC result is stored in a known location.

/Ulf

[toc] | [prev] | [next] | [standalone]


#31799

FromGeorge Neuner <gneuner2@comcast.net>
Date2023-04-20 16:44 -0400
Message-ID<36534il81ipvnhog6980r9ln9tdqn5cbh6@4ax.com>
In reply to#31796
On Thu, 20 Apr 2023 09:45:59 -0700 (PDT), Rick C
<gnuarm.deletethisbit@gmail.com> wrote:

>On Thursday, April 20, 2023 at 11:33:28?AM UTC-4, George Neuner wrote:
>> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C 
>> <gnuarm.del...@gmail.com> wrote: 
>> 
>> >This is a bit of the chicken and egg thing. If you want a embed a 
>> >checksum in a code module to report the checksum, is there a way of 
>> >doing this? It's a bit like being your own grandfather, I think.
>>
>> Take a look at the old xmodem/ymodem CRC. It was designed such that 
>> when the CRC was sent immediately following the data, a receiver 
>> computing CRC over the whole incoming packet (data and CRC both) would 
>> get a result of zero. 
>> 
>> But AFAIK it doesn't work with CCITT equation(s) - you have to use 
>> xmodem/ymodem.
>> >I'm not thinking anything too fancy, like a CRC, but rather a simple 
>> >modulo N addition, maybe N being 2^16.
>> Sorry, I don't know a way to do it with a modular checksum. 
>> YMMV, but I think 16-bit CRC is pretty simple. 
>> 
>> George
>
>CRC is not complicated, but I would not know how to calculate an
>inserted value to force the resulting CRC to zero.  How do you do
>that? 

It's implicit in the equation they chose.  I don't know how it works -
just that it does.


You have some block of data   |....data....|

You compute CRC on the data block and then append the resulting value
to the end of the block.  xmodem CRC is 16-bit, so it adds 2 bytes to
the data.

So now you have a new extended block   |....data....|crc|

Now if you compute a new CRC on the extended block, the resulting
value /should/ come out to zero. If it doesn't, either your data or
the original CRC value appended to it has been changed/corrupted.


>Even so, I'm not trying to validate the file.  I'm trying to come up
>with a substitute for a time stamp or version number.  I don't want
>to have to rely on my consistency in handling the version number
>correctly.  This would be a backup in case there was more than one
>version released, even only within the "lab", that were different.  A
>checksum that could be read by the controlling software would do the
>job.  

I've actually done this: in the early 90s I designed a system that
used a CRC based scheme to identify load modules and track
inter-module code dependencies.  

I computed both 16-bit Xmodem and CCITT CRCs on the modules and
concatenated the two values into a 32-bit identifier.  That identifier
then was used to sign the module and to demand load (or unload) it
when needed.

At the time it worked quite well: the system had quite limited memory,
so code modules were small enough that even a 16-bit CRC could
uniquely identify most/all of them. Combining the two different CRCs
into a 32-bit identifier provided more than enough uniqueness, it was
fast and easy to compute, and it saved a lot of space vs using
something with stronger guarantees like a UUID or crypto-strength
signing hash.
[A lot of the hashing functions available today either didn't exist or
just weren't widely known back then.  And still most of them that even
have 32-bit variants are weak in guarantees for those variants.]


>I have run into this before, where the version number was not a 100%
>indication of the uniqueness of an executable.  The checksum would be
>a second indicator. 

I made it the basis of dependency checking. Version numbers were
secondary and for the benefit of the programmer.


>I should mention that I'm not looking for a solution that relies on
>any specific details of the tools. 

YMMV.
George

[toc] | [prev] | [next] | [standalone]


#31803

FromDon Y <blockedofcourse@foo.invalid>
Date2023-04-20 22:37 -0700
Message-ID<u1t7eb$10gmu$3@dont-email.me>
In reply to#31799
On 4/20/2023 1:44 PM, George Neuner wrote:
> You have some block of data   |....data....|
> 
> You compute CRC on the data block and then append the resulting value
---------------------------------------------^^^^^^
> to the end of the block.  xmodem CRC is 16-bit, so it adds 2 bytes to
> the data.

Exactly.  You *don't* drag the "extra bits" into the initial
CRC calculation but *do* into the CRC *verification*.  Easy
peasy (since forever).

[Think about it:  your performing a division operation
and the residual is the "remainder".]

Note that you want to choose a polynomial that doesn't
give you a "win" result for "obviously" corrupt data.
E.g., if data is all zeros or all 0xFF (as these sorts of
conditions can happen with hardware failures) you probably
wouldn't want a "success" indication!

You can also "salt" the calculation so that the residual
is deliberately nonzero.  So, for example, "success" is
indicated by a residual of 0x474E.  :>

> So now you have a new extended block   |....data....|crc|
> 
> Now if you compute a new CRC on the extended block, the resulting
> value /should/ come out to zero. If it doesn't, either your data or
> the original CRC value appended to it has been changed/corrupted.

As there is usually a lack of originality in the algorithms
chosen, you have to consider if you are also hoping to use
this to safeguard the *integrity* of your image (i.e.,
against intentional modification).

I have an old Compaq Portable 386 (lunchbox) that obviously
wasn't designed to support the disk drives that I would
*later* install in it.  So, I patched the BIOS ROMs to
add another disk type to the Disk Parameter Table.  Then,
made compensating changes to other parts of the ROM
(that I knew would not be referenced) to ensure the original
checksum -- WHEREVER IT MAY HAVE BEEN "STORED" -- would remain
intact.

[I could similarly have altered the boot message to show
a copyright of "Don Y" in place of "Compaq" -- as it would
be pretty easy to locate the "Compaq" string (in plaintext)
in the image.]

[toc] | [prev] | [next] | [standalone]


#31805

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-21 12:43 +0200
Message-ID<u1tpcr$1377f$1@dont-email.me>
In reply to#31803
On 21/04/2023 07:37, Don Y wrote:
> On 4/20/2023 1:44 PM, George Neuner wrote:
>> You have some block of data   |....data....|
>>
>> You compute CRC on the data block and then append the resulting value
> ---------------------------------------------^^^^^^
>> to the end of the block.  xmodem CRC is 16-bit, so it adds 2 bytes to
>> the data.
> 
> Exactly.  You *don't* drag the "extra bits" into the initial
> CRC calculation but *do* into the CRC *verification*.  Easy
> peasy (since forever).
> 
> [Think about it:  your performing a division operation
> and the residual is the "remainder".]

George's earlier posts made it look like the algorithm was inserting 
("embedding") a value somewhere inside the image, so that the CRC over 
the modified image was zero.  This is easy to do for simple checksums 
such as XOR's or a sum-of-bytes checksum, but infeasible for CRC's.

It is a much easier matter when appending the checksum.  Depending 
somewhat on the details of the CRC (such as bit/byte reversals, 
inversions, starting values, etc.) it is typically the case that for a 
binary blob A, crc(A ++ crc(A)) = 0.  i.e., if you append the CRC of 
your data to the data, the CRC of the whole thing is 0.

Of course, this is pretty much irrelevant - whether you check the 
integrity of the final image by running CRC over it all and comparing to 
0, or running it over all but the last word and comparing to the last 
word is a minor matter.

> 
> Note that you want to choose a polynomial that doesn't
> give you a "win" result for "obviously" corrupt data.
> E.g., if data is all zeros or all 0xFF (as these sorts of
> conditions can happen with hardware failures) you probably
> wouldn't want a "success" indication!

No, that is pointless for something like a code image.  It just adds 
needless complexity to your CRC algorithm.

You should already have checks that would eliminate an all-zero image or 
other "obviously corrupt" data.  You'll be checking the image for a key 
or "magic number" that identifies the image as "program image for board 
X, project Y".  You'll be checking version numbers.  You'll be reading 
the length of the image so you know the range for your CRC function, and 
where to find the appended CRC check.  You might not have all of these 
in a given system, but you'll have some kind of check which would fail 
on an all-zero image.

> 
> You can also "salt" the calculation so that the residual
> is deliberately nonzero.  So, for example, "success" is
> indicated by a residual of 0x474E.  :>
> 

Again, pointless.

Salt is important for security-related hashes (like password hashes), 
not for integrity checks.

>> So now you have a new extended block   |....data....|crc|
>>
>> Now if you compute a new CRC on the extended block, the resulting
>> value /should/ come out to zero. If it doesn't, either your data or
>> the original CRC value appended to it has been changed/corrupted.
> 
> As there is usually a lack of originality in the algorithms
> chosen, you have to consider if you are also hoping to use
> this to safeguard the *integrity* of your image (i.e.,
> against intentional modification).
> 

"Integrity" has nothing to do with the motivation for change. 
/Security/ is concerned with intentional modifications that deliberately 
attempt to defeat /integrity/ checks.  Integrity is about detecting any 
changes.

If you are concerned about the possibility of intentional malicious 
changes, CRC's alone are useless.  All the attacker needs to do after 
modifying the image is calculate the CRC themselves, and replace the 
original checksum with their own.

Using non-standard algorithms for security is a simple way to get things 
completely wrong.  "Security by obscurity" is very rarely the right 
answer.  In reality, good security algorithms, and good implementations, 
are difficult and specialised tasks, best left to people who know what 
they are doing.

To make something secure, you have to ensure that the check algorithms 
depend on a key that you know, but that the attacker does not have. 
That's the basis of digital signatures (though you use a secure hash 
algorithm rather than a simple CRC).


[toc] | [prev] | [next] | [standalone]


#31806

FromDon Y <blockedofcourse@foo.invalid>
Date2023-04-21 04:39 -0700
Message-ID<u1tsm3$13ook$1@dont-email.me>
In reply to#31805
On 4/21/2023 3:43 AM, David Brown wrote:
>> Note that you want to choose a polynomial that doesn't
>> give you a "win" result for "obviously" corrupt data.
>> E.g., if data is all zeros or all 0xFF (as these sorts of
>> conditions can happen with hardware failures) you probably
>> wouldn't want a "success" indication!
> 
> No, that is pointless for something like a code image.  It just adds needless 
> complexity to your CRC algorithm.

Perhaps you've forgotten that you don't just use CRCs (secure hashes, etc.)
on "code images"?

> You should already have checks that would eliminate an all-zero image or other 
> "obviously corrupt" data.  You'll be checking the image for a key or "magic 
> number" that identifies the image as "program image for board X, project Y".  
> You'll be checking version numbers.  You'll be reading the length of the image 
> so you know the range for your CRC function, and where to find the appended CRC 
> check.  You might not have all of these in a given system, but you'll have some 
> kind of check which would fail on an all-zero image.

See above.

>> You can also "salt" the calculation so that the residual
>> is deliberately nonzero.  So, for example, "success" is
>> indicated by a residual of 0x474E.  :>
> 
> Again, pointless.
> 
> Salt is important for security-related hashes (like password hashes), not for 
> integrity checks.

You've missed the point.  The correct "sum" can be anything.
Why is "0" more special than any other value?  As the value is
typically meaningless to anything other than the code that verifies
it, you couldn't look at an image (or the output of the verifier)
and gain anything from seeing that obscure value.

OTOH, if the CRC yields something familiar -- or useful -- then
it can tell you something about the image.  E.g., salt the algorithm
with the product code, version number, your initials, 0xDEADBEEF, etc.

>>> So now you have a new extended block   |....data....|crc|
>>>
>>> Now if you compute a new CRC on the extended block, the resulting
>>> value /should/ come out to zero. If it doesn't, either your data or
>>> the original CRC value appended to it has been changed/corrupted.
>>
>> As there is usually a lack of originality in the algorithms
>> chosen, you have to consider if you are also hoping to use
>> this to safeguard the *integrity* of your image (i.e.,
>> against intentional modification).
> 
> "Integrity" has nothing to do with the motivation for change. /Security/ is 
> concerned with intentional modifications that deliberately attempt to defeat 
> /integrity/ checks.  Integrity is about detecting any changes.
> 
> If you are concerned about the possibility of intentional malicious changes, 

Changes don't have to be malicious.  I altered the test procedure for a
piece of military gear we were building simply to skip some lengthy tests that 
I *knew* would pass (I don't want to inject an extra 20 minutes of wait time
just to get through a lengthy test I already know works before I can get
to the test of interest to me, now.

I failed to undo the change before the official signoff on the device.

The only evidence of this was the fact that I had also patched the
startup message to say "Go for coffee..." -- which remained on the
screen for the duration of the lengthy (even with the long test
elided) procedure...

..which alerted folks to the fact that this *probably* wasn't the
original image.  (The computer running the test suite on the DUT had
no problem accepting my patched binary)

> CRC's alone are useless.  All the attacker needs to do after modifying the 
> image is calculate the CRC themselves, and replace the original checksum with 
> their own.

That assumes the "alterer" knows how to replace the checksum, how it
is computed, where it is embedded in the image, etc.  I modified the Compaq
portable mentioned without ever knowing where the checksum was store
or *if* it was explicitly stored.  I had no desire to disassemble the
BIOS ROMs (though could obviously do so as there was no "proprietary
hardware" limiting access to their contents and the instruction set of
the processor is well known!).

Instead, I did this by *guessing* how they would implement such a check
in a bit of kit from that era (ERPOMs aren't easily modified by malware
so it wasn't likely that they would go to great lengths to "protect" the
image).  And, if my guess had been incorrect, I could always reinstall
the original EPROMs -- nothing lost, nothing gained.

Had much experience with folks counterfeiting your products and making
"simple" changes to the binaries?  Like changing the copyright notice
or splash screen?

Then, bringing the (accused) counterfeit of YOUR product into a courtroom
and revealing the *hidden* checksum that the counterfeiter wasn't aware of?

"Gee, why does YOUR (alleged) device have *my* name in it -- in addition
to behaving exactly like mine??"

[I guess obscurity has its place!]

Use a non-secret approach and you invite folks to alter it, as well.

> Using non-standard algorithms for security is a simple way to get things 
> completely wrong.  "Security by obscurity" is very rarely the right answer.  In 
> reality, good security algorithms, and good implementations, are difficult and 
> specialised tasks, best left to people who know what they are doing.
> 
> To make something secure, you have to ensure that the check algorithms depend 
> on a key that you know, but that the attacker does not have. That's the basis 
> of digital signatures (though you use a secure hash algorithm rather than a 
> simple CRC).

If you can remove the check, then what value the key's secrecy?  By your
criteria, the adversary KNOWS how you are implementing your security
so he knows exactly what to remove to bypass your checks and allow his
altered image to operate in its place.

Ever notice how manufacturers don't PUBLICLY disclose their security
hooks (without an NDA)?  If "security by obscurity" was not important,
they would publish these details INVITING challenges (instead of
trying to limit the knowledge to people with whom they've officially
contracted).

[If it was so good and they were trying to rely on trade secret, why
not just PATENT their approach, also disclosing it in the process?
Surely, the details will "leak" from one of the NDA signers long
before patent protection would expire...  And, presumably, these
are "people who know what they are doing"...]

Sign all the binaries and all I have to do is remove the *test* for
those signatures and the images can be as corrupted as I choose.

You need to "secure" the test if you want the image to be securable.
This is why it is so hard to use "open" security protocols on
hardware devices (cuz there are almost always ways to subvert the
verification process/hardware).  Having physical access to a device
usually means it can be compromised -- if worth your effort.

[The trick is to make the effort great enough to be on a par with
just copying the *functionality*, from scratch, and not bothering
trying to alter the executable in a way that is not detectable]

[[There are companies who's business models are exactly that -- cloning
other products (e.g., from folks like big blue) at the functional
level -- yet steering clear of any copyright issues.]]

[toc] | [prev] | [next] | [standalone]


#31808

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-21 16:50 +0200
Message-ID<u1u7ro$2poss$1@dont-email.me>
In reply to#31806
On 21/04/2023 13:39, Don Y wrote:
> On 4/21/2023 3:43 AM, David Brown wrote:
>>> Note that you want to choose a polynomial that doesn't
>>> give you a "win" result for "obviously" corrupt data.
>>> E.g., if data is all zeros or all 0xFF (as these sorts of
>>> conditions can happen with hardware failures) you probably
>>> wouldn't want a "success" indication!
>>
>> No, that is pointless for something like a code image.  It just adds 
>> needless complexity to your CRC algorithm.
> 
> Perhaps you've forgotten that you don't just use CRCs (secure hashes, etc.)
> on "code images"?

No - but "code images" is the topic here.

However, in almost every case where CRC's might be useful, you have 
additional checks of the sanity of the data, and an all-zero or all-one 
data block would be rejected.  For example, Ethernet packets use CRC for 
integrity checking, but an attempt to send a packet type 0 from MAC 
address 00:00:00:00:00:00 to address 00:00:00:00:00:00, of length 0, 
would be rejected anyway.

I can't think of any use-cases where you would be passing around a block 
of "pure" data that could reasonably take absolutely any value, without 
any type of "envelope" information, and where you would think a CRC 
check is appropriate.

> 
>> You should already have checks that would eliminate an all-zero image 
>> or other "obviously corrupt" data.  You'll be checking the image for a 
>> key or "magic number" that identifies the image as "program image for 
>> board X, project Y". You'll be checking version numbers.  You'll be 
>> reading the length of the image so you know the range for your CRC 
>> function, and where to find the appended CRC check.  You might not 
>> have all of these in a given system, but you'll have some kind of 
>> check which would fail on an all-zero image.
> 
> See above.

See above.

> 
>>> You can also "salt" the calculation so that the residual
>>> is deliberately nonzero.  So, for example, "success" is
>>> indicated by a residual of 0x474E.  :>
>>
>> Again, pointless.
>>
>> Salt is important for security-related hashes (like password hashes), 
>> not for integrity checks.
> 
> You've missed the point.  The correct "sum" can be anything.
> Why is "0" more special than any other value?  As the value is
> typically meaningless to anything other than the code that verifies
> it, you couldn't look at an image (or the output of the verifier)
> and gain anything from seeing that obscure value.

Do you actually know what is meant by "salt" in the context of hashes, 
and why it is useful in some circumstances?  Do you understand that 
"salt" is added (usually prepended, or occasionally mixed in in some 
other way) to the data /before/ the hash is calculated?

I have not given the slightest indication to suggest that "0" is a 
special value.  I fully agree that the value you get from the checking 
algorithm does not have to be 0 - I already suggested it could be 
compared to the stored value.  I.e., your build your image file as "data 
++ crc(data)", at check it by re-calculating "crc(data)" on the received 
image and comparing the result to the received crc.  There is no 
necessity or benefit in having a crc run calculated over the received 
data plus the received crc being 0.

"Salt" is used in cases where the original data must be kept secret, and 
only the hashes are transmitted or accessible - by adding salt to the 
original data before hashing it, you avoid a direct correspondence 
between the hash and the original data.  The prime use-case is to stop 
people being able to figure out a password by looking up the hash in a 
list of pre-computed hashes of common passwords.

> 
> OTOH, if the CRC yields something familiar -- or useful -- then
> it can tell you something about the image.  E.g., salt the algorithm
> with the product code, version number, your initials, 0xDEADBEEF, etc.
> 

You are making no sense at all.  Are you suggesting that it would be a 
good idea to add some value to the start of the image so that the 
resulting crc calculation gives a nice recognisable product code?  This 
"salt" would be different for each program image, and calculated by 
trial and error.  If you want a product code, version number, etc., in 
the program image (and it's a good idea), just put these in the program 
image!


>>>> So now you have a new extended block   |....data....|crc|
>>>>
>>>> Now if you compute a new CRC on the extended block, the resulting
>>>> value /should/ come out to zero. If it doesn't, either your data or
>>>> the original CRC value appended to it has been changed/corrupted.
>>>
>>> As there is usually a lack of originality in the algorithms
>>> chosen, you have to consider if you are also hoping to use
>>> this to safeguard the *integrity* of your image (i.e.,
>>> against intentional modification).
>>
>> "Integrity" has nothing to do with the motivation for change. 
>> /Security/ is concerned with intentional modifications that 
>> deliberately attempt to defeat /integrity/ checks.  Integrity is about 
>> detecting any changes.
>>
>> If you are concerned about the possibility of intentional malicious 
>> changes, 
> 
> Changes don't have to be malicious.  


Accidental changes (such as human error, noise during data transfer, 
memory cell errors, etc.) do not pass integrity tests unnoticed.  To be 
more accurate, the chances of them passing unnoticed are of the order of 
1 in 2^n, for a good n-bit check such as a CRC check.  Certain types of 
error are always detectable, such as single and double bit errors.  That 
is the point of using a checksum or hash for integrity checking.

/Intentional/ changes are a different matter.  If a hacker changes the 
program image, they can change the transmitted hash to their own 
calculated hash.  Or for a small CRC, they could change a different part 
of the image until the original checksum matched - for a 16-bit CRC, 
that only takes 65,535 attempts in the worst case.

That is why you need to distinguish between the two possibilities.  If 
you don't have to worry about malicious attacks, a 32-bit CRC takes a 
dozen lines of C code and a 1 KB table, all running extremely 
efficiently.  If security is an issue, you need digital signatures - an 
RSA-based signature system is orders of magnitude more effort in both 
development time and in run time.


> I altered the test procedure for a
> piece of military gear we were building simply to skip some lengthy 
> tests that I *knew* would pass (I don't want to inject an extra 20 
> minutes of wait time
> just to get through a lengthy test I already know works before I can get
> to the test of interest to me, now.
> 
> I failed to undo the change before the official signoff on the device.
> 
> The only evidence of this was the fact that I had also patched the
> startup message to say "Go for coffee..." -- which remained on the
> screen for the duration of the lengthy (even with the long test
> elided) procedure...
> 
> ..which alerted folks to the fact that this *probably* wasn't the
> original image.  (The computer running the test suite on the DUT had
> no problem accepting my patched binary)

And what, exactly, do you think that anecdote tells us about CRC checks 
for image files?  It reminds us that we are all fallible, but does no 
more than that.


> 
>> CRC's alone are useless.  All the attacker needs to do after modifying 
>> the image is calculate the CRC themselves, and replace the original 
>> checksum with their own.
> 
> That assumes the "alterer" knows how to replace the checksum, how it
> is computed, where it is embedded in the image, etc.  I modified the Compaq
> portable mentioned without ever knowing where the checksum was store
> or *if* it was explicitly stored.  I had no desire to disassemble the
> BIOS ROMs (though could obviously do so as there was no "proprietary
> hardware" limiting access to their contents and the instruction set of
> the processor is well known!).
> 
> Instead, I did this by *guessing* how they would implement such a check
> in a bit of kit from that era (ERPOMs aren't easily modified by malware
> so it wasn't likely that they would go to great lengths to "protect" the
> image).  And, if my guess had been incorrect, I could always reinstall
> the original EPROMs -- nothing lost, nothing gained.
> 
> Had much experience with folks counterfeiting your products and making
> "simple" changes to the binaries?  Like changing the copyright notice
> or splash screen?
> 
> Then, bringing the (accused) counterfeit of YOUR product into a courtroom
> and revealing the *hidden* checksum that the counterfeiter wasn't aware of?
> 
> "Gee, why does YOUR (alleged) device have *my* name in it -- in addition
> to behaving exactly like mine??"
> 
> [I guess obscurity has its place!]

Security by obscurity is not security.  Having a hidden signature or 
other mark can be useful for proving ownership (making an intentional 
mistake is another common tactic - such as commercial maps having a few 
subtle spelling errors).  But that is not security.

> 
> Use a non-secret approach and you invite folks to alter it, as well.
> 
>> Using non-standard algorithms for security is a simple way to get 
>> things completely wrong.  "Security by obscurity" is very rarely the 
>> right answer.  In reality, good security algorithms, and good 
>> implementations, are difficult and specialised tasks, best left to 
>> people who know what they are doing.
>>
>> To make something secure, you have to ensure that the check algorithms 
>> depend on a key that you know, but that the attacker does not have. 
>> That's the basis of digital signatures (though you use a secure hash 
>> algorithm rather than a simple CRC).
> 
> If you can remove the check, then what value the key's secrecy?  By your
> criteria, the adversary KNOWS how you are implementing your security
> so he knows exactly what to remove to bypass your checks and allow his
> altered image to operate in its place.
> 
> Ever notice how manufacturers don't PUBLICLY disclose their security
> hooks (without an NDA)?  If "security by obscurity" was not important,
> they would publish these details INVITING challenges (instead of
> trying to limit the knowledge to people with whom they've officially
> contracted).
> 

Any serious manufacturer /does/ invite challenges to their security.

There are multiple reasons why a manufacturer (such as a semiconductor 
manufacturer) might be guarded about the details of their security 
systems.  They can be avoiding giving hints to competitors.  Maybe they 
know their systems aren't really very secure, because their keys are too 
short or they can be read out in some way.

But I think the main reasons are often:

They want to be able to change the details, and that's far easier if 
there are only a few people who have read the information.

They don't want endless support questions from amateurs.

They are limited by idiotic government export restrictions made by 
ignorant politicians who don't understand cryptography.



Some things benefit from being kept hidden, or under restricted access. 
The details of the CRC algorithm you use to catch accidental errors in 
your image file is /not/ one of them.  If you think hiding it has the 
remotest hint of a benefit, you are doing things wrong - you need a 
/security/ check, not a simple /integrity/ check.

And then once you have switched to a security check - a digital 
signature - there's no need to keep that choice hidden either, because 
it is the /key/ that is important, not the type of lock.

[toc] | [prev] | [next] | [standalone]


#31814

FromDon Y <blockedofcourse@foo.invalid>
Date2023-04-21 17:29 -0700
Message-ID<u1v9q8$2v4d5$1@dont-email.me>
In reply to#31808
On 4/21/2023 7:50 AM, David Brown wrote:
> On 21/04/2023 13:39, Don Y wrote:
>> On 4/21/2023 3:43 AM, David Brown wrote:
>>>> Note that you want to choose a polynomial that doesn't
>>>> give you a "win" result for "obviously" corrupt data.
>>>> E.g., if data is all zeros or all 0xFF (as these sorts of
>>>> conditions can happen with hardware failures) you probably
>>>> wouldn't want a "success" indication!
>>>
>>> No, that is pointless for something like a code image.  It just adds 
>>> needless complexity to your CRC algorithm.
>>
>> Perhaps you've forgotten that you don't just use CRCs (secure hashes, etc.)
>> on "code images"?
> 
> No - but "code images" is the topic here.

So, anything unrelated to CRC's as applied to code images is off limits...
per order of the Internet Police"?

If *all* you use CRCs for is checking *a* code image at POST, you're
wasting a valuable resource.

Do you not think data/parameters need to be safeguarded?  Program images?
Communication protocols?

Or, do you develop yet another technique for *each* of those?

> However, in almost every case where CRC's might be useful, you have additional 
> checks of the sanity of the data, and an all-zero or all-one data block would 
> be rejected.  For example, Ethernet packets use CRC for integrity checking, but 
> an attempt to send a packet type 0 from MAC address 00:00:00:00:00:00 to 
> address 00:00:00:00:00:00, of length 0, would be rejected anyway.

Why look at "data" -- which may be suspect -- and *then* check its CRC?
Run the CRC first.  If it fails, decide how you are going to proceed
or recover.

["Data" can be code or parameters]

I treat blocks of "data" (carefully arranged) with individual CRCs,
based on their relative importance to the operation.  If the CRC is
corrupt, I have no idea *where* the error lies -- as it could
be anything in the checked block.  So, one has to (typically)
restore some defaults (or, invoke a reconfigure operation) which
recreates *a* valid dataset.

This is particularly useful when power to a device can be
removed at arbitrary points in time (or, some other abrupt
crash).  Before altering anything in a block, take deliberate
steps to invalidate the CRC, make your changes, then "fix"
the CRC.  So, an interrupted process causes the CRC to fail
and remedial action taken.

Note that replacing a FLASH image (mostly code) falls under
such a mechanism.

> I can't think of any use-cases where you would be passing around a block of 
> "pure" data that could reasonably take absolutely any value, without any type 
> of "envelope" information, and where you would think a CRC check is appropriate.

I append a *version specific* CRC to each packet of marshalled data
in my RMIs.  If the data is corrupted in transit *or* if the
wrong version API ends up targeted, the operation will abend
because we know the data "isn't right".

I *could* put a header saying "this is version 4.2".  And, that
tells me nothing about the integrity of the rest of the data.
OTOH, ensuring the CRC reflects "4.2" does -- it the recipient
expects it to be so.

>>>> You can also "salt" the calculation so that the residual
>>>> is deliberately nonzero.  So, for example, "success" is
>>>> indicated by a residual of 0x474E.  :>
>>>
>>> Again, pointless.
>>>
>>> Salt is important for security-related hashes (like password hashes), not 
>>> for integrity checks.
>>
>> You've missed the point.  The correct "sum" can be anything.
>> Why is "0" more special than any other value?  As the value is
>> typically meaningless to anything other than the code that verifies
>> it, you couldn't look at an image (or the output of the verifier)
>> and gain anything from seeing that obscure value.
> 
> Do you actually know what is meant by "salt" in the context of hashes, and why 
> it is useful in some circumstances?  Do you understand that "salt" is added 
> (usually prepended, or occasionally mixed in in some other way) to the data 
> /before/ the hash is calculated?

What term would you have me use to indicate a "bias" applied to a CRC
algorithm?

> I have not given the slightest indication to suggest that "0" is a special 
> value.  I fully agree that the value you get from the checking algorithm does 
> not have to be 0 - I already suggested it could be compared to the stored 
> value.  I.e., your build your image file as "data ++ crc(data)", at check it by 
> re-calculating "crc(data)" on the received image and comparing the result to 
> the received crc.  There is no necessity or benefit in having a crc run 
> calculated over the received data plus the received crc being 0.
> 
> "Salt" is used in cases where the original data must be kept secret, and only 
> the hashes are transmitted or accessible - by adding salt to the original data 
> before hashing it, you avoid a direct correspondence between the hash and the 
> original data.  The prime use-case is to stop people being able to figure out a 
> password by looking up the hash in a list of pre-computed hashes of common 
> passwords.

See above.

>> OTOH, if the CRC yields something familiar -- or useful -- then
>> it can tell you something about the image.  E.g., salt the algorithm
>> with the product code, version number, your initials, 0xDEADBEEF, etc.
> 
> You are making no sense at all.  Are you suggesting that it would be a good 
> idea to add some value to the start of the image so that the resulting crc 
> calculation gives a nice recognisable product code?  This "salt" would be 
> different for each program image, and calculated by trial and error.  If you 
> want a product code, version number, etc., in the program image (and it's a 
> good idea), just put these in the program image!

Again, that tells you nothing about the rest of the image!
See the RMI desciption.

[Note that the OP is expecting the checksum to help *him*
identify versions:  "Just put these in the program image!"  Eh?]

>>>>> So now you have a new extended block   |....data....|crc|
>>>>>
>>>>> Now if you compute a new CRC on the extended block, the resulting
>>>>> value /should/ come out to zero. If it doesn't, either your data or
>>>>> the original CRC value appended to it has been changed/corrupted.
>>>>
>>>> As there is usually a lack of originality in the algorithms
>>>> chosen, you have to consider if you are also hoping to use
>>>> this to safeguard the *integrity* of your image (i.e.,
>>>> against intentional modification).
>>>
>>> "Integrity" has nothing to do with the motivation for change. /Security/ is 
>>> concerned with intentional modifications that deliberately attempt to defeat 
>>> /integrity/ checks.  Integrity is about detecting any changes.
>>>
>>> If you are concerned about the possibility of intentional malicious changes, 
>>
>> Changes don't have to be malicious. 
> 
> Accidental changes (such as human error, noise during data transfer, memory 
> cell errors, etc.) do not pass integrity tests unnoticed.

That's not true.  The role of the 8test* is to notice these.  If the test
is blind to the types of errors that are likely to occur, then it CAN'T
notice them.

A CRC (hash, etc.) reduces a large block of data to a small bit of
data.  So, by definition, there are multiple DIFFERENT sets of data that
map to the same CRC/hash/etc.  (2^(data_size-CRC-size))

E.g., simply summing the values in a block of memory will yield "0"
for ANY condition that results in the block having identical values
for ALL members, if the block size is a power of 2.  So, a block
of 0xFF, 0x00, 0xFE, 0x27, 0x88, etc. will all yield the same sum.
Clearly a bad choice of test!

OTOH, "salting" the calculation so that it is expected to yield
a value of 0x13 means *those* situations will be flagged as errors
(and a different set of situations will sneak by, undetected).
The trick (engineering) is to figure out which types of
failures/faults/errors are most common to occur and guard
against them.

> To be more accurate, 
> the chances of them passing unnoticed are of the order of 1 in 2^n, for a good 
> n-bit check such as a CRC check.  Certain types of error are always detectable, 
> such as single and double bit errors.  That is the point of using a checksum or 
> hash for integrity checking.
> 
> /Intentional/ changes are a different matter.  If a hacker changes the program 
> image, they can change the transmitted hash to their own calculated hash.  Or 
> for a small CRC, they could change a different part of the image until the 
> original checksum matched - for a 16-bit CRC, that only takes 65,535 attempts 
> in the worst case.

If the approach used is "typical", then you need far fewer attempts to
produce a correct image -- without EVER knowing where the CRC is stored.

> That is why you need to distinguish between the two possibilities.  If you 
> don't have to worry about malicious attacks, a 32-bit CRC takes a dozen lines 
> of C code and a 1 KB table, all running extremely efficiently.  If security is 
> an issue, you need digital signatures - an RSA-based signature system is orders 
> of magnitude more effort in both development time and in run time.

It's considerably more expensive AND not fool-proof -- esp if the
attacker knows you are signing binaries.  "OK, now I need to find
WHERE the signature is verified and just patch that "CALL" out
of the code".

>> I altered the test procedure for a
>> piece of military gear we were building simply to skip some lengthy tests 
>> that I *knew* would pass (I don't want to inject an extra 20 minutes of wait 
>> time
>> just to get through a lengthy test I already know works before I can get
>> to the test of interest to me, now.
>>
>> I failed to undo the change before the official signoff on the device.
>>
>> The only evidence of this was the fact that I had also patched the
>> startup message to say "Go for coffee..." -- which remained on the
>> screen for the duration of the lengthy (even with the long test
>> elided) procedure...
>>
>> ..which alerted folks to the fact that this *probably* wasn't the
>> original image.  (The computer running the test suite on the DUT had
>> no problem accepting my patched binary)
> 
> And what, exactly, do you think that anecdote tells us about CRC checks for 
> image files?  It reminds us that we are all fallible, but does no more than that.

That *was* the point.  Because the folks who designed the test computer
relied on common techniques to safeguard the image.

The counterfeiting example I cited indicates how "obscurity/secrecy"
is far more effective (yet you dismiss it out-of-hand).

>>> CRC's alone are useless.  All the attacker needs to do after modifying the 
>>> image is calculate the CRC themselves, and replace the original checksum 
>>> with their own.
>>
>> That assumes the "alterer" knows how to replace the checksum, how it
>> is computed, where it is embedded in the image, etc.  I modified the Compaq
>> portable mentioned without ever knowing where the checksum was store
>> or *if* it was explicitly stored.  I had no desire to disassemble the
>> BIOS ROMs (though could obviously do so as there was no "proprietary
>> hardware" limiting access to their contents and the instruction set of
>> the processor is well known!).
>>
>> Instead, I did this by *guessing* how they would implement such a check
>> in a bit of kit from that era (ERPOMs aren't easily modified by malware
>> so it wasn't likely that they would go to great lengths to "protect" the
>> image).  And, if my guess had been incorrect, I could always reinstall
>> the original EPROMs -- nothing lost, nothing gained.
>>
>> Had much experience with folks counterfeiting your products and making
>> "simple" changes to the binaries?  Like changing the copyright notice
>> or splash screen?
>>
>> Then, bringing the (accused) counterfeit of YOUR product into a courtroom
>> and revealing the *hidden* checksum that the counterfeiter wasn't aware of?
>>
>> "Gee, why does YOUR (alleged) device have *my* name in it -- in addition
>> to behaving exactly like mine??"
>>
>> [I guess obscurity has its place!]
> 
> Security by obscurity is not security.  Having a hidden signature or other mark 
> can be useful for proving ownership (making an intentional mistake is another 
> common tactic - such as commercial maps having a few subtle spelling errors).  
> But that is not security.

Of course it is!  If *you* check the "hidden signature" at runtime
and then alter "your" operation such that an altered copy fails
to perform properly, then then you have secured it.

Would you want to use a check-writing program if the account
balances it maintains were subtly (but not consistently)
incorrect?

OTOH, if the (altered) program threw up a splash screen and
said "Unlicensed copy detected" and refused to operate, the
"program" is still "secured" -- but, now you've provided an
easy indicator of whether or not the security has been
defeated.

We started doing this in the heyday of video (arcade) gaming;
a counterfeiter would have a clone of YOUR game on the market
(at substantially reduced prices) in a matter of *weeks*.
As Operators have no foreknowledge of which games will be
moneymakers and which will be "90 day wonders" (literally,
no longer played after 90 days of exposure!), what incentive
to pay for a genuine article?

If all a counterfeiter had to do was alter the copyright
notice (even if it was stored in some coded form), or alter
some graphics (name of game, colors/shapes of characters)
that's *no* impediment -- given how often and quickly
it could be done.

Games would not just look at their images during POST
but, also, verify that routineX() had some particular
side-effect that could be tested, etc.  Counterfeiters
would go to lengths to ensure even THESE tests would pass.

Because the game would *complain*, otherwise!  (so, keep
looking for more tests until the game stops throwing an
alarm).

OTOH, if you *hide* the checks in the runtime and alter
the game's performance subtly by folding expected values
into key calculations such that values derived from
altered code differ, you can annoy the player:  "why did
my guy just turn blue and run off the edge of the screen?"
An annoyed player stops putting money into a game.
A game that doesn't earn money -- regardless of how
inexpensive it was to purchase -- quickly teaches the
Owner not to invest in such "buggy" games.

This is much better than taking the counterfeiter to court and
proving the code is a copy of yours!  (and, "FlyByNight
Games Counterfeiters" simply closes up shop and opens up,
next door)

And, because there is no "drop dead" point in the code or
the games behavior, the counterfeiter never knows when
he's found all the protection mechanisms.

Checking signatures, CRCs, licensing schemes, etc. all are used
in a "drop dead" fashion so considerably easier to defeat.
Witness the number of "products" available as warez...

>> Use a non-secret approach and you invite folks to alter it, as well.
>>
>>> Using non-standard algorithms for security is a simple way to get things 
>>> completely wrong.  "Security by obscurity" is very rarely the right answer.  
>>> In reality, good security algorithms, and good implementations, are 
>>> difficult and specialised tasks, best left to people who know what they are 
>>> doing.
>>>
>>> To make something secure, you have to ensure that the check algorithms 
>>> depend on a key that you know, but that the attacker does not have. That's 
>>> the basis of digital signatures (though you use a secure hash algorithm 
>>> rather than a simple CRC).
>>
>> If you can remove the check, then what value the key's secrecy?  By your
>> criteria, the adversary KNOWS how you are implementing your security
>> so he knows exactly what to remove to bypass your checks and allow his
>> altered image to operate in its place.
>>
>> Ever notice how manufacturers don't PUBLICLY disclose their security
>> hooks (without an NDA)?  If "security by obscurity" was not important,
>> they would publish these details INVITING challenges (instead of
>> trying to limit the knowledge to people with whom they've officially
>> contracted).
> 
> Any serious manufacturer /does/ invite challenges to their security.
> 
> There are multiple reasons why a manufacturer (such as a semiconductor 
> manufacturer) might be guarded about the details of their security systems.  
> They can be avoiding giving hints to competitors.  Maybe they know their 
> systems aren't really very secure, because their keys are too short or they can 
> be read out in some way.
> 
> But I think the main reasons are often:
> 
> They want to be able to change the details, and that's far easier if there are 
> only a few people who have read the information.

So, a legitimate customer is subjected to arbitrary changes in
the product's implementation?

> They don't want endless support questions from amateurs.

Only answer with a support contract.

> They are limited by idiotic government export restrictions made by ignorant 
> politicians who don't understand cryptography.

Protections don't always have to be cryptographic.  The
"Fortress" payphone is remarkably well hardened to direct
physical (brute force) attacks -- money is involved.
Ditto many slot machines (again, CASH money).  Yet, all
have vulnerabilities.  "Expose this portion of the die
to ultraviolet light to reset the memory protection bits"
Etc.

> Some things benefit from being kept hidden, or under restricted access. The 
> details of the CRC algorithm you use to catch accidental errors in your image 
> file is /not/ one of them.  If you think hiding it has the remotest hint of a 
> benefit, you are doing things wrong - you need a /security/ check, not a simple 
> /integrity/ check.
> 
> And then once you have switched to a security check - a digital signature - 
> there's no need to keep that choice hidden either, because it is the /key/ that 
> is important, not the type of lock.

Again, meaningless if the attacker can interfere with the *enforcement*
of that check.  Using something "well known" just means he already knows
what to look for in your code.  Or, how to interfere with your
intended implementation in ways that you may have not anticipated
(confident that your "security" can't be MATHEMATICALLY broken).

I had a discussion with a friend who knew just enough about "computers"
to THINK he understood that world.  I mentioned my NOT using ecommerce.
He laughed at me as "naive":  "There's 40 bit encryption on those
connections!  No one is going to eavesdrop on your financial data!"

[Really, Jerry?  You think, as an OLD accountant, you know more
than I do as a young engineer practicing in that field?  Ok...]

"Yeah, and are you 100% sure something isn't already *on* your computer
looking at your keystrokes BEFORE they head down that encrypted tunnel?"

Guess he hadn't really thought out the problem to that level of detail
as his confidence quickly melted away to one of worry ("I wonder if
I've already been hacked??")

People implementing security almost always focus on the wrong
aspects of the problem and walk away THINKING they can rest easy.
Vulnerabilities are often so blatantly obvious, after the fact,
as to be embarassing:  "You're not supposed to do that!"
"Then, why did your product LET ME?"

I use *many* layers of security in my current design and STILL
expect them (at least the ones that are accessible) to all
be subverted.  So, ultimately rely on controlling *what*
the devices can do so that, even compromised, they can't
cause undetectable failures or information leaks.

"Here's my source code.  Here are my schematics.  Here's the
name of the guy who oversees production (bribe him to gain
access to the keys stored in the TPM).  Now, what are you
gonna *do* with all that?"

[toc] | [prev] | [next] | [standalone]


#31820

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-22 16:57 +0200
Message-ID<u20sli$3ag7l$1@dont-email.me>
In reply to#31814
On 22/04/2023 02:29, Don Y wrote:
> On 4/21/2023 7:50 AM, David Brown wrote:
>> On 21/04/2023 13:39, Don Y wrote:
>>> On 4/21/2023 3:43 AM, David Brown wrote:
>>>>> Note that you want to choose a polynomial that doesn't
>>>>> give you a "win" result for "obviously" corrupt data.
>>>>> E.g., if data is all zeros or all 0xFF (as these sorts of
>>>>> conditions can happen with hardware failures) you probably
>>>>> wouldn't want a "success" indication!
>>>>
>>>> No, that is pointless for something like a code image.  It just adds 
>>>> needless complexity to your CRC algorithm.
>>>
>>> Perhaps you've forgotten that you don't just use CRCs (secure hashes, 
>>> etc.)
>>> on "code images"?
>>
>> No - but "code images" is the topic here.
> 
> So, anything unrelated to CRC's as applied to code images is off limits...
> per order of the Internet Police"?
> 

No, it's fine to discuss them - threads on Usenet often wander, and 
that's often good.  (At least, that's my opinion - some people get their 
knickers in a twist if people stray from answering their original question.)

But you have to assume that people are on topic unless it's clear that 
the topic is being expanded.  We were discussing CRC's for code images, 
and so it is appropriate to take advantage of the features of code 
images.  If you want to expand and talk about other uses of CRC's, I've 
no problem with that - but you need to say so.

> If *all* you use CRCs for is checking *a* code image at POST, you're
> wasting a valuable resource.
> 
> Do you not think data/parameters need to be safeguarded?  Program images?
> Communication protocols?

Sure.  Many things need integrity checks.  And CRC's are flexible enough 
to be useful in many circumstances.

> 
> Or, do you develop yet another technique for *each* of those?

Sometimes, yes.  CRC's are, as I wrote, flexible.  But they don't cover 
everything.  Maybe you need a specific type of check to match existing 
protocols or requirements.  Maybe you want forward error correction, not 
just error detection.  Maybe you are guarding against malicious 
interference.  Maybe you are guarding against different kinds of errors 
- CRC's are great for spotting a few damaged bits, but a poor choice if 
the risk is dropped bytes in transmission.

But often CRC's will be a first choice, because they are simple and 
effective in a wide range of uses.

> 
>> However, in almost every case where CRC's might be useful, you have 
>> additional checks of the sanity of the data, and an all-zero or 
>> all-one data block would be rejected.  For example, Ethernet packets 
>> use CRC for integrity checking, but an attempt to send a packet type 0 
>> from MAC address 00:00:00:00:00:00 to address 00:00:00:00:00:00, of 
>> length 0, would be rejected anyway.
> 
> Why look at "data" -- which may be suspect -- and *then* check its CRC?
> Run the CRC first.  If it fails, decide how you are going to proceed
> or recover.
> 

That is usually the order, yes.  Sometimes you want "fail fast", such as 
dropping a packet that was not addressed to you (it doesn't matter if it 
was received correctly but for someone else, or it was addressed to you 
but the receiver address was corrupted - you are dropping the packet 
either way).  But usually you will run the CRC then look at the data.

But the order doesn't matter - either way, you are still checking for 
valid data, and if the data is invalid, it does not matter if the CRC 
only passed by luck or by all zeros.

> ["Data" can be code or parameters]
> 
> I treat blocks of "data" (carefully arranged) with individual CRCs,
> based on their relative importance to the operation.  If the CRC is
> corrupt, I have no idea *where* the error lies -- as it could
> be anything in the checked block.  So, one has to (typically)
> restore some defaults (or, invoke a reconfigure operation) which
> recreates *a* valid dataset.
> 
> This is particularly useful when power to a device can be
> removed at arbitrary points in time (or, some other abrupt
> crash).  Before altering anything in a block, take deliberate
> steps to invalidate the CRC, make your changes, then "fix"
> the CRC.  So, an interrupted process causes the CRC to fail
> and remedial action taken.
> 
> Note that replacing a FLASH image (mostly code) falls under
> such a mechanism.
> 

That's all standard stuff.  (Maybe it's new to some people in this group 
- although most of the regular posters here are experienced embedded 
developers, it's nice to think there might be some people reading these 
posts and learning!)

If you have the space in your flash, eeprom, etc., then it is also 
common to have two slots for your configuration data or code.  You don't 
"invalidate" anything - you keep a version counter with your data, and 
write your new data to the slot with the oldest version.  When your 
system starts, it checks both slots - and uses the one with the newest 
version for which the CRC check passes.

>> I can't think of any use-cases where you would be passing around a 
>> block of "pure" data that could reasonably take absolutely any value, 
>> without any type of "envelope" information, and where you would think 
>> a CRC check is appropriate.
> 
> I append a *version specific* CRC to each packet of marshalled data
> in my RMIs.  If the data is corrupted in transit *or* if the
> wrong version API ends up targeted, the operation will abend
> because we know the data "isn't right".

Using a version-specific CRC sounds silly.  Put the version information 
in the packet.

> 
> I *could* put a header saying "this is version 4.2".  And, that
> tells me nothing about the integrity of the rest of the data.
> OTOH, ensuring the CRC reflects "4.2" does -- it the recipient
> expects it to be so.

Now you don't know if the data is corrupted, or for the wrong version - 
or occasionally, corrupted /and/ the wrong version but passing the CRC 
anyway.

Unless you are absolutely desperate to save every bit you can, your 
system will be simpler, clearer, and more reliable if you separate your 
purposes.

> 
>>>>> You can also "salt" the calculation so that the residual
>>>>> is deliberately nonzero.  So, for example, "success" is
>>>>> indicated by a residual of 0x474E.  :>
>>>>
>>>> Again, pointless.
>>>>
>>>> Salt is important for security-related hashes (like password 
>>>> hashes), not for integrity checks.
>>>
>>> You've missed the point.  The correct "sum" can be anything.
>>> Why is "0" more special than any other value?  As the value is
>>> typically meaningless to anything other than the code that verifies
>>> it, you couldn't look at an image (or the output of the verifier)
>>> and gain anything from seeing that obscure value.
>>
>> Do you actually know what is meant by "salt" in the context of hashes, 
>> and why it is useful in some circumstances?  Do you understand that 
>> "salt" is added (usually prepended, or occasionally mixed in in some 
>> other way) to the data /before/ the hash is calculated?
> 
> What term would you have me use to indicate a "bias" applied to a CRC
> algorithm?

Well, first I'd note that any kind of modification to the basic CRC 
algorithm is pointless from the viewpoint of its use as an integrity 
check.  (There have been, mostly historically, some justifications in 
terms of implementation efficiency.  For example, bit and byte 
re-ordering could be done to suit hardware bit-wise implementations.)

Otherwise I'd say you are picking a specific initial value if that is 
what you are doing, or modifying the final value (inverting it or 
xor'ing it with a fixed value).  There is, AFAIK, no specific terms for 
these - and I don't see any benefit in having one.  Misusing the term 
"salt" from cryptography is certainly not helpful.


> 
>> I have not given the slightest indication to suggest that "0" is a 
>> special value.  I fully agree that the value you get from the checking 
>> algorithm does not have to be 0 - I already suggested it could be 
>> compared to the stored value.  I.e., your build your image file as 
>> "data ++ crc(data)", at check it by re-calculating "crc(data)" on the 
>> received image and comparing the result to the received crc.  There is 
>> no necessity or benefit in having a crc run calculated over the 
>> received data plus the received crc being 0.
>>
>> "Salt" is used in cases where the original data must be kept secret, 
>> and only the hashes are transmitted or accessible - by adding salt to 
>> the original data before hashing it, you avoid a direct correspondence 
>> between the hash and the original data.  The prime use-case is to stop 
>> people being able to figure out a password by looking up the hash in a 
>> list of pre-computed hashes of common passwords.
> 
> See above.
> 
>>> OTOH, if the CRC yields something familiar -- or useful -- then
>>> it can tell you something about the image.  E.g., salt the algorithm
>>> with the product code, version number, your initials, 0xDEADBEEF, etc.
>>
>> You are making no sense at all.  Are you suggesting that it would be a 
>> good idea to add some value to the start of the image so that the 
>> resulting crc calculation gives a nice recognisable product code?  
>> This "salt" would be different for each program image, and calculated 
>> by trial and error.  If you want a product code, version number, etc., 
>> in the program image (and it's a good idea), just put these in the 
>> program image!
> 
> Again, that tells you nothing about the rest of the image!

Again, you are making no sense - not to me, anyway.  If you want 
something in the image to tell you about the image, add such metadata - 
versions, dates, whatever.  If you want an integrity check of the image, 
make one - such as appending a CRC.  Trying to combine these two 
orthogonal tasks into one is not going to be good for either purpose.

> See the RMI desciption.

I'm sorry, I have no idea what "RMI" is or where it is described. 
You've mentioned that abbreviation twice, but I can't figure it out.

> 
> [Note that the OP is expecting the checksum to help *him*
> identify versions:  "Just put these in the program image!"  Eh?]

No.  The OP is looking for a way to be sure that two program images are 
the same.  He wants to be sure that if he (or whoever makes the image) 
forgets to update the version number when making a change to the 
software, the difference between the images is easily detectable or 
identifiable without doing a byte-for-byte compare of the images.  The 
answer to that is a hash of some sort - and a CRC of appropriate size is 
a simple hash that will work well against mistakes (but not necessarily 
malicious changes).  But a hash will not give you a version number.  It 
will let you see that two images are different, but it will not tell you 
that one of them is version 1.20.304 and the other is 1.21.308.  What he 
will see is that if two files say they are version 1.20.304, but are 
actually different, someone has screwed up - the CRC hash makes such 
checks possible without having to read through the entire images.

> 
>>>>>> So now you have a new extended block   |....data....|crc|
>>>>>>
>>>>>> Now if you compute a new CRC on the extended block, the resulting
>>>>>> value /should/ come out to zero. If it doesn't, either your data or
>>>>>> the original CRC value appended to it has been changed/corrupted.
>>>>>
>>>>> As there is usually a lack of originality in the algorithms
>>>>> chosen, you have to consider if you are also hoping to use
>>>>> this to safeguard the *integrity* of your image (i.e.,
>>>>> against intentional modification).
>>>>
>>>> "Integrity" has nothing to do with the motivation for change. 
>>>> /Security/ is concerned with intentional modifications that 
>>>> deliberately attempt to defeat /integrity/ checks.  Integrity is 
>>>> about detecting any changes.
>>>>
>>>> If you are concerned about the possibility of intentional malicious 
>>>> changes, 
>>>
>>> Changes don't have to be malicious. 
>>
>> Accidental changes (such as human error, noise during data transfer, 
>> memory cell errors, etc.) do not pass integrity tests unnoticed.
> 
> That's not true.  The role of the 8test* is to notice these.  If the test
> is blind to the types of errors that are likely to occur, then it CAN'T
> notice them.

I assumed it was unnecessary to say that an integrity test needs to be 
appropriate for the type of data and transfer in question.

> 
> A CRC (hash, etc.) reduces a large block of data to a small bit of
> data.  So, by definition, there are multiple DIFFERENT sets of data that
> map to the same CRC/hash/etc.  (2^(data_size-CRC-size))

Correct.

That's why you need to pick an appropriate size for your CRC.  For a 
telegram of a dozen bytes, an 8-bit CRC is probably fine.  For a program 
image, a 32-bit CRC is usually more appropriate - a one in four billion 
chance of an undetected error is reasonable for most uses.  If you want 
to be more paranoid, go for 64-bit CRC - you should now be far more 
worried about meteors wiping out humanity than undetected errors.  (More 
commonly, if a 32-bit CRC is not enough, it's because you have security 
concerns - so switch to a SHA hash.)

> 
> E.g., simply summing the values in a block of memory will yield "0"
> for ANY condition that results in the block having identical values
> for ALL members, if the block size is a power of 2.  So, a block
> of 0xFF, 0x00, 0xFE, 0x27, 0x88, etc. will all yield the same sum.
> Clearly a bad choice of test!
> 

Correct.

That's why simple sums are not usually considered very good integrity tests.

A CRC has a spreading effect.  Every bit in the data contributes with 
approximately equal weight to every bit in the CRC.  This is a common 
feature for good hash functions.

> OTOH, "salting" the calculation so that it is expected to yield
> a value of 0x13 means *those* situations will be flagged as errors
> (and a different set of situations will sneak by, undetected).

And that gives you exactly /zero/ benefit.

You run your hash algorithm, and check for the single value that 
indicates no errors.  It does not matter if that number is 0, 0x13, or - 
often more conveniently - the number attached at the end of the image as 
the expected result of the hash of the rest of the data.

> The trick (engineering) is to figure out which types of
> failures/faults/errors are most common to occur and guard
> against them.

Yes, that is absolutely the case.  And CRC's have the convenience of 
being particularly good at certain kinds of errors that are feasible in 
a lot of data transmissions.  But they are not ideal for everything, and 
other kinds of checks can be better when you know more about the 
realistic errors.

> 
>> To be more accurate, the chances of them passing unnoticed are of the 
>> order of 1 in 2^n, for a good n-bit check such as a CRC check.  
>> Certain types of error are always detectable, such as single and 
>> double bit errors.  That is the point of using a checksum or hash for 
>> integrity checking.
>>
>> /Intentional/ changes are a different matter.  If a hacker changes the 
>> program image, they can change the transmitted hash to their own 
>> calculated hash.  Or for a small CRC, they could change a different 
>> part of the image until the original checksum matched - for a 16-bit 
>> CRC, that only takes 65,535 attempts in the worst case.
> 
> If the approach used is "typical", then you need far fewer attempts to
> produce a correct image -- without EVER knowing where the CRC is stored.
> 

It is difficult to know what you are trying to say here, but if you 
believe that different initial values in a CRC algorithm makes it harder 
to modify an image to make it pass the integrity test, you are simply wrong.

>> That is why you need to distinguish between the two possibilities.  If 
>> you don't have to worry about malicious attacks, a 32-bit CRC takes a 
>> dozen lines of C code and a 1 KB table, all running extremely 
>> efficiently.  If security is an issue, you need digital signatures - 
>> an RSA-based signature system is orders of magnitude more effort in 
>> both development time and in run time.
> 
> It's considerably more expensive AND not fool-proof -- esp if the
> attacker knows you are signing binaries.  "OK, now I need to find
> WHERE the signature is verified and just patch that "CALL" out
> of the code".

I'm not sure if that is a straw-man argument, or just showing your 
ignorance of the topic.  Do you really think security checks are done by 
the program you are trying to send securely?  That would be like trying 
to have building security where people entering the building look at 
their own security cards.

> 
>>> I altered the test procedure for a
>>> piece of military gear we were building simply to skip some lengthy 
>>> tests that I *knew* would pass (I don't want to inject an extra 20 
>>> minutes of wait time
>>> just to get through a lengthy test I already know works before I can get
>>> to the test of interest to me, now.
>>>
>>> I failed to undo the change before the official signoff on the device.
>>>
>>> The only evidence of this was the fact that I had also patched the
>>> startup message to say "Go for coffee..." -- which remained on the
>>> screen for the duration of the lengthy (even with the long test
>>> elided) procedure...
>>>
>>> ..which alerted folks to the fact that this *probably* wasn't the
>>> original image.  (The computer running the test suite on the DUT had
>>> no problem accepting my patched binary)
>>
>> And what, exactly, do you think that anecdote tells us about CRC 
>> checks for image files?  It reminds us that we are all fallible, but 
>> does no more than that.
> 
> That *was* the point.  Because the folks who designed the test computer
> relied on common techniques to safeguard the image.

There was a human error - procedures were not good enough, or were not 
followed.  It happens, and you learn from it and make better procedures. 
  The fault was in what people did, not in an automated integrity check. 
  It is completely unrelated.

> 
> The counterfeiting example I cited indicates how "obscurity/secrecy"
> is far more effective (yet you dismiss it out-of-hand).

No, it does nothing of the sort.  There is no connection at all.

> 
>>>> CRC's alone are useless.  All the attacker needs to do after 
>>>> modifying the image is calculate the CRC themselves, and replace the 
>>>> original checksum with their own.
>>>
>>> That assumes the "alterer" knows how to replace the checksum, how it
>>> is computed, where it is embedded in the image, etc.  I modified the 
>>> Compaq
>>> portable mentioned without ever knowing where the checksum was store
>>> or *if* it was explicitly stored.  I had no desire to disassemble the
>>> BIOS ROMs (though could obviously do so as there was no "proprietary
>>> hardware" limiting access to their contents and the instruction set of
>>> the processor is well known!).
>>>
>>> Instead, I did this by *guessing* how they would implement such a check
>>> in a bit of kit from that era (ERPOMs aren't easily modified by malware
>>> so it wasn't likely that they would go to great lengths to "protect" the
>>> image).  And, if my guess had been incorrect, I could always reinstall
>>> the original EPROMs -- nothing lost, nothing gained.
>>>
>>> Had much experience with folks counterfeiting your products and making
>>> "simple" changes to the binaries?  Like changing the copyright notice
>>> or splash screen?
>>>
>>> Then, bringing the (accused) counterfeit of YOUR product into a 
>>> courtroom
>>> and revealing the *hidden* checksum that the counterfeiter wasn't 
>>> aware of?
>>>
>>> "Gee, why does YOUR (alleged) device have *my* name in it -- in addition
>>> to behaving exactly like mine??"
>>>
>>> [I guess obscurity has its place!]
>>
>> Security by obscurity is not security.  Having a hidden signature or 
>> other mark can be useful for proving ownership (making an intentional 
>> mistake is another common tactic - such as commercial maps having a 
>> few subtle spelling errors). But that is not security.
> 
> Of course it is!  If *you* check the "hidden signature" at runtime
> and then alter "your" operation such that an altered copy fails
> to perform properly, then then you have secured it.
> 

That is not security.  "Security" means that the program that starts the 
updated program checks the /entire/ image according to its digital 
signature, and rejects it /entirely/ if it does not match.

What you are talking about here is the sort of cat-and-mouse nonsense 
computer games producers did with intentional disk errors to stop 
copying.  It annoys legitimate users and does almost nothing to hinder 
the bad guys.

> Would you want to use a check-writing program if the account
> balances it maintains were subtly (but not consistently)
> incorrect?

Again, you make no sense.  What has this got to do with integrity checks 
or security?

> 
> OTOH, if the (altered) program threw up a splash screen and
> said "Unlicensed copy detected" and refused to operate, the
> "program" is still "secured" -- but, now you've provided an
> easy indicator of whether or not the security has been
> defeated.
> 
> We started doing this in the heyday of video (arcade) gaming;
> a counterfeiter would have a clone of YOUR game on the market
> (at substantially reduced prices) in a matter of *weeks*.
> As Operators have no foreknowledge of which games will be
> moneymakers and which will be "90 day wonders" (literally,
> no longer played after 90 days of exposure!), what incentive
> to pay for a genuine article?
> 
> If all a counterfeiter had to do was alter the copyright
> notice (even if it was stored in some coded form), or alter
> some graphics (name of game, colors/shapes of characters)
> that's *no* impediment -- given how often and quickly
> it could be done.
> 
> Games would not just look at their images during POST
> but, also, verify that routineX() had some particular
> side-effect that could be tested, etc.  Counterfeiters
> would go to lengths to ensure even THESE tests would pass.
> 
> Because the game would *complain*, otherwise!  (so, keep
> looking for more tests until the game stops throwing an
> alarm).
> 
> OTOH, if you *hide* the checks in the runtime and alter
> the game's performance subtly by folding expected values
> into key calculations such that values derived from
> altered code differ, you can annoy the player:  "why did
> my guy just turn blue and run off the edge of the screen?"
> An annoyed player stops putting money into a game.
> A game that doesn't earn money -- regardless of how
> inexpensive it was to purchase -- quickly teaches the
> Owner not to invest in such "buggy" games.
> 
> This is much better than taking the counterfeiter to court and
> proving the code is a copy of yours!  (and, "FlyByNight
> Games Counterfeiters" simply closes up shop and opens up,
> next door)
> 
> And, because there is no "drop dead" point in the code or
> the games behavior, the counterfeiter never knows when
> he's found all the protection mechanisms.
> 
> Checking signatures, CRCs, licensing schemes, etc. all are used
> in a "drop dead" fashion so considerably easier to defeat.
> Witness the number of "products" available as warez...
> 

Look, it is all /really/ simple.  And the year is 2023, not 1973.

If you want to check the integrity of a file against accidental changes, 
a CRC is usually fine.

If you want security, and to protect against malicious changes, use a 
digital signature.  This must be checked by the program that /starts/ 
the updated code, or that downloaded and stored it - not by the program 
itself!


>>> Use a non-secret approach and you invite folks to alter it, as well.
>>>
>>>> Using non-standard algorithms for security is a simple way to get 
>>>> things completely wrong.  "Security by obscurity" is very rarely the 
>>>> right answer. In reality, good security algorithms, and good 
>>>> implementations, are difficult and specialised tasks, best left to 
>>>> people who know what they are doing.
>>>>
>>>> To make something secure, you have to ensure that the check 
>>>> algorithms depend on a key that you know, but that the attacker does 
>>>> not have. That's the basis of digital signatures (though you use a 
>>>> secure hash algorithm rather than a simple CRC).
>>>
>>> If you can remove the check, then what value the key's secrecy?  By your
>>> criteria, the adversary KNOWS how you are implementing your security
>>> so he knows exactly what to remove to bypass your checks and allow his
>>> altered image to operate in its place.
>>>
>>> Ever notice how manufacturers don't PUBLICLY disclose their security
>>> hooks (without an NDA)?  If "security by obscurity" was not important,
>>> they would publish these details INVITING challenges (instead of
>>> trying to limit the knowledge to people with whom they've officially
>>> contracted).
>>
>> Any serious manufacturer /does/ invite challenges to their security.
>>
>> There are multiple reasons why a manufacturer (such as a semiconductor 
>> manufacturer) might be guarded about the details of their security 
>> systems. They can be avoiding giving hints to competitors.  Maybe they 
>> know their systems aren't really very secure, because their keys are 
>> too short or they can be read out in some way.
>>
>> But I think the main reasons are often:
>>
>> They want to be able to change the details, and that's far easier if 
>> there are only a few people who have read the information.
> 
> So, a legitimate customer is subjected to arbitrary changes in
> the product's implementation?
> 

Yes.  It may come as a shock to you, but welcome to the real world.

>> They don't want endless support questions from amateurs.
> 
> Only answer with a support contract.

Oh, sure - the amateurs who have some of the information but not enough 
details, skill or knowledge to get things working will /never/ fill 
forums with questions, complaints or bad reviews that bother your 
support staff or scare away real sales.

> 
>> They are limited by idiotic government export restrictions made by 
>> ignorant politicians who don't understand cryptography.
> 
> Protections don't always have to be cryptographic.  

Correct, but - as with a lot of what you write - completely irrelevant 
to the subject at hand.

Why can't companies give out information about the security systems used 
in their microcontrollers (for example) ?  Because some geriatric 
ignoramuses think banning "export" of such information to certain 
countries will stop those countries knowing about security and cryptography.

> The
> "Fortress" payphone is remarkably well hardened to direct
> physical (brute force) attacks -- money is involved.
> Ditto many slot machines (again, CASH money).  Yet, all
> have vulnerabilities.  "Expose this portion of the die
> to ultraviolet light to reset the memory protection bits"
> Etc.
> 
>> Some things benefit from being kept hidden, or under restricted 
>> access. The details of the CRC algorithm you use to catch accidental 
>> errors in your image file is /not/ one of them.  If you think hiding 
>> it has the remotest hint of a benefit, you are doing things wrong - 
>> you need a /security/ check, not a simple /integrity/ check.
>>
>> And then once you have switched to a security check - a digital 
>> signature - there's no need to keep that choice hidden either, because 
>> it is the /key/ that is important, not the type of lock.
> 
> Again, meaningless if the attacker can interfere with the *enforcement*
> of that check.  Using something "well known" just means he already knows
> what to look for in your code.  Or, how to interfere with your
> intended implementation in ways that you may have not anticipated
> (confident that your "security" can't be MATHEMATICALLY broken).
> 

If the attacker can interfere with the enforcement of the check, then it 
doesn't matter what checks you have.  Keeping the design of a building's 
locks secret does not help you if the bad guys have bribed the security 
guard /inside/ the building!

> I had a discussion with a friend who knew just enough about "computers"
> to THINK he understood that world.  I mentioned my NOT using ecommerce.
> He laughed at me as "naive":  "There's 40 bit encryption on those
> connections!  No one is going to eavesdrop on your financial data!"
> 
> [Really, Jerry?  You think, as an OLD accountant, you know more
> than I do as a young engineer practicing in that field?  Ok...]
> 
> "Yeah, and are you 100% sure something isn't already *on* your computer
> looking at your keystrokes BEFORE they head down that encrypted tunnel?"
> 
> Guess he hadn't really thought out the problem to that level of detail
> as his confidence quickly melted away to one of worry ("I wonder if
> I've already been hacked??")
> 
> People implementing security almost always focus on the wrong
> aspects of the problem and walk away THINKING they can rest easy.
> Vulnerabilities are often so blatantly obvious, after the fact,
> as to be embarassing:  "You're not supposed to do that!"
> "Then, why did your product LET ME?"
> 
> I use *many* layers of security in my current design and STILL
> expect them (at least the ones that are accessible) to all
> be subverted.  So, ultimately rely on controlling *what*
> the devices can do so that, even compromised, they can't
> cause undetectable failures or information leaks.
> 
> "Here's my source code.  Here are my schematics.  Here's the
> name of the guy who oversees production (bribe him to gain
> access to the keys stored in the TPM).  Now, what are you
> gonna *do* with all that?"
> 

The first two should be fine - if people can break your security after 
looking at your source code or schematics, your security is /bad/.  As 
for the third one, if they can break your security by going through the 
production guy, your production procedures are bad.

[toc] | [prev] | [next] | [standalone]


Page 1 of 5  [1] 2 3 4 5  Next page →

Back to top | Article view | comp.arch.embedded


csiph-web