Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.arch.embedded > #31788 > unrolled thread

Embedding a Checksum in an Image File

Started byRick C <gnuarm.deletethisbit@gmail.com>
First post2023-04-19 19:06 -0700
Last post2023-04-27 18:27 +0200
Articles 20 on this page of 85 — 14 participants

Back to article view | Back to comp.arch.embedded


Contents

  Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-19 19:06 -0700
    Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-20 12:14 +0300
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:18 -0700
    Re: Embedding a Checksum in an Image File "Peter Heitzer" <peter.heitzer@rz.uni-regensburg.de> - 2023-04-20 11:30 +0000
    Re: Embedding a Checksum in an Image File dalai lamah <antonio12358@hotmail.com> - 2023-04-20 13:47 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 06:04 -0700
    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 16:46 +0200
    Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 11:33 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 09:45 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-20 22:26 +0200
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:36 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:12 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:35 +0200
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:44 -0400
          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-20 22:37 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 12:43 +0200
              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 04:39 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 16:50 +0200
                  Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-21 17:29 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 16:57 +0200
                      Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-04-24 00:32 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 16:37 +0200
                          Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-05-03 00:15 -0700
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-03 14:48 +0200
                              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:42 +0200
                                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:06 +0200
                                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-10 12:03 +0200
                                    Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:48 -0700
                              Re: Embedding a Checksum in an Image File Don Y <blockedofcourse@foo.invalid> - 2023-08-05 01:42 -0700
          Re: Embedding a Checksum in an Image File Stefan Reuther <stefan.news@arcor.de> - 2023-04-21 19:40 +0200
      Re: Embedding a Checksum in an Image File Tauno Voipio <tauno.voipio@notused.fi.invalid> - 2023-04-20 20:17 +0300
        Re: Embedding a Checksum in an Image File George Neuner <gneuner2@comcast.net> - 2023-04-20 16:49 -0400
    Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-20 22:09 -0400
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-20 19:41 -0700
        Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-21 19:30 -0400
    Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 01:53 -0700
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 05:12 -0700
        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-21 17:02 +0200
          Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:56 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:01 +0200
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:14 -0700
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 17:13 +0200
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:56 -0700
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-22 19:54 +0200
                  Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-22 20:05 +0000
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 17:37 +0200
                      Re: Embedding a Checksum in an Image File Grant Edwards <invalid@invalid.invalid> - 2023-04-23 17:37 +0000
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:45 +0200
                          Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-23 18:16 -0400
                            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:13 +0200
                  Re: Embedding a Checksum in an Image File boB <boB@K7IQ.com> - 2023-04-22 13:41 -0700
                  Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 10:34 -0700
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-23 23:58 +0200
                      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-23 15:24 -0700
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-24 09:17 +0200
                          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-24 01:07 -0700
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:42 +0200
              Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:20 +0200
                Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:44 +0200
        Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-21 16:52 -0700
          Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-21 20:23 -0700
            Re: Embedding a Checksum in an Image File Brian Cockburn <brian.cockburn.1959@gmail.com> - 2023-04-22 07:07 -0700
              Re: Embedding a Checksum in an Image File Richard Damon <Richard@Damon-Family.org> - 2023-04-22 10:31 -0400
              Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-22 09:54 -0700
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:26 +0200
      Re: Embedding a Checksum in an Image File Rick C <gnuarm.deletethisbit@gmail.com> - 2023-04-27 10:09 -0700
        Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:29 +0300
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-27 21:39 +0300
          Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:44 +0200
            Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:38 +0200
              Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:50 +0200
                Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:04 +0200
                  Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:03 +0200
                    Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-30 16:19 +0200
                      Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-05-09 20:34 +0200
                        Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-05-10 10:18 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:33 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 22:36 +0200
          Re: Embedding a Checksum in an Image File Niklas Holsti <niklas.holsti@tidorum.invalid> - 2023-04-28 01:10 +0300
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:54 +0200
      Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 09:24 +0200
        Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-28 10:56 +0200
          Re: Embedding a Checksum in an Image File David Brown <david.brown@hesbynett.no> - 2023-04-28 15:09 +0200
            Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-29 23:02 +0200
    Re: Embedding a Checksum in an Image File Ulf Samuelsson <ulf.r.samuelsson@gmail.com> - 2023-04-27 18:27 +0200

Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →


#31837

FromDon Y <blockedofcourse@foo.invalid>
Date2023-04-24 00:32 -0700
Message-ID<u25bas$8fh0$2@dont-email.me>
In reply to#31820
On 4/22/2023 7:57 AM, David Brown wrote:
>>> However, in almost every case where CRC's might be useful, you have 
>>> additional checks of the sanity of the data, and an all-zero or all-one data 
>>> block would be rejected.  For example, Ethernet packets use CRC for 
>>> integrity checking, but an attempt to send a packet type 0 from MAC address 
>>> 00:00:00:00:00:00 to address 00:00:00:00:00:00, of length 0, would be 
>>> rejected anyway.
>>
>> Why look at "data" -- which may be suspect -- and *then* check its CRC?
>> Run the CRC first.  If it fails, decide how you are going to proceed
>> or recover.
> 
> That is usually the order, yes.  Sometimes you want "fail fast", such as 
> dropping a packet that was not addressed to you (it doesn't matter if it was 
> received correctly but for someone else, or it was addressed to you but the 
> receiver address was corrupted - you are dropping the packet either way).  But 
> usually you will run the CRC then look at the data.
> 
> But the order doesn't matter - either way, you are still checking for valid 
> data, and if the data is invalid, it does not matter if the CRC only passed by 
> luck or by all zeros.

You're assuming the CRC is supposed to *vouch* for the data.
The CRC can be there simply to vouch for the *transport* of a
datagram.

>>> I can't think of any use-cases where you would be passing around a block of 
>>> "pure" data that could reasonably take absolutely any value, without any 
>>> type of "envelope" information, and where you would think a CRC check is 
>>> appropriate.
>>
>> I append a *version specific* CRC to each packet of marshalled data
>> in my RMIs.  If the data is corrupted in transit *or* if the
>> wrong version API ends up targeted, the operation will abend
>> because we know the data "isn't right".
> 
> Using a version-specific CRC sounds silly.  Put the version information in the 
> packet.

The packet routed to a particular interface is *supposed* to
conform to "version X" of an interface.  There are different stubs
generated for different versions of EACH interface.  The OCL for
the interface defines (and is used to check) the form of that
interface to that service/mechanism.

The parameters are checked on the client side -- why tie up the
transport medium with data that is inappropriate (redundant)
to THAT interface?  Why tie up the server verifying that data?
The stub generator can perform all of those checks automatically
and CONSISTENTLY based on the OCL definition of that version
of that interface (because developers make mistakes).

So, at the instant you schedule the marshalled data for transmission,
you *know* the parameters are "appropriate" and compliant with
the constraints of THAT version of THAT interface.

Now, you have to ensure the packet doesn't get corrupted (altered) in
transmission.  If it remains intact, then there is no need to check
the parameters on the server side.

NONE OF THE PARAMETERS... including the (implied) "interface version" field!

Yet, folks make mistakes.  So, you want some additional reassurance
that this is at least intended for this version of the interface,
ESPECIALLY IF THAT CAN BE MADE AVAILABLE FOR ZERO COST (i.e., check
to see if the residual is 0xDEADBEEF instead of 0xB16B00B5).

Why burden the packet with a "protocol version" parameter?

So, use a version-specific CRC on the packet.  If it fails, then
either the data in the packet has been corrupted (which could just
as easily have involved an embedded "interface version" parameter);
or the packet was formed with the wrong CRC.

If the CRC is correct FOR THAT VERSION OF THE PROTOCOL, then
why bother looking at a "protocol version" parameter?  Would
you ALSO want to verify all the rest of the parameters?

>> I *could* put a header saying "this is version 4.2".  And, that
>> tells me nothing about the integrity of the rest of the data.
>> OTOH, ensuring the CRC reflects "4.2" does -- it the recipient
>> expects it to be so.
> 
> Now you don't know if the data is corrupted, or for the wrong version - or 
> occasionally, corrupted /and/ the wrong version but passing the CRC anyway.

You don't know if the parameters have been corrupted in a manner that
allows a packet intended for the correct interface to appear as correct.
What's your point?

> Unless you are absolutely desperate to save every bit you can, your system will 
> be simpler, clearer, and more reliable if you separate your purposes.

Yes.  You verify the correct interface at the client side -- where
it is invoked by the client and enforced in the OCL generated stub.
Thereafter, the server is concerned with corruption during transport
and the version specific CRC just gives another reassurance of
correct version without adding another cost.

[Imagine EVERY subroutine function call in your system having
such overhead.  Would you want to push an "interface version"
onto the stack along with all of the arguments for that
subr/ftn?  Or, would you just hope everything was intact?]

>>>>>> You can also "salt" the calculation so that the residual
>>>>>> is deliberately nonzero.  So, for example, "success" is
>>>>>> indicated by a residual of 0x474E.  :>
>>>>>
>>>>> Again, pointless.
>>>>>
>>>>> Salt is important for security-related hashes (like password hashes), not 
>>>>> for integrity checks.
>>>>
>>>> You've missed the point.  The correct "sum" can be anything.
>>>> Why is "0" more special than any other value?  As the value is
>>>> typically meaningless to anything other than the code that verifies
>>>> it, you couldn't look at an image (or the output of the verifier)
>>>> and gain anything from seeing that obscure value.
>>>
>>> Do you actually know what is meant by "salt" in the context of hashes, and 
>>> why it is useful in some circumstances?  Do you understand that "salt" is 
>>> added (usually prepended, or occasionally mixed in in some other way) to the 
>>> data /before/ the hash is calculated?
>>
>> What term would you have me use to indicate a "bias" applied to a CRC
>> algorithm?
> 
> Well, first I'd note that any kind of modification to the basic CRC algorithm 
> is pointless from the viewpoint of its use as an integrity check.  (There have 
> been, mostly historically, some justifications in terms of implementation 
> efficiency.  For example, bit and byte re-ordering could be done to suit 
> hardware bit-wise implementations.)
> 
> Otherwise I'd say you are picking a specific initial value if that is what you 
> are doing, or modifying the final value (inverting it or xor'ing it with a 
> fixed value).  There is, AFAIK, no specific terms for these - and I don't see 
> any benefit in having one.  Misusing the term "salt" from cryptography is 
> certainly not helpful.

Salt just ensures that you can differentiate between functionally identical
values.  I.e., in a CRC, it differentiates between the "0x0000" that CRC-1
generates from the "0x0000" that CRC-2 generates.

You don't see the parallel to ensuring that *my* use of "Passw0rd" is
encoded in a different manner than *your* use of "Passw0rd"?

>> See the RMI desciption.
> 
> I'm sorry, I have no idea what "RMI" is or where it is described. You've 
> mentioned that abbreviation twice, but I can't figure it out.

<https://en.wikipedia.org/wiki/RMI>
<https://en.wikipedia.org/wiki/OCL>

Nothing magical with either term.

>> OTOH, "salting" the calculation so that it is expected to yield
>> a value of 0x13 means *those* situations will be flagged as errors
>> (and a different set of situations will sneak by, undetected).
> 
> And that gives you exactly /zero/ benefit.

See above.

> You run your hash algorithm, and check for the single value that indicates no 
> errors.  It does not matter if that number is 0, 0x13, or - often more 
-----------^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

As you've admitted, it doesn't matter.  So, why wouldn't I opt to have
an algorithm for THIS interface give me a result that is EXPECTED
for this protocol?  What value picking "0"?

> conveniently - the number attached at the end of the image as the expected 
> result of the hash of the rest of the data.

>>> To be more accurate, the chances of them passing unnoticed are of the order 
>>> of 1 in 2^n, for a good n-bit check such as a CRC check. Certain types of 
>>> error are always detectable, such as single and double bit errors.  That is 
>>> the point of using a checksum or hash for integrity checking.
>>>
>>> /Intentional/ changes are a different matter.  If a hacker changes the 
>>> program image, they can change the transmitted hash to their own calculated 
>>> hash.  Or for a small CRC, they could change a different part of the image 
>>> until the original checksum matched - for a 16-bit CRC, that only takes 
>>> 65,535 attempts in the worst case.
>>
>> If the approach used is "typical", then you need far fewer attempts to
>> produce a correct image -- without EVER knowing where the CRC is stored.
> 
> It is difficult to know what you are trying to say here, but if you believe 
> that different initial values in a CRC algorithm makes it harder to modify an 
> image to make it pass the integrity test, you are simply wrong.

Of course it does!  You don't KNOW what to expect -- unless you've identified
where the test is performed in the code and the result stored/checked.  If
you assume the residual will be 0 and make an attempt to generate a new
checksum that yields 0 and it doesn't work FIRST TIME, then, by definition,
it is HARDER (more work is required -- even if not *conceptually* more
*difficult*).

*My* example use of the different salt is for a different purpose.
And, isn't meant as a deterrent to any developer/attacker but, rather,
simply to ensure the transmission of the packet is intact AND carries
some reassurance that it is in the correct format.

>>> That is why you need to distinguish between the two possibilities.  If you 
>>> don't have to worry about malicious attacks, a 32-bit CRC takes a dozen 
>>> lines of C code and a 1 KB table, all running extremely efficiently.  If 
>>> security is an issue, you need digital signatures - an RSA-based signature 
>>> system is orders of magnitude more effort in both development time and in 
>>> run time.
>>
>> It's considerably more expensive AND not fool-proof -- esp if the
>> attacker knows you are signing binaries.  "OK, now I need to find
>> WHERE the signature is verified and just patch that "CALL" out
>> of the code".
> 
> I'm not sure if that is a straw-man argument, or just showing your ignorance of 
> the topic.  Do you really think security checks are done by the program you are 
> trying to send securely?  That would be like trying to have building security 
> where people entering the building look at their own security cards.

Do YOU really think we all design applications that run in PCs where some
CLOSED OS performs these tests in a manner that can't be subverted?
*WE* (tend to) write ALL the code in the products developed, here.
So, whether it's the POST WE wrote that is performing the test or
the loader WE wrote, it's still *our* program.

Yes, we ARE looking at our own security cards!

Manufacturers *try* to hide ("obscurity") details of these mechanisms
in an attempt to improve effective security.  But, there's nothing
that makes these guarantees.

Give me the sources for Windows (Linux, *BSD, etc.) and I can
subvert all the state-of-the-art digital signing used to ensure
binaries aren't altered.  Nothing *outside* the box is involved
so, by definition, everything I need has to reside *in* the box.

DataI/O was always paranoid about their software/firmware.
They rely on custom silicon to "protect" their investment.
But, use a COTS CPU to execute the code!

So, pull the MC68K out of its socket and plug in an emulator.
Capture the execution trace and you know exactly what the
instruction  stream is/was -- despite it's encoding on the
distribution media.

>>>> I altered the test procedure for a
>>>> piece of military gear we were building simply to skip some lengthy tests 
>>>> that I *knew* would pass (I don't want to inject an extra 20 minutes of 
>>>> wait time
>>>> just to get through a lengthy test I already know works before I can get
>>>> to the test of interest to me, now.
>>>>
>>>> I failed to undo the change before the official signoff on the device.
>>>>
>>>> The only evidence of this was the fact that I had also patched the
>>>> startup message to say "Go for coffee..." -- which remained on the
>>>> screen for the duration of the lengthy (even with the long test
>>>> elided) procedure...
>>>>
>>>> ..which alerted folks to the fact that this *probably* wasn't the
>>>> original image.  (The computer running the test suite on the DUT had
>>>> no problem accepting my patched binary)
>>>
>>> And what, exactly, do you think that anecdote tells us about CRC checks for 
>>> image files?  It reminds us that we are all fallible, but does no more than 
>>> that.
>>
>> That *was* the point.  Because the folks who designed the test computer
>> relied on common techniques to safeguard the image.
> 
> There was a human error - procedures were not good enough, or were not 
> followed.  It happens, and you learn from it and make better procedures.  The 
> fault was in what people did, not in an automated integrity check.  It is 
> completely unrelated.

It shows that the check was designed without consideration of how
it might be subverted.  This is the most common flaw in all
security schemes -- failing to consider an attack/fault vector.

The vendor assumed no one would deliberately alter the test
procedure.  That anyone running it would willingly sit through an
extra half hour of tests ALREADY KNOWN TO PASS instead of opting
to find a way to skip that (because the test designer only consider
"sell off" when designing the test and not *debug* and the test
platform didn't provide hooks to facilitate that, either!!)

I unplugged a cable between two pieces of equipment that I had
never seen before to subvert a security mechanism in a product.
Because the designers never considered the fact that someone
might do that!

Security is no different from any other "solution".  You test
the divisor before a calculation because you reasonably expect to
encounter "unfortunate" values and don't want the operation to
fail.

>> The counterfeiting example I cited indicates how "obscurity/secrecy"
>> is far more effective (yet you dismiss it out-of-hand).
> 
> No, it does nothing of the sort.  There is no connection at all.

The counterfeiter lost the lawsuit because he was unaware (obscurity)
of the hidden SECURITY measures in the product design.  This proven
by his attempts to defeat the OBVIOUS ones!

>>>>> CRC's alone are useless.  All the attacker needs to do after modifying the 
>>>>> image is calculate the CRC themselves, and replace the original checksum 
>>>>> with their own.
>>>>
>>>> That assumes the "alterer" knows how to replace the checksum, how it
>>>> is computed, where it is embedded in the image, etc.  I modified the Compaq
>>>> portable mentioned without ever knowing where the checksum was store
>>>> or *if* it was explicitly stored.  I had no desire to disassemble the
>>>> BIOS ROMs (though could obviously do so as there was no "proprietary
>>>> hardware" limiting access to their contents and the instruction set of
>>>> the processor is well known!).
>>>>
>>>> Instead, I did this by *guessing* how they would implement such a check
>>>> in a bit of kit from that era (ERPOMs aren't easily modified by malware
>>>> so it wasn't likely that they would go to great lengths to "protect" the
>>>> image).  And, if my guess had been incorrect, I could always reinstall
>>>> the original EPROMs -- nothing lost, nothing gained.
>>>>
>>>> Had much experience with folks counterfeiting your products and making
>>>> "simple" changes to the binaries?  Like changing the copyright notice
>>>> or splash screen?
>>>>
>>>> Then, bringing the (accused) counterfeit of YOUR product into a courtroom
>>>> and revealing the *hidden* checksum that the counterfeiter wasn't aware of?
>>>>
>>>> "Gee, why does YOUR (alleged) device have *my* name in it -- in addition
>>>> to behaving exactly like mine??"
>>>>
>>>> [I guess obscurity has its place!]
>>>
>>> Security by obscurity is not security.  Having a hidden signature or other 
>>> mark can be useful for proving ownership (making an intentional mistake is 
>>> another common tactic - such as commercial maps having a few subtle spelling 
>>> errors). But that is not security.
>>
>> Of course it is!  If *you* check the "hidden signature" at runtime
>> and then alter "your" operation such that an altered copy fails
>> to perform properly, then then you have secured it.
> 
> That is not security.  "Security" means that the program that starts the 
> updated program checks the /entire/ image according to its digital signature, 
> and rejects it /entirely/ if it does not match.

No, that's *your* naive assumption of security.  It's why such attempts
invariably fail; they are "drop dead" implementations that make it
clear to anyone trying to subvert that security that their
efforts have not (yet) succeeded.

The goal is to prevent the program/device from being used without
authorization/compensation.  If it KILLS the user as a result of some
hidden feature, it has met its goal -- even if a draconian approach.
If it *pretends* to be doing what you want --- and then fails to
complete some later step -- it is similarly preventing unauthorized
use (and tying up a lot of your time, in the process).

If you want to ensure the image isn't *corrupt* (which could
lead to failures that could invite lawsuits, etc.), then you
are concerned with INTEGRITY.

> What you are talking about here is the sort of cat-and-mouse nonsense computer 
> games producers did with intentional disk errors to stop copying.  It annoys 
> legitimate users and does almost nothing to hinder the bad guys.

Because it was a bad solution that was fairly obvious in its presence:
"I can't copy this disk!  Let me buy Copy2PC..."

The same applies to most licensing schemes and other "tamper
detection" mechanisms.

>> Would you want to use a check-writing program if the account
>> balances it maintains were subtly (but not consistently)
>> incorrect?
> 
> Again, you make no sense.  What has this got to do with integrity checks or 
> security?

If you;re selling check writing software and want to prevent
FlyByNight Accounting Software, Inc. from stealing your
product and reselling it as your own, a great way to prevent that
is to ensure THEIR copy of the product causes accounting errors
that are hard to notice. Their customers will (eventually)
complain that THEIR product is buggy.  But, yours isn't!

If your goal is to track your checks accurately, you're
likely not going to want to wonder what yet-to-be-discovered
errors exist in the "books" that THEIR software has been
maintaining for you.

The original vendor has secured his product against tampering.

>> Checking signatures, CRCs, licensing schemes, etc. all are used
>> in a "drop dead" fashion so considerably easier to defeat.
>> Witness the number of "products" available as warez...
> 
> Look, it is all /really/ simple.  And the year is 2023, not 1973.

Yes!  And it is considerably easier to subvert naive mechanisms
AND SHARE YOUR HACKS!

> If you want to check the integrity of a file against accidental changes, a CRC 
> is usually fine.

As is a CRC on a network packet.  Without having to double-check the
contents of that packet after it has been verified on the sending side!

> If you want security, and to protect against malicious changes, use a digital 
> signature.  This must be checked by the program that /starts/ the updated code, 
> or that downloaded and stored it - not by the program itself!

And who wrote THAT program?  Where is it, physically?  Is there some device
OUTSIDE of the device that you've built that securely performs these
checks?

>> Only answer with a support contract.
> 
> Oh, sure - the amateurs who have some of the information but not enough 
> details, skill or knowledge to get things working will /never/ fill forums with 
> questions, complaints or bad reviews that bother your support staff or scare 
> away real sales.

A forum doesn't have to be "public".  FUD can scare off real sales
even in the total absence of information (or knowledge).

Your goal should always be to produce a good product that does
what it claims to do.  And, rely on the happiness of your
customers to directly (or indirectly) generate additional sales.

I've never "advertised" my services.  I'm actually pretty hard to get
in touch with!  Yet, clients never had a hard time finding me -- through
other clients who were happy with my work.  As they likely weren't
direct competitors to the original clients, they had nothing to
fear (lose) from sharing me, as a resource.

Similarly, a customer making widgets that employ some feature of
your device likely has little to lose by sharing his (good or bad)
experiences with another POTENTIAL customer (making wodjets).
And, likely can benefit from the goodwill he receives from that
other customer as well as from *you* ("Thanks for recommending
us to him!").  And, ensures a continued demand for your products
so you continue to be available for HIS needs!

>>> They are limited by idiotic government export restrictions made by ignorant 
>>> politicians who don't understand cryptography.
>>
>> Protections don't always have to be cryptographic. 
> 
> Correct, but - as with a lot of what you write - completely irrelevant to the 
> subject at hand.
> 
> Why can't companies give out information about the security systems used in 
> their microcontrollers (for example) ?  Because some geriatric ignoramuses 
> think banning "export" of such information to certain countries will stop those 
> countries knowing about security and cryptography.

Do you really think that's the sole reason for all the "secrecy" and NDAs?
I've had to sit with gummit folks and sort out what parts of our technology
could LEGALLY be exported.  Even to our partners in the UK!  Some of it
makes sense ("Nothing goes to Libya!").  Some is bogus.

And, thinking that you can put up a wall that is impermeable is a joke.
Just like printing PGP in book form and selling books overseas.

Or, hiring someone who worked for Company X.  Or, bribing someone
to make a photocopy of <whatever>.

But, this doesn't mean one should ENCOURAGE dissemination of things
that may have special security/economic value.  "Delay" often has
as much value as "deter".

A friend who designed arcade pieces recounted how he was contacted by a guy
who had disassembled ~40KB of (hand-written) code in one of his products.
He had even uncovered latent bugs (!) in the code.

But, his efforts were so "late" that the product had long ago lost
commercial value.  So, it may have been flattering that someone
would invest that much time in such an endeavor.  But, little else.

Nowadays, tools would make that a trivial undertaking.  And, the
possibility of easily enlisting others in the effort (without
resorting to clandestine channels).  OTOH, projects are now
considerably larger (orders of magnitude).  OToOH, much current
work in done in HLLs (so tools can recognize their code genrator
patterns) and with "standard" libraries; I can recognize a call to
printf without decompiling any 9of the code -- folks aren't
likely going to replace "%d" with "?b" just to obscure functionality!

>>> Some things benefit from being kept hidden, or under restricted access. The 
>>> details of the CRC algorithm you use to catch accidental errors in your 
>>> image file is /not/ one of them.  If you think hiding it has the remotest 
>>> hint of a benefit, you are doing things wrong - you need a /security/ check, 
>>> not a simple /integrity/ check.
>>>
>>> And then once you have switched to a security check - a digital signature - 
>>> there's no need to keep that choice hidden either, because it is the /key/ 
>>> that is important, not the type of lock.
>>
>> Again, meaningless if the attacker can interfere with the *enforcement*
>> of that check.  Using something "well known" just means he already knows
>> what to look for in your code.  Or, how to interfere with your
>> intended implementation in ways that you may have not anticipated
>> (confident that your "security" can't be MATHEMATICALLY broken).
>>
> 
> If the attacker can interfere with the enforcement of the check, then it 
> doesn't matter what checks you have.  Keeping the design of a building's locks 
> secret does not help you if the bad guys have bribed the security guard 
> /inside/ the building!

But, if that's the only way to subvert the secrets of those locks,
then you only have to worry about keeping that security guard "happy".

>> "Here's my source code.  Here are my schematics.  Here's the
>> name of the guy who oversees production (bribe him to gain
>> access to the keys stored in the TPM).  Now, what are you
>> gonna *do* with all that?"
> 
> The first two should be fine - if people can break your security after looking 
> at your source code or schematics, your security is /bad/.  As for the third 
> one, if they can break your security by going through the production guy, your 
> production procedures are bad.

You can change your production procedures without having to redesign your
product.  You don't want to embrace a solution/technology that may soon/later
be subverted (e.g., SHA1) and have to redesign portions of your product
(which may already be deployed) to "fix".

IMO, this is the downside of modern cryptography -- if you have a product
with any significant lifespan and "exposure".  You never know when the
next "uncrackable" algorithm will fall.  And, when someone might opt to
marshall a community's resources to attack a particular implementation.

Attacks that used to be considered "nation-state scale" are quickly
becoming "big business scale" and even "network of workstations scale".
So, any implementation that *shares* a key across a product line
is vulnerable to the entire product line being compromised when/if
that key is disclosed/broken.

[I generate unique keys for each device on the customer's site
using a dedicated (physically secure) interface so even the manufacturer
doesn't know what they are.  Crack one (possibly by physically attacking
the device and microprobing the die) and all you get it that one
device -- and whatever *its* role in the system may have been.]

[toc] | [prev] | [next] | [standalone]


#31839

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-24 16:37 +0200
Message-ID<u2646q$clg2$1@dont-email.me>
In reply to#31837
On 24/04/2023 09:32, Don Y wrote:
> On 4/22/2023 7:57 AM, David Brown wrote:
>>>> However, in almost every case where CRC's might be useful, you have 
>>>> additional checks of the sanity of the data, and an all-zero or 
>>>> all-one data block would be rejected.  For example, Ethernet packets 
>>>> use CRC for integrity checking, but an attempt to send a packet type 
>>>> 0 from MAC address 00:00:00:00:00:00 to address 00:00:00:00:00:00, 
>>>> of length 0, would be rejected anyway.
>>>
>>> Why look at "data" -- which may be suspect -- and *then* check its CRC?
>>> Run the CRC first.  If it fails, decide how you are going to proceed
>>> or recover.
>>
>> That is usually the order, yes.  Sometimes you want "fail fast", such 
>> as dropping a packet that was not addressed to you (it doesn't matter 
>> if it was received correctly but for someone else, or it was addressed 
>> to you but the receiver address was corrupted - you are dropping the 
>> packet either way).  But usually you will run the CRC then look at the 
>> data.
>>
>> But the order doesn't matter - either way, you are still checking for 
>> valid data, and if the data is invalid, it does not matter if the CRC 
>> only passed by luck or by all zeros.
> 
> You're assuming the CRC is supposed to *vouch* for the data.
> The CRC can be there simply to vouch for the *transport* of a
> datagram.

I am assuming that the CRC is there to determine the integrity of the 
data in the face of possible unintentional errors.  That's what CRC 
checks are for.  They have nothing to do with the content of the data, 
or the type of the data package or image.

As an example of the use of CRC's in messaging, look at Ethernet frames:

<https://en.wikipedia.org/wiki/Ethernet_frame>

The CRC  does not care about the content of the data it protects.

> 
> So, use a version-specific CRC on the packet.  If it fails, then
> either the data in the packet has been corrupted (which could just
> as easily have involved an embedded "interface version" parameter);
> or the packet was formed with the wrong CRC.
> 
> If the CRC is correct FOR THAT VERSION OF THE PROTOCOL, then
> why bother looking at a "protocol version" parameter?  Would
> you ALSO want to verify all the rest of the parameters?
> 

I'm sorry, I simply cannot see your point.  Identifying the version of a 
protocol, or other protocol type information, is a totally orthogonal 
task to ensuring the integrity of the data.  The concepts should be 
handled separately.


>>> What term would you have me use to indicate a "bias" applied to a CRC
>>> algorithm?
>>
>> Well, first I'd note that any kind of modification to the basic CRC 
>> algorithm is pointless from the viewpoint of its use as an integrity 
>> check.  (There have been, mostly historically, some justifications in 
>> terms of implementation efficiency.  For example, bit and byte 
>> re-ordering could be done to suit hardware bit-wise implementations.)
>>
>> Otherwise I'd say you are picking a specific initial value if that is 
>> what you are doing, or modifying the final value (inverting it or 
>> xor'ing it with a fixed value).  There is, AFAIK, no specific terms 
>> for these - and I don't see any benefit in having one.  Misusing the 
>> term "salt" from cryptography is certainly not helpful.
> 
> Salt just ensures that you can differentiate between functionally identical
> values.  I.e., in a CRC, it differentiates between the "0x0000" that CRC-1
> generates from the "0x0000" that CRC-2 generates.

Can we agree that this is called an "initial value", not "salt" ?

> 
> You don't see the parallel to ensuring that *my* use of "Passw0rd" is
> encoded in a different manner than *your* use of "Passw0rd"?

No.  They are different things.

An important difference is that adding "salt" to a password hash is an 
important security feature.  Picking a different initial value for a CRC 
instead of having appropriate protocol versioning in the data (or a 
surrounding envelope) is a misfeature.

The second difference is the purpose of the hashing.  The CRC here is 
for data integrity - spotting mistakes in the data during transfer or 
storage.  The hash in a password is for security, avoiding the password 
ever being transmitted or stored in plain text.

Any coincidence in the the way these might be implemented is just that - 
coincidence.


> 
>>> See the RMI desciption.
>>
>> I'm sorry, I have no idea what "RMI" is or where it is described. 
>> You've mentioned that abbreviation twice, but I can't figure it out.
> 
> <https://en.wikipedia.org/wiki/RMI>
> <https://en.wikipedia.org/wiki/OCL>
> 
> Nothing magical with either term.

I looked up RMI on Wikipedia before asking, and saw nothing of relevance 
to CRC's or checksums.  I noticed no mention of "OCL" in your posts, and 
looking it up on Wikipedia gives no clues.

So for now, I'll assume you don't want anyone to know what you meant and 
I can safely ignore anything you write in connection with the terms.

> 
>>> OTOH, "salting" the calculation so that it is expected to yield
>>> a value of 0x13 means *those* situations will be flagged as errors
>>> (and a different set of situations will sneak by, undetected).
>>
>> And that gives you exactly /zero/ benefit.
> 
> See above.

I did.  Zero benefit.

Actually, it is worse than useless - it makes it harder to identify the 
protocol, and reduces the information content of the CRC check.

> 
>> You run your hash algorithm, and check for the single value that 
>> indicates no errors.  It does not matter if that number is 0, 0x13, or 
>> - often more 
> -----------^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> 
> As you've admitted, it doesn't matter.  So, why wouldn't I opt to have
> an algorithm for THIS interface give me a result that is EXPECTED
> for this protocol?  What value picking "0"?
> 

A /single/ result does not matter (other than needlessly complicating 
things).  Having multiple different valid results /does/ matter.

>>>> That is why you need to distinguish between the two possibilities.  
>>>> If you don't have to worry about malicious attacks, a 32-bit CRC 
>>>> takes a dozen lines of C code and a 1 KB table, all running 
>>>> extremely efficiently.  If security is an issue, you need digital 
>>>> signatures - an RSA-based signature system is orders of magnitude 
>>>> more effort in both development time and in run time.
>>>
>>> It's considerably more expensive AND not fool-proof -- esp if the
>>> attacker knows you are signing binaries.  "OK, now I need to find
>>> WHERE the signature is verified and just patch that "CALL" out
>>> of the code".
>>
>> I'm not sure if that is a straw-man argument, or just showing your 
>> ignorance of the topic.  Do you really think security checks are done 
>> by the program you are trying to send securely?  That would be like 
>> trying to have building security where people entering the building 
>> look at their own security cards.
> 
> Do YOU really think we all design applications that run in PCs where some
> CLOSED OS performs these tests in a manner that can't be subverted?

Do you bother to read my posts at all?  Or do you prefer to make up 
things that you imagine I write, so that you can make nonsensical 
attacks on them?  Certainly there is no sane reading of my posts 
(written and sent from an /open/ OS) where "do not rely on security by 
obscurity" could be taken to mean "rely on obscured and closed platforms".

> *WE* (tend to) write ALL the code in the products developed, here.
> So, whether it's the POST WE wrote that is performing the test or
> the loader WE wrote, it's still *our* program.
> 
> Yes, we ARE looking at our own security cards!
> 
> Manufacturers *try* to hide ("obscurity") details of these mechanisms
> in an attempt to improve effective security.  But, there's nothing
> that makes these guarantees.

Why are you trying to "persuade" me that manufacturer obscurity is a bad 
thing?  You have been promoting obscurity of algorithms as though it 
were helpful for security - I have made clear that it is not.  Are you 
getting your own position mixed up with mine?

> 
> Give me the sources for Windows (Linux, *BSD, etc.) and I can
> subvert all the state-of-the-art digital signing used to ensure
> binaries aren't altered.  Nothing *outside* the box is involved
> so, by definition, everything I need has to reside *in* the box.

No, you can't.  The sources for Linux and *BSD /are/ all freely 
available.  The private signing keys used by, for example, Red Hat or 
Debian, are /not/ freely available.  You cannot make changes to a Red 
Hat or Debian package that will pass the security checks - you are 
unable to sign the packages.

This is precisely because something /outside/ the box /is/ involved - 
the private half of the public/private key used for signing.  The public 
half - and all the details of the algorithms - is easily available to 
let people verify the signature, but the private half is kept secret.


(Sorry, but I've skipped and snipped the rest.  I simply don't have time 
to go through it in detail.  If others find it useful or interesting, 
that's great, but there has to be limits somewhere.)


[toc] | [prev] | [next] | [standalone]


#31901

FromDon Y <blockedofcourse@foo.invalid>
Date2023-05-03 00:15 -0700
Message-ID<u2t1mh$16mg5$1@dont-email.me>
In reply to#31839
On 4/24/2023 7:37 AM, David Brown wrote:
> On 24/04/2023 09:32, Don Y wrote:
>> On 4/22/2023 7:57 AM, David Brown wrote:
>>>>> However, in almost every case where CRC's might be useful, you have 
>>>>> additional checks of the sanity of the data, and an all-zero or all-one 
>>>>> data block would be rejected.  For example, Ethernet packets use CRC for 
>>>>> integrity checking, but an attempt to send a packet type 0 from MAC 
>>>>> address 00:00:00:00:00:00 to address 00:00:00:00:00:00, of length 0, would 
>>>>> be rejected anyway.
>>>>
>>>> Why look at "data" -- which may be suspect -- and *then* check its CRC?
>>>> Run the CRC first.  If it fails, decide how you are going to proceed
>>>> or recover.
>>>
>>> That is usually the order, yes.  Sometimes you want "fail fast", such as 
>>> dropping a packet that was not addressed to you (it doesn't matter if it was 
>>> received correctly but for someone else, or it was addressed to you but the 
>>> receiver address was corrupted - you are dropping the packet either way).  
>>> But usually you will run the CRC then look at the data.
>>>
>>> But the order doesn't matter - either way, you are still checking for valid 
>>> data, and if the data is invalid, it does not matter if the CRC only passed 
>>> by luck or by all zeros.
>>
>> You're assuming the CRC is supposed to *vouch* for the data.
>> The CRC can be there simply to vouch for the *transport* of a
>> datagram.
> 
> I am assuming that the CRC is there to determine the integrity of the data in 
> the face of possible unintentional errors.  That's what CRC checks are for.  
> They have nothing to do with the content of the data, or the type of the data 
> package or image.

Exactly.  And, a CRC on *a* protocol can use ANY ALGORITHM that the protocol
defines.  Not some "canned one-size fits all" approach.

> As an example of the use of CRC's in messaging, look at Ethernet frames:
> 
> <https://en.wikipedia.org/wiki/Ethernet_frame>
> 
> The CRC  does not care about the content of the data it protects.

AND, if the packet yielded an incorrect CRC, you can assume the
data was corrupt... OR, you are looking at a different protocol
and MISTAKING it for something that you *think* it might be.

If I produce a stream of data, can you tell me what the checksum
for THAT stream *should* be?  You have to either be told what
it is (and have a way of knowing what the checksum SHOULD be)
*or* have to make some assumptions about it.

If you have assumed wrong *or* if the data has been corrupt, then
the CRC should fail.  You don't care why it failed -- because you
can't do anything about it.  You just know that you can't use the data
in the way you THOUGHT it could be used.

>> So, use a version-specific CRC on the packet.  If it fails, then
>> either the data in the packet has been corrupted (which could just
>> as easily have involved an embedded "interface version" parameter);
>> or the packet was formed with the wrong CRC.
>>
>> If the CRC is correct FOR THAT VERSION OF THE PROTOCOL, then
>> why bother looking at a "protocol version" parameter?  Would
>> you ALSO want to verify all the rest of the parameters?
> 
> I'm sorry, I simply cannot see your point.  Identifying the version of a 
> protocol, or other protocol type information, is a totally orthogonal task to 
> ensuring the integrity of the data.  The concepts should be handled separately.

It is.  A packet using protocol XYZ is delivered to port ABC.
Port ABC *only* handles protocol XYZ.  Anything else arriving there,
with a potentially different checksum, is invalid.  Even if, for example,
byte number 27 happens to have the correct "magic number" for that
protocol.

Because the message doesn't obey the rules defined by the protocol
FOR THAT PORT.  What do I gain by insisting that byte number 27 must
be 0x5A that the CRC doesn't already tell me?

You are assuming the CRC has to identify the protocol.  I didn't say that.
All I said was the CRC has to be correct for THAT protocol.

You likely don't use the same algorithm to compute the checksum of
a boot image as you do to verify the integrity of a ethernet datagram.
So, if you were presented with a stream of data, you wouldn't
arbitrarily decide to try different CRCs to see which yielded correct
results and, from that, *infer* the nature of the message.

Why would you think I wouldn't expect *a* particular protocol to use
a particular CRC?

>>>> What term would you have me use to indicate a "bias" applied to a CRC
>>>> algorithm?
>>>
>>> Well, first I'd note that any kind of modification to the basic CRC 
>>> algorithm is pointless from the viewpoint of its use as an integrity check.  
>>> (There have been, mostly historically, some justifications in terms of 
>>> implementation efficiency.  For example, bit and byte re-ordering could be 
>>> done to suit hardware bit-wise implementations.)
>>>
>>> Otherwise I'd say you are picking a specific initial value if that is what 
>>> you are doing, or modifying the final value (inverting it or xor'ing it with 
>>> a fixed value).  There is, AFAIK, no specific terms for these - and I don't 
>>> see any benefit in having one.  Misusing the term "salt" from cryptography 
>>> is certainly not helpful.
>>
>> Salt just ensures that you can differentiate between functionally identical
>> values.  I.e., in a CRC, it differentiates between the "0x0000" that CRC-1
>> generates from the "0x0000" that CRC-2 generates.
> 
> Can we agree that this is called an "initial value", not "salt" ?

It depends on how you implement it.  The point is to produce
different results for the same polynmomial.

>> You don't see the parallel to ensuring that *my* use of "Passw0rd" is
>> encoded in a different manner than *your* use of "Passw0rd"?
> 
> No.  They are different things.
> 
> An important difference is that adding "salt" to a password hash is an 
> important security feature.  Picking a different initial value for a CRC 
> instead of having appropriate protocol versioning in the data (or a surrounding 
> envelope) is a misfeature.

And you don't see that verifying that a packet of data received at
port ABC that should only see the checksum associated with protocol
XYZ as being similarly related?

Why not just assume the lower level protocols are sufficient to
guarantee reliable delivery and, if something arrives at port ABC
then, by definition, it must be intact (not corrupt) and, as
nothing other than protocol XYZ *should* target that port, why
even bother checking magic numbers in a protocol packet?

You build these *superfluous* tests into products to ensure their
integrity -- by catching ANYTHING that "can't happen" (yet
somehow does)

> The second difference is the purpose of the hashing.  The CRC here is for data 
> integrity - spotting mistakes in the data during transfer or storage.  The hash 
> in a password is for security, avoiding the password ever being transmitted or 
> stored in plain text.
> 
> Any coincidence in the the way these might be implemented is just that - 
> coincidence.
> 
>>>> See the RMI desciption.
>>>
>>> I'm sorry, I have no idea what "RMI" is or where it is described. You've 
>>> mentioned that abbreviation twice, but I can't figure it out.
>>
>> <https://en.wikipedia.org/wiki/RMI>
>> <https://en.wikipedia.org/wiki/OCL>
>>
>> Nothing magical with either term.
> 
> I looked up RMI on Wikipedia before asking, and saw nothing of relevance to 
> CRC's or checksums.

How do you think the marshalled arguments get from device A to (remote)
device B?  And, the result(s) from device B back to device A?

Obviously *some* form of communication medium.  So, some potential for
data to be corrupted (or altered!) in transit.  Along with other
data streams to compete for those endpoints.

Imagine invoking a function and, between the actual construction of the
stack frame and the first line of code in the targeted function, "something"
can interfere with the data you're trying to pass (and results you're
hoping to eventually receive) as well as the actual function being targeted!

You don't worry about this because the compiler handles all of the machinery
AND it relies on the CPU being well-behaved; nothing can sneak in and
disturb the address/data -busses or alter register contents during this
process.

If, OTOH, such a possibility existed (as is the case with RPC/RMI), then
you would want the compiler to generate the machinery to ensure the
arguments get to the correct function and for the function to be able to
ensure that the arguments are actually intended for it.

If any of these things failed to happen, you'd panic() -- because there's
nothing you can do, at that point.  You certainly can't fix any corrupted
values and can't deduce where they were intended to go (given that all
of that information can be just as corrupt).

With RPC/RMI, you can at least *know* that the "function linkage" failed
to operate as expected ON THIS INVOCATION.  Because the RPC/RMI can
return a result indicating whether the linkage was intact *and*, if
so, the result of the actual function invocation.

If you deliver every packet to a single port, then the process listening
to that port has to demultiplex incoming messages to determine the server-side
stub to invoke for that message instance.  You would likely use a standardized
protocol because you don't know anything about the incoming message -- except
that it is *supposed* to target a "remote procedure" (*local* to this node).

OTOH, if you target each particular remote function/procedure/method to
a function/procedure/method-SPECIFIC port, then how you handle "messages"
for one function need have no bearing on how you handle them for others.
And, you can exploit this as an added test to ensure the message you
are receiving at port JKL actually *appears* to be intended for port
JKL and not an accidental misdirect of a message intended for some
other port.

> I noticed no mention of "OCL" in your posts, and looking 

You need to read more carefully.

---8<---8<---
 >>>> I can't think of any use-cases where you would be passing around a block of
 >>>> "pure" data that could reasonably take absolutely any value, without any
 >>>> type of "envelope" information, and where you would think a CRC check is
 >>>> appropriate.
 >>>
 >>> I append a *version specific* CRC to each packet of marshalled data
 >>> in my RMIs.  If the data is corrupted in transit *or* if the
 >>> wrong version API ends up targeted, the operation will abend
 >>> because we know the data "isn't right".
 >>
 >> Using a version-specific CRC sounds silly.  Put the version information in
 >> the packet.
 >
 > The packet routed to a particular interface is *supposed* to
 > conform to "version X" of an interface.  There are different stubs
 > generated for different versions of EACH interface.  The OCL for
 > the interface defines (and is used to check) the form of that
 > interface to that service/mechanism.
 >
 > The parameters are checked on the client side -- why tie up the
 > transport medium with data that is inappropriate (redundant)
 > to THAT interface?  Why tie up the server verifying that data?
 > The stub generator can perform all of those checks automatically
 > and CONSISTENTLY based on the OCL definition of that version
 > of that interface (because developers make mistakes).
 >
 > So, at the instant you schedule the marshalled data for transmission,
 > you *know* the parameters are "appropriate" and compliant with
 > the constraints of THAT version of THAT interface.
 >
 > Now, you have to ensure the packet doesn't get corrupted (altered) in
 > transmission.  If it remains intact, then there is no need to check
 > the parameters on the server side.
 >
 > NONE OF THE PARAMETERS... including the (implied) "interface version" field!
 >
 > Yet, folks make mistakes.  So, you want some additional reassurance
 > that this is at least intended for this version of the interface,
 > ESPECIALLY IF THAT CAN BE MADE AVAILABLE FOR ZERO COST (i.e., check
 > to see if the residual is 0xDEADBEEF instead of 0xB16B00B5).
 >
 > Why burden the packet with a "protocol version" parameter?
---8<---8<---

> it up on Wikipedia gives no clues.

As I said, above:

    "If, OTOH, such a possibility existed (as is the case with RPC/RMI),
    then you would want the compiler to generate the machinery to ensure
    the arguments get to the correct function and for the function to be
    able to ensure that the arguments are actually intended for it."

You would want the IDL (Interface Definition Language) compiler to
generate stubs (client- and server-side) that enforced the constraints
specified in the IDL and OCL.

Again, in a perfect world, you'd not need any of these mechanisms.
Data wouldn't be corrupted on the wire.  Hostiles wouldn't try to
subvert those messages.  Developers would always ensure they
adhered to the contracts laid out for each API.  etc.

"Yet, folks make mistakes."

> So for now, I'll assume you don't want anyone to know what you meant and I can 
> safely ignore anything you write in connection with the terms.

Perhaps other folks were more careful in their reading (of the quoted passage,
above).

>>>> OTOH, "salting" the calculation so that it is expected to yield
>>>> a value of 0x13 means *those* situations will be flagged as errors
>>>> (and a different set of situations will sneak by, undetected).
>>>
>>> And that gives you exactly /zero/ benefit.
>>
>> See above.
> 
> I did.  Zero benefit.

Perhaps your reading was as deficient there as you've admitted it to
be elsewhere?

> Actually, it is worse than useless - it makes it harder to identify the 
> protocol, and reduces the information content of the CRC check.
> 
>>> You run your hash algorithm, and check for the single value that indicates 
>>> no errors.  It does not matter if that number is 0, 0x13, or - often more 
>> -----------^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>>
>> As you've admitted, it doesn't matter.  So, why wouldn't I opt to have
>> an algorithm for THIS interface give me a result that is EXPECTED
>> for this protocol?  What value picking "0"?
> 
> A /single/ result does not matter (other than needlessly complicating things).  
> Having multiple different valid results /does/ matter.

For any CRC calculation instance, you *know* what the result is expected to be.
How many different "check" algorithms do you think are operating in your
PC as you type/read (i.e., all of the protocols between devices running
in the box, all of the ROMs in those devices, the media accessed by them,
etc.)?  Has EVERY developer who needed a CRC settled on the "Holy Grail"
of CRCs... because it's easiest?  Or, have they each chosen schemes that
they consider appropriate to their needs?

I compute hashes of individual memory pages during reschedule()s.
And, verify that they are intact when next accessed (because they
may have been corrupted by a side-channel attack while not
actively being accessed -- by the owning task -- despite the
protections afforded by the MMU).  Should I use the same "check"
algorithm that I do when sending a message to another node?
Or, that I use on the wire?

Should I use the same algorithm when checking 4K pages as I would
when checking 16MB pages?  The goal isn't to *correct* errors so
I'd want one that detects the greatest number of errors LIKELY
INDUCED BY SUCH AN ATTACK (which can differ from the types of
*burst* errors that corrupt packets on the wire or lead to
read/write disturb errors in FLASH...)

As I said, up-thread:  "... you don't just use CRCs (secure hashes, etc.)
on 'code images'"

>>>>> That is why you need to distinguish between the two possibilities. If you 
>>>>> don't have to worry about malicious attacks, a 32-bit CRC takes a dozen 
>>>>> lines of C code and a 1 KB table, all running extremely efficiently.  If 
>>>>> security is an issue, you need digital signatures - an RSA-based signature 
>>>>> system is orders of magnitude more effort in both development time and in 
>>>>> run time.
>>>>
>>>> It's considerably more expensive AND not fool-proof -- esp if the
>>>> attacker knows you are signing binaries.  "OK, now I need to find
>>>> WHERE the signature is verified and just patch that "CALL" out
>>>> of the code".
>>>
>>> I'm not sure if that is a straw-man argument, or just showing your ignorance 
>>> of the topic.  Do you really think security checks are done by the program 
>>> you are trying to send securely?  That would be like trying to have building 
>>> security where people entering the building look at their own security cards.
>>
>> Do YOU really think we all design applications that run in PCs where some
>> CLOSED OS performs these tests in a manner that can't be subverted?
> 
> Do you bother to read my posts at all?  Or do you prefer to make up things that 
> you imagine I write, so that you can make nonsensical attacks on them?  
> Certainly there is no sane reading of my posts (written and sent from an /open/ 
> OS) where "do not rely on security by obscurity" could be taken to mean "rely 
> on obscured and closed platforms".

"Do you really think security checks are done by the program you are trying
to send securely?  That would be like trying to have building security where
people entering the building look at their own security cards."

Who *else* is involved in the acceptance/verification of a code image
in an embedded product?  (Not all "run Linux")

>> *WE* (tend to) write ALL the code in the products developed, here.
>> So, whether it's the POST WE wrote that is performing the test or
>> the loader WE wrote, it's still *our* program.
>>
>> Yes, we ARE looking at our own security cards!
>>
>> Manufacturers *try* to hide ("obscurity") details of these mechanisms
>> in an attempt to improve effective security.  But, there's nothing
>> that makes these guarantees.
> 
> Why are you trying to "persuade" me that manufacturer obscurity is a bad 
> thing?  You have been promoting obscurity of algorithms as though it were 
> helpful for security - I have made clear that it is not.  Are you getting your 
> own position mixed up with mine?

If the manufacturer saw no benefit to obscurity, then why embrace it?

>> Give me the sources for Windows (Linux, *BSD, etc.) and I can
>> subvert all the state-of-the-art digital signing used to ensure
>> binaries aren't altered.  Nothing *outside* the box is involved
>> so, by definition, everything I need has to reside *in* the box.
> 
> No, you can't.  The sources for Linux and *BSD /are/ all freely available.  The 
> private signing keys used by, for example, Red Hat or Debian, are /not/ freely 
> available.  You cannot make changes to a Red Hat or Debian package that will 
> pass the security checks - you are unable to sign the packages.

Sure I can!  If you are just signing a package to verify that it hasn't
been tampered with BUT THE CONTENTS ARE NOT ENCRYPTED, then all you have
to do is remove the signature check -- leaving the signature in the
(unchecked) executable.

This is different than *encrypting* the package (the OP said nothing
about encrypting his executable).

> This is precisely because something /outside/ the box /is/ involved - the 
> private half of the public/private key used for signing.  The public half - and 
> all the details of the algorithms - is easily available to let people verify 
> the signature, but the private half is kept secret.

And, if I eliminate the check that verifies the signature, then what
value signing?  "Yes, I assume the risk of running an allegedly signed
executable (THAT MAY HAVE BEEN TAMPERED WITH)."

> (Sorry, but I've skipped and snipped the rest.  I simply don't have time to go 
> through it in detail.  If others find it useful or interesting, that's great, 
> but there has to be limits somewhere.)

The limits seem to be in your imagination.  You believe there's *a* way
of doing things instead of a multitude of ways, each with different
tradeoffs.  And, think you'll always have <whatever> is needed (resources,
time, staff, expertise, etc.) to get exactly those things.  The "box"
surrounding you limits what you can see.

Sad in an engineer.  But, must be incredibly comforting!

Bye, David.

[toc] | [prev] | [next] | [standalone]


#31902

FromDavid Brown <david.brown@hesbynett.no>
Date2023-05-03 14:48 +0200
Message-ID<u2tl7k$19ji0$1@dont-email.me>
In reply to#31901
On 03/05/2023 09:15, Don Y wrote:
> On 4/24/2023 7:37 AM, David Brown wrote:
>> On 24/04/2023 09:32, Don Y wrote:
>>> On 4/22/2023 7:57 AM, David Brown wrote:
>>>>>> However, in almost every case where CRC's might be useful, you 
>>>>>> have additional checks of the sanity of the data, and an all-zero 
>>>>>> or all-one data block would be rejected.  For example, Ethernet 
>>>>>> packets use CRC for integrity checking, but an attempt to send a 
>>>>>> packet type 0 from MAC address 00:00:00:00:00:00 to address 
>>>>>> 00:00:00:00:00:00, of length 0, would be rejected anyway.
>>>>>
>>>>> Why look at "data" -- which may be suspect -- and *then* check its 
>>>>> CRC?
>>>>> Run the CRC first.  If it fails, decide how you are going to proceed
>>>>> or recover.
>>>>
>>>> That is usually the order, yes.  Sometimes you want "fail fast", 
>>>> such as dropping a packet that was not addressed to you (it doesn't 
>>>> matter if it was received correctly but for someone else, or it was 
>>>> addressed to you but the receiver address was corrupted - you are 
>>>> dropping the packet either way). But usually you will run the CRC 
>>>> then look at the data.
>>>>
>>>> But the order doesn't matter - either way, you are still checking 
>>>> for valid data, and if the data is invalid, it does not matter if 
>>>> the CRC only passed by luck or by all zeros.
>>>
>>> You're assuming the CRC is supposed to *vouch* for the data.
>>> The CRC can be there simply to vouch for the *transport* of a
>>> datagram.
>>
>> I am assuming that the CRC is there to determine the integrity of the 
>> data in the face of possible unintentional errors.  That's what CRC 
>> checks are for. They have nothing to do with the content of the data, 
>> or the type of the data package or image.
> 
> Exactly.  And, a CRC on *a* protocol can use ANY ALGORITHM that the 
> protocol
> defines.  Not some "canned one-size fits all" approach.

It makes sense to use an 8-bit CRC on small telegrams, 16-bit CRC on 
bigger things, 32-bit CRC on flash images, and 64-bit CRC when you want 
to use the CRC as an identifying hash (and malicious tampering is 
non-existent).  There can also be benefits of particular choices of CRC 
for particular use-cases, in terms of detection of certain error 
patterns for certain lengths of data.


What I don't see any point in is using variations, such as different 
initial values.  I've already said why I think pathological cases such 
as all zero data are normally irrelevant - but I can accept that there 
may be occasions when they could happen, and thus a /single/ non-zero 
initial value would be useful.

> 
>> As an example of the use of CRC's in messaging, look at Ethernet frames:
>>
>> <https://en.wikipedia.org/wiki/Ethernet_frame>
>>
>> The CRC  does not care about the content of the data it protects.
> 
> AND, if the packet yielded an incorrect CRC, you can assume the
> data was corrupt... OR, you are looking at a different protocol
> and MISTAKING it for something that you *think* it might be.

If the CRC does not match, you reject the packet or data.  End of story. 
  You don't know or care /why/ - because you cannot be sure of any reason.

> 
> If I produce a stream of data, can you tell me what the checksum
> for THAT stream *should* be?  You have to either be told what
> it is (and have a way of knowing what the checksum SHOULD be)
> *or* have to make some assumptions about it.

If you are transmitting some data then both sides need to agree on the 
CRC algorithm (size, polynomial, initial value, etc.), and on whether a 
check is "CRC of everything gives 0" or "CRC of everything except the 
pre-calculated CRC equals the transmitted pre-calculated CRC".

> 
> If you have assumed wrong *or* if the data has been corrupt, then
> the CRC should fail.  You don't care why it failed -- because you
> can't do anything about it.  You just know that you can't use the data
> in the way you THOUGHT it could be used.
> 

Well, yes.  Obviously.

If you are making incorrect assumptions here, someone is doing a pretty 
poor job at designing, describing or implementing the communications 
system.  It is just like getting the baud rate wrong on a UART link.


>>> So, use a version-specific CRC on the packet.  If it fails, then
>>> either the data in the packet has been corrupted (which could just
>>> as easily have involved an embedded "interface version" parameter);
>>> or the packet was formed with the wrong CRC.
>>>
>>> If the CRC is correct FOR THAT VERSION OF THE PROTOCOL, then
>>> why bother looking at a "protocol version" parameter?  Would
>>> you ALSO want to verify all the rest of the parameters?
>>
>> I'm sorry, I simply cannot see your point.  Identifying the version of 
>> a protocol, or other protocol type information, is a totally 
>> orthogonal task to ensuring the integrity of the data.  The concepts 
>> should be handled separately.
> 
> It is.  A packet using protocol XYZ is delivered to port ABC.
> Port ABC *only* handles protocol XYZ.  Anything else arriving there,
> with a potentially different checksum, is invalid.  Even if, for example,
> byte number 27 happens to have the correct "magic number" for that
> protocol.
> 
> Because the message doesn't obey the rules defined by the protocol
> FOR THAT PORT.  What do I gain by insisting that byte number 27 must
> be 0x5A that the CRC doesn't already tell me?
> 

A CRC failure doesn't tell you that the telegram type is wrong.  It 
tells you that the data is corrupted.

If there can be different protocols, or telegram types, or whatever, 
then identify them.  Stop playing silly buggers with abuse of different 
concepts that have different roles in the communication system.


>>> Salt just ensures that you can differentiate between functionally 
>>> identical
>>> values.  I.e., in a CRC, it differentiates between the "0x0000" that 
>>> CRC-1
>>> generates from the "0x0000" that CRC-2 generates.
>>
>> Can we agree that this is called an "initial value", not "salt" ?
> 
> It depends on how you implement it.  The point is to produce
> different results for the same polynmomial.

It is called an "initial value" - it is not "salt".  It doesn't matter 
if you want to pick different initial values for your CRC, or why you 
want to do that.  You are still not talking about salt.

If you insist on using your own terminology, you will be left talking to 
yourself.

> 
>>> You don't see the parallel to ensuring that *my* use of "Passw0rd" is
>>> encoded in a different manner than *your* use of "Passw0rd"?
>>
>> No.  They are different things.
>>
>> An important difference is that adding "salt" to a password hash is an 
>> important security feature.  Picking a different initial value for a 
>> CRC instead of having appropriate protocol versioning in the data (or 
>> a surrounding envelope) is a misfeature.
> 
> And you don't see that verifying that a packet of data received at
> port ABC that should only see the checksum associated with protocol
> XYZ as being similarly related?

No.  They are different things.

Look, I /do/ understand what you are doing, and I appreciate that you 
think it is a good idea.  To me, it is an unpleasant mix of orthogonal 
concepts that needlessly complicates things.  Just because something is 
/possible/, does not mean it is a good idea.


>>
>>>>> See the RMI desciption.
>>>>
>>>> I'm sorry, I have no idea what "RMI" is or where it is described. 
>>>> You've mentioned that abbreviation twice, but I can't figure it out.
>>>
>>> <https://en.wikipedia.org/wiki/RMI>
>>> <https://en.wikipedia.org/wiki/OCL>
>>>
>>> Nothing magical with either term.
>>
>> I looked up RMI on Wikipedia before asking, and saw nothing of 
>> relevance to CRC's or checksums.
> 

I've snipped the ramblings that have nothing to do with the question I 
asked.  I assume you don't want to answer me.

> 
>> I noticed no mention of "OCL" in your posts, and looking 
> 
> You need to read more carefully.

I've looked.  You did not mention "OCL" anywhere before giving the URL 
to the wikipedia page.  You only mentioned it /afterwards/ - without any 
context that suggests what you meant.  (Here's a hint for you - if you 
want to refer to a wikipedia page, put a link to the /relevant/ page.)


Presumably "RMI" and "OCL" have particular meanings that are relevant 
for projects you work on, and are so familiar to you that they are part 
of your language.  No one else knows or cares what they are, and they 
are irrelevant in this thread.  So let's leave them there.

> 
>>> Give me the sources for Windows (Linux, *BSD, etc.) and I can
>>> subvert all the state-of-the-art digital signing used to ensure
>>> binaries aren't altered.  Nothing *outside* the box is involved
>>> so, by definition, everything I need has to reside *in* the box.
>>
>> No, you can't.  The sources for Linux and *BSD /are/ all freely 
>> available.  The private signing keys used by, for example, Red Hat or 
>> Debian, are /not/ freely available.  You cannot make changes to a Red 
>> Hat or Debian package that will pass the security checks - you are 
>> unable to sign the packages.
> 
> Sure I can!  If you are just signing a package to verify that it hasn't
> been tampered with BUT THE CONTENTS ARE NOT ENCRYPTED, then all you have
> to do is remove the signature check -- leaving the signature in the
> (unchecked) executable.

Woah, you /really/ don't understand this stuff, do you?  Here's a clue - 
ask yourself what is being signed, and what is doing the checking.

Perhaps also ask yourself if /all/ the people involved in security for 
Linux or BSD - all the companies such as Red Hat, IBM, Intel, etc., - 
ask if /all/ of them have got it wrong, and only /you/ realise that 
digital signatures on open source software is useless?  /Very/ 
occasionally, there is a lone genius that understands something while 
all the other experts are wrong - but in most cases, the loner is the 
one that is wrong.

[toc] | [prev] | [next] | [standalone]


#31936

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-05-09 20:42 +0200
Message-ID<u3e46p$bcui$2@dont-email.me>
In reply to#31902
Den 2023-05-03 kl. 14:48, skrev David Brown:
> On 03/05/2023 09:15, Don Y wrote:
>> On 4/24/2023 7:37 AM, David Brown wrote:
>>> On 24/04/2023 09:32, Don Y wrote:
>>>> On 4/22/2023 7:57 AM, David Brown wrote:
>>>>>>> However, in almost every case where CRC's might be useful, you 
>>>>>>> have additional checks of the sanity of the data, and an all-zero 
>>>>>>> or all-one data block would be rejected.  For example, Ethernet 
>>>>>>> packets use CRC for integrity checking, but an attempt to send a 
>>>>>>> packet type 0 from MAC address 00:00:00:00:00:00 to address 
>>>>>>> 00:00:00:00:00:00, of length 0, would be rejected anyway.
>>>>>>
>>>>>> Why look at "data" -- which may be suspect -- and *then* check its 
>>>>>> CRC?
>>>>>> Run the CRC first.  If it fails, decide how you are going to proceed
>>>>>> or recover.
>>>>>
>>>>> That is usually the order, yes.  Sometimes you want "fail fast", 
>>>>> such as dropping a packet that was not addressed to you (it doesn't 
>>>>> matter if it was received correctly but for someone else, or it was 
>>>>> addressed to you but the receiver address was corrupted - you are 
>>>>> dropping the packet either way). But usually you will run the CRC 
>>>>> then look at the data.
>>>>>
>>>>> But the order doesn't matter - either way, you are still checking 
>>>>> for valid data, and if the data is invalid, it does not matter if 
>>>>> the CRC only passed by luck or by all zeros.
>>>>
>>>> You're assuming the CRC is supposed to *vouch* for the data.
>>>> The CRC can be there simply to vouch for the *transport* of a
>>>> datagram.
>>>
>>> I am assuming that the CRC is there to determine the integrity of the 
>>> data in the face of possible unintentional errors.  That's what CRC 
>>> checks are for. They have nothing to do with the content of the data, 
>>> or the type of the data package or image.
>>
>> Exactly.  And, a CRC on *a* protocol can use ANY ALGORITHM that the 
>> protocol
>> defines.  Not some "canned one-size fits all" approach.
> 
> It makes sense to use an 8-bit CRC on small telegrams, 16-bit CRC on 
> bigger things, 32-bit CRC on flash images, and 64-bit CRC when you want 
> to use the CRC as an identifying hash (and malicious tampering is 
> non-existent).  There can also be benefits of particular choices of CRC 
> for particular use-cases, in terms of detection of certain error 
> patterns for certain lengths of data.

Flash images larger than X kB may need a 64-bit CRC.
I don't remember exactly when to start considering it,
but something between 64kB-256kB is probably correct.

It is all to do with Hamming Distance, and this is also affected by the 
polynome.
/Ulf


> 
> 
> What I don't see any point in is using variations, such as different 
> initial values.  I've already said why I think pathological cases such 
> as all zero data are normally irrelevant - but I can accept that there 
> may be occasions when they could happen, and thus a /single/ non-zero 
> initial value would be useful.
> 
>>
>>> As an example of the use of CRC's in messaging, look at Ethernet frames:
>>>
>>> <https://en.wikipedia.org/wiki/Ethernet_frame>
>>>
>>> The CRC  does not care about the content of the data it protects.
>>
>> AND, if the packet yielded an incorrect CRC, you can assume the
>> data was corrupt... OR, you are looking at a different protocol
>> and MISTAKING it for something that you *think* it might be.
> 
> If the CRC does not match, you reject the packet or data.  End of story. 
>   You don't know or care /why/ - because you cannot be sure of any reason.
> 
>>
>> If I produce a stream of data, can you tell me what the checksum
>> for THAT stream *should* be?  You have to either be told what
>> it is (and have a way of knowing what the checksum SHOULD be)
>> *or* have to make some assumptions about it.
> 
> If you are transmitting some data then both sides need to agree on the 
> CRC algorithm (size, polynomial, initial value, etc.), and on whether a 
> check is "CRC of everything gives 0" or "CRC of everything except the 
> pre-calculated CRC equals the transmitted pre-calculated CRC".
> 
>>
>> If you have assumed wrong *or* if the data has been corrupt, then
>> the CRC should fail.  You don't care why it failed -- because you
>> can't do anything about it.  You just know that you can't use the data
>> in the way you THOUGHT it could be used.
>>
> 
> Well, yes.  Obviously.
> 
> If you are making incorrect assumptions here, someone is doing a pretty 
> poor job at designing, describing or implementing the communications 
> system.  It is just like getting the baud rate wrong on a UART link.
> 
> 
>>>> So, use a version-specific CRC on the packet.  If it fails, then
>>>> either the data in the packet has been corrupted (which could just
>>>> as easily have involved an embedded "interface version" parameter);
>>>> or the packet was formed with the wrong CRC.
>>>>
>>>> If the CRC is correct FOR THAT VERSION OF THE PROTOCOL, then
>>>> why bother looking at a "protocol version" parameter?  Would
>>>> you ALSO want to verify all the rest of the parameters?
>>>
>>> I'm sorry, I simply cannot see your point.  Identifying the version 
>>> of a protocol, or other protocol type information, is a totally 
>>> orthogonal task to ensuring the integrity of the data.  The concepts 
>>> should be handled separately.
>>
>> It is.  A packet using protocol XYZ is delivered to port ABC.
>> Port ABC *only* handles protocol XYZ.  Anything else arriving there,
>> with a potentially different checksum, is invalid.  Even if, for example,
>> byte number 27 happens to have the correct "magic number" for that
>> protocol.
>>
>> Because the message doesn't obey the rules defined by the protocol
>> FOR THAT PORT.  What do I gain by insisting that byte number 27 must
>> be 0x5A that the CRC doesn't already tell me?
>>
> 
> A CRC failure doesn't tell you that the telegram type is wrong.  It 
> tells you that the data is corrupted.
> 
> If there can be different protocols, or telegram types, or whatever, 
> then identify them.  Stop playing silly buggers with abuse of different 
> concepts that have different roles in the communication system.
> 
> 
>>>> Salt just ensures that you can differentiate between functionally 
>>>> identical
>>>> values.  I.e., in a CRC, it differentiates between the "0x0000" that 
>>>> CRC-1
>>>> generates from the "0x0000" that CRC-2 generates.
>>>
>>> Can we agree that this is called an "initial value", not "salt" ?
>>
>> It depends on how you implement it.  The point is to produce
>> different results for the same polynmomial.
> 
> It is called an "initial value" - it is not "salt".  It doesn't matter 
> if you want to pick different initial values for your CRC, or why you 
> want to do that.  You are still not talking about salt.
> 
> If you insist on using your own terminology, you will be left talking to 
> yourself.
> 
>>
>>>> You don't see the parallel to ensuring that *my* use of "Passw0rd" is
>>>> encoded in a different manner than *your* use of "Passw0rd"?
>>>
>>> No.  They are different things.
>>>
>>> An important difference is that adding "salt" to a password hash is 
>>> an important security feature.  Picking a different initial value for 
>>> a CRC instead of having appropriate protocol versioning in the data 
>>> (or a surrounding envelope) is a misfeature.
>>
>> And you don't see that verifying that a packet of data received at
>> port ABC that should only see the checksum associated with protocol
>> XYZ as being similarly related?
> 
> No.  They are different things.
> 
> Look, I /do/ understand what you are doing, and I appreciate that you 
> think it is a good idea.  To me, it is an unpleasant mix of orthogonal 
> concepts that needlessly complicates things.  Just because something is 
> /possible/, does not mean it is a good idea.
> 
> 
>>>
>>>>>> See the RMI desciption.
>>>>>
>>>>> I'm sorry, I have no idea what "RMI" is or where it is described. 
>>>>> You've mentioned that abbreviation twice, but I can't figure it out.
>>>>
>>>> <https://en.wikipedia.org/wiki/RMI>
>>>> <https://en.wikipedia.org/wiki/OCL>
>>>>
>>>> Nothing magical with either term.
>>>
>>> I looked up RMI on Wikipedia before asking, and saw nothing of 
>>> relevance to CRC's or checksums.
>>
> 
> I've snipped the ramblings that have nothing to do with the question I 
> asked.  I assume you don't want to answer me.
> 
>>
>>> I noticed no mention of "OCL" in your posts, and looking 
>>
>> You need to read more carefully.
> 
> I've looked.  You did not mention "OCL" anywhere before giving the URL 
> to the wikipedia page.  You only mentioned it /afterwards/ - without any 
> context that suggests what you meant.  (Here's a hint for you - if you 
> want to refer to a wikipedia page, put a link to the /relevant/ page.)
> 
> 
> Presumably "RMI" and "OCL" have particular meanings that are relevant 
> for projects you work on, and are so familiar to you that they are part 
> of your language.  No one else knows or cares what they are, and they 
> are irrelevant in this thread.  So let's leave them there.
> 
>>
>>>> Give me the sources for Windows (Linux, *BSD, etc.) and I can
>>>> subvert all the state-of-the-art digital signing used to ensure
>>>> binaries aren't altered.  Nothing *outside* the box is involved
>>>> so, by definition, everything I need has to reside *in* the box.
>>>
>>> No, you can't.  The sources for Linux and *BSD /are/ all freely 
>>> available.  The private signing keys used by, for example, Red Hat or 
>>> Debian, are /not/ freely available.  You cannot make changes to a Red 
>>> Hat or Debian package that will pass the security checks - you are 
>>> unable to sign the packages.
>>
>> Sure I can!  If you are just signing a package to verify that it hasn't
>> been tampered with BUT THE CONTENTS ARE NOT ENCRYPTED, then all you have
>> to do is remove the signature check -- leaving the signature in the
>> (unchecked) executable.
> 
> Woah, you /really/ don't understand this stuff, do you?  Here's a clue - 
> ask yourself what is being signed, and what is doing the checking.
> 
> Perhaps also ask yourself if /all/ the people involved in security for 
> Linux or BSD - all the companies such as Red Hat, IBM, Intel, etc., - 
> ask if /all/ of them have got it wrong, and only /you/ realise that 
> digital signatures on open source software is useless?  /Very/ 
> occasionally, there is a lone genius that understands something while 
> all the other experts are wrong - but in most cases, the loner is the 
> one that is wrong.
> 

[toc] | [prev] | [next] | [standalone]


#31940

FromDavid Brown <david.brown@hesbynett.no>
Date2023-05-10 10:06 +0200
Message-ID<u3fjaa$kadg$1@dont-email.me>
In reply to#31936
On 09/05/2023 20:42, Ulf Samuelsson wrote:
> Den 2023-05-03 kl. 14:48, skrev David Brown:

>> It makes sense to use an 8-bit CRC on small telegrams, 16-bit CRC on 
>> bigger things, 32-bit CRC on flash images, and 64-bit CRC when you 
>> want to use the CRC as an identifying hash (and malicious tampering is 
>> non-existent).  There can also be benefits of particular choices of 
>> CRC for particular use-cases, in terms of detection of certain error 
>> patterns for certain lengths of data.
> 
> Flash images larger than X kB may need a 64-bit CRC.
> I don't remember exactly when to start considering it,
> but something between 64kB-256kB is probably correct.
> 
> It is all to do with Hamming Distance, and this is also affected by the 
> polynome.
> /Ulf
> 
> 
"Need" is too strong a word here.  A CRC will guarantee detection of 
certain kinds of error (such as a single bit error), regardless of the 
length of the data.  Some kinds of error are limited by length.  If you 
plot a graph with guaranteed Hamming distance on the vertical scale and 
length of data on the horizontal scale, each CRC will drop off in steps. 
  For the same CRC size, some will hold a high Hamming distance for 
longer and then drop off sharply, others will hold a lower Hamming 
distance for very large data.  And in general, a bigger CRC will be 
better here.

But Hamming distance is not everything.  It is important in situations 
where there is an approximately independent risk of corruption for each 
bit individually - such as during radio transmission.  Programming 
images into flash has a completely different error risk pattern.  A 
little Hamming is nice to guarantee that any single cell failure in the 
flash will be be found, but the more realistic flash problems involve 
large scale effects - failure to erase a block fully, or software flaws. 
  For this kind of thing, pretty much any valid CRC polynomial works the 
same - a 32-bit polynomial gives you a 1 in 2 ^ 32 chance of the error 
going undetected.  Yes, a 1 in 2 ^ 64 chance is better, but it's rarely 
something to get excited about.

Note that if you are sending the image to a board via a potentially 
flawed mechanism, you'll want appropriate checks during the transfers. 
Ethernet, Wifi, Bluetooth, USB - they will all have suitable checksums 
for each packet.  And for some of those, Hamming distance and particular 
choice of polynomial /is/ an important consideration.

[toc] | [prev] | [next] | [standalone]


#31942

FromUlf Samuelsson <ulf.r.samuelsson@gmail.com>
Date2023-05-10 12:03 +0200
Message-ID<u3fq5b$l8kc$1@dont-email.me>
In reply to#31940
Den 2023-05-10 kl. 10:06, skrev David Brown:
> On 09/05/2023 20:42, Ulf Samuelsson wrote:
>> Den 2023-05-03 kl. 14:48, skrev David Brown:
> 
>>> It makes sense to use an 8-bit CRC on small telegrams, 16-bit CRC on 
>>> bigger things, 32-bit CRC on flash images, and 64-bit CRC when you 
>>> want to use the CRC as an identifying hash (and malicious tampering 
>>> is non-existent).  There can also be benefits of particular choices 
>>> of CRC for particular use-cases, in terms of detection of certain 
>>> error patterns for certain lengths of data.
>>
>> Flash images larger than X kB may need a 64-bit CRC.
>> I don't remember exactly when to start considering it,
>> but something between 64kB-256kB is probably correct.
>>
>> It is all to do with Hamming Distance, and this is also affected by 
>> the polynome.
>> /Ulf
>>
>>
> "Need" is too strong a word here.  A CRC will guarantee detection of 
> certain kinds of error (such as a single bit error), regardless of the 
> length of the data.  Some kinds of error are limited by length.  If you 
> plot a graph with guaranteed Hamming distance on the vertical scale and 
> length of data on the horizontal scale, each CRC will drop off in steps. 
>   For the same CRC size, some will hold a high Hamming distance for 
> longer and then drop off sharply, others will hold a lower Hamming 
> distance for very large data.  And in general, a bigger CRC will be 
> better here.
> 
> But Hamming distance is not everything.  It is important in situations 
> where there is an approximately independent risk of corruption for each 
> bit individually - such as during radio transmission.  Programming 
> images into flash has a completely different error risk pattern.  A 
> little Hamming is nice to guarantee that any single cell failure in the 
> flash will be be found, but the more realistic flash problems involve 
> large scale effects - failure to erase a block fully, or software flaws. 
>   For this kind of thing, pretty much any valid CRC polynomial works the 
> same - a 32-bit polynomial gives you a 1 in 2 ^ 32 chance of the error 
> going undetected.  Yes, a 1 in 2 ^ 64 chance is better, but it's rarely 
> something to get excited about.

Programming a flash memory can flip bits in parts of the flash memory 
which is not programmed.
Bit errors can also be introduced by radiation.
Some applications require better security than others.
Functional Safety may require CRC size based on code size.
/Ulf


> 
> Note that if you are sending the image to a board via a potentially 
> flawed mechanism, you'll want appropriate checks during the transfers. 
> Ethernet, Wifi, Bluetooth, USB - they will all have suitable checksums 
> for each packet.  And for some of those, Hamming distance and particular 
> choice of polynomial /is/ an important consideration.
> 
> 

[toc] | [prev] | [next] | [standalone]


#31980

FromDon Y <blockedofcourse@foo.invalid>
Date2023-08-05 01:48 -0700
Message-ID<ual2ds$1md5g$2@dont-email.me>
In reply to#31942
On 5/10/2023 3:03 AM, Ulf Samuelsson wrote:
> Programming a flash memory can flip bits in parts of the flash memory which is 
> not programmed.
> Bit errors can also be introduced by radiation.

Executing code can also introduce write *and* read disturb events.

Ask yourself how to protect a design that allows arbitrary code to
be executed (even if in a sandbox) in the presence of potential
side-channel exploits.

Or, as a "simpler" problem:  how to detect if such an exploit
has been invoked (even possibly unintentionally)!

[Imagine devices that "run forever"...]

> Some applications require better security than others.

Exactly.

> Functional Safety may require CRC size based on code size.

[toc] | [prev] | [next] | [standalone]


#31979

FromDon Y <blockedofcourse@foo.invalid>
Date2023-08-05 01:42 -0700
Message-ID<ual22t$1md7b$1@dont-email.me>
In reply to#31902
On 5/3/2023 5:48 AM, David Brown wrote:

>>>> Give me the sources for Windows (Linux, *BSD, etc.) and I can
>>>> subvert all the state-of-the-art digital signing used to ensure
>>>> binaries aren't altered.  Nothing *outside* the box is involved
>>>> so, by definition, everything I need has to reside *in* the box.
>>>
>>> No, you can't.  The sources for Linux and *BSD /are/ all freely available.  
>>> The private signing keys used by, for example, Red Hat or Debian, are /not/ 
>>> freely available.  You cannot make changes to a Red Hat or Debian package 
>>> that will pass the security checks - you are unable to sign the packages.
>>
>> Sure I can!  If you are just signing a package to verify that it hasn't
>> been tampered with BUT THE CONTENTS ARE NOT ENCRYPTED, then all you have
>> to do is remove the signature check -- leaving the signature in the
>> (unchecked) executable.
> 
> Woah, you /really/ don't understand this stuff, do you?  Here's a clue - ask 
> yourself what is being signed, and what is doing the checking.

Exactly.  You don't attack the signature or the keys.  You BUILD A NEW KERNEL 
THAT DOESN'T CHECK THE SIGNATURE.  You attack (replace if you have access
to the sources -- as I stipulated above) the "what is doing the checking".
This is c.a.e; you likely have physical access and control of the device
(unlike trying to attack a remote system)

The binary is exposed UNENCRYPTED in the signed executable (please note my
stipulation to that, too, above).  The only thing preventing its execution
(if tampered -- or unlicensed!) is the signature check.  Bypass that in any
way and the code executes AS IF signed.

I design "devices".  Don't you think if there was a foolproof way (by resorting
to "school boy techniques") to protect them from counterfeiting and tampering
that I would have already embraced that?  That EVERY computer-based product
would be INHERENTLY SECURED??

[There are ways that are far from theoretical yet considerably more
effective.  A signature check is easy to detect in an executing device
and, thus, elided.  Lather, rinse, repeat for each precursor level
of such protection.]

Please read what I've written more carefully, lest you look foolish.  Or,
spend a few years playing red-blue games and actually trying to subvert
hardware and software protection mechanisms in REAL products.  (Hint:
you will need to think down BELOW the hardware level to do so successfully
so you can bypass the hardware mechanisms that folks keep trying to embed
in their products)

> Perhaps also ask yourself if /all/ the people involved in security for Linux or 
> BSD - all the companies such as Red Hat, IBM, Intel, etc., - ask if /all/ of 
> them have got it wrong, and only /you/ realise that digital signatures on open 
> source software is useless?

The signature is only of use if the mechanism verifying it is tamperproof.
That's not possible on most (all?) devices sold.  SOMEONE has physical
access to the device so all of the mechanisms you put in place can be
subverted.

*Ask* the Linux and BSD crowds if they can GUARANTEE that ALTERED signed code
can't be executed on a system where the adversary can build and install their
own kernel.  Or, probe the innerworkings of such a device AT THEIR LEISURE/

> /Very/ occasionally, there is a lone genius that 
> understands something while all the other experts are wrong - but in most 
> cases, the loner is the one that is wrong.

In this case, you have clearly failed to understand what was being said.
So, don't count yourself in with the "experts".

If the kernel loading the executable doesn't contain code to validate the
signature  (and, if I have the sources for said kernel/OS then I can
easily *make* such a kernel) then the signature is just another unused
"section" in the BLOB.  Just like debug symbols or copyright information.

[toc] | [prev] | [next] | [standalone]


#31810

FromStefan Reuther <stefan.news@arcor.de>
Date2023-04-21 19:40 +0200
Message-ID<u1uoqs.17k.1@stefan.msgid.phost.de>
In reply to#31799
Am 20.04.2023 um 22:44 schrieb George Neuner:
> On Thu, 20 Apr 2023 09:45:59 -0700 (PDT), Rick C
>> CRC is not complicated, but I would not know how to calculate an
>> inserted value to force the resulting CRC to zero.  How do you do
>> that? 
> 
> It's implicit in the equation they chose.  I don't know how it works -
> just that it does.

It works for any CRC that starts with zero and does not invert.

CRC is based on polynomial division remainders. Basically, the CRC is a
division remainder of the input interpreted as a polynomial, and if you
add that remainder back into the equation, result is zero.

https://crccalc.com/?crc=12345678&method=CRC-16/AUG-CCITT&datatype=hex&outtype=0
-> result is 0xBA3C

https://crccalc.com/?crc=12345678BA3C&method=CRC-16/AUG-CCITT&datatype=hex&outtype=0
-> result is 0x0000

(Need to be careful with byte orders; for some CRCs on that page, you
need to swap the bytes before appending.)


  Stefan

[toc] | [prev] | [next] | [standalone]


#31797

FromTauno Voipio <tauno.voipio@notused.fi.invalid>
Date2023-04-20 20:17 +0300
Message-ID<u1rs2m$m90s$1@dont-email.me>
In reply to#31795
On 20.4.2023 18.33, George Neuner wrote:
> On Wed, 19 Apr 2023 19:06:33 -0700 (PDT), Rick C
> <gnuarm.deletethisbit@gmail.com> wrote:
> 
>> This is a bit of the chicken and egg thing.  If you want a embed a
>> checksum in a code module to report the checksum, is there a way of
>> doing this?  It's a bit like being your own grandfather, I think.
> 
> Take a look at the old xmodem/ymodem CRC.  It was designed such that
> when the CRC was sent immediately following the data, a receiver
> computing CRC over the whole incoming packet (data and CRC both) would
> get a result of zero.
> 
> But AFAIK it doesn't work with CCITT equation(s) - you have to use
> xmodem/ymodem.
> 
> 
>> I'm not thinking anything too fancy, like a CRC, but rather a simple
>> modulo N addition, maybe N being 2^16.
> 
> Sorry, I don't know a way to do it with a modular checksum.
> YMMV, but I think 16-bit CRC is pretty simple.
> 
> George


The method to check for a proper constant value after the whole
block and CRC are received and put through the generator works
with the CRC-CCITT (actually ITU-T). The proper final value
depends on the initial CRC and whether the CRC is inverted before
sending. The limitation is that the CRC has to be sent least
significant octet first.

For a reference, see RFC1662, Appendix C.

-- 

-TV

[toc] | [prev] | [next] | [standalone]


#31800

FromGeorge Neuner <gneuner2@comcast.net>
Date2023-04-20 16:49 -0400
Message-ID<l2934i9u72lgk252grte75h9q5echomb0h@4ax.com>
In reply to#31797
On Thu, 20 Apr 2023 20:17:07 +0300, Tauno Voipio
<tauno.voipio@notused.fi.invalid> wrote:

>The method to check for a proper constant value after the whole
>block and CRC are received and put through the generator works
>with the CRC-CCITT (actually ITU-T). The proper final value
>depends on the initial CRC and whether the CRC is inverted before
>sending. The limitation is that the CRC has to be sent least
>significant octet first.
>
>For a reference, see RFC1662, Appendix C.

I remember seeing an explanation of it decades ago, but I never would
have been able to find it again.

Thanks,
George

[toc] | [prev] | [next] | [standalone]


#31801

FromRichard Damon <Richard@Damon-Family.org>
Date2023-04-20 22:09 -0400
Message-ID<tfm0M.2629804$GNG9.1468222@fx18.iad>
In reply to#31788
On 4/19/23 10:06 PM, Rick C wrote:
> This is a bit of the chicken and egg thing.  If you want a embed a checksum in a code module to report the checksum, is there a way of doing this?  It's a bit like being your own grandfather, I think.
> 
> I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16.
> 
> I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way.  Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum.
> 
> I'm not thinking of any special checksum generator that excludes the checksum data.  That would be too messy.
> 
> I keep thinking there is a different way of looking at this to achieve the result I want...
> 
> Maybe I can prove it is impossible.  Assume the file checksums to X when the checksum data is zero.  The goal would then be to include the checksum data value Y in the file, that would change X to Y.  Given the properties of the module N checksum, this would appear to be impossible for the general case, unless...  Add another data value, called, checksum normalizer.  This data value checksums with the original checksum to give the result zero.  Then, when the checksum is also added, the resulting checksum is, in fact, the checksum.  Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact.
> 
> This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think.  At least, this could work in software, where data can be included in an image file as itself.  In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted.  Memory might have data that is stored as itself.  I'll need to look into that.
> 


IF I understand you correctly, what you want is for the file to compute 
to some "checksum" that comes from the basic contents of the file, and 
then you want to add the "checksum" into the file so the program itself 
can print its checksum.

One fact to remember, is that "cryptographic hashes" were invented 
because it was too easy to create a faked file that matches a 
non-crptographic hash/checksum, so that couldn't be a key to make sure 
you really had the right file in the presence of a determined enemy, but 
the checksums were good enough to catch "random" errors.

This means that you can add the checksum into the file, and some 
additional bytes (likely at the end) and by knowing the propeties of the 
checksum algorithm, compute a value for those extra bytes such that the 
"undo" the changes caused by adding the checksum bytes to file.

I'm not sure exactly how to computes these, but the key is that you add 
something at the end of the file to get the checksum back to what the 
original file had before you added the checksum into the file.

[toc] | [prev] | [next] | [standalone]


#31802

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-20 19:41 -0700
Message-ID<cd494e7f-9952-40ac-98da-4745dff32493n@googlegroups.com>
In reply to#31801
On Thursday, April 20, 2023 at 10:09:35 PM UTC-4, Richard Damon wrote:
> On 4/19/23 10:06 PM, Rick C wrote: 
> > This is a bit of the chicken and egg thing. If you want a embed a checksum in a code module to report the checksum, is there a way of doing this? It's a bit like being your own grandfather, I think. 
> > 
> > I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16. 
> > 
> > I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way. Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum. 
> > 
> > I'm not thinking of any special checksum generator that excludes the checksum data. That would be too messy. 
> > 
> > I keep thinking there is a different way of looking at this to achieve the result I want... 
> > 
> > Maybe I can prove it is impossible. Assume the file checksums to X when the checksum data is zero. The goal would then be to include the checksum data value Y in the file, that would change X to Y. Given the properties of the module N checksum, this would appear to be impossible for the general case, unless... Add another data value, called, checksum normalizer. This data value checksums with the original checksum to give the result zero. Then, when the checksum is also added, the resulting checksum is, in fact, the checksum. Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact. 
> > 
> > This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think. At least, this could work in software, where data can be included in an image file as itself. In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted. Memory might have data that is stored as itself. I'll need to look into that. 
> >
> IF I understand you correctly, what you want is for the file to compute 
> to some "checksum" that comes from the basic contents of the file, and 
> then you want to add the "checksum" into the file so the program itself 
> can print its checksum. 
> 
> One fact to remember, is that "cryptographic hashes" were invented 
> because it was too easy to create a faked file that matches a 
> non-crptographic hash/checksum, so that couldn't be a key to make sure 
> you really had the right file in the presence of a determined enemy, but 
> the checksums were good enough to catch "random" errors. 
> 
> This means that you can add the checksum into the file, and some 
> additional bytes (likely at the end) and by knowing the propeties of the 
> checksum algorithm, compute a value for those extra bytes such that the 
> "undo" the changes caused by adding the checksum bytes to file. 
> 
> I'm not sure exactly how to computes these, but the key is that you add 
> something at the end of the file to get the checksum back to what the 
> original file had before you added the checksum into the file.

Yeah, for a simple checksum, I think that would be easy, at least if "checksum" means a bitwise XOR operation.  If the checksum and extra bytes are both 16 bits, this would also work for an arithmetic checksum where each 16 bit word were added into the checksum.  All the carries would cascade out of the upper 16 bits from adding the inserted checksum and it's 2's complement.  

I don't even want to think about using a CRC to try to do this. 

-- 

  Rick C.

  +- Get 1,000 miles of free Supercharging
  +- Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31811

FromRichard Damon <Richard@Damon-Family.org>
Date2023-04-21 19:30 -0400
Message-ID<20F0M.291598$b7Kc.5484@fx39.iad>
In reply to#31802
On 4/20/23 10:41 PM, Rick C wrote:
> On Thursday, April 20, 2023 at 10:09:35 PM UTC-4, Richard Damon wrote:
>> On 4/19/23 10:06 PM, Rick C wrote:
>>> This is a bit of the chicken and egg thing. If you want a embed a checksum in a code module to report the checksum, is there a way of doing this? It's a bit like being your own grandfather, I think.
>>>
>>> I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16.
>>>
>>> I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way. Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum.
>>>
>>> I'm not thinking of any special checksum generator that excludes the checksum data. That would be too messy.
>>>
>>> I keep thinking there is a different way of looking at this to achieve the result I want...
>>>
>>> Maybe I can prove it is impossible. Assume the file checksums to X when the checksum data is zero. The goal would then be to include the checksum data value Y in the file, that would change X to Y. Given the properties of the module N checksum, this would appear to be impossible for the general case, unless... Add another data value, called, checksum normalizer. This data value checksums with the original checksum to give the result zero. Then, when the checksum is also added, the resulting checksum is, in fact, the checksum. Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact.
>>>
>>> This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think. At least, this could work in software, where data can be included in an image file as itself. In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted. Memory might have data that is stored as itself. I'll need to look into that.
>>>
>> IF I understand you correctly, what you want is for the file to compute
>> to some "checksum" that comes from the basic contents of the file, and
>> then you want to add the "checksum" into the file so the program itself
>> can print its checksum.
>>
>> One fact to remember, is that "cryptographic hashes" were invented
>> because it was too easy to create a faked file that matches a
>> non-crptographic hash/checksum, so that couldn't be a key to make sure
>> you really had the right file in the presence of a determined enemy, but
>> the checksums were good enough to catch "random" errors.
>>
>> This means that you can add the checksum into the file, and some
>> additional bytes (likely at the end) and by knowing the propeties of the
>> checksum algorithm, compute a value for those extra bytes such that the
>> "undo" the changes caused by adding the checksum bytes to file.
>>
>> I'm not sure exactly how to computes these, but the key is that you add
>> something at the end of the file to get the checksum back to what the
>> original file had before you added the checksum into the file.
> 
> Yeah, for a simple checksum, I think that would be easy, at least if "checksum" means a bitwise XOR operation.  If the checksum and extra bytes are both 16 bits, this would also work for an arithmetic checksum where each 16 bit word were added into the checksum.  All the carries would cascade out of the upper 16 bits from adding the inserted checksum and it's 2's complement.
> 
> I don't even want to think about using a CRC to try to do this.
> 

Its is a bit of work, but even a 32-bit CRC will be solvable to find the 
reverse equation. You can do the work once generically, and get a 
formula that computes the value you need to put into the final bytes to 
get the CRC of the file back to the CRC it was before adding the CRC and 
the extra bytes. It wouldn't surprise me if the formula isn't published 
somewhere for the common CRCs.

[toc] | [prev] | [next] | [standalone]


#31804

FromBrian Cockburn <brian.cockburn.1959@gmail.com>
Date2023-04-21 01:53 -0700
Message-ID<66e9fb2a-b9e6-4597-9afa-6572caaa8ca1n@googlegroups.com>
In reply to#31788
On Thursday, April 20, 2023 at 12:06:36 PM UTC+10, Rick C wrote:
> This is a bit of the chicken and egg thing. If you want a embed a checksum in a code module to report the checksum, is there a way of doing this? It's a bit like being your own grandfather, I think. 
> 
> I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16. 
> 
> I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way. Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum. 
> 
> I'm not thinking of any special checksum generator that excludes the checksum data. That would be too messy. 
> 
> I keep thinking there is a different way of looking at this to achieve the result I want... 
> 
> Maybe I can prove it is impossible. Assume the file checksums to X when the checksum data is zero. The goal would then be to include the checksum data value Y in the file, that would change X to Y. Given the properties of the module N checksum, this would appear to be impossible for the general case, unless... Add another data value, called, checksum normalizer. This data value checksums with the original checksum to give the result zero. Then, when the checksum is also added, the resulting checksum is, in fact, the checksum. Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact. 
> 
> This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think. At least, this could work in software, where data can be included in an image file as itself. In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted. Memory might have data that is stored as itself. I'll need to look into that. 
> 
> -- 
> 
> Rick C. 
> 
> - Get 1,000 miles of free Supercharging 
> - Tesla referral code - https://ts.la/richard11209
Rick, What is the purpose of this?  Is it (1) to be able to externally identify a binary, as one might a ROM image by computing a checksum?  Is it (2) for a run-able binary to be able to check itself?  This would of course only be able to detect corruption, not tampering.  Is it (3) for the loader (whatever that might be) to be able to say 'this binary has the correct checksum' and only jump to it if it does?  Again this would only be able to detect corruption, not tampering.  Are you hoping for more than corruption detection?

[toc] | [prev] | [next] | [standalone]


#31807

FromRick C <gnuarm.deletethisbit@gmail.com>
Date2023-04-21 05:12 -0700
Message-ID<a4495c87-c68d-4c87-94aa-701c21bdd19cn@googlegroups.com>
In reply to#31804
On Friday, April 21, 2023 at 4:53:18 AM UTC-4, Brian Cockburn wrote:
> On Thursday, April 20, 2023 at 12:06:36 PM UTC+10, Rick C wrote: 
> > This is a bit of the chicken and egg thing. If you want a embed a checksum in a code module to report the checksum, is there a way of doing this? It's a bit like being your own grandfather, I think. 
> > 
> > I'm not thinking anything too fancy, like a CRC, but rather a simple modulo N addition, maybe N being 2^16. 
> > 
> > I keep thinking of using a placeholder, but that doesn't seem to work out in any useful way. Even if you try to anticipate the impact of adding the checksum, that only gives you a different checksum, that you then need to anticipate further... ad infinitum. 
> > 
> > I'm not thinking of any special checksum generator that excludes the checksum data. That would be too messy. 
> > 
> > I keep thinking there is a different way of looking at this to achieve the result I want... 
> > 
> > Maybe I can prove it is impossible. Assume the file checksums to X when the checksum data is zero. The goal would then be to include the checksum data value Y in the file, that would change X to Y. Given the properties of the module N checksum, this would appear to be impossible for the general case, unless... Add another data value, called, checksum normalizer. This data value checksums with the original checksum to give the result zero. Then, when the checksum is also added, the resulting checksum is, in fact, the checksum. Another way of looking at this is to add a value that combines with the added checksum, to be zero, leaving the original checksum intact. 
> > 
> > This might be inordinately hard for a CRC, but a simple checksum would not be an issue, I think. At least, this could work in software, where data can be included in an image file as itself. In a device like an FPGA, it might not be included in the bit stream file so directly... but that might depend on where in the device it is inserted. Memory might have data that is stored as itself. I'll need to look into that. 
> > 
> > -- 
> > 
> > Rick C. 
> > 
> > - Get 1,000 miles of free Supercharging 
> > - Tesla referral code - https://ts.la/richard11209
> Rick, What is the purpose of this? Is it (1) to be able to externally identify a binary, as one might a ROM image by computing a checksum? Is it (2) for a run-able binary to be able to check itself? This would of course only be able to detect corruption, not tampering. Is it (3) for the loader (whatever that might be) to be able to say 'this binary has the correct checksum' and only jump to it if it does? Again this would only be able to detect corruption, not tampering. Are you hoping for more than corruption detection?

This is simply to be able to say this version is unique, regardless of what the version number says.  Version numbers are set manually and not always done correctly.  I'm looking for something as a backup so that if the checksums are different, I can be sure the versions are not the same.  

The less work involved, the better. 

-- 

  Rick C.

  ++ Get 1,000 miles of free Supercharging
  ++ Tesla referral code - https://ts.la/richard11209

[toc] | [prev] | [next] | [standalone]


#31809

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-21 17:02 +0200
Message-ID<u1u8hu$2ps79$1@dont-email.me>
In reply to#31807
On 21/04/2023 14:12, Rick C wrote:
> 
> This is simply to be able to say this version is unique, regardless
> of what the version number says.  Version numbers are set manually
> and not always done correctly.  I'm looking for something as a backup
> so that if the checksums are different, I can be sure the versions
> are not the same.
> 
> The less work involved, the better.
> 

Run a simple 32-bit crc over the image.  The result is a hash of the 
image.  Any change in the image will show up as a change in the crc.

[toc] | [prev] | [next] | [standalone]


#31813

FromBrian Cockburn <brian.cockburn.1959@gmail.com>
Date2023-04-21 16:56 -0700
Message-ID<04d4cbda-216d-4d0d-8db3-f9decc6e4142n@googlegroups.com>
In reply to#31809
On Saturday, April 22, 2023 at 1:02:28 AM UTC+10, David Brown wrote:
> On 21/04/2023 14:12, Rick C wrote: 
> > 
> > This is simply to be able to say this version is unique, regardless 
> > of what the version number says. Version numbers are set manually 
> > and not always done correctly. I'm looking for something as a backup 
> > so that if the checksums are different, I can be sure the versions 
> > are not the same. 
> > 
> > The less work involved, the better. 
> >
> Run a simple 32-bit crc over the image. The result is a hash of the 
> image. Any change in the image will show up as a change in the crc.
David, a hash and a CRC are not the same thing.  They both produce a reasonably unique result though.  Any change would show in either (unless as a result of intentional tampering).

[toc] | [prev] | [next] | [standalone]


#31821

FromDavid Brown <david.brown@hesbynett.no>
Date2023-04-22 17:01 +0200
Message-ID<u20srm$3ag7l$2@dont-email.me>
In reply to#31813
On 22/04/2023 01:56, Brian Cockburn wrote:
> On Saturday, April 22, 2023 at 1:02:28 AM UTC+10, David Brown wrote:
>> On 21/04/2023 14:12, Rick C wrote:
>>> 
>>> This is simply to be able to say this version is unique,
>>> regardless of what the version number says. Version numbers are
>>> set manually and not always done correctly. I'm looking for
>>> something as a backup so that if the checksums are different, I
>>> can be sure the versions are not the same.
>>> 
>>> The less work involved, the better.
>>> 
>> Run a simple 32-bit crc over the image. The result is a hash of
>> the image. Any change in the image will show up as a change in the
>> crc.
> David, a hash and a CRC are not the same thing. 

A CRC is a type of hash - but hash is a more generic term.

> They both produce a
> reasonably unique result though.  Any change would show in either
> (unless as a result of intentional tampering).

Exactly.  Thus a CRC is a hash.

It is not a cryptographically secure hash, and is not suitable for 
protecting against intentional tampering.  But it /is/ a hash.

[toc] | [prev] | [next] | [standalone]


Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →

Back to top | Article view | comp.arch.embedded


csiph-web