Path: csiph.com!news.mixmin.net!eternal-september.org!reader01.eternal-september.org!.POSTED!not-for-mail From: Tim Rentsch Newsgroups: comp.lang.c Subject: Re: Beginner....Decimal/Octal converter help Date: Mon, 26 Dec 2022 06:09:54 -0800 Organization: A noiseless patient Spider Lines: 167 Message-ID: <86tu1iw91p.fsf@linuxsc.com> References: <8ffd982c-2e82-4caa-9256-ec17be2efe56n@googlegroups.com> <86leqsr091.fsf@linuxsc.com> <875yhv30nu.fsf@bsb.me.uk> <87zgf714zc.fsf@bsb.me.uk> <8735cy0wt4.fsf@bsb.me.uk> <87r10hzn19.fsf@bsb.me.uk> <86v8ptpa7u.fsf@linuxsc.com> <867d26ntiv.fsf@linuxsc.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Injection-Info: reader01.eternal-september.org; posting-host="b017513bf0c3c2565b748ac271e0e545"; logging-data="3458102"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX19W2Tue3zLJ8ZlGRf/wSom2QmjKTH9izyQ=" User-Agent: Gnus/5.11 (Gnus v5.11) Emacs/22.4 (gnu/linux) Cancel-Lock: sha1:GWUSyWMmvFPhw9dEx5mWD91Co88= sha1:KerWwDesxdOPpE5uZVnYPhVzkAU= Xref: csiph.com comp.lang.c:168636 Richard Damon writes: > On 9/14/22 6:24 AM, Tim Rentsch wrote: > >> Richard Damon writes: >> >>> On 9/11/22 11:01 PM, Tim Rentsch wrote: >>> >>>> Ben Bacarisse writes: >>>> >>>>> [...] I worry that, except when matching externally defined >>>>> structures, the use of the intN_t (and maybe even the uintN_t) >>>>> types is almost always knee-jerk programming. >>>> >>>> Completely agree. >>> >>> I would just slightly disagree. >>> >>> Many programs are not designed to be "highly portable", and it may >>> be a reasonable assumption that a given program will only be run >>> on machines which is based on 8 bit bytes. >> >> I am taking 8-bit bytes as being the crucial property in the >> remarks below. There are obvious counterparts for other >> properties, such as having a two's complement representation. >> >>> Since [u]intN_t is the shortest way to specify a size that will >>> likely migrate over modereate architecture variations for variable >>> where you do want to specify the "size" of the variable, using >>> them makes sense. >> >> First off, I wouldn't say using int8_t or uint8_t is the shortest >> way to ensure 8-bit bytes. The way such types are typically >> used, there will be lots of 'int8_t's or 'uint8_t's scattered >> through the source code, with a correspondingly large footprint >> in the program source. It is shorter to write in just one header >> file in the program >> >> #include >> #if CHAR_BIT != 8 >> # error oops.. this program relies on chars being 8 bits... >> #endif >> >> typedef unsigned char U8; >> typedef signed char S8; >> >> which also can be used in implementations that don't support >> C99. > > Yes, but then you get the problem that stdint.h was made to solve, > that of many variations of this naming possible existing in the > program, and the risk of duplication of names. That's your imagination talking, not the C standard. > I suspect that one reason int8_t is slightly ugly is that it was a > name the standard could fairly safely grab and not get into too > many conflicts. > > Yes, a final application code could do something like that, and > not have issues, but library code can't afford the name pollution. The [u]intN_t types are not used elsewhere in the C library. (I believe that also holds true for the fast and least-width types.) Some people might imagine that they are, but they are not. >> Second, ignoring the question of which technique is shortest, >> using [u]int8_t is often not a good way. A problem with the >> exact width types is that they are both an under-specification and >> an over-specification: over-specification because they often >> include properties that aren't important (such as having no trap >> representations), and under-specification because they might not >> have properties that are otherwise desirable (such as being >> exempt from anti-alias rules, as the character types are). >> Granted, such problems are not likely to occur, but there is no >> good reason to take the risk when it can be easily avoided. > > Readability of simple STANDARD names is important. The looser > specified types will have almost exactly properties. Remember, > the criteria was that the programmer KNEW that the machine was a > standard 8-bit byte two's complement machine due to other > dependencies in the code, thus KNOWS that the "over-specified" > properties WILL hold. > > As far as exempt from the anti-alias rules, I don't think there > has been a "standard" byte sized type that doesn't have that > exemption (I think they are adding one in the newest Standard, but > that is a very new addition. Remember, int_least8_t is likely a > typedef to signed char so still gets that attribute. > >> Third, using exact width types can have consequences for parts of >> the code other than declarations. In effect using exact width >> types in one part of a program pushes the program into a mode >> where exact width types more or less have to be used everywhere, >> even when they are needed in only a few places. Even if exact >> width types are an acceptable choice, in most cases they are not >> the best choice. > > As I pointed out, this method is used when you can assume the > architecture supports those types. Except for buffers and the > like passed by pointers (which likely should be void* or the exact > size likely IS important) since these types are just typedefs, > calling a routine defined to take a int_least16_t can easily be > called with a int16_t parameter with NO problem. > >> In summary, I think a better statement is that [u]intN_t is the >> laziest way to achieve what may not be a good result, and often >> is a suboptimal result. The sad thing is that better techniques >> don't need a lot of overhead; there is a small upfront cost, but >> once that initial cost has been paid the incremental cost is >> basically zero. > > That exact same arguement is what causes some people to do things > like > > #define F for > #define R return > ... > and get obfuscated code. I don't think so. Similar arguments maybe, but certainly not exactly the same. > The big issue with user type like u8 is that this becomes a risk in > library code that needs to be able to move to other projects. > > I would disagree that they "often" result is "suboptimal" results, as > it looks like you methods end up with EXACTLY the same fundamental > types being uses in the normal case, so there as no suboptimal. > > Yes, if your goal is "maximum portability" to unusual architectures, > you need to be a bit more careful. > >>> Perhaps instead of making the type least_intN_t/fast_intN_t/intN_t >>> they were defined instead as intN_t/fast_intN_t/exact_intN_t, or >>> least_intN_t/intN_t/exact_intN_t so the simple name was one of the >>> promised to exist types that might be bigger if needed, the short name >>> bias would have lead to better code, but that ship has sailed. >> >> Something that hasn't come up (I think) in the discussion of >> various type choices is the distinction between types and the >> names of types. In the code that I posted, the types used are >> unsigned char and unsigned int. To help the discussion I put in >> some typedefs to give the names UC and UI to these types, but >> that doesn't change what the underlying types are. One problem >> with the types in is we don't know whether these names >> refer to existing (standard) types, or whether they refer to >> non-standard extended types. Are they fish or fowl? We don't >> know. That uncertainty can cause difficulties in code that uses >> some types. For the most part I usually avoid using >> any types, for this reason and also for the reasons >> explained above. > > Does it matter if they are "standard" types of not? > > Again, my premise was that I KNEW my code was going to run on a > "Standard" 8 bit byte, twos complement machine with all the normal > assumptions, because other parts were something somewhat machine > specific. > > Yes, if you are writting maximally compatible code that might end up > on something like a DSP where that doesn't hold, don't do it that way. ISTM your argument boils down to, because a bunch of other people used the exact-width types wrongly, that I should also. Other people may find such an argument convincing, but I do not.