Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.os.msdos.programmer > #1086 > unrolled thread

Smaller C compiler

Started by"Alexei A. Frounze" <alexfrunews@gmail.com>
First post2013-12-01 08:15 -0800
Last post2015-09-06 15:55 -0700
Articles 20 on this page of 85 — 5 participants

Back to article view | Back to comp.os.msdos.programmer


Contents

  Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-01 08:15 -0800
    Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-01 18:42 +0000
      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-01 19:14 -0800
        Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-02 11:13 +0000
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-02 03:21 -0800
            Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-03 06:21 +0000
              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-02 23:17 -0800
                Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-03 13:31 +0000
                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-03 06:05 -0800
                  Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-03 10:47 -0500
                    Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-03 21:18 -0800
    Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-03 12:23 -0500
      Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-03 16:26 -0500
        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-03 23:58 -0800
          Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-04 04:31 -0500
            Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-04 05:01 -0500
              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-04 02:20 -0800
                Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-05 06:33 -0500
                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-05 21:37 -0800
                    Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-06 12:19 -0500
                      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-06 22:23 -0800
                        Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-07 13:48 -0500
                          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-07 16:34 -0800
                            Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-08 02:12 -0500
                              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-13 03:32 -0800
                                Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-15 04:47 -0500
                                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-15 02:14 -0800
                                    Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-15 13:21 -0500
            Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-04 23:39 -0800
              Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-05 06:38 -0500
                Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-05 21:40 -0800
                Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-06 04:19 -0800
                  Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-06 13:40 -0500
                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-14 21:07 -0800
      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-03 22:22 -0800
        Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-04 04:30 -0500
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-04 23:21 -0800
            Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-05 04:23 -0500
              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-05 21:31 -0800
    Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-10 04:59 +0000
      Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-10 06:17 +0000
        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-11 02:36 -0800
          Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-11 13:44 +0000
    Re: Smaller C compiler Harry Potter <rose.joseph12@yahoo.com> - 2013-12-10 10:41 -0800
      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-11 03:06 -0800
        Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-11 10:47 -0500
          Re: Smaller C compiler Harry Potter <rose.joseph12@yahoo.com> - 2013-12-11 08:01 -0800
            Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-11 13:40 -0500
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-11 22:07 -0800
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-25 02:33 -0800
            Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-29 03:24 -0800
    Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-21 15:55 -0800
    Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-29 17:16 +0000
      Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2013-12-29 21:54 -0500
        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-29 19:38 -0800
      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-29 19:36 -0800
        Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2013-12-30 05:07 +0000
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2013-12-29 21:58 -0800
          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-01-05 19:43 -0800
            Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2014-01-06 15:27 +0000
              Re: Smaller C compiler "Rod Pemberton" <dont_use_email@xnohavenotit.cnm> - 2014-01-06 19:51 -0500
                Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2014-01-07 04:27 +0000
                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-01-06 21:52 -0800
            Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-02-16 22:57 -0800
              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-02-25 00:17 -0800
                Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-03-01 21:31 -0800
                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-03-10 01:46 -0700
                    Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-04-20 20:19 -0700
                      Re: Smaller C compiler Harry Potter <rose.joseph12@yahoo.com> - 2014-04-23 11:20 -0700
                        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-04-23 11:41 -0700
                          Re: Smaller C compiler Harry Potter <rose.joseph12@yahoo.com> - 2014-04-25 06:08 -0700
                      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-09-14 03:06 -0700
                        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-11-09 03:26 -0800
                          Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-11-28 04:06 -0800
                            Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2014-12-21 01:56 -0800
                              Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-01-10 10:37 -0800
                                Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-04-19 23:55 -0700
                                  Re: Smaller C compiler "Auric__" <not.my.real@email.address> - 2015-04-22 06:13 +0000
                                    Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-04-22 00:12 -0700
                                  Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-04-25 21:11 -0700
                                    Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-05-16 17:49 -0700
                                      Re: Smaller C compiler "Bill Buckels" <bbuckels@mts.net> - 2015-05-19 20:01 -0500
                                        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-05-19 18:18 -0700
                                      Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-08-15 02:13 -0700
                                        Re: Smaller C compiler "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-06 15:55 -0700

Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →


#1124

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-06 22:23 -0800
Message-ID<b2b9f18d-6cde-44d0-97ff-eb02a2a83e67@googlegroups.com>
In reply to#1119
On Friday, December 6, 2013 9:19:10 AM UTC-8, Rod Pemberton wrote:
> On Fri, 06 Dec 2013 00:37:26 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> > On Thursday, December 5, 2013 3:33:51 AM UTC-8, Rod Pemberton wrote:
> >> On Wed, 04 Dec 2013 05:20:43 -0500, Alexei A. Frounze
> >> <...@gmail.com> wrote:
> >> > On Wednesday, December 4, 2013 2:01:02 AM UTC-8, Rod Pemberton wrote:
> >> >> On Wed, 04 Dec 2013 04:31:56 -0500, Rod Pemberton
> 
> >> >> At the "Constant too big for %d-bit type" line in smlrc.c.
> >> >>
> >> >> SizeOfWord is 4, for 32-bits.
> >> >> sizeof(n) is 4, where n is an 'unsigned'.
> >> >> n >> 8 >> 12 >> 12 is 31.
> >> >> n >> 32 is 0.
> >> >
> >> > Precisely. The idea is to make sure that if n (IOW, unsigned)
> >> > is larger than 32 bits, I can see if it's value doesn't fit
> >> > into 32 bits.
> >> >
> >> >> Could OW v1.3 be having problems with the multiple shifts?
> >> >
> >> > I don't know for sure, but I bet it very well could [...]
> >>
> >> I did the following:
> >>
> >>    changed multiple >> to single in the source
> >>    changed multiple << to single in the source
> >>
> >> After that, I get the following error (32-bit):
> >>
> >>    Error in "smlrc.c" (374:14)
> >>
> >>    Variable(s) take(s) too much space
> >>
> >> This seems to be from errorVarSize().  It's the errorVarSize()
> >> call with near "(size != truncUint(size))" in case '(' of
> >> GetDeclSize().  size is 4.  truncUint(size) is 0.  32-bits.
> >> You use ~0u in truncUint().
> >>
> >> OpenWatcom v1.3 results of ~0u and with shifts:
> >>
> >> wcl/l=dos
> >>
> >> ~0u         8000ffff
> >> ~0u<<15     80000000
> >> ~0u<<16         0000
> >> ~0u<<8<<7       0002
> >> ~0u<<8<<8       6322 or 6333
> >
> > How come 16-bit ints look like 32-bit? Where's the garbage coming from?
> > I mean "8000" in "8000ffff".
> >
> 
> It's coming from the program when it prints the value.
> 
> Specifically, it's from a "%04lx" to a printf().  It overflowed.  E.g.,
> 
>    printf("%04lx\n",~0u);
> 
> Yes, in theory, it should truncate to only four digits for display...

In theory and in practice you shouldn't lie to printf() about the type of the optional parameters. If you write "%lx" but then supply something that's not a long integer (unsigned, AFAIR), you get undefined behavior from printf(). 0 and 0u are an int and an unsigned int, and not any kind of long int.

> Most likely, the compiler is using a larger intermediate representation,
> for some reason, i.e., an "implicit" cast or size conversion.

Nope, it isn't using that. Or, if it is, at least, it must be completely transparent. You have a type mismatch.

> FYI, a "%08lx" was used for 32-bit values, which displayed 8 digits
> for each.

That's because in 32-bit mode on x86 typically (or at least in DOS and Windows) int and long are both defined as 32-bit. There are exceptions, however. If I'm not mistaken, in x86 Linux' gcc long is 64-bit.

> Personally, if I was attempting to do something like this, which I
> generally wouldn't because of possible differences in representation,
> I would cast to a known size for each compiler, with my preference being
> "unsigned long", and logically and & with 0xFFFF or 0xFFFFFFFF, etc.

That'd work. C89 lacks a type specifier in format strings for size_t (%z in C99, AFAIR), so I may cast size_t/sizeof/etc to unsigned long and print it as unsigned long. Ditto for missing support for unit32_t and the like (I conditionally define uint32 as an unsigned int or an unsigned long, one of them must have 32 bits, and then, when printing, I cast thusly defined uint32 to unsigned long and print it as unsigned long).

> > Also, what does "6322 or 6333" mean?
> >
> 
> That depends on the call to truncUint() where I inserted a line to print
> all those values.  truncUint() is called multiple times when smrlc compiles
> itself.  It seems to print 6322 sometimes and 6333 other times, i.e.,
> appears it's being corrupted to me.  The values would tend to imply by
> characters overwriting to me, but I have no idea what the values actually
> indicate.

I'm guessing your OW 1.3 is badly broken.

> >> wcl386/l=dos4g
> >>
> >> ~0u             ffffffff
> >> ~0u<<31         80000000
> >> ~0u<<32         ffffffff
> >> ~0u<<8<<12<<11  80000000
> >> ~0u<<8<<12<<12  00000000
> >>
> >> Obviously, some things are not correct or computing as expected.
> >
> > For one thing, you shouldn't be expecting much from shifting a
> > 16-bit int by 16 bit positions or from shifting a 32-bit int by
> > 32 bit positions. Both invoke undefined behavior.
> >
> 
> I shouldn't?  You're the one doing so... <<8<<12<<12.  Yes?

Yes and no. <<8<<12<<12 isn't necessarily equivalent to <<(8+12+12). The restriction on the shift count applies to one shift. Hence, with unsigned int being at least 16 bits long, I can shift by however many positions I want, provided that I don't shift by more than 15 at a time. C typically translates << to a single shift instruction, without generating any code to check the shift count at run time. In most CPUs, shift instructions only take the least significant bits of the shift count (4 bits for a shift of a 16-bit value, 5 bits for a shift of a 32-bit value and so on) and ignore the rest. So, effectively, your <<32 reduces to shl reg, 0. That's hardly what you want.

[snip]

> >> It appears that Rugxulo has OpenWatcom versions 1.3 and 1.7
> >> on his website:
> >> https://sites.google.com/site/rugxulo/
> >
> > What's wrong with the official OW website?
> 
> AFAIK, nothing. they might still have 1.7 around, perhaps 1.3 too.
> Of course, they might not, which is what I suspect for 1.3.
> Anyway, the point was you could, if you wanted, check each to
> find out was is going on without needing me.  It's very possible
> 1.3 is broken, but then again, maybe not.  If I had 1.7 or 1.9,
> I'd test it for you, but I have no plans to install them.
> Personally, I decided to move away from using OW, even though
> it produces better code, IMO.

Oh, I know what's wrong with http://www.openwatcom.org/. It's down from time to time. :) It appears to be down right now (and has been recently). Anyway, I remember, I could download previous versions from the website (or the FTP server). IMHO, that's the place to go to. When it's up, of course. :)

Alex

[toc] | [prev] | [next] | [standalone]


#1127

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-07 13:48 -0500
Message-ID<op.w7qjvbdi5zc71u@localhost>
In reply to#1124
On Sat, 07 Dec 2013 01:23:14 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Friday, December 6, 2013 9:19:10 AM UTC-8, Rod Pemberton wrote:
>> On Fri, 06 Dec 2013 00:37:26 -0500, Alexei A. Frounze
>> <...@gmail.com> wrote:
>> > On Thursday, December 5, 2013 3:33:51 AM UTC-8, Rod Pemberton wrote:

>> >> OpenWatcom v1.3 results of ~0u and with shifts:
>> >>
>> >> wcl/l=dos
>> >>
>> >> ~0u         8000ffff
>> >> ~0u<<15     80000000
>> >> ~0u<<16         0000
>> >> ~0u<<8<<7       0002
>> >> ~0u<<8<<8       6322 or 6333
>> >
>> > How come 16-bit ints look like 32-bit? Where's the garbage coming  
>> from?
>> > I mean "8000" in "8000ffff".
>> >
>>
>> It's coming from the program when it prints the value.
>>
>> Specifically, it's from a "%04lx" to a printf().  It overflowed.  E.g.,
>>
>>    printf("%04lx\n",~0u);
>>
>> Yes, in theory, it should truncate to only four digits for display...
>
> In theory and in practice you shouldn't lie to printf() about the type  
> of the optional parameters. If you write "%lx" but then supply something  
> that's not a long integer (unsigned, AFAIR), you get undefined behavior  
> from printf(). 0 and 0u are an int and an unsigned int, and not any kind  
> of long int.
>

As a practical matter, with or without a cast, the argument is converted
to a hex unsigned long for display by the compiler.  Whether that  
conversion
is officially classified as "undefined behavior", or classified as an
"implicit cast", it should still convert types correctly for the situation,
or it should cause compilation to error.

If the OW compiler had determined a type mismatch, it would've
warned about it.  I.e., this implies it views the types as equivalent
or as a valid conversion, or promotion.

If a cast or "%hx" or &0xFFFF had been used, you might not have seen this
issue.  To me, the 8-digit displayed values implies the compiler is  
actually
using a 32-bit signed type for ~0u and etc values instead of 16-bit  
unsigned.
When the converted result fits into the display format "%04lx", which is
defined behavior, we're seeing four digits as expected.  When the converted
result doesn't fit the display format, which is officially undefined  
behavior,
e.g., when signed or larger than 64K, it displays the entire,  
non-truncated,
value using eight digits.

>> >> wcl386/l=dos4g
>> >>
>> >> ~0u             ffffffff
>> >> ~0u<<31         80000000
>> >> ~0u<<32         ffffffff
>> >> ~0u<<8<<12<<11  80000000
>> >> ~0u<<8<<12<<12  00000000
>> >>
>> >> Obviously, some things are not correct or computing as expected.
>> >
>> > For one thing, you shouldn't be expecting much from shifting a
>> > 16-bit int by 16 bit positions or from shifting a 32-bit int by
>> > 32 bit positions. Both invoke undefined behavior.
>> >
>>
>> I shouldn't?  You're the one doing so... <<8<<12<<12.  Yes?
>
> Yes and no. <<8<<12<<12 isn't necessarily equivalent to <<(8+12+12).

Just as you "shouldn't lie to printf", you shouldn't abuse the
poorly, and not widely implemented, ANSI C concept of "sequence
points" in an attempt to avoid undefined behavior.

Alternately, you shouldn't abuse known limits of limits.h to
attempt to reason your way around undefined behavior, either.

A C compiler is not as smart as you.  It does things one way for
all apparently equivalent situations.  It's goal is to get
reasonably C compliant assembly in the most sane way possible.
So, if undefined behavior exists in one instance, that same undefined
behavior situation will likely be valid for other instances which
are effectively the same, but not undefined.  I.e., if <<32 on
32-bit is undefined, then effectively <<8<<12<<12 is too.

This is also an example why you have to consider what is going
to be done by the C compiler at the assembly level.  No matter
how <<8<<12<<12 is implemented, it will zero a 32-bit integer.

> The restriction on the shift count applies to one shift. Hence, with
> unsigned int being at least 16 bits long, I can shift by however many
> positions I want, provided that I don't shift by more than 15 at a time.

That officially avoids undefined behavior in C.

It doesn't avoid overflowing integers at the assembly level,
which someone might claim is the basis for undefined behaviors
in C...  ;-)

> C typically translates << to a single shift instruction,

It may.  The underlying architecture is assumed to be unknown to C.
So, the C compiler is free to implement shifts as whatever it takes.
It can do multiple to single, or single to multiple, or even other
non-shift manipulations.

> without generating any code to check the shift count at run time.

True.  That's checked at compile time.

> In most CPUs,
> shift instructions only take the least significant bits of the shift
> count (4 bits for a shift of a 16-bit value, 5 bits for a shift of a
> 32-bit value and so on)

Well, it seems to be 5 bits on x86, for both regular and double shifts.

> and ignore the rest.

The assembly instruction may ignore the remainder of the bits
effectively imposing a per instruction shift limit.  That just
means more instructions are required to implement larger shifts.

The C compiler is required to implement the total shift count using as
many assembly instructions as is needed.  The C compiler can't just
ignore 3 bits of shift for <<8 by truncating it to <<5 because
the CPU only supports <<5...

> So, effectively, your <<32 reduces to shl reg, 0.

No.

Perhaps, you're thinking of "rol reg" instead of "shl reg"?  ;-)

While the effective shift count for "shl reg, 0" is the same
as <<32 due to wrap around, "shl reg, 0" doesn't actually *clear*
the register.  It leaves it alone.  SHL shifts in a cleared bit
per shift.  ROL rotates one bit per shift.

A shift left of 32 bits or more will zero a 32-bit register.
So, <<32 would effectively reduce to equivalent of 'xor reg, reg',
or 'mov reg, 0'.  Of course, most C compilers would use multiple
shl shifts to shift 32 times, but some might optimize to a register
clear.

> That's hardly what you want.

True.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1132

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-07 16:34 -0800
Message-ID<92e0fc2b-68e7-41c4-a5ae-e1469f7047fa@googlegroups.com>
In reply to#1127
On Saturday, December 7, 2013 10:48:37 AM UTC-8, Rod Pemberton wrote:
> On Sat, 07 Dec 2013 01:23:14 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> > On Friday, December 6, 2013 9:19:10 AM UTC-8, Rod Pemberton wrote:
> >> On Fri, 06 Dec 2013 00:37:26 -0500, Alexei A. Frounze
> >> <...@gmail.com> wrote:
> >> > On Thursday, December 5, 2013 3:33:51 AM UTC-8, Rod Pemberton wrote:
> 
> >> >> OpenWatcom v1.3 results of ~0u and with shifts:
> >> >>
> >> >> wcl/l=dos
> >> >>
> >> >> ~0u         8000ffff
> >> >> ~0u<<15     80000000
> >> >> ~0u<<16         0000
> >> >> ~0u<<8<<7       0002
> >> >> ~0u<<8<<8       6322 or 6333
> >> >
> >> > How come 16-bit ints look like 32-bit? Where's the garbage coming  
> >> from?
> >> > I mean "8000" in "8000ffff".
> >> >
> >>
> >> It's coming from the program when it prints the value.
> >>
> >> Specifically, it's from a "%04lx" to a printf().  It overflowed.  E.g.,
> >>
> >>    printf("%04lx\n",~0u);
> >>
> >> Yes, in theory, it should truncate to only four digits for display...
> >
> > In theory and in practice you shouldn't lie to printf() about the type  
> > of the optional parameters. If you write "%lx" but then supply something  
> > that's not a long integer (unsigned, AFAIR), you get undefined behavior  
> > from printf(). 0 and 0u are an int and an unsigned int, and not any kind  
> > of long int.
> >
> 
> As a practical matter, with or without a cast, the argument is converted
> to a hex unsigned long for display by the compiler.  Whether that  
> conversion
> is officially classified as "undefined behavior", or classified as an
> "implicit cast", it should still convert types correctly for the situation,
> or it should cause compilation to error.

You are wrong. Optional arguments (that fall under '...') do not get converted from int to long, neither when passed to the function nor when already in the function just for your convenience to choose at will between %d and %l. You can only get away with printf'ing ints as longs (or the other way around) when their sizes are the same, as I've already mentioned.

> If the OW compiler had determined a type mismatch, it would've
> warned about it.  I.e., this implies it views the types as equivalent
> or as a valid conversion, or promotion.

In this case a conforming compiler is not required to tell you of undefined behavior when it sees it, it's not even required to look for it. And in other cases it may be simply impossible for the compiler to find UB.

So, if OW doesn't warn you, use a better compiler. See how properly tuned gcc detects this:
http://ideone.com/sOGFFP

and that's despite int and long having the same size in ideone's gcc:
http://ideone.com/iKnXJP

> If a cast or "%hx" or &0xFFFF had been used, you might not have seen this
> issue.  To me, the 8-digit displayed values implies the compiler is  
> actually
> using a 32-bit signed type for ~0u and etc values instead of 16-bit  
> unsigned.

It's not that. It's that when printf() looks at the format and sees %lx in it, it grabs a long int (32-bit in this case) from the stack. And if you've only put there (on the stack) an int (16-bit in this case), you're rightly screwed.

> When the converted result fits into the display format "%04lx", which is
> defined behavior, we're seeing four digits as expected.  When the converted
> result doesn't fit the display format, which is officially undefined  
> behavior,
> e.g., when signed or larger than 64K, it displays the entire,  
> non-truncated,
> value using eight digits.

Wrong again. C89/C99 on 4 in in your %04lx, quote:
In no case does a nonexistent or small field width cause truncation of a field; if the result of a conversion is wider than the field width, the field is expanded to contain the conversion result.

> >> >> wcl386/l=dos4g
> >> >>
> >> >> ~0u             ffffffff
> >> >> ~0u<<31         80000000
> >> >> ~0u<<32         ffffffff
> >> >> ~0u<<8<<12<<11  80000000
> >> >> ~0u<<8<<12<<12  00000000
> >> >>
> >> >> Obviously, some things are not correct or computing as expected.
> >> >
> >> > For one thing, you shouldn't be expecting much from shifting a
> >> > 16-bit int by 16 bit positions or from shifting a 32-bit int by
> >> > 32 bit positions. Both invoke undefined behavior.
> >> >
> >>
> >> I shouldn't?  You're the one doing so... <<8<<12<<12.  Yes?
> >
> > Yes and no. <<8<<12<<12 isn't necessarily equivalent to <<(8+12+12).
> 
> Just as you "shouldn't lie to printf", you shouldn't abuse the
> poorly, and not widely implemented, ANSI C concept of "sequence
> points" in an attempt to avoid undefined behavior.
> 
> Alternately, you shouldn't abuse known limits of limits.h to
> attempt to reason your way around undefined behavior, either.

Irrelevant here.

> A C compiler is not as smart as you.  It does things one way for
> all apparently equivalent situations.  It's goal is to get
> reasonably C compliant assembly in the most sane way possible.
> So, if undefined behavior exists in one instance, that same undefined
> behavior situation will likely be valid for other instances which
> are effectively the same, but not undefined.  I.e., if <<32 on
> 32-bit is undefined, then effectively <<8<<12<<12 is too.

I don't know if you're just trolling me or indeed do not understand (and seem unwilling to learn) how several basic things ought to work in C according to the language standard.

<<8<<12<<12 is defined. It may cause UB if the value being shifted is of some signed integer type. If you remember, I do this with unsigned integers.

> This is also an example why you have to consider what is going
> to be done by the C compiler at the assembly level.  No matter
> how <<8<<12<<12 is implemented, it will zero a 32-bit integer.

It will zero an unsigned 32-bit int. If the unsigned int is 16-bit, it will be zeroes too.

> > The restriction on the shift count applies to one shift. Hence, with
> > unsigned int being at least 16 bits long, I can shift by however many
> > positions I want, provided that I don't shift by more than 15 at a time.
> 
> That officially avoids undefined behavior in C.
> 
> It doesn't avoid overflowing integers at the assembly level,
> which someone might claim is the basis for undefined behaviors
> in C...  ;-)

That does not apply to unsigned int. You're free to overflow it in C however you like, the behavior is defined.

[snip]

> > In most CPUs,
> > shift instructions only take the least significant bits of the shift
> > count (4 bits for a shift of a 16-bit value, 5 bits for a shift of a
> > 32-bit value and so on)
> 
> Well, it seems to be 5 bits on x86, for both regular and double shifts.
> 
> > and ignore the rest.
> 
> The assembly instruction may ignore the remainder of the bits
> effectively imposing a per instruction shift limit.  That just
> means more instructions are required to implement larger shifts.

That's the reason for the restriction appearing in the standard. And, transitively, that's the reason for me doing <<8<<12<<12.

> The C compiler is required to implement the total shift count using as
> many assembly instructions as is needed.  The C compiler can't just
> ignore 3 bits of shift for <<8 by truncating it to <<5 because
> the CPU only supports <<5...

The compiler can ignore bits if the count is too big or negative. If the count is in the valid range, there's nothing to ignore.

> > So, effectively, your <<32 reduces to shl reg, 0.
> 
> No.

Why not? It's UB in this case, sure. And the UB in <<32 often manifests in really doing <<(32&31). Most likely that's exactly how you got this with 32-bit int:

~0u<<32         ffffffff 

It is not the only possible way for UB to manifest itself, but it's a common one.

> Perhaps, you're thinking of "rol reg" instead of "shl reg"?  ;-)
> 
> While the effective shift count for "shl reg, 0" is the same
> as <<32 due to wrap around, "shl reg, 0" doesn't actually *clear*
> the register.  It leaves it alone.  

Precisely. That's your:

~0u<<32         ffffffff 

> SHL shifts in a cleared bit
> per shift.  ROL rotates one bit per shift.
> 
> A shift left of 32 bits or more will zero a 32-bit register.
> So, <<32 would effectively reduce to equivalent of 'xor reg, reg',
> or 'mov reg, 0'.  Of course, most C compilers would use multiple
> shl shifts to shift 32 times, but some might optimize to a register
> clear.

Again, if your unsigned int has 32 bits in it, applying <<32 to it is incorrect, it will cause UB, which may not clear said unsigned int. You have produced and seen it yourself:

~0u<<32         ffffffff 

And here's the exact same UB from gcc with me forcing it to happen:
http://ideone.com/iKnXJP

And without forcing, it just fails to compile:
http://ideone.com/6lB1hh

Alex

[toc] | [prev] | [next] | [standalone]


#1140

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-08 02:12 -0500
Message-ID<op.w7riacro5zc71u@localhost>
In reply to#1132
On Sat, 07 Dec 2013 19:34:58 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Saturday, December 7, 2013 10:48:37 AM UTC-8, Rod Pemberton wrote:
>> On Sat, 07 Dec 2013 01:23:14 -0500, Alexei A. Frounze
>> <...@gmail.com> wrote:
>> > On Friday, December 6, 2013 9:19:10 AM UTC-8, Rod Pemberton wrote:
>> >> On Fri, 06 Dec 2013 00:37:26 -0500, Alexei A. Frounze
>> >> <...@gmail.com> wrote:
>> >> > On Thursday, December 5, 2013 3:33:51 AM UTC-8, Rod Pemberton  
>> wrote:

>> >> >> OpenWatcom v1.3 results of ~0u and with shifts:
>> >> >>
>> >> >> wcl/l=dos
>> >> >>
>> >> >> ~0u         8000ffff
>> >> >> ~0u<<15     80000000
>> >> >> ~0u<<16         0000
>> >> >> ~0u<<8<<7       0002
>> >> >> ~0u<<8<<8       6322 or 6333
>> >> >
>> >> > How come 16-bit ints look like 32-bit? Where's the garbage coming
>> >> from?
>> >> > I mean "8000" in "8000ffff".
>> >> >
>> >>
>> >> It's coming from the program when it prints the value.
>> >>
>> >> Specifically, it's from a "%04lx" to a printf().  It overflowed.  >>  
>> E.g.,
>> >>
>> >>    printf("%04lx\n",~0u);
>> >>
>> >> Yes, in theory, it should truncate to only four digits for display...
>> >
>> > In theory and in practice you shouldn't lie to printf() about the
>> > type of the optional parameters. If you write "%lx" but then supply
>> > something that's not a long integer (unsigned, AFAIR), you get
>> > undefined behavior from printf(). 0 and 0u are an int and an unsigned
>> > int, and not any kind of long int.
>> >
>>
>> As a practical matter, with or without a cast, the argument is converted
>> to a hex unsigned long for display by the compiler.  Whether that
>> conversion is officially classified as "undefined behavior", or
>> classified as an "implicit cast", it should still convert types
>> correctly for the situation, or it should cause compilation to error.
>
> You are wrong. Optional arguments (that fall under '...') do not get  
> converted from int to long, neither when passed to the function nor when  
> already in the function just for your convenience to choose at will  
> between %d and %l. You can only get away with printf'ing ints as longs  
> (or the other way around) when their sizes are the same, as I've already  
> mentioned.
>

See below.

>> If a cast or "%hx" or &0xFFFF had been used, you might not have seen  
>> this issue.  To me, the 8-digit displayed values implies the compiler
>> is actually using a 32-bit signed type for ~0u and etc values instead
>> of 16-bit unsigned.
>
> It's not that. It's that when printf() looks at the format and
> sees %lx in it, it grabs a long int (32-bit in this case) from
> the stack.

C doesn't have a stack officially.  So, how can printf() grab a
long from something that officially doesn't exist?  ;-)  The
specific implementation may have a stack and printf() may obtain
a long from from it.  That doesn't mean there is a mismatch in
type sizes.

Of course, the majority of C compilers use a stack.  It allows
for recursion to be implemented easily.

Anyway, I believe the compiler should've converted the type.

E.g., from Harbison & Steele, "C:A Reference Manual", 3rd, ed.:

"An actual argument to a function may be implicitly converted
to another type prior to the function call."

I.e., even if not always implicitly converted, it may be sometimes.
Meaning one compiler can while another might not.

Also, on integer conversions:

"If the destination type is longer than the source type, then
the only case in which the source value will not be representable
in the result type is when a negative signed value is converted
to a longer, unsigned type."

Of course, the source type to "%04lx" is a 16-bit unsigned and
is converted to a unsigned destination type, likely longer.

So, that and other similar statements in H&S are in the C
specification somewhere, most likely with different wording.

>> When the converted result fits into the display format "%04lx",
>> which is defined behavior, we're seeing four digits as expected.
>> When the converted result doesn't fit the display format, which
>> is officially undefined behavior, e.g., when signed or larger
>> than 64K, it displays the entire, non-truncated, value using
>> eight digits.
>
> Wrong again. C89/C99 on 4 in in your %04lx, quote:
> In no case does a nonexistent or small field width cause truncation
> of a field; if the result of a conversion is wider than the field
> width, the field is expanded to contain the conversion result.
>

Isn't that what I just said?  Try reading it without "non-truncated".
Then, assume the true size is 32-bits - not 16-bits - since that
was my perspective.  Re-read with "non-truncated".

>> Just as you "shouldn't lie to printf", you shouldn't abuse the
>> poorly, and not widely implemented, ANSI C concept of "sequence
>> points" in an attempt to avoid undefined behavior.
>>
>> Alternately, you shouldn't abuse known limits of limits.h to
>> attempt to reason your way around undefined behavior, either.
>
> Irrelevant here.
>

Why? The first is exactly what you're doing, IMO.  Don't you think
that's a possible reason why your code doesn't work?  Personally,
I think it's an more likely error in OW since I found more than
a few, but it's just as plausible that your abuse of the specification
is at fault.  ISTM, you're attempting to bypass behavior which
may not be bypassable on every single compiler.

>> A C compiler is not as smart as you.  It does things one way for
>> all apparently equivalent situations.  It's goal is to get
>> reasonably C compliant assembly in the most sane way possible.
>> So, if undefined behavior exists in one instance, that same undefined
>> behavior situation will likely be valid for other instances which
>> are effectively the same, but not undefined.  I.e., if <<32 on
>> 32-bit is undefined, then effectively <<8<<12<<12 is too.
>
> I don't know if you're just trolling me or indeed do not understand
> (and seem unwilling to learn) how several basic things ought to work
> in C according to the language standard.
>

Yet, even though you know the C's language standard has no defined
stack, you used a stack to justify your earlier logic on how printf()
works and fails...  What did you say about not understanding?  :-)

I'm not trying to troll you.  I'm just not interested in boundary
conditions such as UB, limits.h, or differences of type formats.
I know how the underlying assembly works.  I generally know how
C compilers convert much of the code to assembly.  C has a number
of "dark corners" where real-world C compilers work very differently
 from each other, and where historically implemented things work
differently from the C standards.  I try to avoid those corners.
IMO, that's much better than attempting to master them.

Also, it seems you're not always very clear on what C does and
what assembly does, i.e., sometimes equating/conflating the two.

> <<8<<12<<12 is defined. It may cause UB if the value being shifted
> is of some signed integer type. If you remember, I do this with
> unsigned integers.

Yes, you've technically avoided UB in C by doing small shifts, but
probably haven't avoided UB in reality since the small shifts are
equivalent to large shifts, and you definately haven't avoided
zeroing a 32-bit unsigned integer.

>> This is also an example why you have to consider what is going
>> to be done by the C compiler at the assembly level.  No matter
>> how <<8<<12<<12 is implemented, it will zero a 32-bit integer.
>
> It will zero an unsigned 32-bit int. If the unsigned int is 16-bit,
> it will be zeroes too.
>

Yes.

>> > The restriction on the shift count applies to one shift. Hence, with
>> > unsigned int being at least 16 bits long, I can shift by however many
>> > positions I want, provided that I don't shift by more than 15 at a  
>> time.
>>
>> That officially avoids undefined behavior in C.
>>
>> It doesn't avoid overflowing integers at the assembly level,
>> which someone might claim is the basis for undefined behaviors
>> in C...  ;-)
>
> That does not apply to unsigned int. You're free to overflow
> it in C however you like, the behavior is defined.
>
> [snip]

You seem to be contradicting your earlier statements on UB here.

>> > In most CPUs,
>> > shift instructions only take the least significant bits of the shift
>> > count (4 bits for a shift of a 16-bit value, 5 bits for a shift of a
>> > 32-bit value and so on)
>>
>> Well, it seems to be 5 bits on x86, for both regular and double shifts.
>>
>> > and ignore the rest.
>>
>> The assembly instruction may ignore the remainder of the bits
>> effectively imposing a per instruction shift limit.  That just
>> means more instructions are required to implement larger shifts.
>
> That's the reason for the restriction appearing in the standard. And,  
> transitively, that's the reason for me doing <<8<<12<<12.

Let's start over.  How is <<8<<12<<12 any different from <<32?
Now, exclude C's definition of UB.  Did anything the compiler
does for either change?  Prove it.  Now, prove it for all C
compilers.  You can't.  The C specification doesn't specify
implementation.  The C compiler implements the specification
to the best of the compiler authors' abilities.

>> > So, effectively, your <<32 reduces to shl reg, 0.
>>
>> No.
>
> Why not? It's UB in this case, sure. And the UB in <<32 often manifests  
> in really doing <<(32&31). Most likely that's exactly how you got this  
> with 32-bit int:
>
> ~0u<<32         ffffffff
>
> It is not the only possible way for UB to manifest itself, but it's a  
> common one.
>

"Why not?"

Because what you're effectively claiming is that <<8 or <<12
or <<31 will work but <<32 or <<35 doesn't work simply because
the C specification calls it undefined behavior.  That's really
not how things work.  Do you seriously believe that the compiler
implementor checks for the shift count and decided to stop at 31
or do something special thereafter? No, of course not, they
implemented a generic shift routine independent of shift count.
So, why is <<31 so different from <<32? ...

>> Perhaps, you're thinking of "rol reg" instead of "shl reg"?  ;-)
>>
>> While the effective shift count for "shl reg, 0" is the same
>> as <<32 due to wrap around, "shl reg, 0" doesn't actually *clear*
>> the register.  It leaves it alone.
>
> Precisely.

... which means it's in error.  The behavior is incorrect.
That was stated previously.

>> SHL shifts in a cleared bit
>> per shift.  ROL rotates one bit per shift.
>>
>> A shift left of 32 bits or more will zero a 32-bit register.
>> So, <<32 would effectively reduce to equivalent of 'xor reg, reg',
>> or 'mov reg, 0'.  Of course, most C compilers would use multiple
>> shl shifts to shift 32 times, but some might optimize to a register
>> clear.
>
> Again, if your unsigned int has 32 bits in it, applying <<32 to it is  
> incorrect, it will cause UB, which may not clear said unsigned int.

How is that any different from <<8<<12<<12?  AISI, you could very
well have a C compiler which fails for it too, e.g., converting that
to <<32, or converting both to the same shift sequence in assembly,
if one fails so would the other.  So far, AISI, you've not presented
any justification as to why one sequence *must be* different from
the other in assembly.  My understanding is the compiler is free
to optimized that as long as the code doesn't change semantically.
AFAICT, it doesn't.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1169

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-13 03:32 -0800
Message-ID<68c721c7-4bbb-4fe2-a529-22176d3cf980@googlegroups.com>
In reply to#1140
On Saturday, December 7, 2013 11:12:02 PM UTC-8, Rod Pemberton wrote:
> On Sat, 07 Dec 2013 19:34:58 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> > On Saturday, December 7, 2013 10:48:37 AM UTC-8, Rod Pemberton wrote:
> >> On Sat, 07 Dec 2013 01:23:14 -0500, Alexei A. Frounze
> >> <...@gmail.com> wrote:
> >> > On Friday, December 6, 2013 9:19:10 AM UTC-8, Rod Pemberton wrote:
> >> >> On Fri, 06 Dec 2013 00:37:26 -0500, Alexei A. Frounze
> >> >> <...@gmail.com> wrote:
> >> >> > On Thursday, December 5, 2013 3:33:51 AM UTC-8, Rod Pemberton  
> >> wrote:
...
> > It's not that. It's that when printf() looks at the format and
> > sees %lx in it, it grabs a long int (32-bit in this case) from
> > the stack.
> 
> C doesn't have a stack officially.  So, how can printf() grab a
> long from something that officially doesn't exist?  ;-)  The
> specific implementation may have a stack and printf() may obtain
> a long from from it.

The actual term/name or implementation (stack or abstract black box of sorts) isn't important, so long as it conforms to the standard. What is important is how you use this implied storage (which in practice often materializes as a stack).

> That doesn't mean there is a mismatch in
> type sizes.

Promising printf to give it a long and not giving it a long is a mismatch.

> Anyway, I believe the compiler should've converted the type.

Not the way you're expecting it to happen here.

> E.g., from Harbison & Steele, "C:A Reference Manual", 3rd, ed.:
> 
> "An actual argument to a function may be implicitly converted
> to another type prior to the function call."

It's possible.

> I.e., even if not always implicitly converted, it may be sometimes.
> Meaning one compiler can while another might not.

It's not exactly up to the compiler. There are specific rules for different situations. And this one is not one of those. If you pass a char or a short to printf(), it will first get converted into an int (possibly unsigned) and only then will it be available to printf(). Likewise, if you pass a float to printf(), it will first get converted into a double and will be available to printf() as a double.

See in C99:

6.3.1.1 Boolean, characters, and integers; clause 2 (on integer promotions)

6.5.2.2 Function calls; clauses 6, 7, 8 (on default argument promotions and UB)

7.15.1.1 The va_arg macro; clause 2 (on default argument promotions and UB)

7.19.6.1 The fprintf function; clauses 7, 8 and 9 (on the meaning of length modifiers and conversion specifiers and their relation to the corresponding optional parameters and UB), specifically, quote:

If a conversion specification is invalid, the behavior is undefined. If any argument is not the correct type for the corresponding conversion specification, the behavior is undefined.

I'm leaving you at that. I don't want to restate or copy and paste the standard. You can find and read the relevant text in it. Perhaps, if you own K&R 2nd ed., you can find the same info in it.

> Also, on integer conversions:
> 
> "If the destination type is longer than the source type, then
> the only case in which the source value will not be representable
> in the result type is when a negative signed value is converted
> to a longer, unsigned type."
> 
> Of course, the source type to "%04lx" is a 16-bit unsigned and
> is converted to a unsigned destination type, likely longer.

There's no destination type here. (Just like there's no stack.:)

> So, that and other similar statements in H&S are in the C
> specification somewhere, most likely with different wording.

I have not read or seen that book and I can't tell if it's misleading or just wrong (I've seen poor books on C) or you're not interpreting it correctly. I've mentioned the relevant parts of the C99 standard. It's a click away:

http://www.open-std.org/jtc1/sc22/wg14/www/standards.html:
The latest publicly available version of the C99 standard is the combined C99 + TC1 + TC2 + TC3, WG14 N1256, dated 2007-09-07. This is a WG14 working paper, but it reflects the consolidated standard at the time of issue. 

http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf

> >> When the converted result fits into the display format "%04lx",
> >> which is defined behavior, we're seeing four digits as expected.
> >> When the converted result doesn't fit the display format, which
> >> is officially undefined behavior, e.g., when signed or larger
> >> than 64K, it displays the entire, non-truncated, value using
> >> eight digits.
> >
> > Wrong again. C89/C99 on 4 in in your %04lx, quote:
> > In no case does a nonexistent or small field width cause truncation
> > of a field; if the result of a conversion is wider than the field
> > width, the field is expanded to contain the conversion result.
> >
> 
> Isn't that what I just said?  Try reading it without "non-truncated".
> Then, assume the true size is 32-bits - not 16-bits - since that
> was my perspective.  Re-read with "non-truncated".

The conversion mentioned in the quoted bit of the standard is the conversion of the captured numerical or pointer type (long if %ld, int if %d, void* if %p, etc) into printable text. It's not the intermediate conversion of int to long, after which you have the conversion into a bunch of (hopefully) printable characters, as you seem to be suggesting. It's that final conversion to printable stuff.

> >> Just as you "shouldn't lie to printf", you shouldn't abuse the
> >> poorly, and not widely implemented, ANSI C concept of "sequence
> >> points" in an attempt to avoid undefined behavior.
> >>
> >> Alternately, you shouldn't abuse known limits of limits.h to
> >> attempt to reason your way around undefined behavior, either.
> >
> > Irrelevant here.
> >
> 
> Why? The first is exactly what you're doing, IMO.  Don't you think
> that's a possible reason why your code doesn't work? 

While I have written lots of poor code, I have also learned from some of those mistakes. I'm avoiding UB by not triggering it, not by artful camouflaging of questionable actions. UB is not a fate. You can choose to have it or not.

Your compiler is broken.

And that's a normal occurrence in our non-ideal world.

> Personally,
> I think it's an more likely error in OW since I found more than
> a few, but it's just as plausible that your abuse of the specification
> is at fault.  ISTM, you're attempting to bypass behavior which
> may not be bypassable on every single compiler.

Yeah, you can't know beforehand of all the bugs in the compiler. :)

> >> A C compiler is not as smart as you.  It does things one way for
> >> all apparently equivalent situations.  It's goal is to get
> >> reasonably C compliant assembly in the most sane way possible.
> >> So, if undefined behavior exists in one instance, that same undefined
> >> behavior situation will likely be valid for other instances which
> >> are effectively the same, but not undefined.  I.e., if <<32 on
> >> 32-bit is undefined, then effectively <<8<<12<<12 is too.
> >
> > I don't know if you're just trolling me or indeed do not understand
> > (and seem unwilling to learn) how several basic things ought to work
> > in C according to the language standard.
> >
> 
> Yet, even though you know the C's language standard has no defined
> stack, you used a stack to justify your earlier logic on how printf()
> works and fails...  What did you say about not understanding?  :-)

Exactly how did the word stack itself justify anything? Would the word box justify something better, worse or about the same? Did I say stack overflow? Did I mention the order of things in that stack and use it in the reasoning? I didn't. An abstract storage (or maybe not even storage, but, say, floatage or flyage) and abstract methods of accessing it would do the same as "stack", but sound more foreign.

> I'm not trying to troll you.  I'm just not interested in boundary
> conditions such as UB, limits.h, or differences of type formats.

It's like saying "Judge, I didn't know the law, sorry". :) Really, you have to know UBs if you don't want them in your code. Looks like there's still something to learn.

> I know how the underlying assembly works.  I generally know how
> C compilers convert much of the code to assembly.  C has a number
> of "dark corners" where real-world C compilers work very differently
>  from each other, and where historically implemented things work
> differently from the C standards.  I try to avoid those corners.
> IMO, that's much better than attempting to master them.

The are dark corners. There are different architectures. There are bugs in compilers. Fortunately, in the past 25+ years, enough knowledge and experience has been acquired in the programmer community and there are very few truly dark and unexplored corners, even fewer if you consider only stuff of practical importance and not something as odd as I asked here:
http://stackoverflow.com/questions/20343190/evaluating-accessing-a-structure.

> Also, it seems you're not always very clear on what C does and
> what assembly does, i.e., sometimes equating/conflating the two.

I'm mostly clear, unless I forget something important. Remember that there are multiple interpretations going on: me translating my thoughts into English and you translating from my English into your thoughts. You think that can be always flawless? How about you doing more than one interpretation here and extending the total number to at least three? You are conditioned differently than I am and you interpret through the prism of your experience. I may not say something, but you may think I imply it. It's all possible.

> > <<8<<12<<12 is defined. It may cause UB if the value being shifted
> > is of some signed integer type. If you remember, I do this with
> > unsigned integers.
> 
> Yes, you've technically avoided UB in C by doing small shifts, but
> probably haven't avoided UB in reality since the small shifts are
> equivalent to large shifts, and you definately haven't avoided
> zeroing a 32-bit unsigned integer.

I avoided UB in those shifts technically and legally. Period.

> >> This is also an example why you have to consider what is going
> >> to be done by the C compiler at the assembly level.  No matter
> >> how <<8<<12<<12 is implemented, it will zero a 32-bit integer.
> >
> > It will zero an unsigned 32-bit int. If the unsigned int is 16-bit,
> > it will be zeroes too.
> >
> 
> Yes.
> 
> >> > The restriction on the shift count applies to one shift. Hence, with
> >> > unsigned int being at least 16 bits long, I can shift by however many
> >> > positions I want, provided that I don't shift by more than 15 at a  
> >> time.
> >>
> >> That officially avoids undefined behavior in C.
> >>
> >> It doesn't avoid overflowing integers at the assembly level,
> >> which someone might claim is the basis for undefined behaviors
> >> in C...  ;-)
> >
> > That does not apply to unsigned int. You're free to overflow
> > it in C however you like, the behavior is defined.
> >
> > [snip]
> 
> You seem to be contradicting your earlier statements on UB here.

Please try to read and make some sense of C99. The text is searchable in the PDF, if that helps.

This conversation is becoming meaningless. You have your own ideas how things should work and they don't quite match those expressed in the language standard. You may be finding holes and contradictions in my statements all day long until you've learned enough of standard C.

> >> > In most CPUs,
> >> > shift instructions only take the least significant bits of the shift
> >> > count (4 bits for a shift of a 16-bit value, 5 bits for a shift of a
> >> > 32-bit value and so on)
> >>
> >> Well, it seems to be 5 bits on x86, for both regular and double shifts.
> >>
> >> > and ignore the rest.
> >>
> >> The assembly instruction may ignore the remainder of the bits
> >> effectively imposing a per instruction shift limit.  That just
> >> means more instructions are required to implement larger shifts.
> >
> > That's the reason for the restriction appearing in the standard. And,  
> > transitively, that's the reason for me doing <<8<<12<<12.
> 
> Let's start over.  

Let's not. You have something to read and think through. I mean the standard. I've mentioned some relevant clauses and terms. Start there. Or, you could, read the whole thing from the beginning. It's not that bad.

[snip]

Alex

[toc] | [prev] | [next] | [standalone]


#1171

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-15 04:47 -0500
Message-ID<op.w74n58mu5zc71u@localhost>
In reply to#1169
On Fri, 13 Dec 2013 06:32:43 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:

> It's not exactly up to the compiler.

Yes, it is.

The C specification doesn't specify the low-level implementation.
It's an abstraction of C from the implementation details.

E.g., see early Forth specifications for comparison.  They specify
much of the implementation.

If you want to learn about the expected way to implement C, you need
to read the C Rationale as well as original papers by Kernighan, Ritchie,
Johnson, et. al. on the subject.  At least two current versions of the
C Rationale are online as pdf's.  I've not been able to locate the original
rationale as pdf although I've been searching for it for many years now.
(See "Rationale for American National Standard for Information Systems
- Programming Language - C")

Dennis Ritchie's papers on C
http://cm.bell-labs.com/who/dmr/

> There are specific rules for different situations.

True.

But, there are also contradictions which are part of the standards.
This is because the original C89 standard was derived from working
C compilers which worked differently.

How do you resolve a contradiction? Both are required...

> And this one is not one of those.

C preserves values (required).  There is a type mismatch not fixed
by a promotion, conversion, or explicit cast.   The only way to
preserve the value is to cast it to the parameter's type.  When
this is done by the C compiler it's called an "implicit cast."
This is a logical result of the C specifications, i.e., allowed
by it.

> If you pass a char or a short to printf(), it will first get
> converted into an int (possibly unsigned) and only then will
> it be available to printf().

It must be converted to a type that preserves the value.
Preservation of an integer's value is a fundamental C principle
and required by the specifications.

> Likewise, if you pass a float to printf(), it will first get
> converted into a double and will be available to printf() as
> a double.
>

Yes, there are defined type conversions and promotions, but you're
getting too caught up in the rules of C syntax or semantics and
ignoring how real world C compilers are actually implemented.  You
can't expect everything to work exactly the way the C specification
describes.  That's open to interpretation anyway, and I've seen
quite a few self-proclaimed experts on comp.lang.c. describe things
incorrectly.


IMO, good C compilers have a fundamental set of principles they follow:

  1) integer conversions preserve the value (*)
  2) lossless conversions of pointers (*)
  3) operators work as expected, independent of limits.h or UB
  4) implicit casts convert mismatched types (**)
  5) that which fits best onto the native assembly shall be done

(*) required by C specifications
(**) result of C specifications

>> Also, on integer conversions:
>>
>> "If the destination type is longer than the source type, then
>> the only case in which the source value will not be representable
>> in the result type is when a negative signed value is converted
>> to a longer, unsigned type."
>>
>> Of course, the source type to "%04lx" is a 16-bit unsigned and
>> is converted to a unsigned destination type, likely longer.
>
> There's no destination type here. (Just like there's no stack.:)
>

Yes, there is.

The parameter to printf() for %04lx has a defined type.  It's defined
as: a pointer to an unsigned long, i.e., it's the address of the value
being passed as %04lx which must be of type unsigned long.  Dereferencing
the pointer to convert the value results in a type of unsigned long.

>> So, that and other similar statements in H&S are in the C
>> specification somewhere, most likely with different wording.
>
> I have not read or seen that book and I can't tell if it's
> misleading or just wrong

Are you joking?

Guy Steele Jr. was on the ANSI C X3J11 standards committee.

Guy Steele Jr.
http://en.wikipedia.org/wiki/Guy_Steele

Samuel P. Harbison
http://www.harbison.org/sph3/

Another former X3J11 committee member posts to comp.std.c.
His name is Douglas A. Gwyn.  He developed code for U.S.
Army's BRL-Unix.

Another former X3J11 committee member is P.J. Plauger of
Whitesmiths and Dinkumware and author of "The Standard C Library".

Another former X3J11 committee member is Larry Rosler.  He's
occasionally interviewed on C and C++ since he was an early
teacher of C and also knew Ritchie and Stroustrup.

P.J. Plauger
http://en.wikipedia.org/wiki/P._J._Plauger

> I've mentioned the relevant parts of the C99 standard.

No offense, but it's far more likely "you're not interpreting
[C99 standard] correctly," or even more likely, completely.
There are outcomes not explicitly stated in the standards.
E.g., if the C standards committee hadn't remove implicit
ints from C99, would you have even known there was a conflict
between implicit ints and typedef's in C89 by reading C89?
No, you wouldn't.  Yet, it's right there in the language grammar.

> It's a click away:
>
> http://www.open-std.org/jtc1/sc22/wg14/www/standards.html:

I have copies of C89, C99, and numerous drafts as pdfs.
I also have the C Rationale as a pdfs.  I also have C89 and
the C Rationale in book form from before the modern pdf era.

> The conversion mentioned in the quoted bit of the standard is the
> conversion of the captured numerical or pointer type (long if %ld,
> int if %d, void* if %p, etc) into printable text.

That would print the pointer, not the integer value.

> It's not the intermediate conversion of int to long,

That's needed to pass the parameter correctly.

> While I have written lots of poor code, I have also learned from some of
> those mistakes. I'm avoiding UB by not triggering it, not by artful
> camouflaging of questionable actions. UB is not a fate. You can choose
> to have it or not.
>

You seem to be conflating UB with having a coding error.  They're not the
same.  UB doesn't mean you have a coding error.  It only means the
C specification doesn't define what happens in that situation.  The  
compiler
can and does define what happens.  In many cases, the behavior is well
defined and widespread, e.g., use of pointers to access specific memory
locations, or equivalence of pointers between various types the C  
specification
doesn't require.  The original C89 specification was derived from actual
implementations of C on various computing platforms.  So, there are a
number of very well defined situations for all microprocessors which
nearly all C compilers implement that is officially UB.  Why?  Because
mainframes and certain CPUs don't support that functionality.
Microprocessors since 1974 do though.  So, it's basically never going to
cause an issue to use such functionality in the majority of C code since
the machines where UB exists are obsolete and dead.

> http://stackoverflow.com/questions/20343190/evaluating-accessing-a-structure.
>

 From the code, I'm not sure if it's an issue with 'volatile', optimization,
or re-use of the variable name 's' which could be an issue.  What happens
if you change one 's' to a 't'?

Personally, I only use volatile in *two* very specific situations
which _generally_ work...

> I avoided UB in those shifts technically and legally. Period.
>

You're not responsible if the assembly is incorrect and the code
fails?  Tell your employer *that* someday.  I dare you!  ;-)

> You have your own ideas how things should work

That's true.  Mine matches the reality of actual C compilers,
as well as the numerous books on C that I learned from, as well
as the primary book on C that I still use: Harbison and Steele's
"C: A Reference Manual", 3rd. ed. 1991.

> and they don't quite match those expressed in the language standard.

That's true too.  No C compiler complies with the language
standard entirely.  It can't.  There are too many contradictions
and too many unecessary limitations.  C99 obfuscates the actual
meanings even further, compounding the issue.

My understanding of C corresponds very well with at least one
former member of the X3J11 committee since I've read what he
has written and also discussed C with him.  It's more likely my
view on C corresponds to three X3J11 committee members since
I have and use books by the other two.

> You may be finding holes and contradictions in my statements
> all day long until you've learned enough of standard C.

I started programming in 1981.  I've been programming in C since 1991.

I have a library of C books, probably thirty or more of them (packed
away now), all of which I read and learned from once.  Of all those
books, only two are of any value, IMO.  Harbison and Steele's
"C: A Reference Manual" and P.J. Plauger's "The Standard C Library".

>> Let's start over.
>
> Let's not. You have something to read and think through.

No, I don't.

I think you have much to learn yet, and eventually forget the
unimportant stuff, which you currently think is important.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1172

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-15 02:14 -0800
Message-ID<00697226-e8a2-4831-b369-ad345923d65c@googlegroups.com>
In reply to#1171
On Sunday, December 15, 2013 1:47:58 AM UTC-8, Rod Pemberton wrote:
> On Fri, 13 Dec 2013 06:32:43 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:

[snip]

> > Let's not. You have something to read and think through.
> 
> No, I don't.
> 
> I think you have much to learn yet, and eventually forget the
> unimportant stuff, which you currently think is important.

Yeah, I've just learned that I should not discuss C with you because it's pointless. 

It's pointless because you're turning your deaf ear or blind eye (speaking metaphorically, of course) to the provided references in the language standard and to the garbage that your compiler's printf() prints when you pass it a 16-bit int into %lx and to the warnings generated by gcc about int/long not matching, to all the signs, hints and explicit statements that you're invoking UB in this way. And so on.

If you are unwilling to even explore this topic further, there's nothing we can discuss here.

Alex

[toc] | [prev] | [next] | [standalone]


#1173

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-15 13:21 -0500
Message-ID<op.w75byrjq5zc71u@localhost>
In reply to#1172
On Sun, 15 Dec 2013 05:14:54 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Sunday, December 15, 2013 1:47:58 AM UTC-8, Rod Pemberton wrote:
>> On Fri, 13 Dec 2013 06:32:43 -0500, Alexei A. Frounze
>> <...@gmail.com> wrote:

[snip]

>> > Let's not. You have something to read and think through.
>>
>> No, I don't.
>>
>> I think you have much to learn yet, and eventually forget the
>> unimportant stuff, which you currently think is important.
>
> Yeah, I've just learned that I should not discuss C with you because  
> it's pointless.

Now, I'm wondering if you remember any of the numerous past conversations
we've had...  Think alt.os.development.

You mean you didn't learn that from the numerous conversations we've
had going back to 2006?  Do you remember the one on DSP byte ordering?
You left Usenet for something like four years after that one...

I was having a 100% serious conversation.  So, I thought when you said
you could avoid UB you were just displaying ignorance, but now I realize
you were just "trolling" me.  Anyone who knows UB as well, as you  
supposedly
do, knows that UB can't be avoided:

1) startup of a C program is UB
2) exiting a C program is UB is the application doesn't explicitly close
files and reclaim memory
3) memory mapped devices can't be programmed without using a pointer to
integer conversion which is UB
4) setjmp() and longjmp() have UB
5) system() (and spawn()) have UB
etc.

Instead, you focus on UB which doesn't result in errant code, and
generally can't be avoided.

> It's pointless because you're turning your deaf ear or blind eye  
> (speaking metaphorically, of course) to the provided references in the  
> language standard and to the garbage that your compiler's printf()  
> prints when you pass it a 16-bit int into %lx and to the warnings  
> generated by gcc about int/long not matching, to all the signs, hints  
> and explicit statements that you're invoking UB in this way. And so on.
>

If I wanted to, I could say nearly the same thing about you, see:

It's pointless because you're turning your deaf ear or a blind eye to
valid and widely implemented C concepts of implicit casting and UB
resulting from compounding of operations.

> If you are unwilling to even explore this topic further, there's
> nothing we can discuss here.

Are you disappearing from Usenet for four years like the last time
you became angry with me?  Seriously, I hope not.  Or, was that
because you got caught posting out to Usenet via MS' corporate
network?  Or, was that Bill Gates personal home IP?


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1111

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-04 23:39 -0800
Message-ID<cafe7f44-48a0-4c97-9b74-b6b7116481c9@googlegroups.com>
In reply to#1106
On Wednesday, December 4, 2013 1:31:56 AM UTC-8, Rod Pemberton wrote:
> On Wed, 04 Dec 2013 02:58:30 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> 
> > On Tuesday, December 3, 2013 1:26:27 PM UTC-8, Rod Pemberton wrote:
> >> On Tue, 03 Dec 2013 12:23:45 -0500, Rod Pemberton
> >> <dont_use_email@xnohavenotit.cnm> wrote:
> >> > On Sun, 01 Dec 2013 11:15:37 -0500, Alexei A. Frounze
> 
> >> For 16-bit OpenWatcom:
> >>
> >>    wcl/l=dos -wx smlrc.c
> >>    smlrc -seg16 smlrc.c smlrc.asm
> >>    nasm -f obj -o smlrc.obj smlrc.asm
> >>
> >> That seems to work.  I'm not sure what the wlink
> >> line is for OpenWatcom.  I've been trying this
> >> which fails to link:
> >>
> >>    wlink system dos name smlrc.exe file smlrc.obj
> >
> > It looks like it won't work like that. smlrc-generated code expects  
> > what's known as __cdecl in OW. OTOH, OW's stnadard libary functions  
> > appear to be __watcall. The difference is in underscoring of symbols
> > and in how parameters are passed.
> ...
> 
> > Would you like to research this further?
> 
> Would you? ;-)

Maybe. And I've already done that, some years ago. But since supporting OW's assembler and/or linker is not a priority for me at the moment, I thought maybe you would. Besides, you know how it works in projects like this, if you want something, you volunteer to do it. :)

> Being able to link with OW would be nice.  It's not a necessity in
> my book.  It just means whomever uses OpenWatcom needs an additional
> linker from another source to link .obj's.  If a .com can be made
> then no linker would be needed.  Or, perhaps a raw (no linker)
> DOS exe via NASM as a .bin format could be generated...  IIRC,
> Steve did something like this.  Off hand, I don't recall if you did
> so too in your DOS code.  Since a single .obj file is produced,
> then linking might not be required.  I'm thinking along the lines
> of DJGPP's coff2exe and exe2coff utilities, but these would be
> for obj's.  MS has exe2bin.  Is there a bin2exe?  If not, is
> that a restriction of the OMF/OBJ object?  I.e., why can DJGPP
> style COFF's go from COFF to EXE, but DOS style OMF/OBJ can't go
>  from OMF/OBJ to EXE?  If there is a sound reason, then perhaps a
> COFF object or ELF object could be used instead of OMF/OBJ? ...
> 
> http://www.delorie.com/djgpp/doc/utils/utils_toc.html
> 
> Didn't NASM have a way to link? I know NASM has it's own object
> format called RDOFF.

I've given it a thought and it looks like I might be able to get 16-bit .EXEs straight out of NASM, without using any additional linker. Consider this example (works with newer NASM, not sure if it's going to work with the old one you're keeping around) for the -f bin option:

---------8<---------
bits 16
org 0

exe_header_start:
    db  "MZ"
    dw  0 ; last 512-byte page size
    dw  256 ; exe image size in 512-byte pages, including header
    dw  0 ; number of entries in relocation table
    dw  2 ; size of header in 16-byte paragraphs AKA (relative) image base (will be added to CS and SS)
    dw  0 ; min RAM needed in paragraphs after the image
    dw  0 ; max RAM needed in paragraphs after the image
    dw  4096-2 ; initial SS
    dw  0xFFFE ; initial SP, can't be 0 (if it is, td.exe and debug.exe screw up SS (trying to somehow "correct" the stack?))
    dw  0 ; checksum
    dw  code_start; initial IP
    dw  -2 ; initial CS
    dw  exe_header_relo ; file offset of 1st relocation entry
    dw  0 ; overlay number
exe_header_relo:
    dw  0 ; fake relocation entry, just for header padding to a multiple of 16 bytes
    dw  0

; code start: file offset = 0x20, ip = file offset & 0xFFFF = 0x20
code_start:

    mov ax, ss
    mov ds, ax ; ds=es=ss:0 to point to cs:0 + 64K
    mov es, ax

    mov dx, msg
    mov ah, 9
    int 0x21

    mov ax, 0x4c00
    int 0x21

    mov ax, msg
    mov bx, msg
    mov ax, [msg]
    mov bx, [msg] ; warning: word data exceeds bounds
    mov dx, [msg] ; warning: word data exceeds bounds
    mov al, [msg]

    times (0x10000 - ($ - exe_header_start)) nop

; data start: file offset = 0x10000, ip = file offset & 0xFFFF = 0
data_start:

msg db "HELLO WORLD FROM A FAKE .EXE!",13,10,"$"

    times (0x10000 - ($ - data_start)) db 0

; file end: file offset = 0x20000
---------8<---------

I get a nice 128K .EXE out of the above with NASM and it works in DosBox.

All I need to do is to make my compiler first emit all code and then emit all data or emit them into separate files and then concatenate them, so in the end I can make two 64K segments in such a "fake" .EXE, one with code, another with data and stack.

Alex

[toc] | [prev] | [next] | [standalone]


#1114

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-05 06:38 -0500
Message-ID<op.w7mal4f55zc71u@localhost>
In reply to#1111
On Thu, 05 Dec 2013 02:39:14 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:

> I've given it a thought and it looks like I might be able to get 16-bit  
> .EXEs straight out of NASM, without using any additional linker.  
> Consider this example (works with newer NASM, not sure if it's going to  
> work with the old one you're keeping around) for the -f bin option:
>
> ---------8<---------
> [snip]
> ---------8<---------
>
> I get a nice 128K .EXE out of the above with NASM and it works in DosBox.
>
> All I need to do is to make my compiler first emit all code and then  
> emit all data or emit them into separate files and then concatenate  
> them, so in the end I can make two 64K segments in such a "fake" .EXE,  
> one with code, another with data and stack.
>

I didn't check the size, but it works here.  I checked in real-mode
MS-DOS v7.10 and Windows SE DOS console.  I don't have 6.22.

Most versions of NASM I have generated the same error on lines
40 and 41 as listed in your source:

   ; warning: word data exceeds bounds

0.98.39 generated this warning on lines 39 and 42 instead.
NASM versions 2.06rc8 and 2.08rc9 don't generate any warnings.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1117

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-05 21:40 -0800
Message-ID<f80d6090-236b-4e59-990f-7231dd5940e0@googlegroups.com>
In reply to#1114
On Thursday, December 5, 2013 3:38:18 AM UTC-8, Rod Pemberton wrote:
> On Thu, 05 Dec 2013 02:39:14 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> 
> > I've given it a thought and it looks like I might be able to get 16-bit  
> > .EXEs straight out of NASM, without using any additional linker.  
> > Consider this example (works with newer NASM, not sure if it's going to  
> > work with the old one you're keeping around) for the -f bin option:
> >
> > ---------8<---------
> > [snip]
> > ---------8<---------
> >
> > I get a nice 128K .EXE out of the above with NASM and it works in DosBox.
> >
> > All I need to do is to make my compiler first emit all code and then  
> > emit all data or emit them into separate files and then concatenate  
> > them, so in the end I can make two 64K segments in such a "fake" .EXE,  
> > one with code, another with data and stack.
> >
> 
> I didn't check the size, but it works here.  I checked in real-mode
> MS-DOS v7.10 and Windows SE DOS console.  I don't have 6.22.
> 
> Most versions of NASM I have generated the same error on lines
> 40 and 41 as listed in your source:
> 
>    ; warning: word data exceeds bounds

Are those errors or warnings?

> 0.98.39 generated this warning on lines 39 and 42 instead.
> NASM versions 2.06rc8 and 2.08rc9 don't generate any warnings.

The hack may be problematic if NASM actually errors out instead of simply warning us about something smelling fishy.

Alex

[toc] | [prev] | [next] | [standalone]


#1118

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-06 04:19 -0800
Message-ID<bf6a7a9d-7c5f-41ec-9078-24b74569a826@googlegroups.com>
In reply to#1114
On Thursday, December 5, 2013 3:38:18 AM UTC-8, Rod Pemberton wrote:
> On Thu, 05 Dec 2013 02:39:14 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> 
> > I've given it a thought and it looks like I might be able to get 16-bit  
> > .EXEs straight out of NASM, without using any additional linker.  
> > Consider this example (works with newer NASM, not sure if it's going to  
> > work with the old one you're keeping around) for the -f bin option:
> >
> > ---------8<---------
> > [snip]
> > ---------8<---------
> >
> > I get a nice 128K .EXE out of the above with NASM and it works in DosBox.
> >
> > All I need to do is to make my compiler first emit all code and then  
> > emit all data or emit them into separate files and then concatenate  
> > them, so in the end I can make two 64K segments in such a "fake" .EXE,  
> > one with code, another with data and stack.
> >
> 
> I didn't check the size, but it works here.  I checked in real-mode
> MS-DOS v7.10 and Windows SE DOS console.  I don't have 6.22.
> 
> Most versions of NASM I have generated the same error on lines
> 40 and 41 as listed in your source:
> 
>    ; warning: word data exceeds bounds
> 
> 0.98.39 generated this warning on lines 39 and 42 instead.
> NASM versions 2.06rc8 and 2.08rc9 don't generate any warnings.

Actually, it may be even easier. The following snipped assembles with NASM and works as well, which means minimal changes are required to support linker-less generation of 16-bit DOS .EXEs with just smlrc and NASM:

---------8<---------
bits 16
org 0

section .code

exe_header_start:
    db  "MZ"
    dw  0 ; last 512-byte page size
    dw  256 ; exe image size in 512-byte pages, including header
    dw  0 ; number of entries in relocation table
    dw  2 ; size of header in 16-byte paragraphs AKA (relative) image base (will be added to CS and SS)
    dw  0 ; min RAM needed in paragraphs after the image
    dw  0 ; max RAM needed in paragraphs after the image
    dw  4096-2 ; initial SS
    dw  0xFFFE ; initial SP, can't be 0 (if it is, td.exe and debug.exe screw up SS (trying to somehow "correct" the stack?))
    dw  0 ; checksum
    dw  code_start; initial IP
    dw  -2 ; initial CS
    dw  exe_header_relo ; file offset of 1st relocation entry
    dw  0 ; overlay number
exe_header_relo:
    dw  0 ; fake relocation entry, just for header padding to a multiple of 16 bytes
    dw  0

; code start: file offset = 0x20, ip = file offset & 0xFFFF = 0x20
code_start:

    mov ax, ss
    mov ds, ax ; ds=es=ss:0 to point to cs:0 + 64K
    mov es, ax

    mov dx, msg
    mov ah, 9
    int 0x21

    call f

    mov ax, 0x4c00
    int 0x21

section .data

; data start: file offset = 0x10000, ip = file offset & 0xFFFF = 0
data_start:

msg db "HELLO WORLD FROM A FAKE .EXE!",13,10,"$"

section .code

f:
    mov dx, msg2
    mov ah, 9
    int 0x21
    ret

    mov ax, msg
    mov bx, msg
    mov ax, [msg]
    mov bx, [msg]
    mov dx, [msg]
    mov al, [msg]

section .data

msg2 db "BLAH!",13,10,"$"

section .code

    times (0x10000 - ($ - exe_header_start)) nop

section .data

    times (0x10000 - ($ - data_start)) db 0

; file end: file offset = 0x20000
---------8<---------

Also, I'm not getting any warnings from NASM (v 2.10 Mar 12 2012).

Alex

[toc] | [prev] | [next] | [standalone]


#1120

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-06 13:40 -0500
Message-ID<op.w7ootcf35zc71u@localhost>
In reply to#1118
On Fri, 06 Dec 2013 07:19:00 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Thursday, December 5, 2013 3:38:18 AM UTC-8, Rod Pemberton wrote:
>> On Thu, 05 Dec 2013 02:39:14 -0500, Alexei A. Frounze
>> <...@gmail.com> wrote:

> Actually, it may be even easier. The following snipped assembles with  
> NASM and works as well, which means minimal changes are required to  
> support linker-less generation of 16-bit DOS .EXEs with just smlrc and  
> NASM:
>
> [...]
>
> Also, I'm not getting any warnings from NASM (v 2.10 Mar 12 2012).
>

Works.  No errors.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1170

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-14 21:07 -0800
Message-ID<549b290b-a7f3-4cc5-a0a7-d7dd3d2bb1d6@googlegroups.com>
In reply to#1118
On Friday, December 6, 2013 4:19:00 AM UTC-8, Alexei A. Frounze wrote:
> On Thursday, December 5, 2013 3:38:18 AM UTC-8, Rod Pemberton wrote:
> > On Thu, 05 Dec 2013 02:39:14 -0500, Alexei A. Frounze  
> > <...@gmail.com> wrote:
> > 
> > > I've given it a thought and it looks like I might be able to get 16-bit  
> > > .EXEs straight out of NASM, without using any additional linker.  
> > > Consider this example (works with newer NASM, not sure if it's going to  
> > > work with the old one you're keeping around) for the -f bin option:
> > >
> > > ---------8<---------
> > > [snip]
> > > ---------8<---------
> > >
> > > I get a nice 128K .EXE out of the above with NASM and it works in DosBox.
> > >
> > > All I need to do is to make my compiler first emit all code and then  
> > > emit all data or emit them into separate files and then concatenate  
> > > them, so in the end I can make two 64K segments in such a "fake" .EXE,  
> > > one with code, another with data and stack.
> > >
> > 
> > I didn't check the size, but it works here.  I checked in real-mode
> > MS-DOS v7.10 and Windows SE DOS console.  I don't have 6.22.
> > 
> > Most versions of NASM I have generated the same error on lines
> > 40 and 41 as listed in your source:
> > 
> >    ; warning: word data exceeds bounds
> > 
> > 0.98.39 generated this warning on lines 39 and 42 instead.
> > NASM versions 2.06rc8 and 2.08rc9 don't generate any warnings.
> 
> Actually, it may be even easier. The following snipped assembles
> with NASM and works as well, which means minimal changes are
> required to support linker-less generation of 16-bit DOS .EXEs
> with just smlrc and NASM:
> 
> ---------8<---------

[snip]

> ---------8<---------

I've just updated the compiler on github to support self-compilation into a 16-bit DOS .EXE using this method.

Now you can (re)compile Smaller C for DOS in any environment, where you have NASM, and for which you can compile Smaller C with a regular C compiler just once (buggy Open Watcom C/C++ 1.3 doesn't count:).

This is how you (re)compile it:
smlrc -seg16 -no-externs lbdos.c lbdos.asm
smlrc -seg16 -no-externs -label 1001 smlrc.c smlrc.asm
nasm -f bin smlrcdos.asm -o smlrcdos.exe

Alex

[toc] | [prev] | [next] | [standalone]


#1103

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-03 22:22 -0800
Message-ID<9e4bfa79-327e-4599-a899-4b9810fa4c65@googlegroups.com>
In reply to#1098
On Tuesday, December 3, 2013 9:23:45 AM UTC-8, Rod Pemberton wrote:

[snip]

> Auric mentioned your included functions.  It's
> possible to reduce the quantity used, or implement them
> internally for bootstrapping.  That may have a negative
> impact on speed though.
> 
> exit -> suggested to rework to use return()

That's a bad suggestion for a recursive program with many places where error() is called from (error() in turn calls exit() and thus avoids checking any extra error codes returned from functions all the way 20-levels-deep from main()).

> atoi -> obsolete

Good enough here.

> -usually strtod(), strtol(), strtoul() are recommended

Don't care here.

> -sscanf() or fscanf() can also be used

What's faster and easier to implement and test, atoi() or *scanf()? :)

> strlen
> strcpy
> strchr
> strcmp
> strncmp -> can usually be converted to strcmp(),
> -if you terminate '\0' the strings at 'n'
> memmove -> usually can be changed to use memcpy()
> memcpy
> memset
> 
> Most of the str*() functions and mem*() functions
> are easy to implement.  Most are just loops with one
> simple operation and/or a check for the '\0' null
> terminator.  You could also used reduced versions
> with less than normal functionality, e.g., for strcmp().

Re "easy to implement" and "reduced versions": it is precisely what I did. My *printf() does not support everything, just enough for the most common scenarios.

> Comparing for equal or not equal is much simpler than
> producing -1/false, 0, and 1/true.  If you use files
> instead of memory, you can eliminate most of these also.

Don't care.

> isspace
> isdigit
> isalpha
> isalnum
> 
> If you desire, you can replace is*() functions by using
> strchr() or strrchr() with a string, or a 128 or 256 char
> arrays as a lookup table, or even use character based
> procedures, e.g., with a switch(c) or even if()'s.

I know that as well, thanks.

> fopen -> keep it.
> fclose -> keep it.
> putchar -> can use fputc to stdout
> fputc -> keep it.
> fgetc -> keep it.
> puts -> can use fprintf() to stdout with '\n'
> fputs -> can use fprintf() with '\n'
> sprintf -> sometimes can rework for printf() or fprintf()
> printf -> can use fprintf() to stdin or stdout
> fprintf -> keep it.

Thanks, but no thanks. Where do I get the macros stdin/stdout/stderr from if I don't include the standard stdio.h? :)

> It's hard to eliminate sprintf() and fprintf() when you
> need the formatting options.  They are also long and
> complicated routines.  puts(), of course, is faster.

At the moment I'm the least concerned about complete implementations of formatted I/O functions and about the speed of operation. There are some more pressing areas (cleaning up error checks/messages, implementing struct and casts).

What I'm more concerned about in terms of the standard library is the common functions operating with long ints, like ftell() and fseek(). There's no support for 32-bit integers in 16-bit code and yet long ints can't be shorter than 32 bits and DOS file sizes are 32-bit. Oops! Looks like as a workaround I'll need to implement _ftell() and _fseek() taking and returning 32-bit values as pairs of 16-bit ones. Support for longs isn't coming soon. It'll require reworking of code in a fair number of places (and the most problematic of them is the code generator) and the code and data size of the compiler will significantly grow as a result of supporting longs, likely making it no longer self-compilable in 16-bit. I might need to start using far pointers to address more than 64K of the code and then I'd use far pointers for data as well and to support far pointers, I'll have to support far pointers, meaning even bigger and complex code generator! :) OK, this isn't the only possible option, but the long problem needs some research.

> However, with some work, you can usually eliminate all
> internal code that uses memory by preferencing file
> I/O functions.  So, you can eliminate sprintf(), sscanf(),
> in favor of fprintf(), fscanf(), etc.

Internally, my *printf()'s share most of the same code. I don't gain much (of what I'd care now) by preferring or avoiding one subset or the other. 

> Some compilers
> implement tmpfile() as a memory file.  That can provide
> a large speed boost.  You effectively get random access
> memory - as a file - with convenient to use functions.

I might use that one at some point, especially in 16-bit code.

> vsprintf()
> vprintf()
> vfprintf()
> 
> I've not use v*printf() functions much.  So, it should be
> possible to eliminate these by using other printf functions,
> depending on what you're using them for.  If you're using
> them to handle variadic argument lists, then you could use
> a stack.

Well, the thing is... Take a look at the following declaration in the compiler:

void error(char* format, ...);

How do you think error("Something horrible happened at %d!\n", 404); gets its arguments out? It uses va_list and vprintf(). And vprintf() is underneath printf() anyway.

I do not feel like disfavoring anything from the used subset of the standard library functions.

> > Testers/test-drivers? ;)
> 
> That's coming up next...  I was going to try compiling it
> with DJGPP, perhaps OpenWatcom.  I'll probably need to
> review Aurics post to figure out how to link it.
> 
> It has C++ style comments, so I'm wondering about warnings:
> 
> gcc -Wall -ansi -pedantic
> wcl/l=dos -wx
> wcl386/l=dos4g -wx

You don't need to guess what's not implemented entirely as per C89. I can tell you many things myself. :) But I think only rarely would someone object to the C++ style comment support as it is enabled by default in most current compilers, C89 to C2011 and everything in between. Guess what? Implicit function declaration is also commonly supported, although banned in C99. -std=c99 isn't enough for gcc to disallow it, you also need to add -pedantic on top of it.

Don't care.

You can, btw, do this:

int foo(int bar, int/*(: unused unnamed argument :)*/)
{
  return foo(bar, bar);
}

and this (in file scope):

int a = (1, 2);

neither of which is allowed by C89. :)

Alex

[toc] | [prev] | [next] | [standalone]


#1105

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-04 04:30 -0500
Message-ID<op.w7j901yn5zc71u@localhost>
In reply to#1103
On Wed, 04 Dec 2013 01:22:21 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Tuesday, December 3, 2013 9:23:45 AM UTC-8, Rod Pemberton wrote:

> At the moment I'm the least concerned about complete implementations of  
> formatted I/O functions and about the speed of operation. There are some  
> more pressing areas (cleaning up error checks/messages, implementing  
> struct and casts).

Are you interested in making the program's code ANSI C?

> What I'm more concerned about in terms of the standard library is the  
> common functions operating with long ints, like ftell() and fseek().  
> There's no support for 32-bit integers in 16-bit code and yet long ints  
> can't be shorter than 32 bits and DOS file sizes are 32-bit.

I would think TCC and Borland C would be 16-bit too.

Except for a few pointer construction macro's and calls to DOS functions,
I don't recall much difference between programs compiled by OW as both
16-bit and 32-bit.  But, I haven't done much in-depth with OW in a few
years.

> Looks like as a workaround I'll need to implement _ftell() and _fseek()
> taking and returning 32-bit values as pairs of 16-bit ones.

:-)

There is a reason DJGPP doesn't support 16-bit.
There are reasons LCC stopped supporting 16-bit.

> Support for longs isn't coming soon.

Isn't that the same issue as above?

> It'll require reworking of code in a fair number of places (and the most  
> problematic of them is the code generator) and the code and data size of  
> the compiler will significantly grow as a result of supporting longs,  
> likely making it no longer self-compilable in 16-bit.

Don't they use different integer sizes in C for different modes on x86?

> I might need to start using far pointers to address more than 64K of the  
> code and then I'd use far pointers for data as well and to support far  
> pointers, I'll have to support far pointers, meaning even bigger and  
> complex code generator! :) OK, this isn't the only possible option, but  
> the long problem needs some research.
>

IIRC, far and near pointers and some other incompatibilities with the
C specifications was why LCC stopped supporting 16-bit.  v3.5 and v3.6
supported DOS.  v4.1 and v4.2 stopped.  I never got around to it, but
it looked like the changes were very minor, i.e., like maybe v4.1 or
v4.2 had a good chance of being made to support 16-bits.

>> However, with some work, you can usually eliminate all
>> internal code that uses memory by preferencing file
>> I/O functions.  So, you can eliminate sprintf(), sscanf(),
>> in favor of fprintf(), fscanf(), etc.
>
> Internally, my *printf()'s share most of the same code.

I was looking at your externally declared functions.
I thought you were using an available external library
for them.

>> vsprintf()
>> vprintf()
>> vfprintf()
>>
>> I've not use v*printf() functions much.  So, it should be
>> possible to eliminate these by using other printf functions,
>> depending on what you're using them for.  If you're using
>> them to handle variadic argument lists, then you could use
>> a stack.
>
> Well, the thing is... Take a look at the following declaration
> in the compiler:
>
> void error(char* format, ...);
>
> How do you think error("Something horrible happened at %d!\n", 404);  
> gets its arguments out? It uses va_list and vprintf().

Well, you don't have to do it that way.  You could use printf()
directly.  You could use fprintf() directly.  You could use
sprintf() to a buffer and then a puts().  You could use one printf()
for the string and another for the integer, e.g., string looked up
in table.  You could pass only one string to a custom error
function after formatting.

> And vprintf() is underneath printf() anyway.

In yours? OpenWatcom? DJGPP? ... DJGPP uses a central print routine
which is mostly printf().  I'm not sure about OpenWatcom.

> I do not feel like disfavoring anything from the used subset
> of the standard library functions.

It was just suggestions in case you wanted to reduce the quantity
to make bootstrapping easier.  But, you're bootingstrapping with
modern C compilers with complete C libraries, so none of that is
probably even needed.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1110

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-04 23:21 -0800
Message-ID<0e1f8d5c-505c-490b-a378-cb755658910a@googlegroups.com>
In reply to#1105
On Wednesday, December 4, 2013 1:30:27 AM UTC-8, Rod Pemberton wrote:
> On Wed, 04 Dec 2013 01:22:21 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> > On Tuesday, December 3, 2013 9:23:45 AM UTC-8, Rod Pemberton wrote:
> 
> > At the moment I'm the least concerned about complete implementations of  
> > formatted I/O functions and about the speed of operation. There are some  
> > more pressing areas (cleaning up error checks/messages, implementing  
> > struct and casts).
> 
> Are you interested in making the program's code ANSI C?

Which program's, how and why?

> > What I'm more concerned about in terms of the standard library is the  
> > common functions operating with long ints, like ftell() and fseek().  
> > There's no support for 32-bit integers in 16-bit code and yet long ints  
> > can't be shorter than 32 bits and DOS file sizes are 32-bit.
> 
> I would think TCC and Borland C would be 16-bit too.

And?

> Except for a few pointer construction macro's and calls to DOS functions,
> I don't recall much difference between programs compiled by OW as both
> 16-bit and 32-bit.  But, I haven't done much in-depth with OW in a few
> years.

There are some differences in inputs and outputs to and from interrupt-invoking functions because the registers have different sizes and meanings.

The standard subset of the library is, of course, the same in both cases.

> > Looks like as a workaround I'll need to implement _ftell() and _fseek()
> > taking and returning 32-bit values as pairs of 16-bit ones.
> 
> :-)
> 
> There is a reason DJGPP doesn't support 16-bit.
> There are reasons LCC stopped supporting 16-bit.

There isn't much interest/incentive/money. :)

> > Support for longs isn't coming soon.
> 
> Isn't that the same issue as above?

What do you mean by that?

> > It'll require reworking of code in a fair number of places (and the most  
> > problematic of them is the code generator) and the code and data size of  
> > the compiler will significantly grow as a result of supporting longs,  
> > likely making it no longer self-compilable in 16-bit.
> 
> Don't they use different integer sizes in C for different modes on x86?

Obviously, some do. Not sure if you had any specific *they* in mind. :)

> > I might need to start using far pointers to address more than 64K of the  
> > code and then I'd use far pointers for data as well and to support far  
> > pointers, I'll have to support far pointers, meaning even bigger and  
> > complex code generator! :) OK, this isn't the only possible option, but  
> > the long problem needs some research.
> >
> 
> IIRC, far and near pointers and some other incompatibilities with the
> C specifications was why LCC stopped supporting 16-bit.  v3.5 and v3.6
> supported DOS.  v4.1 and v4.2 stopped.  I never got around to it, but
> it looked like the changes were very minor, i.e., like maybe v4.1 or
> v4.2 had a good chance of being made to support 16-bits.

What happened? Did they not agree on whether the keyword far applies to (associates with) the pointer or to(with) the object/function pointed to?

> >> However, with some work, you can usually eliminate all
> >> internal code that uses memory by preferencing file
> >> I/O functions.  So, you can eliminate sprintf(), sscanf(),
> >> in favor of fprintf(), fscanf(), etc.
> >
> > Internally, my *printf()'s share most of the same code.
> 
> I was looking at your externally declared functions.
> 
> I thought you were using an available external library
> for them.

Yes and no. I have the needed functions to run smlrc on RetroBSD. For DOS they'd need minor changes. In the end, you still need to make a typical system call and pass a file handle/descriptor and a buffer pointer or something like that.

> >> vsprintf()
> >> vprintf()
> >> vfprintf()
> >>
> >> I've not use v*printf() functions much.  So, it should be
> >> possible to eliminate these by using other printf functions,
> >> depending on what you're using them for.  If you're using
> >> them to handle variadic argument lists, then you could use
> >> a stack.
> >
> > Well, the thing is... Take a look at the following declaration
> > in the compiler:
> >
> > void error(char* format, ...);
> >
> > How do you think error("Something horrible happened at %d!\n", 404);  
> > gets its arguments out? It uses va_list and vprintf().
> 
> Well, you don't have to do it that way.  You could use printf()
> directly.  You could use fprintf() directly.  You could use
> sprintf() to a buffer and then a puts().  You could use one printf()
> for the string and another for the integer, e.g., string looked up
> in table.  You could pass only one string to a custom error
> function after formatting.

Don't forget that error() != printf(), error() does more than a mere printf(). If I replace all calls to error() with 2 calls, first, to printf(), second, to something else, error2() or exit2(), then I lose memory on all those additional instructions doing additional calls.

> > And vprintf() is underneath printf() anyway.
> 
> In yours? OpenWatcom? DJGPP? ... DJGPP uses a central print routine
> which is mostly printf().  I'm not sure about OpenWatcom.

In my code, sorry.

> > I do not feel like disfavoring anything from the used subset
> > of the standard library functions.
> 
> It was just suggestions in case you wanted to reduce the quantity
> to make bootstrapping easier.  But, you're bootingstrapping with
> modern C compilers with complete C libraries, so none of that is
> probably even needed.

That's a pretty much solved problem. I've already mentioned that I have enough code to do that on RetroBSD and only relatively small changes/additions would need to be done to run the compiler in DOS (or Linux, for that matter) without depending on a standard library from another compiler.

Alex

[toc] | [prev] | [next] | [standalone]


#1112

From"Rod Pemberton" <dont_use_email@xnohavenotit.cnm>
Date2013-12-05 04:23 -0500
Message-ID<op.w7l4c1ui5zc71u@localhost>
In reply to#1110
On Thu, 05 Dec 2013 02:21:15 -0500, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:
> On Wednesday, December 4, 2013 1:30:27 AM UTC-8, Rod Pemberton wrote:
>> On Wed, 04 Dec 2013 01:22:21 -0500, Alexei A. Frounze
>> <...@gmail.com> wrote:
>> > On Tuesday, December 3, 2013 9:23:45 AM UTC-8, Rod Pemberton wrote:

>> > At the moment I'm the least concerned about [...]
>>
>> Are you interested in making the program's code ANSI C?
>
> Which program's, how and why?
>

Smaller C's source files.  Do you have any plans to make them be ANSI
or comply with GCC's '-ansi' flag?  (No.)  I think you stated in one
of the other posts that the code is C89 exempt C++ style comments.
I thought I saw many more errors and warnings with '-ansi' than that
would or should produce, but I'll double check.  DJGPP v2.03 uses
an old version of GCC. v2.04 has problems on real-mode DOS, but
is fine in dos console windows.  DJGPP project is updating GCC
for v2.04, but not v2.03, AFAIK.

>> > What I'm more concerned about in terms of the standard library is the
>> > common functions operating with long ints, like ftell() and fseek().
>> > There's no support for 32-bit integers in 16-bit code and yet long
>> > ints can't be shorter than 32 bits and DOS file sizes are 32-bit.
>>
>> I would think TCC and Borland C would be 16-bit too.
>
> And?

You could check to see if TCC and Borland C use long ints for ftell()
and fseek() too, or if they use some other solution, e.g., perhaps
they use near pointers, if you don't already know what they do.
I think (don't quote me) OpenWatcom, at least v1.3, used a few
non-standard returns.

>> > Support for longs isn't coming soon.
>>
>> Isn't that the same issue as above?
>
> What do you mean by that?

Isn't implementing longs on 16-bit the same issue as needing
to implement "32-bit values as pairs of 16-bit ones"?

Or, was that longs for 32-bit instead?

>> Don't they use different integer sizes in C for different modes on x86?
>
> Obviously, some do. Not sure if you had any specific *they* in mind. :)

No, just thinking that common sizes for 16-bit might fit better with
native 16-bit x86 assembly assembly size, than using 32-bit sizes.

>> > I might need to start using far pointers to address more than 64K of
>> > the code and then I'd use far pointers for data as well and to support
>> > far pointers, I'll have to support far pointers, meaning even bigger
>> > and complex code generator! :) OK, this isn't the only possible  
>> option,
>> > but the long problem needs some research.
>> >
>>
>> IIRC, far and near pointers and some other incompatibilities with the
>> C specifications was why LCC stopped supporting 16-bit.  v3.5 and v3.6
>> supported DOS.  v4.1 and v4.2 stopped.  I never got around to it, but
>> it looked like the changes were very minor, i.e., like maybe v4.1 or
>> v4.2 had a good chance of being made to support 16-bits.
>
> What happened?

Uh, I don't recall now...


See "4.html" for LCC v4.1's changes from v3.x:
http://zeq2.com/SVN/Source/Engine/tools/lcc/doc/4.html

I vaguely recall reading that there was a problem with pointer
compliance for C89 with DOS versions, so they dropped DOS support.

I also vaguely recall that there were differences in integers.

But, I'm not able to confirm either, at the moment.


I did have a file with a few notes I made:

lcc42 -no DOS support - incompatible w/3.x versions
lcc41 -lists changes from 3.x versions
lcc36 -last DOS version
lcc35 -last DOS version with included binaries

NASM 0.98.39 has support files for LCC for Linux in src/lcc
NASM 0.99.00 has support files for LCC v3.6

LCC 3.6 has x86-dos reference in bind.c
LCC 3.6 has two "dos" dir's in include\x86 and x86
LCC 4.1 and 4.2 don't

x86nasmw.md - NASM machine description for LCC 4.1 on Windows
ns32k093.zip - DJGPP changes for LCC v3.2
Quake3 Virtual Machine specification - modified LCC used for Q3VM

> Did they not agree on whether the keyword far applies to
> (associates with) the pointer or to(with) the object/function
> pointed to?

????

Was that MASM vs. NASM humor?
E.g., MASM dword/tword/qword/fword etc. vs. NASM brackets.

Or, was that C indirection or dereference humor?

>> >> vsprintf()
>> >> vprintf()
>> >> vfprintf()
>> >>
>> >> I've not use v*printf() functions much.  So, it should be
>> >> possible to eliminate these by using other printf functions,
>> >> depending on what you're using them for.  If you're using
>> >> them to handle variadic argument lists, then you could use
>> >> a stack.
>> >
>> > Well, the thing is... Take a look at the following declaration
>> > in the compiler:
>> >
>> > void error(char* format, ...);
>> >
>> > How do you think error("Something horrible happened at %d!\n", 404);
>> > gets its arguments out? It uses va_list and vprintf().
>>
>> Well, you don't have to do it that way.  You could use printf()
>> directly.  You could use fprintf() directly.  You could use
>> sprintf() to a buffer and then a puts().  You could use one printf()
>> for the string and another for the integer, e.g., string looked up
>> in table.  You could pass only one string to a custom error
>> function after formatting.
>
> Don't forget that error() != printf(), error() does more than a mere
> printf(). If I replace all calls to error() with 2 calls, first, to
> printf(), second, to something else, error2() or exit2(), then I lose
> memory on all those additional instructions doing additional calls.

It won't matter for exit().  If error() calls exit(), it won't matter
then either.  Application terminated.  OS recovers memory.  Yes? ...


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#1115

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2013-12-05 21:31 -0800
Message-ID<74deff20-17dd-4979-8003-2fa8c5ae1406@googlegroups.com>
In reply to#1112
On Thursday, December 5, 2013 1:23:15 AM UTC-8, Rod Pemberton wrote:
> On Thu, 05 Dec 2013 02:21:15 -0500, Alexei A. Frounze  
> <...@gmail.com> wrote:
> > On Wednesday, December 4, 2013 1:30:27 AM UTC-8, Rod Pemberton wrote:
> >> On Wed, 04 Dec 2013 01:22:21 -0500, Alexei A. Frounze
> >> <...@gmail.com> wrote:
> >> > On Tuesday, December 3, 2013 9:23:45 AM UTC-8, Rod Pemberton wrote:

[snip]

> You could check to see if TCC and Borland C use long ints for ftell()
> and fseek() too, or if they use some other solution, e.g., perhaps
> they use near pointers, if you don't already know what they do.
> I think (don't quote me) OpenWatcom, at least v1.3, used a few
> non-standard returns.

16-bit Turbo/Borland C/C++ and OW C/C++ use true longs, no hidden pointers to longs. The calling convention specifies that 32-bit values (longs and far pointers) are to be returned in DX:AX.

> >> > Support for longs isn't coming soon.
> >>
> >> Isn't that the same issue as above?
> >
> > What do you mean by that?
> 
> Isn't implementing longs on 16-bit the same issue as needing
> to implement "32-bit values as pairs of 16-bit ones"?
> 
> Or, was that longs for 32-bit instead?

The issue is the same. It's going to complicate and grow the code significantly, to the point of smlrc itself not fitting into a 64K code segment when compiled as 16-bit. I've already mentioned, smlrc compiles itself to about 60K of 16-bit code now (slightly less than that, actually, but there's a noticeable chunk of the standard library coming from Turbo C++ and OW if those compilers' standard libraries are used).

> >> Don't they use different integer sizes in C for different modes on x86?
> >
> > Obviously, some do. Not sure if you had any specific *they* in mind. :)
> 
> No, just thinking that common sizes for 16-bit might fit better with
> native 16-bit x86 assembly assembly size, than using 32-bit sizes.

Of course, 16-bit numbers and calculations are more suitable for 16-bit modes than 32-bit numbers and calculations.

If I use 32-bit registers for longs in 16-bit code, I'll either lose the ability to link with Turbo C++'s library (and OW's too) or I'll need to add special support for a second calling convention (by introducing some keyword like __cdecl and converting from DX:AX into EAX).

[snip]

> > Did they not agree on whether the keyword far applies to
> > (associates with) the pointer or to(with) the object/function
> > pointed to?
> 
> ????
> 
> Was that MASM vs. NASM humor?
> E.g., MASM dword/tword/qword/fword etc. vs. NASM brackets.
> 
> Or, was that C indirection or dereference humor?

It was const volatile C humor. :)

People often confuse

volatile int* p; /* or int volatile * p; */

with

int* volatile p;

far is the same w.r.t. association. I remember I've had problems with putting far in the right place.

[snip]

> >> > Well, the thing is... Take a look at the following declaration
> >> > in the compiler:
> >> >
> >> > void error(char* format, ...);
> >> >
> >> > How do you think error("Something horrible happened at %d!\n", 404);
> >> > gets its arguments out? It uses va_list and vprintf().
> >>
> >> Well, you don't have to do it that way.  You could use printf()
> >> directly.  You could use fprintf() directly.  You could use
> >> sprintf() to a buffer and then a puts().  You could use one printf()
> >> for the string and another for the integer, e.g., string looked up
> >> in table.  You could pass only one string to a custom error
> >> function after formatting.
> >
> > Don't forget that error() != printf(), error() does more than a mere
> > printf(). If I replace all calls to error() with 2 calls, first, to
> > printf(), second, to something else, error2() or exit2(), then I lose
> > memory on all those additional instructions doing additional calls.
> 
> It won't matter for exit().  If error() calls exit(), it won't matter
> then either.  Application terminated.  OS recovers memory.  Yes? ...

Yes. Even better, the OS won't need to reclaim any memory because of smlrc being too big to compile and ultimately because of not having an executable for smlrc. :) Every function call takes several bytes for the call instruction and parameter passing. Replacing one call with two calls will make and I'm already at ~60K of 16-bit self-compiled code of smlrc.

Alex

[toc] | [prev] | [next] | [standalone]


#1143

From"Auric__" <not.my.real@email.address>
Date2013-12-10 04:59 +0000
Message-ID<XnsA291DFC28EE40auricauricauricauric@78.46.70.116>
In reply to#1086
Alexei A. Frounze wrote:

> In case this is of interest to anyone, I'm working on a small and simple
> C compiler currently targeting x86 (and MIPS:), Smaller C. 
[snip]
> Testers/test-drivers? ;)

Just reporting, tcc & lcc both working good with the current revision of the 
source. Same notes re: building apps with smlrc as before. (tcc: no external 
linker; lcc: works good, but compile & link for win32 only.)

I should perhaps mention that lcc has a *nix version as well; I haven't tried 
it, nor do I really intend to.

-- 
There are many, many things that I will never understand...
and they're all women.

[toc] | [prev] | [next] | [standalone]


Page 2 of 5 — ← Prev page 1 [2] 3 4 5  Next page →

Back to top | Article view | comp.os.msdos.programmer


csiph-web