Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > alt.os.development > #8344 > unrolled thread

Re: Smaller C

Started by"Alexei A. Frounze" <alexfrunews@gmail.com>
First post2015-07-11 20:39 -0700
Last post2015-09-16 10:17 -0700
Articles 20 on this page of 88 — 7 participants

Back to article view | Back to alt.os.development

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-07-11 20:39 -0700
    Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-07-12 04:24 -0400
      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-07-12 03:17 -0700
        Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-07-13 03:28 -0400
      Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-07-12 19:12 +0100
        Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-07-12 17:09 -0700
        Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-07-13 11:10 +0100
          Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-07-13 23:13 -0400
            Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-07-14 09:08 +0100
              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-07-15 02:23 -0400
    Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-08-15 02:12 -0700
      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-06 15:54 -0700
        Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-07 12:19 -0700
          Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-07 13:30 -0700
            Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-09 19:38 -0700
              Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-09 19:49 -0700
              Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-09 20:58 -0700
              Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-10 01:45 -0700
                Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-10 10:45 -0700
                  Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-10 23:20 -0700
                    Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-11 09:26 +0200
                      Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-11 09:50 +0200
                      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-11 00:58 -0700
                        Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-11 17:38 -0400
                          Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-11 23:39 +0100
                            Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-11 19:39 -0400
                              Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-12 11:22 +0100
                          Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-12 02:21 -0700
                          Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-12 21:53 +0200
                            Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 03:37 -0400
                              Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 09:49 +0100
                                Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-13 12:58 +0200
                                Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 07:32 -0400
                                  Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 19:05 +0100
                                    Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-14 13:19 +0100
                                      Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-15 00:01 -0400
                                        Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-18 14:48 +0100
                                          Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-26 16:23 -0400
                                            Re: Smaller C James Harris <james.harris.1@gmail.com> - 2015-09-27 00:23 +0100
                                              Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-26 16:37 -0700
                                              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-27 10:17 -0400
                                                Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-27 10:24 -0400
                                                Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-27 10:46 -0400
                                                Re: Smaller C James Harris <james.harris.1@gmail.com> - 2016-01-16 16:14 +0000
                                                  Re: Smaller C Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-23 13:20 -0500
                                                    Re: Smaller C James Harris <james.harris.1@gmail.com> - 2016-01-27 13:52 +0000
                                              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-27 10:28 -0400
                                                Re: Smaller C James Harris <james.harris.1@gmail.com> - 2016-01-16 16:31 +0000
                                                  Re: Smaller C Rod Pemberton <NoHaveNotOne@bcczxcfre.cmm> - 2016-01-23 13:20 -0500
                                      Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-18 16:21 +0100
                                        Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-20 17:25 -0700
                              Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-13 12:43 +0200
                            Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 09:21 +0100
                              Re: Smaller C "wolfgang kern" <nowhere@never.at> - 2015-09-13 13:02 +0200
                                Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 19:09 +0100
                    Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-11 13:32 -0700
                      Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-11 18:51 -0400
                        Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-11 18:16 -0700
                          Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-12 02:36 -0700
                        Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-12 11:56 +0100
                          Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 03:45 -0400
                            Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 10:28 +0100
                              Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 02:53 -0700
                                Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 11:17 -0400
                                  Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 16:54 -0700
                                    Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 21:39 -0400
                                      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 20:31 -0700
                              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 09:03 -0400
                                Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-13 19:48 +0100
                                  Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 21:33 -0400
                                    Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-14 23:17 +0100
                                      Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-20 17:37 -0400
                                        Re: Smaller C "James Harris" <james.harris.1@gmail.com> - 2015-09-20 23:46 +0100
    Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-11 21:37 -0700
      Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-11 22:29 -0700
        Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-12 03:25 -0700
          Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-12 11:18 -0700
            Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-12 11:39 -0700
            Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 03:47 -0400
              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 11:31 -0400
            Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-13 02:00 -0700
              Re: Smaller C "Rod Pemberton" <boo@fasdfrewar.cdm> - 2015-09-13 11:31 -0400
              Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-13 09:15 -0700
      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-12 02:45 -0700
        Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-12 10:55 -0700
    Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-15 19:23 -0700
      Re: Smaller C "Alexei A. Frounze" <alexfrunews@gmail.com> - 2015-09-16 03:03 -0700
        Re: Smaller C "Benjamin David Lunt" <zfysz@fysnet.net> - 2015-09-16 10:17 -0700

Page 1 of 5  [1] 2 3 4 5  Next page →


#8344 — Re: Smaller C

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-07-11 20:39 -0700
SubjectRe: Smaller C
Message-ID<39da71f8-7e21-4fcb-84da-263d5c81059f@googlegroups.com>
One more improvement for Ben's 50+K executable. :)
Actually, I've had a couple of requests for this.

Anyway, the prologue/epilogue parts are now shorter.

This is what Smaller C used to generate:
_foo:
  push bp
  mov  bp, sp
  jmp  L1
L2:
  ...
  leave
  ret
L1:
  sub  sp, size_of_locals
  jmp  L2

This is what it generates now:
_foo:
  push bp
  mov  bp, sp
  sub  sp, size_of_locals
  ...
  leave
  ret

If there are no locals, there's no sub either. This shaves ~500 bytes
off a 64KB 16-bit code section.

The trick? fgetpos()/fsetpos() majic.
The price? Must always supply the output assembly file name to smlrc.

Alex

[toc] | [next] | [standalone]


#8345

From"Rod Pemberton" <boo@fasdfrewar.cdm>
Date2015-07-12 04:24 -0400
Message-ID<op.x1nizmxyyfako5@localhost>
In reply to#8344
On Sat, 11 Jul 2015 23:39:51 -0400, Alexei A. Frounze
<alexfrunews@gmail.com> wrote:

> [...]
>
> The trick? fgetpos()/fsetpos() majic.
> The price? Must always supply the output assembly file name to smlrc.

Oh, I'm sorry to disappoint you, but it's probably not magic Alex.

I ran into the same issue of not being able to output my program's
output directly to stdout in one of my projects some time ago.  The
program would optionally take a file to change the file pointer.
This problem was because of unknown forward references, and I
couldn't figure out an easy and clean and portable method to patch
up stdout.  I could've delayed writing to stdout, but that would've
"complicated" the program.

Did you know that stdin, stdout, stderr aren't static in Linux?  ...

FILE *out=stdout ; not valid with GLIBC, move assignment to main() etc

For most of these projects, I restricted myself to the C language and
only <stdio.h>.  So, I use file I/O with a tmpfile() "database" of saved
information.  printf() writes the formatted data.  scanf() reads the data.
I had to use hex formats, i.e., casts needed, for all of them since DOS
and Linux are incompatible in their implementations of various scanf()
formats.  DOS compilers wouldn't read in the same data they wrote for
some formats either.  ftell() and fseek() are used to move through the
data and save file position info within the tmpfile() where needed.
I think most of my programs only use one tmpfile(), but one uses three
or maybe four ...

I also use the same tmpfile() technique to put strings at the end of my
assembler output, which I think I mentioned to James or Ben recently.

A few weeks ago, I decided to code another BrainFuck (BF) parser to do
some actual transformations to BF code.  I already had one that worked
in memory using malloc(), but it's ad-hoc design wasn't working too well.
So, instead of rewriting it, I decided to implement a new one using only
tmpfile() ...  That was a PIA.  But, it works even though it's slow.
The biggest issue here was that there was no way to reset the EOF position
to the beginning of tmpfile() without closing the file.  This occasionally
left garbage at the end of file when the file pointer was re-used for new
temporary data.  The second biggest issue was having to use fseek() and
ftell() offsets as "memory" pointers.  Linked-list using file I/O anyone?

So, filespace can be used do everything or almost everything that
memory from malloc() can do, but not as easily.


Rod Pemberton

-- 
Tolerance and socialism attracts intolerance and terrorism.  See:
France, United Kingdom, Germany, Denmark, Belgium, Netherlands, ...
UK: Strong encryption, non-issue. GR: Eurozone currency, priority.

[toc] | [prev] | [next] | [standalone]


#8346

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-07-12 03:17 -0700
Message-ID<934a8e92-9474-4a0a-9014-42c706e30002@googlegroups.com>
In reply to#8345
On Sunday, July 12, 2015 at 1:24:36 AM UTC-7, Rod Pemberton wrote:
> On Sat, 11 Jul 2015 23:39:51 -0400, Alexei A. Frounze
> <...@gmail.com> wrote:
> 
> > [...]
> >
> > The trick? fgetpos()/fsetpos() majic.
> > The price? Must always supply the output assembly file name to smlrc.
> 
> Oh, I'm sorry to disappoint you, but it's probably not magic Alex.

:)

> I ran into the same issue of not being able to output my program's
> output directly to stdout in one of my projects some time ago.  The
> program would optionally take a file to change the file pointer.
> This problem was because of unknown forward references, and I
> couldn't figure out an easy and clean and portable method to patch
> up stdout. 

That's exactly the case. It is only when a function has been parsed
that I know how much needs to be subtracted from the stack pointer
to allocate space for automatic variables and it's too late at that
point. And the compiler doesn't buffer its output nor has a way to
perform two passes over the function source code (i.e. by implementing
some kind of dry run that would discard all the info but the cumulative
size of the variables and reset the input to the file and position,
where the function definition started).

> I could've delayed writing to stdout, but that would've
> "complicated" the program.

Yep.

> Did you know that stdin, stdout, stderr aren't static in Linux?  ...
> 
> FILE *out=stdout ; not valid with GLIBC, move assignment to main() etc

Um, it's probably how the library implements the standard streams.

In Smaller C, for example, you have this:
extern FILE *__stdin, *__stdout, *__stderr;
#define stdout __stdout

So, you can only do something like this:
static FILE** pf = &stdout;

Had stdout been implemented like this (exact copy from BSD 2.11):
extern struct _iobuf { ... } _iob[];
#define FILE        struct _iobuf
#define stdout      (&_iob[1])

you would be able to do
static FILE* f = stdout;

C99 doesn't tell whether these stdout's should be constant
expressions, so, it's up to the implementors.

I don't know if there are any other practical reasons for glibc to
not define those as constants (e.g. each thread having its own set
of standard streams, not only errno?).

Anyhow, you still have your standard UNIX file descriptors (0,1,2)
for the underlying I/O. While those can in fact be redirected, you
still know them by these 3 constant numbers.

> For most of these projects, I restricted myself to the C language and
> only <stdio.h>.  So, I use file I/O with a tmpfile() "database" of saved
> information.  printf() writes the formatted data.  scanf() reads the data.

Yep. You have shared this bit with us before.

> I had to use hex formats, i.e., casts needed, for all of them since DOS
> and Linux are incompatible in their implementations of various scanf()
> formats. 

How so? Vastly different versions of the language (pre-ASNI vs ANSI vs
C99?) supported by those DOS and Linux compilers? You used something
non-standard? You had to deal with different integer types sizes on the
two platforms (e.g. 16-bit ints/pointers vs 32-bit ints/pointers)?

> DOS compilers wouldn't read in the same data they wrote for
> some formats either. 

Or even library bugs? Or you didn't use the library in intended ways?

> ftell() and fseek() are used to move through the
> data and save file position info within the tmpfile() where needed.
> I think most of my programs only use one tmpfile(), but one uses three
> or maybe four ...
> 
> I also use the same tmpfile() technique to put strings at the end of my
> assembler output, which I think I mentioned to James or Ben recently.

Yep.

> A few weeks ago, I decided to code another BrainFuck (BF) parser to do
> some actual transformations to BF code.  I already had one that worked
> in memory using malloc(), but it's ad-hoc design wasn't working too well.
> So, instead of rewriting it, I decided to implement a new one using only
> tmpfile() ...  That was a PIA.  But, it works even though it's slow.
> The biggest issue here was that there was no way to reset the EOF position
> to the beginning of tmpfile() without closing the file.  This occasionally
> left garbage at the end of file when the file pointer was re-used for new
> temporary data. 

All practical systems that can delete a file or somehow simulate its
deletion and therefore support C's remove() have the power to truncate
a file or create a new empty one. Standard C having remove() but no
[f]truncate() is a bit of a bummer. The same goes for getting both
the temporary file and its name. A file size function would be
useful as well, even though we can usually get by without it.
I think the decision was to have the bare minimum of functionality
in the language and its library to make C squeezable into every hole.

OTOH, you could probably just store the effective file length
(total - garbage) in the first bytes of the file.

> The second biggest issue was having to use fseek() and
> ftell() offsets as "memory" pointers.  Linked-list using file I/O anyone?
> 
> So, filespace can be used do everything or almost everything that
> memory from malloc() can do, but not as easily.

Awkwardly and slow, yes.

Alex

[toc] | [prev] | [next] | [standalone]


#8352

From"Rod Pemberton" <boo@fasdfrewar.cdm>
Date2015-07-13 03:28 -0400
Message-ID<op.x1pa1mepyfako5@localhost>
In reply to#8346
On Sun, 12 Jul 2015 06:17:10 -0400, Alexei A. Frounze  
<alexfrunews@gmail.com> wrote:

> On Sunday, July 12, 2015 at 1:24:36 AM UTC-7, Rod Pemberton wrote:
>> On Sat, 11 Jul 2015 23:39:51 -0400, Alexei A. Frounze
>> <...@gmail.com> wrote:
>>

>> I had to use hex formats, i.e., casts needed, for all of them since DOS
>> and Linux are incompatible in their implementations of various scanf()
>> formats.
>
> How so?

Some worked on Linux.  Some worked on DOS.

Part of the problem may have been some DOS input formats not
reading in what was output for the same output format.

> Vastly different versions of the language (pre-ASNI vs ANSI vs
> C99?) supported by those DOS and Linux compilers?

DJGPP v2.03 for 32-bit DOS (uses GCC v3.4.1)
GCC 4.7.2 GLIBC 2.16 for 64-bit Linux

> You used something non-standard?

No.  ANSI C.  Supported formats from Harbison & Steele book.

> You had to deal with different integer types sizes on the
> two platforms (e.g. 16-bit ints/pointers vs 32-bit ints/pointers)?

They were 'int' offsets from ftell(), so perhaps ...

>> DOS compilers wouldn't read in the same data they wrote for
>> some formats either.
>
> Or even library bugs?

I think there might be some library bugs for DJGPP, or a difference
perhaps in the handling of the format ...

The 'int' offset from ftell() for DJGPP in DOS was returning
exceptionally large numbers, i.e. apparently not zero-based or
perhaps due to my large input file ...  The output format seemed
to be treating these large values as signed and writing them
out as such as text.  But, the input format wouldn't correctly
scan in the large text values.  It was like the output was signed
but the input was broken or expecting unsigned for large values.

> Or you didn't use the library in intended ways?

How would I know that until I know that?
It's working now with my changes, so ...
I'll probably never know that.


Rod Pemberton

-- 
Tolerance and socialism attracts intolerance and terrorism.  See:
France, United Kingdom, Germany, Denmark, Belgium, Netherlands, ...
UK: Strong encryption, non-issue. GR: Eurozone currency, priority.

[toc] | [prev] | [next] | [standalone]


#8348

From"James Harris" <james.harris.1@gmail.com>
Date2015-07-12 19:12 +0100
Message-ID<mnuak9$f21$1@dont-email.me>
In reply to#8345
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message 
news:op.x1nizmxyyfako5@localhost...

...

> Did you know that stdin, stdout, stderr aren't static in Linux?  ...
>
> FILE *out=stdout ; not valid with GLIBC, move assignment to main() etc

Apparently stdin, stdout and stderr are defined to be macros so it 
should not be possible to copy them. A man page says that the call

  freopen

(i.e. "f reopen" rather then "fre open") is provided for such 
machinations.

James

[toc] | [prev] | [next] | [standalone]


#8349

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-07-12 17:09 -0700
Message-ID<d84b0130-4872-4bf4-b92e-2a5dabc1c279@googlegroups.com>
In reply to#8348
On Sunday, July 12, 2015 at 11:12:49 AM UTC-7, James Harris wrote:
> "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message 
> news:op.x1nizmxyyfako5@localhost...
> 
> ...
> 
> > Did you know that stdin, stdout, stderr aren't static in Linux?  ...
> >
> > FILE *out=stdout ; not valid with GLIBC, move assignment to main() etc
> 
> Apparently stdin, stdout and stderr are defined to be macros so it 
> should not be possible to copy them.

There's no guarantee that the three expand to constant expressions and
are suitable static object initializers.
However, them being macros does not by itself preclude them from bing
constant, just like I showed a BSD implementation.

> A man page says that the call
> 
>   freopen
> 
> (i.e. "f reopen" rather then "fre open") is provided for such 
> machinations.

Such being which? Reopening a stream (or rather, reassigning it to
point to another file) has its own issues (e.g. it first closes
one file, then opens another and you can't do much about closing
errors or retrieving std* once it's been closed).
Using a FILE pointer variable, which you can change, has fewer
associated problems.

Alex

[toc] | [prev] | [next] | [standalone]


#8354

From"James Harris" <james.harris.1@gmail.com>
Date2015-07-13 11:10 +0100
Message-ID<mo02q5$i3g$1@speranza.aioe.org>
In reply to#8348
"Alexei A. Frounze" <alexfrunews@gmail.com> wrote in message
news:d84b0130-4872-4bf4-b92e-2a5dabc1c279@googlegroups.com...
> On Sunday, July 12, 2015 at 11:12:49 AM UTC-7, James Harris wrote:
>> "Rod Pemberton" <...@fasdfrewar.cdm> wrote in message
>> news:op.x1nizmxyyfako5@localhost...
>>
>> ...
>>
>> > Did you know that stdin, stdout, stderr aren't static in Linux?
>> > ...
>> >
>> > FILE *out=stdout ; not valid with GLIBC, move assignment to main()
>> > etc
>>
>> Apparently stdin, stdout and stderr are defined to be macros so it
>> should not be possible to copy them.

That should say copy *to* them.

> There's no guarantee that the three expand to constant expressions and
> are suitable static object initializers.
> However, them being macros does not by itself preclude them from bing
> constant, just like I showed a BSD implementation.

That would be OK in some environments but not others. The text in the
man page I have for, say, stdin (i.e. type "man stdin" to see it) says

"Since the symbols stdin, stdout, and stderr are specified to be macros,
assigning to them is nonportable."

I would have thought that if they were defined that way then gcc would
issue a warning but from testing even with -Wall -Wextra no warning is
issued. There may, however, be another flag that will turn on such a
warning.

As stdin, stdout and stderr are pointers that it should be possible to
change I don't see how they can be constants, at least in a conforming
implementation. There are corresponding file descriptors, however, which
are integers and constants, i.e. STDIN_FILENO, STDOUT_FILENO,
STDERR_FILENO.

>> A man page says that the call
>>
>>   freopen
>>
>> (i.e. "f reopen" rather then "fre open") is provided for such
>> machinations.
>
> Such being which?

Reassigning to the file pointer.

> Reopening a stream (or rather, reassigning it to
> point to another file) has its own issues (e.g. it first closes
> one file, then opens another and you can't do much about closing
> errors or retrieving std* once it's been closed).

Agreed. I think we are *supposed* to use freopen if we want portable
code but you could probably fclose the old one first if you want to
catch errors.

> Using a FILE pointer variable, which you can change, has fewer
> associated problems.

Well, according to the man page that action is not portable.

I have been looking at freopen() and it is a bit of a weird one. The
signature is

  FILE *freopen(const char *path, const char *mode, FILE *stream);

The path and the mode are as in the fopen call but note that freopen has
*two* file pointers.

After some digging I think that the parameter FILE* is the stream that
should be reopened (repointed, if you like). The return FILE* is only
specified to be a FILE pointer

"Upon  successful  completion  fopen(),  fdopen()  and  freopen() return
a FILE pointer"

But it doesn't say what file it will point at. I can only guess that the
freopen call changes the parameter FILE pointer in some way and also
returns it the new value. The return is possibly there just to indicate
success (or failure if NULL).

A valid call, then, might be

  if (freopen("prog.c", "r", stdin)) ....

I still don't understand how that can change stdin unless freopen is a
macro. But it's not important. It is something I may never use.

James

[toc] | [prev] | [next] | [standalone]


#8362

From"Rod Pemberton" <boo@fasdfrewar.cdm>
Date2015-07-13 23:13 -0400
Message-ID<op.x1qtwod5yfako5@localhost>
In reply to#8354
On Mon, 13 Jul 2015 06:10:14 -0400, James Harris
<james.harris.1@gmail.com> wrote:

[...]

> As stdin, stdout and stderr are pointers that it should be possible
> to change

...

> I don't see how they can be constants, at least in a conforming
> implementation.

Most C implementations treat them as constants.  GCC is non-standard
here.  GCC argues that the C specification doesn't require static
pointers for stdin, stdout, stderr.  Apparently, they do this to
allow GCC to reassign streams, e.g., pipes.

> I have been looking at freopen() and it is a bit of a weird one. The
> signature is
>
>   FILE *freopen(const char *path, const char *mode, FILE *stream);
>
> The path and the mode are as in the fopen call but note that freopen
> has *two* file pointers.

It really has only *one* file pointer ...

According to H&S, not the C specifications, the return file pointer
must be either the passed file pointer 'stream' or NULL for failure.
I.e., FILE *stream is passed back as the return FILE * unless it needs
to be set to NULL for failure.

freopen() fclose's stream, fopen's path/mode, associates path/mode
with stream instead of using a new file pointer.

> After some digging I think that the parameter FILE* is the stream that
> should be reopened (repointed, if you like).

Yes.

> The return FILE* is only specified to be a FILE pointer

... and will be either 'stream' or NULL.  See above.

> "Upon  successful  completion  fopen(),  fdopen()  and  freopen() return
> a FILE pointer"
>
> But it doesn't say what file it will point at.

For fopen(), it will be the newly opened file ... or NULL.

For freopen(), it will be the path/mode you passed, using the original
'stream' file pointer, or it'll return NULL for failure.

fdopen() must be POSIX or Linux.

> I can only guess that the
> freopen call changes the parameter FILE pointer in some way and also
> returns it the new value. The return is possibly there just to indicate
> success (or failure if NULL).

Changes?

> A valid call, then, might be
>
>   if (freopen("prog.c", "r", stdin)) ....
>
> I still don't understand how that can change stdin unless freopen is a
> macro. But it's not important. It is something I may never use.

Why would there need to be any change? ...


Rod Pemberton

-- 
Tolerance and socialism attracts intolerance and terrorism.  See:
France, United Kingdom, Germany, Denmark, Belgium, Netherlands, ...
UK: Strong encryption, non-issue. GR: Eurozone currency, priority.

[toc] | [prev] | [next] | [standalone]


#8366

From"James Harris" <james.harris.1@gmail.com>
Date2015-07-14 09:08 +0100
Message-ID<mo2g2d$53h$1@speranza.aioe.org>
In reply to#8362
"Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
news:op.x1qtwod5yfako5@localhost...
> On Mon, 13 Jul 2015 06:10:14 -0400, James Harris
> <james.harris.1@gmail.com> wrote:

>> As stdin, stdout and stderr are pointers that it should be possible
>> to change
>
> ...
>
>> I don't see how they can be constants, at least in a conforming
>> implementation.
>
> Most C implementations treat them as constants.  GCC is non-standard
> here.  GCC argues that the C specification doesn't require static
> pointers for stdin, stdout, stderr.  Apparently, they do this to
> allow GCC to reassign streams, e.g., pipes.

The C89 draft that I have says: "The primary use of the freopen function
is to change the file associated with a standard text stream ( stderr ,
stdin , or stdout ), as those identifiers need not be modifiable lvalues
to which the value returned by the fopen function may be assigned."

So you may not be able to say

  stdout = open("prog.c", "r");

but you can say

  freopen("prog.c", "r", stdout);

>> I have been looking at freopen() and it is a bit of a weird one. The
>> signature is
>>
>>   FILE *freopen(const char *path, const char *mode, FILE *stream);
>>
>> The path and the mode are as in the fopen call but note that freopen
>> has *two* file pointers.
>
> It really has only *one* file pointer ...
>
> According to H&S, not the C specifications, the return file pointer
> must be either the passed file pointer 'stream' or NULL for failure.
> I.e., FILE *stream is passed back as the return FILE * unless it needs
> to be set to NULL for failure.

OK.

> freopen() fclose's stream, fopen's path/mode, associates path/mode
> with stream instead of using a new file pointer.

So 'stream' points to the same FILE struct and it is the FILE struct
that has somehow been changed rather than the pointer? That would make
sense in isolation but it would screw up any other FILE pointer which
pointed to the same struct. If it helps to visualise a FILE struct there
is a sample at

  http://www.c4learn.com/c-programming/c-file-structure-and-file-pointer/

It is as follows.

typedef struct
{
 short level ;
 short token ;
 short bsize ;
 char fd ;
 unsigned flags ;
 unsigned char hold ;
 unsigned char *buffer ;
 unsigned char * curp ;
 unsigned istemp;
}FILE ;

If the pointer is to be unchanged then the FILE struct that it points to
has to change - perhaps by changing the fd field, above, and others -
and, as I was saying, there could be another pointer to the same struct
which would also find that its associated file had changed. It would
seem better for freopen to allow the pointer to change as

  freopen(fname, mode, &stdout);

where the & allows the *pointer* to change rather than the FILE struct.

The freopen call may be a macro which does prepend an ampersand. I
cannot see how it could change the pointer otherwise.

>> After some digging I think that the parameter FILE* is the stream 
>> that
>> should be reopened (repointed, if you like).
>
> Yes.
>
>> The return FILE* is only specified to be a FILE pointer
>
> ... and will be either 'stream' or NULL.  See above.
>
>> "Upon  successful  completion  fopen(),  fdopen()  and  freopen() 
>> return
>> a FILE pointer"
>>
>> But it doesn't say what file it will point at.
>
> For fopen(), it will be the newly opened file ... or NULL.
>
> For freopen(), it will be the path/mode you passed, using the original
> 'stream' file pointer, or it'll return NULL for failure.
>
> fdopen() must be POSIX or Linux.
>
>> I can only guess that the
>> freopen call changes the parameter FILE pointer in some way and also
>> returns it the new value. The return is possibly there just to 
>> indicate
>> success (or failure if NULL).
>
> Changes?

Yes. Detail both above and below.

>> A valid call, then, might be
>>
>>   if (freopen("prog.c", "r", stdin)) ....
>>
>> I still don't understand how that can change stdin unless freopen is 
>> a
>> macro. But it's not important. It is something I may never use.
>
> Why would there need to be any change? ...

Well, stdin is a pointer. It points to a FILE struct. AIUI if it points
to the standard-in FILE struct before the call then I thought the
purpose of the call was to get it to point to a different FILE struct.
So it would have to change.

James

[toc] | [prev] | [next] | [standalone]


#8377

From"Rod Pemberton" <boo@fasdfrewar.cdm>
Date2015-07-15 02:23 -0400
Message-ID<op.x1sxdtm6yfako5@localhost>
In reply to#8366
On Tue, 14 Jul 2015 04:08:43 -0400, James Harris
<james.harris.1@gmail.com> wrote:

> "Rod Pemberton" <boo@fasdfrewar.cdm> wrote in message
> news:op.x1qtwod5yfako5@localhost...
>> On Mon, 13 Jul 2015 06:10:14 -0400, James Harris
>> <james.harris.1@gmail.com> wrote:

>>> As stdin, stdout and stderr are pointers that it should be possible
>>> to change
>>
>> ...
>>
>>> I don't see how they can be constants, at least in a conforming
>>> implementation.
>>
>> Most C implementations treat them as constants.  GCC is non-standard
>> here.  GCC argues that the C specification doesn't require static
>> pointers for stdin, stdout, stderr.  Apparently, they do this to
>> allow GCC to reassign streams, e.g., pipes.
>
> The C89 draft that I have says: "The primary use of the freopen function
> is to change the file associated with a standard text stream ( stderr ,
> stdin , or stdout ), as those identifiers need not be modifiable lvalues
> to which the value returned by the fopen function may be assigned."

Yes.

> So you may not be able to say
>
>   stdout = open("prog.c", "r");

Why would you want/need to do that though? ...

You can assign stdout's value to a variable of type file pointer
and use that file pointer to write to stdout.

> but you can say
>
>   freopen("prog.c", "r", stdout);

Yes.  You can attach or redirect stdout to a file with freopen().

>>> I have been looking at freopen() and it is a bit of a weird one. The
>>> signature is
>>>
>>>   FILE *freopen(const char *path, const char *mode, FILE *stream);
>>>
>>> The path and the mode are as in the fopen call but note that freopen
>>> has *two* file pointers.
>>
>> It really has only *one* file pointer ...
>>
>> According to H&S, not the C specifications, the return file pointer
>> must be either the passed file pointer 'stream' or NULL for failure.
>> I.e., FILE *stream is passed back as the return FILE * unless it needs
>> to be set to NULL for failure.
>
> OK.
>
>> freopen() fclose's stream, fopen's path/mode, associates path/mode
>> with stream instead of using a new file pointer.
>
> So 'stream' points to the same FILE struct

Yes.

> and it is the FILE struct
> that has somehow been changed rather than the pointer?

What?

Again, why are you insisting that something changes in that struct?

> That would make sense in isolation but it would screw up any other
> FILE pointer which pointed to the same struct.

Why would you have additional file pointers pointing to the same struct?

> If it helps to visualise a FILE struct there is a sample at
>
>   [link]
>
> It is as follows.
>
> typedef struct
> {
>  short level ;
>  short token ;
>  short bsize ;
>  char fd ;
>  unsigned flags ;
>  unsigned char hold ;
>  unsigned char *buffer ;
>  unsigned char * curp ;
>  unsigned istemp;
> }FILE ;

...

> If the pointer is to be unchanged then the FILE struct that it points to
> has to change - perhaps by changing the fd field, above, and others -

Why?  ...

1) The 'fd' was 0 for stdin before the call to freopen().  Still is.
2) The filename and file mode flags aren't stored in the struct.

What changed?

> and, as I was saying, there could be another pointer to the same struct
> which would also find that its associated file had changed.

How?

How could there be two file pointers to the same file or standard
I/O stream (stdin, stdout, stderr)?

You would have multiple writers to the file or stream.  That's bad.

> It would seem better for freopen to allow the pointer to change as
>
>   freopen(fname, mode, &stdout);
>
> where the & allows the *pointer* to change rather than the FILE struct.

I think you're confused about that struct needing to change.

> The freopen call may be a macro which does prepend an ampersand.
> I cannot see how it could change the pointer otherwise.

It doesn't or shouldn't.

>> Why would there need to be any change? ...
>
> Well, stdin is a pointer. It points to a FILE struct.

> AIUI if it points to the standard-in FILE struct before the call then
> I thought the purpose of the call was to get it to point to a different
> FILE struct.

No.

freopen() is to associate a different file, i.e., filename, with
the file ... However, the file here isn't a file but is a stream.
freopen() is used for redirecting stdin, stdout, stderr.

...

AISI, this is the everything is a file concept for POSIX.

stdin is a stream, but can be used like an open file.  What is
stdin's filename?  What name was used with fopen() by the OS to
open the file stdin?  It doesn't have one or it's unknown to us
since it's a stream.

I can only describe the "mental model" I have for this since I've
neither implemented it, nor studied it in any depth.  If implementing
a C library, I would need to find a solution that works correctly
according to C's usage.  I did do something in a very limited sense
related to this some years ago for in memory files.

So, in my "mental model," there should be at least *two* structs per
file, but only *one* struct per stream.  One of the file structs
handles the high-level filesystem information.  The other file struct
handles the low-level filesystem information.  In this case, a C FILE
struct would be the low-level information needed for direct or raw
filesystem access to the opened file.  There should also be another
high-level struct which handles the high-level file information, e.g.,
filename, path, drive, file open mode, renamed flag, delete flag,
type of buffering, pointers to buffers, pointer to directory entry,
etc.  All this info is created with fopen() and flushed with fclose().
So, one item in the high-level struct should be a FILE pointer which
points to the low-level struct for the opened file or stream.  In my
mental model, both of these are stored in arrays of the appropriate
type.  E.g., array of high-level structs, and an array of low-level
structs.  Now, 'stdin', 'stdout', and 'stderr' are streams and so
they don't have a high-level struct associated with them, unless
you use freopen() to connect them with a true file.


Rod Pemberton

[toc] | [prev] | [next] | [standalone]


#8498

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-08-15 02:12 -0700
Message-ID<d963a13a-1a1e-4adc-9e29-8fb6be652add@googlegroups.com>
In reply to#8344
What's new in the compiler?
- DPMI support and DPMI binaries of the compiler itself
- support for a.out (OMAGIC) executable format with relocation records
in the linker (that's what DPMI .EXEs really are), which may be
useful for OS dev
https://github.com/alexfru/SmallerC

Alex

[toc] | [prev] | [next] | [standalone]


#8712

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-09-06 15:54 -0700
Message-ID<870c24c2-224c-48e5-9955-fcbe6f67b68c@googlegroups.com>
In reply to#8498
What's new in the compiler?
- it can use FASM instead of NASM when compiling 32-bit protected
mode code (DOS/DPMI32, Windows, Linux)

There's now a wrapper around FASM, n2f.c, which if compiled into
nasm[.exe] will transparently convert the assembly code generated
by smlrc to the syntax and layout that FASM understands and invoke
FASM on it. A kind of NASM replacement/substitution. This tool
isn't general purpose, it converts only a very restricted subset of
assembly code from NASM syntax to FASM syntax and layout.

What does this mean? Well, for one thing, FASM is faster and
smaller than NASM, so you can make a floppy with the compiler
and the assembler and, say, DOS/DPMI32 library and you'll still
have about a half of the space free on the floppy. And you can
run this even in DOSBox without fearing that NASM would take
forever to assemble code with many jump instructions and without
having to give many many more CPU cycles to DOSBox.

Another thing is that this makes Smaller C fully and easily
portable to any x86 OS running apps in 32-bit protected mode.
Smaller C can fully recompile its libraries and itself if
there's NASM or YASM in your system. If porting NASM or YASM
is somehow problematic or undesirable, there's now another
option, FASM. FASM is written in 80386 assembly and can
recompile itself without needing any other tools. So, you
only need to teach FASM to use your OS syscalls and do the
same with the Smaller C library.

Ben, in case you're interested in shaving off another several
hundred of bytes from your bootloader, n2f.c combines section
fragments, which NBASM does not do just like FASM. You can
adapt it for NBASM.

Alex

[toc] | [prev] | [next] | [standalone]


#8714

From"Benjamin David Lunt" <zfysz@fysnet.net>
Date2015-09-07 12:19 -0700
Message-ID<msko0v$rtj$2@speranza.aioe.org>
In reply to#8712
"Alexei A. Frounze" <alexfrunews@gmail.com> wrote in message 
news:870c24c2-224c-48e5-9955-fcbe6f67b68c@googlegroups.com...

<snip>

> Ben, in case you're interested in shaving off another several
> hundred of bytes from your bootloader, n2f.c combines section
> fragments, which NBASM does not do just like FASM. You can
> adapt it for NBASM.

Since I save all of the strings until last, *and* I am only
interested in flat mode compiles, I don't have much interest
in that part of the compile process, but thanks for the note.

On another note, I have done a little work with mine, looking
at your latest and you have changed quite a bit of code.

My 'farE' stuff is now more complicated  :-)

I am working with your ParseBase() code to make it a little
easier on me to include my 'farE' stuff.

Anyway, my code is now very close to yours again, though
I cannot get it to compile my loader stuff yet due to the
'farE' parsing.  I hope a few more hours of work and I will
have that fixed.

How about another comment.  Notice the following:

int main(.....) {

  int array[10] = { 1, 2, 3, 4, .... };

  ....

}

Of course you have to some how fill the stack location
of 'array' with the values indicated.  Hence, the call
at the end of the code.

However, I haven't looked at your code to closely to
see, but does your code compensate for 16-bit/32-bit
compiles?

For example, let's say that _main is within 16-bit
code, and _sub is in 32-bit code.  Does your compiler
then generate two different function calls?  i.e.: does
it set one of two different flags to generate one of
two functions at the close of the compile?

I think mine does (or did).  I need to check to be sure.

Also, is there another way of doing this technique, the
technique of initializing a non-static array within a
function?

I was thinking of compiling some code with various C compilers
to see what they each produce.

My problem is that if I am within the first part of my loader
code, in 16-bit code, and don't generate the function call
within 15 bits of address, my compiler fails to generate
the correct code.  (again, remembering that my compiler
is made to switch between 16-bit and 32-bit code at will).

Anyway, just a thought.  Something I will be working on soon.

Thanks for the update.

Ben

-- 
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-
Forever Young Software
http://www.fysnet.net/index.htm
http://www.fysnet.net/osdesign_book_series.htm
To reply by email, please remove the zzzzzz's

Batteries not included, some Assembly required. 

[toc] | [prev] | [next] | [standalone]


#8716

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-09-07 13:30 -0700
Message-ID<c391e578-3385-4332-b8ba-44c917e3788d@googlegroups.com>
In reply to#8714
On Monday, September 7, 2015 at 12:20:04 PM UTC-7, Benjamin David Lunt wrote:
...
> On another note, I have done a little work with mine, looking
> at your latest and you have changed quite a bit of code.
> 
> My 'farE' stuff is now more complicated  :-)
> 
> I am working with your ParseBase() code to make it a little
> easier on me to include my 'farE' stuff.
> 
> Anyway, my code is now very close to yours again, though
> I cannot get it to compile my loader stuff yet due to the
> 'farE' parsing.  I hope a few more hours of work and I will
> have that fixed.
> 
> How about another comment.  Notice the following:
> 
> int main(.....) {
> 
>   int array[10] = { 1, 2, 3, 4, .... };
> 
>   ....
> 
> }
> 
> Of course you have to some how fill the stack location
> of 'array' with the values indicated.  Hence, the call
> at the end of the code.
> 
> However, I haven't looked at your code to closely to
> see, but does your code compensate for 16-bit/32-bit
> compiles?
> 
> For example, let's say that _main is within 16-bit
> code, and _sub is in 32-bit code.  Does your compiler
> then generate two different function calls?  i.e.: does
> it set one of two different flags to generate one of
> two functions at the close of the compile?

I don't support changing from generating 16-bit code to
generating 32-bit code and back. So, naturally, there's just
one subroutine emitted at the end. I look at SizeOfWord
and OutputFormat to see what code should be emitted there
(16-bit, 32-bit, huge).

> I think mine does (or did).  I need to check to be sure.

I can see that yours can emit both of them together.

> Also, is there another way of doing this technique, the
> technique of initializing a non-static array within a
> function?

You need code to fill a portion of the stack. There are two or
two and a half ways to do it, unless we're talking about
optimizing compilers.
1. Call an equivalent of memcpy() to copy the initializers
   to the stack from some static location
2. Inline the above call: generate code for "memcpy()" in-place
3. Generate code to initialize every array/struct member on the
   stack individually, as if by a series of assignment statements

If your array/struct is very small (a few machine words at most),
option 3 is probably the best. But you need to know the size
before you can make a choice. And you may not know it in
advance, e.g. 

  int array[] = /*Do we know the size yet?*/ { 1, 2, 3, 4, 5 };

So, before you can choose, you can have the initializers already
emitted.

> I was thinking of compiling some code with various C compilers
> to see what they each produce.

An optimizing compiler may eliminate copying (if the array is
never modified) and even the initializers for the array (if the
array is not used or if just one fixed element of it is used).
It depends on how the array is used and how good the compiler is.
If neither of these optimizations is done, there isn't much else
to do other than:
1. call _memcpy
2. rep movsb
3. mov, mov, mov, ...

> My problem is that if I am within the first part of my loader
> code, in 16-bit code, and don't generate the function call
> within 15 bits of address, my compiler fails to generate
> the correct code.  (again, remembering that my compiler
> is made to switch between 16-bit and 32-bit code at will).
> 
> Anyway, just a thought.  Something I will be working on soon.
> 
> Thanks for the update.

Cool!
Alex

[toc] | [prev] | [next] | [standalone]


#8733

From"Benjamin David Lunt" <zfysz@fysnet.net>
Date2015-09-09 19:38 -0700
Message-ID<msqqhg$7m5$1@speranza.aioe.org>
In reply to#8716
"Alexei A. Frounze" <alexfrunews@gmail.com> wrote in message 
news:c391e578-3385-4332-b8ba-44c917e3788d@googlegroups.com...
> On Monday, September 7, 2015 at 12:20:04 PM UTC-7, Benjamin David Lunt 
> wrote:
>>
>> For example, let's say that _main is within 16-bit
>> code, and _sub is in 32-bit code.  Does your compiler
>> then generate two different function calls?  i.e.: does
>> it set one of two different flags to generate one of
>> two functions at the close of the compile?
>
> I don't support changing from generating 16-bit code to
> generating 32-bit code and back. So, naturally, there's just
> one subroutine emitted at the end. I look at SizeOfWord
> and OutputFormat to see what code should be emitted there
> (16-bit, 32-bit, huge).
>
>> I think mine does (or did).  I need to check to be sure.
>
> I can see that yours can emit both of them together.

I remembered that yours did not as soon as I read your
reply :-)

>> Also, is there another way of doing this technique, the
>> technique of initializing a non-static array within a
>> function?

< snip Alex's reply. I will have to come back to this subject
  at a later time >

Alex,

Grab my latest code from
   http://www.fysnet.net/newbasic.htm
 or directly from
   http://www.fysnet.net/zips/nbasm.zip

and have a look.

Please note the major re-write is ParseBase().

It now allows you to have any keyword in any order even my 'farE'
stuff.

The biggest change is that I created a loop that loops until a
non-type keyword is found.

You may or may not want to modify yours to match, though in my
opinion, it will and does make the 'const', 'volatile', etc. keywords
available.  You no longer have to "ignore" them in your GetToken().
You can now do something with them here, in ParseBase(), or ignore
them here, again in ParseBase().  The mechanism for making sure
these keywords are allowed with others still remains from your current
code, though I have not checked against the C specs to verify
the accuracy.  A simple flag change is all that is needed.

I currently set a flag of ident_flags, but plan to combine it with
the passed base->flags field and remove the ident_flags parameter.

Anyway, your main focus should be with ParseBase().  The only other
changes are changes you already have made to yours.

Thanks again for your work.  It has been interesting and enjoyable
to work on mine, with the help of yours.

I should get to other parts of mine soon.  I need to still check
to see if my changes didn't break anything. :-)

Thanks,
Ben

-- 
-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-=-
Forever Young Software
http://www.fysnet.net/index.htm
http://www.fysnet.net/osdesign_book_series.htm
To reply by email, please remove the zzzzzz's

Batteries not included, some Assembly required. 

[toc] | [prev] | [next] | [standalone]


#8735

From"Benjamin David Lunt" <zfysz@fysnet.net>
Date2015-09-09 19:49 -0700
Message-ID<msqr4v$8r4$2@speranza.aioe.org>
In reply to#8733
"Benjamin David Lunt" <zfysz@fysnet.net> wrote in message 
news:msqqhg$7m5$1@speranza.aioe.org...
>
> I should get to other parts of mine soon.  I need to still check
> to see if my changes didn't break anything. :-)

Yep.  It broke a few things.  I will get them fixed and
upload a new version when I get a chance.

Thanks,
Ben

[toc] | [prev] | [next] | [standalone]


#8736

From"Benjamin David Lunt" <zfysz@fysnet.net>
Date2015-09-09 20:58 -0700
Message-ID<msqv7u$f22$1@speranza.aioe.org>
In reply to#8733
"Benjamin David Lunt" <zfysz@fysnet.net> wrote in message 
news:msqqhg$7m5$1@speranza.aioe.org...
>
>
> Grab my latest code from
>   http://www.fysnet.net/newbasic.htm
> or directly from
>   http://www.fysnet.net/zips/nbasm.zip
>

Okay, fixed the few errors.  Try the latest code uploaded
at 8.45pm Arizona USA time (UTC -7 ???)

Anyway, mine still gives the error of

"Can not return a value of type 'void'"

for:

tS f0(tS s) {
  return f0(s);
}

I will have to look a little more.

Anyway, of previous comments that your code still has
comments that it doesn't support returning 'structs',
ignore my mistake.  You have

  #ifndef NO_STRUCT_BY_VAL

in many places, but

  #ifdef NO_STRUCT_BY_VAL

in the two I was talking about.  Notice the difference
in the two 'if' statements.  My mistake, I didn't notice
the difference. :-)

Thanks again,
Ben

[toc] | [prev] | [next] | [standalone]


#8737

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-09-10 01:45 -0700
Message-ID<4837e85f-886b-4db4-8b7e-6005a9f39494@googlegroups.com>
In reply to#8733
On Wednesday, September 9, 2015 at 7:39:48 PM UTC-7, Benjamin David Lunt wrote:
> Alex,
> 
> Grab my latest code from
>    http://www.fysnet.net/newbasic.htm
>  or directly from
>    http://www.fysnet.net/zips/nbasm.zip
> 
> and have a look.
> 
> Please note the major re-write is ParseBase().
> 
> It now allows you to have any keyword in any order even my 'farE'
> stuff.
> 
> The biggest change is that I created a loop that loops until a
> non-type keyword is found.
> 
> You may or may not want to modify yours to match, though in my
> opinion, it will and does make the 'const', 'volatile', etc. keywords
> available.  You no longer have to "ignore" them in your GetToken().
> You can now do something with them here, in ParseBase(), or ignore
> them here, again in ParseBase().  The mechanism for making sure
> these keywords are allowed with others still remains from your current
> code, though I have not checked against the C specs to verify
> the accuracy.  A simple flag change is all that is needed.
> 
> I currently set a flag of ident_flags, but plan to combine it with
> the passed base->flags field and remove the ident_flags parameter.
> 
> Anyway, your main focus should be with ParseBase().  The only other
> changes are changes you already have made to yours.
> 
> Thanks again for your work.  It has been interesting and enjoyable
> to work on mine, with the help of yours.
> 
> I should get to other parts of mine soon.  I need to still check
> to see if my changes didn't break anything. :-)

[Sorry, haven't looked at the code yet.]

But that only takes care of cv modifiers in the base part of the type.
I've mentioned it multiple times now that const and volatile (and,
ideally, your Far* modifiers) may appear in the derived part as well:

int * volatile * const * const volatile pppi;

The base part is the int. Asterisks, parens and brackers further form
a type derived from the base type.

You also need to modify ParseDerived() to take further care of the
modifiers and, if you care about them anymore than I do, also store
them someplace, but that someplace obviously can't be the same as
where you store cvfar-ness of the base type, because the modifiers
can appear almost at any level between the base and the final
derived types.

Alex

[toc] | [prev] | [next] | [standalone]


#8738

From"Benjamin David Lunt" <zfysz@fysnet.net>
Date2015-09-10 10:45 -0700
Message-ID<mssfjm$e3$1@speranza.aioe.org>
In reply to#8737
"Alexei A. Frounze" <alexfrunews@gmail.com> wrote in message 
news:4837e85f-886b-4db4-8b7e-6005a9f39494@googlegroups.com...
> On Wednesday, September 9, 2015 at 7:39:48 PM UTC-7, Benjamin David Lunt 
> wrote:
>
> [Sorry, haven't looked at the code yet.]
>
> But that only takes care of cv modifiers in the base part of the type.
> I've mentioned it multiple times now that const and volatile (and,
> ideally, your Far* modifiers) may appear in the derived part as well:
>
> int * volatile * const * const volatile pppi;
>
> The base part is the int. Asterisks, parens and brackers further form
> a type derived from the base type.
>
> You also need to modify ParseDerived() to take further care of the
> modifiers and, if you care about them anymore than I do, also store
> them someplace, but that someplace obviously can't be the same as
> where you store cvfar-ness of the base type, because the modifiers
> can appear almost at any level between the base and the final
> derived types.

Unless I am missing something, ParseDerived() calls ParseBase()
to do just that.  Maybe I am mistaken and don't quite understand
what you are saying.

Can you provide an example?

My new code will easily parse the following:

long l;
signed long sl;
unsigned long ul;
unsigned long int uli;
const register auto unsigned long volatile farE *ulss;
unsigned long farE ule;
unsigned long farE *ulep;

and how about:

typedef   unsigned char   bit8u;
typedef   unsigned long   bit32u;

struct S_TYPE {
  bit8u  resv[10];
} type;

int main(void) {
  if (* (bit8u farE *) 0x0000FFFF & 0x10) {

  }
}

Please give me an example source file, just a line or two,
about what you are talking about and I will see if my code
compiles it or not.  Then if it does not, I will have a look
to see why not.

Thanks,
Ben

[toc] | [prev] | [next] | [standalone]


#8739

From"Alexei A. Frounze" <alexfrunews@gmail.com>
Date2015-09-10 23:20 -0700
Message-ID<a9110bdf-add3-428c-b6da-2dbd52f45557@googlegroups.com>
In reply to#8738
On Thursday, September 10, 2015 at 10:45:31 AM UTC-7, Benjamin David Lunt wrote:
> "Alexei A. Frounze" <...@gmail.com> wrote in message 
> news:4837e85f-886b-4db4-8b7e-6005a9f39494@googlegroups.com...
> > On Wednesday, September 9, 2015 at 7:39:48 PM UTC-7, Benjamin David Lunt 
> > wrote:
> >
> > [Sorry, haven't looked at the code yet.]
> >
> > But that only takes care of cv modifiers in the base part of the type.
> > I've mentioned it multiple times now that const and volatile (and,
> > ideally, your Far* modifiers) may appear in the derived part as well:
> >
> > int * volatile * const * const volatile pppi;
> >
> > The base part is the int. Asterisks, parens and brackers further form
> > a type derived from the base type.
> >
> > You also need to modify ParseDerived() to take further care of the
> > modifiers and, if you care about them anymore than I do, also store
> > them someplace, but that someplace obviously can't be the same as
> > where you store cvfar-ness of the base type, because the modifiers
> > can appear almost at any level between the base and the final
> > derived types.
> 
> Unless I am missing something, ParseDerived() calls ParseBase()
> to do just that. 

Nope, it does not do just that.

> Maybe I am mistaken and don't quite understand
> what you are saying.

You are, you don't. But I'll explain it in a moment.

> Can you provide an example?

It was in the message to which you just replied:

  int * volatile * const * const volatile pppi;

> My new code will easily parse the following:
> 
> long l;
> signed long sl;
> unsigned long ul;
> unsigned long int uli;
> const register auto unsigned long volatile farE *ulss;
> unsigned long farE ule;
> unsigned long farE *ulep;
> 
> and how about:
> 
> typedef   unsigned char   bit8u;
> typedef   unsigned long   bit32u;
> 
> struct S_TYPE {
>   bit8u  resv[10];
> } type;
> 
> int main(void) {
>   if (* (bit8u farE *) 0x0000FFFF & 0x10) {
> 
>   }
> }
> 
> Please give me an example source file, just a line or two,
> about what you are talking about and I will see if my code
> compiles it or not.  Then if it does not, I will have a look
> to see why not.

I don't know if it was the terminology that confused you or
you rarely used const and volatile in those other places,
where they can legally appear, but here's how it works.
Let's go back to the very basics or close to them.

Consider this declaration declaring multiple things:

  extern int a, *b, c[2], d(void), (*e)(void);

What do we declare here actually?

a: external int
b: external pointer to int
c: external array of 2 ints
d: external function (taking nothing) returning int
e: external pointer to function (taking nothing) returning int

Agreed?

All those types of b, c, d and e are derived types. They are
derived from int by adding asterisks, parens, brackets and stuff
like that. This int is the base type for a through e.

In place of that int you could just as well have an enum or
a struct (union):

  extern enum { EZERO } a, *b, c[2], d(void), (*e)(void);
  extern struct { int foo; } a, *b, c[2], d(void), (*e)(void);

ParseBase() parses base types only: void, all these integer
types (e.g. unsigned short int), enums and structs(union).

It does not parse anything beyond the base type, at least,
not by itself (it an be invoked recursively, though).

All of the following parts are parsed in ParseDerived():

  a
  *b
  c[2]
  d(void)
  (*e)(void)

Are you still with me?

A little guide through the code of ParseDerived():

The following parses the asterisk in "*b" (I'm quoting your
version now):

  while (tok == '*') {
    stars++;
    tok = GetToken();
  }

The following parses the identifier name in "a", "*b", "c[2]",
"d(void)", "*e" (this last one parsed on a recursive call to
ParseDerived()):

  } else if (tok == tokIdent) {
    PushSyntax3(tok, (bit32u) AddIdent(TokenIdentName, ident_flags), flags);
    tok = GetToken();
  }

The following is used when there's no identifier, e.g. you have
an unnamed parameter in a function prototype or a cast (both
look like declarations without an identifier name):

  } else
    PushSyntax3(tokIdent, (bit32u) AddIdent("<something>", ident_flags), flags);


The following parses array dimensions in "c[2]":

  } else if (tok == '[') {
    // allow the first [] without the dimension in function parameters
    int allowEmptyDimension = 1;
    if (SyntaxStack[SyntaxStackCnt - 1].tok == ')')
      error(__LINE__, 0x010F, FALSE, GetTokenName('[')); // function returning array
    while (tok == '[') {
      int oldsp = SyntaxStackCnt;
      PushSyntax(tokVoid); // prevent cases like "int arr[arr];" and "int arr[arr[0]];"
      PushSyntax(tok);
      tok = ParseArrayDimension(allowEmptyDimension);
      if (tok != ']')
        error(__LINE__, 0x010F, FALSE, GetTokenName(tok));
      PushSyntax(']');
      tok = GetToken();
      DeleteSyntax(oldsp, 1);
      allowEmptyDimension = 0;
    }

This takes care of functions and their parameters in "d(void)"
and "(*e)(void)":

  if (params || (tok == '(')) {
    int t = SyntaxStack[SyntaxStackCnt - 1].tok;
    if ((t == ')') | (t == ']'))
      error(__LINE__, 0x010F, FALSE, GetTokenName('(')); // array of functions or function returning function
    if (!params)
      tok = GetToken();
    else
      PushSyntax2(tokIdent, (bit32u) AddIdent("<something>", ident_flags));
    if (isInterrupt)
      PushSyntax2('(', 1);
    else // fallthrough
      PushSyntax('(');
    
    ParseLevel++;
    ParamLevel++;
    ParseFxnParams(tok);
    ParamLevel--;
    ParseLevel--;
    PushSyntax(')');
    tok = GetToken();
  }

And this handles recursion in declarations, e.g. dives inside
of "(*e)" to then parse "*e" just as in the case of "*b":

  if (tok == '(') {
    tok = GetToken();
    if (tok != ')' && !TokenStartsDeclaration(tok, 1)) {
      tok = ParseDerived(tok, flags, ident_flags);
      if (tok != ')')
        error(__LINE__, 0x0111, FALSE, GetTokenName(tok));
      tok = GetToken();
    } else
      params = 1;
  }

Still with me?

Now comes the most important part. You can have const and volatile
in both the base type and in what comes after it and produces the
derived type.

Let me copy that example one more time so it's not lost:

  int * volatile * const * const volatile pppi;

These consts and volatiles aren't redundant. Generally, it's not
just a matter of taste of where to put them.

What does the above declaration declare?

Let's read it right to left:

pppi: pointer (itself being const and volatile) to
      pointer (which is constant) to
      pointer (which is volatile) to
      int.

You can also make that int (the base type) const and/or volatile:

  const volatile int * volatile * const * const volatile pppi;

and in the base type part it is permitted to put these modifiers
before or after the type they modify:

  int const volatile * volatile * const * const volatile pppi;

Still with me?

Well, then there, you have it. You have just been explained why
incorporating support for const and volatile in ParseBase() only
is insufficient.

Agreed? Or was this still somehow confusing? If it was, ask.

Cheers,
Alex

[toc] | [prev] | [next] | [standalone]


Page 1 of 5  [1] 2 3 4 5  Next page →

Back to top | Article view | alt.os.development


csiph-web