Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c > #395686 > unrolled thread
| Started by | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| First post | 2025-12-06 01:05 +0000 |
| Last post | 2025-12-17 00:52 -0600 |
| Articles | 20 on this page of 119 — 23 participants |
Back to article view | Back to comp.lang.c
is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-06 01:05 +0000
Re: is_binary_file() Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2025-12-06 01:41 +0000
Re: is_binary_file() Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2025-12-06 02:00 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 17:40 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 11:35 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-10 15:07 +0000
Re: is_binary_file() Michael S <already5chosen@yahoo.com> - 2025-12-10 19:00 +0200
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-10 17:18 +0000
Re: is_binary_file() Richard Heathfield <rjh@cpax.org.uk> - 2025-12-10 19:42 +0000
Re: is_binary_file() bart <bc@freeuk.com> - 2025-12-10 22:37 +0000
Re: is_binary_file() Paul <nospam@needed.invalid> - 2025-12-10 22:35 -0500
Re: is_binary_file() bart <bc@freeuk.com> - 2025-12-11 11:46 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-11 12:53 +0100
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 18:42 +0000
Re: is_binary_file() Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2025-12-10 15:58 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 18:44 +0000
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-10 12:46 -0500
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 18:45 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 18:41 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 20:57 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-10 22:07 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-11 01:09 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-11 12:33 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-12 19:25 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-12 22:54 +0000
Re: is_binary_file() "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2025-12-12 15:33 -0800
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-13 00:20 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-13 02:32 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-16 00:26 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-16 17:24 +0100
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-17 03:19 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-17 07:57 +0100
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-17 19:35 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-18 08:44 +0100
Re: is_binary_file() bart <bc@freeuk.com> - 2025-12-18 12:49 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-18 14:06 +0100
Re: is_binary_file() gazelle@shell.xmission.com (Kenny McCormack) - 2025-12-18 13:17 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-18 16:03 +0100
Re: is_binary_file() Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2025-12-05 17:42 -0800
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-06 17:37 +0000
Re: is_binary_file() Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2025-12-06 16:05 -0800
Re: is_binary_file() Louis Krupp <lkrupp@invalid.pssw.com.invalid> - 2025-12-07 03:43 -0700
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-07 16:47 +0000
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-27 03:18 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 17:46 +0000
Re: is_binary_file() Kaz Kylheku <046-301-5902@kylheku.com> - 2025-12-06 02:42 +0000
Re: is_binary_file() bart <bc@freeuk.com> - 2025-12-06 12:42 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-06 17:40 +0000
Re: is_binary_file() Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2025-12-06 18:04 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-06 19:06 +0000
Re: is_binary_file() Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2025-12-06 21:16 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 17:48 +0000
Re: is_binary_file() Kaz Kylheku <046-301-5902@kylheku.com> - 2025-12-08 19:26 +0000
Re: is_binary_file() Richard Heathfield <rjh@cpax.org.uk> - 2025-12-08 19:42 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-09 21:49 +0000
Re: is_binary_file() Paul <nospam@needed.invalid> - 2025-12-06 03:14 -0500
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 17:56 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-08 20:16 +0000
Re: is_binary_file() David Brown <david.brown@hesbynett.no> - 2025-12-09 09:03 +0100
Re: is_binary_file() Richard Heathfield <rjh@cpax.org.uk> - 2025-12-09 09:43 +0000
Re: is_binary_file() Richard Harnden <richard.nospam@gmail.invalid> - 2025-12-09 10:17 +0000
Re: is_binary_file() Kaz Kylheku <046-301-5902@kylheku.com> - 2025-12-09 20:15 +0000
Re: is_binary_file() tTh <tth@none.invalid> - 2025-12-09 12:22 +0100
Re: is_binary_file() Paul <nospam@needed.invalid> - 2025-12-09 20:26 -0500
Re: is_binary_file() Paul <nospam@needed.invalid> - 2025-12-09 06:38 -0500
Re: is_binary_file() Michael S <already5chosen@yahoo.com> - 2025-12-09 17:31 +0200
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-28 02:49 +0000
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-28 00:12 +0000
Re: is_binary_file() richard@cogsci.ed.ac.uk (Richard Tobin) - 2025-12-28 00:43 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-06 17:33 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-07 19:04 +0100
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-06 20:37 -0500
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 18:02 +0000
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-09 16:29 -0500
Re: is_binary_file() Michael S <already5chosen@yahoo.com> - 2025-12-10 11:21 +0200
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-10 12:48 -0500
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 11:38 +0000
Re: is_binary_file() antispam@fricas.org (Waldek Hebisch) - 2025-12-07 03:43 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 18:04 +0000
Re: is_binary_file() bart <bc@freeuk.com> - 2025-12-08 18:44 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-09 19:53 +0000
Re: is_binary_file() Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2025-12-09 15:42 -0800
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 11:41 +0000
Re: is_binary_file() Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2025-12-10 15:20 -0800
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-10 23:59 +0000
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-09 16:23 -0500
Re: is_binary_file() Richard Harnden <richard.nospam@gmail.invalid> - 2025-12-07 19:01 +0000
Re: is_binary_file() Richard Heathfield <rjh@cpax.org.uk> - 2025-12-07 21:51 +0000
Re: is_binary_file() Richard Harnden <richard.nospam@gmail.invalid> - 2025-12-07 22:49 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-08 13:51 +0100
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-08 16:04 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-08 19:27 +0100
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-27 05:51 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-29 16:06 +0100
Re: is_binary_file() mjos_examine <m6502x64@gmail.com> - 2025-12-29 11:49 -0500
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-29 20:49 +0100
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-30 01:52 +0000
Re: is_binary_file() scott@slp53.sl.home (Scott Lurndal) - 2025-12-08 16:02 +0000
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 18:07 +0000
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-27 03:13 +0000
Re: is_binary_file() Paul <nospam@needed.invalid> - 2025-12-27 01:28 -0500
Re: is_binary_file() Lawrence D’Oliveiro <ldo@nz.invalid> - 2025-12-27 21:27 +0000
Re: is_binary_file() antispam@fricas.org (Waldek Hebisch) - 2025-12-28 05:46 +0000
Re: is_binary_file() "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2025-12-07 14:42 -0800
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-08 18:09 +0000
Re: is_binary_file() "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2025-12-09 12:45 -0800
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-08 20:36 +0100
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-08 20:50 +0100
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-09 15:09 +0100
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-10 09:18 +0100
Re: is_binary_file() Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2025-12-08 14:43 -0800
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-09 21:38 +0000
Re: is_binary_file() Kaz Kylheku <046-301-5902@kylheku.com> - 2025-12-11 17:33 +0000
Re: is_binary_file() Bonita Montero <Bonita.Montero@gmail.com> - 2025-12-11 19:10 +0100
Re: is_binary_file() "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2025-12-11 14:56 -0800
Re: is_binary_file() James Kuyper <jameskuyper@alumni.caltech.edu> - 2025-12-11 18:15 -0500
Re: is_binary_file() Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2025-12-12 02:19 +0100
Re: is_binary_file() Michael Sanders <porkchop@invalid.foo> - 2025-12-14 08:27 +0000
Re: is_binary_file() Lynn McGuire <lynnmcguire5@gmail.com> - 2025-12-17 00:52 -0600
Page 1 of 6 [1] 2 3 4 5 6 Next page →
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-06 01:05 +0000 |
| Subject | is_binary_file() |
| Message-ID | <10gvvh8$1vv6e$1@dont-email.me> |
Am I close? Missing anything you'd consider to be (or not) needed?
<stdio.h>
/*
* Checks if a file is likely a binary by examining its content
* for NULL bytes (0x00) or unusual control characters.
* Returns 0 if text, 1 if binary or file open failure.
*/
int is_binary_file(const char *path) {
FILE *f = fopen(path, "rb");
if (!f) return 1; // cannot open file, treat as error/fail check
unsigned char buf[65536];
size_t n, i;
while ((n = fread(buf, 1, sizeof(buf), f)) > 0) {
for (i = 0; i < n; i++) {
unsigned char c = buf[i];
// 1. check for the NULL byte (strong indicator of binary data)
if (c == 0x00) {
fclose(f);
return 1; // IS binary
}
// 2. check for C0 control codes (0x01-0x1F), excluding known
// text formatting characters: 0x09 (Tab), 0x0A (LF), 0x0D (CR)
if (c < 0x20) {
if (c != 0x09 && c != 0x0A && c != 0x0D) {
fclose(f);
return 1; // IS binary (contains unexpected control code)
}
}
}
}
fclose(f);
return 0; // NOT binary
}
--
:wq
Mike Sanders
[toc] | [next] | [standalone]
| From | Lew Pitcher <lew.pitcher@digitalfreehold.ca> |
|---|---|
| Date | 2025-12-06 01:41 +0000 |
| Message-ID | <10h01k8$1ta3h$1@dont-email.me> |
| In reply to | #395686 |
On Sat, 06 Dec 2025 01:05:44 +0000, Michael Sanders wrote:
> Am I close? Missing anything you'd consider to be (or not) needed?
>
> <stdio.h>
>
> /*
> * Checks if a file is likely a binary by examining its content
> * for NULL bytes (0x00) or unusual control characters.
> * Returns 0 if text, 1 if binary or file open failure.
> */
First off, until we get computers that store file data in formats
other than binary, /all/ files (text or not) are "binary" files
(meaning that an is_binary_file() function should always return true).
OTOH, "text files" are a distinguishable subset of binary files.
I suggest that this makes an "is_text_file()" function more valuable
and more fitting than an "is_binary_file()" function.
Secondly, ISTM that the function should return a unique failure value
rather than overload the "is binary" return value. After all, you
actually have three return values: is_text, is_not_text, and
is_indeterminate (because of file access failure).
Thirdly, your determination of whether or not the file contains text
seemingly depends only on the existence or absence of certain control
characters. But text isn't just control characters; so you need a test
for invalid non-control characters as well. And, IIRC, not all control
characters occupy the ASCII/Unicode C0 band, so you might have to expand
your "acceptable control character" test to include some of those other
control codes.
Finally, you've hardcoded the binary values for certain acceptable
ASCII/Unicode control characters. However, not all platforms use ASCII
or Unicode, and these tests would fail to test the corresponding character
value correctly (I think here of EBCDIC, where "Line Feed" doesn't exist
but it's equivalent "NewLine" is 0x15 and Horizontal Tab is 0x05). Better
here to use the C equivalent escape characters '\n' and '\t' instead.
You may also consider expanding the control-character test to include other
line-formatting characters (at least as far as C will allow): Vertical Tab
('\v'), Form Feed ('\f'), Carriage Return ('\r') and Backspace ('\b').
> int is_binary_file(const char *path) {
> FILE *f = fopen(path, "rb");
> if (!f) return 1; // cannot open file, treat as error/fail check
>
> unsigned char buf[65536];
> size_t n, i;
>
> while ((n = fread(buf, 1, sizeof(buf), f)) > 0) {
> for (i = 0; i < n; i++) {
> unsigned char c = buf[i];
>
> // 1. check for the NULL byte (strong indicator of binary data)
> if (c == 0x00) {
> fclose(f);
> return 1; // IS binary
> }
>
> // 2. check for C0 control codes (0x01-0x1F), excluding known
> // text formatting characters: 0x09 (Tab), 0x0A (LF), 0x0D (CR)
> if (c < 0x20) {
> if (c != 0x09 && c != 0x0A && c != 0x0D) {
> fclose(f);
> return 1; // IS binary (contains unexpected control code)
> }
> }
> }
> }
>
> fclose(f);
> return 0; // NOT binary
> }
--
Lew Pitcher
"In Skills We Trust"
Not LLM output - I'm just like this.
[toc] | [prev] | [next] | [standalone]
| From | Lew Pitcher <lew.pitcher@digitalfreehold.ca> |
|---|---|
| Date | 2025-12-06 02:00 +0000 |
| Message-ID | <10h02nm$1ta3h$2@dont-email.me> |
| In reply to | #395687 |
On Sat, 06 Dec 2025 01:41:28 +0000, Lew Pitcher wrote: > On Sat, 06 Dec 2025 01:05:44 +0000, Michael Sanders wrote: > >> Am I close? Missing anything you'd consider to be (or not) needed? >> >> <stdio.h> >> >> /* >> * Checks if a file is likely a binary by examining its content >> * for NULL bytes (0x00) or unusual control characters. >> * Returns 0 if text, 1 if binary or file open failure. >> */ > > First off, until we get computers that store file data in formats > other than binary, /all/ files (text or not) are "binary" files > (meaning that an is_binary_file() function should always return true). > OTOH, "text files" are a distinguishable subset of binary files. > I suggest that this makes an "is_text_file()" function more valuable > and more fitting than an "is_binary_file()" function. > > Secondly, ISTM that the function should return a unique failure value > rather than overload the "is binary" return value. After all, you > actually have three return values: is_text, is_not_text, and > is_indeterminate (because of file access failure). [snip] I should have added that I feel that you probably haven't really defined /what/ "text file" means, and that has interfered with the development of this function. As Keith pointed out, the task of distinguishing between a "text" file and a "binary" file is not easy. I'll add that a lot of the difficulty stems from the fact that there are many definitions (some conflicting) of what a "text" file actually contains. The best advice I can give here is that you should pick a definition of what a text file consists of, document /that/ definition, and use /that/ documentation to build your code. If you say that, for instance, EBCDIC is out of scope, then your code does not have to handle EBCDIC (but if you /don't/ say that, then you leave your code open to the ambiguity of whether or not it will work with EBCDIC). Likewise for ASCII or "Extended ASCII" (sic) or Unicode (or 6Bit (multiple different choices here) or Baudot or even Morse). With suitable definitions beforehand, you can write an acceptable "is_text_file()" function and/or a passable "is_binary_file()" function. HTH -- Lew Pitcher "In Skills We Trust" Not LLM output - I'm just like this.
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-08 17:40 +0000 |
| Message-ID | <10h72ja$9q1e$1@dont-email.me> |
| In reply to | #395689 |
On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: > HTH Yes sir it really does. I'll study your post closely & dont think because my reply is brief that I'm not considering your words. Thank you Lew. -- :wq Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 11:35 +0000 |
| Message-ID | <10hbluk$1fc0e$1@dont-email.me> |
| In reply to | #395689 |
On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote:
> I should have added that I feel that you probably haven't really
> defined /what/ "text file" means, and that has interfered with
> the development of this function. As Keith pointed out, the task
> of distinguishing between a "text" file and a "binary" file is not
> easy. I'll add that a lot of the difficulty stems from the fact
> that there are many definitions (some conflicting) of what a "text"
> file actually contains.
Yes. Here's my 2nd attempt following the template (of thinking)
you've suggested...
#include <stdio.h> // FILE, fopen, fread, fclose
#include <stddef.h> // size_t
// is_text_file()
// Returns:
// -1 : could not open file
// 0 : is NOT a text file (binary indicators found)
// 1 : is PROBABLY a text file (no strong binary signatures)
int is_text_file(const char *path) {
// Try opening the file in binary mode,
// required so that bytes are read exact.
FILE *f = fopen(path, "rb");
if (!f) return -1; // Could not open file
unsigned char buf[4096]; // 4KB chunks
size_t n, i;
// Read in file until EOF
while ((n = fread(buf, 1, sizeof(buf), f)) > 0) {
for (i = 0; i < n; i++) {
unsigned char c = buf[i];
// 1. null byte is a very strong indication of binary data.
// Text files virtually never contain 0x00.
if (c == 0x00) {
fclose(f);
return 0; // Contains binary-only byte: NOT text
}
// 2. Check for raw C0 control codes (0x01–0x1F).
// We *allow* \t (09), \n (0A), \r (0D) because they are normal in text.
// Any other control code is highly suspicious and usually means binary.
if (c < 0x20) {
if (c != 0x09 && c != 0x0A && c != 0x0D) {
fclose(f);
return 0; // unexpected control character → NOT text
}
}
// 3. NOTE: We intentionally do *not* reject bytes >= 0x80.
// These occur in UTF-8, extended ASCII, and many local encodings.
// Rejecting them would treat valid multilingual text as binary.
// So we treat high bytes as acceptable for "probably text".
}
}
fclose(f);
return 1; // Probably text (no strong binary signatures found)
}
--
:wq
Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-12-10 15:07 +0000 |
| Message-ID | <S0g_Q.170822$l1A9.2749@fx10.iad> |
| In reply to | #395756 |
Michael Sanders <porkchop@invalid.foo> writes: >On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: > >> I should have added that I feel that you probably haven't really >> defined /what/ "text file" means, and that has interfered with >> the development of this function. As Keith pointed out, the task >> of distinguishing between a "text" file and a "binary" file is not >> easy. I'll add that a lot of the difficulty stems from the fact >> that there are many definitions (some conflicting) of what a "text" >> file actually contains. > >Yes. Here's my 2nd attempt following the template (of thinking) >you've suggested... The problem with all of your attempts is the performance issue. Success requires reading every single byte of the file, one byte at a time. The word 'slow' is not sufficient to describe how bad the performance will be for a very large file. At a minimum, dump the stdio double-buffered byte-by-byte algorithm and use mmap(). In reality, I still don't see any benefit to this type of heuristic-based approach.
[toc] | [prev] | [next] | [standalone]
| From | Michael S <already5chosen@yahoo.com> |
|---|---|
| Date | 2025-12-10 19:00 +0200 |
| Message-ID | <20251210190038.0000617a@yahoo.com> |
| In reply to | #395759 |
On Wed, 10 Dec 2025 15:07:30 GMT scott@slp53.sl.home (Scott Lurndal) wrote: > Michael Sanders <porkchop@invalid.foo> writes: > >On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: > > > >> I should have added that I feel that you probably haven't really > >> defined /what/ "text file" means, and that has interfered with > >> the development of this function. As Keith pointed out, the task > >> of distinguishing between a "text" file and a "binary" file is not > >> easy. I'll add that a lot of the difficulty stems from the fact > >> that there are many definitions (some conflicting) of what a "text" > >> file actually contains. > > > >Yes. Here's my 2nd attempt following the template (of thinking) > >you've suggested... > > The problem with all of your attempts is the performance > issue. Success requires reading every single byte of the > file, one byte at a time. The word 'slow' is not sufficient > to describe how bad the performance will be for a very large > file. > > At a minimum, dump the stdio double-buffered byte-by-byte > algorithm and use mmap(). > I suggest to do actual speed measurements before making bold claims like above. Don't trust your intuition! > In reality, I still don't see any benefit to this type of > heuristic-based approach. > Neither do I. But OP is not doing it for us, but for himself.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2025-12-10 17:18 +0000 |
| Message-ID | <VXh_Q.7622$Dklb.6363@fx17.iad> |
| In reply to | #395761 |
Michael S <already5chosen@yahoo.com> writes: >On Wed, 10 Dec 2025 15:07:30 GMT >scott@slp53.sl.home (Scott Lurndal) wrote: > >> Michael Sanders <porkchop@invalid.foo> writes: >> >On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: >> > >> >> I should have added that I feel that you probably haven't really >> >> defined /what/ "text file" means, and that has interfered with >> >> the development of this function. As Keith pointed out, the task >> >> of distinguishing between a "text" file and a "binary" file is not >> >> easy. I'll add that a lot of the difficulty stems from the fact >> >> that there are many definitions (some conflicting) of what a "text" >> >> file actually contains. >> > >> >Yes. Here's my 2nd attempt following the template (of thinking) >> >you've suggested... >> >> The problem with all of your attempts is the performance >> issue. Success requires reading every single byte of the >> file, one byte at a time. The word 'slow' is not sufficient >> to describe how bad the performance will be for a very large >> file. >> >> At a minimum, dump the stdio double-buffered byte-by-byte >> algorithm and use mmap(). >> > >I suggest to do actual speed measurements before making bold >claims like above. Don't trust your intuition! I have, more than once, done such measurements after mmap() was introduced in SVR4 circa 1989 (ported from SunOS). On a single-user system, running a single job, the difference for smaller files is in the noise. For larger files, or when the system is heavily loaded or multiuser, it can be significant.
[toc] | [prev] | [next] | [standalone]
| From | Richard Heathfield <rjh@cpax.org.uk> |
|---|---|
| Date | 2025-12-10 19:42 +0000 |
| Message-ID | <10hcif0$1njdq$1@dont-email.me> |
| In reply to | #395762 |
On 10/12/2025 17:18, Scott Lurndal wrote: > Michael S <already5chosen@yahoo.com> writes: >> On Wed, 10 Dec 2025 15:07:30 GMT >> scott@slp53.sl.home (Scott Lurndal) wrote: >> >>> Michael Sanders <porkchop@invalid.foo> writes: >>>> On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: >>>> >>>>> I should have added that I feel that you probably haven't really >>>>> defined /what/ "text file" means, and that has interfered with >>>>> the development of this function. As Keith pointed out, the task >>>>> of distinguishing between a "text" file and a "binary" file is not >>>>> easy. I'll add that a lot of the difficulty stems from the fact >>>>> that there are many definitions (some conflicting) of what a "text" >>>>> file actually contains. >>>> >>>> Yes. Here's my 2nd attempt following the template (of thinking) >>>> you've suggested... >>> >>> The problem with all of your attempts is the performance >>> issue. Success requires reading every single byte of the >>> file, one byte at a time. The word 'slow' is not sufficient >>> to describe how bad the performance will be for a very large >>> file. >>> >>> At a minimum, dump the stdio double-buffered byte-by-byte >>> algorithm and use mmap(). >>> >> >> I suggest to do actual speed measurements before making bold >> claims like above. Don't trust your intuition! > > I have, more than once, done such measurements after mmap() > was introduced in SVR4 circa 1989 (ported from SunOS). > > On a single-user system, running a single job, the difference > for smaller files is in the noise. For larger files, or when > the system is heavily loaded or multiuser, it can be significant. 1989 is 36 years ago. Technology has moved on. If reading your file is too slow to read, get yourself a real computer. On my very ordinary desktop machine, I just freq'd[1] a 7,032,963,565-byte file in 12.256 seconds. That's 573,838,410 bytes per second. It's a damn sight faster than I could do by hand. How, exactly, are you using `slow'? [1] Nothing fancy; a getc loop with ++pfm[ch].count written entirely in what used to be called clc-conforming code, and I can see at least one egregious inefficiency in the code that I can't be bothered to fix because half a gig a second is *easily* fast enough for my needs. -- Richard Heathfield Email: rjh at cpax dot org dot uk "Usenet is a strange place" - dmr 29 July 1999 Sig line 4 vacant - apply within
[toc] | [prev] | [next] | [standalone]
| From | bart <bc@freeuk.com> |
|---|---|
| Date | 2025-12-10 22:37 +0000 |
| Message-ID | <10hcsnq$1qc9f$1@dont-email.me> |
| In reply to | #395770 |
On 10/12/2025 19:42, Richard Heathfield wrote: > On 10/12/2025 17:18, Scott Lurndal wrote: >> Michael S <already5chosen@yahoo.com> writes: >>> On Wed, 10 Dec 2025 15:07:30 GMT >>> scott@slp53.sl.home (Scott Lurndal) wrote: >>> >>>> Michael Sanders <porkchop@invalid.foo> writes: >>>>> On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote: >>>>>> I should have added that I feel that you probably haven't really >>>>>> defined /what/ "text file" means, and that has interfered with >>>>>> the development of this function. As Keith pointed out, the task >>>>>> of distinguishing between a "text" file and a "binary" file is not >>>>>> easy. I'll add that a lot of the difficulty stems from the fact >>>>>> that there are many definitions (some conflicting) of what a "text" >>>>>> file actually contains. >>>>> >>>>> Yes. Here's my 2nd attempt following the template (of thinking) >>>>> you've suggested... >>>> >>>> The problem with all of your attempts is the performance >>>> issue. Success requires reading every single byte of the >>>> file, one byte at a time. The word 'slow' is not sufficient >>>> to describe how bad the performance will be for a very large >>>> file. >>>> >>>> At a minimum, dump the stdio double-buffered byte-by-byte >>>> algorithm and use mmap(). >>>> >>> >>> I suggest to do actual speed measurements before making bold >>> claims like above. Don't trust your intuition! >> >> I have, more than once, done such measurements after mmap() >> was introduced in SVR4 circa 1989 (ported from SunOS). >> >> On a single-user system, running a single job, the difference >> for smaller files is in the noise. For larger files, or when >> the system is heavily loaded or multiuser, it can be significant. > > 1989 is 36 years ago. Technology has moved on. If reading your file is > too slow to read, get yourself a real computer. > > On my very ordinary desktop machine, I just freq'd[1] a 7,032,963,565- > byte file in 12.256 seconds. That's 573,838,410 bytes per second. It's a > damn sight faster than I could do by hand. > > How, exactly, are you using `slow'? > A getc loop took 4.3 seconds to read a 192MB file from SSD, on my Windows PC. Under WSL it took 8.4 seconds (8.4/0.5 real/user). However reading it all in one go took 0.14 seconds. I guess not all 'getc' implementations are the same.
[toc] | [prev] | [next] | [standalone]
| From | Paul <nospam@needed.invalid> |
|---|---|
| Date | 2025-12-10 22:35 -0500 |
| Message-ID | <10hde6r$1v5rr$1@dont-email.me> |
| In reply to | #395773 |
On Wed, 12/10/2025 5:37 PM, bart wrote:
> On 10/12/2025 19:42, Richard Heathfield wrote:
>> On 10/12/2025 17:18, Scott Lurndal wrote:
>>> Michael S <already5chosen@yahoo.com> writes:
>>>> On Wed, 10 Dec 2025 15:07:30 GMT
>>>> scott@slp53.sl.home (Scott Lurndal) wrote:
>>>>
>>>>> Michael Sanders <porkchop@invalid.foo> writes:
>>>>>> On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote:
>>>>>>> I should have added that I feel that you probably haven't really
>>>>>>> defined /what/ "text file" means, and that has interfered with
>>>>>>> the development of this function. As Keith pointed out, the task
>>>>>>> of distinguishing between a "text" file and a "binary" file is not
>>>>>>> easy. I'll add that a lot of the difficulty stems from the fact
>>>>>>> that there are many definitions (some conflicting) of what a "text"
>>>>>>> file actually contains.
>>>>>>
>>>>>> Yes. Here's my 2nd attempt following the template (of thinking)
>>>>>> you've suggested...
>>>>>
>>>>> The problem with all of your attempts is the performance
>>>>> issue. Success requires reading every single byte of the
>>>>> file, one byte at a time. The word 'slow' is not sufficient
>>>>> to describe how bad the performance will be for a very large
>>>>> file.
>>>>>
>>>>> At a minimum, dump the stdio double-buffered byte-by-byte
>>>>> algorithm and use mmap().
>>>>>
>>>>
>>>> I suggest to do actual speed measurements before making bold
>>>> claims like above. Don't trust your intuition!
>>>
>>> I have, more than once, done such measurements after mmap()
>>> was introduced in SVR4 circa 1989 (ported from SunOS).
>>>
>>> On a single-user system, running a single job, the difference
>>> for smaller files is in the noise. For larger files, or when
>>> the system is heavily loaded or multiuser, it can be significant.
>>
>> 1989 is 36 years ago. Technology has moved on. If reading your file is too slow to read, get yourself a real computer.
>>
>> On my very ordinary desktop machine, I just freq'd[1] a 7,032,963,565- byte file in 12.256 seconds. That's 573,838,410 bytes per second. It's a damn sight faster than I could do by hand.
>>
>> How, exactly, are you using `slow'?
>>
>
> A getc loop took 4.3 seconds to read a 192MB file from SSD, on my Windows PC.
>
> Under WSL it took 8.4 seconds (8.4/0.5 real/user).
>
> However reading it all in one go took 0.14 seconds.
>
> I guess not all 'getc' implementations are the same.
#include <stdio.h>
#include <stdlib.h>
#include <windows.h>
/* gcc -Wl,--stack,1200000000 -o getcbench.exe getcbench.c */
int main(int argc, char **argv)
{ FILE* source;
int c; /* getc holder */
const int size = 1000*1000*1000;
char keep[size];
int i=0;
printf( "\nWelcome to getcbench.exe\n\n" );
__int64 time1 = 0, time2 = 0, freq = 0; /* code added for timestamp */
if (argc != 2) {
fprintf(stderr, "Usage: %s source_file\n", argv[0]);
return -1;
}
printf( "Array ready, opening file %s\n", argv[1] );
source = fopen(argv[1], "rb");
if (!source) {
fprintf(stderr, "Could not open %s\n", argv[1]);
return -1;
}
QueryPerformanceCounter((LARGE_INTEGER *) &time1); /* clock is running */
QueryPerformanceFrequency((LARGE_INTEGER *)&freq);
printf("time1 = %llX freq = %lld \n", time1, freq);
while ((c = getc(source)) != EOF) {
keep[i++] = c;
if (i >= size) break;
}
QueryPerformanceCounter((LARGE_INTEGER *) &time2);
printf("time2 = %llX \n", time2);
printf("Read %d bytes in %010.6f seconds\n", i, (float)(time2-time1)/freq);
}
$ getcbench.exe D:\test.txt # D: is capable of gigabytes per second speeds
Welcome to getcbench.exe
Array ready, opening file D:test.txt
time1 = 3380876B31 freq = 10000000
time2 = 338D011DCC
Read 1000000000 bytes in 020.930217 seconds # Process Monitor shows that 4096 byte reads are being done
$
***************************************************************
This has additional gubbins.
https://en.cppreference.com/w/c/io/setvbuf
Add some code after the fopen.
if (setvbuf(source, NULL, _IOFBF, 65536) != 0)
{
fprintf(stderr, "setvbuf() failed\n\n" );
return -1;
}
Process Monitor shows the reads now happen in 65536 chunks.
But this does not do a thing for performance (with this style of I/O and no optimization).
$ getcbenchbuf.exe D:\test.txt
Welcome to getcbenchbuf.exe
Array ready, opening file D:test.txt
time1 = 37192A7827 freq = 10000000
time2 = 37256FEAFA
Read 1000000000 bytes in 020.587797 seconds
***************************************************************
If I do this to the original program (-O2), it still is
doing 4096 byte reads, but the performance is better.
$ gcc -O2 -Wl,--stack,1200000000 -o getcbench.exe getcbench.c
$ getcbench.exe D:\\test2.txt
Welcome to getcbench.exe
Array ready, opening file D:\test2.txt
time1 = 3B4D7C1022 freq = 10000000
time2 = 3B4E5EB775
Read 1000000000 bytes in 001.485397 seconds
Busy sum = FFFFFFFFE216FE9C
Extra code was added so keep[] was not optimized away.
for (k = 0; k<i; k++) sum += keep[k];
printf("Busy sum = %llX\n", sum);
That's about 673MB/sec.
The version with the setvbuf, is still reading 65536 byte chunks.
$ gcc -O2 -Wl,--stack,1200000000 -o getcbenchbuf.exe getcbenchbuf.c
$ getcbenchbuf.exe D:\\test2.txt
Welcome to getcbenchbuf.exe
Array ready, opening file D:\test2.txt
time1 = 3C1EA5ACDF freq = 10000000
time2 = 3C1F49CE7D
Read 1000000000 bytes in 001.075651 seconds
Busy sum = FFFFFFFFE216FE9C
That's getting close to a gigabyte per second.
Summary: The -O2 makes a BIG difference.
No idea how it is cheating.
Paul
[toc] | [prev] | [next] | [standalone]
| From | bart <bc@freeuk.com> |
|---|---|
| Date | 2025-12-11 11:46 +0000 |
| Message-ID | <10heau9$25sj9$1@dont-email.me> |
| In reply to | #395780 |
On 11/12/2025 03:35, Paul wrote:
> On Wed, 12/10/2025 5:37 PM, bart wrote:
>> A getc loop took 4.3 seconds to read a 192MB file from SSD, on my Windows PC.
>>
>> Under WSL it took 8.4 seconds (8.4/0.5 real/user).
>>
>> However reading it all in one go took 0.14 seconds.
>>
>> I guess not all 'getc' implementations are the same.
>
> #include <stdio.h>
> #include <stdlib.h>
> #include <windows.h>
>
> /* gcc -Wl,--stack,1200000000 -o getcbench.exe getcbench.c */
>
> int main(int argc, char **argv)
> { FILE* source;
>
> int c; /* getc holder */
> const int size = 1000*1000*1000;
> char keep[size];
I didn't see the point of either keeping the array on the stack, or
using a VLA. I made it static. That also allowed me a choice of
compilers with no special options needed.
> Add some code after the fopen.
>
> if (setvbuf(source, NULL, _IOFBF, 65536) != 0)
> {
> fprintf(stderr, "setvbuf() failed\n\n" );
> return -1;
> }
When I added that, it slowed it down! Maybe it was already using a
bigger buffer.
> Extra code was added so keep[] was not optimized away.
My loop didn't store the characters anywhere; it just bumped a count.
I think it was enough that it was calling an external function, 'getc';
a commpiler can't optimise that away.
> Read 1000000000 bytes in 001.075651 seconds
>
> Busy sum = FFFFFFFFE216FE9C
>
> That's getting close to a gigabyte per second.
>
> Summary: The -O2 makes a BIG difference.
> No idea how it is cheating.
How a look at the generated assembly: is it still making an actual call
to 'getc', or has it been inlined?
In my case -O2 made little difference, and it was still calling getc().
-O2 can't effect such a precompiled function, unless getc() is not
really an external function: either a macro, or a wrapper.
Also, the generated EXE file actually imports getc from msvcrt.dll,
which is a library not known to be performant.
[toc] | [prev] | [next] | [standalone]
| From | Bonita Montero <Bonita.Montero@gmail.com> |
|---|---|
| Date | 2025-12-11 12:53 +0100 |
| Message-ID | <10heb9p$26b6v$1@raubtier-asyl.eternal-september.org> |
| In reply to | #395781 |
Am 11.12.2025 um 12:46 schrieb bart:
> On 11/12/2025 03:35, Paul wrote:
>> On Wed, 12/10/2025 5:37 PM, bart wrote:
>
>>> A getc loop took 4.3 seconds to read a 192MB file from SSD, on my
>>> Windows PC.
>>>
>>> Under WSL it took 8.4 seconds (8.4/0.5 real/user).
>>>
>>> However reading it all in one go took 0.14 seconds.
>>>
>>> I guess not all 'getc' implementations are the same.
>>
>> #include <stdio.h>
>> #include <stdlib.h>
>> #include <windows.h>
>>
>> /* gcc -Wl,--stack,1200000000 -o getcbench.exe getcbench.c */
>>
>> int main(int argc, char **argv)
>> { FILE* source;
>>
>> int c; /* getc holder */
>> const int size = 1000*1000*1000;
>> char keep[size];
>
> I didn't see the point of either keeping the array on the stack, or
> using a VLA. I made it static. That also allowed me a choice of
> compilers with no special options needed.
Yes. Under Linux/x64 the default stack size is 8MiB, unter Windows/x64
one MiB.
That's a stack overflow - or should I call it underflow since it grows
downards
- for sure.
>
>> Add some code after the fopen.
>>
>> if (setvbuf(source, NULL, _IOFBF, 65536) != 0)
>> {
>> fprintf(stderr, "setvbuf() failed\n\n" );
>> return -1;
>> }
>
> When I added that, it slowed it down! Maybe it was already using a
> bigger buffer.
>
>> Extra code was added so keep[] was not optimized away.
>
> My loop didn't store the characters anywhere; it just bumped a count.
>
> I think it was enough that it was calling an external function,
> 'getc'; a commpiler can't optimise that away.
>
>> Read 1000000000 bytes in 001.075651 seconds
>>
>> Busy sum = FFFFFFFFE216FE9C
>>
>> That's getting close to a gigabyte per second.
>>
>> Summary: The -O2 makes a BIG difference.
>> No idea how it is cheating.
>
> How a look at the generated assembly: is it still making an actual
> call to 'getc', or has it been inlined?
>
> In my case -O2 made little difference, and it was still calling
> getc(). -O2 can't effect such a precompiled function, unless getc() is
> not really an external function: either a macro, or a wrapper.
>
> Also, the generated EXE file actually imports getc from msvcrt.dll,
> which is a library not known to be performant.
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 18:42 +0000 |
| Message-ID | <10hcev7$1ledq$2@dont-email.me> |
| In reply to | #395759 |
On Wed, 10 Dec 2025 15:07:30 GMT, Scott Lurndal wrote: > The problem with all of your attempts is the performance > issue. Success requires reading every single byte of the > file, one byte at a time. The word 'slow' is not sufficient > to describe how bad the performance will be for a very large > file. > > At a minimum, dump the stdio double-buffered byte-by-byte > algorithm and use mmap(). > > In reality, I still don't see any benefit to this type of > heuristic-based approach. Yeah agreed, its one of those things... -- :wq Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | Lew Pitcher <lew.pitcher@digitalfreehold.ca> |
|---|---|
| Date | 2025-12-10 15:58 +0000 |
| Message-ID | <10hc5bh$1jj9p$1@dont-email.me> |
| In reply to | #395756 |
On Wed, 10 Dec 2025 11:35:48 +0000, Michael Sanders wrote:
> On Sat, 6 Dec 2025 02:00:22 -0000 (UTC), Lew Pitcher wrote:
>
>> I should have added that I feel that you probably haven't really
>> defined /what/ "text file" means, and that has interfered with
>> the development of this function. As Keith pointed out, the task
>> of distinguishing between a "text" file and a "binary" file is not
>> easy. I'll add that a lot of the difficulty stems from the fact
>> that there are many definitions (some conflicting) of what a "text"
>> file actually contains.
>
> Yes. Here's my 2nd attempt following the template (of thinking)
> you've suggested...
FWIW, my opinion doesn't matter in the measure of whether or not you have
written a competent is_text_file() function; what matters is that it
fits (or does not fit) the use-case you wrote it for. If it were me,
I'd have a hard time writing this function, because I don't know your
use-case, and I'd try to generalize it. I've worked with text files
stored in ASCII, and in EBCDIC, and in various Unicode formats, and
(god help me) in a bunch of other formats as well, and I'd have a hard
time generalizing all that into a universal is_text_file() function.
So, my real advice is to pick your battles, and document exactly what
sort of text file you intend to look for with this function. What
you've wrote might suit your needs exactly, without accounting for
all the variations of what a text file consists of.
> #include <stdio.h> // FILE, fopen, fread, fclose
> #include <stddef.h> // size_t
>
> // is_text_file()
> // Returns:
> // -1 : could not open file
> // 0 : is NOT a text file (binary indicators found)
> // 1 : is PROBABLY a text file (no strong binary signatures)
>
> int is_text_file(const char *path) {
> // Try opening the file in binary mode,
> // required so that bytes are read exact.
> FILE *f = fopen(path, "rb");
> if (!f) return -1; // Could not open file
>
> unsigned char buf[4096]; // 4KB chunks
> size_t n, i;
>
> // Read in file until EOF
> while ((n = fread(buf, 1, sizeof(buf), f)) > 0) {
> for (i = 0; i < n; i++) {
> unsigned char c = buf[i];
>
> // 1. null byte is a very strong indication of binary data.
> // Text files virtually never contain 0x00.
Except for UTF16 and UTF32 text files, of course.
So, part of your definition of what constitutes a text file is that
a text file (at least as far as is_text_file() is concerned) does not
contain any UTF16 or UTF32 characters.
> if (c == 0x00) {
> fclose(f);
> return 0; // Contains binary-only byte: NOT text
> }
>
> // 2. Check for raw C0 control codes (0x01–0x1F).
> // We *allow* \t (09), \n (0A), \r (0D) because they are normal in text.
> // Any other control code is highly suspicious and usually means binary.
> if (c < 0x20) {
> if (c != 0x09 && c != 0x0A && c != 0x0D) {
Except for all the flavours of EBCDIC.
So, another part of your definition of what constitutes a text file is that
a text file (at least as far as is_text_file() is concerned) does not contain EBCDIC
> fclose(f);
> return 0; // unexpected control character → NOT text
> }
> }
>
> // 3. NOTE: We intentionally do *not* reject bytes >= 0x80.
> // These occur in UTF-8, extended ASCII, and many local encodings.
> // Rejecting them would treat valid multilingual text as binary.
> // So we treat high bytes as acceptable for "probably text".
Except for ASCII, which is limited to 7bit characters between 0x00 and 0x7f
(ignoring, of course, those text files that store ASCII with even or odd parity)
So, another part of your definition of what constitutes a text file is that
a text file (at least as far as is_text_file() is concerned) may contain
ASCII, but is not guaranteed to do so.
> }
> }
>
> fclose(f);
> return 1; // Probably text (no strong binary signatures found)
> }
--
Lew Pitcher
"In Skills We Trust"
Not LLM output - I'm just like this.
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 18:44 +0000 |
| Message-ID | <10hcf1v$1ledq$3@dont-email.me> |
| In reply to | #395760 |
On Wed, 10 Dec 2025 15:58:41 -0000 (UTC), Lew Pitcher wrote: > [...] Thanks Lew. I'm stumped, but learned allot. -- :wq Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | James Kuyper <jameskuyper@alumni.caltech.edu> |
|---|---|
| Date | 2025-12-10 12:46 -0500 |
| Message-ID | <10hcbls$1b254$1@dont-email.me> |
| In reply to | #395756 |
On 2025-12-10 06:35, Michael Sanders wrote:
...
> #include <stdio.h> // FILE, fopen, fread, fclose
> #include <stddef.h> // size_t
>
> // is_text_file()
> // Returns:
> // -1 : could not open file
> // 0 : is NOT a text file (binary indicators found)
> // 1 : is PROBABLY a text file (no strong binary signatures)
>
> int is_text_file(const char *path) {
> // Try opening the file in binary mode,
> // required so that bytes are read exact.
> FILE *f = fopen(path, "rb");
> if (!f) return -1; // Could not open file
>
> unsigned char buf[4096]; // 4KB chunks
> size_t n, i;
>
> // Read in file until EOF
> while ((n = fread(buf, 1, sizeof(buf), f)) > 0) {
> for (i = 0; i < n; i++) {
> unsigned char c = buf[i];
I'd recommend against buffering this; C stdio is already buffered, and
it just complicates your code to keep track of a second level of
buffering. Use getc() instead.
> // 1. null byte is a very strong indication of binary data.
> // Text files virtually never contain 0x00.
> if (c == 0x00) {
> fclose(f);
> return 0; // Contains binary-only byte: NOT text
> }
>
> // 2. Check for raw C0 control codes (0x01–0x1F).
> // We *allow* \t (09), \n (0A), \r (0D) because they are normal in text.
> // Any other control code is highly suspicious and usually means binary.
> if (c < 0x20) {
> if (c != 0x09 && c != 0x0A && c != 0x0D) {
> fclose(f);
> return 0; // unexpected control character → NOT text
> }
> }
I would recommend against use of explicit numerical codes for
characters. They make your code dependent upon a particular encoding,
and you're free to make that choice, but for implementations where that
encoding is the default, the corresponding C escape sequences will have
precisely the the correct value, and make it easier to understand what
your code is doing:
0x00 '\0'
0x09 '\t'
0x0A '\n'
0x0D '\r'
0x20 ' '
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 18:45 +0000 |
| Message-ID | <10hcf4j$1ledq$4@dont-email.me> |
| In reply to | #395764 |
On Wed, 10 Dec 2025 12:46:36 -0500, James Kuyper wrote: > I would recommend against use of explicit numerical codes for > characters. They make your code dependent upon a particular encoding, > and you're free to make that choice, but for implementations where that > encoding is the default, the corresponding C escape sequences will have > precisely the the correct value, and make it easier to understand what > your code is doing: > > 0x00 '\0' > 0x09 '\t' > 0x0A '\n' > 0x0D '\r' > 0x20 ' ' Aye, moving towards that (eventually). Thanks for your comments James. -- :wq Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 18:41 +0000 |
| Message-ID | <10hcesh$1ledq$1@dont-email.me> |
| In reply to | #395756 |
On Wed, 10 Dec 2025 11:35:48 -0000 (UTC), Michael Sanders wrote:
> Yes. Here's my 2nd attempt...
>
> [...]
Last version for me (I have to pivot to other things).
Main change is a look up table, ought to provide
optional future extensibility...
Earnest thanks to each & all =)
#include <stdio.h> // FILE, fopen, fread, fclose
#include <stddef.h> // size_t
// is_text_file()
// Returns:
// -1 : could not open file
// 0 : is NOT a text file (binary indicators found)
// 1 : is PROBABLY a text file (no strong binary signatures)
int is_text_file(const char *path) {
FILE *f = fopen(path, "rb");
if (!f) return -1;
unsigned char chunk[4096]; // 4KB
size_t n, i;
// Look Up Table: 1 = allowed in text, 0 = binary indicator
// Allows TAB(0x09), LF(0x0A), CR(0x0D), printable ASCII (0x20–0x7E)
static const unsigned char LUT[128] = {
0,0,0,0,0,0,0,0,0,1,1,0,0,1,0,0, // 0x00–0x0F
0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0, // 0x10–0x1F
1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, // 0x20–0x2F
1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, // 0x30–0x3F
1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, // 0x40–0x4F
1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, // 0x50–0x5F
1,1,1,1,1,1,1,1,1,1,1,1,1,1,1,1, // 0x60–0x6F
1,1,1,1,1,1,1,1,1,1,1,0 // 0x70–0x7F, last 0 = DEL
};
while ((n = fread(chunk, 1, sizeof(chunk), f)) > 0) {
for (i = 0; i < n; i++) {
if (chunk[i] < 128 && !LUT[chunk[i]]) {
fclose(f);
return 0; // binary indicator found
}
// bytes >= 128 are accepted as probably text
}
}
fclose(f);
return 1; // probably text
}
--
:wq
Mike Sanders
[toc] | [prev] | [next] | [standalone]
| From | Michael Sanders <porkchop@invalid.foo> |
|---|---|
| Date | 2025-12-10 20:57 +0000 |
| Message-ID | <10hcmsd$1p4lc$1@dont-email.me> |
| In reply to | #395766 |
On Wed, 10 Dec 2025 18:41:22 -0000 (UTC), Michael Sanders wrote:
> Last version for me (I have to pivot to other things).
>
> [...]
smaller look up table still + bit shifting!
*fastest implantation yet* but virtually unreadable =(
#include <stdio.h>
#include <stddef.h>
#include <stdint.h>
// is_text_file()
// Returns:
// -1 : could not open file
// 0 : is NOT a text file (binary indicators found)
// 1 : is PROBABLY a text file (no strong binary signatures)
int is_text_file(const char *path) {
FILE *f = fopen(path, "rb");
if (!f) return -1;
unsigned char chunk[4096];
size_t n, i;
// 128-bit bitmask (16 bytes × 8 bits / byte), 1=allowed, 0=disallowed
// Allowed bytes: TAB(0x09), LF(0x0A), CR(0x0D), printable ASCII 0x20–0x7E
static const uint8_t MASK[16] = {
0x00, 0x24, 0x00, 0x00, // 0x00–0x0F: TAB(09), LF(0A), CR(0D)
0xFF, 0xFF, 0xFF, 0xFF, // 0x10–0x2F: SPC!"#$%&'()*+,-./
0xFF, 0xFF, 0xFF, 0xFF, // 0x30–0x4F: 0123456789:;<=>?@
0xFF, 0xFF, 0xFF, 0x7F // 0x50–0x7F: ABCDEFGHIJKLMNOPQRSTUVWXYZ[\]^_`abcdef...
};
while ((n = fread(chunk, 1, sizeof(chunk), f)) > 0) {
for (i = 0; i < n; i++) {
if (chunk[i] < 128 && !(MASK[chunk[i] >> 3] & (1 << (chunk[i] & 7)))) {
fclose(f);
return 0; // binary indicator found
}
// bytes >= 128 are accepted as probably text
}
}
fclose(f);
return 1; // probably text
}
--
:wq
Mike Sanders
[toc] | [prev] | [next] | [standalone]
Page 1 of 6 [1] 2 3 4 5 6 Next page →
Back to top | Article view | comp.lang.c
csiph-web