Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c > #400708 > unrolled thread
| Started by | gazelle@shell.xmission.com (Kenny McCormack) |
|---|---|
| First post | 2026-08-02 14:17 +0000 |
| Last post | 2026-08-06 15:33 -0700 |
| Articles | 13 on this page of 133 — 18 participants |
Back to article view | Back to comp.lang.c
Default signedness of 'plain' char. gazelle@shell.xmission.com (Kenny McCormack) - 2026-08-02 14:17 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-03 02:45 +0800
Re: Default signedness of 'plain' char. gazelle@shell.xmission.com (Kenny McCormack) - 2026-08-03 01:47 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-03 23:14 +0800
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-03 17:04 +0100
The Spanish Inquisition (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 00:23 +0800
Re: The Spanish Inquisition bart <bc@freeuk.com> - 2026-08-03 18:03 +0100
Re: The Spanish Inquisition Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 01:26 +0800
Re: Default signedness of 'plain' char. Theo <theom+news@chiark.greenend.org.uk> - 2026-08-04 13:13 +0100
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:30 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:18 +0800
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-05 17:20 +0100
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:57 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 15:29 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-03 09:56 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-03 13:56 -0500
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-03 13:48 +0000
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-03 14:47 +0000
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-03 14:47 +0000
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-03 15:04 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-03 16:58 -0700
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-04 15:10 +0000
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-04 15:50 +0000
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 03:00 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-05 09:01 +0200
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-04 18:10 +0000
Compilers targetting the C64 (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 03:17 +0800
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-04 16:25 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 02:57 +0000
Re: Default signedness of 'plain' char. Lynn McGuire <lynnmcguire5@gmail.com> - 2026-08-05 00:19 -0500
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:57 +0800
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-05 14:33 +0000
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-05 22:38 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-04 21:36 -0700
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-06 00:21 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 13:20 -0700
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-03 22:06 +0100
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 14:25 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-03 18:38 -0500
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 19:34 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-05 04:40 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:22 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-05 04:02 -0500
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-05 12:35 +0200
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:49 +0800
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:39 -0700
Dear Chris (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-06 20:15 +0800
Re: Dear Chris bart <bc@freeuk.com> - 2026-08-06 15:05 +0100
Re: Dear Chris Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-06 23:16 +0800
Re: Dear Chris David Brown <david.brown@hesbynett.no> - 2026-08-06 20:10 +0200
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:54 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-06 04:49 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-07 00:10 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-07 09:56 +0200
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-05 13:33 -0400
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:50 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-06 10:34 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-09 15:19 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-09 20:33 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-09 18:04 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-10 02:45 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-10 02:24 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-09 18:03 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-10 08:58 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 15:17 -0500
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-10 15:19 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 17:31 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-10 23:48 +0000
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-10 20:03 -0400
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 20:44 -0500
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-11 11:17 -0400
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-11 14:54 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-12 03:36 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-12 01:24 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-12 07:47 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 12:52 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 12:56 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 02:16 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 21:04 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-12 23:07 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 05:00 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 04:37 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-13 14:36 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 23:07 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 23:44 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-13 18:50 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 02:41 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 23:47 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 04:59 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 13:47 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-13 23:50 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 05:59 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-14 03:16 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 08:32 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 12:36 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-15 03:16 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 22:38 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-16 00:08 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 12:17 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-11 13:46 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 13:49 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-11 08:58 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-11 04:03 -0500
Re: Default signedness of 'plain' char. steve g <Sgonedes1977@gmail.com> - 2026-08-10 20:29 -0400
Re: Vectors! Quaternions! (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-11 03:58 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-11 09:06 +0200
Re: Default signedness of 'plain' char. Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-10 21:15 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-11 04:52 +0000
Re: Default signedness of 'plain' char. Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-11 06:45 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 14:01 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:36 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:18 +0000
Re: Default signedness of 'plain' char. Lynn McGuire <lynnmcguire5@gmail.com> - 2026-08-05 18:30 -0500
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 16:41 -0700
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-06 01:24 +0100
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 17:46 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-06 22:59 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 19:54 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 11:02 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 10:35 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 18:32 +0000
Re: Default signedness of 'plain' char. Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 20:55 +0200
Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 20:51 +0200
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 13:03 -0700
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 13:32 -0700
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 21:44 +0000
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 23:47 +0200
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 15:26 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-06 22:57 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 18:23 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 11:47 +0000
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-06 14:40 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 15:33 -0700
Page 7 of 7 — ← Prev page 1 2 3 4 5 6 [7]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2026-08-07 18:32 +0000 |
| Message-ID | <11558bh$t9g$1@reader1.panix.com> |
| In reply to | #400918 |
In article <115550j$rpsl$1@kst.eternal-september.org>, Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >cross@spitfire.i.gajendra.net (Dan Cross) writes: >> In article <1153hcv$ba35$1@kst.eternal-september.org>, >> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >>>cross@spitfire.i.gajendra.net (Dan Cross) writes: >>>[...] >>>> Cast the value when using as the index: >>>> >>>> while (*s) ++counts[(unsignd char)*s++]; >>>> >>>> Note that this is already required for the `is*` functions >>>> defined in `ctype.h` (except, IIRC, `isascii`). >>> >>>isascii() is not defined by ISO C, or even by POSIX. >> >> Ah, right you are. `isascii` was marked obsolescent in POSIX >> Issue 7 (2018) and removed in Issue 8 (2024); it never made it >> into standardized C, and was dropped during the initial work >> leading up to ANSI C (the C89 rationale discusses it), though it >> remains broadly implemented, presumably for compatibility with >> older code. >> >>>isblank() is >>>another common extension. For implementations that support them, >>>I don't think they're handled differently from the other is*() >>>functions. Most implementations handle values from SCHAR_MIN to >>>UCHAR_MAX without error, but the behavior is undefined for any >>>value other than EOF outside the range 0..UCHAR_MAX. >> >> Before removal from POSIX, `isascii` was specified as defined >> for all integer values. >> https://pubs.opengroup.org/onlinepubs/9699919799/functions/isascii.html >> >> I imagine that `isblank` would be implemented similarly to the >> others, however. > >Interesting. The 2018 POSIX specification says that "The isascii() >function is defined on all integer values.", but for isblank() and >isdigit() it says "The c argument is an int, the value of which the >application shall ensure is a character representable as an unsigned >char or equal to the value of the macro EOF. If the argument has any >other value, the behavior is undefined.". (I presume the same applies >to the other ISO-C-defined functions, but I haven't checked them all.) I think that's right. I have a vague memory of lore that said one should write, `if (isascii(c) && iswhatever(c))` (or the equivalent for logically negative tests), though I can no longer remember _where_ I saw that. Of course, once localization is in play, let alone portability to EBCDIC or whatever, it's less relevant if not outright wrong. >isascii() is (was?) a special case, probably because it's easier to >implement without using a lookup table. In glibc: > >#define __isascii(c) (((c) & ~0x7f) == 0) Yes, `isascii` is easy. Not so much for other locales, let alone for Unicode generally or UTF-8. As ASCII specifically has slid from relevance, and localization led to widespread adoption of UTF-8 and other encodings, one can see how it was merely an anachronism to be removed. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | Janis Papanagnou <janis_papanagnou+ng@hotmail.com> |
|---|---|
| Date | 2026-08-07 20:55 +0200 |
| Message-ID | <11559md$3fa9a$2@dont-email.me> |
| In reply to | #400919 |
On 2026-08-07 20:32, Dan Cross wrote: >> [...] > > I think that's right. I have a vague memory of lore that said > one should write, `if (isascii(c) && iswhatever(c))` (or the > equivalent for logically negative tests), though I can no longer > remember _where_ I saw that. Of course, once localization is in > play, let alone portability to EBCDIC or whatever, it's less > relevant if not outright wrong. Same memories here. - And you're not wrong. (See my other reply just a few minutes ago.) Janis >> [..]
[toc] | [prev] | [next] | [standalone]
| From | Janis Papanagnou <janis_papanagnou+ng@hotmail.com> |
|---|---|
| Date | 2026-08-07 20:51 +0200 |
| Subject | Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <11559es$3fa9a$1@dont-email.me> |
| In reply to | #400912 |
On 2026-08-07 13:02, Dan Cross wrote: > In article <1153hcv$ba35$1@kst.eternal-september.org>, > Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >> cross@spitfire.i.gajendra.net (Dan Cross) writes: >> [...] >>> Cast the value when using as the index: >>> >>> while (*s) ++counts[(unsignd char)*s++]; >>> >>> Note that this is already required for the `is*` functions >>> defined in `ctype.h` (except, IIRC, `isascii`). >> >> isascii() is not defined by ISO C, or even by POSIX. > > Ah, right you are. `isascii` was marked obsolescent in POSIX > Issue 7 (2018) and removed in Issue 8 (2024); it never made it > into standardized C, and was dropped during the initial work > leading up to ANSI C (the C89 rationale discusses it), though it > remains broadly implemented, presumably for compatibility with > older code. (A side track about 'isascii'...) I seem to have a faint recollection that isascii() once had been a _necessary_ predicate to make the other ctype.h functions provide a *valid* response [in non-ASCII contexts]. (I thought that I might have got that from K&R, but no, there's no mention of isascii() at all in my copy.) - Though a quick search lead me to a man page that says (e.g. for 'isalpha') about 'isascii': "isalpha is a macro which classifies ASCII integer values by table lookup. It is a predicate returning non-zero when c represents an alphabetic ASCII character, and 0 otherwise. It is defined only when isascii(c) is true or c is EOF." (Memory seems to work.) With I18N and localization obviously just a legacy topic meanwhile. Janis >> [...]
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-07 13:03 -0700 |
| Subject | Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <1155dm4$uui4$1@kst.eternal-september.org> |
| In reply to | #400920 |
Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
> On 2026-08-07 13:02, Dan Cross wrote:
>> In article <1153hcv$ba35$1@kst.eternal-september.org>,
>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>> cross@spitfire.i.gajendra.net (Dan Cross) writes:
>>> [...]
>>>> Cast the value when using as the index:
>>>>
>>>> while (*s) ++counts[(unsignd char)*s++];
>>>>
>>>> Note that this is already required for the `is*` functions
>>>> defined in `ctype.h` (except, IIRC, `isascii`).
>>>
>>> isascii() is not defined by ISO C, or even by POSIX.
>> Ah, right you are. `isascii` was marked obsolescent in POSIX
>> Issue 7 (2018) and removed in Issue 8 (2024); it never made it
>> into standardized C, and was dropped during the initial work
>> leading up to ANSI C (the C89 rationale discusses it), though it
>> remains broadly implemented, presumably for compatibility with
>> older code.
>
> (A side track about 'isascii'...)
>
> I seem to have a faint recollection that isascii() once had been a
> _necessary_ predicate to make the other ctype.h functions provide
> a *valid* response [in non-ASCII contexts]. (I thought that I might
> have got that from K&R, but no, there's no mention of isascii() at
> all in my copy.) - Though a quick search lead me to a man page that
> says (e.g. for 'isalpha') about 'isascii':
>
> "isalpha is a macro which classifies ASCII integer values by table
> lookup. It is a predicate returning non-zero when c represents an
> alphabetic ASCII character, and 0 otherwise. It is defined only
> when isascii(c) is true or c is EOF."
>
> (Memory seems to work.)
>
> With I18N and localization obviously just a legacy topic meanwhile.
That must be an old man page. Where did you find it?
isascii() (on systems where it's provided) is true for arguments in the
range 0..127, false for anything else. The above implies that isalpha()
has undefined behavior for values above 127, which contradicts the ISO C
requirement that it's defined for values in the range of unsigned char
(0..255 in almost all implementations).
newlib, the C library implementation used by Cygwin, has similar wording
in its isspace(3) man page:
isspace is a macro which classifies singlebyte charset values by
table lookup. It is a predicate returning non-zero for whitespace
characters, and 0 for other characters. It is defined only when
isascii(c) is true or c is EOF.
In fact isspace() is implemented correctly, returning non-zero
(happens to be 1) for '\t', '\n', '\v', '\f', '\r', and ' ', and
zero for all other values in the range 128..255 and for -1 (EOF).
Apparently the man page hasn't been updated in a long time.
The is*() functions have been defined for EOF and all values from
0 to UCHAR_MAX since C89/C90.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-07 13:32 -0700 |
| Subject | Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <1155fct$uui4$2@kst.eternal-september.org> |
| In reply to | #400922 |
Keith Thompson <Keith.S.Thompson+u@gmail.com> writes:
[...]
> newlib, the C library implementation used by Cygwin, has similar wording
> in its isspace(3) man page:
>
> isspace is a macro which classifies singlebyte charset values by
> table lookup. It is a predicate returning non-zero for whitespace
> characters, and 0 for other characters. It is defined only when
> isascii(c) is true or c is EOF.
>
> In fact isspace() is implemented correctly, returning non-zero
> (happens to be 1) for '\t', '\n', '\v', '\f', '\r', and ' ', and
> zero for all other values in the range 128..255 and for -1 (EOF).
>
> Apparently the man page hasn't been updated in a long time.
> The is*() functions have been defined for EOF and all values from
> 0 to UCHAR_MAX since C89/C90.
This only affects the isspace(3) man page; the other is*(3) man pages
are correct.
I've reported this to the Cygwin mailing list.
https://cygwin.com/pipermail/cygwin/2026-August/259928.html
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2026-08-07 21:44 +0000 |
| Subject | Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <1155jjb$2vj$1@reader1.panix.com> |
| In reply to | #400922 |
In article <1155dm4$uui4$1@kst.eternal-september.org>, Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes: >> [snip] >> I seem to have a faint recollection that isascii() once had been a >> _necessary_ predicate to make the other ctype.h functions provide >> a *valid* response [in non-ASCII contexts]. (I thought that I might >> have got that from K&R, but no, there's no mention of isascii() at >> all in my copy.) - Though a quick search lead me to a man page that >> says (e.g. for 'isalpha') about 'isascii': >> >> "isalpha is a macro which classifies ASCII integer values by table >> lookup. It is a predicate returning non-zero when c represents an >> alphabetic ASCII character, and 0 otherwise. It is defined only >> when isascii(c) is true or c is EOF." >> >> (Memory seems to work.) >> >> With I18N and localization obviously just a legacy topic meanwhile. > >That must be an old man page. Where did you find it? I just looked around my menagerie of old Unix versions, and that language (or similar) was common util through at least 4.3BSD-Tahoe, and retained all the way through 10th Edition Research Unix (which was actually based on ~4.1BSD). >isascii() (on systems where it's provided) is true for arguments in the >range 0..127, false for anything else. The above implies that isalpha() >has undefined behavior for values above 127, which contradicts the ISO C >requirement that it's defined for values in the range of unsigned char >(0..255 in almost all implementations). I think if it is on a system that says one has to use `isascii` prior to one of the other predicates defined in `<ctype.h>`, it is safe to assume it predates standard C. >newlib, the C library implementation used by Cygwin, has similar wording >in its isspace(3) man page: > > isspace is a macro which classifies singlebyte charset values by > table lookup. It is a predicate returning non-zero for whitespace > characters, and 0 for other characters. It is defined only when > isascii(c) is true or c is EOF. > >In fact isspace() is implemented correctly, returning non-zero >(happens to be 1) for '\t', '\n', '\v', '\f', '\r', and ' ', and >zero for all other values in the range 128..255 and for -1 (EOF). > >Apparently the man page hasn't been updated in a long time. >The is*() functions have been defined for EOF and all values from >0 to UCHAR_MAX since C89/C90. I am disappointed, but unsurprised. The art of writing man pages (and troff, for that matter) is quickly becoming a lost art. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | Janis Papanagnou <janis_papanagnou+ng@hotmail.com> |
|---|---|
| Date | 2026-08-07 23:47 +0200 |
| Subject | Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <1155jp2$3fa9a$3@dont-email.me> |
| In reply to | #400922 |
On 2026-08-07 22:03, Keith Thompson wrote: > Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes: >> On 2026-08-07 13:02, Dan Cross wrote: >>> In article <1153hcv$ba35$1@kst.eternal-september.org>, >>> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >>>> cross@spitfire.i.gajendra.net (Dan Cross) writes: >>>> [...] >>>>> Cast the value when using as the index: >>>>> >>>>> while (*s) ++counts[(unsignd char)*s++]; >>>>> >>>>> Note that this is already required for the `is*` functions >>>>> defined in `ctype.h` (except, IIRC, `isascii`). >>>> >>>> isascii() is not defined by ISO C, or even by POSIX. >>> Ah, right you are. `isascii` was marked obsolescent in POSIX >>> Issue 7 (2018) and removed in Issue 8 (2024); it never made it >>> into standardized C, and was dropped during the initial work >>> leading up to ANSI C (the C89 rationale discusses it), though it >>> remains broadly implemented, presumably for compatibility with >>> older code. >> >> (A side track about 'isascii'...) >> >> I seem to have a faint recollection that isascii() once had been a >> _necessary_ predicate to make the other ctype.h functions provide >> a *valid* response [in non-ASCII contexts]. (I thought that I might >> have got that from K&R, but no, there's no mention of isascii() at >> all in my copy.) - Though a quick search lead me to a man page that >> says (e.g. for 'isalpha') about 'isascii': >> >> "isalpha is a macro which classifies ASCII integer values by table >> lookup. It is a predicate returning non-zero when c represents an >> alphabetic ASCII character, and 0 otherwise. It is defined only >> when isascii(c) is true or c is EOF." >> >> (Memory seems to work.) >> >> With I18N and localization obviously just a legacy topic meanwhile. > > That must be an old man page. Yes, likely. - As I've said, that was a legacy thing that I (and as it seems also Dan) remembered. - If I'd have to date that information I'd guess it must have been somewhere around 1985-95 that I've read it in some Unix man page on some of the platforms I used back then.[*] (The quote just backs up our memories as not being pure imaginations.) > Where did you find it? It was the first hit of a Web search. I haven't stored it because it's nowadays meaningless.[**] Janis [*] If someone wants to look up man pages on those platforms, it may be one of UTS (Amdahl), SunOS 4 (Sun), AIX 3.x (IBM), HP-UX 9 (HP). But it (the information) may also stem from another source, but less likely. [**] Wait! I have it still in the cache... - but note that this library was *not* what _we_ used back these days. GNUPro C Library Copyright © 1992-1997 Cygnus Support. https://users.informatik.haw-hamburg.de/~krabat/FH-Labor/gnupro/4_GNUPro_Libraries/a_GNUPro_C_Library/libc.html
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-07 15:26 -0700 |
| Subject | Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) |
| Message-ID | <1155m3b$10uo7$1@kst.eternal-september.org> |
| In reply to | #400925 |
Janis Papanagnou <janis_papanagnou+ng@hotmail.com> writes:
> On 2026-08-07 22:03, Keith Thompson wrote:
[...]
>> That must be an old man page.
>
> Yes, likely. - As I've said, that was a legacy thing that I (and as it
> seems also Dan) remembered. - If I'd have to date that information I'd
> guess it must have been somewhere around 1985-95 that I've read it in
> some Unix man page on some of the platforms I used back then.[*]
>
> (The quote just backs up our memories as not being pure imaginations.)
>
>> Where did you find it?
>
[SNIP]
>
> [**] Wait! I have it still in the cache... - but note that this library
> was *not* what _we_ used back these days.
>
> GNUPro C Library Copyright © 1992-1997 Cygnus Support.
> https://users.informatik.haw-hamburg.de/~krabat/FH-Labor/gnupro/4_GNUPro_Libraries/a_GNUPro_C_Library/libc.html
It seems that Cygwin/Newlib shares some common ancestry with GNUPro.
The history of the newlib git repo shows a fix in 2013 that corrected
this for most of the is*() functions. It looks like isspace() was
simply overlooked.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2026-08-06 22:57 +0000 |
| Message-ID | <11533gk$bt5$1@reader1.panix.com> |
| In reply to | #400883 |
In article <1150hng$3dds0$1@kst.eternal-september.org>, Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >Lynn McGuire <lynnmcguire5@gmail.com> writes: >[...] >> So, you have to check every compiler and every compiler version to see >> what the signedness of char is. That is not good in these days of >> UTF-8. > >Not really. You can usually write code that doesn't care whether >plain char is signed or unsigned. And if it matters, you can check >whether CHAR_MIN==0. > >This evolved from systems like the PDP-11 where character values >ranged from 0 to 127, so the signedness of plain char didn't >matter much. I'd phrase that slightly differently; on systmes like the PDP-11 that used the 7-bit US-ASCII character set (sorry, Europeans), the signedness of `char` was irrelevant for handling character data. The issue arises because early C did not define a generic byte-sized integer type separate from `char`, so `char` got overloaded to serve as a "very small `int`" in lots of places. The situation got somewhat better when `<stdint.h>` was introduced, but by then the die was cast. >It's annoying that, in many implementations, UTF-8 strings can >contain elements with negative values (negative character values >rarely make sense), but in practice it doesn't cause many problems. > >IMHO it would be cleaner to require plain char to be unsigned, >but I don't see that happening. The cleanest thing would be to define `char` to be, specifically, a type designated to hold only character data, with a corresponding `int` type with the usual signed and unsigned variants, and explicit conversion functions to translate between `char` and the underlying representation. But that ain't happenin'; all the reasons you outlined for changing the signedness of `char` among them. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-06 18:23 -0700 |
| Message-ID | <1153c2h$9tat$1@kst.eternal-september.org> |
| In reply to | #400901 |
cross@spitfire.i.gajendra.net (Dan Cross) writes:
> In article <1150hng$3dds0$1@kst.eternal-september.org>,
> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
>>Lynn McGuire <lynnmcguire5@gmail.com> writes:
>>[...]
>>> So, you have to check every compiler and every compiler version to see
>>> what the signedness of char is. That is not good in these days of
>>> UTF-8.
>>
>>Not really. You can usually write code that doesn't care whether
>>plain char is signed or unsigned. And if it matters, you can check
>>whether CHAR_MIN==0.
>>
>>This evolved from systems like the PDP-11 where character values
>>ranged from 0 to 127, so the signedness of plain char didn't
>>matter much.
>
> I'd phrase that slightly differently; on systmes like the PDP-11
> that used the 7-bit US-ASCII character set (sorry, Europeans),
> the signedness of `char` was irrelevant for handling character
> data.
>
> The issue arises because early C did not define a generic
> byte-sized integer type separate from `char`, so `char` got
> overloaded to serve as a "very small `int`" in lots of places.
>
> The situation got somewhat better when `<stdint.h>` was
> introduced, but by then the die was cast.
Agreed, good clarification.
>>It's annoying that, in many implementations, UTF-8 strings can
>>contain elements with negative values (negative character values
>>rarely make sense), but in practice it doesn't cause many problems.
>>
>>IMHO it would be cleaner to require plain char to be unsigned,
>>but I don't see that happening.
>
> The cleanest thing would be to define `char` to be,
> specifically, a type designated to hold only character data,
> with a corresponding `int` type with the usual signed and
> unsigned variants, and explicit conversion functions to
> translate between `char` and the underlying representation.
Other than requiring explicit conversions, that's pretty much what we
have now. signed char and unsigned char are the two standard integer
types, probably narrower than signed short and unsigned short.
char is a special case, "designated to hold only character data",
though that designation is not enforced.
If it were practical, I'd like to see the "char" type either
unsigned, or for its signedness to be irrelevant (i.e., not an
integer type).
(Ada, for example, defines Character as an enumeration type, and
allows character constants as enumeration constants.)
> But that ain't happenin'; all the reasons you outlined for
> changing the signedness of `char` among them.
Yup.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2026-08-07 11:47 +0000 |
| Message-ID | <1154gkf$mnc$1@reader1.panix.com> |
| In reply to | #400905 |
In article <1153c2h$9tat$1@kst.eternal-september.org>, Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >cross@spitfire.i.gajendra.net (Dan Cross) writes: >> In article <1150hng$3dds0$1@kst.eternal-september.org>, >> Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote: >>>[snip] >>>IMHO it would be cleaner to require plain char to be unsigned, >>>but I don't see that happening. >> >> The cleanest thing would be to define `char` to be, >> specifically, a type designated to hold only character data, >> with a corresponding `int` type with the usual signed and >> unsigned variants, and explicit conversion functions to >> translate between `char` and the underlying representation. > >Other than requiring explicit conversions, that's pretty much what we >have now. signed char and unsigned char are the two standard integer >types, probably narrower than signed short and unsigned short. It is close, indeed, though I would argue that the implicit conversions can cause some minor grief. The need to cast arguments to `is*` feels superfluous. In the big scheme of things it's not a huge deal, of course, but still ugly. >char is a special case, "designated to hold only character data", >though that designation is not enforced. I'm not even sure that was the original intent. Sure, `char` was (and is) useful for holding character data, but I believe it was always intended as the byte integer type; it was probably just _most often_ used for representing character data. I think they didn't want to make it separate from other integer types because they felt it was "good enough" and didn't want to add another reserved word to the language (and what would it be?). This business with signed vs unsigned `char` is just historical baggage because for the first decade or so of C's existence, they didn't have to care. >If it were practical, I'd like to see the "char" type either >unsigned, or for its signedness to be irrelevant (i.e., not an >integer type). Agreed. It shouldn't be an integer type. >(Ada, for example, defines Character as an enumeration type, and >allows character constants as enumeration constants.) Even Pascal's `Char` type is distinct from the usual integer types, with `chr` and `ord` operators to convert to and from. Not to beat the Rust drum again, but I think they came up with a fairly nice abstraction: an instance of the `char` type is a multibyte datum designed specifically to hold character data (UNICODE code points, in particular). It is illegal to put anything else into a `char`. Its representation is known to be some primitive integer compatible with `u32`, so it is cheap to copy, pass as an argument to a function, put in a `struct` and so on, but conversion to and from other types is explicit. I wouldn't have expected anyone to do that on a PDP-11/20 in 1972, though. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2026-08-06 14:40 +0000 |
| Message-ID | <M11dS.39949$K4L1.30023@fx12.iad> |
| In reply to | #400882 |
Lynn McGuire <lynnmcguire5@gmail.com> writes:
>On 8/2/2026 9:17 AM, Kenny McCormack wrote:
>> First off, I know the "standards" answer is "Either is correct; you have no
>> right to complain about anything", but I am not interested in the
>> "standards" answer. If this is all you can do, then just click Next and go
>> on.
>>
>> I'm interested in the "why" of why implementations might prefer one or the
>> other.
>>
>> Consider:
>>
>> /* macro 'U' must be defined on the cmd line */
>> #include <stdio.h>
>>
>> int main(void)
>> {
>> U char c = 255;
>>
>> printf("Result of 'c > 0': %d\n",c > 0);
>> }
>>
>> And the following command lines:
>>
>> $ tcc -DU= -run CheckSignedChar.c
>> $ tcc -DU=signed -run CheckSignedChar.c
>> $ tcc -DU=unsigned -run CheckSignedChar.c
>>
>> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the default is
>> signed. I'm interested in what sorts of factors drive the decision-making.
>>
>> Note, BTW, that I first noticed this in a project using gcc, but it is
>> easier to test using tcc, as above.
>>
>> Also, total aside, I'm surprised that one needs to do -DU= instead of just
>> -DU. I thought -DU would define it as an empty string, but that generates
>> a compile error. You need -DU=. Why?
>
>So, you have to check every compiler and every compiler version to see
>what the signedness of char is.
No. It's really simple, just specify which you need (unsigned char or signed char)
directly. Don't rely on unspecified behavior.
Personally, I use uint8_t or int8_t depending on the use case.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-06 15:33 -0700 |
| Message-ID | <1153247$775e$1@kst.eternal-september.org> |
| In reply to | #400893 |
scott@slp53.sl.home (Scott Lurndal) writes:
> Lynn McGuire <lynnmcguire5@gmail.com> writes:
[...]
>>So, you have to check every compiler and every compiler version to see
>>what the signedness of char is.
>
> No. It's really simple, just specify which you need (unsigned char or
> signed char) directly. Don't rely on unspecified behavior.
>
> Personally, I use uint8_t or int8_t depending on the use case.
That can be a good approach in some cases, but char, signed char,
and unsigned char are distinct types, and there are cases where
you have to use plain char. Functions in <string.h> operate on
array of plain char -- and sometimes specify unsigned semantics.
For example, strcmp() operates on arrays of char, but treats them
as unsigned char.
Yeah, it's a bit of a mess.
There are (rare) occasions where you need to know whether plain
char is signed or unsigned. Fortunately, that's easy enough to
determine in code, even in the preprocessor.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [standalone]
Page 7 of 7 — ← Prev page 1 2 3 4 5 6 [7]
Back to top | Article view | comp.lang.c
csiph-web