Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.c > #400708 > unrolled thread
| Started by | gazelle@shell.xmission.com (Kenny McCormack) |
|---|---|
| First post | 2026-08-02 14:17 +0000 |
| Last post | 2026-08-06 15:33 -0700 |
| Articles | 20 on this page of 133 — 18 participants |
Back to article view | Back to comp.lang.c
Default signedness of 'plain' char. gazelle@shell.xmission.com (Kenny McCormack) - 2026-08-02 14:17 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-03 02:45 +0800
Re: Default signedness of 'plain' char. gazelle@shell.xmission.com (Kenny McCormack) - 2026-08-03 01:47 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-03 23:14 +0800
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-03 17:04 +0100
The Spanish Inquisition (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 00:23 +0800
Re: The Spanish Inquisition bart <bc@freeuk.com> - 2026-08-03 18:03 +0100
Re: The Spanish Inquisition Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-04 01:26 +0800
Re: Default signedness of 'plain' char. Theo <theom+news@chiark.greenend.org.uk> - 2026-08-04 13:13 +0100
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:30 +0000
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:18 +0800
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-05 17:20 +0100
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:57 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 15:29 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-03 09:56 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-03 13:56 -0500
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-03 13:48 +0000
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-03 14:47 +0000
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-03 14:47 +0000
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-03 15:04 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-03 16:58 -0700
Re: Default signedness of 'plain' char. Lew Pitcher <lew.pitcher@digitalfreehold.ca> - 2026-08-04 15:10 +0000
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-04 15:50 +0000
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 03:00 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-05 09:01 +0200
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-04 18:10 +0000
Compilers targetting the C64 (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 03:17 +0800
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-04 16:25 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 02:57 +0000
Re: Default signedness of 'plain' char. Lynn McGuire <lynnmcguire5@gmail.com> - 2026-08-05 00:19 -0500
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:57 +0800
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-05 14:33 +0000
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-05 22:38 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-04 21:36 -0700
Re: Default signedness of 'plain' char. antispam@fricas.org (Waldek Hebisch) - 2026-08-06 00:21 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 13:20 -0700
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-03 22:06 +0100
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 14:25 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-03 18:38 -0500
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-03 19:34 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-05 04:40 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:22 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-05 04:02 -0500
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-05 12:35 +0200
Re: Default signedness of 'plain' char. Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-05 18:49 +0800
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:39 -0700
Dear Chris (was: Re: Default signedness of 'plain' char.) Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-06 20:15 +0800
Re: Dear Chris bart <bc@freeuk.com> - 2026-08-06 15:05 +0100
Re: Dear Chris Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> - 2026-08-06 23:16 +0800
Re: Dear Chris David Brown <david.brown@hesbynett.no> - 2026-08-06 20:10 +0200
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:54 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-06 04:49 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-07 00:10 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-07 09:56 +0200
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-05 13:33 -0400
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 21:50 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-06 10:34 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-09 15:19 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-09 20:33 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-09 18:04 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-10 02:45 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-10 02:24 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-09 18:03 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-10 08:58 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 15:17 -0500
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-10 15:19 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 17:31 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-10 23:48 +0000
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-10 20:03 -0400
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-10 20:44 -0500
Re: Default signedness of 'plain' char. James Kuyper <jameskuyper@alumni.caltech.edu> - 2026-08-11 11:17 -0400
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-11 14:54 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-12 03:36 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-12 01:24 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-12 07:47 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 12:52 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 12:56 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 02:16 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-12 21:04 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-12 23:07 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 05:00 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 04:37 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-13 14:36 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 23:07 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-13 23:44 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-13 18:50 -0700
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 02:41 +0000
Re: Quaternions (was Re: Default signedness of 'plain' char.) BGB <cr88192@gmail.com> - 2026-08-13 23:47 -0500
Re: Quaternions (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 04:59 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 13:47 -0700
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-13 23:50 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 05:59 +0000
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-14 03:16 -0500
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-14 08:32 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 12:36 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-15 03:16 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 22:38 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-16 00:08 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-14 12:17 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-11 13:46 +0000
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 13:49 -0700
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-11 08:58 +0200
Re: Default signedness of 'plain' char. BGB <cr88192@gmail.com> - 2026-08-11 04:03 -0500
Re: Default signedness of 'plain' char. steve g <Sgonedes1977@gmail.com> - 2026-08-10 20:29 -0400
Re: Vectors! Quaternions! (was Re: Default signedness of 'plain' char.) Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-11 03:58 +0000
Re: Default signedness of 'plain' char. David Brown <david.brown@hesbynett.no> - 2026-08-11 09:06 +0200
Re: Default signedness of 'plain' char. Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-10 21:15 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-11 04:52 +0000
Re: Default signedness of 'plain' char. Ross Finlayson <ross.a.finlayson@gmail.com> - 2026-08-11 06:45 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-11 14:01 -0700
Re: Default signedness of 'plain' char. "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> - 2026-08-05 12:36 -0700
Re: Default signedness of 'plain' char. Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-05 07:18 +0000
Re: Default signedness of 'plain' char. Lynn McGuire <lynnmcguire5@gmail.com> - 2026-08-05 18:30 -0500
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 16:41 -0700
Re: Default signedness of 'plain' char. bart <bc@freeuk.com> - 2026-08-06 01:24 +0100
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-05 17:46 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-06 22:59 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 19:54 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 11:02 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 10:35 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 18:32 +0000
Re: Default signedness of 'plain' char. Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 20:55 +0200
Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 20:51 +0200
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 13:03 -0700
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 13:32 -0700
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 21:44 +0000
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Janis Papanagnou <janis_papanagnou+ng@hotmail.com> - 2026-08-07 23:47 +0200
Re: Use of isascii() back in early days (was Re: Default signedness of 'plain' char.) Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-07 15:26 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-06 22:57 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 18:23 -0700
Re: Default signedness of 'plain' char. cross@spitfire.i.gajendra.net (Dan Cross) - 2026-08-07 11:47 +0000
Re: Default signedness of 'plain' char. scott@slp53.sl.home (Scott Lurndal) - 2026-08-06 14:40 +0000
Re: Default signedness of 'plain' char. Keith Thompson <Keith.S.Thompson+u@gmail.com> - 2026-08-06 15:33 -0700
Page 2 of 7 — ← Prev page 1 [2] 3 4 5 6 7 Next page →
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-03 16:58 -0700 |
| Message-ID | <114r9vj$1o8k3$1@kst.eternal-september.org> |
| In reply to | #400748 |
Lew Pitcher <lew.pitcher@digitalfreehold.ca> writes:
[...]
> For what it's worth, this was also the reason (prior to Unicode)
> that C did not specify that alphabetic characters would have a
> contiguous sequence in the execution characterset. In EBCDIC,
> the alphabetics group a-i, j-r, s-z and A-I, J-R, S-Z, with
> various other characters (both assigned and unassigned) between
> the groupings.
C still doesn't require Unicode (well, mostly), and still doesn't
require 'i'+1=='j'. C does have UTF-8 string literals, such as
u8"hello", which are encoded as UTF-8, but ordinary string literals
like "hello" are still encoded using the execution character set,
which could be EBCDIC.
There's a proposal to require 'a'..'f' and 'A'..'F' to be contiguous,
but it hasn't appeared in the latest C2y draft.
https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3192.pdf
I suppose that an implementation whose execution character set is
some version of EBCDIC would have to treat "hello" and u8"hello"
very differently.
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | Lew Pitcher <lew.pitcher@digitalfreehold.ca> |
|---|---|
| Date | 2026-08-04 15:10 +0000 |
| Message-ID | <114svdu$289pl$1@dont-email.me> |
| In reply to | #400790 |
On Mon, 03 Aug 2026 16:58:39 -0700, Keith Thompson wrote: > Lew Pitcher <lew.pitcher@digitalfreehold.ca> writes: > [...] >> For what it's worth, this was also the reason (prior to Unicode) >> that C did not specify that alphabetic characters would have a >> contiguous sequence in the execution characterset. In EBCDIC, >> the alphabetics group a-i, j-r, s-z and A-I, J-R, S-Z, with >> various other characters (both assigned and unassigned) between >> the groupings. > > C still doesn't require Unicode (well, mostly), Yes. I mentioned Unicode because it both simplifies /and/ complicates the matter of alphabetic value contiguity; While it ensures that contiguity within an alphabet, it does not ensure contiguity between alphabets (not that I think it should), leading to the same problem that EBCDIC presented in the first place. > and still doesn't require 'i'+1=='j'. I pointed that out because it has become a common programmer misconception; hand rolled isalpha-like functions often express themselves with a range check in the form of is_lower_case = ((some_char >= 'a') && (some_char <= 'z')); is_upper_case = ((some_char >= 'A') && (some_char <= 'Z')); The standard explicitly requires /numeric/ characters to have contiguity, though. is_number = ((some_char >= '0') && (some_char <= '9')); > C does have UTF-8 string literals, such as > u8"hello", which are encoded as UTF-8, but ordinary string literals > like "hello" are still encoded using the execution character set, > which could be EBCDIC. > > There's a proposal to require 'a'..'f' and 'A'..'F' to be contiguous, > but it hasn't appeared in the latest C2y draft. > > https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3192.pdf That's going to complicate the mainframe C compilers a bit :-) Or, as a TV presenter often put it... "Oh no! Anyway ..." > I suppose that an implementation whose execution character set is > some version of EBCDIC would have to treat "hello" and u8"hello" > very differently. -- Lew Pitcher "In Skills We Trust" Not LLM output - I'm just like this.
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2026-08-04 15:50 +0000 |
| Message-ID | <zTncS.5177$bZ8b.3401@fx01.iad> |
| In reply to | #400820 |
Lew Pitcher <lew.pitcher@digitalfreehold.ca> writes: >On Mon, 03 Aug 2026 16:58:39 -0700, Keith Thompson wrote: > >> Lew Pitcher <lew.pitcher@digitalfreehold.ca> writes: >> [...] >>> For what it's worth, this was also the reason (prior to Unicode) >>> that C did not specify that alphabetic characters would have a >>> contiguous sequence in the execution characterset. In EBCDIC, >>> the alphabetics group a-i, j-r, s-z and A-I, J-R, S-Z, with >>> various other characters (both assigned and unassigned) between >>> the groupings. >> >> C still doesn't require Unicode (well, mostly), > >Yes. I mentioned Unicode because it both simplifies /and/ complicates >the matter of alphabetic value contiguity; While it ensures that contiguity >within an alphabet, it does not ensure contiguity between alphabets (not >that I think it should), leading to the same problem that EBCDIC presented >in the first place. > >> and still doesn't require 'i'+1=='j'. > >I pointed that out because it has become a common programmer misconception; >hand rolled isalpha-like functions often express themselves with a range >check in the form of > is_lower_case = ((some_char >= 'a') && (some_char <= 'z')); > is_upper_case = ((some_char >= 'A') && (some_char <= 'Z')); > >The standard explicitly requires /numeric/ characters to have contiguity, >though. > is_number = ((some_char >= '0') && (some_char <= '9')); > > >> C does have UTF-8 string literals, such as >> u8"hello", which are encoded as UTF-8, but ordinary string literals >> like "hello" are still encoded using the execution character set, >> which could be EBCDIC. >> >> There's a proposal to require 'a'..'f' and 'A'..'F' to be contiguous, >> but it hasn't appeared in the latest C2y draft. >> >> https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3192.pdf > >That's going to complicate the mainframe C compilers a bit :-) Not really. Even in EBCDIC, the encodings for both upper and lower-case A-F are contiguous, as required by n3192.
[toc] | [prev] | [next] | [standalone]
| From | Lawrence D’Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2026-08-05 03:00 +0000 |
| Message-ID | <114u90m$2l3pd$3@dont-email.me> |
| In reply to | #400820 |
On Tue, 4 Aug 2026 15:10:55 -0000 (UTC), Lew Pitcher wrote: > I mentioned Unicode because it both simplifies /and/ complicates the > matter of alphabetic value contiguity; While it ensures that > contiguity within an alphabet, it does not ensure contiguity between > alphabets (not that I think it should), leading to the same problem > that EBCDIC presented in the first place. Such simplistic arithmetic-based notions of character classification only worked in ASCII (and related encodings) for very limited character sets anyway. The international nature of the present-day computer market forces you to bite the bullet and admit that proper localization handling requires nontrivial libraries to implement properly. Luckily, such libraries are widely available in the open-source world -- for Unicode.
[toc] | [prev] | [next] | [standalone]
| From | David Brown <david.brown@hesbynett.no> |
|---|---|
| Date | 2026-08-05 09:01 +0200 |
| Message-ID | <114un3i$2npj3$1@dont-email.me> |
| In reply to | #400820 |
On 04/08/2026 17:10, Lew Pitcher wrote: > On Mon, 03 Aug 2026 16:58:39 -0700, Keith Thompson wrote: > >> Lew Pitcher <lew.pitcher@digitalfreehold.ca> writes: >> [...] >>> For what it's worth, this was also the reason (prior to Unicode) >>> that C did not specify that alphabetic characters would have a >>> contiguous sequence in the execution characterset. In EBCDIC, >>> the alphabetics group a-i, j-r, s-z and A-I, J-R, S-Z, with >>> various other characters (both assigned and unassigned) between >>> the groupings. >> >> C still doesn't require Unicode (well, mostly), > > Yes. I mentioned Unicode because it both simplifies /and/ complicates > the matter of alphabetic value contiguity; While it ensures that contiguity > within an alphabet, it does not ensure contiguity between alphabets (not > that I think it should), leading to the same problem that EBCDIC presented > in the first place. > >> and still doesn't require 'i'+1=='j'. > > I pointed that out because it has become a common programmer misconception; > hand rolled isalpha-like functions often express themselves with a range > check in the form of > is_lower_case = ((some_char >= 'a') && (some_char <= 'z')); > is_upper_case = ((some_char >= 'A') && (some_char <= 'Z')); > > The standard explicitly requires /numeric/ characters to have contiguity, > though. > is_number = ((some_char >= '0') && (some_char <= '9')); > > >> C does have UTF-8 string literals, such as >> u8"hello", which are encoded as UTF-8, but ordinary string literals >> like "hello" are still encoded using the execution character set, >> which could be EBCDIC. >> >> There's a proposal to require 'a'..'f' and 'A'..'F' to be contiguous, >> but it hasn't appeared in the latest C2y draft. >> >> https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3192.pdf > > That's going to complicate the mainframe C compilers a bit :-) > Or, as a TV presenter often put it... > "Oh no! Anyway ..." > No, it is not going to be an issue for any real-world character sets (including EBCDIC). Still, I don't think it is going to be particularly useful for anyone. The only purpose I can see of knowing that "A" - "F" and "a" - "f" are contiguous is for convenience when converting to and from hex characters. And the kind of system where you would find that useful (rather than just using "printf" and friends) is for small embedded systems. In such cases, you already know the character set, and you know you are not coding for a mainframe. So this proposal is simply documenting something you already know - even on mainframes and dinosaurs. There's nothing wrong with that, and it can be good to have things written out explicitly in the standards. But I don't think this particular change is going to make things easier for anyone.
[toc] | [prev] | [next] | [standalone]
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-08-04 18:10 +0000 |
| Message-ID | <114t9ur$eoee$1@paganini.bofh.team> |
| In reply to | #400748 |
Lew Pitcher <lew.pitcher@digitalfreehold.ca> wrote:
> On Mon, 03 Aug 2026 14:47:21 +0000, Lew Pitcher wrote:
>
>> On Sun, 02 Aug 2026 14:17:45 +0000, Kenny McCormack wrote:
>>
>>> First off, I know the "standards" answer is "Either is correct; you have no
>>> right to complain about anything", but I am not interested in the
>>> "standards" answer. If this is all you can do, then just click Next and go
>>> on.
>>>
>>> I'm interested in the "why" of why implementations might prefer one or the
>>> other.
>>
>> Consider the effects of the integer promotion rules on a system with an 8-bit
>> execution characterset (CHAR_BIT == 8) that has significant characters in the
>> 0x80 through 0xff range[1], and how it affects the return results of functions
>> like getchar(), getc(), and fgetc().
> [snip]
>> [1] Not as hypothetical as you might think; Some of the earliest C compilers
>> (and current compilers as well) targetted the IBM EBCDIC systems, where much
>> of the basic execution characterset resides between 0x80 and 0xff, with the
>> numeric characters residing between 0xf0 and 0xf9. A signed <<char>> would
>> not work here.
>
> For what it's worth, this was also the reason (prior to Unicode) that C
> did not specify that alphabetic characters would have a contiguous sequence
> in the execution characterset. In EBCDIC, the alphabetics group a-i, j-r, s-z
> and A-I, J-R, S-Z, with various other characters (both assigned and unassigned)
> between the groupings.
Unless you are payed specifically to do so I see no reason to support
EBCDIC. Of course, IBM have enough influence to keep C standard
as it is regarding character set, but it does not mean that anybody
else should take is seriously.
--
Waldek Hebisch
[toc] | [prev] | [next] | [standalone]
| From | Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> |
|---|---|
| Date | 2026-08-05 03:17 +0800 |
| Subject | Compilers targetting the C64 (was: Re: Default signedness of 'plain' char.) |
| Message-ID | <lVqcS.94640$jNNe.7616@fx15.ams4> |
| In reply to | #400825 |
On 05/08/2026 2:10 AM, Waldek Hebisch wrote:
> Lew Pitcher <lew.pitcher@digitalfreehold.ca> wrote:
>> On Mon, 03 Aug 2026 14:47:21 +0000, Lew Pitcher wrote:
>>
>>> On Sun, 02 Aug 2026 14:17:45 +0000, Kenny McCormack wrote:
>>>
>>>> First off, I know the "standards" answer is "Either is correct; you have no
>>>> right to complain about anything", but I am not interested in the
>>>> "standards" answer. If this is all you can do, then just click Next and go
>>>> on.
>>>>
>>>> I'm interested in the "why" of why implementations might prefer one or the
>>>> other.
>>>
>>> Consider the effects of the integer promotion rules on a system with an 8-bit
>>> execution characterset (CHAR_BIT == 8) that has significant characters in the
>>> 0x80 through 0xff range[1], and how it affects the return results of functions
>>> like getchar(), getc(), and fgetc().
>> [snip]
>>> [1] Not as hypothetical as you might think; Some of the earliest C compilers
>>> (and current compilers as well) targetted the IBM EBCDIC systems, where much
>>> of the basic execution characterset resides between 0x80 and 0xff, with the
>>> numeric characters residing between 0xf0 and 0xf9. A signed <<char>> would
>>> not work here.
>>
>> For what it's worth, this was also the reason (prior to Unicode) that C
>> did not specify that alphabetic characters would have a contiguous sequence
>> in the execution characterset. In EBCDIC, the alphabetics group a-i, j-r, s-z
>> and A-I, J-R, S-Z, with various other characters (both assigned and unassigned)
>> between the groupings.
>
> Unless you are payed specifically to do so I see no reason to support
> EBCDIC. Of course, IBM have enough influence to keep C standard
> as it is regarding character set, but it does not mean that anybody
> else should take is seriously.
>
On the other tentacle, I believe everyone creating a C compiler for
Commodore 64 should take PETSCII seriously; if I'm reading the Wiki-
pedia page right, and remember the C graphic characters correctly,
you'll need these two digraphs,
<% for {
%> for }
and everything else seems to be in place; just different from ASCII
according to
https://www.c64os.com/post/petsciiasciiconversion
but you're welcome to force everyone on C64s to use ASCII anyway, for
your C compiler.
Have a nice C64 day!
--
Johann | email: invalid -> com | http://www.myrkraverk.com/blog/
I'm not from the Internet, I just work there. | via Easynews.com
https://bsky.app/profile/myrkraverk.bsky.social
[toc] | [prev] | [next] | [standalone]
| From | BGB <cr88192@gmail.com> |
|---|---|
| Date | 2026-08-04 16:25 -0500 |
| Message-ID | <114tlkt$2g79h$1@dont-email.me> |
| In reply to | #400825 |
On 8/4/2026 1:10 PM, Waldek Hebisch wrote:
> Lew Pitcher <lew.pitcher@digitalfreehold.ca> wrote:
>> On Mon, 03 Aug 2026 14:47:21 +0000, Lew Pitcher wrote:
>>
>>> On Sun, 02 Aug 2026 14:17:45 +0000, Kenny McCormack wrote:
>>>
>>>> First off, I know the "standards" answer is "Either is correct; you have no
>>>> right to complain about anything", but I am not interested in the
>>>> "standards" answer. If this is all you can do, then just click Next and go
>>>> on.
>>>>
>>>> I'm interested in the "why" of why implementations might prefer one or the
>>>> other.
>>>
>>> Consider the effects of the integer promotion rules on a system with an 8-bit
>>> execution characterset (CHAR_BIT == 8) that has significant characters in the
>>> 0x80 through 0xff range[1], and how it affects the return results of functions
>>> like getchar(), getc(), and fgetc().
>> [snip]
>>> [1] Not as hypothetical as you might think; Some of the earliest C compilers
>>> (and current compilers as well) targetted the IBM EBCDIC systems, where much
>>> of the basic execution characterset resides between 0x80 and 0xff, with the
>>> numeric characters residing between 0xf0 and 0xf9. A signed <<char>> would
>>> not work here.
>>
>> For what it's worth, this was also the reason (prior to Unicode) that C
>> did not specify that alphabetic characters would have a contiguous sequence
>> in the execution characterset. In EBCDIC, the alphabetics group a-i, j-r, s-z
>> and A-I, J-R, S-Z, with various other characters (both assigned and unassigned)
>> between the groupings.
>
> Unless you are payed specifically to do so I see no reason to support
> EBCDIC. Of course, IBM have enough influence to keep C standard
> as it is regarding character set, but it does not mean that anybody
> else should take is seriously.
>
Practically speaking, unless one is targeting a machine that uses EBCDIC
or some other nonstandard character set, better advised to mostly ignore
it, as ASCII has made a decisive win here...
Well, and UTF-8...
Had in my projects partly adopted Unicode, but not without fudging.
Basic character-set mostly limited to a few blocks:
Latin-1 range;
Also went and added Greek and Cyrillic characters and similar.
Except 0600..07FF: Reclaimed / Reused in 8x8 console fonts.
Most characters in this range can't be represented in 8x8 pixels.
Was more useful to use 0600..06FF for 00..FF dense hexadecimal.
0..9, A..F: Can be represented nicely in 4x8 pixels.
In some cases, it is nice to be able to display twice the hexadecimal in
half the space (can also be used for decimal by treating it as BCD).
Also for reasons was nicer if it could fit in the UTF-8 2-byte range.
In this case, 0700..07FF can be used for some patterns related to UI
drawing and representing images via color-cells.
Say, for example (6 bits):
Vsgn,Hsgn,Vfrq2,Hfrq2
Which effectively specifies a sine-wave pattern at 1 of 4 frequencies
with a sign; Both horizontal and vertical.
This can be used to generate a series of 64 patterns that are useful in
approximating images as color-cells.
Well, then some rounded curves and dithered gradient patterns (for more
useful image approximation); and some basic tilesets for UI elements.
Well, some of these can be useful if one is storing graphics data in a
form like, say:
1b: Escape (0=Normal, 1=Skip/RLE/etc)
7b: Cell-Index
4b: ColorA, 16-color / RGBI
4b: ColorB: 16-color / RGBI
Skip might be used for blocks that are skipped over;
RLE for blocks repeating the same pattern or a flat-color region.
...
Though, not exactly high fidelity; but when it works OK, may be hard to
beat (and can reuse text-console mechanics).
...
[toc] | [prev] | [next] | [standalone]
| From | Lawrence D’Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2026-08-05 02:57 +0000 |
| Message-ID | <114u8qd$2l3pd$2@dont-email.me> |
| In reply to | #400825 |
On Tue, 4 Aug 2026 18:10:37 -0000 (UTC), Waldek Hebisch wrote: > Unless you are payed specifically to do so I see no reason to > support EBCDIC. Of course, IBM have enough influence to keep C > standard as it is regarding character set, but it does not mean that > anybody else should take is seriously. Interesting that IBM’s excuse for creating EBCDIC (for the System/360 range) was that the ASCII standard wasn’t quite “mature” enough for production use at the time. Given that both came out in 1964, the difference could only have been a few months at most.
[toc] | [prev] | [next] | [standalone]
| From | Lynn McGuire <lynnmcguire5@gmail.com> |
|---|---|
| Date | 2026-08-05 00:19 -0500 |
| Message-ID | <114uh60$2nab4$1@dont-email.me> |
| In reply to | #400841 |
On 8/4/2026 9:57 PM, Lawrence D’Oliveiro wrote: > On Tue, 4 Aug 2026 18:10:37 -0000 (UTC), Waldek Hebisch wrote: > >> Unless you are payed specifically to do so I see no reason to >> support EBCDIC. Of course, IBM have enough influence to keep C >> standard as it is regarding character set, but it does not mean that >> anybody else should take is seriously. > > Interesting that IBM’s excuse for creating EBCDIC (for the System/360 > range) was that the ASCII standard wasn’t quite “mature” enough for > production use at the time. > > Given that both came out in 1964, the difference could only have been > a few months at most. IBM probably had a hundred people working on EBCDIC for a couple of years. Plus, wasn't the System/360 the first 8 bit byte / 32 bit word machine? Lynn
[toc] | [prev] | [next] | [standalone]
| From | Johann 'Myrkraverk' Oskarsson <johann@myrkraverk.invalid> |
|---|---|
| Date | 2026-08-05 18:57 +0800 |
| Message-ID | <FGEcS.3205$s8Z7.63@fx09.ams4> |
| In reply to | #400844 |
On 05/08/2026 1:19 PM, Lynn McGuire wrote: > On 8/4/2026 9:57 PM, Lawrence D’Oliveiro wrote: >> On Tue, 4 Aug 2026 18:10:37 -0000 (UTC), Waldek Hebisch wrote: >> >>> Unless you are payed specifically to do so I see no reason to >>> support EBCDIC. Of course, IBM have enough influence to keep C >>> standard as it is regarding character set, but it does not mean that >>> anybody else should take is seriously. >> >> Interesting that IBM’s excuse for creating EBCDIC (for the System/360 >> range) was that the ASCII standard wasn’t quite “mature” enough for >> production use at the time. >> >> Given that both came out in 1964, the difference could only have been >> a few months at most. > > IBM probably had a hundred people working on EBCDIC for a couple of years. > > Plus, wasn't the System/360 the first 8 bit byte / 32 bit word machine? > > Lynn > If I remember correctly, there were several incompatible versions of the ASCII standard in that time frame. Before ASCII solidified, it was pro- bably the correct choice to ignore it. Of course, I didn't dig up sources, so I don't remember if ASCII was solid before or after 1964. I wouldn't know about the world's first 8bit byte/32bit word machine, but weren't there machines being developed in different countries too? BPCL was not invented for an American computer, yet its descendant, C comes from America, and is used everywhere now. Happy C coding in EBCDIC! -- Johann | email: invalid -> com | http://www.myrkraverk.com/blog/ I'm not from the Internet, I just work there. | via Easynews.com https://bsky.app/profile/myrkraverk.bsky.social
[toc] | [prev] | [next] | [standalone]
| From | scott@slp53.sl.home (Scott Lurndal) |
|---|---|
| Date | 2026-08-05 14:33 +0000 |
| Message-ID | <BQHcS.4944$wI58.2578@fx44.iad> |
| In reply to | #400844 |
Lynn McGuire <lynnmcguire5@gmail.com> writes: >On 8/4/2026 9:57 PM, Lawrence D’Oliveiro wrote: >> On Tue, 4 Aug 2026 18:10:37 -0000 (UTC), Waldek Hebisch wrote: >> >>> Unless you are payed specifically to do so I see no reason to >>> support EBCDIC. Of course, IBM have enough influence to keep C >>> standard as it is regarding character set, but it does not mean that >>> anybody else should take is seriously. >> >> Interesting that IBM’s excuse for creating EBCDIC (for the System/360 >> range) was that the ASCII standard wasn’t quite “mature” enough for >> production use at the time. >> >> Given that both came out in 1964, the difference could only have been >> a few months at most. > >IBM probably had a hundred people working on EBCDIC for a couple of years. Unlikely. IBM used 6-bit BCDIC for systems prior to the 360 family and it was simply logical to extend it to 8-bits and maintain compatability with prior generations of IBM systems.
[toc] | [prev] | [next] | [standalone]
| From | cross@spitfire.i.gajendra.net (Dan Cross) |
|---|---|
| Date | 2026-08-05 22:38 +0000 |
| Message-ID | <1150e0v$1t3$1@reader1.panix.com> |
| In reply to | #400844 |
In article <114uh60$2nab4$1@dont-email.me>, Lynn McGuire <lynnmcguire5@gmail.com> wrote: >On 8/4/2026 9:57 PM, Lawrence D’Oliveiro wrote: >> On Tue, 4 Aug 2026 18:10:37 -0000 (UTC), Waldek Hebisch wrote: >> >>> Unless you are payed specifically to do so I see no reason to >>> support EBCDIC. Of course, IBM have enough influence to keep C >>> standard as it is regarding character set, but it does not mean that >>> anybody else should take is seriously. >> >> Interesting that IBM’s excuse for creating EBCDIC (for the System/360 >> range) was that the ASCII standard wasn’t quite “mature” enough for >> production use at the time. >> >> Given that both came out in 1964, the difference could only have been >> a few months at most. > >IBM probably had a hundred people working on EBCDIC for a couple of years. > >Plus, wasn't the System/360 the first 8 bit byte / 32 bit word machine? I don't know if it was the first, but it was certainly the first _successful_ machine with those properties. As the tale goes, Fred Brooks kicked Gene Amdahl out of his office and told him not to come back until he had power-of-two data sizes. - Dan C.
[toc] | [prev] | [next] | [standalone]
| From | Keith Thompson <Keith.S.Thompson+u@gmail.com> |
|---|---|
| Date | 2026-08-04 21:36 -0700 |
| Message-ID | <114uel4$2mhdg$1@kst.eternal-september.org> |
| In reply to | #400825 |
antispam@fricas.org (Waldek Hebisch) writes:
[...]
> Unless you are payed specifically to do so I see no reason to support
> EBCDIC. Of course, IBM have enough influence to keep C standard
> as it is regarding character set, but it does not mean that anybody
> else should take is seriously.
What kind of "support" are you talking about?
Most of the time, it's just as easy to write code that will work
correctly regardless of the target system's character set, as long
as the implementation is conforming. You don't need to write
('a' <= c && c <= 'z') when you can write islower((unsigned char)c).
Though it can make a difference if you need to deal with multi-byte
characters; Unicode, which is based on ASCII, is just about the only
realistic option (I don't know that anyone actually uses UTF-EBCDIC).
--
Keith Thompson (The_Other_Keith) Keith.S.Thompson+u@gmail.com
void Void(void) { Void(); } /* The recursive call of the void */
[toc] | [prev] | [next] | [standalone]
| From | antispam@fricas.org (Waldek Hebisch) |
|---|---|
| Date | 2026-08-06 00:21 +0000 |
| Message-ID | <1150k2k$vbia$1@paganini.bofh.team> |
| In reply to | #400843 |
Keith Thompson <Keith.S.Thompson+u@gmail.com> wrote:
> antispam@fricas.org (Waldek Hebisch) writes:
> [...]
>> Unless you are payed specifically to do so I see no reason to support
>> EBCDIC. Of course, IBM have enough influence to keep C standard
>> as it is regarding character set, but it does not mean that anybody
>> else should take is seriously.
>
> What kind of "support" are you talking about?
>
> Most of the time, it's just as easy to write code that will work
> correctly regardless of the target system's character set, as long
> as the implementation is conforming. You don't need to write
> ('a' <= c && c <= 'z') when you can write islower((unsigned char)c).
The two are not that same: assuming ASCII based encoding first
detects ASCII lowercase letter, the second is locale dependent.
In my use cases the second is usually wrong, so the first is
better.
> Though it can make a difference if you need to deal with multi-byte
> characters; Unicode, which is based on ASCII, is just about the only
> realistic option (I don't know that anyone actually uses UTF-EBCDIC).
>
--
Waldek Hebisch
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2026-08-03 13:20 -0700 |
| Message-ID | <114qt5o$1k97b$4@dont-email.me> |
| In reply to | #400708 |
On 8/2/2026 7:17 AM, Kenny McCormack wrote:
> First off, I know the "standards" answer is "Either is correct; you have no
> right to complain about anything", but I am not interested in the
> "standards" answer. If this is all you can do, then just click Next and go
> on.
>
> I'm interested in the "why" of why implementations might prefer one or the
> other.
>
> Consider:
>
> /* macro 'U' must be defined on the cmd line */
> #include <stdio.h>
>
> int main(void)
> {
> U char c = 255;
>
> printf("Result of 'c > 0': %d\n",c > 0);
> }
>
> And the following command lines:
>
> $ tcc -DU= -run CheckSignedChar.c
> $ tcc -DU=signed -run CheckSignedChar.c
> $ tcc -DU=unsigned -run CheckSignedChar.c
>
> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the default is
> signed. I'm interested in what sorts of factors drive the decision-making.
>
> Note, BTW, that I first noticed this in a project using gcc, but it is
> easier to test using tcc, as above.
>
> Also, total aside, I'm surprised that one needs to do -DU= instead of just
> -DU. I thought -DU would define it as an empty string, but that generates
> a compile error. You need -DU=. Why?
>
The sign of char is just what the underlying system needs to do its
thing. If you want a signed char, just signed char. ;^)
fwiw, I personally prefer unsigned char for all of my raw buffers and
such, but that's just me.
[toc] | [prev] | [next] | [standalone]
| From | bart <bc@freeuk.com> |
|---|---|
| Date | 2026-08-03 22:06 +0100 |
| Message-ID | <114qvst$1latt$1@dont-email.me> |
| In reply to | #400773 |
On 03/08/2026 21:20, Chris M. Thomasson wrote:
> On 8/2/2026 7:17 AM, Kenny McCormack wrote:
>> First off, I know the "standards" answer is "Either is correct; you
>> have no
>> right to complain about anything", but I am not interested in the
>> "standards" answer. If this is all you can do, then just click Next
>> and go
>> on.
>>
>> I'm interested in the "why" of why implementations might prefer one or
>> the
>> other.
>>
>> Consider:
>>
>> /* macro 'U' must be defined on the cmd line */
>> #include <stdio.h>
>>
>> int main(void)
>> {
>> U char c = 255;
>>
>> printf("Result of 'c > 0': %d\n",c > 0);
>> }
>>
>> And the following command lines:
>>
>> $ tcc -DU= -run CheckSignedChar.c
>> $ tcc -DU=signed -run CheckSignedChar.c
>> $ tcc -DU=unsigned -run CheckSignedChar.c
>>
>> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the
>> default is
>> signed. I'm interested in what sorts of factors drive the decision-
>> making.
>>
>> Note, BTW, that I first noticed this in a project using gcc, but it is
>> easier to test using tcc, as above.
>>
>> Also, total aside, I'm surprised that one needs to do -DU= instead of
>> just
>> -DU. I thought -DU would define it as an empty string, but that
>> generates
>> a compile error. You need -DU=. Why?
>>
>
> The sign of char is just what the underlying system needs to do its
> thing. If you want a signed char, just signed char. ;^)
It's not that simple. Very many libraries including the standard library
make use of char* for strings for example. And string literals will be
char* too.
So you have to play along, you can't just use signed char* or unsigned
char*; compilers will complain.
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2026-08-03 14:25 -0700 |
| Message-ID | <114r0vj$1lmn5$1@dont-email.me> |
| In reply to | #400776 |
On 8/3/2026 2:06 PM, bart wrote:
> On 03/08/2026 21:20, Chris M. Thomasson wrote:
>> On 8/2/2026 7:17 AM, Kenny McCormack wrote:
>>> First off, I know the "standards" answer is "Either is correct; you
>>> have no
>>> right to complain about anything", but I am not interested in the
>>> "standards" answer. If this is all you can do, then just click Next
>>> and go
>>> on.
>>>
>>> I'm interested in the "why" of why implementations might prefer one
>>> or the
>>> other.
>>>
>>> Consider:
>>>
>>> /* macro 'U' must be defined on the cmd line */
>>> #include <stdio.h>
>>>
>>> int main(void)
>>> {
>>> U char c = 255;
>>>
>>> printf("Result of 'c > 0': %d\n",c > 0);
>>> }
>>>
>>> And the following command lines:
>>>
>>> $ tcc -DU= -run CheckSignedChar.c
>>> $ tcc -DU=signed -run CheckSignedChar.c
>>> $ tcc -DU=unsigned -run CheckSignedChar.c
>>>
>>> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the
>>> default is
>>> signed. I'm interested in what sorts of factors drive the decision-
>>> making.
>>>
>>> Note, BTW, that I first noticed this in a project using gcc, but it is
>>> easier to test using tcc, as above.
>>>
>>> Also, total aside, I'm surprised that one needs to do -DU= instead of
>>> just
>>> -DU. I thought -DU would define it as an empty string, but that
>>> generates
>>> a compile error. You need -DU=. Why?
>>>
>>
>> The sign of char is just what the underlying system needs to do its
>> thing. If you want a signed char, just signed char. ;^)
>
> It's not that simple. Very many libraries including the standard library
> make use of char* for strings for example. And string literals will be
> char* too.
>
> So you have to play along, you can't just use signed char* or unsigned
> char*; compilers will complain.
>
>
I use unsigned char for my personal buffers. If a char is signed or not
is up to the impl. C std besides the point here. If I want to use a C
function, I know how to do it.
[toc] | [prev] | [next] | [standalone]
| From | BGB <cr88192@gmail.com> |
|---|---|
| Date | 2026-08-03 18:38 -0500 |
| Message-ID | <114r8vp$1nop2$1@dont-email.me> |
| In reply to | #400778 |
On 8/3/2026 4:25 PM, Chris M. Thomasson wrote:
> On 8/3/2026 2:06 PM, bart wrote:
>> On 03/08/2026 21:20, Chris M. Thomasson wrote:
>>> On 8/2/2026 7:17 AM, Kenny McCormack wrote:
>>>> First off, I know the "standards" answer is "Either is correct; you
>>>> have no
>>>> right to complain about anything", but I am not interested in the
>>>> "standards" answer. If this is all you can do, then just click Next
>>>> and go
>>>> on.
>>>>
>>>> I'm interested in the "why" of why implementations might prefer one
>>>> or the
>>>> other.
>>>>
>>>> Consider:
>>>>
>>>> /* macro 'U' must be defined on the cmd line */
>>>> #include <stdio.h>
>>>>
>>>> int main(void)
>>>> {
>>>> U char c = 255;
>>>>
>>>> printf("Result of 'c > 0': %d\n",c > 0);
>>>> }
>>>>
>>>> And the following command lines:
>>>>
>>>> $ tcc -DU= -run CheckSignedChar.c
>>>> $ tcc -DU=signed -run CheckSignedChar.c
>>>> $ tcc -DU=unsigned -run CheckSignedChar.c
>>>>
>>>> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the
>>>> default is
>>>> signed. I'm interested in what sorts of factors drive the decision-
>>>> making.
>>>>
>>>> Note, BTW, that I first noticed this in a project using gcc, but it is
>>>> easier to test using tcc, as above.
>>>>
>>>> Also, total aside, I'm surprised that one needs to do -DU= instead
>>>> of just
>>>> -DU. I thought -DU would define it as an empty string, but that
>>>> generates
>>>> a compile error. You need -DU=. Why?
>>>>
>>>
>>> The sign of char is just what the underlying system needs to do its
>>> thing. If you want a signed char, just signed char. ;^)
>>
>> It's not that simple. Very many libraries including the standard
>> library make use of char* for strings for example. And string literals
>> will be char* too.
>>
>> So you have to play along, you can't just use signed char* or unsigned
>> char*; compilers will complain.
>>
>>
>
> I use unsigned char for my personal buffers. If a char is signed or not
> is up to the impl. C std besides the point here. If I want to use a C
> function, I know how to do it.
I typically do:
typedef unsigned char byte; //often
typedef signed char sbyte; //sometimes
Then often u16/u32/u64, s16/s32/s64, ...
But, mostly because even with C99, "uint64_t" and similar are enough
typing to be more annoying (whenever one feels a need for an exact-width
type). Had started gradually shifting to using the C99 types as a
reference point, as I am no longer actively using compilers that don't
support the C99 "stdint.h" stuff (though last I checked, MSVC still
doesn't fully support C99; eg, still no VLAs or _Complex).
...
[toc] | [prev] | [next] | [standalone]
| From | "Chris M. Thomasson" <chris.m.thomasson.1@gmail.com> |
|---|---|
| Date | 2026-08-03 19:34 -0700 |
| Message-ID | <114rj38$1qif9$1@dont-email.me> |
| In reply to | #400788 |
On 8/3/2026 4:38 PM, BGB wrote:
> On 8/3/2026 4:25 PM, Chris M. Thomasson wrote:
>> On 8/3/2026 2:06 PM, bart wrote:
>>> On 03/08/2026 21:20, Chris M. Thomasson wrote:
>>>> On 8/2/2026 7:17 AM, Kenny McCormack wrote:
>>>>> First off, I know the "standards" answer is "Either is correct; you
>>>>> have no
>>>>> right to complain about anything", but I am not interested in the
>>>>> "standards" answer. If this is all you can do, then just click
>>>>> Next and go
>>>>> on.
>>>>>
>>>>> I'm interested in the "why" of why implementations might prefer one
>>>>> or the
>>>>> other.
>>>>>
>>>>> Consider:
>>>>>
>>>>> /* macro 'U' must be defined on the cmd line */
>>>>> #include <stdio.h>
>>>>>
>>>>> int main(void)
>>>>> {
>>>>> U char c = 255;
>>>>>
>>>>> printf("Result of 'c > 0': %d\n",c > 0);
>>>>> }
>>>>>
>>>>> And the following command lines:
>>>>>
>>>>> $ tcc -DU= -run CheckSignedChar.c
>>>>> $ tcc -DU=signed -run CheckSignedChar.c
>>>>> $ tcc -DU=unsigned -run CheckSignedChar.c
>>>>>
>>>>> On 32 bit RpiOS, the default is unsigned, but on x64 Ubuntu, the
>>>>> default is
>>>>> signed. I'm interested in what sorts of factors drive the
>>>>> decision- making.
>>>>>
>>>>> Note, BTW, that I first noticed this in a project using gcc, but it is
>>>>> easier to test using tcc, as above.
>>>>>
>>>>> Also, total aside, I'm surprised that one needs to do -DU= instead
>>>>> of just
>>>>> -DU. I thought -DU would define it as an empty string, but that
>>>>> generates
>>>>> a compile error. You need -DU=. Why?
>>>>>
>>>>
>>>> The sign of char is just what the underlying system needs to do its
>>>> thing. If you want a signed char, just signed char. ;^)
>>>
>>> It's not that simple. Very many libraries including the standard
>>> library make use of char* for strings for example. And string
>>> literals will be char* too.
>>>
>>> So you have to play along, you can't just use signed char* or
>>> unsigned char*; compilers will complain.
>>>
>>>
>>
>> I use unsigned char for my personal buffers. If a char is signed or
>> not is up to the impl. C std besides the point here. If I want to use
>> a C function, I know how to do it.
>
> I typically do:
> typedef unsigned char byte; //often
> typedef signed char sbyte; //sometimes
I also remember using the word, word a lot... ;^)
Not for bytes, but for things like uintptr_t. Always found it useful.
A double word struct is comprised of two adjacent words.
struct anchor
{
word m_part_0;
word m_part_1;
};
make sure with a static assert or something that the sizeof(struct
anchor) == (sizeof(word) * 2)
Fwiw, it works well with the CMPXCH8B or CMPXCHG16B instructions on x86.
Double width atomic CAS.
> Then often u16/u32/u64, s16/s32/s64, ...
>
> But, mostly because even with C99, "uint64_t" and similar are enough
> typing to be more annoying (whenever one feels a need for an exact-width
> type). Had started gradually shifting to using the C99 types as a
> reference point, as I am no longer actively using compilers that don't
> support the C99 "stdint.h" stuff (though last I checked, MSVC still
> doesn't fully support C99; eg, still no VLAs or _Complex).
Its been a while since I used c99 on windows, but iirc, their (MSVC)
support for complex numbers was total crap. I think there was a way to
get it working, but it was not std at all. Iirc the last GCC I used had
good support for std complex numbers in C99.
[toc] | [prev] | [next] | [standalone]
Page 2 of 7 — ← Prev page 1 [2] 3 4 5 6 7 Next page →
Back to top | Article view | comp.lang.c
csiph-web