Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #194373 > unrolled thread
| Started by | mess-mate <mess-mate@gmx.com> |
|---|---|
| First post | 2018-04-01 16:10 +0200 |
| Last post | 2018-04-04 14:20 +0200 |
| Articles | 20 on this page of 101 — 21 participants |
Back to article view | Back to linux.debian.user
utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200
Page 4 of 6 — ← Prev page 1 2 3 [4] 5 6 Next page →
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 21:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAWgx-2ha-9@gated-at.bofh.it> |
| In reply to | #194500 |
On Wednesday, April 04, 2018 02:10:16 PM Jonathan de Boyne Pollard wrote: > rhkramer: > > The reason I wanted such a byte was to use it as a record separator in > > a set of text files (that I use as an askSam "workalike" (or > > "worksimilar") so that I could use msort (which depends on a 1 byte > > record separator to --separate the records ;-) while sorting. Some of > > the files already include UTF-8, and, in the future, I anticpate all > > will be in UTFF-8. > > Note that ISO 646, hence ISO 8859, hence ISO 10646, has had a > single-byte Record Separator character since the 1960s. (-: Ok, thanks, I see that is Dec 30, Hex 1e. A quick look at the UTF-8 table in the Wikipedia article on UTF-8 seems to indicate that byte is a valid UTF-9 byte, which makes it unsuitabe for my use.
[toc] | [prev] | [next] | [standalone]
| From | Ben Caradoc-Davies <ben@transient.nz> |
|---|---|
| Date | 2018-04-02 23:40 +0200 |
| Message-ID | <vAflg-6VY-9@gated-at.bofh.it> |
| In reply to | #194388 |
On 02/04/18 19:39, Andre Majorel wrote: > On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote: >> Why? UTF (especially UTF-8) is vastly superior for all purposes: > I wouldn't say that. UTF-8 breaks a number of assumptions. For > instance, > 1) every character has the same size, > 2) every byte sequence is a valid character, > 3) the equality or inequality of two characters comes down to > the equality or inequality of the bytes they encode to. > With ASCII and the many encodings based on it, most things can > be done without having knowledge of the encoding. With UTF-8, > even basic operations like determining the length of a string or > reporting at what column an error occurred require knowledge of > the encoding. That is entirely correct. The problem is that ASCII has the worst assumption of all: everyone uses US English. This not valid in a Universal Operating System. Kind regards, -- Ben Caradoc-Davies <ben@transient.nz> Director Transient Software Limited <https://transient.nz/> New Zealand
[toc] | [prev] | [next] | [standalone]
| From | Darac Marjal <mailinglist@darac.org.uk> |
|---|---|
| Date | 2018-04-03 11:00 +0200 |
| Message-ID | <vApXj-5wU-1@gated-at.bofh.it> |
| In reply to | #194388 |
[Multipart message — attachments visible in raw view] — view raw
On Mon, Apr 02, 2018 at 09:39:05AM +0200, Andre Majorel wrote: >On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote: >> On 02/04/18 02:05, mess-mate wrote: >> >howto change the system utf to eu character set ? >> >> Why? UTF (especially UTF-8) is vastly superior for all purposes: > >I wouldn't say that. UTF-8 breaks a number of assumptions. For >instance, >1) every character has the same size, >2) every byte sequence is a valid character, >3) the equality or inequality of two characters comes down to > the equality or inequality of the bytes they encode to. If these things matter to you, it's better to convert from UTF-8 to Unicode, first. I tend to think of Unicode as an arbitrarily large code page. Each character maps to a number, but that number could be 1, 1000 or 500_000 (Unicode seems to be growing without might end in sight). Internally, you might store those code points as Integers or QUad Words or whatever you like. Only once you're ready to transfer the text to another process (print on screen, save to a file, stream across a network), do you convert the Unicode back into UTF-8. Basically, you consider UTF-8 to be a transfer-only format (like Base64). If you want to do anything non-trivial with it, decode it into Unicode. > >With ASCII and the many encodings based on it, most things can >be done without having knowledge of the encoding. With UTF-8, >even basic operations like determining the length of a string or >reporting at what column an error occurred require knowledge of >the encoding. > >-- >André Majorel <http://www.teaser.fr/~amajorel/> >Imagine what would happen if the Debian project disclosed the email >addresses of their users. Spambots would harvest them and Debian >users would be inundated with spam. Good thing they don't, eh ? > -- For more information, please reread.
[toc] | [prev] | [next] | [standalone]
| From | Richard Hector <richard@walnut.gen.nz> |
|---|---|
| Date | 2018-04-03 11:20 +0200 |
| Message-ID | <vAqgG-5VJ-23@gated-at.bofh.it> |
| In reply to | #194433 |
[Multipart message — attachments visible in raw view] — view raw
On 03/04/18 20:55, Darac Marjal wrote: > If these things matter to you, it's better to convert from UTF-8 to > Unicode, first. I tend to think of Unicode as an arbitrarily large code > page. Each character maps to a number, but that number could be 1, 1000 > or 500_000 (Unicode seems to be growing without might end in sight). > Internally, you might store those code points as Integers or QUad Words > or whatever you like. Only once you're ready to transfer the text to > another process (print on screen, save to a file, stream across a > network), do you convert the Unicode back into UTF-8. > > Basically, you consider UTF-8 to be a transfer-only format (like > Base64). If you want to do anything non-trivial with it, decode it into > Unicode. Eh? UTF-8 is an encoding of Unicode. You can't "convert UTF-8 to Unicode" - it already is Unicode. You could convert it to another encoding, eg UTF-16 or UTF-32. Perhaps UTF-32 is what you mean, being fixed-width. Richard
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-03 11:30 +0200 |
| Message-ID | <vAqql-5Zj-7@gated-at.bofh.it> |
| In reply to | #194434 |
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1
On Tue, Apr 03, 2018 at 09:14:22PM +1200, Richard Hector wrote:
> On 03/04/18 20:55, Darac Marjal wrote:
> > If these things matter to you, it's better to convert from UTF-8 to
> > Unicode, first. I tend to think of Unicode as an arbitrarily large code
> > page. Each character maps to a number, but that number could be 1, 1000
> > or 500_000 (Unicode seems to be growing without might end in sight).
> > Internally, you might store those code points as Integers or QUad Words
> > or whatever you like. Only once you're ready to transfer the text to
> > another process (print on screen, save to a file, stream across a
> > network), do you convert the Unicode back into UTF-8.
> >
> > Basically, you consider UTF-8 to be a transfer-only format (like
> > Base64). If you want to do anything non-trivial with it, decode it into
> > Unicode.
>
> Eh? UTF-8 is an encoding of Unicode. You can't "convert UTF-8 to
> Unicode" - it already is Unicode. You could convert it to another
> encoding, eg UTF-16 or UTF-32. Perhaps UTF-32 is what you mean, being
> fixed-width.
I think Darac was talking about UTF-32 [1], which is a fixed-width encoding
of Unicode. Yes, Unicode is strictly speaking the abstract "mapping" between
integers ("code points") and characters. A computer has no integers...
What's curious is that there's no UTF-24 (although Unicode currently has
all its code points below 2^21). That would make for a slightly more
compact fixed-width encoding.
I think these days fixed-width encodings are losing their charm a bit,
since memory access is getting much more expensive than CPU power.
Things might change once again when the Chinese dominate culturally,
since UTF-8 plays its advantage only with ASCII dominated text.
But perhaps then, another encoding will make more sense. Or just UTF-24
is born, for a 25% savings :-)
Cheers
[1] https://en.wikipedia.org/wiki/UTF-32
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)
iEYEARECAAYFAlrDST0ACgkQBcgs9XrR2kadzQCeO8N1Kjua/p0aOdfE8QQTvv6R
PisAn0+0DbwWXb+cWYsUmMwqSqQN+BKQ
=y+Fu
-----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-03 12:20 +0200 |
| Message-ID | <vArcJ-6wn-1@gated-at.bofh.it> |
| In reply to | #194433 |
[Multipart message — attachments visible in raw view] — view raw
> On Mon, Apr 02, 2018 at 09:39:05AM +0200, Andre Majorel wrote: > >I wouldn't say that. UTF-8 breaks a number of assumptions. For > >instance, > >1) every character has the same size, > >2) every byte sequence is a valid character, > >3) the equality or inequality of two characters comes down to > > the equality or inequality of the bytes they encode to. I am sure you do not realize that none of these assumptions are really met by any encoding and none of these assumptions actually bring something. They are just rehashed poor arguments to rationalize a fear of change by people afraid their long-earned knowledge will become obsolete. I do not know if you fit in that category. Odds are you have just been misinformed after having normal trouble during the transition. Darac Marjal (2018-04-03): > If these things matter to you, it's better to convert from UTF-8 to Unicode, > first. "Convert to Unicode" does not mean anything. Unicode is not a format, and therefore you cannot convert something to it. Unicode is a catalog of "infolinguistic entities". I do not say characters, because they are not all. Most of Unicode is characters, but not all. As is stands, the principle of Unicode is that any "pure" text can be represented as a sequence of Unicode code points. For storing in a file or sending on the network, this sequence of code points must be converted into a sequence of octets. UTF-8 is by far the best choice for that, because it has many interesting properties. Other encodings suffer from being incomplete, being incompatible with ASCII, being sensible to endianness problems, being subject to corruption or all of the above. > I tend to think of Unicode as an arbitrarily large code page. Each > character maps to a number, but that number could be 1, 1000 or 500_000 > (Unicode seems to be growing without might end in sight). The twentieth century just called, it wanks the "code page" idiom back. > Internally, you > might store those code points as Integers or QUad Words or whatever you > like. Only once you're ready to transfer the text to another process (print > on screen, save to a file, stream across a network), do you convert the > Unicode back into UTF-8. Internally, this is a reasonable choice in some cases, but not actually that useful in most. Your discourse is based on the assumption that accessing a single Unicode code point in the string would be useful. Most of the time, it is not. Remember that an Unicode code point is not necessarily a character. You need to decide the data structure based on the treatments you intend to perform on the text. And actually, most of the time using an array of octets in UTF-8 is the best choice for internal representation too. > Basically, you consider UTF-8 to be a transfer-only format (like Base64). If > you want to do anything non-trivial with it, decode it into Unicode. No, definitely not. If somebody wants to do anything non-trivial with text, then either they already know what they are doing better than this, and do not need that advice, or they do not and they will get it wrong. Use a library. And use whatever text format that library use. The problem does not come from UTF-8 or Unicode or anything computer-related, the problem comes from the principle of written human text: writing systems are insanely complex. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Ben Caradoc-Davies <ben@transient.nz> |
|---|---|
| Date | 2018-04-03 22:50 +0200 |
| Message-ID | <vAB2p-4vD-9@gated-at.bofh.it> |
| In reply to | #194433 |
On 03/04/18 20:55, Darac Marjal wrote: > If these things matter to you, it's better to convert from UTF-8 to > Unicode, first. Fixed length encodings like UTF-32 will not fix broken assumptions about some relationship between byte length and number of characters because Unicode contains things like combining characters. What is the length of a string? Are you trying to count the number of glyphs? I do not think that you can do this by naïvely counting code points, regardless of encoding. Because there is more than one way to represent an accented character, Unicode string comparison is nontrivial: https://en.wikipedia.org/wiki/Unicode_equivalence Kind regards, -- Ben Caradoc-Davies <ben@transient.nz> Director Transient Software Limited <https://transient.nz/> New Zealand
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-03 23:00 +0200 |
| Message-ID | <vABc6-4A2-7@gated-at.bofh.it> |
| In reply to | #194463 |
[Multipart message — attachments visible in raw view] — view raw
Ben Caradoc-Davies (2018-04-04): > What is the length of a > string? When is that relevant? > Are you trying to count the number of glyphs? What for? > I do not think that > you can do this by naïvely counting code points, regardless of encoding. Fortunately, this is almost never useful. > Because there is more than one way to represent an accented character, > Unicode string comparison is nontrivial: Fortunately, it is also rather rarely useful in all that complexity. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-03 23:00 +0200 |
| Message-ID | <vABc6-4A2-9@gated-at.bofh.it> |
| In reply to | #194464 |
On Tue, Apr 03, 2018 at 10:51:43PM +0200, Nicolas George wrote: > Ben Caradoc-Davies (2018-04-04): > > What is the length of a > > string? > > When is that relevant? When you're trying to display one on a screen, or print one on paper. When you've been asked to find the longest/shortest string from a list. When you've been asked to sort a list of strings by length. Or in other words, basically every time you do anything with the string at all other than blindly byte-copying it to a different place in memory.
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-03 23:30 +0200 |
| Message-ID | <vABFa-50o-59@gated-at.bofh.it> |
| In reply to | #194465 |
[Multipart message — attachments visible in raw view] — view raw
Greg Wooledge (2018-04-03): > When you're trying to display one on a screen, or print one on paper. With just the length? You will not get anything done. To properly display a string, you need to handle ligatures, right-to-left, kerning, etc. The length of the string is barely relevant. > When you've been asked to find the longest/shortest string from a list. Why would that be useful? > When you've been asked to sort a list of strings by length. Why would that be useful? > Or in other words, basically every time you do anything with the string > at all other than blindly byte-copying it to a different place in memory. No, the length of the string is hardly relevant, and when it is it is not enough anyway. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | deloptes <deloptes@gmail.com> |
|---|---|
| Date | 2018-04-04 19:00 +0200 |
| Message-ID | <vATVo-Ax-11@gated-at.bofh.it> |
| In reply to | #194466 |
Nicolas George wrote: > No, the length of the string is hardly relevant, and when it is it is > not enough anyway. @Nicolas, I think OP does not understand you - perhaps it is not worth the effort. My impression is that you refer to a string (properly) as sequence of bytes and other refer to it as number of chars, which is not consistant with utf. >From my work with UTF, it is possible but not satisfying to guess encoding. I wonder why no one suggested a kind of markup (xml) instead of byte delimiter. And regarding the mbox thing, well mbox was depreciated for many reasons. I guess if it was that good it wouldn't be depreciated. @OP, at some point of time everyone has to redesign and reimplement because technology evolves and all the tools listed can be updated to the new format. Recent example of redesign I worked with is gnupg - huge changes from v1.x to v2.1 regards
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 19:10 +0200 |
| Message-ID | <vAU54-Tb-21@gated-at.bofh.it> |
| In reply to | #194486 |
[Multipart message — attachments visible in raw view] — view raw
deloptes (2018-04-04): > @Nicolas, I think OP does not understand you - perhaps it is not worth the > effort. My impression is that you refer to a string (properly) as sequence > of bytes and other refer to it as number of chars, which is not consistant > with utf. Not at all, I am well speaking of text string, made up of characters, expressed as sequences of Unicode code point, and encoded as any data structure convenient. What I am trying to explain (not to the OP who is focussing on another thing entirely and obviously beyond help anyway), is that if you are thinking of stings in terms of "access the n-th char", "find the length of the string in chars", etc., then you completely have missed the point of them. Find me a case where you need to access the n-th char of a string, with n completely out of the blue, and I will explain how somebody botched their design. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-04 19:30 +0200 |
| Message-ID | <vAUoq-11h-11@gated-at.bofh.it> |
| In reply to | #194487 |
On Wed, Apr 04, 2018 at 07:07:01PM +0200, Nicolas George wrote: > Find me a case where you need to access the n-th char of a string, with > n completely out of the blue, and I will explain how somebody botched > their design. Does it count if we want the 1st char, then the 2nd char, then the 3rd char, then the 4th char, and so on? Or is that not blue enough? How about the last char? Or the last two chars? Or all chars starting just after the last slash or period? How about performing a checksum like the <https://en.wikipedia.org/wiki/Luhn_algorithm> on a user input string which is supposed to be a 10-digit <https://en.wikipedia.org/wiki/National_Provider_Identifier> ? Might it be useful to check the length of an input string before bothering to decompose it into individual digits and perform the arithmetic? And here, by "length", I mean "number of characters". You can see how that might be a handy thing, right? Length as in "number of bytes required to store it" is also an important value, of course. Character length is also useful when displaying strings on CHARACTER-ORIENTED OUTPUT MEDIA. Like terminals. You know, those things that Unix-like systems use all the time? How else are you going to space-pad the fields so that the output columns line up, if you don't know how many extra spaces you need, because you don't know the length of the string? (It's frankly disturbing to me that when I talked about length being relevant when printing strings, you immediately jumped to "pixels" and "fonts". This tells me that you no longer accept the terminal as your lord and savior. If you ever did.) All of these things matter, and are real, and don't necessarily indicate "botched design".
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 19:40 +0200 |
| Message-ID | <vAUy6-14B-11@gated-at.bofh.it> |
| In reply to | #194488 |
[Multipart message — attachments visible in raw view] — view raw
Greg Wooledge (2018-04-04): > Does it count if we want the 1st char, then the 2nd char, then the 3rd > char, then the 4th char, and so on? Or is that not blue enough? It is not out of the blue, it is in sequence. > How about the last char? Or the last two chars? Ditto. > Or all chars starting > just after the last slash or period? There is no n here, you have to first look for the last slash. > How about performing a checksum like the > <https://en.wikipedia.org/wiki/Luhn_algorithm> on a user input string > which is supposed to be a 10-digit > <https://en.wikipedia.org/wiki/National_Provider_Identifier> ? Again, in sequence. > Might it be useful to check the length of an input string before > bothering to decompose it into individual digits and perform the > arithmetic? And to check it is made of all digits. Checking the length is only a byproduct. > And here, by "length", I mean "number of characters". > You can see how that might be a handy thing, right? Still, no. > Length as in "number of bytes required to store it" is also an > important value, of course. You have to compute it to store it, indeed. > Character length is also useful when displaying strings on > CHARACTER-ORIENTED OUTPUT MEDIA. Like terminals. You know, those > things that Unix-like systems use all the time? How else are you > going to space-pad the fields so that the output columns line up, > if you don't know how many extra spaces you need, because you don't > know the length of the string? I do not know. Please tell me, how do you handle control characters, escape sequence, double-width characters, etc., without walking the string in sequence? > (It's frankly disturbing to me that when I talked about length being > relevant when printing strings, you immediately jumped to "pixels" and > "fonts". This tells me that you no longer accept the terminal as your > lord and savior. If you ever did.) I do not have a "lord and savior". I use the terminal a lot, but I am aware of the hidden complexities. > All of these things matter, and are real, and don't necessarily indicate > "botched design". All these things matter, but they do not require random access in a string by char number. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | deloptes <deloptes@gmail.com> |
|---|---|
| Date | 2018-04-04 19:40 +0200 |
| Message-ID | <vAUy5-14B-1@gated-at.bofh.it> |
| In reply to | #194487 |
Nicolas George wrote: > Find me a case where you need to access the n-th char of a string, with > n completely out of the blue, and I will explain how somebody botched > their design. ok, thanks. I understood the part above, but not sure if I understand this part. A standard text editing operation is find and replace, where you get the start and end point in the string. Of course it is not "n completely out of the blue". regards
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 19:40 +0200 |
| Message-ID | <vAUy6-14B-9@gated-at.bofh.it> |
| In reply to | #194489 |
[Multipart message — attachments visible in raw view] — view raw
deloptes (2018-04-04): > ok, thanks. I understood the part above, but not sure if I understand this > part. A standard text editing operation is find and replace, where you get > the start and end point in the string. Of course it is not "n completely > out of the blue". I am not sure exactly what is your example, but you got its flaw right: n is not out of the blue, it was obtained by previously walking the string. And in that case, you have all freedom to express n as a more convenient entity than an index expressed in terms of chars. A pointer maybe, or a pair with both the char index and the octet offset. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-04 19:50 +0200 |
| Message-ID | <vAUHM-17T-5@gated-at.bofh.it> |
| In reply to | #194492 |
On Wed, Apr 04, 2018 at 07:35:37PM +0200, Nicolas George wrote: > I am not sure exactly what is your example, but you got its flaw right: > n is not out of the blue, it was obtained by previously walking the > string. And in that case, you have all freedom to express n as a more > convenient entity than an index expressed in terms of chars. A pointer > maybe, or a pair with both the char index and the octet offset. The problem is, you reject every single example that everyone gives you. I don't know what you expect from us. You just seem to have Decided, for reasons known only to you, that The Character Length Of A String Is Not Useful. Despite literally decades of programs that have used strlen() in various ways. Have you never been given ANY kind of problem that involves analysis of character strings? Ever? At all? What if the question is "Find all the English words that have an E in the 5th position and a U in the 7th"? I mean, seriously, at some point you either have to accept that one of our examples is good enough to justify the existence of strlen() and character-based string indexing, or we just label you a loon and ignore everything you say henceforth. I think most of the REST of us can agree that the character length and the byte length of a string are BOTH useful quantities, important in countless ways to countless programs. And that sometimes your algorithm wants to treat a string as an indexed array of characters, and retrieve the nth character. And that we shouldn't have to dig up an entire textbook worth of examples to explain this.
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 20:00 +0200 |
| Message-ID | <vAURs-1cT-7@gated-at.bofh.it> |
| In reply to | #194494 |
[Multipart message — attachments visible in raw view] — view raw
Greg Wooledge (2018-04-04): > The problem is, you reject every single example that everyone gives > you. I do not reject them, I refute them. > I don't know what you expect from us. Acknowledge that I am right once I have refuted all your examples and you have eventually understood my point. At this time, you have not yet understood. > You just seem to have Decided, for reasons known only to you, that > The Character Length Of A String Is Not Useful. Despite literally > decades of programs that have used strlen() in various ways. Decades of programs that were variously limited or flawed. Most of them working only with a subset of English and English-like languages. > Have you never been given ANY kind of problem that involves analysis > of character strings? Ever? At all? Analysis? Yes, of course. Tons of them. They are all about SCANNING the string, not jumping randomly in it. > What if the question is "Find all the English words that have an E > in the 5th position and a U in the 7th"? Yes, what? Who would ever ask such a question? What is the point of such a question? The point of such a question is only to try and disprove my point, but my point is about useful operations, and therefore artificial questions like that will not dent it. > I mean, seriously, at some point you either have to accept that one > of our examples is good enough to justify the existence of strlen() > and character-based string indexing, or we just label you a loon and > ignore everything you say henceforth. To be honest, I do not care much what "you" think about me. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | deloptes <deloptes@gmail.com> |
|---|---|
| Date | 2018-04-04 23:20 +0200 |
| Message-ID | <vAXZ0-3zG-9@gated-at.bofh.it> |
| In reply to | #194497 |
Nicolas George wrote: >> What if the question is "Find all the English words that have an E >> in the 5th position and a U in the 7th"? > > Yes, what? Who would ever ask such a question? What is the point of such > a question? > > The point of such a question is only to try and disprove my point, but > my point is about useful operations, and therefore artificial questions > like that will not dent it. I agree with you, I get the point and it is correct. In the above example it is not clear first of all where do we look for those English words. Assume you have them in string, you again need the offset - start of word and then check the fifth and seventh position, so further iteration. I never bothered to look in stdc++ or libc how it is implemented - for example c++ string at() operation. Can someone enlight us pls?
[toc] | [prev] | [next] | [standalone]
| From | Richard Hector <richard@walnut.gen.nz> |
|---|---|
| Date | 2018-04-05 02:30 +0200 |
| Message-ID | <vB0WR-5Hj-1@gated-at.bofh.it> |
| In reply to | #194497 |
[Multipart message — attachments visible in raw view] — view raw
On 05/04/18 05:53, Nicolas George wrote: >> What if the question is "Find all the English words that have an E >> in the 5th position and a U in the 7th"? > > Yes, what? Who would ever ask such a question? What is the point of such > a question? Solving a crossword puzzle? Richard
[toc] | [prev] | [next] | [standalone]
Page 4 of 6 — ← Prev page 1 2 3 [4] 5 6 Next page →
Back to top | Article view | linux.debian.user
csiph-web