Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #194373 > unrolled thread

utf

Started bymess-mate <mess-mate@gmx.com>
First post2018-04-01 16:10 +0200
Last post2018-04-04 14:20 +0200
Articles 20 on this page of 101 — 21 participants

Back to article view | Back to linux.debian.user


Contents

  utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
    Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
    Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
    Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
    Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
      Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
        Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
      Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
        Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
                              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
                                Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
                                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
                          mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re:  utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
                            Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was:  Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
                    Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
                      Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
        Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
        Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
            Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
          Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
          Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
            Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
              Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
                Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
                    Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
                      Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
                            Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
                              Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
                              Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
                                Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
                            Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
                    Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
                      Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
                        Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
                          Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
                            Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
                              Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
                                Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
                                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
                                  Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
                                    Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
                                      Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
                                      Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
                              Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
                            Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
                Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
          Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200

Page 4 of 6 — ← Prev page 1 2 3 [4] 5 6  Next page →


#194508 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 21:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAWgx-2ha-9@gated-at.bofh.it>
In reply to#194500
On Wednesday, April 04, 2018 02:10:16 PM Jonathan de Boyne Pollard wrote:
> rhkramer:
> > The reason I wanted such a byte was to use it as a record separator in
> > a set of text files (that I use as an askSam "workalike" (or
> > "worksimilar") so that I could use msort (which depends on a 1 byte
> > record separator to --separate the records ;-) while sorting. Some of
> > the files already include UTF-8, and, in the future, I anticpate all
> > will be in UTFF-8.
> 
> Note that ISO 646, hence ISO 8859, hence ISO 10646, has had a
> single-byte Record Separator character since the 1960s.  (-:

Ok, thanks, I see that is Dec 30, Hex 1e.  

A quick look at the UTF-8 table in the Wikipedia article on UTF-8 seems to 
indicate that byte is a valid UTF-9 byte, which makes it unsuitabe for my use.

[toc] | [prev] | [next] | [standalone]


#194422

FromBen Caradoc-Davies <ben@transient.nz>
Date2018-04-02 23:40 +0200
Message-ID<vAflg-6VY-9@gated-at.bofh.it>
In reply to#194388
On 02/04/18 19:39, Andre Majorel wrote:
> On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote:
>> Why? UTF (especially UTF-8) is vastly superior for all purposes:
> I wouldn't say that. UTF-8 breaks a number of assumptions. For
> instance,
> 1) every character has the same size,
> 2) every byte sequence is a valid character,
> 3) the equality or inequality of two characters comes down to
>     the equality or inequality of the bytes they encode to.
> With ASCII and the many encodings based on it, most things can
> be done without having knowledge of the encoding. With UTF-8,
> even basic operations like determining the length of a string or
> reporting at what column an error occurred require knowledge of
> the encoding.

That is entirely correct. The problem is that ASCII has the worst 
assumption of all: everyone uses US English. This not valid in a 
Universal Operating System.

Kind regards,

-- 
Ben Caradoc-Davies <ben@transient.nz>
Director
Transient Software Limited <https://transient.nz/>
New Zealand

[toc] | [prev] | [next] | [standalone]


#194433

FromDarac Marjal <mailinglist@darac.org.uk>
Date2018-04-03 11:00 +0200
Message-ID<vApXj-5wU-1@gated-at.bofh.it>
In reply to#194388

[Multipart message — attachments visible in raw view] — view raw

On Mon, Apr 02, 2018 at 09:39:05AM +0200, Andre Majorel wrote:
>On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote:
>> On 02/04/18 02:05, mess-mate wrote:
>> >howto change the system utf to eu character set ?
>>
>> Why? UTF (especially UTF-8) is vastly superior for all purposes:
>
>I wouldn't say that. UTF-8 breaks a number of assumptions. For
>instance,
>1) every character has the same size,
>2) every byte sequence is a valid character,
>3) the equality or inequality of two characters comes down to
>   the equality or inequality of the bytes they encode to.

If these things matter to you, it's better to convert from UTF-8 to 
Unicode, first. I tend to think of Unicode as an arbitrarily large code 
page. Each character maps to a number, but that number could be 1, 1000 
or 500_000 (Unicode seems to be growing without might end in sight). 
Internally, you might store those code points as Integers or QUad Words 
or whatever you like. Only once you're ready to transfer the text to 
another process (print on screen, save to a file, stream across a 
network), do you convert the Unicode back into UTF-8.

Basically, you consider UTF-8 to be a transfer-only format (like 
Base64). If you want to do anything non-trivial with it, decode it into 
Unicode.

>
>With ASCII and the many encodings based on it, most things can
>be done without having knowledge of the encoding. With UTF-8,
>even basic operations like determining the length of a string or
>reporting at what column an error occurred require knowledge of
>the encoding.
>
>-- 
>André Majorel <http://www.teaser.fr/~amajorel/>
>Imagine what would happen if the Debian project disclosed the email
>addresses of their users. Spambots would harvest them and Debian
>users would be inundated with spam. Good thing they don't, eh ?
>

-- 
For more information, please reread.

[toc] | [prev] | [next] | [standalone]


#194434

FromRichard Hector <richard@walnut.gen.nz>
Date2018-04-03 11:20 +0200
Message-ID<vAqgG-5VJ-23@gated-at.bofh.it>
In reply to#194433

[Multipart message — attachments visible in raw view] — view raw

On 03/04/18 20:55, Darac Marjal wrote:
> If these things matter to you, it's better to convert from UTF-8 to
> Unicode, first. I tend to think of Unicode as an arbitrarily large code
> page. Each character maps to a number, but that number could be 1, 1000
> or 500_000 (Unicode seems to be growing without might end in sight).
> Internally, you might store those code points as Integers or QUad Words
> or whatever you like. Only once you're ready to transfer the text to
> another process (print on screen, save to a file, stream across a
> network), do you convert the Unicode back into UTF-8.
> 
> Basically, you consider UTF-8 to be a transfer-only format (like
> Base64). If you want to do anything non-trivial with it, decode it into
> Unicode.

Eh? UTF-8 is an encoding of Unicode. You can't "convert UTF-8 to
Unicode" - it already is Unicode. You could convert it to another
encoding, eg UTF-16 or UTF-32. Perhaps UTF-32 is what you mean, being
fixed-width.

Richard


[toc] | [prev] | [next] | [standalone]


#194435

From<tomas@tuxteam.de>
Date2018-04-03 11:30 +0200
Message-ID<vAqql-5Zj-7@gated-at.bofh.it>
In reply to#194434
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Tue, Apr 03, 2018 at 09:14:22PM +1200, Richard Hector wrote:
> On 03/04/18 20:55, Darac Marjal wrote:
> > If these things matter to you, it's better to convert from UTF-8 to
> > Unicode, first. I tend to think of Unicode as an arbitrarily large code
> > page. Each character maps to a number, but that number could be 1, 1000
> > or 500_000 (Unicode seems to be growing without might end in sight).
> > Internally, you might store those code points as Integers or QUad Words
> > or whatever you like. Only once you're ready to transfer the text to
> > another process (print on screen, save to a file, stream across a
> > network), do you convert the Unicode back into UTF-8.
> > 
> > Basically, you consider UTF-8 to be a transfer-only format (like
> > Base64). If you want to do anything non-trivial with it, decode it into
> > Unicode.
> 
> Eh? UTF-8 is an encoding of Unicode. You can't "convert UTF-8 to
> Unicode" - it already is Unicode. You could convert it to another
> encoding, eg UTF-16 or UTF-32. Perhaps UTF-32 is what you mean, being
> fixed-width.

I think Darac was talking about UTF-32 [1], which is a fixed-width encoding
of Unicode. Yes, Unicode is strictly speaking the abstract "mapping" between
integers ("code points") and characters. A computer has no integers...

What's curious is that there's no UTF-24 (although Unicode currently has
all its code points below 2^21). That would make for a slightly more
compact fixed-width encoding.

I think these days fixed-width encodings are losing their charm a bit,
since memory access is getting much more expensive than CPU power.

Things might change once again when the Chinese dominate culturally,
since UTF-8 plays its advantage only with ASCII dominated text.

But perhaps then, another encoding will make more sense. Or just UTF-24
is born, for a 25% savings :-)

Cheers

[1] https://en.wikipedia.org/wiki/UTF-32
- -- tomás


-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrDST0ACgkQBcgs9XrR2kadzQCeO8N1Kjua/p0aOdfE8QQTvv6R
PisAn0+0DbwWXb+cWYsUmMwqSqQN+BKQ
=y+Fu
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194437

FromNicolas George <george@nsup.org>
Date2018-04-03 12:20 +0200
Message-ID<vArcJ-6wn-1@gated-at.bofh.it>
In reply to#194433

[Multipart message — attachments visible in raw view] — view raw

> On Mon, Apr 02, 2018 at 09:39:05AM +0200, Andre Majorel wrote:
> >I wouldn't say that. UTF-8 breaks a number of assumptions. For
> >instance,
> >1) every character has the same size,
> >2) every byte sequence is a valid character,
> >3) the equality or inequality of two characters comes down to
> >  the equality or inequality of the bytes they encode to.

I am sure you do not realize that none of these assumptions are really
met by any encoding and none of these assumptions actually bring
something. They are just rehashed poor arguments to rationalize a fear
of change by people afraid their long-earned knowledge will become
obsolete. I do not know if you fit in that category. Odds are you have
just been misinformed after having normal trouble during the transition.

Darac Marjal (2018-04-03):
> If these things matter to you, it's better to convert from UTF-8 to Unicode,
> first.

"Convert to Unicode" does not mean anything. Unicode is not a format,
and therefore you cannot convert something to it.

Unicode is a catalog of "infolinguistic entities". I do not say
characters, because they are not all. Most of Unicode is characters, but
not all. As is stands, the principle of Unicode is that any "pure" text
can be represented as a sequence of Unicode code points.

For storing in a file or sending on the network, this sequence of code
points must be converted into a sequence of octets. UTF-8 is by far the
best choice for that, because it has many interesting properties. Other
encodings suffer from being incomplete, being incompatible with ASCII,
being sensible to endianness problems, being subject to corruption or
all of the above.

>	 I tend to think of Unicode as an arbitrarily large code page. Each
> character maps to a number, but that number could be 1, 1000 or 500_000
> (Unicode seems to be growing without might end in sight).

The twentieth century just called, it wanks the "code page" idiom back.

>							    Internally, you
> might store those code points as Integers or QUad Words or whatever you
> like. Only once you're ready to transfer the text to another process (print
> on screen, save to a file, stream across a network), do you convert the
> Unicode back into UTF-8.

Internally, this is a reasonable choice in some cases, but not actually
that useful in most. Your discourse is based on the assumption that
accessing a single Unicode code point in the string would be useful.
Most of the time, it is not. Remember that an Unicode code point is not
necessarily a character.

You need to decide the data structure based on the treatments you intend
to perform on the text. And actually, most of the time using an array of
octets in UTF-8 is the best choice for internal representation too.

> Basically, you consider UTF-8 to be a transfer-only format (like Base64). If
> you want to do anything non-trivial with it, decode it into Unicode.

No, definitely not. If somebody wants to do anything non-trivial with
text, then either they already know what they are doing better than
this, and do not need that advice, or they do not and they will get it
wrong.

Use a library. And use whatever text format that library use.

The problem does not come from UTF-8 or Unicode or anything
computer-related, the problem comes from the principle of written human
text: writing systems are insanely complex.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194463

FromBen Caradoc-Davies <ben@transient.nz>
Date2018-04-03 22:50 +0200
Message-ID<vAB2p-4vD-9@gated-at.bofh.it>
In reply to#194433
On 03/04/18 20:55, Darac Marjal wrote:
> If these things matter to you, it's better to convert from UTF-8 to 
> Unicode, first.

Fixed length encodings like UTF-32 will not fix broken assumptions about 
some relationship between byte length and number of characters because 
Unicode contains things like combining characters. What is the length of 
a string? Are you trying to count the number of glyphs? I do not think 
that you can do this by naïvely counting code points, regardless of 
encoding.

Because there is more than one way to represent an accented character, 
Unicode string comparison is nontrivial:
https://en.wikipedia.org/wiki/Unicode_equivalence

Kind regards,

-- 
Ben Caradoc-Davies <ben@transient.nz>
Director
Transient Software Limited <https://transient.nz/>
New Zealand

[toc] | [prev] | [next] | [standalone]


#194464

FromNicolas George <george@nsup.org>
Date2018-04-03 23:00 +0200
Message-ID<vABc6-4A2-7@gated-at.bofh.it>
In reply to#194463

[Multipart message — attachments visible in raw view] — view raw

Ben Caradoc-Davies (2018-04-04):
>						     What is the length of a
> string?

When is that relevant?

>	  Are you trying to count the number of glyphs?

What for?

>							I do not think that
> you can do this by naïvely counting code points, regardless of encoding.

Fortunately, this is almost never useful.

> Because there is more than one way to represent an accented character,
> Unicode string comparison is nontrivial:

Fortunately, it is also rather rarely useful in all that complexity.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194465

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-03 23:00 +0200
Message-ID<vABc6-4A2-9@gated-at.bofh.it>
In reply to#194464
On Tue, Apr 03, 2018 at 10:51:43PM +0200, Nicolas George wrote:
> Ben Caradoc-Davies (2018-04-04):
> >						     What is the length of a
> > string?
> 
> When is that relevant?

When you're trying to display one on a screen, or print one on paper.

When you've been asked to find the longest/shortest string from a list.

When you've been asked to sort a list of strings by length.

Or in other words, basically every time you do anything with the string
at all other than blindly byte-copying it to a different place in memory.

[toc] | [prev] | [next] | [standalone]


#194466

FromNicolas George <george@nsup.org>
Date2018-04-03 23:30 +0200
Message-ID<vABFa-50o-59@gated-at.bofh.it>
In reply to#194465

[Multipart message — attachments visible in raw view] — view raw

Greg Wooledge (2018-04-03):
> When you're trying to display one on a screen, or print one on paper.

With just the length? You will not get anything done. To properly
display a string, you need to handle ligatures, right-to-left, kerning,
etc. The length of the string is barely relevant.

> When you've been asked to find the longest/shortest string from a list.

Why would that be useful?

> When you've been asked to sort a list of strings by length.

Why would that be useful?

> Or in other words, basically every time you do anything with the string
> at all other than blindly byte-copying it to a different place in memory.

No, the length of the string is hardly relevant, and when it is it is
not enough anyway.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194486

Fromdeloptes <deloptes@gmail.com>
Date2018-04-04 19:00 +0200
Message-ID<vATVo-Ax-11@gated-at.bofh.it>
In reply to#194466
Nicolas George wrote:

> No, the length of the string is hardly relevant, and when it is it is
> not enough anyway.

@Nicolas, I think OP does not understand you - perhaps it is not worth the
effort. My impression is that you refer to a string (properly) as sequence
of bytes and other refer to it as number of chars, which is not consistant
with utf.

>From my work with UTF, it is possible but not satisfying to guess encoding.
I wonder why no one suggested a kind of markup (xml) instead of byte
delimiter.

And regarding the mbox thing, well mbox was depreciated for many reasons. I
guess if it was that good it wouldn't be depreciated.

@OP, at some point of time everyone has to redesign and reimplement because
technology evolves and all the tools listed can be updated to the new
format.
Recent example of redesign I worked with is gnupg - huge changes from v1.x
to v2.1

regards

[toc] | [prev] | [next] | [standalone]


#194487

FromNicolas George <george@nsup.org>
Date2018-04-04 19:10 +0200
Message-ID<vAU54-Tb-21@gated-at.bofh.it>
In reply to#194486

[Multipart message — attachments visible in raw view] — view raw

deloptes (2018-04-04):
> @Nicolas, I think OP does not understand you - perhaps it is not worth the
> effort. My impression is that you refer to a string (properly) as sequence
> of bytes and other refer to it as number of chars, which is not consistant
> with utf.

Not at all, I am well speaking of text string, made up of characters,
expressed as sequences of Unicode code point, and encoded as any data
structure convenient.

What I am trying to explain (not to the OP who is focussing on another
thing entirely and obviously beyond help anyway), is that if you are
thinking of stings in terms of "access the n-th char", "find the length
of the string in chars", etc., then you completely have missed the point
of them.

Find me a case where you need to access the n-th char of a string, with
n completely out of the blue, and I will explain how somebody botched
their design.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194488

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-04 19:30 +0200
Message-ID<vAUoq-11h-11@gated-at.bofh.it>
In reply to#194487
On Wed, Apr 04, 2018 at 07:07:01PM +0200, Nicolas George wrote:
> Find me a case where you need to access the n-th char of a string, with
> n completely out of the blue, and I will explain how somebody botched
> their design.

Does it count if we want the 1st char, then the 2nd char, then the 3rd
char, then the 4th char, and so on?  Or is that not blue enough?

How about the last char?  Or the last two chars?  Or all chars starting
just after the last slash or period?

How about performing a checksum like the
<https://en.wikipedia.org/wiki/Luhn_algorithm> on a user input string
which is supposed to be a 10-digit
<https://en.wikipedia.org/wiki/National_Provider_Identifier> ?

Might it be useful to check the length of an input string before
bothering to decompose it into individual digits and perform the
arithmetic?  And here, by "length", I mean "number of characters".
You can see how that might be a handy thing, right?

Length as in "number of bytes required to store it" is also an
important value, of course.

Character length is also useful when displaying strings on
CHARACTER-ORIENTED OUTPUT MEDIA.  Like terminals.  You know, those
things that Unix-like systems use all the time?  How else are you
going to space-pad the fields so that the output columns line up,
if you don't know how many extra spaces you need, because you don't
know the length of the string?

(It's frankly disturbing to me that when I talked about length being
relevant when printing strings, you immediately jumped to "pixels" and
"fonts".  This tells me that you no longer accept the terminal as your
lord and savior.  If you ever did.)

All of these things matter, and are real, and don't necessarily indicate
"botched design".

[toc] | [prev] | [next] | [standalone]


#194491

FromNicolas George <george@nsup.org>
Date2018-04-04 19:40 +0200
Message-ID<vAUy6-14B-11@gated-at.bofh.it>
In reply to#194488

[Multipart message — attachments visible in raw view] — view raw

Greg Wooledge (2018-04-04):
> Does it count if we want the 1st char, then the 2nd char, then the 3rd
> char, then the 4th char, and so on?  Or is that not blue enough?

It is not out of the blue, it is in sequence.

> How about the last char?  Or the last two chars?

Ditto.

>						    Or all chars starting
> just after the last slash or period?

There is no n here, you have to first look for the last slash.


> How about performing a checksum like the
> <https://en.wikipedia.org/wiki/Luhn_algorithm> on a user input string
> which is supposed to be a 10-digit
> <https://en.wikipedia.org/wiki/National_Provider_Identifier> ?

Again, in sequence.

> Might it be useful to check the length of an input string before
> bothering to decompose it into individual digits and perform the
> arithmetic?

And to check it is made of all digits. Checking the length is only a
byproduct.

>	       And here, by "length", I mean "number of characters".
> You can see how that might be a handy thing, right?

Still, no.

> Length as in "number of bytes required to store it" is also an
> important value, of course.

You have to compute it to store it, indeed.

> Character length is also useful when displaying strings on
> CHARACTER-ORIENTED OUTPUT MEDIA.  Like terminals.  You know, those
> things that Unix-like systems use all the time?  How else are you
> going to space-pad the fields so that the output columns line up,
> if you don't know how many extra spaces you need, because you don't
> know the length of the string?

I do not know. Please tell me, how do you handle control characters,
escape sequence, double-width characters, etc., without walking the
string in sequence?

> (It's frankly disturbing to me that when I talked about length being
> relevant when printing strings, you immediately jumped to "pixels" and
> "fonts".  This tells me that you no longer accept the terminal as your
> lord and savior.  If you ever did.)

I do not have a "lord and savior". I use the terminal a lot, but I am
aware of the hidden complexities.

> All of these things matter, and are real, and don't necessarily indicate
> "botched design".

All these things matter, but they do not require random access in a
string by char number.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194489

Fromdeloptes <deloptes@gmail.com>
Date2018-04-04 19:40 +0200
Message-ID<vAUy5-14B-1@gated-at.bofh.it>
In reply to#194487
Nicolas George wrote:

> Find me a case where you need to access the n-th char of a string, with
> n completely out of the blue, and I will explain how somebody botched
> their design.

ok, thanks. I understood the part above, but not sure if I understand this
part. A standard text editing operation is find and replace, where you get
the start and end point in the string. Of course it is not "n completely
out of the blue".

regards

[toc] | [prev] | [next] | [standalone]


#194492

FromNicolas George <george@nsup.org>
Date2018-04-04 19:40 +0200
Message-ID<vAUy6-14B-9@gated-at.bofh.it>
In reply to#194489

[Multipart message — attachments visible in raw view] — view raw

deloptes (2018-04-04):
> ok, thanks. I understood the part above, but not sure if I understand this
> part. A standard text editing operation is find and replace, where you get
> the start and end point in the string. Of course it is not "n completely
> out of the blue".

I am not sure exactly what is your example, but you got its flaw right:
n is not out of the blue, it was obtained by previously walking the
string. And in that case, you have all freedom to express n as a more
convenient entity than an index expressed in terms of chars. A pointer
maybe, or a pair with both the char index and the octet offset.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194494

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-04 19:50 +0200
Message-ID<vAUHM-17T-5@gated-at.bofh.it>
In reply to#194492
On Wed, Apr 04, 2018 at 07:35:37PM +0200, Nicolas George wrote:
> I am not sure exactly what is your example, but you got its flaw right:
> n is not out of the blue, it was obtained by previously walking the
> string. And in that case, you have all freedom to express n as a more
> convenient entity than an index expressed in terms of chars. A pointer
> maybe, or a pair with both the char index and the octet offset.

The problem is, you reject every single example that everyone gives
you.  I don't know what you expect from us.

You just seem to have Decided, for reasons known only to you, that
The Character Length Of A String Is Not Useful.  Despite literally
decades of programs that have used strlen() in various ways.

Have you never been given ANY kind of problem that involves analysis
of character strings?  Ever?  At all?

What if the question is "Find all the English words that have an E
in the 5th position and a U in the 7th"?

I mean, seriously, at some point you either have to accept that one
of our examples is good enough to justify the existence of strlen()
and character-based string indexing, or we just label you a loon and
ignore everything you say henceforth.

I think most of the REST of us can agree that the character length
and the byte length of a string are BOTH useful quantities, important
in countless ways to countless programs.  And that sometimes your
algorithm wants to treat a string as an indexed array of characters,
and retrieve the nth character.  And that we shouldn't have to dig
up an entire textbook worth of examples to explain this.

[toc] | [prev] | [next] | [standalone]


#194497

FromNicolas George <george@nsup.org>
Date2018-04-04 20:00 +0200
Message-ID<vAURs-1cT-7@gated-at.bofh.it>
In reply to#194494

[Multipart message — attachments visible in raw view] — view raw

Greg Wooledge (2018-04-04):
> The problem is, you reject every single example that everyone gives
> you.

I do not reject them, I refute them.

> I don't know what you expect from us.

Acknowledge that I am right once I have refuted all your examples and
you have eventually understood my point.

At this time, you have not yet understood.

> You just seem to have Decided, for reasons known only to you, that
> The Character Length Of A String Is Not Useful.  Despite literally
> decades of programs that have used strlen() in various ways.

Decades of programs that were variously limited or flawed. Most of them
working only with a subset of English and English-like languages.

> Have you never been given ANY kind of problem that involves analysis
> of character strings?  Ever?  At all?

Analysis? Yes, of course. Tons of them. They are all about SCANNING the
string, not jumping randomly in it.

> What if the question is "Find all the English words that have an E
> in the 5th position and a U in the 7th"?

Yes, what? Who would ever ask such a question? What is the point of such
a question?

The point of such a question is only to try and disprove my point, but
my point is about useful operations, and therefore artificial questions
like that will not dent it.

> I mean, seriously, at some point you either have to accept that one
> of our examples is good enough to justify the existence of strlen()
> and character-based string indexing, or we just label you a loon and
> ignore everything you say henceforth.

To be honest, I do not care much what "you" think about me.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194511

Fromdeloptes <deloptes@gmail.com>
Date2018-04-04 23:20 +0200
Message-ID<vAXZ0-3zG-9@gated-at.bofh.it>
In reply to#194497
Nicolas George wrote:

>> What if the question is "Find all the English words that have an E
>> in the 5th position and a U in the 7th"?
> 
> Yes, what? Who would ever ask such a question? What is the point of such
> a question?
> 
> The point of such a question is only to try and disprove my point, but
> my point is about useful operations, and therefore artificial questions
> like that will not dent it.

I agree with you, I get the point and it is correct.
In the above example it is not clear first of all where do we look for those
English words. Assume you have them in string, you again need the offset -
start of word and then check the fifth and seventh position, so further
iteration. 
I never bothered to look in stdc++ or libc how it is implemented - for
example c++ string at() operation. Can someone enlight us pls?

[toc] | [prev] | [next] | [standalone]


#194518

FromRichard Hector <richard@walnut.gen.nz>
Date2018-04-05 02:30 +0200
Message-ID<vB0WR-5Hj-1@gated-at.bofh.it>
In reply to#194497

[Multipart message — attachments visible in raw view] — view raw

On 05/04/18 05:53, Nicolas George wrote:
>> What if the question is "Find all the English words that have an E
>> in the 5th position and a U in the 7th"?
>
> Yes, what? Who would ever ask such a question? What is the point of such
> a question?

Solving a crossword puzzle?

Richard

[toc] | [prev] | [next] | [standalone]


Page 4 of 6 — ← Prev page 1 2 3 [4] 5 6  Next page →

Back to top | Article view | linux.debian.user


csiph-web