Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #194373 > unrolled thread

utf

Started bymess-mate <mess-mate@gmx.com>
First post2018-04-01 16:10 +0200
Last post2018-04-04 14:20 +0200
Articles 20 on this page of 101 — 21 participants

Back to article view | Back to linux.debian.user


Contents

  utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
    Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
    Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
    Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
    Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
      Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
        Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
      Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
        Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
                              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
                                Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
                                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
                          mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re:  utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
                            Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was:  Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
                    Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
                      Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
        Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
        Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
            Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
          Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
          Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
            Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
              Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
                Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
                    Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
                      Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
                            Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
                              Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
                              Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
                                Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
                            Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
                    Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
                      Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
                        Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
                          Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
                            Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
                              Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
                                Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
                                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
                                  Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
                                    Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
                                      Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
                                      Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
                              Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
                            Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
                Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
          Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200

Page 1 of 6  [1] 2 3 4 5 6  Next page →


#194373 — utf

Frommess-mate <mess-mate@gmx.com>
Date2018-04-01 16:10 +0200
Subjectutf
Message-ID<vzLQd-4cg-1@gated-at.bofh.it>
Hi,

howto change the system utf to eu character set ?

regards

[toc] | [next] | [standalone]


#194377

FromCurt <curty@free.fr>
Date2018-04-01 17:50 +0200
Message-ID<vzNoZ-5bx-1@gated-at.bofh.it>
In reply to#194373
On 2018-04-01, mess-mate <mess-mate@gmx.com> wrote:
> Hi,
>
> howto change the system utf to eu character set ?

https://wiki.debian.org/Locale

(if it's not wrong or hopelessly out of date).

The last I checked--and as far as I know--(with administrative privileges),

  dpkg-reconfigure locales
 
was the way to go.

Now you tell us how you solved the 3 screens problem (you know, for
posterity).

> regards
>
>


-- 
The time-state of attainment eliminates so accurately the time-state of
aspiration, that the actual seems the inevitable, and, all conscious
intellectual effort to reconstitute the invisible and unthinkable as a reality
being fruitless, we are incapable of appreciating our joy by comparing it with
our sorrow.  --Samuel Beckett

[toc] | [prev] | [next] | [standalone]


#194378

FromIonel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl>
Date2018-04-01 18:30 +0200
Message-ID<vzO1H-5Qv-1@gated-at.bofh.it>
In reply to#194373
On  1-04-2018, at 16h 05'41", mess-mate wrote about "utf"
> Hi,
> 
> howto change the system utf to eu character set ?
> 
> regards


What is "eu character set"? 


None of the Latin-x or iso8859-y can cover all the gliphs of all the
countries in Europe. So, without knowing what do you think is "eu
character set", I would say that you must stay with UTF-8. I really
mean this.

Ionel

[toc] | [prev] | [next] | [standalone]


#194379

FromAndre Majorel <aym-naibed@teaser.fr>
Date2018-04-01 18:40 +0200
Message-ID<vzObn-5UZ-3@gated-at.bofh.it>
In reply to#194373
On 2018-04-01 16:05 +0200, mess-mate wrote:

> howto change the system utf to eu character set ?

If, by "eu", you mean ISO 8859-1 or -15, here is the procedure
that works for me :

1) Run

     dpkg-reconfigure locales

   and select some appropriate non-UTF-8 locales. When asked
   what the default locale should be, select a non-UTF-8 one.

2) In your ~/.profile, ~/.bashrc or whatever is relevant, make
   sure that none of LANG, LC_ALL, LC_CTYPE, LC_MESSAGES and
   friends name a UTF-8 locale. You may not have to be so
   drastic but then you will have to understand the somewhat
   counter-intuitive rules of precedence among those variables.

3) Run

     dpkg-reconfigure console-setup

   and make some sensible choices. Annoyingly, the scripts or
   programs that configure the console (/dev/tty1 .. /dev/tty6)
   in Debian set the keyboard to UTF-8 despite your preferences.
   The incantation to (temporarily) fix that is

     setupcon -k -v
     stty -iutf8

Note that I don't use any fancy desktop environments. I wouldn't
be surprised to learn that those who do must follow different
procedures.

-- 
André Majorel <http://www.teaser.fr/~amajorel/>
lists.debian.org, an essential online resource for spammers.

[toc] | [prev] | [next] | [standalone]


#194384

FromBen Caradoc-Davies <ben@transient.nz>
Date2018-04-01 22:10 +0200
Message-ID<vzRsC-8m1-3@gated-at.bofh.it>
In reply to#194373
On 02/04/18 02:05, mess-mate wrote:
> howto change the system utf to eu character set ?

Why? UTF (especially UTF-8) is vastly superior for all purposes:
http://utf8everywhere.org/

What are you trying to do, and why do you thing a non-UTF encoding might 
help you to do it? You will likely be able to achieve your goals with a 
UTF locale.

Kind regards,

-- 
Ben Caradoc-Davies <ben@transient.nz>
Director
Transient Software Limited <https://transient.nz/>
New Zealand

[toc] | [prev] | [next] | [standalone]


#194386

FromCindy-Sue Causey <butterflybytes@gmail.com>
Date2018-04-02 00:50 +0200
Message-ID<vzTXs-1iJ-1@gated-at.bofh.it>
In reply to#194384
On 4/1/18, Ben Caradoc-Davies <ben@transient.nz> wrote:
> On 02/04/18 02:05, mess-mate wrote:
>> howto change the system utf to eu character set ?
>
> Why? UTF (especially UTF-8) is vastly superior for all purposes:
> http://utf8everywhere.org/
>
> What are you trying to do, and why do you thing a non-UTF encoding might
> help you to do it? You will likely be able to achieve your goals with a
> UTF locale.


Sometimes it takes seeing certain words together to trigger a thought.
Ben's words made me wonder if maybe OP is missing a font package. That
thought occurred because I use UTF-8 and find fonts missing sometimes
(too?)....

Cindy :)
-- 
Cindy-Sue Causey
Talking Rock, Pickens County, Georgia, USA

* hops with duct tape *

[toc] | [prev] | [next] | [standalone]


#194387

FromCurt <curty@free.fr>
Date2018-04-02 09:50 +0200
Message-ID<vA2o1-6N4-1@gated-at.bofh.it>
In reply to#194386
On 2018-04-01, Cindy-Sue Causey <butterflybytes@gmail.com> wrote:
> On 4/1/18, Ben Caradoc-Davies <ben@transient.nz> wrote:
>> On 02/04/18 02:05, mess-mate wrote:
>>> howto change the system utf to eu character set ?
>>
>> Why? UTF (especially UTF-8) is vastly superior for all purposes:
>> http://utf8everywhere.org/
>>
>> What are you trying to do, and why do you thing a non-UTF encoding might
>> help you to do it? You will likely be able to achieve your goals with a
>> UTF locale.
>
>
> Sometimes it takes seeing certain words together to trigger a thought.
> Ben's words made me wonder if maybe OP is missing a font package. That
> thought occurred because I use UTF-8 and find fonts missing sometimes
> (too?)....
>
> Cindy :)


The thought provoked in my neurological matter was why there are other
locales at all if UTF8 (the locale of this here .homie machine, BTW) is
"vastly superior for all purposes".

That leaves no purposes remaining whatsoever for the myriad other
locales.

If this is indeed so, let's get rid of them, then, (the superfluous
locales) thus lightening our loads in computer life. 

Or are there legacy, corner, arcane purposes for these locales, of which
the hoi polloi (of which I am a card-carrying member) has not to concern
itself? 

-- 
The time-state of attainment eliminates so accurately the time-state of
aspiration, that the actual seems the inevitable, and, all conscious
intellectual effort to reconstitute the invisible and unthinkable as a reality
being fruitless, we are incapable of appreciating our joy by comparing it with
our sorrow.  --Samuel Beckett

[toc] | [prev] | [next] | [standalone]


#194390

FromRichard Hector <richard@walnut.gen.nz>
Date2018-04-02 11:40 +0200
Message-ID<vA46u-7WZ-7@gated-at.bofh.it>
In reply to#194387

[Multipart message — attachments visible in raw view] — view raw

On 02/04/18 19:43, Curt wrote:
> The thought provoked in my neurological matter was why there are other
> locales at all if UTF8 (the locale of this here .homie machine, BTW) is
> "vastly superior for all purposes".

There's more to the locale than the character set - things like default
language, thousands separators, currency symbols, collation order ...

Richard

[toc] | [prev] | [next] | [standalone]


#194396

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-02 15:30 +0200
Message-ID<vA7H3-1Tr-3@gated-at.bofh.it>
In reply to#194387
On Mon, Apr 02, 2018 at 07:43:23AM +0000, Curt wrote:
> The thought provoked in my neurological matter was why there are other
> locales at all if UTF8 (the locale of this here .homie machine, BTW) is
> "vastly superior for all purposes".
> 
> That leaves no purposes remaining whatsoever for the myriad other
> locales.
> 
> If this is indeed so, let's get rid of them, then, (the superfluous
> locales) thus lightening our loads in computer life. 
> 
> Or are there legacy, corner, arcane purposes for these locales, of which
> the hoi polloi (of which I am a card-carrying member) has not to concern
> itself? 

Yes, there are other computer systems on the planet, and some of them
still use various single-byte character sets (ISO 8859-* or WIN1252
or similar).

In heterogeneous environments, it is often important for a Debian server
to support multiple locales, to fit the needs of legacy client systems.

For example, if you're sitting at a legacy Unix system which uses the
ISO 8859-1 character set, and you ssh to a Debian system that hard-codes
"LANG=en_US.UTF-8" in /etc/default/locale , this will clobber the locale
variables that are sent by the ssh client.  Then, commands like "man"
will generate output containing UTF-8 characters, which will seriously
mess up the display on an ISO 8859-1 terminal emulator.

(Why "man", or groff or whatever, feels the need to generate non-ASCII
characters is beyond me.)

This is why <https://wiki.debian.org/Locale> has complicated instructions
that DO NOT force an override LANG=... in /etc/default/locale (which is
destructive and wrong), but instead use fallback values in /etc/profile
and /etc/csh/login.d/lang .

This is also why I was probing login.conf and pam_env.conf for the ability
to support fallback/default environment variable values in that other
thread last week.  (Sadly, there is still no known alternative to the
login-shell-dependent hackery.)

[toc] | [prev] | [next] | [standalone]


#194388

FromAndre Majorel <aym-naibed@teaser.fr>
Date2018-04-02 09:50 +0200
Message-ID<vA2o1-6N4-3@gated-at.bofh.it>
In reply to#194384
On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote:
> On 02/04/18 02:05, mess-mate wrote:
> >howto change the system utf to eu character set ?
> 
> Why? UTF (especially UTF-8) is vastly superior for all purposes:

I wouldn't say that. UTF-8 breaks a number of assumptions. For
instance,
1) every character has the same size,
2) every byte sequence is a valid character,
3) the equality or inequality of two characters comes down to
   the equality or inequality of the bytes they encode to.

With ASCII and the many encodings based on it, most things can
be done without having knowledge of the encoding. With UTF-8,
even basic operations like determining the length of a string or
reporting at what column an error occurred require knowledge of
the encoding.

-- 
André Majorel <http://www.teaser.fr/~amajorel/>
Imagine what would happen if the Debian project disclosed the email
addresses of their users. Spambots would harvest them and Debian
users would be inundated with spam. Good thing they don't, eh ?

[toc] | [prev] | [next] | [standalone]


#194391 — Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-02 14:40 +0200
SubjectInvalid UTF-8 byte? (was: Re: utf)
Message-ID<vA6UG-1kM-1@gated-at.bofh.it>
In reply to#194388
On Monday, April 02, 2018 03:39:05 AM Andre Majorel wrote:
> > Why? UTF (especially UTF-8) is vastly superior for all purposes:
> I wouldn't say that. UTF-8 breaks a number of assumptions. For
> instance,
> 1) every character has the same size,
> 2) every byte sequence is a valid character,

A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally 
invalid byte (not an invalid sequence of bytes).  At the time, I tried some 
googling, but it looked rather hopeless (maybe it was my googling that was 
hopeless).

I know that your statement does not imply there is such a byte, but maybe you 
(or someone else reading this) know(s)?

(The reason I wanted such a byte was to use it as a record separator in a set 
of text files (that I use as an askSam "workalike" (or "worksimilar") so that I 
could use msort (which depends on a 1 byte record separator to --separate the 
records ;-) while sorting.)  (Some of the files already include UTF-8, and, in 
the future, I anticpate all will be in UTFF-8.)



> 3) the equality or inequality of two characters comes down to
>    the equality or inequality of the bytes they encode to.

[toc] | [prev] | [next] | [standalone]


#194392 — Re: Invalid UTF-8 byte? (was: Re: utf)

From<tomas@tuxteam.de>
Date2018-04-02 15:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vA7nH-1KQ-7@gated-at.bofh.it>
In reply to#194391
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Mon, Apr 02, 2018 at 08:37:54AM -0400, rhkramer@gmail.com wrote:
> On Monday, April 02, 2018 03:39:05 AM Andre Majorel wrote:
> > > Why? UTF (especially UTF-8) is vastly superior for all purposes:
> > I wouldn't say that. UTF-8 breaks a number of assumptions. For
> > instance,
> > 1) every character has the same size,
> > 2) every byte sequence is a valid character,
> 
> A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally 
> invalid byte (not an invalid sequence of bytes).

If you look at man utf-8 (7), you'll find, for the encoding:

   Encoding
       The following byte sequences are used to represent a character.
       The sequence to be used depends on the UCS code number of the character:

       0x00000000 - 0x0000007F:
           0xxxxxxx

       0x00000080 - 0x000007FF:
           110xxxxx 10xxxxxx

       0x00000800 - 0x0000FFFF:
           1110xxxx 10xxxxxx 10xxxxxx

       0x00010000 - 0x001FFFFF:
           11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

       0x00200000 - 0x03FFFFFF:
           111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx

       0x04000000 - 0x7FFFFFFF:
           1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx

This means that 1111111x are the only (two) illegal bytes in UTF-8,
at least currently. I don't know what will happen once we need code
points beyond 0x7fffffff -- perhaps Klingon has an ideographic
variant our linguists don't know of yet (the alphabetical variant
is in the private area and an official place seems to be in the
works [1]).

So I wouldn't build my house on it. And oh, there are better
serialization "protocols" than using an arbitrary record
separator. Why not use some of the lower ASCII thingies plus
an escape mechanism?

Cheers

[1] http://www.klingonwiki.net/En/Unicode
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrCKdUACgkQBcgs9XrR2kbBnQCfe5c4WVNYCcpZbsgg5dDZwBHR
XNkAn2CSUkMY59VE2zLciII/3kUz8W45
=aPFG
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194394 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromHenrique de Moraes Holschuh <hmh@debian.org>
Date2018-04-02 15:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vA7nH-1KQ-9@gated-at.bofh.it>
In reply to#194391
On Mon, 02 Apr 2018, rhkramer@gmail.com wrote:
> A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally 
> invalid byte (not an invalid sequence of bytes).  At the time, I tried some 
> googling, but it looked rather hopeless (maybe it was my googling that was 
> hopeless).

0xff should work.  But any of those in RED on the wikipedia article
about UTF-8 would do for Unicode text:

https://en.wikipedia.org/wiki/UTF-8

-- 
  Henrique Holschuh

[toc] | [prev] | [next] | [standalone]


#194411 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-02 19:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAbKF-4vW-3@gated-at.bofh.it>
In reply to#194394
Thanks to tomas and Henrique!

The wikipedia article is rather interesting, in a quick skim, I learned some 
interesting things about UTF-8, especially the property of self-
synchronization.

I had trouble reading that large table--but if I simply take the red boxes at 
face value, maybe there are 10 or so bytes that are not valid UTF-8.  I'll 
probably first consider the bytes that tomas also mentions, i.e., decimal 254 
and 255).

I guess I have a followup question--are those two bytes (or either one of 
them) also unused in all possible "code pages"?  

The problem is that I copy snippets of text from all kinds of sources into 
those text files (which are formatted like mbox files), so I might find one or 
both of those bytes in the file already.

I guess it's not a big deal as, I will either:

   * search the file (with a hex editor, I guess) to see if decimal 254 or 255 
is in use already--if only one or a few cases, I might replace it with 
something else before adding additional instances to serve as a temporary 
record separator (for use by msort), or

   * use one of the other utilities that I've since found which can apparently 
sort mbox files while keeping emails intact (I have to read up on those (or 
it?) again as there were, iirc, also some limitations there that might not let 
me accomplish what I want (usually, sorting the emails by the title in the 
mbox "From " header (and ususally not by the email From: header).

Thanks again!


On Monday, April 02, 2018 09:05:52 AM Henrique de Moraes Holschuh wrote:
> On Mon, 02 Apr 2018, rhkramer@gmail.com wrote:
> > A few weeks ago, I was looking for a byte that, in UTF-8, would be a
> > totally invalid byte (not an invalid sequence of bytes).  At the time, I
> > tried some googling, but it looked rather hopeless (maybe it was my
> > googling that was hopeless).
> 
> 0xff should work.  But any of those in RED on the wikipedia article
> about UTF-8 would do for Unicode text:
> 
> https://en.wikipedia.org/wiki/UTF-8

[toc] | [prev] | [next] | [standalone]


#194413 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromHenrique de Moraes Holschuh <hmh@debian.org>
Date2018-04-02 20:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAcdI-4Wt-9@gated-at.bofh.it>
In reply to#194411
On Mon, 02 Apr 2018, rhkramer@gmail.com wrote:
> The wikipedia article is rather interesting, in a quick skim, I learned some 
> interesting things about UTF-8, especially the property of self-
> synchronization.

Yes, UTF-8 is a brilliant design.

> I had trouble reading that large table--but if I simply take the red boxes at 
> face value, maybe there are 10 or so bytes that are not valid UTF-8.  I'll 
> probably first consider the bytes that tomas also mentions, i.e., decimal 254 
> and 255).

On that table, columns are the least significant bits (second hex
digit), and rows are the most significant bits (first hex digit) of a
byte.  As in C1 is row C, column 1.

The "2-byte", "3-byte" and "4-byte" are comments that remind you of the
self-sinchronizing nature of UTF-8, and that these bytes would be
invalid outside of that position in an UTF-8 sequence that encodes a
single code point (but they would be valid in the correct position).

The stuff in "red" on that table is always invalid for Unicode: if you
find one of those in a data file, that file is *not* valid UTF-8 (but it
could be valid UTF-16, valid UTF-32, or valid ISO-8859-*, etc).

> I guess I have a followup question--are those two bytes (or either one of 
> them) also unused in all possible "code pages"?  

For Unicode, yes, because Unicode can't go past code point 0x10ffff.
And that isn't about to change anytime soon (lots of stuff hardcode it
somehow, e.g., by limiting the number of UTF-8 bytes that can be used to
encode a single code point...).  I have not read the Unicode standard to
check what it says about future expansions related to the valid code
point range, though.

> The problem is that I copy snippets of text from all kinds of sources into 
> those text files (which are formatted like mbox files), so I might find one or 
> both of those bytes in the file already.

Then it isn't a valid unicode text file in UTF-8 format, and it needs to
be converted (or fixed) first to be encoded in UTF-8 :-)

-- 
  Henrique Holschuh

[toc] | [prev] | [next] | [standalone]


#194415 — Re: Invalid UTF-8 byte? (was: Re: utf)

From<tomas@tuxteam.de>
Date2018-04-02 20:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAcGJ-58A-11@gated-at.bofh.it>
In reply to#194413
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote:
> On Mon, 02 Apr 2018, rhkramer@gmail.com wrote:
> > The wikipedia article is rather interesting, in a quick skim, I learned some 
> > interesting things about UTF-8, especially the property of self-
> > synchronization.
> 
> Yes, UTF-8 is a brilliant design.

Possibly relevant, definitely entertaining, Rob Pike's account
of UTF-8's gestation [1]

Yeah. Elegant design. Until the Unicode Consortium left Microsoft
near it (Byte Order Mark, I'm looking at you!).

[...]

> > I guess I have a followup question--are those two bytes (or either one of 
> > them) also unused in all possible "code pages"?  

I'm not sure what you mean here: there are two layers at work (at least
if you have UTF-8 encoded Unicode). As Henrique says, if you assume
both to be "correct" then you get more illegal things. But sometimes
UTF-8 encoding is used for other things (notably Emacs encodes a superset
of Unicode, to be able to express "raw byte values" next to "Unicode
characters".

> > The problem is that I copy snippets of text from all kinds of sources into 
> > those text files (which are formatted like mbox files), so I might find one or 
> > both of those bytes in the file already.
> 
> Then it isn't a valid unicode text file in UTF-8 format, and it needs to
> be converted (or fixed) first to be encoded in UTF-8 :-)

Agreed: if you don't know what's coming in, you better plan for anything :)

Cheers
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrCeTcACgkQBcgs9XrR2kbRtgCfaRHoodlkFFt8Gm0Oq438ymvg
0oMAn2NkpsqMJ3Tcy5BvAJIpTvfG8mdj
=iVqF
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194417 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromtomas@tuxteam.de
Date2018-04-02 21:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAcQp-5c2-3@gated-at.bofh.it>
In reply to#194415
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Mon, Apr 02, 2018 at 08:40:55PM +0200, tomas@tuxteam.de wrote:
> On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote:
> > On Mon, 02 Apr 2018, rhkramer@gmail.com wrote:
> > > The wikipedia article is rather interesting, in a quick skim, I learned some 
> > > interesting things about UTF-8, especially the property of self-
> > > synchronization.
> > 
> > Yes, UTF-8 is a brilliant design.
> 
> Possibly relevant, definitely entertaining, Rob Pike's account
> of UTF-8's gestation [1]

Oops: here's the ref:

[1] http://doc.cat-v.org/bell_labs/utf-8_history
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrCe8EACgkQBcgs9XrR2kY5mQCfQpQE1lniIWWWMXNm79oETOhP
IoQAn3Sm9i6NFwPyPe1XlFrqNTMFyB8M
=oE1V
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194418 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-02 21:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAd06-5v7-19@gated-at.bofh.it>
In reply to#194415
Thanks, again, to Henrique and tomas for the followups!

On Monday, April 02, 2018 02:40:55 PM tomas@tuxteam.de wrote:
> On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote:

[toc] | [prev] | [next] | [standalone]


#194424 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromMichael Lange <klappnase@freenet.de>
Date2018-04-03 00:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAgqZ-7Cc-3@gated-at.bofh.it>
In reply to#194391
Hi,

On Mon, 2 Apr 2018 08:37:54 -0400
rhkramer@gmail.com wrote:

> A few weeks ago, I was looking for a byte that, in UTF-8, would be a
> totally invalid byte (not an invalid sequence of bytes).  At the time,
> I tried some googling, but it looked rather hopeless (maybe it was my
> googling that was hopeless).
> 
> I know that your statement does not imply there is such a byte, but
> maybe you (or someone else reading this) know(s)?
> 
> (The reason I wanted such a byte was to use it as a record separator in
> a set of text files (that I use as an askSam "workalike" (or
> "worksimilar") so that I could use msort (which depends on a 1 byte
> record separator to --separate the records ;-) while sorting.)  (Some
> of the files already include UTF-8, and, in the future, I anticpate all
> will be in UTFF-8.)

maybe you could use the null byte?

Regards

Michael


.-.. .. ...- .   .-.. --- -. --.   .- -. -..   .--. .-. --- ... .--. . .-.

War is never imperative.
		-- McCoy, "Balance of Terror", stardate 1709.2

[toc] | [prev] | [next] | [standalone]


#194441 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-03 13:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAsBP-7m1-11@gated-at.bofh.it>
In reply to#194424
On Monday, April 02, 2018 06:43:28 PM Michael Lange wrote:
> On Mon, 2 Apr 2018 08:37:54 -0400
> 
> rhkramer@gmail.com wrote:
> > A few weeks ago, I was looking for a byte that, in UTF-8, would be a
> > totally invalid byte (not an invalid sequence of bytes).  At the time,
> > I tried some googling, but it looked rather hopeless (maybe it was my
> > googling that was hopeless).
> > 
> > I know that your statement does not imply there is such a byte, but
> > maybe you (or someone else reading this) know(s)?
> > 
> > (The reason I wanted such a byte was to use it as a record separator in
> > a set of text files (that I use as an askSam "workalike" (or
> > "worksimilar") so that I could use msort (which depends on a 1 byte
> > record separator to --separate the records ;-) while sorting.)  (Some
> > of the files already include UTF-8, and, in the future, I anticpate all
> > will be in UTFF-8.)
> 
> maybe you could use the null byte?

Thanks!

Surprisingly (to me), this (and maybe several other of the control characters 
might work--I did a search of one of the files, and there are no null bytes.

Next I'll have to refresh my memory on how to replace the existing From with 
From preceded by the null character, i.e., something like:

Find: \n\nFrom 
Replace with \n\n0x00\nFrom

I'll probably look into doing that with something like Awk or Perl.  I'll have 
to review how to represent hex 00 in the Awk or Perl statement.

(I didn't check to see if any 0xff bytes are present in the file, I suspect 
there aren't, and I could use that as well.

[toc] | [prev] | [next] | [standalone]


Page 1 of 6  [1] 2 3 4 5 6  Next page →

Back to top | Article view | linux.debian.user


csiph-web