Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #194373 > unrolled thread
| Started by | mess-mate <mess-mate@gmx.com> |
|---|---|
| First post | 2018-04-01 16:10 +0200 |
| Last post | 2018-04-04 14:20 +0200 |
| Articles | 20 on this page of 101 — 21 participants |
Back to article view | Back to linux.debian.user
utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200
Page 1 of 6 [1] 2 3 4 5 6 Next page →
| From | mess-mate <mess-mate@gmx.com> |
|---|---|
| Date | 2018-04-01 16:10 +0200 |
| Subject | utf |
| Message-ID | <vzLQd-4cg-1@gated-at.bofh.it> |
Hi, howto change the system utf to eu character set ? regards
[toc] | [next] | [standalone]
| From | Curt <curty@free.fr> |
|---|---|
| Date | 2018-04-01 17:50 +0200 |
| Message-ID | <vzNoZ-5bx-1@gated-at.bofh.it> |
| In reply to | #194373 |
On 2018-04-01, mess-mate <mess-mate@gmx.com> wrote: > Hi, > > howto change the system utf to eu character set ? https://wiki.debian.org/Locale (if it's not wrong or hopelessly out of date). The last I checked--and as far as I know--(with administrative privileges), dpkg-reconfigure locales was the way to go. Now you tell us how you solved the 3 screens problem (you know, for posterity). > regards > > -- The time-state of attainment eliminates so accurately the time-state of aspiration, that the actual seems the inevitable, and, all conscious intellectual effort to reconstitute the invisible and unthinkable as a reality being fruitless, we are incapable of appreciating our joy by comparing it with our sorrow. --Samuel Beckett
[toc] | [prev] | [next] | [standalone]
| From | Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> |
|---|---|
| Date | 2018-04-01 18:30 +0200 |
| Message-ID | <vzO1H-5Qv-1@gated-at.bofh.it> |
| In reply to | #194373 |
On 1-04-2018, at 16h 05'41", mess-mate wrote about "utf" > Hi, > > howto change the system utf to eu character set ? > > regards What is "eu character set"? None of the Latin-x or iso8859-y can cover all the gliphs of all the countries in Europe. So, without knowing what do you think is "eu character set", I would say that you must stay with UTF-8. I really mean this. Ionel
[toc] | [prev] | [next] | [standalone]
| From | Andre Majorel <aym-naibed@teaser.fr> |
|---|---|
| Date | 2018-04-01 18:40 +0200 |
| Message-ID | <vzObn-5UZ-3@gated-at.bofh.it> |
| In reply to | #194373 |
On 2018-04-01 16:05 +0200, mess-mate wrote:
> howto change the system utf to eu character set ?
If, by "eu", you mean ISO 8859-1 or -15, here is the procedure
that works for me :
1) Run
dpkg-reconfigure locales
and select some appropriate non-UTF-8 locales. When asked
what the default locale should be, select a non-UTF-8 one.
2) In your ~/.profile, ~/.bashrc or whatever is relevant, make
sure that none of LANG, LC_ALL, LC_CTYPE, LC_MESSAGES and
friends name a UTF-8 locale. You may not have to be so
drastic but then you will have to understand the somewhat
counter-intuitive rules of precedence among those variables.
3) Run
dpkg-reconfigure console-setup
and make some sensible choices. Annoyingly, the scripts or
programs that configure the console (/dev/tty1 .. /dev/tty6)
in Debian set the keyboard to UTF-8 despite your preferences.
The incantation to (temporarily) fix that is
setupcon -k -v
stty -iutf8
Note that I don't use any fancy desktop environments. I wouldn't
be surprised to learn that those who do must follow different
procedures.
--
André Majorel <http://www.teaser.fr/~amajorel/>
lists.debian.org, an essential online resource for spammers.
[toc] | [prev] | [next] | [standalone]
| From | Ben Caradoc-Davies <ben@transient.nz> |
|---|---|
| Date | 2018-04-01 22:10 +0200 |
| Message-ID | <vzRsC-8m1-3@gated-at.bofh.it> |
| In reply to | #194373 |
On 02/04/18 02:05, mess-mate wrote: > howto change the system utf to eu character set ? Why? UTF (especially UTF-8) is vastly superior for all purposes: http://utf8everywhere.org/ What are you trying to do, and why do you thing a non-UTF encoding might help you to do it? You will likely be able to achieve your goals with a UTF locale. Kind regards, -- Ben Caradoc-Davies <ben@transient.nz> Director Transient Software Limited <https://transient.nz/> New Zealand
[toc] | [prev] | [next] | [standalone]
| From | Cindy-Sue Causey <butterflybytes@gmail.com> |
|---|---|
| Date | 2018-04-02 00:50 +0200 |
| Message-ID | <vzTXs-1iJ-1@gated-at.bofh.it> |
| In reply to | #194384 |
On 4/1/18, Ben Caradoc-Davies <ben@transient.nz> wrote: > On 02/04/18 02:05, mess-mate wrote: >> howto change the system utf to eu character set ? > > Why? UTF (especially UTF-8) is vastly superior for all purposes: > http://utf8everywhere.org/ > > What are you trying to do, and why do you thing a non-UTF encoding might > help you to do it? You will likely be able to achieve your goals with a > UTF locale. Sometimes it takes seeing certain words together to trigger a thought. Ben's words made me wonder if maybe OP is missing a font package. That thought occurred because I use UTF-8 and find fonts missing sometimes (too?).... Cindy :) -- Cindy-Sue Causey Talking Rock, Pickens County, Georgia, USA * hops with duct tape *
[toc] | [prev] | [next] | [standalone]
| From | Curt <curty@free.fr> |
|---|---|
| Date | 2018-04-02 09:50 +0200 |
| Message-ID | <vA2o1-6N4-1@gated-at.bofh.it> |
| In reply to | #194386 |
On 2018-04-01, Cindy-Sue Causey <butterflybytes@gmail.com> wrote: > On 4/1/18, Ben Caradoc-Davies <ben@transient.nz> wrote: >> On 02/04/18 02:05, mess-mate wrote: >>> howto change the system utf to eu character set ? >> >> Why? UTF (especially UTF-8) is vastly superior for all purposes: >> http://utf8everywhere.org/ >> >> What are you trying to do, and why do you thing a non-UTF encoding might >> help you to do it? You will likely be able to achieve your goals with a >> UTF locale. > > > Sometimes it takes seeing certain words together to trigger a thought. > Ben's words made me wonder if maybe OP is missing a font package. That > thought occurred because I use UTF-8 and find fonts missing sometimes > (too?).... > > Cindy :) The thought provoked in my neurological matter was why there are other locales at all if UTF8 (the locale of this here .homie machine, BTW) is "vastly superior for all purposes". That leaves no purposes remaining whatsoever for the myriad other locales. If this is indeed so, let's get rid of them, then, (the superfluous locales) thus lightening our loads in computer life. Or are there legacy, corner, arcane purposes for these locales, of which the hoi polloi (of which I am a card-carrying member) has not to concern itself? -- The time-state of attainment eliminates so accurately the time-state of aspiration, that the actual seems the inevitable, and, all conscious intellectual effort to reconstitute the invisible and unthinkable as a reality being fruitless, we are incapable of appreciating our joy by comparing it with our sorrow. --Samuel Beckett
[toc] | [prev] | [next] | [standalone]
| From | Richard Hector <richard@walnut.gen.nz> |
|---|---|
| Date | 2018-04-02 11:40 +0200 |
| Message-ID | <vA46u-7WZ-7@gated-at.bofh.it> |
| In reply to | #194387 |
[Multipart message — attachments visible in raw view] — view raw
On 02/04/18 19:43, Curt wrote: > The thought provoked in my neurological matter was why there are other > locales at all if UTF8 (the locale of this here .homie machine, BTW) is > "vastly superior for all purposes". There's more to the locale than the character set - things like default language, thousands separators, currency symbols, collation order ... Richard
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-02 15:30 +0200 |
| Message-ID | <vA7H3-1Tr-3@gated-at.bofh.it> |
| In reply to | #194387 |
On Mon, Apr 02, 2018 at 07:43:23AM +0000, Curt wrote: > The thought provoked in my neurological matter was why there are other > locales at all if UTF8 (the locale of this here .homie machine, BTW) is > "vastly superior for all purposes". > > That leaves no purposes remaining whatsoever for the myriad other > locales. > > If this is indeed so, let's get rid of them, then, (the superfluous > locales) thus lightening our loads in computer life. > > Or are there legacy, corner, arcane purposes for these locales, of which > the hoi polloi (of which I am a card-carrying member) has not to concern > itself? Yes, there are other computer systems on the planet, and some of them still use various single-byte character sets (ISO 8859-* or WIN1252 or similar). In heterogeneous environments, it is often important for a Debian server to support multiple locales, to fit the needs of legacy client systems. For example, if you're sitting at a legacy Unix system which uses the ISO 8859-1 character set, and you ssh to a Debian system that hard-codes "LANG=en_US.UTF-8" in /etc/default/locale , this will clobber the locale variables that are sent by the ssh client. Then, commands like "man" will generate output containing UTF-8 characters, which will seriously mess up the display on an ISO 8859-1 terminal emulator. (Why "man", or groff or whatever, feels the need to generate non-ASCII characters is beyond me.) This is why <https://wiki.debian.org/Locale> has complicated instructions that DO NOT force an override LANG=... in /etc/default/locale (which is destructive and wrong), but instead use fallback values in /etc/profile and /etc/csh/login.d/lang . This is also why I was probing login.conf and pam_env.conf for the ability to support fallback/default environment variable values in that other thread last week. (Sadly, there is still no known alternative to the login-shell-dependent hackery.)
[toc] | [prev] | [next] | [standalone]
| From | Andre Majorel <aym-naibed@teaser.fr> |
|---|---|
| Date | 2018-04-02 09:50 +0200 |
| Message-ID | <vA2o1-6N4-3@gated-at.bofh.it> |
| In reply to | #194384 |
On 2018-04-02 08:00 +1200, Ben Caradoc-Davies wrote: > On 02/04/18 02:05, mess-mate wrote: > >howto change the system utf to eu character set ? > > Why? UTF (especially UTF-8) is vastly superior for all purposes: I wouldn't say that. UTF-8 breaks a number of assumptions. For instance, 1) every character has the same size, 2) every byte sequence is a valid character, 3) the equality or inequality of two characters comes down to the equality or inequality of the bytes they encode to. With ASCII and the many encodings based on it, most things can be done without having knowledge of the encoding. With UTF-8, even basic operations like determining the length of a string or reporting at what column an error occurred require knowledge of the encoding. -- André Majorel <http://www.teaser.fr/~amajorel/> Imagine what would happen if the Debian project disclosed the email addresses of their users. Spambots would harvest them and Debian users would be inundated with spam. Good thing they don't, eh ?
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-02 14:40 +0200 |
| Subject | Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vA6UG-1kM-1@gated-at.bofh.it> |
| In reply to | #194388 |
On Monday, April 02, 2018 03:39:05 AM Andre Majorel wrote: > > Why? UTF (especially UTF-8) is vastly superior for all purposes: > I wouldn't say that. UTF-8 breaks a number of assumptions. For > instance, > 1) every character has the same size, > 2) every byte sequence is a valid character, A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally invalid byte (not an invalid sequence of bytes). At the time, I tried some googling, but it looked rather hopeless (maybe it was my googling that was hopeless). I know that your statement does not imply there is such a byte, but maybe you (or someone else reading this) know(s)? (The reason I wanted such a byte was to use it as a record separator in a set of text files (that I use as an askSam "workalike" (or "worksimilar") so that I could use msort (which depends on a 1 byte record separator to --separate the records ;-) while sorting.) (Some of the files already include UTF-8, and, in the future, I anticpate all will be in UTFF-8.) > 3) the equality or inequality of two characters comes down to > the equality or inequality of the bytes they encode to.
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-02 15:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vA7nH-1KQ-7@gated-at.bofh.it> |
| In reply to | #194391 |
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1
On Mon, Apr 02, 2018 at 08:37:54AM -0400, rhkramer@gmail.com wrote:
> On Monday, April 02, 2018 03:39:05 AM Andre Majorel wrote:
> > > Why? UTF (especially UTF-8) is vastly superior for all purposes:
> > I wouldn't say that. UTF-8 breaks a number of assumptions. For
> > instance,
> > 1) every character has the same size,
> > 2) every byte sequence is a valid character,
>
> A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally
> invalid byte (not an invalid sequence of bytes).
If you look at man utf-8 (7), you'll find, for the encoding:
Encoding
The following byte sequences are used to represent a character.
The sequence to be used depends on the UCS code number of the character:
0x00000000 - 0x0000007F:
0xxxxxxx
0x00000080 - 0x000007FF:
110xxxxx 10xxxxxx
0x00000800 - 0x0000FFFF:
1110xxxx 10xxxxxx 10xxxxxx
0x00010000 - 0x001FFFFF:
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
0x00200000 - 0x03FFFFFF:
111110xx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
0x04000000 - 0x7FFFFFFF:
1111110x 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx 10xxxxxx
This means that 1111111x are the only (two) illegal bytes in UTF-8,
at least currently. I don't know what will happen once we need code
points beyond 0x7fffffff -- perhaps Klingon has an ideographic
variant our linguists don't know of yet (the alphabetical variant
is in the private area and an official place seems to be in the
works [1]).
So I wouldn't build my house on it. And oh, there are better
serialization "protocols" than using an arbitrary record
separator. Why not use some of the lower ASCII thingies plus
an escape mechanism?
Cheers
[1] http://www.klingonwiki.net/En/Unicode
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)
iEYEARECAAYFAlrCKdUACgkQBcgs9XrR2kbBnQCfe5c4WVNYCcpZbsgg5dDZwBHR
XNkAn2CSUkMY59VE2zLciII/3kUz8W45
=aPFG
-----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | Henrique de Moraes Holschuh <hmh@debian.org> |
|---|---|
| Date | 2018-04-02 15:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vA7nH-1KQ-9@gated-at.bofh.it> |
| In reply to | #194391 |
On Mon, 02 Apr 2018, rhkramer@gmail.com wrote: > A few weeks ago, I was looking for a byte that, in UTF-8, would be a totally > invalid byte (not an invalid sequence of bytes). At the time, I tried some > googling, but it looked rather hopeless (maybe it was my googling that was > hopeless). 0xff should work. But any of those in RED on the wikipedia article about UTF-8 would do for Unicode text: https://en.wikipedia.org/wiki/UTF-8 -- Henrique Holschuh
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-02 19:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAbKF-4vW-3@gated-at.bofh.it> |
| In reply to | #194394 |
Thanks to tomas and Henrique! The wikipedia article is rather interesting, in a quick skim, I learned some interesting things about UTF-8, especially the property of self- synchronization. I had trouble reading that large table--but if I simply take the red boxes at face value, maybe there are 10 or so bytes that are not valid UTF-8. I'll probably first consider the bytes that tomas also mentions, i.e., decimal 254 and 255). I guess I have a followup question--are those two bytes (or either one of them) also unused in all possible "code pages"? The problem is that I copy snippets of text from all kinds of sources into those text files (which are formatted like mbox files), so I might find one or both of those bytes in the file already. I guess it's not a big deal as, I will either: * search the file (with a hex editor, I guess) to see if decimal 254 or 255 is in use already--if only one or a few cases, I might replace it with something else before adding additional instances to serve as a temporary record separator (for use by msort), or * use one of the other utilities that I've since found which can apparently sort mbox files while keeping emails intact (I have to read up on those (or it?) again as there were, iirc, also some limitations there that might not let me accomplish what I want (usually, sorting the emails by the title in the mbox "From " header (and ususally not by the email From: header). Thanks again! On Monday, April 02, 2018 09:05:52 AM Henrique de Moraes Holschuh wrote: > On Mon, 02 Apr 2018, rhkramer@gmail.com wrote: > > A few weeks ago, I was looking for a byte that, in UTF-8, would be a > > totally invalid byte (not an invalid sequence of bytes). At the time, I > > tried some googling, but it looked rather hopeless (maybe it was my > > googling that was hopeless). > > 0xff should work. But any of those in RED on the wikipedia article > about UTF-8 would do for Unicode text: > > https://en.wikipedia.org/wiki/UTF-8
[toc] | [prev] | [next] | [standalone]
| From | Henrique de Moraes Holschuh <hmh@debian.org> |
|---|---|
| Date | 2018-04-02 20:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAcdI-4Wt-9@gated-at.bofh.it> |
| In reply to | #194411 |
On Mon, 02 Apr 2018, rhkramer@gmail.com wrote: > The wikipedia article is rather interesting, in a quick skim, I learned some > interesting things about UTF-8, especially the property of self- > synchronization. Yes, UTF-8 is a brilliant design. > I had trouble reading that large table--but if I simply take the red boxes at > face value, maybe there are 10 or so bytes that are not valid UTF-8. I'll > probably first consider the bytes that tomas also mentions, i.e., decimal 254 > and 255). On that table, columns are the least significant bits (second hex digit), and rows are the most significant bits (first hex digit) of a byte. As in C1 is row C, column 1. The "2-byte", "3-byte" and "4-byte" are comments that remind you of the self-sinchronizing nature of UTF-8, and that these bytes would be invalid outside of that position in an UTF-8 sequence that encodes a single code point (but they would be valid in the correct position). The stuff in "red" on that table is always invalid for Unicode: if you find one of those in a data file, that file is *not* valid UTF-8 (but it could be valid UTF-16, valid UTF-32, or valid ISO-8859-*, etc). > I guess I have a followup question--are those two bytes (or either one of > them) also unused in all possible "code pages"? For Unicode, yes, because Unicode can't go past code point 0x10ffff. And that isn't about to change anytime soon (lots of stuff hardcode it somehow, e.g., by limiting the number of UTF-8 bytes that can be used to encode a single code point...). I have not read the Unicode standard to check what it says about future expansions related to the valid code point range, though. > The problem is that I copy snippets of text from all kinds of sources into > those text files (which are formatted like mbox files), so I might find one or > both of those bytes in the file already. Then it isn't a valid unicode text file in UTF-8 format, and it needs to be converted (or fixed) first to be encoded in UTF-8 :-) -- Henrique Holschuh
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-02 20:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAcGJ-58A-11@gated-at.bofh.it> |
| In reply to | #194413 |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote: > On Mon, 02 Apr 2018, rhkramer@gmail.com wrote: > > The wikipedia article is rather interesting, in a quick skim, I learned some > > interesting things about UTF-8, especially the property of self- > > synchronization. > > Yes, UTF-8 is a brilliant design. Possibly relevant, definitely entertaining, Rob Pike's account of UTF-8's gestation [1] Yeah. Elegant design. Until the Unicode Consortium left Microsoft near it (Byte Order Mark, I'm looking at you!). [...] > > I guess I have a followup question--are those two bytes (or either one of > > them) also unused in all possible "code pages"? I'm not sure what you mean here: there are two layers at work (at least if you have UTF-8 encoded Unicode). As Henrique says, if you assume both to be "correct" then you get more illegal things. But sometimes UTF-8 encoding is used for other things (notably Emacs encodes a superset of Unicode, to be able to express "raw byte values" next to "Unicode characters". > > The problem is that I copy snippets of text from all kinds of sources into > > those text files (which are formatted like mbox files), so I might find one or > > both of those bytes in the file already. > > Then it isn't a valid unicode text file in UTF-8 format, and it needs to > be converted (or fixed) first to be encoded in UTF-8 :-) Agreed: if you don't know what's coming in, you better plan for anything :) Cheers - -- t -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) iEYEARECAAYFAlrCeTcACgkQBcgs9XrR2kbRtgCfaRHoodlkFFt8Gm0Oq438ymvg 0oMAn2NkpsqMJ3Tcy5BvAJIpTvfG8mdj =iVqF -----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | tomas@tuxteam.de |
|---|---|
| Date | 2018-04-02 21:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAcQp-5c2-3@gated-at.bofh.it> |
| In reply to | #194415 |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On Mon, Apr 02, 2018 at 08:40:55PM +0200, tomas@tuxteam.de wrote: > On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote: > > On Mon, 02 Apr 2018, rhkramer@gmail.com wrote: > > > The wikipedia article is rather interesting, in a quick skim, I learned some > > > interesting things about UTF-8, especially the property of self- > > > synchronization. > > > > Yes, UTF-8 is a brilliant design. > > Possibly relevant, definitely entertaining, Rob Pike's account > of UTF-8's gestation [1] Oops: here's the ref: [1] http://doc.cat-v.org/bell_labs/utf-8_history - -- t -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) iEYEARECAAYFAlrCe8EACgkQBcgs9XrR2kY5mQCfQpQE1lniIWWWMXNm79oETOhP IoQAn3Sm9i6NFwPyPe1XlFrqNTMFyB8M =oE1V -----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-02 21:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAd06-5v7-19@gated-at.bofh.it> |
| In reply to | #194415 |
Thanks, again, to Henrique and tomas for the followups! On Monday, April 02, 2018 02:40:55 PM tomas@tuxteam.de wrote: > On Mon, Apr 02, 2018 at 03:18:38PM -0300, Henrique de Moraes Holschuh wrote:
[toc] | [prev] | [next] | [standalone]
| From | Michael Lange <klappnase@freenet.de> |
|---|---|
| Date | 2018-04-03 00:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAgqZ-7Cc-3@gated-at.bofh.it> |
| In reply to | #194391 |
Hi, On Mon, 2 Apr 2018 08:37:54 -0400 rhkramer@gmail.com wrote: > A few weeks ago, I was looking for a byte that, in UTF-8, would be a > totally invalid byte (not an invalid sequence of bytes). At the time, > I tried some googling, but it looked rather hopeless (maybe it was my > googling that was hopeless). > > I know that your statement does not imply there is such a byte, but > maybe you (or someone else reading this) know(s)? > > (The reason I wanted such a byte was to use it as a record separator in > a set of text files (that I use as an askSam "workalike" (or > "worksimilar") so that I could use msort (which depends on a 1 byte > record separator to --separate the records ;-) while sorting.) (Some > of the files already include UTF-8, and, in the future, I anticpate all > will be in UTFF-8.) maybe you could use the null byte? Regards Michael .-.. .. ...- . .-.. --- -. --. .- -. -.. .--. .-. --- ... .--. . .-. War is never imperative. -- McCoy, "Balance of Terror", stardate 1709.2
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-03 13:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAsBP-7m1-11@gated-at.bofh.it> |
| In reply to | #194424 |
On Monday, April 02, 2018 06:43:28 PM Michael Lange wrote: > On Mon, 2 Apr 2018 08:37:54 -0400 > > rhkramer@gmail.com wrote: > > A few weeks ago, I was looking for a byte that, in UTF-8, would be a > > totally invalid byte (not an invalid sequence of bytes). At the time, > > I tried some googling, but it looked rather hopeless (maybe it was my > > googling that was hopeless). > > > > I know that your statement does not imply there is such a byte, but > > maybe you (or someone else reading this) know(s)? > > > > (The reason I wanted such a byte was to use it as a record separator in > > a set of text files (that I use as an askSam "workalike" (or > > "worksimilar") so that I could use msort (which depends on a 1 byte > > record separator to --separate the records ;-) while sorting.) (Some > > of the files already include UTF-8, and, in the future, I anticpate all > > will be in UTFF-8.) > > maybe you could use the null byte? Thanks! Surprisingly (to me), this (and maybe several other of the control characters might work--I did a search of one of the files, and there are no null bytes. Next I'll have to refresh my memory on how to replace the existing From with From preceded by the null character, i.e., something like: Find: \n\nFrom Replace with \n\n0x00\nFrom I'll probably look into doing that with something like Awk or Perl. I'll have to review how to represent hex 00 in the Awk or Perl statement. (I didn't check to see if any 0xff bytes are present in the file, I suspect there aren't, and I could use that as well.
[toc] | [prev] | [next] | [standalone]
Page 1 of 6 [1] 2 3 4 5 6 Next page →
Back to top | Article view | linux.debian.user
csiph-web