Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #194373 > unrolled thread

utf

Started bymess-mate <mess-mate@gmx.com>
First post2018-04-01 16:10 +0200
Last post2018-04-04 14:20 +0200
Articles 20 on this page of 101 — 21 participants

Back to article view | Back to linux.debian.user


Contents

  utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
    Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
    Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
    Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
    Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
      Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
        Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
      Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
        Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
                              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
                                Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
                                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
                          mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re:  utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
                            Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was:  Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
                    Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
                      Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
        Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
        Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
            Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
          Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
          Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
            Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
              Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
                Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
                    Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
                      Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
                            Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
                              Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
                              Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
                                Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
                            Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
                    Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
                      Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
                        Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
                          Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
                            Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
                              Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
                                Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
                                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
                                  Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
                                    Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
                                      Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
                                      Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
                              Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
                            Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
                Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
          Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200

Page 2 of 6 — ← Prev page 1 [2] 3 4 5 6  Next page →


#194443 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromNicolas George <george@nsup.org>
Date2018-04-03 14:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAsLw-7pK-9@gated-at.bofh.it>
In reply to#194441

[Multipart message — attachments visible in raw view] — view raw

rhkramer@gmail.com (2018-04-03):
> Next I'll have to refresh my memory on how to replace the existing From with 
> From preceded by the null character, i.e., something like:
> 
> Find: \n\nFrom 
> Replace with \n\n0x00\nFrom

This is a very bad idea, and you are obviously about tu reproduce the
errors of the past.

You need to change your design. You obviously have free-form text, with
possibly some rigid syntax but not enough for your needs. Therefore, you
cannot use a delimiter inside the text. You have to put your structure
outside the text.

It is very typical of the "problems" people have with UTF-8: the problem
resides not in the properties of UTF-8 but in the unwritten assumptions
about the way they should be implementing things.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194448 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-03 14:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAtex-7RB-5@gated-at.bofh.it>
In reply to#194443
On Tuesday, April 03, 2018 07:54:35 AM Nicolas George wrote:
> rhkramer@gmail.com (2018-04-03):
> > Next I'll have to refresh my memory on how to replace the existing From
> > with From preceded by the null character, i.e., something like:
> > 
> > Find: \n\nFrom
> > Replace with \n\n0x00\nFrom
> 
> This is a very bad idea, and you are obviously about tu reproduce the
> errors of the past.
> 
> You need to change your design. 

I don't understand the errors of the past nor is it feasible to change my 
design.   Two points:

   * the other half of the above replacement is deleting the 0x00 after 
sorting (unless it is totally innocuous, which it might be)

   * my design is what I'd call a mashup of existing programs that use and 
require a specified file format (basically, mbox)--the programs that I use 
include kmail, nail, recol, kate, and, in the future, any editor which uses 
Scintilla, and, I would hope to be able to use any email program that can use 
mbox files.






> You obviously have free-form text, with
> possibly some rigid syntax but not enough for your needs. Therefore, you
> cannot use a delimiter inside the text. You have to put your structure
> outside the text.
> 
> It is very typical of the "problems" people have with UTF-8: the problem
> resides not in the properties of UTF-8 but in the unwritten assumptions
> about the way they should be implementing things.
> 
> Regards,

[toc] | [prev] | [next] | [standalone]


#194445 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromMichael Lange <klappnase@freenet.de>
Date2018-04-03 14:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAsVb-7IQ-7@gated-at.bofh.it>
In reply to#194441
Hi,

On Tue, 3 Apr 2018 07:43:02 -0400
rhkramer@gmail.com wrote:

> > maybe you could use the null byte?
> 
> Thanks!
> 
> Surprisingly (to me), this (and maybe several other of the control
> characters might work--I did a search of one of the files, and there
> are no null bytes.

I believe (please anyone correct me if I am wrong) that "text" files
won't contain any null byte; many text editors even refuse to open such a
file, I guess since they assume it is a "binary" file.
Probably it is the same with some other control characters like 04 (End
of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII
it seems like 1C (File Separator) or 1E (Record Separator) might be 
appropriate choices for you. I'm no expert on this, though.

Best regards

Michael


.-.. .. ...- .   .-.. --- -. --.   .- -. -..   .--. .-. --- ... .--. . .-.

Oh, that sound of male ego.  You travel halfway across the galaxy and
it's still the same song.
		-- Eve McHuron, "Mudd's Women", stardate 1330.1

[toc] | [prev] | [next] | [standalone]


#194446 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromMichael Lange <klappnase@freenet.de>
Date2018-04-03 14:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAt4S-7NM-7@gated-at.bofh.it>
In reply to#194445
On Tue, 3 Apr 2018 13:58:33 +0200
Michael Lange <klappnase@freenet.de> wrote:

> I believe (please anyone correct me if I am wrong) that "text" files
> won't contain any null byte; many text editors even refuse to open such
> a file, I guess since they assume it is a "binary" file.
> Probably it is the same with some other control characters like 04 (End
> of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII
> it seems like 1C (File Separator) or 1E (Record Separator) might be 
> appropriate choices for you. I'm no expert on this, though.

Addendum: iirc (again please correct me if I am wrong) unix file names
may contain (at least in theory) any byte except 2F (the slash) and the
null byte. So if your text files might contain arbitrary file names there
may be (at least in theory) a (admittedly very small) chance that such a
file name actually might contain any control character except the null
byte.

Regards

Michael

.-.. .. ...- .   .-.. --- -. --.   .- -. -..   .--. .-. --- ... .--. . .-.

Live long and prosper.
		-- Spock, "Amok Time", stardate 3372.7

[toc] | [prev] | [next] | [standalone]


#194449 — Re: Invalid UTF-8 byte? (was: Re: utf)

From<tomas@tuxteam.de>
Date2018-04-03 14:40 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAtoe-7V4-21@gated-at.bofh.it>
In reply to#194446
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Tue, Apr 03, 2018 at 02:14:07PM +0200, Michael Lange wrote:
> On Tue, 3 Apr 2018 13:58:33 +0200
> Michael Lange <klappnase@freenet.de> wrote:
> 
> > I believe (please anyone correct me if I am wrong) that "text" files
> > won't contain any null byte; many text editors even refuse to open such
> > a file, I guess since they assume it is a "binary" file.

Emacs and vi(m) do open, edit and save such a file (and make no fuss
about it). By default, both depict those NULL characters as ^@.

Are there other editors? (yah, that was a bit snarky, I know ;-)

> > Probably it is the same with some other control characters like 04 (End
> > of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII
> > it seems like 1C (File Separator) or 1E (Record Separator) might be 
> > appropriate choices for you. I'm no expert on this, though.

Just assuming "this won't happen" is a sure recipe for some debugging
fun. But perhaps the OP is looking for such fun. (S)he seems set on
trying...

> Addendum: iirc (again please correct me if I am wrong) unix file names
> may contain (at least in theory) any byte except 2F (the slash) and the
> null byte. So if your text files might contain arbitrary file names there
> may be (at least in theory) a (admittedly very small) chance that such a
> file name actually might contain any control character except the null
> byte.

You are correct on file names.

Cheers
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrDdEgACgkQBcgs9XrR2kYbPgCeIA04MoYQleL5IDw5wwmerx0o
bqEAnA24L1+etC0tlCH2ExSdNigEPMDU
=iV6I
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194457 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromMichael Lange <klappnase@freenet.de>
Date2018-04-03 21:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAA6l-3PH-1@gated-at.bofh.it>
In reply to#194449
Hi,

On Tue, 3 Apr 2018 14:32:08 +0200
<tomas@tuxteam.de> wrote:

> > > Probably it is the same with some other control characters like 04
> > > (End of Transmission). When I look at
> > > https://en.wikipedia.org/wiki/ASCII it seems like 1C (File
> > > Separator) or 1E (Record Separator) might be appropriate choices
> > > for you. I'm no expert on this, though.
> 
> Just assuming "this won't happen" is a sure recipe for some debugging
> fun. But perhaps the OP is looking for such fun. (S)he seems set on
> trying...
> 
> > Addendum: iirc (again please correct me if I am wrong) unix file names
> > may contain (at least in theory) any byte except 2F (the slash) and
> > the null byte. So if your text files might contain arbitrary file
> > names there may be (at least in theory) a (admittedly very small)
> > chance that such a file name actually might contain any control
> > character except the null byte.
> 
> You are correct on file names.

Thanks for the clarification.

>From what i have understood I think the OP should certainly at least,
whatever the files they want to include exactly look like and whichever
byte they choose as delimiter, scan the file first for such a byte and if
it is actually found replace it with either an empty string or
(probably better) some sort of "tag" before applying the contents to the
new database. This way they could at least be sure that their chosen
delimiter does not split one record into halves.

I don't think it makes really any difference which of the bytes that
aren't supposed to be present in "text" files is used, one can never be
100% sure that these bytes don't show up nonetheless.

I have no idea what these "text files" look like of course. It just seemed
-to me - that the fact that the null byte cannot ever be part of a file
name might make it slightly more appropriate for this purpose than other
candidate bytes. Of course, it depends...

Best regards

Michael

.-.. .. ...- .   .-.. --- -. --.   .- -. -..   .--. .-. --- ... .--. . .-.

Suffocating together ... would create heroic camaraderie.
		-- Khan Noonian Singh, "Space Seed", stardate 3142.8

[toc] | [prev] | [next] | [standalone]


#194458 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-03 21:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAA6l-3PH-3@gated-at.bofh.it>
In reply to#194457
On Tue, Apr 03, 2018 at 09:36:42PM +0200, Michael Lange wrote:
> >From what i have understood I think the OP should certainly at least,
> whatever the files they want to include exactly look like and whichever
> byte they choose as delimiter, scan the file first for such a byte and if
> it is actually found replace it with either an empty string or
> (probably better) some sort of "tag" before applying the contents to the
> new database. This way they could at least be sure that their chosen
> delimiter does not split one record into halves.

Or abort the program with an error message.

> I have no idea what these "text files" look like of course. It just seemed
> -to me - that the fact that the null byte cannot ever be part of a file
> name might make it slightly more appropriate for this purpose than other
> candidate bytes. Of course, it depends...

NUL bytes are an excellent choice for delimiter in lots of situations.

Of course, we still need the OP to tell us what the actual situation is.

[toc] | [prev] | [next] | [standalone]


#194460 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromMichael Lange <klappnase@freenet.de>
Date2018-04-03 22:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAApH-4eF-3@gated-at.bofh.it>
In reply to#194458
On Tue, 3 Apr 2018 15:47:57 -0400
Greg Wooledge <wooledg@eeg.ccf.org> wrote:

> On Tue, Apr 03, 2018 at 09:36:42PM +0200, Michael Lange wrote:
> > >From what i have understood I think the OP should certainly at least,
> > whatever the files they want to include exactly look like and
> > whichever byte they choose as delimiter, scan the file first for such
> > a byte and if it is actually found replace it with either an empty
> > string or (probably better) some sort of "tag" before applying the
> > contents to the new database. This way they could at least be sure
> > that their chosen delimiter does not split one record into halves.
> 
> Or abort the program with an error message.

Yes, or ask the user what to do with that particular file :)

Regards

Michael

> 
> > I have no idea what these "text files" look like of course. It just
> > seemed -to me - that the fact that the null byte cannot ever be part
> > of a file name might make it slightly more appropriate for this
> > purpose than other candidate bytes. Of course, it depends...
> 
> NUL bytes are an excellent choice for delimiter in lots of situations.
> 
> Of course, we still need the OP to tell us what the actual situation is.
> 



.-.. .. ...- .   .-.. --- -. --.   .- -. -..   .--. .-. --- ... .--. . .-.

Peace was the way.
		-- Kirk, "The City on the Edge of Forever", stardate
unknown

[toc] | [prev] | [next] | [standalone]


#194450 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-03 14:40 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAtod-7V4-15@gated-at.bofh.it>
In reply to#194446
> Addendum: iirc (again please correct me if I am wrong) unix file names
> may contain (at least in theory) any byte except 2F (the slash) and the
> null byte. So if your text files might contain arbitrary file names there
> may be (at least in theory) a (admittedly very small) chance that such a
> file name actually might contain any control character except the null
> byte.

One might question whether a file that contains a list of filenames is
really a "text file".  It sounds more like a broken data file.

The real question here (for the OP) is:

WHAT ARE YOU TRYING TO DO?

There was a glimpse a few messages back that looked like you were trying
to parse information out of an mbox-format mail folder.  (I.e. a flat
file that has a concatenated series of mbox-format mail messages in it,
with all the silliness and problems inherent in this format, like having
to prefix body lines with ">" if they begin with the word "From".)

"I want to write a shell script to parse an mbox folder..." is enough
to send most people running away screaming.  What other horrors are we
in store for next?

Of course, that might be a red herring, since you didn't actually tell
us what your goal is, or what your inputs are, and we're having to
guess at the moment based on tiny hints and information leaks.

[toc] | [prev] | [next] | [standalone]


#194468 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 03:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAF61-7kF-1@gated-at.bofh.it>
In reply to#194450
On Tuesday, April 03, 2018 08:30:04 AM Greg Wooledge wrote:
> > Addendum: iirc (again please correct me if I am wrong) unix file names
> > may contain (at least in theory) any byte except 2F (the slash) and the
> > null byte. So if your text files might contain arbitrary file names there
> > may be (at least in theory) a (admittedly very small) chance that such a
> > file name actually might contain any control character except the null
> > byte.
> 
> One might question whether a file that contains a list of filenames is
> really a "text file".  It sounds more like a broken data file.
> 
> The real question here (for the OP) is:
> 
> WHAT ARE YOU TRYING TO DO?
> 
> There was a glimpse a few messages back that looked like you were trying
> to parse information out of an mbox-format mail folder.  (I.e. a flat
> file that has a concatenated series of mbox-format mail messages in it,
> with all the silliness and problems inherent in this format, like having
> to prefix body lines with ">" if they begin with the word "From".)
> 
> "I want to write a shell script to parse an mbox folder..." is enough
> to send most people running away screaming.  What other horrors are we
> in store for next?
> 
> Of course, that might be a red herring, since you didn't actually tell
> us what your goal is, or what your inputs are, and we're having to
> guess at the moment based on tiny hints and information leaks.

[toc] | [prev] | [next] | [standalone]


#194469 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 03:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAFfH-7nW-1@gated-at.bofh.it>
In reply to#194450
On Tuesday, April 03, 2018 08:30:04 AM Greg Wooledge wrote:
> WHAT ARE YOU TRYING TO DO?

I am building (have built several iterations) of a free format database to 
work something like askSam.  It is a mashup of several applications, things 
like recol, kmail, nail, kate and the data is stored in mbox formatted files.

Each record is treated as an email.

(Earlier iterations of the mashup used nedit, some nedit macros, and a custome 
file format using a 4 byte sequence (0x80 0x81 0x82 0x83) as the record 
separator.)

I occasionally want to sort the data in order by what I call the record title, 
which is stored in quotation marks ("") after the "From" in the mbox header.

I have seen some utilities that might sort the mbox files, but may require that 
I somehow move them into IMAP or some other manipulations that might be 
inconvenient.

msort was mentioned on this list within the last few months and might do the 
job for me if I can insert a temporary one byte record separator into the file 
(in addition to the mbox From line).  

Most likely this would be only a temporary addition, and I would need to do 
things like make sure that one byte will be unique in the file.  It sounds like 
there are at least a few candidates.


> 
> There was a glimpse a few messages back that looked like you were trying
> to parse information out of an mbox-format mail folder.  (I.e. a flat
> file that has a concatenated series of mbox-format mail messages in it,
> with all the silliness and problems inherent in this format, like having
> to prefix body lines with ">" if they begin with the word "From".)
> 
> "I want to write a shell script to parse an mbox folder..." is enough
> to send most people running away screaming.  What other horrors are we
> in store for next?
> 
> Of course, that might be a red herring, since you didn't actually tell
> us what your goal is, or what your inputs are, and we're having to
> guess at the moment based on tiny hints and information leaks.

[toc] | [prev] | [next] | [standalone]


#194473 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromNicolas George <george@nsup.org>
Date2018-04-04 13:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAOM2-5EV-9@gated-at.bofh.it>
In reply to#194469

[Multipart message — attachments visible in raw view] — view raw

rhkramer@gmail.com (2018-04-03):
>				and the data is stored in mbox formatted files.

DO NOT DO THAT.

This is the only good advice you can have for that project. Store your
data in a decent format.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194474 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 14:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAPoJ-6dv-9@gated-at.bofh.it>
In reply to#194473
Sorry, I already have 300 MB plus stored in that format.  Where were you in 
2000 when I started the project?

On Wednesday, April 04, 2018 07:23:25 AM Nicolas George wrote:
> rhkramer@gmail.com (2018-04-03):
> >				and the data is stored in mbox formatted files.
> 
> DO NOT DO THAT.
> 
> This is the only good advice you can have for that project. Store your
> data in a decent format.
> 
> Regards,

[toc] | [prev] | [next] | [standalone]


#194475 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromNicolas George <george@nsup.org>
Date2018-04-04 14:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAPoK-6dv-19@gated-at.bofh.it>
In reply to#194474

[Multipart message — attachments visible in raw view] — view raw

rhkramer@gmail.com (2018-04-04):
> Sorry, I already have 300 MB plus stored in that format.

Then convert. Small extra work now. Many less headaches later.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194478 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 14:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAQ1r-6sw-1@gated-at.bofh.it>
In reply to#194475
I'll convert the file format after you convert the programs to work with the 
different file format.  Those programs include kmail, nail, (essentially all 
email programs that use mbox as the file format), recoll (conversion should not 
be difficult), various editors (nedit, kate, for which I've written syntax 
highlighters / folders for the current format), and all scintilla based 
editors, for which I'm working on a highlighter / folder.

Let me know when you're almost finished, so I can make the conversion.

On Wednesday, April 04, 2018 08:09:11 AM Nicolas George wrote:
> rhkramer@gmail.com (2018-04-04):
> > Sorry, I already have 300 MB plus stored in that format.
> 
> Then convert. Small extra work now. Many less headaches later.
> 
> Regards,

[toc] | [prev] | [next] | [standalone]


#194479 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromNicolas George <george@nsup.org>
Date2018-04-04 15:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAQb8-6wq-5@gated-at.bofh.it>
In reply to#194478

[Multipart message — attachments visible in raw view] — view raw

rhkramer@gmail.com (2018-04-04):
> I'll convert the file format after you convert the programs to work with the 
> different file format.  Those programs include kmail, nail, (essentially all 
> email programs that use mbox as the file format), recoll (conversion should not 
> be difficult), various editors (nedit, kate, for which I've written syntax 
> highlighters / folders for the current format), and all scintilla based 
> editors, for which I'm working on a highlighter / folder.
> 
> Let me know when you're almost finished, so I can make the conversion.

I have given you advice (for free), you are not taking it. Too bad for
you. Good day.

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194482 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromAndre Majorel <aym-naibed@teaser.fr>
Date2018-04-04 16:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vARqx-7yA-3@gated-at.bofh.it>
In reply to#194479
On 2018-04-04 14:55 +0200, Nicolas George wrote:

> I have given you advice (for free), you are not taking it. Too bad for
> you. Good day.

Is advice that comes with condescension truly free ?

-- 
André Majorel <http://www.teaser.fr/~amajorel/>
I trust bugs.debian.org to not publish my email address for
spammers to harvest.

[toc] | [prev] | [next] | [standalone]


#194483 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-04 16:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vARAd-7D1-3@gated-at.bofh.it>
In reply to#194482
On Wed, Apr 04, 2018 at 04:15:48PM +0200, Andre Majorel wrote:
> On 2018-04-04 14:55 +0200, Nicolas George wrote:
> 
> > I have given you advice (for free), you are not taking it. Too bad for
> > you. Good day.
> 
> Is advice that comes with condescension truly free ?

Any advice that stops the OP from storing structured data in mbox format
flat files is a bargain at any price.

Sadly, it seems we've failed to find the magic argument to achieve
that result.  At some point, you just have to walk away.

[toc] | [prev] | [next] | [standalone]


#194484 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 18:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vASZk-8sv-7@gated-at.bofh.it>
In reply to#194483
On Wednesday, April 04, 2018 10:24:06 AM Greg Wooledge wrote:
> On Wed, Apr 04, 2018 at 04:15:48PM +0200, Andre Majorel wrote:
> > On 2018-04-04 14:55 +0200, Nicolas George wrote:
> > > I have given you advice (for free), you are not taking it. Too bad for
> > > you. Good day.
> > 
> > Is advice that comes with condescension truly free ?
> 
> Any advice that stops the OP from storing structured data in mbox format
> flat files is a bargain at any price.

Why do you call it "structured data"--it is free format data, no structure, or 
at least no regular / consistent structure.  It is anything that I want to put 
into it.


> 
> Sadly, it seems we've failed to find the magic argument to achieve
> that result.  At some point, you just have to walk away.

[toc] | [prev] | [next] | [standalone]


#194485 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 18:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vASZk-8sv-9@gated-at.bofh.it>
In reply to#194482
On Wednesday, April 04, 2018 10:15:48 AM Andre Majorel wrote:
> On 2018-04-04 14:55 +0200, Nicolas George wrote:
> > I have given you advice (for free), you are not taking it. Too bad for
> > you. Good day.
> 
> Is advice that comes with condescension truly free ?

Thank you!

[toc] | [prev] | [next] | [standalone]


Page 2 of 6 — ← Prev page 1 [2] 3 4 5 6  Next page →

Back to top | Article view | linux.debian.user


csiph-web