Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #194373 > unrolled thread
| Started by | mess-mate <mess-mate@gmx.com> |
|---|---|
| First post | 2018-04-01 16:10 +0200 |
| Last post | 2018-04-04 14:20 +0200 |
| Articles | 20 on this page of 101 — 21 participants |
Back to article view | Back to linux.debian.user
utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200
Page 2 of 6 — ← Prev page 1 [2] 3 4 5 6 Next page →
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-03 14:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAsLw-7pK-9@gated-at.bofh.it> |
| In reply to | #194441 |
[Multipart message — attachments visible in raw view] — view raw
rhkramer@gmail.com (2018-04-03): > Next I'll have to refresh my memory on how to replace the existing From with > From preceded by the null character, i.e., something like: > > Find: \n\nFrom > Replace with \n\n0x00\nFrom This is a very bad idea, and you are obviously about tu reproduce the errors of the past. You need to change your design. You obviously have free-form text, with possibly some rigid syntax but not enough for your needs. Therefore, you cannot use a delimiter inside the text. You have to put your structure outside the text. It is very typical of the "problems" people have with UTF-8: the problem resides not in the properties of UTF-8 but in the unwritten assumptions about the way they should be implementing things. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-03 14:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAtex-7RB-5@gated-at.bofh.it> |
| In reply to | #194443 |
On Tuesday, April 03, 2018 07:54:35 AM Nicolas George wrote: > rhkramer@gmail.com (2018-04-03): > > Next I'll have to refresh my memory on how to replace the existing From > > with From preceded by the null character, i.e., something like: > > > > Find: \n\nFrom > > Replace with \n\n0x00\nFrom > > This is a very bad idea, and you are obviously about tu reproduce the > errors of the past. > > You need to change your design. I don't understand the errors of the past nor is it feasible to change my design. Two points: * the other half of the above replacement is deleting the 0x00 after sorting (unless it is totally innocuous, which it might be) * my design is what I'd call a mashup of existing programs that use and require a specified file format (basically, mbox)--the programs that I use include kmail, nail, recol, kate, and, in the future, any editor which uses Scintilla, and, I would hope to be able to use any email program that can use mbox files. > You obviously have free-form text, with > possibly some rigid syntax but not enough for your needs. Therefore, you > cannot use a delimiter inside the text. You have to put your structure > outside the text. > > It is very typical of the "problems" people have with UTF-8: the problem > resides not in the properties of UTF-8 but in the unwritten assumptions > about the way they should be implementing things. > > Regards,
[toc] | [prev] | [next] | [standalone]
| From | Michael Lange <klappnase@freenet.de> |
|---|---|
| Date | 2018-04-03 14:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAsVb-7IQ-7@gated-at.bofh.it> |
| In reply to | #194441 |
Hi, On Tue, 3 Apr 2018 07:43:02 -0400 rhkramer@gmail.com wrote: > > maybe you could use the null byte? > > Thanks! > > Surprisingly (to me), this (and maybe several other of the control > characters might work--I did a search of one of the files, and there > are no null bytes. I believe (please anyone correct me if I am wrong) that "text" files won't contain any null byte; many text editors even refuse to open such a file, I guess since they assume it is a "binary" file. Probably it is the same with some other control characters like 04 (End of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII it seems like 1C (File Separator) or 1E (Record Separator) might be appropriate choices for you. I'm no expert on this, though. Best regards Michael .-.. .. ...- . .-.. --- -. --. .- -. -.. .--. .-. --- ... .--. . .-. Oh, that sound of male ego. You travel halfway across the galaxy and it's still the same song. -- Eve McHuron, "Mudd's Women", stardate 1330.1
[toc] | [prev] | [next] | [standalone]
| From | Michael Lange <klappnase@freenet.de> |
|---|---|
| Date | 2018-04-03 14:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAt4S-7NM-7@gated-at.bofh.it> |
| In reply to | #194445 |
On Tue, 3 Apr 2018 13:58:33 +0200 Michael Lange <klappnase@freenet.de> wrote: > I believe (please anyone correct me if I am wrong) that "text" files > won't contain any null byte; many text editors even refuse to open such > a file, I guess since they assume it is a "binary" file. > Probably it is the same with some other control characters like 04 (End > of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII > it seems like 1C (File Separator) or 1E (Record Separator) might be > appropriate choices for you. I'm no expert on this, though. Addendum: iirc (again please correct me if I am wrong) unix file names may contain (at least in theory) any byte except 2F (the slash) and the null byte. So if your text files might contain arbitrary file names there may be (at least in theory) a (admittedly very small) chance that such a file name actually might contain any control character except the null byte. Regards Michael .-.. .. ...- . .-.. --- -. --. .- -. -.. .--. .-. --- ... .--. . .-. Live long and prosper. -- Spock, "Amok Time", stardate 3372.7
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-03 14:40 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAtoe-7V4-21@gated-at.bofh.it> |
| In reply to | #194446 |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On Tue, Apr 03, 2018 at 02:14:07PM +0200, Michael Lange wrote: > On Tue, 3 Apr 2018 13:58:33 +0200 > Michael Lange <klappnase@freenet.de> wrote: > > > I believe (please anyone correct me if I am wrong) that "text" files > > won't contain any null byte; many text editors even refuse to open such > > a file, I guess since they assume it is a "binary" file. Emacs and vi(m) do open, edit and save such a file (and make no fuss about it). By default, both depict those NULL characters as ^@. Are there other editors? (yah, that was a bit snarky, I know ;-) > > Probably it is the same with some other control characters like 04 (End > > of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII > > it seems like 1C (File Separator) or 1E (Record Separator) might be > > appropriate choices for you. I'm no expert on this, though. Just assuming "this won't happen" is a sure recipe for some debugging fun. But perhaps the OP is looking for such fun. (S)he seems set on trying... > Addendum: iirc (again please correct me if I am wrong) unix file names > may contain (at least in theory) any byte except 2F (the slash) and the > null byte. So if your text files might contain arbitrary file names there > may be (at least in theory) a (admittedly very small) chance that such a > file name actually might contain any control character except the null > byte. You are correct on file names. Cheers - -- tomás -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) iEYEARECAAYFAlrDdEgACgkQBcgs9XrR2kYbPgCeIA04MoYQleL5IDw5wwmerx0o bqEAnA24L1+etC0tlCH2ExSdNigEPMDU =iV6I -----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | Michael Lange <klappnase@freenet.de> |
|---|---|
| Date | 2018-04-03 21:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAA6l-3PH-1@gated-at.bofh.it> |
| In reply to | #194449 |
Hi, On Tue, 3 Apr 2018 14:32:08 +0200 <tomas@tuxteam.de> wrote: > > > Probably it is the same with some other control characters like 04 > > > (End of Transmission). When I look at > > > https://en.wikipedia.org/wiki/ASCII it seems like 1C (File > > > Separator) or 1E (Record Separator) might be appropriate choices > > > for you. I'm no expert on this, though. > > Just assuming "this won't happen" is a sure recipe for some debugging > fun. But perhaps the OP is looking for such fun. (S)he seems set on > trying... > > > Addendum: iirc (again please correct me if I am wrong) unix file names > > may contain (at least in theory) any byte except 2F (the slash) and > > the null byte. So if your text files might contain arbitrary file > > names there may be (at least in theory) a (admittedly very small) > > chance that such a file name actually might contain any control > > character except the null byte. > > You are correct on file names. Thanks for the clarification. >From what i have understood I think the OP should certainly at least, whatever the files they want to include exactly look like and whichever byte they choose as delimiter, scan the file first for such a byte and if it is actually found replace it with either an empty string or (probably better) some sort of "tag" before applying the contents to the new database. This way they could at least be sure that their chosen delimiter does not split one record into halves. I don't think it makes really any difference which of the bytes that aren't supposed to be present in "text" files is used, one can never be 100% sure that these bytes don't show up nonetheless. I have no idea what these "text files" look like of course. It just seemed -to me - that the fact that the null byte cannot ever be part of a file name might make it slightly more appropriate for this purpose than other candidate bytes. Of course, it depends... Best regards Michael .-.. .. ...- . .-.. --- -. --. .- -. -.. .--. .-. --- ... .--. . .-. Suffocating together ... would create heroic camaraderie. -- Khan Noonian Singh, "Space Seed", stardate 3142.8
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-03 21:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAA6l-3PH-3@gated-at.bofh.it> |
| In reply to | #194457 |
On Tue, Apr 03, 2018 at 09:36:42PM +0200, Michael Lange wrote: > >From what i have understood I think the OP should certainly at least, > whatever the files they want to include exactly look like and whichever > byte they choose as delimiter, scan the file first for such a byte and if > it is actually found replace it with either an empty string or > (probably better) some sort of "tag" before applying the contents to the > new database. This way they could at least be sure that their chosen > delimiter does not split one record into halves. Or abort the program with an error message. > I have no idea what these "text files" look like of course. It just seemed > -to me - that the fact that the null byte cannot ever be part of a file > name might make it slightly more appropriate for this purpose than other > candidate bytes. Of course, it depends... NUL bytes are an excellent choice for delimiter in lots of situations. Of course, we still need the OP to tell us what the actual situation is.
[toc] | [prev] | [next] | [standalone]
| From | Michael Lange <klappnase@freenet.de> |
|---|---|
| Date | 2018-04-03 22:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAApH-4eF-3@gated-at.bofh.it> |
| In reply to | #194458 |
On Tue, 3 Apr 2018 15:47:57 -0400 Greg Wooledge <wooledg@eeg.ccf.org> wrote: > On Tue, Apr 03, 2018 at 09:36:42PM +0200, Michael Lange wrote: > > >From what i have understood I think the OP should certainly at least, > > whatever the files they want to include exactly look like and > > whichever byte they choose as delimiter, scan the file first for such > > a byte and if it is actually found replace it with either an empty > > string or (probably better) some sort of "tag" before applying the > > contents to the new database. This way they could at least be sure > > that their chosen delimiter does not split one record into halves. > > Or abort the program with an error message. Yes, or ask the user what to do with that particular file :) Regards Michael > > > I have no idea what these "text files" look like of course. It just > > seemed -to me - that the fact that the null byte cannot ever be part > > of a file name might make it slightly more appropriate for this > > purpose than other candidate bytes. Of course, it depends... > > NUL bytes are an excellent choice for delimiter in lots of situations. > > Of course, we still need the OP to tell us what the actual situation is. > .-.. .. ...- . .-.. --- -. --. .- -. -.. .--. .-. --- ... .--. . .-. Peace was the way. -- Kirk, "The City on the Edge of Forever", stardate unknown
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-03 14:40 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAtod-7V4-15@gated-at.bofh.it> |
| In reply to | #194446 |
> Addendum: iirc (again please correct me if I am wrong) unix file names > may contain (at least in theory) any byte except 2F (the slash) and the > null byte. So if your text files might contain arbitrary file names there > may be (at least in theory) a (admittedly very small) chance that such a > file name actually might contain any control character except the null > byte. One might question whether a file that contains a list of filenames is really a "text file". It sounds more like a broken data file. The real question here (for the OP) is: WHAT ARE YOU TRYING TO DO? There was a glimpse a few messages back that looked like you were trying to parse information out of an mbox-format mail folder. (I.e. a flat file that has a concatenated series of mbox-format mail messages in it, with all the silliness and problems inherent in this format, like having to prefix body lines with ">" if they begin with the word "From".) "I want to write a shell script to parse an mbox folder..." is enough to send most people running away screaming. What other horrors are we in store for next? Of course, that might be a red herring, since you didn't actually tell us what your goal is, or what your inputs are, and we're having to guess at the moment based on tiny hints and information leaks.
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 03:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAF61-7kF-1@gated-at.bofh.it> |
| In reply to | #194450 |
On Tuesday, April 03, 2018 08:30:04 AM Greg Wooledge wrote: > > Addendum: iirc (again please correct me if I am wrong) unix file names > > may contain (at least in theory) any byte except 2F (the slash) and the > > null byte. So if your text files might contain arbitrary file names there > > may be (at least in theory) a (admittedly very small) chance that such a > > file name actually might contain any control character except the null > > byte. > > One might question whether a file that contains a list of filenames is > really a "text file". It sounds more like a broken data file. > > The real question here (for the OP) is: > > WHAT ARE YOU TRYING TO DO? > > There was a glimpse a few messages back that looked like you were trying > to parse information out of an mbox-format mail folder. (I.e. a flat > file that has a concatenated series of mbox-format mail messages in it, > with all the silliness and problems inherent in this format, like having > to prefix body lines with ">" if they begin with the word "From".) > > "I want to write a shell script to parse an mbox folder..." is enough > to send most people running away screaming. What other horrors are we > in store for next? > > Of course, that might be a red herring, since you didn't actually tell > us what your goal is, or what your inputs are, and we're having to > guess at the moment based on tiny hints and information leaks.
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 03:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAFfH-7nW-1@gated-at.bofh.it> |
| In reply to | #194450 |
On Tuesday, April 03, 2018 08:30:04 AM Greg Wooledge wrote:
> WHAT ARE YOU TRYING TO DO?
I am building (have built several iterations) of a free format database to
work something like askSam. It is a mashup of several applications, things
like recol, kmail, nail, kate and the data is stored in mbox formatted files.
Each record is treated as an email.
(Earlier iterations of the mashup used nedit, some nedit macros, and a custome
file format using a 4 byte sequence (0x80 0x81 0x82 0x83) as the record
separator.)
I occasionally want to sort the data in order by what I call the record title,
which is stored in quotation marks ("") after the "From" in the mbox header.
I have seen some utilities that might sort the mbox files, but may require that
I somehow move them into IMAP or some other manipulations that might be
inconvenient.
msort was mentioned on this list within the last few months and might do the
job for me if I can insert a temporary one byte record separator into the file
(in addition to the mbox From line).
Most likely this would be only a temporary addition, and I would need to do
things like make sure that one byte will be unique in the file. It sounds like
there are at least a few candidates.
>
> There was a glimpse a few messages back that looked like you were trying
> to parse information out of an mbox-format mail folder. (I.e. a flat
> file that has a concatenated series of mbox-format mail messages in it,
> with all the silliness and problems inherent in this format, like having
> to prefix body lines with ">" if they begin with the word "From".)
>
> "I want to write a shell script to parse an mbox folder..." is enough
> to send most people running away screaming. What other horrors are we
> in store for next?
>
> Of course, that might be a red herring, since you didn't actually tell
> us what your goal is, or what your inputs are, and we're having to
> guess at the moment based on tiny hints and information leaks.
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 13:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAOM2-5EV-9@gated-at.bofh.it> |
| In reply to | #194469 |
[Multipart message — attachments visible in raw view] — view raw
rhkramer@gmail.com (2018-04-03): > and the data is stored in mbox formatted files. DO NOT DO THAT. This is the only good advice you can have for that project. Store your data in a decent format. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 14:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAPoJ-6dv-9@gated-at.bofh.it> |
| In reply to | #194473 |
Sorry, I already have 300 MB plus stored in that format. Where were you in 2000 when I started the project? On Wednesday, April 04, 2018 07:23:25 AM Nicolas George wrote: > rhkramer@gmail.com (2018-04-03): > > and the data is stored in mbox formatted files. > > DO NOT DO THAT. > > This is the only good advice you can have for that project. Store your > data in a decent format. > > Regards,
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 14:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAPoK-6dv-19@gated-at.bofh.it> |
| In reply to | #194474 |
[Multipart message — attachments visible in raw view] — view raw
rhkramer@gmail.com (2018-04-04): > Sorry, I already have 300 MB plus stored in that format. Then convert. Small extra work now. Many less headaches later. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 14:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAQ1r-6sw-1@gated-at.bofh.it> |
| In reply to | #194475 |
I'll convert the file format after you convert the programs to work with the different file format. Those programs include kmail, nail, (essentially all email programs that use mbox as the file format), recoll (conversion should not be difficult), various editors (nedit, kate, for which I've written syntax highlighters / folders for the current format), and all scintilla based editors, for which I'm working on a highlighter / folder. Let me know when you're almost finished, so I can make the conversion. On Wednesday, April 04, 2018 08:09:11 AM Nicolas George wrote: > rhkramer@gmail.com (2018-04-04): > > Sorry, I already have 300 MB plus stored in that format. > > Then convert. Small extra work now. Many less headaches later. > > Regards,
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 15:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAQb8-6wq-5@gated-at.bofh.it> |
| In reply to | #194478 |
[Multipart message — attachments visible in raw view] — view raw
rhkramer@gmail.com (2018-04-04): > I'll convert the file format after you convert the programs to work with the > different file format. Those programs include kmail, nail, (essentially all > email programs that use mbox as the file format), recoll (conversion should not > be difficult), various editors (nedit, kate, for which I've written syntax > highlighters / folders for the current format), and all scintilla based > editors, for which I'm working on a highlighter / folder. > > Let me know when you're almost finished, so I can make the conversion. I have given you advice (for free), you are not taking it. Too bad for you. Good day. -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Andre Majorel <aym-naibed@teaser.fr> |
|---|---|
| Date | 2018-04-04 16:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vARqx-7yA-3@gated-at.bofh.it> |
| In reply to | #194479 |
On 2018-04-04 14:55 +0200, Nicolas George wrote: > I have given you advice (for free), you are not taking it. Too bad for > you. Good day. Is advice that comes with condescension truly free ? -- André Majorel <http://www.teaser.fr/~amajorel/> I trust bugs.debian.org to not publish my email address for spammers to harvest.
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-04 16:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vARAd-7D1-3@gated-at.bofh.it> |
| In reply to | #194482 |
On Wed, Apr 04, 2018 at 04:15:48PM +0200, Andre Majorel wrote: > On 2018-04-04 14:55 +0200, Nicolas George wrote: > > > I have given you advice (for free), you are not taking it. Too bad for > > you. Good day. > > Is advice that comes with condescension truly free ? Any advice that stops the OP from storing structured data in mbox format flat files is a bargain at any price. Sadly, it seems we've failed to find the magic argument to achieve that result. At some point, you just have to walk away.
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 18:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vASZk-8sv-7@gated-at.bofh.it> |
| In reply to | #194483 |
On Wednesday, April 04, 2018 10:24:06 AM Greg Wooledge wrote: > On Wed, Apr 04, 2018 at 04:15:48PM +0200, Andre Majorel wrote: > > On 2018-04-04 14:55 +0200, Nicolas George wrote: > > > I have given you advice (for free), you are not taking it. Too bad for > > > you. Good day. > > > > Is advice that comes with condescension truly free ? > > Any advice that stops the OP from storing structured data in mbox format > flat files is a bargain at any price. Why do you call it "structured data"--it is free format data, no structure, or at least no regular / consistent structure. It is anything that I want to put into it. > > Sadly, it seems we've failed to find the magic argument to achieve > that result. At some point, you just have to walk away.
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 18:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vASZk-8sv-9@gated-at.bofh.it> |
| In reply to | #194482 |
On Wednesday, April 04, 2018 10:15:48 AM Andre Majorel wrote: > On 2018-04-04 14:55 +0200, Nicolas George wrote: > > I have given you advice (for free), you are not taking it. Too bad for > > you. Good day. > > Is advice that comes with condescension truly free ? Thank you!
[toc] | [prev] | [next] | [standalone]
Page 2 of 6 — ← Prev page 1 [2] 3 4 5 6 Next page →
Back to top | Article view | linux.debian.user
csiph-web