Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #194373 > unrolled thread
| Started by | mess-mate <mess-mate@gmx.com> |
|---|---|
| First post | 2018-04-01 16:10 +0200 |
| Last post | 2018-04-04 14:20 +0200 |
| Articles | 20 on this page of 101 — 21 participants |
Back to article view | Back to linux.debian.user
utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200
Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6 Next page →
| From | Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> |
|---|---|
| Date | 2018-04-04 20:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVDP-1KU-11@gated-at.bofh.it> |
| In reply to | #194474 |
rhkramer: > Where were you in 2000 when I started the project? > I cannot speak for anyone else, but I was probably once again giving a frequently given answer that I eventually put up on a WWW page. http://jdebp.eu./FGA/mail-mbox-formats.html
[toc] | [prev] | [next] | [standalone]
| From | Greg Wooledge <wooledg@eeg.ccf.org> |
|---|---|
| Date | 2018-04-04 14:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAPI6-6l6-13@gated-at.bofh.it> |
| In reply to | #194473 |
On Wed, Apr 04, 2018 at 01:23:25PM +0200, Nicolas George wrote: > rhkramer@gmail.com (2018-04-03): > > and the data is stored in mbox formatted files. > > DO NOT DO THAT. > > This is the only good advice you can have for that project. Store your > data in a decent format. Perhaps an sqlite database. At least, that is my first thought. You might come up with a better solution depending on your specific needs. That solution won't be "an mbox folder".
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 15:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAQkN-6Pk-1@gated-at.bofh.it> |
| In reply to | #194477 |
On Wednesday, April 04, 2018 08:26:41 AM Greg Wooledge wrote: > On Wed, Apr 04, 2018 at 01:23:25PM +0200, Nicolas George wrote: > > rhkramer@gmail.com (2018-04-03): > > > and the data is stored in mbox formatted files. > > > > DO NOT DO THAT. > > > > This is the only good advice you can have for that project. Store your > > data in a decent format. > > Perhaps an sqlite database. At least, that is my first thought. > You might come up with a better solution depending on your specific > needs. That solution won't be "an mbox folder". Past experience with "databases" (before I switched to LInux)--things like dBase (III, III+, IV), Microsoft Access, and others that I can't recall atm made the fixed (and even variable) length fields. The key thing for me was making something that worked reasonably like askSam, which is / was free format, fully searchable, and other things I can't recall atm. I don't know enough about sqlite to know what capabilities it has for variable length fields (if any--how many are possible per record, is the content fully searchable, etc.) But, it really doesn't matter, I am not interested in changing the data format. (Besides, with the current design, almost any email client is a client for my mashup.)
[toc] | [prev] | [next] | [standalone]
| From | Don Armstrong <don@debian.org> |
|---|---|
| Date | 2018-04-04 19:40 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAUy5-14B-5@gated-at.bofh.it> |
| In reply to | #194469 |
On Tue, 03 Apr 2018, rhkramer@gmail.com wrote: > I am building (have built several iterations) of a free format > database to work something like askSam. It is a mashup of several > applications, things like recol, kmail, nail, kate and the data is > stored in mbox formatted files. > > Each record is treated as an email. You should consider looking at using Maildir with notmuch and using things which integrate notmuch.[1] > Most likely this would be only a temporary addition, and I would need > to do things like make sure that one byte will be unique in the file. > It sounds like there are at least a few candidates. Maildir is the solution to this. While you *can* handle mbox and do all the escape rules properly (From to >From and back), it's a pain. Let your filesystem handle it for you. [I'm speaking from experience; I currently maintain debbugs, which basically stores everything in a custom format mbox. This inevitably makes things slow, as you have to search through the mbox linearly to find any message in the mbox unless you also write indexes for the mailbox.] 1: Notmuch itself uses xapian to do the heavy lifting. -- Don Armstrong https://www.donarmstrong.com Cheop's Law: Nothing ever gets built on schedule or within budget. -- Robert Heinlein _Time Enough For Love_ p242
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 19:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAUHM-17T-3@gated-at.bofh.it> |
| In reply to | #194490 |
[Multipart message — attachments visible in raw view] — view raw
Don Armstrong (2018-04-04): > You should consider looking at using Maildir with notmuch and using > things which integrate notmuch.[1] Maildir is not that much better than mbox. Sure, it eliminates most of its worse flaws, but it brings flaws of its own, like trashing the inode and dentries caches, requiring extra disk reads and cache due to partial file ends, or causing much more seeking (granted, this one is becoming less of an issue with non-mechanical storage). The filesystem is really the least common factor of database systems (no, I did not mix LCM and GCD). There is a reason people designed more advanced and optimized formats on top of it. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | Don Armstrong <don@debian.org> |
|---|---|
| Date | 2018-04-04 20:00 +0200 |
| Subject | mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] |
| Message-ID | <vAURr-1cT-3@gated-at.bofh.it> |
| In reply to | #194493 |
On Wed, 04 Apr 2018, Nicolas George wrote: > Don Armstrong (2018-04-04): > > You should consider looking at using Maildir with notmuch and using > > things which integrate notmuch.[1] > > Maildir is not that much better than mbox. Sure, it eliminates most of > its worse flaws, but it brings flaws of its own, like trashing the > inode and dentries caches, requiring extra disk reads and cache due to > partial file ends, or causing much more seeking (granted, this one is > becoming less of an issue with non-mechanical storage). There are definitely better formats than Maildir, like Dovecot's multi-dbox.[1] These issues are why almost everyone who uses Maildir just uses it as the backing message store and uses the index on top to do avoid ever reading all of the messages in the Maildir. 1: https://wiki2.dovecot.org/MailboxFormat/dbox -- Don Armstrong https://www.donarmstrong.com A Bill of Rights that means what the majority wants it to mean is worthless. -- U.S. Supreme Court Justice Antonin Scalia
[toc] | [prev] | [next] | [standalone]
| From | Nicolas George <george@nsup.org> |
|---|---|
| Date | 2018-04-04 20:10 +0200 |
| Subject | Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)] |
| Message-ID | <vAV18-1vZ-9@gated-at.bofh.it> |
| In reply to | #194496 |
[Multipart message — attachments visible in raw view] — view raw
Don Armstrong (2018-04-04): > There are definitely better formats than Maildir, like Dovecot's > multi-dbox.[1] > > These issues are why almost everyone who uses Maildir just uses it as > the backing message store and uses the index on top to do avoid ever > reading all of the messages in the Maildir. Glad to read this. There are too many maildir zealots out there. Regards, -- Nicolas George
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 20:40 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVu9-1Hd-1@gated-at.bofh.it> |
| In reply to | #194490 |
On Wednesday, April 04, 2018 01:36:15 PM Don Armstrong wrote: > On Tue, 03 Apr 2018, rhkramer@gmail.com wrote: > > I am building (have built several iterations) of a free format > > database to work something like askSam. It is a mashup of several > > applications, things like recol, kmail, nail, kate and the data is > > stored in mbox formatted files. > > > > Each record is treated as an email. > > You should consider looking at using Maildir with notmuch and using > things which integrate notmuch.[1] > > > Most likely this would be only a temporary addition, and I would need > > to do things like make sure that one byte will be unique in the file. > > It sounds like there are at least a few candidates. > > Maildir is the solution to this. While you *can* handle mbox and do all > the escape rules properly (From to >From and back), it's a pain. Let > your filesystem handle it for you. > > [I'm speaking from experience; I currently maintain debbugs, which > basically stores everything in a custom format mbox. This inevitably > makes things slow, as you have to search through the mbox linearly to > find any message in the mbox unless you also write indexes for the > mailbox.] > > 1: Notmuch itself uses xapian to do the heavy lifting. I'll probably look into notmuch, just for kicks. I've considered maildir--it meets some of my requirements (that is, to make something close to an askSam workalike), but one drawback is that it is essentially one email (i.e., my "record"). One of the desirable features of askSam is that you did not have to create a new file to add a new note / record, you just start typing in an existing open record and then, as time or other constraints allow, you can add more "tags" or a record separator. (It's been so long since I've used askSam I actually forget what had to be done (f anything) to separate a new record from the previous record). askSam basically stores all it's records in one file, although it is (of course) possible to separate them.
[toc] | [prev] | [next] | [standalone]
| From | Don Armstrong <don@debian.org> |
|---|---|
| Date | 2018-04-04 20:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVDP-1KU-9@gated-at.bofh.it> |
| In reply to | #194501 |
On Wed, 04 Apr 2018, rhkramer@gmail.com wrote: > I've considered maildir--it meets some of my requirements (that is, to > make something close to an askSam workalike), but one drawback is that > it is essentially one email (i.e., my "record"). One of the desirable > features of askSam is that you did not have to create a new file to > add a new note / record, you just start typing in an existing open > record and then, as time or other constraints allow, you can add more > "tags" or a record separator. (It's been so long since I've used > askSam I actually forget what had to be done (f anything) to separate > a new record from the previous record). You might want to consider looking at org-mode too.[1] There are even integrations for notmuch+mutt+Maildir there. 1: https://www.youtube.com/watch?v=oJTwQvgfgMM -- Don Armstrong https://www.donarmstrong.com I really wanted to talk to her. I just couldn't find an algorithm that fit. -- Peter Watts _Blindsight_ p294
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-05 14:40 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vBclj-4wA-1@gated-at.bofh.it> |
| In reply to | #194503 |
On Wednesday, April 04, 2018 02:45:49 PM Don Armstrong wrote: > On Wed, 04 Apr 2018, rhkramer@gmail.com wrote: > > I've considered maildir--it meets some of my requirements (that is, to > > make something close to an askSam workalike), but one drawback is that > > it is essentially one email (i.e., my "record"). One of the desirable > > features of askSam is that you did not have to create a new file to > > add a new note / record, you just start typing in an existing open > > record and then, as time or other constraints allow, you can add more > > "tags" or a record separator. (It's been so long since I've used > > askSam I actually forget what had to be done (f anything) to separate > > a new record from the previous record). > > You might want to consider looking at org-mode too.[1] There are > even integrations for notmuch+mutt+Maildir there. > > 1: https://www.youtube.com/watch?v=oJTwQvgfgMM Thanks for the link / pointer to org-mode. From the little I've looked at, it has several similarities to what I'm doing--for example, the collapsible outlining / folding, and the use of a mark-up language (I use the TWiki markup language with a few extensions / modifications--oh, and that reminds me (if anyone is keeping track of my mashup)--TWiki / Foswiki are other programs that are part of the mashup--records in my mashup are maintained in a form that would work as a page (they use a different word--ohh, topic) of a TWiki / Foswiki, and someday I'd like to have an automatic import / export facility-- in my home mashup. With that automatic import / export facility, I could designate certain records to be "public" (or something like that) which would mean: (1) if they were not on my public TWiki / Foswiki, they would be automatically exported (when I was connected to the Internet and designated that it was appropriate to "sync" my home storage with the public wiki), and (2) any public wiki pages that might have been modified (by others) would be copied to the mashup as backup. (Also, if a mashup page was designated as "public", after I made changes on the mashup they would (at an appropriate time) be uploaded to the public wiki. I could discuss my hate-hate relationship with Emacs--I tried a few times to learn Emacs, and had difficulty for a variety of reasons. Part of it was my hate-hate relationship with Lisp, another part was that Emacs (at the time) seemed much less GUI / mouse friendly than the editors I had grown used to in my DOS / Windows days (even though I learned (and liked) a lot of shortcut keys--I guess maybe an early encounter (with shortcut keys) was with Wordstar, and then I used a shareware editor (that I paid for--high praise indeed from me) that used the same set of shortcut keys--I can't remember the name of that editor atm. Anyway, thanks for prompting me to reminisce. I know that at the times I looked at Emacs (and Xemacs) they had outline mode, I'm not sure I recall org mode. I expect I will spend a little more time looking further into org mode, although I think my mashup has or will have the features I've seen there so far. (Aside: My mashup uses kate as the editor, but because of some long standing bugs in kate (which, finally after several years were supposedly fixed a few years ago) I had to add closing markup to the opening markup of TWiki / Foswiki. In the meantime, I decided I wanted to switch to using any Scintilla based editor (there are a lot) as the editor, but so far, have not got a folder / highlighter written for Scintilla. (And, the one time I tried to test the fixed Kate, it didn't seem to work as promised, and I just left it as is, still requiring the ending markup.)) Thanks again for the response!
[toc] | [prev] | [next] | [standalone]
| From | deloptes <deloptes@gmail.com> |
|---|---|
| Date | 2018-04-05 00:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAYV3-4f0-3@gated-at.bofh.it> |
| In reply to | #194501 |
rhkramer@gmail.com wrote: > I'll probably look into notmuch, just for kicks. > > I've considered maildir--it meets some of my requirements (that is, to > make something close to an askSam workalike), but one drawback is that it > is essentially one email (i.e., my "record"). One of the desirable > features of askSam is that you did not have to create a new file to add a > new note / record, you just start typing in an existing open record and > then, as time or other constraints allow, you can add more "tags" or a > record separator. (It's been so long since I've used askSam I actually > forget what had to be done (f anything) to separate a new record from the > previous record). > > askSam basically stores all it's records in one file, although it is (of > course) possible to separate them. I still don't understand why not use XML. If file is not getting too big (which is also a problem with mbox). You may need to adapt your applications, but it won't be more effort then making shit out of shit - sorry for my language. regards
[toc] | [prev] | [next] | [standalone]
| From | rhkramer@gmail.com |
|---|---|
| Date | 2018-04-04 21:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAWgx-2ha-13@gated-at.bofh.it> |
| In reply to | #194490 |
On Wednesday, April 04, 2018 01:36:15 PM Don Armstrong wrote: > On Tue, 03 Apr 2018, rhkramer@gmail.com wrote: > > I am building (have built several iterations) of a free format > > database to work something like askSam. It is a mashup of several > > applications, things like recol, kmail, nail, kate and the data is > > stored in mbox formatted files. > > > > Each record is treated as an email. > > You should consider looking at using Maildir with notmuch and using > things which integrate notmuch.[1] Ahh, OK, notmuch looks like it could be an alternative to recol as part of my mashup. Thanks, I'll probably dig deeper as time goes on.
[toc] | [prev] | [next] | [standalone]
| From | Henrique de Moraes Holschuh <hmh@debian.org> |
|---|---|
| Date | 2018-04-04 13:20 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAOCl-5AM-7@gated-at.bofh.it> |
| In reply to | #194445 |
On Tue, 03 Apr 2018, Michael Lange wrote: > I believe (please anyone correct me if I am wrong) that "text" files > won't contain any null byte; many text editors even refuse to open such a Depends on the encoding. For ASCII, ISO-8859-* and UTF-8 (and any other modern encoding AFAIK, other than modified UTF-8), any zero bytes map one-to-one to the NUL character/code point. I don't recall how it is on other common encodings of the 80's and 90's, though. Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use bytes with the value of zero when encoding characters, so NUL is encoded by a different sequence, and you can safely use a byte with the value of zero for some out-of-band control (like zero-terminated strings that can contain NULs, etc) -- note that NUL is a character, and it might be represented by a sequence of bytes that has nothing to do with zeroes on a particular encoding... (in fact, C strings are *zero-terminated*, not NUL-terminated, but most of the time this is irrelevant :p). Also, a text file MAY contain NULs (the character), it is just considered bad practice (nowadays?). Don't assume you won't see any. For example, received e-mail is *more* likely to have NULs in it than normal text due to the quality of some mail agents out there. I recall postfix would reject a *lot* of crap when we configured it to refuse to accept NULs outside of 8-bit bodies, because Cyrus-IMAPd *refuses* any such crap, and we wanted it bounced as early as possible. (note that NULs are forbidden in MIME-compliant email text and ESMTP, unless encoded or guarded by a 8-bit transfer area of known size, so there you have it: NULs in one text format that actually forbids them!). > Probably it is the same with some other control characters like 04 (End > of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII > it seems like 1C (File Separator) or 1E (Record Separator) might be > appropriate choices for you. I'm no expert on this, though. Well, ASCII control characters were inherited by ISO-8859-* and Unicode, so yes, you can use them. But so could the data file. It would be perfectly ok for a text data file to use the record separator control characters to delimit records in a table, for example... Here's a good definition of them (follow the hyperlinks for the definition of each control character): https://en.wikipedia.org/wiki/Basic_Latin_(Unicode_block) Here is also a proper solution: use modified UTF-8 (which encodes NUL so that zero bytes are *never* present in the stream): encode every input format to modified UTF-8, then add the zero-byte separators you want. You'll have to normalize the input data set into known charset/encodings and then recode them to modified UTF-8, of course. You can't blindly call any random data "UTF-8" (let alone modified UTF-8) and expect things not to break horribly. -- Henrique Holschuh
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-04 16:10 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vARgR-7uD-3@gated-at.bofh.it> |
| In reply to | #194472 |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On Wed, Apr 04, 2018 at 08:18:23AM -0300, Henrique de Moraes Holschuh wrote: > On Tue, 03 Apr 2018, Michael Lange wrote: > > I believe (please anyone correct me if I am wrong) that "text" files > > won't contain any null byte; many text editors even refuse to open such a > > Depends on the encoding. For ASCII, ISO-8859-* and UTF-8 (and any other > modern encoding AFAIK, other than modified UTF-8), any zero bytes map > one-to-one to the NUL character/code point. I don't recall how it is on > other common encodings of the 80's and 90's, though. Try UTF-16, what Microsoft (and a couple of years ago Apple) love to call "Unicode": in more "Western" contexts every second byte is NULL! > Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use > bytes with the value of zero when encoding characters, so NUL is encoded > by a different sequence, and you can safely use a byte with the value of > zero for some out-of-band control [...] Yes, the problem is that someone else before you could have been doing exactly that. I'd guard against that. It's not exactly difficult, the traditional "escape" mechanism (aka character stuffing) does it pretty well... Cheers - -- tomás -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) iEYEARECAAYFAlrE3JcACgkQBcgs9XrR2kYW7ACeMG0SQB23RSySoeSJBItB+Eji QEgAnipwAcoVJuzynJVBO1CR2rrLeuFs =xhja -----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | Henrique de Moraes Holschuh <hmh@debian.org> |
|---|---|
| Date | 2018-04-04 20:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVDP-1KU-7@gated-at.bofh.it> |
| In reply to | #194481 |
On Wed, 04 Apr 2018, tomas@tuxteam.de wrote: > On Wed, Apr 04, 2018 at 08:18:23AM -0300, Henrique de Moraes Holschuh wrote: > > On Tue, 03 Apr 2018, Michael Lange wrote: > > > I believe (please anyone correct me if I am wrong) that "text" files > > > won't contain any null byte; many text editors even refuse to open such a > > > > Depends on the encoding. For ASCII, ISO-8859-* and UTF-8 (and any other > > modern encoding AFAIK, other than modified UTF-8), any zero bytes map > > one-to-one to the NUL character/code point. I don't recall how it is on > > other common encodings of the 80's and 90's, though. > > Try UTF-16, what Microsoft (and a couple of years ago Apple) love to > call "Unicode": in more "Western" contexts every second byte is NULL! Ah, yes. I forgot about them, indeed. UTF-16BE and UTF-16LE will have zero bytes in the resulting byte stream. And I suppose one could call them "modern encodings", even if they are horrifying to use when compared to UTF-8 (UTF-16 has byte-order issues) or UTF-32 (UTF-16 has surrogate pairs). > > Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use > > bytes with the value of zero when encoding characters, so NUL is encoded > > by a different sequence, and you can safely use a byte with the value of > > zero for some out-of-band control [...] > > Yes, the problem is that someone else before you could have been doing > exactly that. You can modified-UTF-8 bit-packing to encode anything, and the result will be zero-free (and it will restore the zeroes when decoded). The price is a size increase (it is a variant of UTF-8 that uses two bytes to encode NUL, which would take just one byte in normal UTF-8). There are much better bit packing schemes if you just need to escape zeroes ;-) That said, it is always safe to break valid "modified UTF-8" into records using zeroes, as long as you don't expect the result to be valid UTF-8 (it isn't valid UTF-8 because NULs will be encoded using a non-minimal byte sequence that *will* decode to a zero even if it is invalid) or valid modified UTF-8 (it isn't valid modified UTF-8 because 0 is not valid as an encoding for NUL in modified UTF-8). But a lax UTF-8 or modified UTF-8 *would* parse "modified UTF-8 with zero as record separators" and reconstruct the unicode text properly (but it would read the record separators as NULs, so you'd get extra NULs in the resulting text). That, of course, assumes you have unicode text as the input (encoding doesn't matter, as long as you know it), and recode it to modified UTF-8 before you add the zeroes as end-of-record marks. This is not about bit-packing generic binary data. > I'd guard against that. It's not exactly difficult, the traditional > "escape" mechanism (aka character stuffing) does it pretty well... Yes, any bitstuffing/escape-based wrapping would do. -- Henrique Holschuh
[toc] | [prev] | [next] | [standalone]
| From | <tomas@tuxteam.de> |
|---|---|
| Date | 2018-04-04 21:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAWgx-2ha-3@gated-at.bofh.it> |
| In reply to | #194502 |
-----BEGIN PGP SIGNED MESSAGE----- Hash: SHA1 On Wed, Apr 04, 2018 at 03:44:23PM -0300, Henrique de Moraes Holschuh wrote: [...] > That said, it is always safe to break valid "modified UTF-8" into > records using zeroes, as long as you don't expect the result to be valid > UTF-8 (it isn't valid UTF-8 because NULs will be encoded using a > non-minimal byte sequence that *will* decode to a zero even if it is > invalid) or valid modified UTF-8 (it isn't valid modified UTF-8 because > 0 is not valid as an encoding for NUL in modified UTF-8). But a lax > UTF-8 or modified UTF-8 *would* parse "modified UTF-8 with zero as > record separators" and reconstruct the unicode text properly (but it > would read the record separators as NULs, so you'd get extra NULs in the > resulting text). You are a nasty guy, aren't you ;-) Pretty cunning... Cheers - -- t -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.4.12 (GNU/Linux) iEYEARECAAYFAlrFJngACgkQBcgs9XrR2kZqLgCdEuap+rqSU6HCrXpkL6XHl3Az lRUAnjwGhiMNNlY+SXwIxpd/kfnvst1z =kHBa -----END PGP SIGNATURE-----
[toc] | [prev] | [next] | [standalone]
| From | Ben Caradoc-Davies <ben@transient.nz> |
|---|---|
| Date | 2018-04-04 23:50 +0200 |
| Subject | Re: Invalid UTF-8 byte? |
| Message-ID | <vAYs1-3MP-3@gated-at.bofh.it> |
| In reply to | #194481 |
On 05/04/18 02:09, tomas@tuxteam.de wrote: > Try UTF-16, what Microsoft (and a couple of years ago Apple) love to > call "Unicode": in more "Western" contexts every second byte is NULL! The Java platform uses UTF-16 internally: "The char data type (and therefore the value that a Character object encapsulates) are based on the original Unicode specification, which defined characters as fixed-width 16-bit entities." https://docs.oracle.com/javase/8/docs/api/java/lang/Character.html Kind regards, -- Ben Caradoc-Davies <ben@transient.nz> Director Transient Software Limited <https://transient.nz/> New Zealand
[toc] | [prev] | [next] | [standalone]
| From | Michael Stone <mstone@debian.org> |
|---|---|
| Date | 2018-04-05 01:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? |
| Message-ID | <vAZxL-4ze-1@gated-at.bofh.it> |
| In reply to | #194513 |
On Thu, Apr 05, 2018 at 09:42:19AM +1200, Ben Caradoc-Davies wrote: >On 05/04/18 02:09, tomas@tuxteam.de wrote: >>Try UTF-16, what Microsoft (and a couple of years ago Apple) love to >>call "Unicode": in more "Western" contexts every second byte is NULL! > >The Java platform uses UTF-16 internally: Yes, many people thought UCS-2 was the answer back when 16 bits was enough for anybody. Mike Stone
[toc] | [prev] | [next] | [standalone]
| From | Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> |
|---|---|
| Date | 2018-04-04 21:00 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVNB-1On-15@gated-at.bofh.it> |
| In reply to | #194472 |
Henrique de Moraes Holschuh: > Also, a text file MAY contain NULs (the character), it is just > considered bad practice (nowadays?). Don't assume you won't see any. > For example, received e-mail is *more* likely to have NULs in it than > normal text due to the quality of some mail agents out there. > I suspect not as likely as anything that was in the process of being appended to on a not-fully-journalling filesystem when a dirty shutdown happens. (-: * https://askubuntu.com/questions/356981/ Or anything that "rotates" output files by truncating them and pulls the rug out from underneath an old-style simplistic indefinitely-running text output writer. * http://jdebp.eu./FGA/do-not-use-logrotate.html#Background
[toc] | [prev] | [next] | [standalone]
| From | Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> |
|---|---|
| Date | 2018-04-04 20:30 +0200 |
| Subject | Re: Invalid UTF-8 byte? (was: Re: utf) |
| Message-ID | <vAVku-1DA-13@gated-at.bofh.it> |
| In reply to | #194391 |
rhkramer: > The reason I wanted such a byte was to use it as a record separator in > a set of text files (that I use as an askSam "workalike" (or > "worksimilar") so that I could use msort (which depends on a 1 byte > record separator to --separate the records ;-) while sorting. Some of > the files already include UTF-8, and, in the future, I anticpate all > will be in UTFF-8. > Note that ISO 646, hence ISO 8859, hence ISO 10646, has had a single-byte Record Separator character since the 1960s. (-:
[toc] | [prev] | [next] | [standalone]
Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6 Next page →
Back to top | Article view | linux.debian.user
csiph-web