Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #194373 > unrolled thread

utf

Started bymess-mate <mess-mate@gmx.com>
First post2018-04-01 16:10 +0200
Last post2018-04-04 14:20 +0200
Articles 20 on this page of 101 — 21 participants

Back to article view | Back to linux.debian.user


Contents

  utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
    Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
    Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
    Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
    Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
      Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
        Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
      Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
        Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
                              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
                                Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
                                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
                          mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re:  utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
                            Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was:  Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
                    Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
                      Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
        Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
        Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
            Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
          Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
          Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
            Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
              Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
                Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
                    Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
                      Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
                            Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
                              Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
                              Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
                                Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
                            Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
                    Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
                      Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
                        Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
                          Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
                            Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
                              Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
                                Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
                                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
                                  Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
                                    Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
                                      Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
                                      Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
                              Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
                            Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
                Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
          Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200

Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6  Next page →


#194504 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromJonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM>
Date2018-04-04 20:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVDP-1KU-11@gated-at.bofh.it>
In reply to#194474
rhkramer:

> Where were you in 2000 when I started the project?
>
I cannot speak for anyone else, but I was probably once again giving a 
frequently given answer that I eventually put up on a WWW page.

http://jdebp.eu./FGA/mail-mbox-formats.html

[toc] | [prev] | [next] | [standalone]


#194477 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromGreg Wooledge <wooledg@eeg.ccf.org>
Date2018-04-04 14:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAPI6-6l6-13@gated-at.bofh.it>
In reply to#194473
On Wed, Apr 04, 2018 at 01:23:25PM +0200, Nicolas George wrote:
> rhkramer@gmail.com (2018-04-03):
> >				and the data is stored in mbox formatted files.
> 
> DO NOT DO THAT.
> 
> This is the only good advice you can have for that project. Store your
> data in a decent format.

Perhaps an sqlite database.  At least, that is my first thought.
You might come up with a better solution depending on your specific
needs.  That solution won't be "an mbox folder".

[toc] | [prev] | [next] | [standalone]


#194480 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 15:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAQkN-6Pk-1@gated-at.bofh.it>
In reply to#194477
On Wednesday, April 04, 2018 08:26:41 AM Greg Wooledge wrote:
> On Wed, Apr 04, 2018 at 01:23:25PM +0200, Nicolas George wrote:
> > rhkramer@gmail.com (2018-04-03):
> > >				and the data is stored in mbox formatted files.
> > 
> > DO NOT DO THAT.
> > 
> > This is the only good advice you can have for that project. Store your
> > data in a decent format.
> 
> Perhaps an sqlite database.  At least, that is my first thought.
> You might come up with a better solution depending on your specific
> needs.  That solution won't be "an mbox folder".

Past experience with "databases" (before I switched to LInux)--things like 
dBase (III, III+, IV), Microsoft Access, and others that I can't recall atm 
made the fixed (and even variable) length fields.

The key thing for me was making something that worked reasonably like askSam, 
which is / was free format, fully searchable, and other things I can't recall 
atm.

I don't know enough about sqlite to know what capabilities it has for variable 
length fields (if any--how many are possible per record, is the content fully 
searchable, etc.)

But, it really doesn't matter, I am not interested in changing the data 
format.

(Besides, with the current design, almost any email client is a client for my 
mashup.)

[toc] | [prev] | [next] | [standalone]


#194490 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromDon Armstrong <don@debian.org>
Date2018-04-04 19:40 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAUy5-14B-5@gated-at.bofh.it>
In reply to#194469
On Tue, 03 Apr 2018, rhkramer@gmail.com wrote:
> I am building (have built several iterations) of a free format
> database to work something like askSam. It is a mashup of several
> applications, things like recol, kmail, nail, kate and the data is
> stored in mbox formatted files.
> 
> Each record is treated as an email.

You should consider looking at using Maildir with notmuch and using
things which integrate notmuch.[1]

> Most likely this would be only a temporary addition, and I would need
> to do things like make sure that one byte will be unique in the file.
> It sounds like there are at least a few candidates.

Maildir is the solution to this. While you *can* handle mbox and do all
the escape rules properly (From to >From and back), it's a pain. Let
your filesystem handle it for you.

[I'm speaking from experience; I currently maintain debbugs, which
basically stores everything in a custom format mbox. This inevitably
makes things slow, as you have to search through the mbox linearly to
find any message in the mbox unless you also write indexes for the
mailbox.]

1: Notmuch itself uses xapian to do the heavy lifting.
-- 
Don Armstrong                      https://www.donarmstrong.com

Cheop's Law: Nothing ever gets built on schedule or within budget.
 -- Robert Heinlein _Time Enough For Love_ p242

[toc] | [prev] | [next] | [standalone]


#194493 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromNicolas George <george@nsup.org>
Date2018-04-04 19:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAUHM-17T-3@gated-at.bofh.it>
In reply to#194490

[Multipart message — attachments visible in raw view] — view raw

Don Armstrong (2018-04-04):
> You should consider looking at using Maildir with notmuch and using
> things which integrate notmuch.[1]

Maildir is not that much better than mbox. Sure, it eliminates most of
its worse flaws, but it brings flaws of its own, like trashing the inode
and dentries caches, requiring extra disk reads and cache due to partial
file ends, or causing much more seeking (granted, this one is becoming
less of an issue with non-mechanical storage).

The filesystem is really the least common factor of database systems
(no, I did not mix LCM and GCD). There is a reason people designed more
advanced and optimized formats on top of it.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194496 — mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)]

FromDon Armstrong <don@debian.org>
Date2018-04-04 20:00 +0200
Subjectmbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)]
Message-ID<vAURr-1cT-3@gated-at.bofh.it>
In reply to#194493
On Wed, 04 Apr 2018, Nicolas George wrote:
> Don Armstrong (2018-04-04):
> > You should consider looking at using Maildir with notmuch and using
> > things which integrate notmuch.[1]
> 
> Maildir is not that much better than mbox. Sure, it eliminates most of
> its worse flaws, but it brings flaws of its own, like trashing the
> inode and dentries caches, requiring extra disk reads and cache due to
> partial file ends, or causing much more seeking (granted, this one is
> becoming less of an issue with non-mechanical storage).

There are definitely better formats than Maildir, like Dovecot's
multi-dbox.[1]

These issues are why almost everyone who uses Maildir just uses it as
the backing message store and uses the index on top to do avoid ever
reading all of the messages in the Maildir.

1: https://wiki2.dovecot.org/MailboxFormat/dbox
-- 
Don Armstrong                      https://www.donarmstrong.com

A Bill of Rights that means what the majority wants it to mean is worthless. 
 -- U.S. Supreme Court Justice Antonin Scalia

[toc] | [prev] | [next] | [standalone]


#194498 — Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)]

FromNicolas George <george@nsup.org>
Date2018-04-04 20:10 +0200
SubjectRe: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re: utf)]
Message-ID<vAV18-1vZ-9@gated-at.bofh.it>
In reply to#194496

[Multipart message — attachments visible in raw view] — view raw

Don Armstrong (2018-04-04):
> There are definitely better formats than Maildir, like Dovecot's
> multi-dbox.[1]
> 
> These issues are why almost everyone who uses Maildir just uses it as
> the backing message store and uses the index on top to do avoid ever
> reading all of the messages in the Maildir.

Glad to read this. There are too many maildir zealots out there.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194501 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 20:40 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVu9-1Hd-1@gated-at.bofh.it>
In reply to#194490
On Wednesday, April 04, 2018 01:36:15 PM Don Armstrong wrote:
> On Tue, 03 Apr 2018, rhkramer@gmail.com wrote:
> > I am building (have built several iterations) of a free format
> > database to work something like askSam. It is a mashup of several
> > applications, things like recol, kmail, nail, kate and the data is
> > stored in mbox formatted files.
> > 
> > Each record is treated as an email.
> 
> You should consider looking at using Maildir with notmuch and using
> things which integrate notmuch.[1]
> 
> > Most likely this would be only a temporary addition, and I would need
> > to do things like make sure that one byte will be unique in the file.
> > It sounds like there are at least a few candidates.
> 
> Maildir is the solution to this. While you *can* handle mbox and do all
> the escape rules properly (From to >From and back), it's a pain. Let
> your filesystem handle it for you.
> 
> [I'm speaking from experience; I currently maintain debbugs, which
> basically stores everything in a custom format mbox. This inevitably
> makes things slow, as you have to search through the mbox linearly to
> find any message in the mbox unless you also write indexes for the
> mailbox.]
> 
> 1: Notmuch itself uses xapian to do the heavy lifting.

I'll probably look into notmuch, just for kicks.

I've considered maildir--it meets some of my requirements (that is, to make 
something close to an askSam workalike), but one drawback is that it is 
essentially one email (i.e., my "record").  One of the desirable features of 
askSam is that you did not have to create a new file to add a new note / 
record, you just start typing in an existing open record and then, as time or 
other constraints allow, you can add more "tags" or a record separator.  (It's 
been so long since I've used askSam I actually forget what had to be done (f 
anything) to separate a new record from the previous record).

askSam basically stores all it's records in one file, although it is (of 
course) possible to separate them.

[toc] | [prev] | [next] | [standalone]


#194503 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromDon Armstrong <don@debian.org>
Date2018-04-04 20:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVDP-1KU-9@gated-at.bofh.it>
In reply to#194501
On Wed, 04 Apr 2018, rhkramer@gmail.com wrote:
> I've considered maildir--it meets some of my requirements (that is, to
> make something close to an askSam workalike), but one drawback is that
> it is essentially one email (i.e., my "record"). One of the desirable
> features of askSam is that you did not have to create a new file to
> add a new note / record, you just start typing in an existing open
> record and then, as time or other constraints allow, you can add more
> "tags" or a record separator. (It's been so long since I've used
> askSam I actually forget what had to be done (f anything) to separate
> a new record from the previous record).

You might want to consider looking at org-mode too.[1] There are
even integrations for notmuch+mutt+Maildir there.

1: https://www.youtube.com/watch?v=oJTwQvgfgMM

-- 
Don Armstrong                      https://www.donarmstrong.com

I really wanted to talk to her.
I just couldn't find an algorithm that fit.
 -- Peter Watts _Blindsight_ p294

[toc] | [prev] | [next] | [standalone]


#194522 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-05 14:40 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vBclj-4wA-1@gated-at.bofh.it>
In reply to#194503
On Wednesday, April 04, 2018 02:45:49 PM Don Armstrong wrote:
> On Wed, 04 Apr 2018, rhkramer@gmail.com wrote:
> > I've considered maildir--it meets some of my requirements (that is, to
> > make something close to an askSam workalike), but one drawback is that
> > it is essentially one email (i.e., my "record"). One of the desirable
> > features of askSam is that you did not have to create a new file to
> > add a new note / record, you just start typing in an existing open
> > record and then, as time or other constraints allow, you can add more
> > "tags" or a record separator. (It's been so long since I've used
> > askSam I actually forget what had to be done (f anything) to separate
> > a new record from the previous record).
> 
> You might want to consider looking at org-mode too.[1] There are
> even integrations for notmuch+mutt+Maildir there.
> 
> 1: https://www.youtube.com/watch?v=oJTwQvgfgMM

Thanks for the link / pointer to org-mode.  From the little I've looked at, it 
has several similarities to what I'm doing--for example, the collapsible 
outlining / folding, and the use of a mark-up language (I use the TWiki markup 
language with a few extensions / modifications--oh, and that reminds me (if 
anyone is keeping track of my mashup)--TWiki / Foswiki are other programs that 
are part of the mashup--records in my mashup are maintained in a form that 
would work as a page (they use a different word--ohh, topic) of a TWiki / 
Foswiki, and someday I'd like to have an automatic import / export facility--
in my home mashup.

With that automatic import / export facility, I could designate certain 
records to be "public" (or something like that) which would mean: (1) if they 
were not on my public TWiki / Foswiki, they would be automatically exported 
(when I was connected to the Internet and designated that it was appropriate 
to "sync" my home storage with the public wiki), and (2) any public wiki pages 
that might have been modified (by others) would be copied to the mashup as 
backup.  (Also, if a mashup page was designated as "public", after I made 
changes on the mashup they would (at an appropriate time) be uploaded to the 
public wiki.

I could discuss my hate-hate relationship with Emacs--I tried a few times to 
learn Emacs, and had difficulty for a variety of reasons.   Part of it was my 
hate-hate relationship with Lisp, another part was that Emacs (at the time) 
seemed much less GUI / mouse friendly than the editors I had grown used to in 
my DOS / Windows days (even though I learned (and liked) a lot of shortcut 
keys--I guess maybe an early encounter (with shortcut keys) was with Wordstar, 
and then I used a shareware editor (that I paid for--high praise indeed from 
me) that used the same set of shortcut keys--I can't remember the name of that 
editor atm.

Anyway, thanks for prompting me to reminisce.  

I know that at the times I looked at Emacs (and Xemacs) they had outline mode, 
I'm not sure I recall org mode.  I expect I will spend a little more time 
looking further into org mode, although I think my mashup has or will have the 
features I've seen there so far.

(Aside: My mashup uses kate as the editor, but because of some long standing 
bugs in kate (which, finally after several years were supposedly fixed a few 
years ago) I had to add closing markup to the opening markup of TWiki / 
Foswiki.  In the meantime, I decided I wanted to switch to using any Scintilla 
based editor (there are a lot) as the editor, but so far, have not got a 
folder / highlighter written for Scintilla.  (And, the one time I tried to 
test the fixed Kate, it didn't seem to work as promised, and I just left it as 
is, still requiring the ending markup.))

Thanks again for the response! 

[toc] | [prev] | [next] | [standalone]


#194516 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromdeloptes <deloptes@gmail.com>
Date2018-04-05 00:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAYV3-4f0-3@gated-at.bofh.it>
In reply to#194501
rhkramer@gmail.com wrote:

> I'll probably look into notmuch, just for kicks.
> 
> I've considered maildir--it meets some of my requirements (that is, to
> make something close to an askSam workalike), but one drawback is that it
> is essentially one email (i.e., my "record").  One of the desirable
> features of askSam is that you did not have to create a new file to add a
> new note / record, you just start typing in an existing open record and
> then, as time or other constraints allow, you can add more "tags" or a
> record separator.  (It's been so long since I've used askSam I actually
> forget what had to be done (f anything) to separate a new record from the
> previous record).
> 
> askSam basically stores all it's records in one file, although it is (of
> course) possible to separate them.

I still don't understand why not use XML. If file is not getting too big
(which is also a problem with mbox). You may need to adapt your
applications, but it won't be more effort then making shit out of shit -
sorry for my language.

regards

[toc] | [prev] | [next] | [standalone]


#194509 — Re: Invalid UTF-8 byte? (was: Re: utf)

Fromrhkramer@gmail.com
Date2018-04-04 21:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAWgx-2ha-13@gated-at.bofh.it>
In reply to#194490
On Wednesday, April 04, 2018 01:36:15 PM Don Armstrong wrote:
> On Tue, 03 Apr 2018, rhkramer@gmail.com wrote:
> > I am building (have built several iterations) of a free format
> > database to work something like askSam. It is a mashup of several
> > applications, things like recol, kmail, nail, kate and the data is
> > stored in mbox formatted files.
> > 
> > Each record is treated as an email.
> 
> You should consider looking at using Maildir with notmuch and using
> things which integrate notmuch.[1]

Ahh, OK, notmuch looks like it could be an alternative to recol as part of my 
mashup.  Thanks, I'll probably dig deeper as time goes on.

[toc] | [prev] | [next] | [standalone]


#194472 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromHenrique de Moraes Holschuh <hmh@debian.org>
Date2018-04-04 13:20 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAOCl-5AM-7@gated-at.bofh.it>
In reply to#194445
On Tue, 03 Apr 2018, Michael Lange wrote:
> I believe (please anyone correct me if I am wrong) that "text" files
> won't contain any null byte; many text editors even refuse to open such a

Depends on the encoding.  For ASCII, ISO-8859-* and UTF-8 (and any other
modern encoding AFAIK, other than modified UTF-8), any zero bytes map
one-to-one to the NUL character/code point.  I don't recall how it is on
other common encodings of the 80's and 90's, though.

Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use
bytes with the value of zero when encoding characters, so NUL is encoded
by a different sequence, and you can safely use a byte with the value of
zero for some out-of-band control (like zero-terminated strings that can
contain NULs, etc) -- note that NUL is a character, and it might be
represented by a sequence of bytes that has nothing to do with zeroes on
a particular encoding...

(in fact, C strings are *zero-terminated*, not NUL-terminated, but most
of the time this is irrelevant :p).

Also, a text file MAY contain NULs (the character), it is just
considered bad practice (nowadays?).  Don't assume you won't see any.
For example, received e-mail is *more* likely to have NULs in it than
normal text due to the quality of some mail agents out there.  I recall
postfix would reject a *lot* of crap when we configured it to refuse to
accept NULs outside of 8-bit bodies, because Cyrus-IMAPd *refuses* any
such crap, and we wanted it bounced as early as possible.

(note that NULs are forbidden in MIME-compliant email text and ESMTP,
unless encoded or guarded by a 8-bit transfer area of known size, so
there you have it: NULs in one text format that actually forbids them!).

> Probably it is the same with some other control characters like 04 (End
> of Transmission). When I look at https://en.wikipedia.org/wiki/ASCII
> it seems like 1C (File Separator) or 1E (Record Separator) might be 
> appropriate choices for you. I'm no expert on this, though.

Well, ASCII control characters were inherited by ISO-8859-* and Unicode,
so yes, you can use them.  But so could the data file.  It would be
perfectly ok for a text data file to use the record separator control
characters to delimit records in a table, for example...

Here's a good definition of them (follow the hyperlinks for the
definition of each control character):
https://en.wikipedia.org/wiki/Basic_Latin_(Unicode_block)


Here is also a proper solution: use modified UTF-8 (which encodes NUL so
that zero bytes are *never* present in the stream): encode every input
format to modified UTF-8, then add the zero-byte separators you want.

You'll have to normalize the input data set into known charset/encodings
and then recode them to modified UTF-8, of course.  You can't blindly
call any random data "UTF-8" (let alone modified UTF-8) and expect
things not to break horribly.

-- 
  Henrique Holschuh

[toc] | [prev] | [next] | [standalone]


#194481 — Re: Invalid UTF-8 byte? (was: Re: utf)

From<tomas@tuxteam.de>
Date2018-04-04 16:10 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vARgR-7uD-3@gated-at.bofh.it>
In reply to#194472
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Wed, Apr 04, 2018 at 08:18:23AM -0300, Henrique de Moraes Holschuh wrote:
> On Tue, 03 Apr 2018, Michael Lange wrote:
> > I believe (please anyone correct me if I am wrong) that "text" files
> > won't contain any null byte; many text editors even refuse to open such a
> 
> Depends on the encoding.  For ASCII, ISO-8859-* and UTF-8 (and any other
> modern encoding AFAIK, other than modified UTF-8), any zero bytes map
> one-to-one to the NUL character/code point.  I don't recall how it is on
> other common encodings of the 80's and 90's, though.

Try UTF-16, what Microsoft (and a couple of years ago Apple) love to
call "Unicode": in more "Western" contexts every second byte is NULL!

> Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use
> bytes with the value of zero when encoding characters, so NUL is encoded
> by a different sequence, and you can safely use a byte with the value of
> zero for some out-of-band control [...]

Yes, the problem is that someone else before you could have been doing
exactly that.

I'd guard against that. It's not exactly difficult, the traditional
"escape" mechanism (aka character stuffing) does it pretty well...

Cheers
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrE3JcACgkQBcgs9XrR2kYW7ACeMG0SQB23RSySoeSJBItB+Eji
QEgAnipwAcoVJuzynJVBO1CR2rrLeuFs
=xhja
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194502 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromHenrique de Moraes Holschuh <hmh@debian.org>
Date2018-04-04 20:50 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVDP-1KU-7@gated-at.bofh.it>
In reply to#194481
On Wed, 04 Apr 2018, tomas@tuxteam.de wrote:
> On Wed, Apr 04, 2018 at 08:18:23AM -0300, Henrique de Moraes Holschuh wrote:
> > On Tue, 03 Apr 2018, Michael Lange wrote:
> > > I believe (please anyone correct me if I am wrong) that "text" files
> > > won't contain any null byte; many text editors even refuse to open such a
> > 
> > Depends on the encoding.  For ASCII, ISO-8859-* and UTF-8 (and any other
> > modern encoding AFAIK, other than modified UTF-8), any zero bytes map
> > one-to-one to the NUL character/code point.  I don't recall how it is on
> > other common encodings of the 80's and 90's, though.
> 
> Try UTF-16, what Microsoft (and a couple of years ago Apple) love to
> call "Unicode": in more "Western" contexts every second byte is NULL!

Ah, yes.  I forgot about them, indeed.  UTF-16BE and UTF-16LE will have
zero bytes in the resulting byte stream.  And I suppose one could call
them "modern encodings", even if they are horrifying to use when
compared to UTF-8 (UTF-16 has byte-order issues) or UTF-32 (UTF-16 has
surrogate pairs).

> > Some even-more-modern encodings (modified UTF-8 :p) simply do NOT use
> > bytes with the value of zero when encoding characters, so NUL is encoded
> > by a different sequence, and you can safely use a byte with the value of
> > zero for some out-of-band control [...]
> 
> Yes, the problem is that someone else before you could have been doing
> exactly that.

You can modified-UTF-8 bit-packing to encode anything, and the result
will be zero-free (and it will restore the zeroes when decoded).  The
price is a size increase (it is a variant of UTF-8 that uses two bytes
to encode NUL, which would take just one byte in normal UTF-8).  There
are much better bit packing schemes if you just need to escape zeroes
;-)

That said, it is always safe to break valid "modified UTF-8" into
records using zeroes, as long as you don't expect the result to be valid
UTF-8 (it isn't valid UTF-8 because NULs will be encoded using a
non-minimal byte sequence that *will* decode to a zero even if it is
invalid) or valid modified UTF-8 (it isn't valid modified UTF-8 because
0 is not valid as an encoding for NUL in modified UTF-8).  But a lax
UTF-8 or modified UTF-8 *would* parse "modified UTF-8 with zero as
record separators" and reconstruct the unicode text properly (but it
would read the record separators as NULs, so you'd get extra NULs in the
resulting text).

That, of course, assumes you have unicode text as the input (encoding
doesn't matter, as long as you know it), and recode it to modified UTF-8
before you add the zeroes as end-of-record marks.  This is not about
bit-packing generic binary data.

> I'd guard against that. It's not exactly difficult, the traditional
> "escape" mechanism (aka character stuffing) does it pretty well...

Yes, any bitstuffing/escape-based wrapping would do.

-- 
  Henrique Holschuh

[toc] | [prev] | [next] | [standalone]


#194506 — Re: Invalid UTF-8 byte? (was: Re: utf)

From<tomas@tuxteam.de>
Date2018-04-04 21:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAWgx-2ha-3@gated-at.bofh.it>
In reply to#194502
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Wed, Apr 04, 2018 at 03:44:23PM -0300, Henrique de Moraes Holschuh wrote:

[...]

> That said, it is always safe to break valid "modified UTF-8" into
> records using zeroes, as long as you don't expect the result to be valid
> UTF-8 (it isn't valid UTF-8 because NULs will be encoded using a
> non-minimal byte sequence that *will* decode to a zero even if it is
> invalid) or valid modified UTF-8 (it isn't valid modified UTF-8 because
> 0 is not valid as an encoding for NUL in modified UTF-8).  But a lax
> UTF-8 or modified UTF-8 *would* parse "modified UTF-8 with zero as
> record separators" and reconstruct the unicode text properly (but it
> would read the record separators as NULs, so you'd get extra NULs in the
> resulting text).

You are a nasty guy, aren't you ;-)

Pretty cunning...

Cheers
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrFJngACgkQBcgs9XrR2kZqLgCdEuap+rqSU6HCrXpkL6XHl3Az
lRUAnjwGhiMNNlY+SXwIxpd/kfnvst1z
=kHBa
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194513 — Re: Invalid UTF-8 byte?

FromBen Caradoc-Davies <ben@transient.nz>
Date2018-04-04 23:50 +0200
SubjectRe: Invalid UTF-8 byte?
Message-ID<vAYs1-3MP-3@gated-at.bofh.it>
In reply to#194481
On 05/04/18 02:09, tomas@tuxteam.de wrote:
> Try UTF-16, what Microsoft (and a couple of years ago Apple) love to
> call "Unicode": in more "Western" contexts every second byte is NULL!

The Java platform uses UTF-16 internally:

"The char data type (and therefore the value that a Character object 
encapsulates) are based on the original Unicode specification, which 
defined characters as fixed-width 16-bit entities."
https://docs.oracle.com/javase/8/docs/api/java/lang/Character.html

Kind regards,

-- 
Ben Caradoc-Davies <ben@transient.nz>
Director
Transient Software Limited <https://transient.nz/>
New Zealand

[toc] | [prev] | [next] | [standalone]


#194517 — Re: Invalid UTF-8 byte?

FromMichael Stone <mstone@debian.org>
Date2018-04-05 01:00 +0200
SubjectRe: Invalid UTF-8 byte?
Message-ID<vAZxL-4ze-1@gated-at.bofh.it>
In reply to#194513
On Thu, Apr 05, 2018 at 09:42:19AM +1200, Ben Caradoc-Davies wrote:
>On 05/04/18 02:09, tomas@tuxteam.de wrote:
>>Try UTF-16, what Microsoft (and a couple of years ago Apple) love to
>>call "Unicode": in more "Western" contexts every second byte is NULL!
>
>The Java platform uses UTF-16 internally:

Yes, many people thought UCS-2 was the answer back when 16 bits was 
enough for anybody.

Mike Stone

[toc] | [prev] | [next] | [standalone]


#194505 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromJonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM>
Date2018-04-04 21:00 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVNB-1On-15@gated-at.bofh.it>
In reply to#194472
Henrique de Moraes Holschuh:

> Also, a text file MAY contain NULs (the character), it is just 
> considered bad practice (nowadays?). Don't assume you won't see any. 
> For example, received e-mail is *more* likely to have NULs in it than 
> normal text due to the quality of some mail agents out there.
>
I suspect not as likely as anything that was in the process of being 
appended to on a not-fully-journalling filesystem when a dirty shutdown 
happens.  (-:

* https://askubuntu.com/questions/356981/

Or anything that "rotates" output files by truncating them and pulls the 
rug out from underneath an old-style simplistic indefinitely-running 
text output writer.

* http://jdebp.eu./FGA/do-not-use-logrotate.html#Background

[toc] | [prev] | [next] | [standalone]


#194500 — Re: Invalid UTF-8 byte? (was: Re: utf)

FromJonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM>
Date2018-04-04 20:30 +0200
SubjectRe: Invalid UTF-8 byte? (was: Re: utf)
Message-ID<vAVku-1DA-13@gated-at.bofh.it>
In reply to#194391
rhkramer:

> The reason I wanted such a byte was to use it as a record separator in 
> a set of text files (that I use as an askSam "workalike" (or 
> "worksimilar") so that I could use msort (which depends on a 1 byte 
> record separator to --separate the records ;-) while sorting. Some of 
> the files already include UTF-8, and, in the future, I anticpate all 
> will be in UTFF-8.
>
Note that ISO 646, hence ISO 8859, hence ISO 10646, has had a 
single-byte Record Separator character since the 1960s.  (-:

[toc] | [prev] | [next] | [standalone]


Page 3 of 6 — ← Prev page 1 2 [3] 4 5 6  Next page →

Back to top | Article view | linux.debian.user


csiph-web