Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #194373 > unrolled thread

utf

Started bymess-mate <mess-mate@gmx.com>
First post2018-04-01 16:10 +0200
Last post2018-04-04 14:20 +0200
Articles 20 on this page of 101 — 21 participants

Back to article view | Back to linux.debian.user


Contents

  utf mess-mate <mess-mate@gmx.com> - 2018-04-01 16:10 +0200
    Re: utf Curt <curty@free.fr> - 2018-04-01 17:50 +0200
    Re: utf Ionel Mugurel Ciobîcă <I.M.Ciobica@upcmail.nl> - 2018-04-01 18:30 +0200
    Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-01 18:40 +0200
    Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-01 22:10 +0200
      Re: utf Cindy-Sue Causey <butterflybytes@gmail.com> - 2018-04-02 00:50 +0200
        Re: utf Curt <curty@free.fr> - 2018-04-02 09:50 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-02 11:40 +0200
          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-02 15:30 +0200
      Re: utf Andre Majorel <aym-naibed@teaser.fr> - 2018-04-02 09:50 +0200
        Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 14:40 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 15:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 15:10 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 19:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-02 20:20 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-02 20:50 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) tomas@tuxteam.de - 2018-04-02 21:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-02 21:10 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 00:50 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 13:50 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-03 14:00 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-03 14:30 +0200
              Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:10 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 14:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 21:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 21:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Michael Lange <klappnase@freenet.de> - 2018-04-03 22:10 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 14:40 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 03:20 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 13:30 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:10 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 14:10 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 14:50 +0200
                              Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 15:00 +0200
                                Re: Invalid UTF-8 byte? (was: Re: utf) Andre Majorel <aym-naibed@teaser.fr> - 2018-04-04 16:20 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 16:30 +0200
                                    Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                                  Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 18:00 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:50 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 14:30 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 15:10 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 19:40 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) Nicolas George <george@nsup.org> - 2018-04-04 19:50 +0200
                          mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was: Re:  utf)] Don Armstrong <don@debian.org> - 2018-04-04 20:00 +0200
                            Re: mbox vs maildir vs better formats [Re: Invalid UTF-8 byte? (was:  Re: utf)] Nicolas George <george@nsup.org> - 2018-04-04 20:10 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 20:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) Don Armstrong <don@debian.org> - 2018-04-04 20:50 +0200
                            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-05 14:40 +0200
                          Re: Invalid UTF-8 byte? (was: Re: utf) deloptes <deloptes@gmail.com> - 2018-04-05 00:20 +0200
                        Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
                Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 13:20 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 16:10 +0200
                    Re: Invalid UTF-8 byte? (was: Re: utf) Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 20:50 +0200
                      Re: Invalid UTF-8 byte? (was: Re: utf) <tomas@tuxteam.de> - 2018-04-04 21:30 +0200
                    Re: Invalid UTF-8 byte? Ben Caradoc-Davies <ben@transient.nz> - 2018-04-04 23:50 +0200
                      Re: Invalid UTF-8 byte? Michael Stone <mstone@debian.org> - 2018-04-05 01:00 +0200
                  Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 21:00 +0200
          Re: Invalid UTF-8 byte? (was: Re: utf) Jonathan de Boyne Pollard <J.deBoynePollard-newsgroups@NTLWorld.COM> - 2018-04-04 20:30 +0200
            Re: Invalid UTF-8 byte? (was: Re: utf) rhkramer@gmail.com - 2018-04-04 21:30 +0200
        Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-02 23:40 +0200
        Re: utf Darac Marjal <mailinglist@darac.org.uk> - 2018-04-03 11:00 +0200
          Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-03 11:20 +0200
            Re: utf <tomas@tuxteam.de> - 2018-04-03 11:30 +0200
          Re: utf Nicolas George <george@nsup.org> - 2018-04-03 12:20 +0200
          Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-03 22:50 +0200
            Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:00 +0200
              Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-03 23:00 +0200
                Re: utf Nicolas George <george@nsup.org> - 2018-04-03 23:30 +0200
                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:00 +0200
                    Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:10 +0200
                      Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:30 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 19:40 +0200
                        Re: utf Nicolas George <george@nsup.org> - 2018-04-04 19:40 +0200
                          Re: utf Greg Wooledge <wooledg@eeg.ccf.org> - 2018-04-04 19:50 +0200
                            Re: utf Nicolas George <george@nsup.org> - 2018-04-04 20:00 +0200
                              Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:20 +0200
                              Re: utf Richard Hector <richard@walnut.gen.nz> - 2018-04-05 02:30 +0200
                                Re: utf Nicolas George <george@nsup.org> - 2018-04-05 14:10 +0200
                            Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 00:00 +0200
                    Re: utf rhkramer@gmail.com - 2018-04-04 20:30 +0200
                      Re: utf Joel Roth <joelz@pobox.com> - 2018-04-04 21:30 +0200
                      Re: utf deloptes <deloptes@gmail.com> - 2018-04-04 23:40 +0200
                        Re: utf <tomas@tuxteam.de> - 2018-04-05 08:30 +0200
                          Re: utf rhkramer@gmail.com - 2018-04-05 14:50 +0200
                            Re: utf <tomas@tuxteam.de> - 2018-04-05 15:00 +0200
                              Re: utf Nicolas George <george@nsup.org> - 2018-04-05 15:00 +0200
                                Re: utf tomas@tuxteam.de - 2018-04-05 15:10 +0200
                                  Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 21:40 +0200
                                  Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 23:30 +0200
                                    Re: utf deloptes <deloptes@gmail.com> - 2018-04-05 23:40 +0200
                                      Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-06 03:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:00 +0200
                                      Re: utf Ben Caradoc-Davies <ben@transient.nz> - 2018-04-06 04:20 +0200
                                        Re: utf <tomas@tuxteam.de> - 2018-04-06 09:10 +0200
                              Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-05 19:00 +0200
                            Re: utf rhkramer@gmail.com - 2018-04-05 20:00 +0200
                Re: utf Stefan Monnier <monnier@iro.umontreal.ca> - 2018-04-03 23:40 +0200
          Re: utf Henrique de Moraes Holschuh <hmh@debian.org> - 2018-04-04 14:20 +0200

Page 5 of 6 — ← Prev page 1 2 3 4 [5] 6  Next page →


#194521

FromNicolas George <george@nsup.org>
Date2018-04-05 14:10 +0200
Message-ID<vBbSi-4n5-11@gated-at.bofh.it>
In reply to#194518

[Multipart message — attachments visible in raw view] — view raw

Richard Hector (2018-04-05):
> >> What if the question is "Find all the English words that have an E
> >> in the 5th position and a U in the 7th"?
> > Yes, what? Who would ever ask such a question? What is the point of such
> > a question?
> Solving a crossword puzzle?

This is a good example, thanks, but I think it goes eventually for my
point.

Words in a crossword puzzle are not generic text. You cannot do them in
Chinese, for example, AFAIK. Even in English, if you just have a list of
words, you cannot use it directly for crossword puzzles, you first need
to filter it to remove the few diacritics that have seeped from other
languages, because the "é" in "précising", for example, can be crossed
with a normal "e" in any word.

Starting from a list in "pure" text, you need to build the data
structure that is convenient for the particular problem. And indeed, for
crossword puzzles, accessing the n-th letter of a de-diacriticized word
is a necessary operation. But it is not accessing the n-th letter of a
generic string.

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194514

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2018-04-05 00:00 +0200
Message-ID<vAYBH-3RC-1@gated-at.bofh.it>
In reply to#194494
> You just seem to have Decided, for reasons known only to you, that
> The Character Length Of A String Is Not Useful.  Despite literally
> decades of programs that have used strlen() in various ways.

strlen was mostly used in a context where char-length = byte-length =
display-width.  Most of those calls to strlen have nothing to do with
char-length but are more interested in display-width or byte-length.

In the context of Unicode, using utf-8 doesn't make byte-length any
harder than with ASCII.  And in the context of Unicode, display-width
is a lot more complex than strlen regardless of which encoding you use
because any given Unicode char can have a display-width of 0, 1, or
2 (even if you disregard proportional fonts and other fancy rendering
tricks).  So utf-8 doesn't make the computation of display-width any
more complex than utf-32.

> What if the question is "Find all the English words that have an E
> in the 5th position and a U in the 7th"?

That can be answered just as easily and efficiently from a utf-8
representation of the string as from a utf-32 representation.


        Stefan

[toc] | [prev] | [next] | [standalone]


#194499

Fromrhkramer@gmail.com
Date2018-04-04 20:30 +0200
Message-ID<vAVkt-1DA-9@gated-at.bofh.it>
In reply to#194486
On Wednesday, April 04, 2018 12:58:57 PM deloptes wrote:
> And regarding the mbox thing, well mbox was depreciated for many reasons. I
> guess if it was that good it wouldn't be depreciated.

Oh, I wasn't aware that mbox was deprecated--can you shed more light on that.  
AFAIK, it is not defined in an RFC and is used by quite a few email programs.

[toc] | [prev] | [next] | [standalone]


#194507

FromJoel Roth <joelz@pobox.com>
Date2018-04-04 21:30 +0200
Message-ID<vAWgx-2ha-5@gated-at.bofh.it>
In reply to#194499
On Wed, Apr 04, 2018 at 02:20:17PM -0400, rhkramer@gmail.com wrote:
> On Wednesday, April 04, 2018 12:58:57 PM deloptes wrote:
> > And regarding the mbox thing, well mbox was depreciated for many reasons. I
> > guess if it was that good it wouldn't be depreciated.

> Oh, I wasn't aware that mbox was deprecated--can you shed more light on that.  
> AFAIK, it is not defined in an RFC and is used by quite a few email programs.
 
Not exactly deprecated, but it's considered a less reliable storage format, 
because of potential problems. 

https://en.wikipedia.org/wiki/Mbox

I converted to Maildir for better compatibility with the mu
indexing programs (package maildir-utils). 

cheers,
 

-- 
Joel Roth
  

[toc] | [prev] | [next] | [standalone]


#194512

Fromdeloptes <deloptes@gmail.com>
Date2018-04-04 23:40 +0200
Message-ID<vAYim-3HY-5@gated-at.bofh.it>
In reply to#194499
rhkramer@gmail.com wrote:

> Oh, I wasn't aware that mbox was deprecated--can you shed more light on
> that. AFAIK, it is not defined in an RFC and is used by quite a few email
> programs.

yes but Maildir format was introduced for couple of reasons (as well as
other formats). I wouldn't store my mail in mbox anyway. For local
system/user mails as a simple default storage perhaps yes - it might be OK,
but for public mail, where you have 1000+ mails and perhaps multiple
interfaces ... no chance.

https://wiki2.dovecot.org/MailboxFormat

https://wiki2.dovecot.org/MailboxFormat/mbox
https://wiki2.dovecot.org/MailboxFormat/Maildir

I have worked on cloud mail solution using dovecot with mysql backend for
18mil customers. Another company was using dbmail with mysql with very good
results. 
But this goes somehow off topic in regards of original UTF

The only advantage I see with mbox is that it is really simple.

regards

[toc] | [prev] | [next] | [standalone]


#194519

From<tomas@tuxteam.de>
Date2018-04-05 08:30 +0200
Message-ID<vB6zf-Sw-1@gated-at.bofh.it>
In reply to#194512
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Wed, Apr 04, 2018 at 11:33:13PM +0200, deloptes wrote:

[...]
> other formats). I wouldn't store my mail in mbox anyway. For local
> system/user mails as a simple default storage perhaps yes - it might be OK,
> but for public mail, where you have 1000+ mails and perhaps multiple
> interfaces ... no chance.

Increase that by 2-3 orders of magnitude and I'd agree (FWIW, I'm working
off an mbox with ~15K mails. Works fine).

Cheers
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrFwXkACgkQBcgs9XrR2kaAsACfa1vlRYeQJDH/mFkn8NWViU5J
JKMAn1AvJGRmB1tGHKxWiFT15z9NqnQy
=K1gZ
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194523

Fromrhkramer@gmail.com
Date2018-04-05 14:50 +0200
Message-ID<vBcuZ-4Bk-5@gated-at.bofh.it>
In reply to#194519
On Thursday, April 05, 2018 02:26:01 AM tomas@tuxteam.de wrote:
> On Wed, Apr 04, 2018 at 11:33:13PM +0200, deloptes wrote:
> 
> [...]
> 
> > other formats). I wouldn't store my mail in mbox anyway. For local
> > system/user mails as a simple default storage perhaps yes - it might be
> > OK, but for public mail, where you have 1000+ mails and perhaps multiple
> > interfaces ... no chance.
> 
> Increase that by 2-3 orders of magnitude and I'd agree (FWIW, I'm working
> off an mbox with ~15K mails. Works fine).

I'm laughing (at myself)--I just checked my mail directory, I have at least 4 
mbox files (and then I stopped looking) greater than 175 MB.  One of them, my 
inbox, is 2.8 GB--no problems.

I do need to compact my inbox, and I did, but maybe the actual file isn't 
changed until I quit kmail--I'll try that later.

[toc] | [prev] | [next] | [standalone]


#194524

From<tomas@tuxteam.de>
Date2018-04-05 15:00 +0200
Message-ID<vBcEF-4Fa-1@gated-at.bofh.it>
In reply to#194523
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Thu, Apr 05, 2018 at 08:42:39AM -0400, rhkramer@gmail.com wrote:
> On Thursday, April 05, 2018 02:26:01 AM tomas@tuxteam.de wrote:

[...]

> > Increase that by 2-3 orders of magnitude [...]

> I'm laughing (at myself)--I just checked my mail directory, I have at least 4 
> mbox files (and then I stopped looking) greater than 175 MB.  One of them, my 
> inbox, is 2.8 GB--no problems.
> 
> I do need to compact my inbox, and I did, but maybe the actual file isn't 
> changed until I quit kmail--I'll try that later.

Actually people saying mbox is a bad database are in principle right
(I never liked maildir either: dumping metadata into file names seemed
to me a bit disgusting too, but I disgress). But there's something
special about mail databases which eases that a bit: records (i.e.
mails) are *mostly* immutable (save for some metadata), so cleverly
written libs (mutt, dovecot) can be suprisingly good despite all.

But then when I see people proposing XML as structured data
representation, I suddenly grow very sad...

Cheers
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrGHKQACgkQBcgs9XrR2kbeJQCfS1sQhck1kmoysI4bBR2gRUtn
+LMAnAjbK9bFWygOtoA1OwmS9TLqulHU
=BvB3
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194525

FromNicolas George <george@nsup.org>
Date2018-04-05 15:00 +0200
Message-ID<vBcEG-4Fa-17@gated-at.bofh.it>
In reply to#194524

[Multipart message — attachments visible in raw view] — view raw

tomas@tuxteam.de (2018-04-05):
> But then when I see people proposing XML as structured data
> representation, I suddenly grow very sad...

Isn't it?

Regards,

-- 
  Nicolas George

[toc] | [prev] | [next] | [standalone]


#194526

Fromtomas@tuxteam.de
Date2018-04-05 15:10 +0200
Message-ID<vBcOm-4XV-21@gated-at.bofh.it>
In reply to#194525
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Thu, Apr 05, 2018 at 02:56:47PM +0200, Nicolas George wrote:
> tomas@tuxteam.de (2018-04-05):
> > But then when I see people proposing XML as structured data
> > representation, I suddenly grow very sad...
> 
> Isn't it?

For the last 6 years I've seen it done a lot, and yes, I speak
of painful experience.

XML is a baroque, but passable document serialization language.

But (mis-)using it as a data serialization language must be one
of the worst (and ugliest) misunderstandings IT has had the last
20 years.

Cheers
- -- t
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrGHfEACgkQBcgs9XrR2kZLOgCdHbxd3ACh+/5BSLAPtijkkZra
UaUAn2nkEXU74FXWntUunIwzAFAL6VwJ
=FkWM
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194530

Fromdeloptes <deloptes@gmail.com>
Date2018-04-05 21:40 +0200
Message-ID<vBiTM-gQ-7@gated-at.bofh.it>
In reply to#194526
tomas@tuxteam.de wrote:

> On Thu, Apr 05, 2018 at 02:56:47PM +0200, Nicolas George wrote:
>> tomas@tuxteam.de (2018-04-05):
>> > But then when I see people proposing XML as structured data
>> > representation, I suddenly grow very sad...
>> 
>> Isn't it?
> 
> For the last 6 years I've seen it done a lot, and yes, I speak
> of painful experience.
> 
> XML is a baroque, but passable document serialization language.
> 
> But (mis-)using it as a data serialization language must be one
> of the worst (and ugliest) misunderstandings IT has had the last
> 20 years.
> 

Don't know about your experience - in the past 10+y I worked a lot with XML.
It might be not suitable for the specific case (I am not sure I understood
correctly what OPs goal is), but XML does not care what kind of data you
will embed as far as it is typed and as a machine exchange language it is
exactly what it is intended to be.
In terms of serialization it also depends - but the point was to split text
and process it.

[toc] | [prev] | [next] | [standalone]


#194531

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2018-04-05 23:30 +0200
Message-ID<vBkCd-1rd-11@gated-at.bofh.it>
In reply to#194526
> But (mis-)using it as a data serialization language must be one
> of the worst (and ugliest) misunderstandings IT has had the last
> 20 years.

UUIC that's partly why it's finally losing popularity and being replaced
with json for that use.  I'm not familiar enough with json to know if
it's really a good replacement, but it does look like an improvement.


        Stefan

[toc] | [prev] | [next] | [standalone]


#194532

Fromdeloptes <deloptes@gmail.com>
Date2018-04-05 23:40 +0200
Message-ID<vBkLT-1uD-1@gated-at.bofh.it>
In reply to#194531
Stefan Monnier wrote:

> UUIC that's partly why it's finally losing popularity and being replaced
> with json for that use.  I'm not familiar enough with json to know if
> it's really a good replacement, but it does look like an improvement.

that is simply not true. JSON might be more simple, and might be a
replacement to XML in some cases, but XML has the features of SGML and per
definition
XML is not a single Markup Language. It is a metalanguage to let users
design their own markup language.

According wikipedia JSON (JavaScript Object Notation) "is a very common data
format used for asynchronous browser–server communication".

regards

[toc] | [prev] | [next] | [standalone]


#194538

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2018-04-06 03:20 +0200
Message-ID<vBocN-46c-1@gated-at.bofh.it>
In reply to#194532
>> UUIC that's partly why it's finally losing popularity and being replaced
>> with json for that use.  I'm not familiar enough with json to know if
>> it's really a good replacement, but it does look like an improvement.
> that is simply not true.

Did you read the text to which I was responding?  Because your reply
does not seem to contradict mine: I was talking about uses of XML for
"data serialization".


        Stefan

[toc] | [prev] | [next] | [standalone]


#194540

From<tomas@tuxteam.de>
Date2018-04-06 09:00 +0200
Message-ID<vBtvQ-7xO-3@gated-at.bofh.it>
In reply to#194538
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Thu, Apr 05, 2018 at 09:14:51PM -0400, Stefan Monnier wrote:
> >> UUIC that's partly why it's finally losing popularity and being replaced
> >> with json for that use.  I'm not familiar enough with json to know if
> >> it's really a good replacement, but it does look like an improvement.
> > that is simply not true.
> 
> Did you read the text to which I was responding?  Because your reply
> does not seem to contradict mine: I was talking about uses of XML for
> "data serialization".

Exactly. And that's the "misunderstanding" part I was talking about.
XML was intended as a document serialization language (as was SGML).
As such, it's passable, although (by far!) not pretty. It has been
misused as a data serialization language, and that's the problem.

Now JSON *is* a much better data serialization language (it does have
its problems, mind you: among other things having "integers" and not
telling you whether there's a max int or what happens when you pass
that possible limit).

The JSON folks don't help, in that they stubbornly talk about a
JSON "document" and pit "JSON vs XML". They are different beasts.

Try writing your next letter in JSON to see what I mean (this is
more directed at deloptes: you, Stefan and me are in violent agreement,
I think).

Cheers

- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrHGU8ACgkQBcgs9XrR2kYGXwCfYGe5n5NBhoCeH/VMJ/WSmPNN
kacAnjGk/KjRr+/0tkVz79ra0rPFbefq
=RJ22
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194539

FromBen Caradoc-Davies <ben@transient.nz>
Date2018-04-06 04:20 +0200
Message-ID<vBp8R-4Qn-5@gated-at.bofh.it>
In reply to#194532
On 06/04/18 09:33, deloptes wrote:
> Stefan Monnier wrote:
>> UUIC that's partly why it's finally losing popularity and being replaced
>> with json for that use.  I'm not familiar enough with json to know if
>> it's really a good replacement, but it does look like an improvement.
> that is simply not true. JSON might be more simple, and might be a
> replacement to XML in some cases, but XML has the features of SGML and per
> definition
> XML is not a single Markup Language. It is a metalanguage to let users
> design their own markup language.

Indeed. XML has W3C XML Schema, and stronger validation can be specified 
with Schematron (an ISO standard). JSON Schema is only an IETF draft. 
JSONP allows avoidance of JavaScript XMLHttpRequest same-origin rules 
and has increased the popularity of JSON, but the modern solution is CORS.

Kind regards,

-- 
Ben Caradoc-Davies <ben@transient.nz>
Director
Transient Software Limited <https://transient.nz/>
New Zealand

[toc] | [prev] | [next] | [standalone]


#194541

From<tomas@tuxteam.de>
Date2018-04-06 09:10 +0200
Message-ID<vBtFv-7QW-15@gated-at.bofh.it>
In reply to#194539
-----BEGIN PGP SIGNED MESSAGE-----
Hash: SHA1

On Fri, Apr 06, 2018 at 02:12:22PM +1200, Ben Caradoc-Davies wrote:
> On 06/04/18 09:33, deloptes wrote:
> >Stefan Monnier wrote:
> >>UUIC that's partly why it's finally losing popularity and being replaced
> >>with json for that use.  I'm not familiar enough with json to know if
> >>it's really a good replacement, but it does look like an improvement.
> >that is simply not true. JSON might be more simple, and might be a
> >replacement to XML in some cases, but XML has the features of SGML and per
> >definition
> >XML is not a single Markup Language. It is a metalanguage to let users
> >design their own markup language.
> 
> Indeed. XML has W3C XML Schema, and stronger validation can be
> specified with Schematron (an ISO standard). JSON Schema is only an
> IETF draft. JSONP allows avoidance of JavaScript XMLHttpRequest
> same-origin rules and has increased the popularity of JSON, but the
> modern solution is CORS.

Well, I think here we're leaving the document/vs data thing: if you
are already serializing whole objects (with methods), that's a whole
other kettle of fish (with more interesting consequences: the Java
community has lots of experience with that [1])

Cheers

[1] https://it.slashdot.org/story/18/03/27/2041225/atlanta-hit-by-ransomware-attack-also-fell-victim-to-leaked-nsa-exploits
- -- tomás
-----BEGIN PGP SIGNATURE-----
Version: GnuPG v1.4.12 (GNU/Linux)

iEYEARECAAYFAlrHGwcACgkQBcgs9XrR2kbeXQCfSNIO5s234lVFZHVy/+NFnGYR
JBsAn3geK1MkkWpX3OxJOp7+Aixo+FFJ
=2u7X
-----END PGP SIGNATURE-----

[toc] | [prev] | [next] | [standalone]


#194527

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2018-04-05 19:00 +0200
Message-ID<vBgoW-73k-5@gated-at.bofh.it>
In reply to#194524
> Actually people saying mbox is a bad database are in principle right
> (I never liked maildir either: dumping metadata into file names seemed
> to me a bit disgusting too, but I disgress).  But there's something
> special about mail databases which eases that a bit: records (i.e.
> mails) are *mostly* immutable (save for some metadata), so cleverly
> written libs (mutt, dovecot) can be suprisingly good despite all.

Actually, I think the reason it works is unrelated: it's just that
people have put enough engineering efforts into making it work for large
mailboxes despite its inadequate format.

Caching, auxiliary indexes, batched-rewrites, etc... can go a long way.


        Stefan

[toc] | [prev] | [next] | [standalone]


#194529

Fromrhkramer@gmail.com
Date2018-04-05 20:00 +0200
Message-ID<vBhl0-7Dz-23@gated-at.bofh.it>
In reply to#194523
On Thursday, April 05, 2018 08:42:39 AM rhkramer@gmail.com wrote:
> I'm laughing (at myself)--I just checked my mail directory, I have at least
> 4 mbox files (and then I stopped looking) greater than 175 MB.  One of
> them, my inbox, is 2.8 GB--no problems.
> 
> I do need to compact my inbox, and I did, but maybe the actual file isn't
> changed until I quit kmail--I'll try that later.

Just an update, in case someone is trying to follow along--it turns out (and I 
remember now) that kmail disables compaction on large mbox files "for safety 
reasons".

I forget how I've done it in the past--I probably copied the still valid 
emails into a temporary mail folder, deleted the current inbox, then created a 
new inbox and moved the emails from the temporary folder back to the new 
inbox.  I'll do all that with kmail not running.

(Getting old is aggravating.)

I might also look for a utility that can compact mbox files--at first glance, it 
looks like archivemail can do that and is in the repository for Wheezy.

[toc] | [prev] | [next] | [standalone]


#194467

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2018-04-03 23:40 +0200
Message-ID<vABON-53T-7@gated-at.bofh.it>
In reply to#194465
>> > What is the length of a string?
>> When is that relevant?
> When you're trying to display one on a screen, or print one on paper.

To display a string you don't just need its length, you need the actual
bitmap representation, and getting info such as length is trivial once
you've rendered the string into a bitmap (finding relevant fonts, etc..).

Also, when it comes to display you never really care about the length of
the string: you typically care about its pixel-width and
pixel-height instead.

> When you've been asked to find the longest/shortest string from a list.
> When you've been asked to sort a list of strings by length.

Again, if it's for display purposes, you'll want to use the actually
pixel-width instead, which again introduces the need to figure out which
font (or set of fonts since many/most fonts only cover a subset of
Unicode) to use, etc...

> Or in other words, basically every time you do anything with the string
> at all other than blindly byte-copying it to a different place in memory.

Your experience is quite different from mine.


        Stefan

[toc] | [prev] | [next] | [standalone]


Page 5 of 6 — ← Prev page 1 2 3 4 [5] 6  Next page →

Back to top | Article view | linux.debian.user


csiph-web